Blueprints Linted, living reference architectures for agentic systems.

Payment dispute handling with human decisions

Classifies a payment dispute, assesses the evidence against policy and prepares a recommendation that a case worker decides.

router risk high v1 financial services seed

Why this topology

  1. high-risk-no-autonomy: autonomous topology removed (risk is high)
  2. prefer-workflow: agentic topologies removed (the task is not open-ended)
  3. autonomous-needs-consent: autonomous topology removed (autonomy is act_with_approval)
  4. multiple-task-kinds: router (2 request kinds)

Mandatory controls: human_gate input_filter audit_log output_filter retrieval

Flow

  1. intake → Dispute intake · 1200 in / 200 out
  2. assess → Evidence analyst · 6000 in / 900 out
Diagram source (Mermaid)
flowchart LR
    io_in(["request"])
    io_out(["response"])
    c_intake_model(["Dispute intake"])
    c_analyst_model(["Evidence analyst"])
    c_policy_kb[("Policy and precedent index")]
    c_case_tool[["Case system"]]
    c_ledger_tool[["Ledger"]]
    c_decision_gate{{"Case decision"}}
    g_redact>"input_filter"]
    g_disclosure_check>"output_filter"]
    g_audit>"audit_log"]
    io_in --> c_intake_model
    c_intake_model --> c_analyst_model
    c_analyst_model --> io_out
    g_redact -.-> c_intake_model
    g_redact -.-> c_analyst_model
    g_disclosure_check -.-> c_analyst_model
    g_audit -.-> c_case_tool
    g_audit -.-> c_ledger_tool

Components

idnamekindrolemodel / tool / framework
intake-modelDispute intakemodelclassify the dispute type and extract the claimed facts claude-haiku-4-5
analyst-modelEvidence analystmodelassess the evidence against policy and draft a recommendation claude-opus-5-5
policy-kbPolicy and precedent indexretrievaldispute rules and past decisions
case-toolCase systemtoolread and update the dispute case
ledger-toolLedgertoolpost a refund or a chargeback to the ledger
decision-gateCase decisionhuman_gatea case worker decides every outcome

Guardrails

idkindprotectswhat it does
redactinput_filterintake-model, analyst-modelAccount and card numbers are tokenised before any model sees the case.
disclosure-checkoutput_filteranalyst-modelBlocks recommendations that quote personal data or internal fraud rules.
auditaudit_logcase-tool, ledger-toolEvery read, update and posting is logged with the case id and the decider.

Failure modes

failurecomponentmitigated by
A dispute is classified under the wrong rule set.Dispute intakedecision-gate
The assessment ignores a document that changes the outcome.Evidence analystdecision-gate
A recommendation exposes personal data or internal rules.Evidence analystdisclosure-check
A posting cannot be traced to a person and a case.Ledgeraudit

Threats

OWASPcomponentmitigated bynote
LLM01 Prompt InjectionDispute intakeredactcustomer-submitted text is untrusted
LLM02 Sensitive Information DisclosureEvidence analystdisclosure-checksensitive data in outputs
LLM06 Excessive AgencyLedgerdecision-gatemoney movement
LLM09 MisinformationEvidence analystdecision-gatea person checks every recommendation

Eval plan

metricmethodthreshold
classification accuracy500 labelled historical disputes>= 0.95
recommendation agreementcase workers compare recommendations with their own decision on a blind sample>= 0.85
gate coveragereplay; no posting without a decision record100%

Observability and deployment

service inside the regulated network · No outbound calls except to the model gateway.

Estimated cost

$0.0442 per request · $1060.80 / month at 800 requests/day · ~18.03 s end to end

Prices as of 2026-10-01 · source · latency is assumed.

Runnable starter

No starter is offered: no API data for langgraph from Radar yet.

Changelog