Payment dispute handling with human decisions
Classifies a payment dispute, assesses the evidence against policy and prepares a recommendation that a case worker decides.
router risk high v1 financial services seed
Why this topology
- high-risk-no-autonomy: autonomous topology removed (risk is high)
- prefer-workflow: agentic topologies removed (the task is not open-ended)
- autonomous-needs-consent: autonomous topology removed (autonomy is act_with_approval)
- multiple-task-kinds: router (2 request kinds)
Mandatory controls: human_gate input_filter audit_log output_filter retrieval
Flow
- intake → Dispute intake · 1200 in / 200 out
- assess → Evidence analyst · 6000 in / 900 out
Diagram source (Mermaid)
flowchart LR
io_in(["request"])
io_out(["response"])
c_intake_model(["Dispute intake"])
c_analyst_model(["Evidence analyst"])
c_policy_kb[("Policy and precedent index")]
c_case_tool[["Case system"]]
c_ledger_tool[["Ledger"]]
c_decision_gate{{"Case decision"}}
g_redact>"input_filter"]
g_disclosure_check>"output_filter"]
g_audit>"audit_log"]
io_in --> c_intake_model
c_intake_model --> c_analyst_model
c_analyst_model --> io_out
g_redact -.-> c_intake_model
g_redact -.-> c_analyst_model
g_disclosure_check -.-> c_analyst_model
g_audit -.-> c_case_tool
g_audit -.-> c_ledger_tool
Components
| id | name | kind | role | model / tool / framework |
|---|---|---|---|---|
| intake-model | Dispute intake | model | classify the dispute type and extract the claimed facts | claude-haiku-4-5 |
| analyst-model | Evidence analyst | model | assess the evidence against policy and draft a recommendation | claude-opus-5-5 |
| policy-kb | Policy and precedent index | retrieval | dispute rules and past decisions | |
| case-tool | Case system | tool | read and update the dispute case | |
| ledger-tool | Ledger | tool | post a refund or a chargeback to the ledger | |
| decision-gate | Case decision | human_gate | a case worker decides every outcome |
Guardrails
| id | kind | protects | what it does |
|---|---|---|---|
| redact | input_filter | intake-model, analyst-model | Account and card numbers are tokenised before any model sees the case. |
| disclosure-check | output_filter | analyst-model | Blocks recommendations that quote personal data or internal fraud rules. |
| audit | audit_log | case-tool, ledger-tool | Every read, update and posting is logged with the case id and the decider. |
Failure modes
| failure | component | mitigated by |
|---|---|---|
| A dispute is classified under the wrong rule set. | Dispute intake | decision-gate |
| The assessment ignores a document that changes the outcome. | Evidence analyst | decision-gate |
| A recommendation exposes personal data or internal rules. | Evidence analyst | disclosure-check |
| A posting cannot be traced to a person and a case. | Ledger | audit |
Threats
| OWASP | component | mitigated by | note |
|---|---|---|---|
| LLM01 Prompt Injection | Dispute intake | redact | customer-submitted text is untrusted |
| LLM02 Sensitive Information Disclosure | Evidence analyst | disclosure-check | sensitive data in outputs |
| LLM06 Excessive Agency | Ledger | decision-gate | money movement |
| LLM09 Misinformation | Evidence analyst | decision-gate | a person checks every recommendation |
Eval plan
| metric | method | threshold |
|---|---|---|
| classification accuracy | 500 labelled historical disputes | >= 0.95 |
| recommendation agreement | case workers compare recommendations with their own decision on a blind sample | >= 0.85 |
| gate coverage | replay; no posting without a decision record | 100% |
Observability and deployment
- traces per case
- gate decisions and overrides
- disclosure-check blocks
- time to decision
service inside the regulated network · No outbound calls except to the model gateway.
Estimated cost
$0.0442 per request · $1060.80 / month at 800 requests/day · ~18.03 s end to end
Prices as of 2026-10-01 · source · latency is assumed.
- token counts per step are estimates written with the blueprint and approved by the owner, not measurements
- steps run one after another, so parallel branches make the latency an overestimate
- latency uses an assumed time to first token and output speed per model
- 30 days per month and the stated requests per day
Runnable starter
No starter is offered: no API data for langgraph from Radar yet.
Changelog
- v1: seed