Documentation Q&A with cited answers
Rewrites the question, retrieves documentation passages and answers only with citations to them.
workflow risk low v1 developer documentation seed
Why this topology
- interactive-latency: only workflow or router (interactive latency)
- prefer-workflow: agentic topologies removed (the task is not open-ended)
- autonomous-needs-consent: autonomous topology removed (autonomy is suggest)
- default-workflow: workflow (one kind of task with fixed steps)
Mandatory controls: retrieval rate_limit
Flow
- rewrite → Query rewriter · 400 in / 40 out
- answer → Answer writer · 3000 in / 350 out
Diagram source (Mermaid)
flowchart LR
io_in(["request"])
io_out(["response"])
c_rewrite_model(["Query rewriter"])
c_answer_model(["Answer writer"])
c_docs_index[("Documentation index")]
g_rate>"rate_limit"]
g_citations>"output_filter"]
io_in --> c_rewrite_model
c_rewrite_model --> c_answer_model
c_answer_model --> io_out
g_rate -.-> c_rewrite_model
g_rate -.-> c_answer_model
g_citations -.-> c_answer_model
Components
| id | name | kind | role | model / tool / framework |
|---|---|---|---|---|
| rewrite-model | Query rewriter | model | turn the question into a search query | claude-haiku-4-5 |
| answer-model | Answer writer | model | answer from the retrieved passages and cite them | claude-haiku-4-5 |
| docs-index | Documentation index | retrieval | hybrid keyword and embedding index of the documentation |
Guardrails
| id | kind | protects | what it does |
|---|---|---|---|
| rate | rate_limit | rewrite-model, answer-model | Per-user and global request caps keep spend bounded. |
| citations | output_filter | answer-model | Drops any answer sentence that does not cite a retrieved passage. |
Failure modes
| failure | component | mitigated by |
|---|---|---|
| Retrieval returns passages that do not contain the answer. | Documentation index | citations |
| The model answers from memory instead of the passages. | Answer writer | citations |
| A traffic burst multiplies model spend. | Answer writer | rate |
Threats
| OWASP | component | mitigated by | note |
|---|---|---|---|
| LLM01 Prompt Injection | Answer writer | citations | documentation pages can carry instructions |
| LLM08 Vector and Embedding Weaknesses | Documentation index | citations | poisoned or stale passages |
| LLM09 Misinformation | Answer writer | citations | uncited claims are removed |
| LLM10 Unbounded Consumption | Answer writer | rate | unbounded consumption |
Eval plan
| metric | method | threshold |
|---|---|---|
| retrieval recall at 5 | 300 questions with known source pages | >= 0.85 |
| faithfulness | sampled answers graded against their cited passages | >= 0.95 |
Observability and deployment
- traces per question
- retrieval scores
- uncited sentences removed
- cost per question
stateless service with a managed index · Reindex on every documentation release.
Estimated cost
$0.0053 per request · $3210.00 / month at 20000 requests/day · ~3.6 s end to end
Over the stated budget of $3000 per month.
Prices as of 2026-10-01 · source · latency is assumed.
- token counts per step are estimates written with the blueprint and approved by the owner, not measurements
- steps run one after another, so parallel branches make the latency an overestimate
- latency uses an assumed time to first token and output speed per model
- 30 days per month and the stated requests per day
Runnable starter
No starter is offered: no API data for langgraph from Radar yet.
Changelog
- v1: seed