Blueprints Linted, living reference architectures for agentic systems.

Documentation Q&A with cited answers

Rewrites the question, retrieves documentation passages and answers only with citations to them.

workflow risk low v1 developer documentation seed

Why this topology

  1. interactive-latency: only workflow or router (interactive latency)
  2. prefer-workflow: agentic topologies removed (the task is not open-ended)
  3. autonomous-needs-consent: autonomous topology removed (autonomy is suggest)
  4. default-workflow: workflow (one kind of task with fixed steps)

Mandatory controls: retrieval rate_limit

Flow

  1. rewrite → Query rewriter · 400 in / 40 out
  2. answer → Answer writer · 3000 in / 350 out
Diagram source (Mermaid)
flowchart LR
    io_in(["request"])
    io_out(["response"])
    c_rewrite_model(["Query rewriter"])
    c_answer_model(["Answer writer"])
    c_docs_index[("Documentation index")]
    g_rate>"rate_limit"]
    g_citations>"output_filter"]
    io_in --> c_rewrite_model
    c_rewrite_model --> c_answer_model
    c_answer_model --> io_out
    g_rate -.-> c_rewrite_model
    g_rate -.-> c_answer_model
    g_citations -.-> c_answer_model

Components

idnamekindrolemodel / tool / framework
rewrite-modelQuery rewritermodelturn the question into a search query claude-haiku-4-5
answer-modelAnswer writermodelanswer from the retrieved passages and cite them claude-haiku-4-5
docs-indexDocumentation indexretrievalhybrid keyword and embedding index of the documentation

Guardrails

idkindprotectswhat it does
raterate_limitrewrite-model, answer-modelPer-user and global request caps keep spend bounded.
citationsoutput_filteranswer-modelDrops any answer sentence that does not cite a retrieved passage.

Failure modes

failurecomponentmitigated by
Retrieval returns passages that do not contain the answer.Documentation indexcitations
The model answers from memory instead of the passages.Answer writercitations
A traffic burst multiplies model spend.Answer writerrate

Threats

OWASPcomponentmitigated bynote
LLM01 Prompt InjectionAnswer writercitationsdocumentation pages can carry instructions
LLM08 Vector and Embedding WeaknessesDocumentation indexcitationspoisoned or stale passages
LLM09 MisinformationAnswer writercitationsuncited claims are removed
LLM10 Unbounded ConsumptionAnswer writerrateunbounded consumption

Eval plan

metricmethodthreshold
retrieval recall at 5300 questions with known source pages>= 0.85
faithfulnesssampled answers graded against their cited passages>= 0.95

Observability and deployment

stateless service with a managed index · Reindex on every documentation release.

Estimated cost

$0.0053 per request · $3210.00 / month at 20000 requests/day · ~3.6 s end to end

Over the stated budget of $3000 per month.

Prices as of 2026-10-01 · source · latency is assumed.

Runnable starter

No starter is offered: no API data for langgraph from Radar yet.

Changelog