Patterns
Prompt chaining draft
When to use. The task splits into fixed sequential steps and each step is easier than the whole.
When not to. The steps depend on what earlier steps find, so the sequence cannot be fixed in advance.
Failure modes.
- An early mistake is carried through every later step.
- Latency adds up across steps.
Cost profile: 1.0x calls, 1.0x latency. One model call per step, run in order.
Routing draft
When to use. Requests fall into distinct kinds that need different prompts, tools or models.
When not to. There is one kind of request, or the kinds cannot be told apart reliably.
Failure modes.
- A request is sent to the wrong handler.
- A new request kind has no route.
Cost profile: 1.0x calls, 1.0x latency. One cheap classification call, then the chosen path.
Parallelization draft
When to use. Independent subtasks can run at the same time, or several attempts are voted on.
When not to. Subtasks depend on each other, or the budget cannot absorb the extra calls.
Failure modes.
- Results conflict and nothing merges them.
- Cost multiplies with the number of branches.
Cost profile: 2.0x calls, 1.0x latency. Calls scale with branches; latency stays near one branch.
Orchestrator and workers draft
When to use. The subtasks are not known in advance and a lead model must plan and delegate them.
When not to. The subtasks are fixed, or latency is tight.
Failure modes.
- The orchestrator plans badly or loops.
- Worker output is not checked before use.
Cost profile: 3.0x calls, 2.0x latency. A planning call, several worker calls, and a merge.
Evaluator and optimizer draft
When to use. There are clear criteria and a second pass measurably improves the answer.
When not to. There is no objective criterion, or latency is tight.
Failure modes.
- The evaluator and the generator agree on a wrong answer.
- The loop never converges.
Cost profile: 3.0x calls, 3.0x latency. A draft and at least one evaluate and revise round.
ReAct draft
When to use. The next step depends on tool results and the path cannot be planned up front.
When not to. A fixed workflow would do, or actions are costly or irreversible.
Failure modes.
- The agent loops on a failing tool.
- A tool result steers the agent off task.
Cost profile: 4.0x calls, 4.0x latency. A model call per reasoning and action step, unbounded without a cap.
Plan and execute draft
When to use. A long task benefits from an explicit plan that is then carried out step by step.
When not to. The plan would be invalidated by the first result.
Failure modes.
- The plan is wrong and execution follows it anyway.
- Replanning is never triggered.
Cost profile: 3.0x calls, 3.0x latency. One planning call plus a call per plan step.
Reflection draft
When to use. A model can catch its own mistakes when asked to critique with explicit criteria.
When not to. The model cannot verify the claim without outside information.
Failure modes.
- The critique rubber-stamps the draft.
- Extra rounds add cost without improving quality.
Cost profile: 2.0x calls, 2.0x latency. A draft plus a critique and rewrite.
Handoff and swarm draft
When to use. Specialist agents own different domains and the conversation moves between them.
When not to. One agent with the right tools can do the job.
Failure modes.
- Context is lost at a handoff.
- Agents hand the task back and forth.
Cost profile: 2.0x calls, 2.0x latency. Each handoff adds a call and repeats context.
Human approval gate draft
When to use. An action is irreversible, costly or regulated, and a person can decide quickly.
When not to. Nobody is available to approve in time, or approvals would be rubber-stamped at volume.
Failure modes.
- Approvers rubber-stamp under load.
- The queue stalls the whole flow.
Cost profile: 1.0x calls, 1.0x latency. No extra model calls; human latency is outside the model budget.
Retrieval-augmented generation draft
When to use. Answers must be grounded in documents that change or are too large for the prompt.
When not to. The knowledge is small and static enough to put in the prompt.
Failure modes.
- Retrieval returns the wrong passages.
- The model answers from memory instead of the passages.
- The index goes stale.
Cost profile: 1.0x calls, 1.0x latency. One embedding lookup plus a longer prompt for the answer call.