Most agent workflows are one of six shapes, and two of them, routing and verification, give you the most for the least effort. Teams often wire agents together ad hoc: one prompt calls another, a retry gets bolted on, a reviewer shows up after the first bad output. It works until it doesn’t, and then nobody can say why.
Naming the shape first makes the failure modes predictable. Below is the full set at a glance, then a deep dive with code on the two I’d build first, then a fuller look at the other four. The code is illustrative pseudo-Python: helpers like llm() and generate_sql() stand in for your own implementations.
The six at a glance
| PatternUse it whenMain risk | ||
| Classify-And-Act | Tasks need different specialists | Misrouting |
| Fanout-And-Synthesize | Work splits into independent parts | Weak merge, duplicated work |
| Adversarial Verification | A wrong answer is expensive | Verifiers that agree by default |
| Generate-And-Filter | Good answers are rare | A vague or harsh rubric |
| Tournament | Comparing is easier than scoring | Judge bias |
| Loop Until Done | The amount of work is unknown | A loop with no stop condition |
Deep dive 1: Classify-And-Act as a question router
A router is the cheapest pattern to add and the one that pays off soonest. A small classifier reads the request, picks one specialist, and only that specialist runs. Cost stays close to a single call, and each specialist can be tuned for one job.
Take an analytics assistant that has to answer from two very different sources: uploaded documents and a SQL database. “What does the policy say about refunds?” is a retrieval question. “How many refunds did we issue last month?” is a data question. One agent doing both tends to do both badly, so route first:
ROUTES = {"docs": answer_with_rag, "data": answer_with_sql}
def route(question: str) -> str: label = llm( "Classify the question as exactly one of: docs, data, unclear. " "Reply with the label only.\n" f"Question: {question}" ).strip().lower() return label if label in ROUTES else "unclear" # anything unexpected is unclear
def answer(question: str) -> str: label = route(question) log_route(question, label) # keep every routing decision if label == "unclear": return ask_clarifying_question(question) return ROUTES[label](question)CopyThe classifier is the whole risk. If it misroutes, the right agent never sees the question, and the wrong agent answers confidently. Four habits keep that manageable:
- Keep the label set small and mutually exclusive.
- Include an explicit
unclearlabel and ask a follow-up instead of guessing. - Log every routing decision so you can review mistakes.
- Build a small labelled set of real questions and measure routing accuracy before you tune anything else.
Skip the pattern when nearly every request goes to one path anyway, or when you can’t tell which specialist is needed without doing most of the work first.
To know it works, collect a set of real questions, label each with the route it should take, and measure routing accuracy on that set before you tune prompts. Re-run it whenever you change the classifier or add a route.
Deep dive 2: Adversarial Verification for generated SQL
Verification is worth its extra cost when a wrong answer looks right. Generated SQL is the classic case: a query can run cleanly, return rows, and still answer a different question than the one asked. Nobody notices until a number ends up in a report.
The pattern is to let a worker produce the result, then have independent checks try to break it. Mix cheap deterministic checks with model-based ones, and only accept the answer if all of them pass:
def static_issue(sql: str) -> str | None: """Deterministic checks: plain code, no LLM.""" if not is_read_only(sql): return "Use a single read-only SELECT statement." if missing := unknown_columns(sql): return f"These columns do not exist: {missing}" return None
def semantic_issue(question: str, sql: str, rows) -> str | None: """Model-based checks, prompted to hunt for faults. Returns a reason or None.""" return ( llm_find_flaw(question, sql) or llm_check_rows_plausible(question, rows) )
def answer_with_sql(question: str, max_attempts: int = 2) -> str: feedback = None for _ in range(max_attempts): # hard cap on attempts sql = generate_sql(question, feedback) issue = static_issue(sql) if issue is None: # gate BEFORE executing rows = run_read_only(sql) # also use a read-only DB role issue = semantic_issue(question, sql, rows) if issue is None: return summarize(rows) feedback = issue # retry with the reason return "I couldn't verify an answer to that."CopyThe order matters. The static checks run before the query executes, so an unsafe statement never reaches the database, and the cheap checks run before the model-based ones, so most failures cost almost nothing.
Three details decide whether this works:
- Verifiers must be adversarial. Prompt them to find faults, not to confirm. A verifier told to “check this looks fine” mostly says yes.
- Avoid shared blind spots. If the same model with the same prompt writes and checks the SQL, it will often repeat its own mistake. Deterministic checks, a different prompt, or a different model break that correlation.
- A failed check is information. Retry with the failure reason, cap the retries, and return an honest “couldn’t verify” instead of a confident guess.
Skip it for low-stakes answers or tight latency budgets. Every verifier adds cost and delay, so spend them where an error would actually hurt.
Track how often the verifier rejects an answer and why. If it never rejects anything, it probably isn’t adversarial enough; if it rejects everything, the checks are too strict or the generator needs work.
The other four patterns
Fanout-And-Synthesize
Use it when a task breaks into independent parts, so running them in parallel saves time. Several workers run at once and one synthesizer merges what they return. Latency is roughly the slowest worker plus the merge, not the sum of all workers.
Example: researching a competitor. Separate workers cover pricing, product, hiring and news, and a fifth agent writes the brief. Give each worker a distinct slice and a fixed output format, or you’ll pay for overlapping answers that are hard to merge. Put most of your effort into the synthesizer: it has to reconcile contradictions, flag gaps, and say which worker a claim came from. Skip it when the parts depend on each other, because a later step then needs an earlier result.
Generate-And-Filter
Use it when good answers are rare but easy to recognize. Generators produce many candidates, and a filter applies a written rubric and removes duplicates.
Example: generate 30 product names, drop repeats and anything too long or hard to spell, and keep the top three. Write the rubric before you generate, with pass/fail criteria where you can, because a vague rubric turns the filter into one more opinion. Vary the generators, with different prompts or settings, so they don’t all produce the same ideas. Log the discarded pile too: it shows whether the rubric is too harsh. Skip it when there is one correct answer to find, and use verification instead.
Tournament
Use it when you can’t score an answer on its own but can say which of two is better. Attempts run independently, pairwise judges compare them two at a time, and winners advance until one remains. Comparing two answers is usually easier and more consistent than giving each an absolute score.
Example: five drafts of a release note. Judges compare pairs, and the winners face off. Cost grows with the number of attempts, since eight attempts need seven comparisons, so keep brackets small. Judges often favor whichever answer they see first, so randomize the order or judge each pair in both orders. Skip it when a simple test or rubric can score each attempt directly, because that is cheaper.
Loop Until Done
Use it when you don’t know in advance how much work there is. An agent works, then a check asks whether it found anything new. If yes, spawn another agent; if no, stop.
Example: a bug-hunting agent rescans a repo until a round comes back empty. The stop condition is the whole design, so set a hard cap on rounds or budget, and define “new” precisely, such as a finding not already on the running list.
def hunt(repo, max_rounds: int = 5) -> dict: findings = {} # stable key -> finding for _ in range(max_rounds): # hard cap results = run_agent(repo, known=list(findings)) new = {f.key: f for f in results if f.key not in findings} if not new: # nothing new: stop break findings.update(new) return findingsCopySkip it when the amount of work is known up front, because a plain fanout is simpler.
Start with one
The patterns compose: a router can send work to a fanout, a fanout can end in verification, and a loop can spawn a fanout each round. But start with one. Pick the step in your pipeline that fails most often and match it to a pattern.
Log each agent’s input and output so you can see where quality drops, and add a second pattern only when the first one’s failure mode shows up in those logs. For most systems, that first pattern will be a router or a verifier.



