Four graph edges replaced three paragraphs of prompt

A 220-line system prompt had grown three separate paragraphs defending one step's ordering. Moving the sequence into four graph edges deleted all three, and cut a pipeline run from four model round trips to one.

Four nodes connected in a fixed sequence, with a bypass route drawn as a dashed line ending in a stop block

Three paragraphs of a 220-line system prompt were dedicated to a single rule: do not skip the third step. They were written on three separate occasions, after three separate incidents, in three different places in the file. The rule is not a heuristic or a style preference. It is the order four stages have to execute in, and it was being enforced by asking.

The problem

The pipeline sizes a facility in four stages. Convert plant capacity into annual demand, search outward for a procurement radius that meets it, retrieve every harvestable cluster inside that radius, then price and rank them into a supply curve. Each stage consumes the previous one's output.

The model chose which tool to call. The prompt told it which order to choose them in. This is the default design for a tool-using agent and it is fine until a sequence contains a step whose absence is invisible.

Retrieval is that step. It loads a pool of forest clusters into memory, and the map is drawn from that pool. The pool is not persisted between requests, because it runs to tens of thousands of records and the session store keeps a summary instead.

So the failure has a specific shape. On a later turn the model still holds the previous scenario summary in context: the radius, the average delivered cost, how many clusters were selected. It can answer a follow-up from those numbers alone, fluently and with the right figures, while the map beside the answer renders empty because the pool backing it is gone.

Nothing errors. The call that would have failed is the one that was never made.

Each incident produced another paragraph. That is worth stating plainly, because the pattern is easy to be inside of without noticing: a prompt that grows a new rule after each failure is not converging on a guarantee. It is accumulating a record of the times the last rule did not hold.

What moving control flow into edges does

The four stages became four nodes in a LangGraph state graph, with fixed edges between them, and the graph is exposed to the model as one tool rather than four.

builder.add_edge(START, "estimate_demand")

for source, following in (
    ("estimate_demand", "estimate_radius"),
    ("estimate_radius", "retrieve_clusters"),
    ("retrieve_clusters", "build_supply_curve"),
):
    builder.add_conditional_edges(
        source, _continue_or_stop, {"continue": following, "stop": END}
    )

The conditional edge carries the failure path, so a stage that returns an error stops the run and names the stage that stopped it, rather than feeding a broken result forward.

What actually changed is where the ordering lives.

MechanismEnforces orderFails howCosts
Prompt instructionNo, requests itSilently, and only sometimesTokens on every turn
Tool result telling the model to go backNo, requests it againSilently, one turn laterAn extra round trip
Runtime assertion in each toolYesLoudly, as a tool error the model must recover fromAn error path per tool
Graph edgeYesNot applicable, there is no ordering to get wrongA dependency

The row that matters is the third against the fourth. An assertion inside build_supply_curve that refuses to run without a pool would also make the bug impossible, and needs no new dependency. It converts a silent wrong answer into a loud failure the model has to handle, which is a real improvement. The graph goes further by removing the decision instead of checking it, and that is the difference worth paying for: the model no longer sequences four calls correctly on every turn, it makes one.

The four stages shown twice, once joined by dashed advisory arrows and once by solid enforced edges
The stages did not change. What changed is whether the arrows between them are advice or structure.

The pattern is not specific to agents. Anywhere a sequence must hold and one step's absence produces a plausible output rather than an error, the same move applies: take the ordering out of the thing that can forget it. A checkout flow where the payment authorisation is skipped but the confirmation still renders. An ETL job where a late-arriving join is missed and the downstream aggregate is merely smaller, not wrong-looking. The test is not whether the step is important, it is whether skipping it is visible.

All three paragraphs came out of the prompt.

The chat progress rail showing four completed stages under a single tool call, each with its own result summary
The four stages still report individually. The model only made one call.

What it costs

A dependency, and not a small one. LangGraph pulls in twenty-one transitive packages onto a backend whose requirements file had twenty lines. For four nodes and one conditional, a hand-written state machine would be shorter than the import. The justification is the second and third graph, and the checkpointing underneath, not this one.

A layering bug the tests did not catch. Progress spinners in the chat interface open on a tool_start event and close on a tool_done naming the same tool. The wrapper that emitted the closing event sat above the graph, so once the stages ran as nodes, four spinners opened and none closed. The frontend drops a close event naming no open step, so the failure would have been silent: four spinners turning forever, and any error reported against a tool with no spinner at all. Moving work down a layer moves it out from under whatever the layer above was doing for it.

No speed gain, and none was expected. Consolidating four model round trips into one is a real saving in latency and tokens, but the compute underneath is identical. The concurrency the graph made safe turned out to be worth 1.05x, because the scoring it would parallelise is CPU-bound Python and the GIL serialises it. That is a separate result and it is not a flattering one.

Where it breaks

The guarantee is narrower than it looks. Edges enforce the order of the stages. They do not make any stage correct, and they do not compel the model to call the tool at all. What changed is the size of the ask, from sequencing four calls correctly on every turn to choosing one tool. That is a much smaller surface. It is not zero.

The old tools still exist. Two other features read the cluster pool directly, so the individual stage tools remain callable and anything reaching them bypasses the graph. The ordering guarantee holds on the path through the graph rather than across the codebase, which is a weaker claim than it first appears.

Nothing reports when the model declines to use it. If the model answers a siting question without calling the tool, the graph never runs and there is no signal that it did not. The class of bug this fixes is invisible by construction, and so is this residue of it.

Failure inside a node is not modelled as richly as it should be. A stage returning an error stops the run, which is correct. But there is no retry, no partial result, and no way for the model to supply a corrected parameter and resume from the failed stage. In practice a failure means starting the pipeline again.

More on how a siting question is answered and how to read an answer and its map together. FRED itself is at biofred.us if you would rather ask it something than read about it.

References

  1. LangChain. LangGraph documentation.

  2. Python Software Foundation. Global interpreter lock. Python glossary.

  3. Yao, S. et al. (2023). ReAct: Synergizing Reasoning and Acting in Language Models. ICLR 2023.