Using agents to update the UI

An agent that draws on a map does not draw on the map. It mutates state, and the view is derived from that state, so the prose and the pixels have different producers. Keeping them describing the same thing is the actual work.

One source node feeding two separate downstream paths, one ending in a text block and one in a map frame

Ask the system where to site a facility and two things come back: a written answer and a map with several thousand forest clusters on it, coloured by whether they made the supply curve. It looks like the model drew the map. It did not, and the difference is the whole architecture.

The problem

The model produces text. That is all it produces. It also calls tools, and those tools mutate a session object: the facility location, the retrieved cluster pool, the current scenario, the last supply curve result.

The map is built afterwards, from that session object, by a function the model never invokes:

def build_spatial_response(state: StateManager) -> Optional[SpatialState]:
    facility = state.state.facility
    scenario = state.state.current_scenario
    cached = state.state.cached_clusters

So there are two producers of one answer. The prose comes from the model. The view comes from the state the model's tool calls left behind. Nothing reconciles them.

The appeal of this is real. The model cannot draw a cluster that does not exist, cannot invent a coordinate, and cannot colour a stand as selected when the optimiser did not select it. Every mark on the map traces to a database row. If instead the model emitted drawing instructions, every one of those would be a thing it could get wrong, and it would get them wrong in the confident, plausible way models do.

The cost arrives immediately: the two can describe different things, and nothing detects it.

What that looks like when it fails

The failure has a specific shape, and it is not hypothetical.

The cluster pool is held in memory and not persisted between requests, because it runs to tens of thousands of records. The scenario summary is persisted: radius, average delivered cost, how many clusters were selected.

On a later turn the model still has that summary in its context. It can answer a follow-up question from those numbers alone, fluently, with correct figures, and never call the retrieval tool. The prose is right. The map is empty, because the pool it derives from is gone.

Nothing errors. Both halves did exactly what they were built to do.

That specific case is now prevented by making the retrieval step structural rather than optional, which is a separate post. The general shape is not prevented, and it cannot be while the two paths are independent.

What constrains the view

Between the state and the map there are two small pure functions, and both encode a judgement that is easy to get wrong in a component.

Layers are mutually exclusive, and the rule is explicit. A statewide search paints hexagons and candidate markers. A completed analysis paints clusters. Both at once is unreadable, so when a result carries clusters the coarser layers are dropped:

export function withoutCoarseLayers(spatial: SpatialState): SpatialState {
  return { ...spatial, hexes: [], hex_resolution: null, candidates: [] };
}

Intent has to be separated from noise. Follow-up turns return the same facility coordinates, and floating point makes "the same" unreliable. So a move is only a move past a tolerance:

// Roughly 100m.
export const FACILITY_MOVE_TOLERANCE_DEG = 0.001;

This matters because only a genuine relocation may discard the multi-year projection. Without the tolerance, arithmetic noise on an unchanged site would throw away a twenty-year simulation. A hundred metres is far below the resolution of any siting decision here and far above float error, so there is a wide band to be wrong in. That band is what makes the constant defensible, not the particular value.

Both are ordinary functions taking a value and returning one, tested without a map or a browser. The rules that decide what a user sees are the ones worth being able to test in isolation.

The loop runs both ways

The map is not only an output. Selecting a candidate location does not call an API; it composes a message and sends it as if the user had typed it:

chatRef.current?.sendMessage(
  `I've selected a new location at ${lat.toFixed(4)}°N, ${lng.toFixed(4)}°W.`
);

A click becomes a sentence, and the agent handles it with the same path as anything typed. There is one way into the system, so the map needs no privileged channel and the conversation stays a complete record of what happened. A user scrolling back sees why the analysis moved.

It is also unmistakably a hack. The interaction is structured, it is being encoded into English so it can be parsed back out, and the phrasing is load-bearing in a way nothing enforces. Reword that string carelessly and the agent may not understand its own UI.

Where it breaks

Nothing reports disagreement. No check compares what the answer claims against what the view contains. The two are assembled independently and shipped together, and every consistency guarantee comes from the tools that mutated the state, not from anything watching the pair.

Deriving the payload costs seconds at scale. For each cluster in the pool, the builder linearly scans the selected clusters to attach its cost and score. At a small radius that is 47 milliseconds and invisible. At 47,000 clusters against 5,500 selected it is 2.9 seconds of pure Python on every turn, before anything is sent. Indexing the selected set by cluster number first makes the same output in 2.6 milliseconds. The set for the selected flag is already built two lines above; the other two fields just never got the same treatment.

Rendering is deferred by a heuristic. A large marker build blocks the frame, so it is pushed past two animation frames to let the answer paint first, with a spinner cleared on render completion. It works. It is also a guess about frame budget with no measurement behind it, and it will be wrong on some device.

Structured intent is encoded as prose. The synthetic message is parsed by a model, so a UI action reaches the agent through natural language and the coupling is invisible to every tool that would normally catch a broken interface: no type, no test, no compiler.

More on how a siting question is answered and how to read an answer and its map together. FRED itself is at biofred.us if you would rather drive the map than read about it.

References

  1. Yao, S. et al. (2023). ReAct: Synergizing Reasoning and Acting in Language Models. ICLR 2023.

  2. Leaflet. Layer groups and layers control.

  3. MDN. Window.requestAnimationFrame().