One book I found both interesting and fun to read is Imran Ahmad’s 30 Agents Every AI Engineer Must Build. In chapter 7, the one about Chain-of-Agents, he uses a trip planner to explain the pattern, and does it in a genuinely educational way: an orchestrator that does no domain work, specialists that hand work to each other, a shared memory, and a mechanism for when two results don’t add up. I liked the idea, and set out to build it with one of those patterns as the foundation, using some concepts that differ from the book and that address some of the loose ends in Mr. Ahmad’s demo.

The loose end that caught my attention most was the conflict. In the code that accompanies the chapter, the conflict is detected: a score is computed, compared against a threshold, and the report recommends further investigation. Arbitration, consensus and renegotiation stay in the text of the book. The specialists, on top of that, return simulated data. That’s reasonable for a teaching notebook that covers three patterns in a few cells, but it left me with the question I actually cared about: what happens when the conflict has to be resolved, with a real model that gets things wrong and real data that sometimes doesn’t arrive?

The result is a demo where you type something like “Quito, 3 days, $600, food and historic center” and watch, live, as a chain of agents builds the itinerary: which agent is working, what it decides, with what probability, where a budget conflict appears, and how the chain resolves it without anyone stepping in. The pretty itinerary isn’t the point. The point is making the orchestration visible.

And the same thing I said in previous posts: I built it together with a coding assistant, over long sessions of decisions and measurements, with a devlog that records every problem and how it was solved.

The constraints defined the design (again)

This demo lives on the same ephemeral infrastructure as the previous ones, so it inherited their rules before I wrote a single line:

  • No paid API. A local LLM, served by Ollama. No per-token bills.
  • CPU only, 2 vCPU / 4 GiB for the whole pod, model included.
  • Stateless. The environment can die at any moment.
  • A small, open model. I evaluated Qwen, Llama and Phi, and settled on Qwen 2.5 1.5B.

A 1.5B model on two cores is slow and gets facts wrong. That combination designed everything else.

JEV: reasoning without generating

One of the ideas I wanted in this project was JEV, from TypeSafe, which was released recently (in early access since September 15). It’s what they call a System One model: instead of generating text token by token, it takes a program’s state and a set of questions whose shape you declare up front —which fields you want and which values each one can take— and returns those answers filled in, each with its probability. I found it interesting to include reasoning ability without having to pay for generation to get the answer, just an array of predefined answers. If the answer can only be one of four options, there’s no way for it to invent a fifth.

The problem was practical: JEV is a hosted API, in early access with a waitlist and no option to run it locally, so it didn’t meet the “everything local” rule. The solution was to keep the shape and make the engine swappable. The demo defines an interface shaped like JEV: state plus typed questions, which return answers with probabilities.

TypedQuestion(
    id="harm:cheaper_lodging",
    family="action_harm",          # to dispatch rules
    kind="score",                  # "choice" with options, or "score" -> P(yes)
    text="Does downgrading the lodging hurt the traveler's interests?",
    params={"action_id": "cheaper_lodging"},
)
# -> Answer(value=0.1, probabilities={"yes": 0.1, "no": 0.9}, source="rules")

Today it’s implemented by a rule-based engine: a keyword taxonomy in Spanish and English plus explicit heuristics. It’s deterministic, instant and works offline. Each question carries a family for the rules and a natural-language text, which is exactly what JEV would evaluate. The day I get access, plugging it in means implementing one method.

(I’ll admit that writing this made me think again about Agentic Racing. In the end, the team boss only chose between attack, defend or push, with low, medium or high aggression. A driver doesn’t need a paragraph over the radio: they need one of three words, and fast. We’ll see.)

Generate once, decide many times

The first version regenerated the itinerary on every round of the budget conflict. On CPU, that meant several long JSON outputs per trip, and it didn’t fit in a reasonable time. That’s where the principle that organizes the whole project came from: each step uses the cheapest mechanism that solves it well.

StepMechanism
Intake: classify each interestTyped decision
Live data: weather, places, exchange ratePublic APIs, no LLM
Destination researchCurated catalog + live data (with luck, 0 tokens)
ItineraryGeneration, only once
BudgetPython arithmetic
Conflict resolutionTyped decisions + arithmetic
Itinerary revisionCode, no regeneration
SynthesisGeneration (the narrative) + code (the figures)

The model writes twice per trip; everything else is a decision or a calculation. A full trip uses between 850 and 1,200 tokens, against a 6,000 cap, and takes 25 to 45 seconds within the pod’s limits. The conflict loop uses zero.

Resolving the conflict, not just detecting it

This is what I’d wanted to build since I read the chapter. When the total exceeds the budget, the conflict agent doesn’t ask the model to “make the trip cheaper”. It has a fixed catalog of actions: swap paid activities for free walks, downgrade the lodging, eat at markets, buy a transit pass. For each one:

  1. Code computes how much it saves. It’s arithmetic over the itinerary’s cost assumptions.
  2. The decision engine estimates how much it hurts the traveler’s interests, as a probability. For someone who asked for museums, free walks hurt more than for someone who asked to walk.
  3. Code chooses by savings × (1 − P(harm)), applies the changes to the itinerary and re-runs the budget.

The loop has a hard cap of two iterations. If the trip still doesn’t fit after that, the result says so. Lisbon for two people on $50 is my favorite test case: two rounds of cuts, free walks in every district, and an honest verdict that it isn’t enough. A planner that always says your trip fits isn’t a planner, it’s a salesman.

Every decision shows up in the trace with its probability.

When the model doesn’t know where Shibuya is

The first trips with the real model were educational. Qwen 1.5B put Shibuya, which is in Tokyo, as a district of Kyoto. It invented districts in Hanoi, a “Museo del Ámbito” that doesn’t exist, and priced Hanoi higher than Lisbon.

Prompts improved the prices, but not the geography. Then I ran a benchmark of 1.5B against 3B, with the same trips and the same CPU limits:

Trip1.5B3B
Kyoto (catalog)31 s71 s
Lisbon, $50 (catalog)21 s68 s
Hanoi (catalog)25 s56 s
Valparaíso (model only)23 s75 s

The 3B took more than twice as long and didn’t improve the geography: it put Kinkaku-ji in Gion, “Ribeira” (which is in Porto) in Lisbon, and “Paseo Alcorta” (which is in Buenos Aires) in Valparaíso.

A prompt doesn’t teach a small model what it doesn’t know, and a slightly bigger model doesn’t either. Facts that have to be right are stored as data. The demo has a curated catalog of 45 cities, each with three districts, and each district lists the highlights that are physically in it. Code decides which district each day visits and passes the model only that district’s highlights, so a temple can’t end up on the wrong day or in the wrong district. If the model returns a different district, normalization enforces the planned one. The model still writes all the prose; what it no longer decides is the geography.

The same thing happened with the cuts: the model wrote “Explore Alfama’s streets” on all three days in Lisbon. The free alternative moved to code too: “Free walk around Alfama: Miradouro de Santa Luzia”.

The honest part is written by code

The bug that taught me the most came up in the first e2e test with the real model. The final narrative said:

“You can expect to spend 366,whichiswithinyourbudgetof366, which is within your budget of 50.”

The chain had correctly computed that the trip exceeded the budget. The model, while drafting the summary, decided the opposite, with all the confidence of a well-built sentence. It’s the same mistake the RAG chatbot made with the World Cup: a guarantee that depends on a 1.5B model having discipline is not a guarantee.

The solution was to split the synthesis in two. The model writes the narrative without seeing any numbers and is forbidden from talking about money; a filter in code removes any sentence that does anyway. The budget paragraph —total, cuts applied, whether it fits or not— is written by code. The CI e2e test fails if the summary contradicts the budget verdict, so the rule doesn’t depend on me remembering to check it.

The same filter drops sentences that name places that aren’t in the plan (“Parque Lagoa”). It’s a heuristic with known limits: “Kyoto Central station” gets through, because “Kyoto” is a known name.

LangGraph, only for orchestration

The first version of the orchestrator was a hand-written loop. It worked, but to showcase the pattern, a recognizable framework and a diagram generated from the real graph were worth the migration. The orchestrator became a LangGraph StateGraph whose state is the shared memory itself:

  • One node per agent, which returns only its own section; it never mutates the state.
  • Reducers for the trace and the token count, so the memory stays append-only.
  • Routing in code functions, with no LLM: after the budget, to conflict if it’s over and iterations remain; otherwise, to synthesis.
  • Live events through LangGraph’s custom stream, relayed to the browser over SSE.

LangGraph is used only for orchestration. The model is still called with the app’s own Ollama client, without LangChain wrappers. The proof that the migration went well was boring, which is how these proofs should be: the 56 existing tests passed unchanged.

Live data: the API that wouldn’t respond

The last phase added facts that change: the real weather for the trip dates, the exchange rate to the local currency and, for cities outside the catalog, real districts and places. All from free sources with no API key: Open-Meteo, European Central Bank rates, Wikidata and OpenStreetMap. None of them is required: if one fails, the chain continues with the catalog or the model and says so in the trace.

For places I started with OpenStreetMap’s Overpass API. Over five e2e runs in CI, I went through this:

  1. The public instance returned 504. I added a second, fallback instance.
  2. Both timed out. If two independent servers fail, the problem isn’t load, it’s the query: it asked for nodes, ways and relations, and resolving relation geometry is expensive. I rewrote it with a bounding box and only nodes and ways.
  3. With the lighter query, 504 and timeout again: about 28 extra seconds per trip.
  4. I switched the primary source to Wikidata, with Overpass as fallback. It responded in 0.5 s.

Valparaíso went from districts invented by the model (“Casa Blanca”, “Playa de Coquimbo”) to real places —the Cathedral, the Open-Air Museum—, with 5 of 5 highlights on their day, and the live step dropped from 28 to 1.3 seconds. When a public dependency fails consistently, switching dependencies is usually cheaper than continuing to negotiate with it.

Real weather also brought an effect I didn’t expect: for catalog cities, the weather note is now written by code from the forecast numbers, and destination research stopped generating. Zero tokens. Real data didn’t just make the demo more honest, it made it cheaper.

What I deliberately didn’t do

  • No human in the loop. The book includes one in its workflows, and it’s a real gap compared to the pattern. For a public demo meant to show automatic resolution, I left it out.
  • No real hotel or flight prices. The APIs that offer them require a key and a commercial agreement. Weather and exchange rates are live; prices are still estimates.
  • No JEV, yet. The interface is ready; access is what’s missing.

Conclusion

Looking back, almost every problem in this project had the same solution: take a decision away from the model. Geography moved to a catalog, the budget to Python, the free alternative to code, the summary’s figures to code, the weather to an API. What’s left for the model is what a 1.5B model does well: writing prose around facts that are already correct.

And that, I think, is what the chapter 7 pattern gains when it leaves the notebook. A deterministic orchestrator, a shared memory that only grows, and a conflict resolved with visible decisions aren’t just a tidy way to coordinate agents. They’re how you know which part of the system is allowed to make things up, and make sure the answer is “only the prose”.

The demo is available at alexisalulema.com/projects: it’s provisioned when you ask for it, runs for a few minutes and destroys itself. The code, with the full log of decisions, measurements and mistakes, is at github.com/alulema/trip-planner.

If JEV gives me access, this post will get a part two.