Skip to content
AI-Native PM
Field Notes

Field Note

How We Budget an Agent Fan-Out Before We Run It

After we let a Fable 5 loop use up most of a week's allowance, every fleet we launch now gets an estimate built from a measured unit, a hard cap set in the same message as the approval, and a pilot before the rest runs. On our largest run so far the unit held and the agent count is what the estimate missed. This is the routine, and what we changed at each miss.

· 6 min read

In July a Fable 5 session offered to test three skills we had written, described the test as "a few dozen small Claude calls," and used most of a week's model allowance before we stopped it. We wrote that up in How We Spent a Week's Tokens on Fable 5's Great Idea, and we have not launched a fleet without a budget since. The largest one so far was the rewrite of the first part of our course in September: seven chapters, seven slide decks, about 200 agents over two days, and roughly 18 million tokens against a cap we set before the first agent ran. This note is the routine we used, including the two places it missed and what we changed each time.

A fan-out is a job split into slices that many agents work on at once. It is the cheapest way to get a large job done, and the fastest way to spend a week's allowance. The routine below is how we keep the second from happening, and every step of it came from a miss.

Ask for the estimate before you approve the plan

Every fleet starts in plan mode, and the plan has to answer these questions before we say yes:

  • What will the tool do, and in what stages?
  • How many agents will each stage launch?
  • How many tokens does that come to?
  • What is the hard cap?

The estimate cannot be a guess. It has to be built from a unit we have measured on this repo, and the unit is tokens per agent. Every agent loads the project instructions, the content rules, and its brief before it does any work, and on this repo that comes to about 90,000 tokens per agent. Agents per stage times 90,000, summed across the stages, is the estimate. The cap sits about a quarter above it, and it is a hard stop, not a number to check afterward.

For the rewrite, the tool put the choice in front of us as a question, printed here as it appeared:

Which budget for the agent fan-out? (Estimates from measured costs on this
repo; the cap is a hard stop, and spending is gated: exemplar audit, then
your review, then a one-chapter pilot, then chapters 3-7.)
Full pipeline: ~13M tokens, 16M cap (Recommended)

The other option was a lean pipeline, one simulated reader instead of two and fewer skeptics on minor findings, at about 9 million with an 11 million cap. We chose the full one.

Gate the spending, so the first miss is a small one

Approval does not release the whole budget. It releases the first gate. For the rewrite, the first gate was the audit of one hand-written exemplar chapter, and the second was our own review of it. The third was a pilot of one more chapter through the full pipeline, and only after that did the remaining chapters run. Nothing past the exemplar ran until we had reviewed it, and nothing past the pilot ran until we had read its cost report. This is the pilot-first rule from Failure modes: how fleets go wrong together, applied to the budget rather than the output.

Spend against the cap at every gateA line chart of cumulative tokens in millions across six gates: plan approved at zero, exemplar audit at 5.47, pilot chapter at 8.19, second pilot at 11.69, chapters 5 to 7 at 17.8, and finished by hand at 18.3. A clay cap line sits at 16.3 from the plan to the second pilot, then steps up to 19 for the rest. A dashed line at 12.8 marks the original estimate, and a gold mark at the second pilot shows the 16.5 re-projection that crossed the cap and forced the decision. Caption: spending stayed under the cap at every gate, and the cap moved once, when we chose to raise it.SPEND AGAINST THE CAP, GATE BY GATE0M5M10M15M20Mestimate 12.8Mcap 16.3Mcap raised to 19M, by decisionre-projected 16.5Mplanapproved5.47Mexemplaraudit8.19Mpilotchapter11.69Msecondpilot17.8Mchapters5 to 718.3Mfinishedby handSpending stayed under the cap at every gate; the cap moved once, when we chose to raise it.

Every run reports its cost, with the reason for the gap

Each gate ends with a cost paragraph: what the run used, what it was estimated at, the running total against the cap, and why the two differ. The exemplar audit used 3.23 million tokens against an estimate of 2.5 million. The simulated readers produced far more findings than planned, every finding went to two skeptics, and each skeptic was another agent. The pilot chapter used 2.72 million against a projection of 1.4 million, for the same reason: the audit still ran 30 agents. Across every run, the cost per agent stayed between about 70,000 and 90,000 tokens. The unit held. What the estimates got wrong, both times, was the number of agents the audits would spawn.

Tokens per agent, four runsFour bars showing thousands of tokens per agent: the exemplar audit at 77 thousand across 42 agents and 3.23 million in total; the pilot chapter at 91 thousand across 30 agents and 2.72 million; the second pilot at 90 thousand across 39 agents and 3.50 million; chapters 5 to 7 at 71 thousand across 55 agents and 3.91 million. All four sit inside a shaded band from 70 to 90 thousand. Caption: the cost per agent held in a narrow band, and the agent count is what the estimates missed.THE UNIT BEHIND THE ESTIMATEtokens per agent, on this repo, across the four audit runs0k25k50k75k100kthe unit: 70k to 90k per agent77kexemplar audit42 agents, 3.23M91kpilot chapter30 agents, 2.72M90ksecond pilot39 agents, 3.50M71kchapters 5 to 755 agents, 3.91MThe cost per agent held in a narrow band; the agent count is what the estimates missed.

Estimate in agents, not tokens. On one repo the cost per agent is stable, and the agent count is what a pilot exists to measure.

When the estimate misses, cut agents, not quality, and pilot again

After the pilot, the tool proposed a leaner audit rather than a bigger budget, and it listed what each cut would drop:

  • No separate read-aloud pass on the slide decks, since two other audits already covered the same lines.
  • The diagram check folded into the fidelity audit.
  • One skeptic lens per batch instead of two, and larger batches.
  • No separate checking agent, since the author already runs the checkers.

Both simulated readers stayed, because they had produced every finding that actually changed a chapter. Then it ran a second pilot on two chapters to measure the leaner configuration before committing the rest. Those used 1.75 million tokens per chapter against a projection of 1.5 million, so the estimate now covered 86 percent of the bill instead of half. A second pilot is what makes the estimate honest again after the first one breaks it.

Crossing the cap is our decision, made from priced options

With the leaner cost measured, the projection for the whole job came to about 16.5 million against the 16.3 million cap, with no margin for a chapter that needed rework. The tool did not quietly continue, and it did not quietly raise the cap. It stopped and gave us three options, each with a price:

  • Raise the cap to 19 million and keep both readers and the recheck for the remaining chapters.
  • Hold at 16.3 million by dropping one reader and the recheck, at about 1.1 million per chapter.
  • Pause, and decide after the second pilot reported.

It said it would not launch the remaining chapters until we chose. We raised the cap, because the two readers had disagreed usefully on both chapters so far and that was worth more than the margin.

Guards for the things that break anyway

The workflow script carries a budget guard: if what remains cannot cover the next chapter, it drops that chapter and says so, rather than running it partway. When the session limit stopped a run in the middle of an audit, we resumed it from its cache. The eighteen agents that had finished were replayed at no cost, and only the unfinished ones ran. When the weekly allowance ended with two passes left, we did those by hand inside the cap, and the hand pass found two story contradictions across chapters that no agent had flagged. The rewrite closed at about 18.3 million tokens against the 19 million cap.

The routine

  • Ask in plan mode: stages, agents per stage, tokens, and a hard cap, before you approve anything.
  • Build the estimate from a measured unit, tokens per agent on this repo, not from a feeling about the job.
  • Set the cap in the same message as the approval, about a quarter above the estimate, as a hard stop.
  • Gate the spend: exemplar, your review, one pilot, then the rest.
  • Read the cost paragraph after every gate: actual, estimate, running total, and the reason for the gap.
  • When it misses, cut agents, not quality, and run a second pilot before you commit the rest.
  • Crossing the cap is the owner's call, made from options with prices on them.
  • Resume from cache, and finish by hand if the allowance ends. Read what any interrupted run returns before you trust it.

Sources