← all posts

Opus is cheaper than Haiku?!

This is exactly why we should do evals 🧪

TL;DR: Sometimes. On small tasks, Opus reading the code itself was cheaper than handing the reading to Haiku. Opus re-reads its cached context at $0.20 per million tokens, while Haiku's brief comes back as new tokens at $8 per million. On a 453k-line repo it flipped, and Haiku readers cut the bill by 48%.

In my hype report I told you to delegate reading to cheaper agents. A day later I ran evals on that tip, and it didn't survive as written. This post is the full lab notebook: what I tested, what broke, and the rules I kept.

The cheap part was cheap

The setup

I had a routing skill, beholdr-prism-mini, that does exactly what the tip said: the main model hands search and reading to a cheaper model, which sends back a short brief with file paths and quotes.

  • Harness: Claude Code, headless, one fresh checkout per run.
  • Models: Claude Opus 5.5 as the main model, Claude Haiku 4.5 as the reader.
  • Repositories: a 29k-line Rust/TypeScript project of mine, and Payload, about 453k lines of TypeScript.1
  • Grading: a fact list written from the code before the first run. Claude wrote the fact lists and graded the answers, not blind.
  • Honesty box: single runs. Two runs of the same setup came out $1.80 and $1.52, so read every difference below with about ±15% noise in mind. 🎲

Finding 1: count dollars, not tokens

Prompt caching changes the maths. On Opus 5.5 with the one-hour cache, a new token entering the context costs $8 per million once; re-reading it later costs $0.20.2 Here is one task, split by token type:

Opus, four-file task Output Cache writes Cache reads Total
Reading alone $0.19 $0.54 $0.08 $0.81
With Haiku readers $0.19 $0.60 $0.06 $0.85

Two-thirds of the bill was cache writes. Delegation cut the cheap part (re-reads) and added to the expensive part, because the brief, the skill's instructions and the main model's double-checking are all new tokens.

The cheap model was cheap. The brief was not. 🥲

Finding 2: small tasks lose, big repositories win

Task Repo size Opus alone Opus + Haiku readers Cost
One file 29k lines $0.30 $0.34 +12%
Four files 29k lines $0.81 $0.91 +12%
Whole repo, diagrams 29k lines $0.97 $0.76 −22%
REST request, end to end 453k lines $1.80 $0.93 −48%

On the big repository, Opus alone made 45 tool calls, and every turn re-read a growing context: 2.9 million cached tokens, a third of its bill. So the cheap re-reads from Finding 1 stop being cheap when there are enough turns. The reader moves that whole loop out of the main context. That is where delegation earns its keep.

The first version of my instructions lost on every task, the whole-repo one included. What fixed it: a tool-call budget for the readers, briefs that quote each cited line, and a rule that the main model re-opens only the claims it doubts.

Finding 3: cheaper is not deeper

On the big repository, both answers covered all 15 facts on the list. But Opus reading alone also found two real bugs along the way; the delegated run found none and was half as long. Briefs are good at maps ("where does this flow go?") and bad at scrutiny ("what is subtly wrong here?").

Delegate the map. Keep the magnifying glass. 🔍

Finding 4: more reading didn't help

What if the readers get more room? I removed the budget.

Big repo Budget of ~15 calls No budget
Reader tool calls 15, 15, 22 48, 34, 42
Reader cost $0.22 $0.60
Total cost $0.93 $1.18
Facts covered 15 / 15 14.5 / 15

Three times the reading, nearly three times the reader cost, no better answer. The briefs are capped anyway, and the misses came from what nobody looked at, not from running out of calls.

Finding 5: the hook that couldn't

A popular alternative: skip subagents and compress tool output with a hook before the model sees it.3 Claude Code documents a PostToolUse field, updatedToolOutput, that replaces a tool's result.4 In my tests on v2.1.289 the hook ran, and Claude Code ignored the replacement for both Read and Bash. Only additionalContext worked, and that adds tokens instead of removing them.

The workaround that does work is rewriting the tool's input before it runs: send large reads through Haiku into a compressed copy, and pipe read-only shell commands through a compressor. Result on the big repository:

Big repo Opus alone Opus + compression hook Opus + Haiku readers
Cost $1.80 $1.82 $0.93
Time 219 s 239 s 171 s
Main-model tool calls 45 60 14

Only two outputs were big enough to compress, because Opus already reads in small slices. After seeing a compressed view, it went back for the exact lines, so it made more calls. Each compression also blocked for about 15 seconds. The hook shrinks single outputs; the cost lives in the number of turns.

The rules I kept

  • Don't delegate small tasks, even when asked: a few known files, one grep, or anything the main model has already read.
  • Don't delegate audits, reviews or bug hunts. They need the detail briefs drop.
  • Delegate wide sweeps across large or unfamiliar codebases, and research that would take many fetches.
  • Give readers a budget, about 15 tool calls, and a method: grep for structure, then read line ranges.
  • Briefs quote what they cite, and the main model re-opens only what it doubts. If a brief has a gap, ask again; never fill it with a guess.

Finding 6: same story on Codex

Was this a Claude Code quirk? I ran the big-repository task again in Codex, with GPT-6.1 Sol as the main model and GPT-6 Luna as the reader. The skill and its worker definitions carried over with two fixes: Codex agent names allow only underscores, and Codex ignored the model in the agent file, so the main model now names it when spawning the worker.5

Big repo Cost Time Facts covered
Claude Code: Opus alone $1.80 219 s 15 / 15
Claude Code: Opus + Haiku readers $0.93 (−48%) 171 s 15 / 15
Codex: Sol alone $0.60 382 s 15 / 15
Codex: Sol + Luna readers $0.36 (−40%) 310 s 15 / 15

Codex costs are API list prices; on a subscription you pay in weekly quota instead.6 Luna did all its reading for $0.05.

Two more things fell out of it:

  • A main model can invent its workers. When a spawn failed during setup, Sol still answered "the agent reports…", with no agent involved. The skill now says: only report what a worker actually returned.
  • My fact list wasn't the truth either. Both Sol answers noted that Payload gives collections a default access rule that only lets admin users create; both Opus answers said any logged-in user may. Sol was right, and my list never asked. 🥲

What I don't know yet

These are single runs on one task per repository. Research-heavy and edit-heavy work, where the skill now has its own worker modes, haven't been measured at all.

And if anyone tells you their setup "cuts agent costs by 66%," ask three things: tokens or dollars, how they checked the answers stayed correct, and how many runs. Then run your own evals, just don't be lazy do it 🤣🧪

If you want to try the skill yourself, beholdr-prism-mini is out on GitHub.

Beholdr Prism Mini: tiny but powerful

Footnotes

  1. payloadcms/payload, pinned at commit 8001944. ↩

  2. Claude Docs, “Prompt caching”. ↩

  3. precc-cli/precc, a PreToolUse hook that rewrites shell commands before they run, with optional output compressors. I found it through a Facebook post claiming 40–66% fewer input tokens; the README doesn't give that figure. ↩

  4. Claude Code Docs, “Hooks reference”. ↩

  5. Codex Docs, “Subagents”. Tested with Codex CLI 0.160. ↩

  6. OpenAI, GPT-6.1 Sol and GPT-6 Luna model pages. ↩

🤖Ask my AI