We measured a week of assistant work at roughly $2,262 at list rates, and 63% of it was cache reads — not thinking, just carrying long sessions forward. We found that out afterwards, in a log. That is the actual problem: not that the number is large, but that it is invisible until it is historical.
Measured across one week: ~$2,262 at list rates. A live spike in the pale segment means the cache broke and that turn is being paid for twice.
Three lines at the bottom of the window
The status line in this playbook shows context fill as a percentage and a token count, colour-shifting as it climbs — grey under 40%, yellow to 60, orange to 80, red past that. Then it splits the spend two ways.
Cached versus new. Cached tokens are served from the prompt
cache at roughly a tenth of the normal input price, and they stay
cached for five minutes after they are last touched. New tokens are
full price. So a sudden spike on the new line means your
cache broke, and you are paying rack rate for the rest of that turn.
Then by source. Skills, MCP servers, and hook output, each with its own count. That last breakdown is the interesting one, because every MCP server you have connected costs 200–500 tokens simply by being registered, whether you use it that day or not.
Why we care about the hook line specifically
Our own configuration file is big enough to be a line item on its own, and two of our hooks fire at session start and inject open work straight into context. That is precisely the pattern the playbook names as expensive.
We believe that is our largest single line item. We cannot currently prove it, because nothing here reports it. A byte count and an estimate is not a measurement, and this is the instrument that would turn one into the other — and then show whether an audit actually worked.
The risk is genuinely low, which is unusual
One Python file, standard library only, no network, no writes. It reads the session transcript and prints a string. If it fails, the status line renders nothing and your session carries on.
What we have not done
We have not installed it. Nothing like it is configured on our machine — we checked both settings files and neither has a status line key.
Two things would need changing first: a hardcoded macOS path and a personal tag in the script, both trivially removable. And one thing genuinely needs testing. The script reads specific usage field names out of the transcript. If those names have changed, it will not fail loudly — it will quietly display zero, which is worse than displaying nothing at all, because a meter reading zero looks like good news.
We have been caught by that shape before: generated code whose comments were right and whose constants were wrong. Only running it proves it.
What to do
If you use these tools for more than an hour at a time, put a meter
somewhere you can see it — this one or any other. Then watch the
new line. When it spikes, something broke your cache, and
that turn cost about ten times what the previous one did.