Useful tools · Watching the spend

Would the cheaper one be worse here?

One question sorts most of your AI bill, and it takes four seconds to ask: would the cheaper model actually be worse on this job? Usually the expensive one is running your small, routine tasks only because nobody ever chose otherwise. Ask it per task, and the default stops quietly costing you money.

Your most expensive model is running your searches, your renames and your three small edits, because nobody ever decided otherwise. One question sorts most of it, and it takes four seconds to ask: would the cheaper model's output actually be worse here?

If no, route it down. If yes, that is a real job for the expensive one. That misallocation is where most people quietly bleed a subscription — not on the hard problems, on the hundred small ones.

Three switches that are already in the box

None of these needs installing.

Run the built-in diagnostic in a fresh session, before real work. It flags old skills, forgotten servers, hooks and bloated config. The framing that makes it click: anything it flags is a cost you pay on every single message, whether you use that thing or not. A connected server you touched once in June is still loading its instructions into every message you send.

Write a routing table into your configuration file. Two dials — which model, and how hard it thinks before answering. Search, quick edits and writing on the cheap tier. Building something new on the strong one. Architecture, hard debugging and security review on the strong one at maximum effort. The walkthrough's own advice is that maximum effort burns a great deal for a small gain on most work, and it is right.

Override live when a session turns. You can change the model mid-session without restarting. Bump down for a stretch of easy edits, bump up for the one hard bug, then come back down — instead of paying top rate for the ninety minutes after the hard part is over.

What we found running the audit by hand

We could not invoke the diagnostic — it is a slash command, and our assistant works through a different path — so we read the actual configuration instead.

A pile of hooks, a handful of skills, and a configuration file large enough to be a line item on its own.

The hooks are not the waste, and cutting them would be a mistake. They fire on events rather than sitting in context, and every one of ours is load-bearing — a generic "trim your hooks" pass would take out the safety rails and leave the clutter.

The skills are a fairer target. Several of ours are variations on the same integration, and whether all of them earn their place in every session is a reasonable question. It is cheap to check and we have not checked.

Where we go further than the source

The playbook's default is to keep everything on high effort. Ours flips the burden the other way: the top tier only with a stated reason, and silence means step down. Being confidently wrong about security, money or client-facing words is expensive in a way a slow rename is not — so those get the strong model, and the rename does not.

12 top tier each with a stated reason on file
10 mid tier stepped down from the top tier
1 cheapest gather-and-report only

The rule this bar exists to prove: top tier only with a stated reason — silence means step down. A specialist that drifts into the gold segment without one written down is the failure this chart is built to catch.

One number we are not repeating: the bonus section claims roughly seventy times fewer tokens per search using a knowledge graph. We have that tool installed and we have never measured the claim on our own machine. Do not take that figure from us.

What to do

Ask the question before the next task, out loud: would the cheaper one actually be worse here? Then write the answer into your config as a table, so you only have to decide it once.

Want this running in your own practice? Let's talk.