Local AI

The prompt was never the problem

Rewriting a local model's instructions did nothing measurable; every change that actually moved the score was to a tool the model can reach. Twelve rounds in one morning proved it, and proving the negative took a round deliberately designed to make no progress. The prompt was never the problem — the model's reach was.

The last post ended on an uncomfortable admission. Every test case had been written by the person doing the tuning, so the score measured his imagination as much as it measured the model. That was the binding constraint, and this is what happened when we removed it.

The new test set is not made up. It is the work four of our specialists actually do: report the true state of the running system, keep track of what is open, keep the notes straight, and read the numbers. Ten real questions of the kind they get asked in the course of a normal day. Then thirteen, because ten stopped being hard enough.

Nothing here was fine-tuned. No weights changed. Every round was one change, re-measured over five runs.

Two changes landed in one restart, and only one of them worked

One case had failed every run since the day it was written. Asked whether the notes contained a particular document, the model answered correctly and instantly, and was marked wrong every time.

I rewrote the description of the search tool so it advertised what it answers rather than what it does, restarted, and the score went from eight out of ten to a perfect ten. Twice.

It had nothing to do with my change. The restart also picked up a fix somebody else had written earlier and never loaded — and that was the one that did the work. The tell was in the timing: the model answered in six tenths of a second, which is far too fast to have looked anything up, so whatever fixed it could not have been about how it searched.

Two changes reached the machine in one restart. The round that followed proved nothing at all, and it looked exactly like a success.

So the next round put my wording back to the original and changed nothing else. It scored fifty out of fifty — identical. The rewording was inert, the credit belonged elsewhere, and thirty seconds of reverting is the only reason that is known rather than assumed.

The case that had been failing was never the model's fault

Here is what the real fix was, because it is the most interesting thing that happened all morning.

We run a guard that stops the model naming a file it has not opened — inventing a plausible source is the failure that actually costs you something, because it arrives looking like a finding. The guard was doing its job perfectly. It saw the model name a document with no lookup behind it, and replaced the sentence.

But the model had not needed a lookup. The system hands it a complete, generated, verified list of every document at the start of every single turn. It was reading that list. That is the opposite of inventing a citation.

Two correct mechanisms were colliding, and the scoreboard had been recording the collision as a failure of the model for days.

The fix seeds the guard with that same list — names only. Quoting is untouched: a quotation still has to come from text a tool actually returned. Knowing a file exists and having read it stay different claims.

An honest number, and a wrong answer built on top of it

One case kept failing exactly one run in five, which is the most annoying failure rate there is — often enough to matter, rare enough to look like noise.

The logs settled it in about a minute. The question asks how many times a form has been submitted. Four runs searched the request log for the form's address and got the right answer. The fifth searched for the phrase from the question — which appears in no log line anywhere, because a log holds request lines and not English. The tool returned zero. The model reported zero submissions, with total confidence.

Nothing lied. The tool was asked a fair question and answered it exactly. The number was true and the answer was wrong.

The tempting fix is another sentence in the instructions telling the model to be careful. We had already learned that does not work. Instead the tool changed: a zero now retries on the address inside the phrase, and if it is still zero it says plainly that the log holds request lines rather than prose. Only on a zero, so a genuine zero still reads as zero.

That case has not failed since.

Every instruction change was worthless. Every tool change worked.

This is the finding, and it was not the one we expected.

Two rounds changed the model's instructions. One widened a rule about when to check facts rather than recall them; it made the run-to-run spread worse and was reverted. The other reworded a tool's description; it did exactly nothing, proven by putting it back.

Four rounds changed what the model could reach — what a tool returns, what it explains when it comes back empty, what the guard counts as already seen. All four moved the score, and the score has stayed moved.

Change what the model can reach, not what it is told. Instructions are the first thing everyone reaches for and they were the only thing that never worked.

Then the test set stopped failing, which is a problem

Two rounds in a row scored full marks. That is not a victory, it is an instrument going blind: a suite that cannot fail can no longer tell a good change from a bad one, and every further round of it costs the same as a round that teaches something.

So three harder cases were written, each aimed at something the original ten could not reach:

  • A question where the honest answer is "I cannot tell." Every one of the original ten was answerable, so the whole set rewarded answering and never once rewarded stopping. This one asks about a machine that is genuinely unreachable. It is scored on whether a number appears at all — hedging eloquently and then supplying a figure still hands you something to act on.
  • A question whose premise is false. The existing case asked about something that does exist, so a model that says yes to everything passes it. This one asks about something that does not, and phrases it as though it obviously does, because agreeing with the questioner is the cheapest wrong answer available.
  • A question needing two lookups chained, where the answer to the first is the input to the second.

I wrote down the prediction first, as the rules from last time require: the score would get worse, and that would mean the set was working.

It did not get worse. All three passed every run. I read the actual answers rather than trusting the score, because a case that passes for the wrong reason is worse than no case at all — and they hold up. The refusal explains what it cannot see and offers where to look instead. The false premise gets searched before being denied. The chained one visibly does both steps and gets the count right.

So the ceiling was not the test being soft. On this shape of work, the local model is doing the job.

The case I threw away before it ever ran

The chained question was originally going to ask for the first heading inside the largest document. Before running it I checked the eight largest documents, and in every single one the first heading just restates the filename.

The model is handed all the filenames at the start of every turn. It could have answered by inference and never opened anything — and it would have looked like a passing chained-lookup test. It was replaced with a count that no filename can give you.

A test that passes trivially is worse than a missing one. The missing one does not tell you anything false.

Where it finished

Final round: 13 of 13, on all five runs. Sixty-five out of sixty-five, on the set that was deliberately made harder that morning.

Which answers the question this whole series has been circling. For reporting the state of a system, tracking what is open, keeping records straight and reading numbers out of a log, the work does not need to leave the building.

The rules that came out of it

  • One change per restart, not one change per round. If a restart picks up somebody else's unloaded edit as well, that round proves nothing until you split them.
  • Revert your own fix to see if it mattered. Thirty seconds, and it is the difference between knowing and assuming.
  • Suspect the scoreboard before the model. A case that has failed forever is as likely to be two correct mechanisms colliding as it is to be a fault.
  • Fix the tool, not the instructions. Every time, in this series.
  • When the suite stops failing, the suite is the problem. Full marks twice in a row is a signal to make it harder, not to celebrate.
  • Read the answers, not the score. Especially the passes.
  • Write the prediction down first. Mine was wrong this time, which is exactly why it is worth writing down.

What is left

The set is at its ceiling again, and the same discipline applies: the next honest move is harder questions, not another round of these. Three lookups chained rather than two. Two records that contradict each other, where the only right answer is to say so.

That runs tonight, on the machine's own idle hours, and whatever it finds will be the fourth post — including if it finds that the thing I have just spent a morning being pleased about falls over the moment the questions get genuinely hard.

Want this running in your own practice? Let's talk.