The first post was about measuring before you tune. This one is about the harder question that comes next: how do you know when to stop?
Nothing here was fine-tuned. No weights changed. Every change was to the instructions, the file layout, or one configuration value, and every one was re-measured over multiple runs.
Two changes made it worse, and one of them was my most confident idea
The sampling temperature had been sitting slightly above zero all day. That means the model is designed to answer a little differently every time. I was confident this was the source of the run-to-run wobble we had spent hours carefully measuring, said so out loud, and set it to zero.
The score fell from a median of seven out of eight to five. Two cases that had been nearly perfect began failing every single run. And it did not even make the runs identical — five of eight cases still produced different answers with sampling switched off entirely.
A little randomness was helping. With greedy decoding the model locked onto the same wrong first move every time and never recovered.
It was reverted the same night, with the result written into the config file beside the setting so nobody re-fixes it in six months. The reason that round was worth running is that the theory was extremely plausible and completely wrong, and twenty minutes of measurement was the only thing standing between it and a permanent change.
Where the variance actually was, and how one script found it
Three rounds had been spent theorising about it. What answered the question was one small script over data already on disk: for each case, compare the runs against each other and ask what differed.
Three of eight cases varied in which files the model chose to open. Two varied only in wording with identical tool calls — that is the hardware's own non-determinism and no prompt will remove it. Three were completely stable.
The first group mapped one-to-one onto pass and fail. The model was not choosing badly. It was inconsistently choosing to look things up at all.
The scoring tool was wrong fifteen times, always the same way
Fifteen times in one day the grader reported a failure that was not one, and every single time it made the model look worse than it was.
A correct refusal missed because the model wrote a curly apostrophe. A correct answer missed because the proof we looked for was a number and the model, being a voice system, spelled it out in words. A perfectly good answer counted as a refusal because it contained the phrase "there is no" in the middle of a real sentence.
Two of them are worth singling out. One fix caused the next fault — tightening a pattern to stop it over-matching immediately blinded it to a different correct refusal, and only a round that changed nothing found that. And one fault was caught by a test rather than by a bad score, which was the first time all day the instrument was corrected before it lied rather than after.
If your evaluation says the model got worse, suspect the evaluation first. Ours was wrong more often than the thing it was measuring.
A claim I had to withdraw four hours after making it
Partway through I reported that the model produced zero fabrications across three runs. Across six runs it was not zero.
One of them was textbook. Asked what our notes said about a client that does not exist, the model found a real sentence about retainers in a real file and served it up as that client's line. Every ingredient true; the answer invented. That is the failure that actually costs something, because it arrives looking exactly like a finding.
Three runs was too few to have claimed otherwise, and it is the same sample-size mistake this project made at three separate points in one day — every time in the flattering direction. The bias is not random. It is always toward the answer you were hoping for, which is why the rule has to be mechanical rather than a matter of judgement.
The most useful round changed nothing
The last round re-ran the identical configuration with no changes at all, to find out how much of what remained was irreducible.
It found two grader faults, the withdrawn claim above, and the answer to the real question: the median had sat at seven for five consecutive rounds. The spread moved exactly once — when the model itself changed — and never again.
Nothing left was an instruction problem. Which means more rounds would have been spending real money to chase a floor, and the only reason we knew is that one round deliberately produced no progress.
The rules that came out of it
- Write down what each result would mean before you run it. It costs a minute and it is the only reliable defence against explaining a disappointing result away afterwards.
- Change one thing per round. Two at once and you learn nothing about either.
- Re-derive every old score when you fix the grader. A number from this morning and a number from tonight were measured with different rulers.
- Read the logs before theorising. Three rounds died to a diagnosis that one script over existing data would have killed in a minute.
- Run a round that changes nothing. It is the only one that can tell you the difference between a floor and a problem.
- Keep the failures in the write-up. A record where every idea worked is not a record.
What is actually left
One thing, and it is not technical. Every test case in the suite was written by the person doing the tuning. They test what he thought of and they cannot test what he did not.
That is now the binding constraint — not the model, not the prompt, and not the hardware. Until the cases come from real questions asked in the course of real work, the score measures the imagination behind the test set as much as it measures anything else.
Which is a fittingly humbling place for a day of measurement to end.