The goal was narrow: get a model running on our own hardware to pick the right reference document for a job by itself, open it, and answer out of it — reliably enough to trust when the internet is not involved.
We did not fine-tune anything. No weights changed, there is no dataset, and nothing was trained. Every improvement below came from fixing the instructions and the file layout around the model, one variable at a time, and re-measuring after each one.
Measure before you tune. The obvious version is not the useful one
Everyone nods at "measure first." The version that actually matters is narrower and much more annoying: measure how much your measurement wobbles before you believe any of it.
We ran the identical test three times with nothing changed. It scored four, then one, then three. So the run-to-run noise was three points wide on an eight-point scale — which meant several changes we had already congratulated ourselves for were coin flips, and at least one thing we had "fixed" had never been broken.
That single check reframed the whole day. It is why the chart above shows bands and not a line, and it is the step almost everyone skips.
The instrument was wrong more often than the model
Ten times during the day the scoring tool reported a failure that was not one. Every time, it made the model look worse than it was. Two examples, because they generalise:
A correct refusal was scored as a fabrication because the model wrote its apostrophe as a curly quote and the pattern looking for it used a straight one. And a perfectly correct answer scored as never having read the file, because the proof we were looking for was a number — and this model speaks out loud, so it wrote "one thousand five hundred eighty-four" instead of the digits.
Neither is exotic. Both would have sent us tuning a model that was already right. If your evaluation says the model got worse, suspect the evaluation first — and go and read the actual answer before you believe a score of any kind.
Two changes out of five did nothing, and that is the honest part
We rewrote the index the model reads. No effect. We renamed the file it kept grabbing by mistake. No effect, and it pushed the error onto a neighbouring case instead.
Both were aimed at a diagnosis that turned out to be half right. The index really was badly written — but the logs showed the model had opened it exactly once in twelve runs. We had spent two rounds carefully rewording a page it was not reading.
What finally worked was four sentences added to the model's own instructions: read the index first, and do not pick a file just because its name looks like a word in the question. That fixed the main failure, and promptly created a new one — it started answering from the index instead of opening the file the index named. One more sentence closed that.
Prove the document changed the answer, not just that it was opened
Picking the right file is a claim about behaviour. It is not proof the model used what was in it, because it might have known the answer anyway.
So we ran the same eight jobs twice — once with the reference documents reachable, once with the folder moved out of its reach entirely.
Seven of the eight cases pass with the documents and fail every single time without them. That is the difference between "it opened the file" and "the file is where the answer came from," and it is worth the extra ten minutes it costs.
The best practices, as short as they go
- Run the same test three times before you trust one number. If you do not know your noise floor, you cannot tell an improvement from a coin flip.
- Score the two failure modes separately. Making things up, and refusing things it should have answered, are opposite faults. Averaged into one accuracy figure, a model that refuses everything looks excellent.
- Re-score old runs with the current rules. Every time you fix the grader you change what a score means, so a number printed last week and a number printed today were measured with different rulers.
- Write down what each result would mean before you run it. It takes a minute and it is the only reliable defence against explaining a disappointing result away afterwards.
- Change one thing per round. Two changes at once and you learn nothing about either.
- Keep the failures in the write-up. Two of our five changes did nothing. A record where every idea worked is not a record, it is a pitch.
Why this is the first step to going offline, not a detour
Running a model on your own hardware is the easy part. It is an afternoon, and the instructions are all over the internet.
The hard part is the question you get asked immediately afterwards, and the one you should be asking yourself: how do you know it is any good? With a hosted model you are borrowing somebody else's answer to that. The moment you take the work in-house, that borrowed confidence goes away, and nothing replaces it except your own numbers.
Offline is not a switch you flip. It is a threshold you cross when you can say what your model gets right, how often, and how much that varies — and the only way to be able to say it is to have measured it. That is why measurement comes first and not last. It is not the boring prelude to the interesting work. It is the thing that makes going offline a decision rather than a hope.
What is next, and what we have deliberately not done yet
Three things are queued, and none of them has been touched, because doing them today would have tangled them with the results above.
The first is embarrassing and it is going first. The sampling temperature has been set slightly above zero the whole time — which means the model was designed to answer a little differently each run. A good part of the noise we spent the day carefully measuring is a configuration default nobody questioned. Setting it to zero is a one-line change. There is a real catch, though: at zero the repeated-run noise check stops working, because the runs become identical by construction. So the answer is not "turn it off" — it is zero for real use and a small deliberately-varied set kept aside for checking robustness.
The second is trying a larger model that is already sitting on the same machine. That is also a one-line change, and it invalidates every number here, so it means re-running the whole series to compare honestly.
The third is the real ceiling. All eight test cases came out of one head. They test what we thought to test and they cannot test what we did not think of. Until the cases come from real questions and real corrections, the score measures our imagination as much as the model.
Those are the next post.