Local AI

The pause is the work

An assistant that never makes you wait is one that never stops to check. Across five hundred and ten measured answers, the wrong ones came back twice as fast as the right ones — the pause is not the assistant being slow, it is the assistant doing the work of looking something up before it speaks.

This started as a complaint, and a fair one: why do you pause a lot when inquiring?

The honest answer is that some of that pause is the assistant going and looking something up, and some of it is waste. Those are different problems, and it is worth being able to tell them apart before anyone optimises the wrong one — because the obvious fix for the first one is to stop checking.

The fast answers were the wrong ones

Four horizontal bars, one per response-time bucket, each showing
                the share of answers that were correct. Under 0.8 seconds: 64
                percent correct, 30 wrong of 83. 0.8 to 1.2 seconds: 94 percent,
                4 wrong of 67. 1.2 to 2.0 seconds: 92 percent, 22 wrong of 288.
                Over 2.0 seconds: 99 percent, 1 wrong of 72.
Every answer is a real timed run. The top bar is the one that matters: the answers that came back too fast to have checked anything.

We have been measuring a local model on the jobs four of our specialists actually do — report the state of the running system, track what is open, keep the records straight, read the numbers. Every answer is timed.

510 measured answers.
Correct answers: median 1.20 seconds.
Wrong answers: median 0.60 seconds.
Answered in under a second: 66% were right.
Answered in over 1.2 seconds: 94% were right.

That gap is not a coincidence and it is not about the model thinking harder. It is the difference between an answer recalled and an answer looked up. A tool call — opening the file, reading the log, asking the machine — costs about a second. The answers that skipped it came back fast, and were wrong three times as often.

Speed was the tell. Not the wording, not the confidence — the clock. An answer that arrived too quickly to have been checked usually had not been.

One of the fast answers was right, and still wrong

The case that made this concrete asked whether our notes contained a particular document. The model answered in six tenths of a second, every single time, and named the correct file.

It had not looked. Six tenths of a second is not enough time to look. It knew, or it guessed well, and on that day it happened to be right.

We count that as a failure, and the reason is the whole point of this post: a right answer that was not checked is the same behaviour as a wrong one. It will be right until the day the file is renamed, and nothing about how it arrives will change on that day. You cannot tell the two apart from the outside, which means you cannot trust either.

So what is happening during the pause?

Roughly: working out what would actually answer the question, opening the real file or running the real command, reading what came back, and only then saying something. On a question about a running system that might be four or five separate looks — what is running, what does the log say, what changed, does the claim in the notes still hold.

Each of those is a round trip, and the round trip is the silence.

There is a version of this that never pauses. It is the version that answers from memory, and it is confident, fluent, fast and occasionally inventing. That failure has a particular shape worth naming: it does not look like a guess. It looks like a finding — a real filename, a plausible number, a sentence with the right shape. It is much more expensive than an obvious error, because you act on it.

The tempting shortcut, and why we turned it down

There is a class of model built to remove this problem entirely. Full duplex speech: it listens and talks at the same time, no turn-taking gap at all, sub-second responses. We looked at NVIDIA's, seriously, this week.

It generates its own answers end to end. There is no documented way to hand it text from another system and have it simply speak. So adopting it does not make a careful assistant faster — it replaces the careful assistant with a smaller one that cannot read your files, cannot call a specialist, and cannot hand a hard question to something stronger.

It never pauses because it never checks. That is not the same product with better latency; it is a different product.

Which is a fair trade for some jobs. It is not a fair trade for the job of telling you what is true about your own business.

The pause that IS worth removing

None of that excuses the waste, and there is real waste.

Two examples from our own pipeline. It waits for you to stop talking before it starts transcribing, when it could transcribe while you speak. And it waits for the entire answer to exist before it speaks the first word of it, when the first sentence is usually finished long before the last one.

Neither of those is checking. Both are just queueing, and both are being fixed.

There is a third, and it is the interesting one: a lot of questions never needed the expensive brain at all. Is that service running. How many notes are in that folder. What does the log say. Those are lookups, and a small model on your own hardware does them well — that is what the five hundred and ten measurements above were measuring. Sending those to a large cloud model adds a network round trip to a question that a machine in the room could answer.

The rule this leaves us with

  • Remove the queueing, keep the checking. Every millisecond spent waiting for a turn to end is waste. Every second spent opening the file is the product.
  • Route by what the question needs. A lookup can be answered beside you. A judgement call should not be, and the test for which is which has to be cheap — if deciding needs its own model call, you have rebuilt the pause somewhere less visible.
  • Treat a suspiciously fast answer as suspicious. This is the one that transfers to anyone using any of these tools. If it came back instantly and it involves a fact about your files, your numbers or your systems, ask it what it opened.

What we are not going to do

We are not going to make it faster by making it check less. That is available, it is easy, and it would not be visible to you for weeks — which is exactly what makes it the wrong trade.

The pause is going to get shorter. It is not going to get shorter because the assistant stopped looking things up.

Want this running in your own practice? Let's talk.