The machine is a laptop with an RTX 5090 Laptop GPU and 24 GB of video memory. That last number is the one that decides everything, so hold onto it. The card is not empty when we start, either: our voice setup is already sitting in about 4.3 GB of it, which leaves about 19.7 GB free. You cannot spend all of that on a model's weights, though — running it also needs working memory for the context window, and that comes out of the same pool — so the practical ceiling before things get tight is closer to 18 GB. That is a realistic starting point, not a lab with the card scrubbed clean.
Every model was asked the same thing — a 256-token prompt, run through ollama — and we recorded how fast it produced words, how much memory it took, and how long it made us wait for the first one. Speed is the median of several runs, not a lucky one.
How to read a number like "160 tokens a second"
Four numbers describe a local model in daily use. None of them mean anything until you say them out loud in human terms, so here is the translation we use.
Tokens per second is how fast it feels as it types. A person reads at roughly five words a second, which is about seven tokens a second. So seven is reading speed. Anything over 100 tok/s is far faster than you can keep up with — the answer is simply there. Even 29 tok/s is about four times reading speed; it still beats you, but slowly enough that you watch it think.
Resident VRAM is whether it fits on your card — and what is left for everything else. This is the footprint while the model is running — its weights plus the working buffers the context window needs — not the size of the file on disk. It is why a model that is about 17 GB on disk can show 19 GB resident: the extra is the room the conversation itself takes up. A model that sits at 18 GB fits in our free space with a little air. One that climbs to 21 GB is already borrowing from the CPU, and what that costs is finding two.
Load time is the cold wait before the first word. The seconds between "go" and the model being ready. Small models are ready in about four seconds; the larger ones take seven to thirteen. You pay it once per session, not per question, but it is the difference between a tool that feels instant and one you have to wake up.
Active versus total parameters is why a bigger model can be the faster one. Some models are "mixture of experts": they are large on disk but only fire a small slice of themselves for each word. A 35-billion-parameter model that only activates 3 billion per token does far less work per word than an 8-billion-parameter model that fires all of itself every time. Keep that in mind for the first finding, because it is the one that looks wrong until you know this.
| Model | Design | Params (total) | VRAM | On GPU | Load | Speed |
|---|---|---|---|---|---|---|
| Qwen3.6-35B-A3B | MoE | 34.7B (~3B active) | 18 GB | 100% | ~11.6s | ~160 tok/s |
| gpt-oss 20B | MoE | 20.9B | 12 GB | 100% | ~7.1s | ~152 tok/s |
| Llama 3.1 8B | Dense | 8B | 9.7 GB | 100% | ~3.9s | ~129 tok/s |
| GLM-4.7-flash | MoE | 29.9B | 21 GB | ~89% | ~11.7s | ~114 tok/s |
| Qwen3.8-27B | Dense | 27.3B | 19 GB | ~93% | ~12.8s | ~29 tok/s |
Two things in that table look like typos and are not. The biggest model is the fastest. And the slowest one is not the biggest. Both are the same lesson told twice.
Finding one: the biggest model was the fastest
The 35B mixture-of-experts model ran at about 160 tok/s. The 8B dense model — less than a quarter of its size on paper — ran at about 129. The large model won, and it was not close.
This is the active-versus-total point from earlier, in the wild. The 35B model only fires around 3 billion parameters for each word it writes. The 8B model fires all 8 billion, every word. So the "small" model is doing more work per token than the "big" one, and it produces words more slowly despite being a fraction of the size. "Bigger" and "slower" are not the same axis, and anyone shopping on parameter count alone will pick wrong.
Finding two: the memory wall is real, but smaller and sneakier than it looks
It is tempting to pin that whole 160-to-29 gap on memory — the fast model fits, the slow one spilled onto the CPU, case closed. That is mostly wrong, and it is worth getting right, because it is the mistake the spec sheets encourage. Most of the gap is the thing we just covered: architecture. The 35B model is a mixture-of-experts firing 3 billion parameters a word; the 27B is dense and fires all of itself. The memory wall is real too, but it is a smaller, quieter penalty layered on top. Here is how we pulled the two apart.
We ran the same 27B dense model twice, changing only one thing: how much conversation we asked it to hold in mind.
Same model, same machine, same prompt. At a small 8k context it fits entirely on the GPU and runs at about 39 tok/s. Open it up to a 64k context and its running footprint climbs to 19 GB, about 7% of it spills onto the CPU and system RAM — far slower for this work — and it drops to about 29. So the spill itself costs roughly a quarter of the speed. That is real, and you get no warning it happened: the model just quietly feels slower, and you would swear something was broken when nothing is. But a quarter is not "off a cliff."
Notice what it took to make that model fit in the first place: we had to cut its context from 64k down to 8k. Fitting is not free. You buy it with the memory the conversation would otherwise use, which means a shorter effective memory. "Does it fit" and "how much can it keep in mind at once" turn out to be the same dial.
The clearest proof that spill is not the main lever is GLM. It spills more than the 27B did — about 11% onto the CPU — and still runs at 114 tok/s. If overflow were what tanked the dense model, the one overflowing harder would be slower, not four times faster. It is faster because it is built to do less work per word. So: architecture is the big lever, worth most of the difference; the memory wall is a real but modest tax, around a quarter, that arrives silently. If you are choosing a model by its specs, weigh both — pick a design that does less work per word for speed, and keep its running footprint, context and all, inside your free VRAM so you never pay the quiet tax by accident.
(One practical footnote from setup: gpt-oss needed reasoning mode turned off in our stack to behave. Small operational quirks like that are normal with local models and worth budgeting an afternoon for.)
Big model or small model? Both give something up
The honest answer to "which should I run" is that it depends on the job, and each end of the range costs you something real.
The large models are hungry. They eat most of a 24 GB card, leaving little for anything else you want running alongside them. They take seven to thirteen seconds to wake up, which is overkill if all you needed was to reformat a list. And being large does not make them honest — they still make things up, confidently, exactly like every other language model. Their advantage is headroom: more nuance, longer and harder chains of reasoning, better behaviour on prompts that would trip a smaller one.
The small models — the 7-to-8B class — are the opposite trade. They fit almost anywhere, wake up in about four seconds, leave most of your card free, and are genuinely "good enough" for a huge share of real work: drafting, summarising, tidying, routing, first-pass triage. What you give up is at the edges. They reason less deeply, hold less of a long conversation in mind, and get brittle on the genuinely hard prompts — the place a bigger model's headroom earns its keep. A small model that is right 90% of the time and instant can be worth more than a large one that is right 94% of the time and makes you wait, or the reverse, depending entirely on what the wrong 6–10% costs you.
Neither end is "better." One is a scalpel you keep in your pocket; the other is a workshop you walk into. Most people who run models locally end up with both and route the easy jobs to the small one.
A note on the specialists
We measured a few task-specific models too, and they behave the way you would now predict. The coding models — a 7B coder at about 136 tok/s in 5.7 GB, and a 30B coder at about 115 tok/s that needs 22 GB and spills — and a 7B vision model at about 134 tok/s in 7.9 GB. Same rules apply: the small ones are quick and fit easily, the large one runs into the same memory wall as everything else. They are a side note here because most people's daily driver is a general-purpose chat model, and that is where we spent the measuring.
What these numbers are, and what they are not
These are our numbers, on one machine, from one prompt, with specific 4-bit quantizations. Say that back to yourself before you lean on any of them. Change the prompt or the amount of context you feed in and the tokens-per-second figures will move. A different card, a different quant, a different amount of memory already in use, and the whole table shifts.
And the biggest caveat of all: we measured speed and footprint, not whether the answers were any good. Fast and small tells you what a model costs to run. It tells you nothing about whether it is right. Answer quality is a separate measurement — a harder and slower one — and we wrote about how we do that in an earlier post. Do not let a speed table pick your model for you.
What we hope you take away is not a ranking. It is a way to read the specs the next time you see them: does it fit on my card with room to spare, how fast does that make it feel, how long is the cold wait, and is it firing all of itself or just a slice? Answer those four for your own hardware and your own work, and you can make the call yourself — which is the only call worth making.