Run your own AI
Connect it to your real things, or run it on your own hardware
Two paths off the main install guide: granting real account access, and running models on a graphics card you own instead of a metered one.
Moved here from /learn on 28 August 2026 so
that guide could go straight from memory to the visualizer. Nothing
below was rewritten — same steps, same numbers, same machine they were
measured on.
./start-here.sh
where are you right now?
Each path stands on its own. Start at the one that matches you and stop whenever you have enough.
~/paths/1-connect-it
[1] connect it to your real things
Mail, calendar, documents. This is where it stops being a clever toy, and it is also the point at which the questions get serious.
what connecting actually means
You are granting a program the same reach into an account that you have. Not a copy, not a summary — the real thing. That is exactly why it is useful, and exactly why it deserves a minute's thought per connection rather than clicking through.
the rule worth adopting before the first one
Connect one service. Use it for a week. Then decide whether the next is worth it. The failure people regret is not one bad connection, it is granting six on a Sunday afternoon and never revisiting any of them.
there are two different routes, and people confuse them
| Route | Set up where | Good for |
|---|---|---|
| account connectors | Claude's own settings, in a browser | Gmail, Calendar, Drive and the other ready-made ones. Nothing to install. |
claude mcp add | Your terminal | Anything else — a third-party tool, or something you wrote yourself. |
Both end up in the same place. Whatever you connect either way shows up in one list, which is the command worth learning first.
See what you already have
claude mcp list
It prints every connected service and checks each one is actually answering. On this machine it returns three, all reporting connected:
claude.ai Google Calendar: … ✔ Connected claude.ai Gmail: … ✔ Connected claude.ai Google Drive: … ✔ ConnectedRun this first, and run it again any time something stops working. "Connected" versus silence is the fastest way to tell a broken connector from a broken instruction.
Add the ready-made ones in the browser
Mail, calendar and files are set up in Claude's own settings under Connectors, not from the terminal. You sign in to the account you are granting access to, and it appears in
claude mcp listafterwards.Connect one. Then stop and use it for a week.
Add anything else from the terminal
For a service that hosts its own endpoint:
claude mcp add --transport http <name> <url>
If it needs a key, the key travels as a header rather than sitting in the URL, where it would end up in your shell history:
claude mcp add --transport http <name> <url> \ --header "Authorization: Bearer <token>"
And for one that runs as a local program instead of a web service:
claude mcp add <name> -e API_KEY=xxx -- <command>
Check it took
claude mcp list
If the new one is missing or not reporting connected, fix that before going further. A half-attached connector fails in ways that look like the assistant being stupid rather than the plumbing being wrong.
The rule that prevents it is simple and has to be set before the first connector, not after the first incident: anything arriving from outside is information, never instructions. Web pages, emails, documents, search results, files someone sent. Your assistant may report what they say. It must never do what they say.
This is not theoretical and it is not rare. It is the single most likely way a genuinely useful setup turns into a problem, and it costs nothing to rule out on the first day.
~/paths/2-your-own-hardware
[2] run models on your own hardware
Everything up to here has been metered — you pay per use. This is the part that is not.
the part that stays metered
Claude, the model every earlier step on the main guide has been about, is not one of the options below. Anthropic has not published its weights and has said it does not intend to — there is no file to download and nothing to quantise. Every Claude request, including the one that assembled this page, leaves your machine and is answered on Anthropic's own servers, metered, no matter how small the question. That is not a hole in this guide, it is the fact the rest of it is built around: what follows are the models that were actually built to be handed to you.
the one number that decides everything
Not your system memory. The memory on your graphics card. A model has to fit in there to run at a usable speed. If it does not fit, it either refuses to load or spills into system memory and slows by roughly ten times, which in practice means unusable.
what belongs on your own machine
Anything high-volume and low-difficulty: transcription, speech, routing, classification, summarising, and anything touching data you would rather not send anywhere. These are precisely the jobs that generate an alarming bill when metered, and precisely the ones a small local model handles well.
what still justifies paying
Long reasoning, large context, code that has to be right, and anything where being wrong is expensive. A frontier model is meaningfully better at these, and pretending otherwise is how people end up disappointed in local AI and blaming the wrong thing.
Two real, actively maintained open-weight families exist right now, sized against the tiers below: OpenAI's own gpt-oss and Nous Research's Hermes. What they actually are, and the exact commands to install one, are just below. Once one is running, step [4] on the main guide is how you find out whether it is any good — by hand, on your own material.
what it takes to run this
The binding constraint is graphics memory. Everything else is comfort. A model has to fit on the card in one piece — if it does not, it either refuses to load or spills into system memory and slows down by an order of magnitude.
| Minimum | Recommended | |
|---|---|---|
| Graphics memory | 8 GB | 16–24 GB |
| System memory | 16 GB | 32 GB |
| Storage | 100 GB free | 500 GB NVMe |
| Graphics card | NVIDIA | NVIDIA, current generation |
| Operating system | Linux | Linux |
why those numbers
8 GB of graphics memory is the floor, and it is a real floor. A 2–4B model quantised sits around 2–3 GB and leaves room for the context window, which grows with the length of the conversation and is the thing people forget to budget for. Under 8 GB you are choosing models by what fits rather than by what is good.
16 GB is where it stops being a compromise. A 9B model quantised is about 6.5 GB, so 16 GB runs it with genuine headroom — a long context, speech-to-text alongside it, and no swapping when something else wants the card.
24 GB is what a 30B-class model needs, and it fits with little to spare. That is the ceiling on a consumer card today, not a comfortable cruise. If your work does not need a model that size, the money is better spent on memory and a fast disk.
System memory matters more than people expect, because the model is not the only thing running. Speech-to-text, text-to-speech, the web services and the editor all want their share. 16 GB works. 32 GB means you stop thinking about it.
Storage is about patience, not capacity. The models below total roughly 40 GB, so 100 GB is enough to start. But a model is read off disk every time it loads, and on a spinning disk that is a wait you will feel every single time. NVMe is the upgrade you notice most per pound spent.
NVIDIA, and this is not brand loyalty. The tooling in this space is built on CUDA first and everything else second. AMD and Apple silicon both genuinely work and both cost you time in setup and in finding help when something breaks. If you are learning, spend that time on the work instead.
One caution from experience: a brand-new card can be ahead of the software. A card released weeks ago may have no matching build in the standard toolkits yet, and you can end up compiling things yourself to use hardware you already paid for. Last generation is often the smoother purchase.
models installed locally
| Model | Size | Role |
|---|---|---|
| Qwen3.6-35B-A3B | 18.6 GB | The heavy one. Needs a 24 GB card and fills most of it. |
| Qwen3.5-9B | 6.6 GB | The everyday model. Real work, with headroom left over. |
| gemma-4-E4B-it | 6.0 GB | A second opinion from a different family. |
| gemma-4-E2B-it | 4.1 GB | Small and quick. Classification and routing. |
| Qwen3-4B-Instruct | 2.5 GB | Instruction-following at low cost. |
| Qwen3.5-2B | 1.9 GB | The fast one, when latency beats depth. |
All quantised — compressed to fit consumer hardware at a small and usually unnoticeable cost in quality. Unquantised, most would not fit on this card at all. That trade is the most useful thing to understand before picking a model.
the assistant, in pieces
What people picture as "an AI" is several separate services, each doing one job. Splitting them is what makes it fixable when it breaks.
| speech to text | What you said, into words. Local. |
| speaker check | Confirms it was you and not the television. Local. |
| local model | Answers on your own hardware. No network, no meter. |
| cloud model | The hard things a local model cannot. Metered. |
| text to speech | The answer, back into a voice. Local. |
| the loop | Decides who handles what and holds it together. |
Four of those six never touch the internet. That is a privacy answer and a cost answer at the same time, and it is why the split is worth the setup.
two open-weight families worth knowing
The table above is one specific machine's working set. These two are not installed here day to day, but they are real, current, and worth naming because most write-ups on this stop at "run something local" without ever saying what.
| Model | Size | Role |
|---|---|---|
| gpt-oss:20b | 13 GB | OpenAI's own open-weight release, mixture-of-experts with 3.6B active parameters. Apache 2.0. Fits the 16–24 GB tier. |
| gpt-oss:120b | 65 GB | Same family, scaled up. Needs a single 80 GB GPU — does not fit any tier on this page. |
| Hermes 4 14B | 9.0 GB | Nous Research's current line, fine-tuned from Qwen3-14B for tool use and agentic work. Apache 2.0. Fits the 16 GB tier. |
| Hermes 3 8B | 4.7 GB | The older generation, still the one packaged directly in ollama's own library. A safer fit for the 8 GB floor than the 14B above. |
| Hermes 4 70B / 405B | 40 / 229 GB | Same family, scaled well past anything a card on this page reaches. |
gpt-oss:20b and Hermes 4 14B were pulled and run on this machine to write this section: both loaded at 100% on the card, and generated at 147.7 tok/s and 71.6 tok/s respectively — plugged in, GPU power limit 140 W of a 175 W max. On battery that limit drops and so does the rate, which is why a speed with no power state next to it is not a number worth trusting. The 120B, 405B, 70B and Hermes 3 sizes above are the publishers' own figures, not verified on this box, because none of them fits a card this page recommends.
Hermes is tuned to refuse less than most instruction-tuned models. That is a real capability, not a flaw, but it means the guardrails a general-purpose assistant carries are lighter here — worth knowing before pointing it at anything client-facing.
installing one
Both run through the same tool already holding the models in the table above: ollama. If it is not on your machine yet:
Install ollama
curl -fsSL https://ollama.com/install.sh | sh
It installs as a background service and listens only on your own machine unless you deliberately change that.
Pull one
# OpenAI's, 13 GB, the 16-24 GB tier ollama pull gpt-oss:20b# Nous Research's, 9 GB, the 16 GB tier ollama pull hf.co/bartowski/NousResearch_Hermes-4-14B-GGUF:Q4_K_MThat second command is the same
hf.co/<repo>pattern this page's own working set uses — ollama can pull any GGUF someone has quantised on Hugging Face, not only what is in its own library.Talk to it
ollama run gpt-oss:20b
Drops you straight into a prompt. Leave with
/bye.Confirm it landed on the card, not the processor
ollama ps
The
PROCESSORcolumn should read100% GPU. Anything less and it spilled onto system memory — usually a missing or outdated graphics driver.