A 125-billion-parameter model that "usually needs a server" now runs on a normal gaming PC. Strata is an MIT-licensed inference engine that runs Qwen3.8-Flash-Next — a 125B mixture-of-experts model from the Qwen team — on consumer GPUs with as little as 12 GB of VRAM, and it publishes the token/s numbers rather than a vague claim about speed.
The measured numbers on 12 GB and 16 GB cards
The README's benchmark table was run on two ordinary gaming PCs, and it separates "writes answers" (decode) from "reads your prompt" (prefill, measured with a 32K-token document):
| Size | RTX 5070 12GB writes | RTX 5070 reads | RX 9070 XT 16GB writes | RX 9070 XT reads |
|---|---|---|---|---|
| Q2_0 | 94 tok/s | 2,650 tok/s | 60 tok/s | 1,160 tok/s |
| IQ2_XS | 79 tok/s | 2,090 tok/s | 52 tok/s | 1,110 tok/s |
| IQ3_XXS | 62 tok/s | 1,750 tok/s | not stated | not stated |
| IQ3_S | 53 tok/s | 1,620 tok/s | not stated | not stated |
| Coder | 55 tok/s | 2,180 tok/s | 44 tok/s | 1,420 tok/s |
The NVIDIA host was a Ryzen 5 7600 with 64 GB RAM; the AMD host a Ryzen 9 3900X with 47 GB RAM. Most rows were measured on engine 0.1.26, the Q2_0 row on 0.1.36, with 4K answers and 32K prompts. A 24 GB card is faster still: the README puts an RTX 3090 at roughly 100–140 tok/s.
Long-prompt prefill lands at "over 1,000 tokens per second" because the engine reads text in pieces of up to 8,192 tokens at a time. The first message of a chat is read in full — about a minute per 30,000 tokens — while follow-ups start in seconds.
Where the 125B Actually Lives
The trick is not cramming the weights into VRAM. Qwen3.8-Flash-Next has 24,576 experts, and each token needs only about 10 of them. Strata keeps the most-used few thousand experts on the GPU, all of them in RAM, runs the remainder on the CPU, and keeps a lookup table on the SSD. It also runs speculative decoding: a small draft model guesses the next few tokens and the big model verifies them in one pass, which the README credits with a 1.6–1.8× speedup at the same output.
The hardware floor: 12 GB of VRAM (RTX 20/30/40/50 series, or AMD RX 6800 through RX 9070 XT and Radeon AI PRO R9700), 32 GB of RAM, and about 80 GB of free disk — the model download is around 70 GB. Expect the PC to be slow or unresponsive for 1–3 minutes on first start while 35–55 GB is loaded into RAM and locked for the GPU.
Picking a Size Changes What You Get
The same model ships in several compression levels, and the choice is a real quality tradeoff, not just speed:
| Your RAM | Take | Why |
|---|---|---|
| 32 GB | Coder | Fits 32 GB; a coding variant with half the experts removed. Reaches 91% of the full model's SWE-bench Verified score (per its authors), but is weaker outside code, including CJK text |
| 48 GB | IQ2_XS (or Q2_0) | Larger sizes do not fit |
| 64 GB | IQ2_XS, or IQ3_XXS / IQ3_S | Every size fits; IQ3_S is best and slowest |
| 96 GB+ | IQ3_S, or Unsloth UD-IQ4_XS | Room for the near-4-bit sizes |
Unsloth's UD-Q4_K_XL, the closest to full precision, is experimental and streams most weights from the SSD as it answers — 7–8.5 tok/s on a 64 GB machine. OrcaRouter's Uncensored IQ3_XXS is manual setup and not in the installer menu.
The Two Localhost Endpoints Are the Point
This is where I would actually reach for it. Once running, it serves an OpenAI-compatible API at http://127.0.0.1:8080/v1 — any API key, any model name — plus an Anthropic-compatible path at http://127.0.0.1:8080/v1/messages, a Responses API path at /v1/responses for Codex CLI, and an MCP server so an AI assistant can start and stop it. Claude Code points at it with ANTHROPIC_BASE_URL=http://127.0.0.1:8080. There is also a web app on the same port with Chat, a live Monitor of model and GPU/CPU/RAM, and About showing addresses and settings.
You can drive the whole install from an agent: paste a prompt pointing at docs/AI_SETUP.md, and it checks the card, RAM and disk, picks the size, installs, and starts it.
The Catch: One Request, Then Wait
By default Strata answers one request at a time; the rest queue. Setting "parallel": 2 lets it answer several at once, but on a 12 GB card that slows every individual answer. AMD image input works on Linux through the processor and does not work on Windows yet. Exposing it beyond localhost means passing --setup --host 0.0.0.0 --api-key <secret>, and the README is blunt that you should always set a key.
Running a 125B model on a 12 GB card is not the same as running it on a 24 GB card or a server. The quantized sizes are compressed hard, the top end of quality wants 96 GB of RAM, and the default single-request mode is a desktop-shaped assumption. But for a solo engineer or small team that wants a capable local model with a drop-in OpenAI/Anthropic endpoint, the numbers here are specific enough to plan against.
Sources: