A 125B Model Now Writes 94 tok/s on a 12GB RTX 5070
Strata runs Qwen3.8-Flash-Next, a 125B MoE, on 12GB consumer cards: 94 tok/s write speed at Q2_0 on an RTX 5070, with an OpenAI/Anthropic-compatible server on localhost.
7 posts
Strata runs Qwen3.8-Flash-Next, a 125B MoE, on 12GB consumer cards: 94 tok/s write speed at Q2_0 on an RTX 5070, with an OpenAI/Anthropic-compatible server on localhost.
Qwen 3.8-27B is the new dense flagship you can run on a single 24-32 GB GPU. It jumps to 73.0 on Terminal Bench 2.1 (up from 63.4), adds native image and hour-scale video understanding, and ships under Apache 2.0. What changed versus Qwen3.6-27B, how to run it at 4-bit, and the catches I would watch.
The --spec-draft-p-min filter in llama.cpp PR #22397 rescues MTP for Qwen3.6-27B: 48.9 tok/s vs 29 tok/s at 2000 tokens on a 24GB card.
llama.cpp renamed the MTP flag on May 13. The old --spec-type mtp is silently ignored. If your tok/s dropped from 140 to 70 you are likely running without speculative decoding.
Three model releases in three weeks moved local AI from 'good enough for hobbies' to 'good enough for production'. Here's what changed and why it matters.
A practical guide to running Qwen3.6-27B on consumer hardware in 2026 — memory requirements per quant level, recommended runners, and the MTP trick that doubles your tokens per second.
Qwen3.6-27B running locally now scores within 10 points of frontier closed models on SWE-bench Verified. The benchmark table, lined up side by side.