125 billion parameters on a 12 GB graphics card: what Strata does differently
Summary
Strata brings Qwen3.8-Flash-Next, a language model with 125 billion parameters, to an ordinary gaming PC with 12 GB of graphics memory and 32 to 64 GB of RAM. Version 0.1.39 was released on 4 October 2026 for the project, which has only been public since 24 September and already had around 11,000 stars on GitHub on 5 October. This article explains what Strata is, how it works technically, what you need it for and where its limits are.
What is Strata?
Strata is an inference engine with an installer for Windows and Linux, released under the MIT licence. It downloads a large language model, starts it on your own computer and provides it at http://127.0.0.1:8080: in the browser as a chat with a monitor for graphics card, processor and RAM, and for other programs as an interface in the OpenAI and Anthropic formats. Coding agents such as Claude Code or Codex can thus be redirected to the local model with a changed base URL.
The model behind it is Qwen3.8-Flash-Next from the Qwen team. It has 125 billion parameters, of which only 6 billion compute per token, plus an n-gram embedding with 51 billion parameters. The context natively covers 262,144 tokens. Models of this size normally run on servers with hundreds of gigabytes of graphics memory.
How does it fit on a small card?
The trick is sharing the work across the whole PC. The model consists of 24,576 small specialists, the so-called experts, and each word needs only 10 of them. The graphics card handles the parts that compute for every word and fills its remaining memory with the most frequently requested experts. Strata learns which ones those are while it is being used. All experts are also held in RAM. If one is missing on the card, the processor computes it, at the same time as the graphics card. The 28.8 GB n-gram table stays on the SSD; only a few rows are read from it per word.
On top of that comes the principle “guess, then check”: a small helper layer built into the model suggests up to three next tokens, and the large model checks all suggestions in a single pass. Because the large model confirms every word itself, the answer stays the same. According to the project, however, it arrives 1.6 to 1.8 times faster.
This is the logical continuation of what I tried in September with a 35-billion-parameter model on a 16 GB card, described in What I learned from 2.9 billion tokens. With mixture-of-experts models, it is not the model size alone that matters but how cleverly the experts are distributed between graphics card, RAM and SSD.
What you can take away: With such models, more graphics memory brings more than a faster GPU. According to the project, every additional gigabyte holds around 700 more experts on the card. Which model size runs at all, on the other hand, is decided by the RAM.
Why do you need it?
- Privacy: prompts, source code and documents do not leave your own computer.
- Cost: anyone who works a lot with agents quickly uses billions of tokens, because the context is reread at every step. Locally, every token costs only electricity.
- Independence: no provider can change prices, limits or model versions as long as the local model runs.
- Compatibility: existing tools only need a different base URL, no new code.
The added value in numbers
The project measured on an RTX 5070 with 12 GB, a Ryzen 5 7600 and 64 GB of RAM. The fastest size, Q2_0, writes 94 tokens per second, the recommended size IQ2_XS 79 and the best size IQ3_S 53. Strata reads long inputs at 1,600 to 2,650 tokens per second. On an AMD RX 9070 XT with 16 GB, Q2_0 writes 60 and IQ2_XS 52 tokens per second. For comparison: a token is about three quarters of a word; 60 tokens per second is faster than you can read along.
Where the limits are
- Compression costs quality: the fastest values apply to the most heavily compressed sizes with around 2 bits per weight. According to the project, only IQ3_S matches the full model on the published tests.
- The Coder variant keeps only 256 of the 512 experts per layer so that it fits into 32 GB. The quoted 91 percent of the SWE-bench score was not measured by Strata but comes from the variant’s authors. Outside code it is weaker.
- Resources: around 70 GB of download and about 80 GB of free space. At start-up Strata loads 35 to 55 GB into RAM, and the PC may not respond for 1 to 3 minutes.
- One request at a time is the default. Several parallel requests are possible but make each answer slower on a 12 GB card.
- Installation by your own AI: the project suggests letting a coding agent do the setup. That is convenient, but the agent then runs third-party scripts with your rights. Read what setup.sh does first.
- Network access: with --host 0.0.0.0 Strata can be reached from the whole network. Then always set an API key, as the project itself stresses.
- Licences: Strata is under the MIT licence, the model under the Qwen Community License 1.0. Before commercial use, its terms are worth a look.
- Young and fast-moving: the project has been public since 24 September 2026 and releases a new version almost daily. Numbers and operation can change quickly.
What you can take away
- Before installing, free up around 80 GB on an SSD and close memory-hungry programs such as browsers.
- Start with IQ2_XS and only switch to IQ3_S if you have enough RAM.
- Only expose the service to the network with an API key.
- Combine local and cloud: the local model for confidential and high-volume tasks, the cloud for the hardest cases.
Numbers at a glance
- Strata 0.1.39 was released on 4 October 2026; the project has been public since 24 September 2026.
- Qwen3.8-Flash-Next has 125 billion parameters, 6 billion of them active per token, plus 51 billion in the n-gram embedding.
- The model consists of 24,576 experts; each token uses 10 of them and one shared expert.
- At least 12 GB of graphics memory and 32 GB of RAM are required; with 64 GB all sizes run.
- NVIDIA GeForce RTX 20 to 50 and current AMD Radeon cards are supported on Windows and Linux.
References
- Strata on GitHub
- Strata: How does Strata work?
- Strata: model sizes and measurements
- Qwen3.8-Flash-Next on Hugging Face
- llama.cpp / ggml
Remarks
- All speeds are the project’s measurements, not my own. I have not run Strata myself so far.
Links to the original source and the Web Archive open in a new tab.