A 125B-parameter model on a consumer GPU sounds like a misprint. A dense model that size needs about 250GB just to hold its weights at 16-bit precision, before you count the KV cache. Qwen3.8-Flash-Next gets away with it because it's a mixture-of-experts model, and Strata, a new C++ and CUDA inference engine, packages it as a one-click install for Windows and Linux. The "8GB" in the headline needs an asterisk, and I'll get to it. First, why this works at all.

Why MoE Makes the Math Work

In a dense model, every parameter takes part in every token. In a mixture-of-experts model, each layer holds a pool of small feed-forward networks (the experts) and a router that picks a few of them for each token. The headline parameter count covers every expert in the pool. The compute per token only covers the ones the router picked.

Qwen's model card puts it at 125B parameters with 6B activated. Each of its 48 layers has 512 experts, and every token goes through 10 routed experts plus 1 shared one. The card also lists 51B of n-gram embeddings and a 4B multi-token-prediction layer on top of the 125B, so the files on disk are bigger than the name suggests.

That 6B is what sets your speed, and it's the part of MoE I care most about after running big models at home. Generating tokens on a large model is limited by memory bandwidth, because every new token means reading the weights it uses out of memory. When I benchmarked DocQuery on one DGX Spark, dense Llama 3 70B fit easily in 128GB of unified memory and still only generated 5.8 tokens per second, almost exactly what the bandwidth math predicts. A MoE model reads only its active experts for each token, so a model with 6B active parameters generates at roughly small-model speed while the full 125B stays available for the router to pick from.

The catch is that the idle experts still have to live somewhere. Active parameters set your speed, and total parameters set your memory bill. Two Sparks were the only way I could fit GLM-5.3-Flash, a 320B MoE with 18B active, at 4-bit: 164 GiB of weights against 128GB per node. Strata attacks the same problem from the other end, with one consumer card and a lot of system RAM.

Where Strata Puts 24,576 Experts

Forty-eight layers of 512 experts comes to 24,576 experts, and they don't fit on a 12GB card at any quantization. Strata's design notes split the model three ways:

That design explains the hardware line. The README asks for an RTX 20-series or newer card with 12 GB of VRAM or more, and the design notes add "(8 GB runs, slowly)." It recommends 64 GB of system RAM (37.6 to 54.8 GB depending on the quant) and about 80 GB of free disk. So 8GB is the floor. The machine that runs this well is a 12GB-plus GPU in a desktop with 64GB of RAM, which is a well-specced gaming PC.

Strata's published numbers come from that kind of box: an RTX 5070 with 12GB, a Ryzen 5 7600, and 64GB of DDR5. On the smallest quant it generates 93 tokens per second at a 4K context and 73.7 at 128K. For a 125B model on one consumer GPU, that's fast. I haven't run it myself, and I'd expect an 8GB card to land well below those numbers, since a smaller expert cache sends more of every token to the CPU.

The engine serves an OpenAI-compatible API at http://127.0.0.1:8080/v1 and an Anthropic-compatible one at /v1/messages, and image input is an option you turn on during setup. The API compatibility is what makes it cheap to try. When I brought GLM-5.3-Flash up on the Sparks, I connected GitHub Copilot Chat to it as a custom OpenAI-compatible endpoint, and once a macOS network permission was sorted out, the client side came down to an endpoint URL and a model name. Strata should be the same amount of work for any tool that already speaks either API.

What I'd Check Before Relying on It

Local inference keeps your prompts on your own machine, and that's the main reason I'd look at this. A 125B-class model on one consumer GPU puts that within reach of a developer workstation, which matters for anyone who can't send client data or proprietary code to a hosted API. I started DocQuery fully local for a smaller reason (no API keys, no cloud bills, no rate limits), and its benchmark is why I'm optimistic here: on focused retrieval over a small corpus, a free 8B model on my MacBook held its own against Azure's gpt-5-mini, with zero hallucinations from either.

Concurrency. Strata answers one request at a time, which rules it out as a drop-in backend for anything that fans out. AgentReview runs three review agents concurrently, which is how 31 seconds of sequential agent work lands in 14 seconds of wall clock. Pointed at Strata, those agents would queue and the review would take about as long as running them one by one. AgentReview also calls Claude through the Anthropic SDK with tool use in every agent loop, so I'd confirm the /v1/messages endpoint handles tool calls before swapping anything. For a single-user chat or a coding assistant, one request at a time is fine.

Quantization. Strata's quants (Q2_0, IQ2_XS, IQ3_XXS, IQ3_S) are all in the 2-to-3-bit range. That's what makes 125B fit in 64GB of RAM, and it's where I'd be most careful. When I picked a recipe for GLM-5.3-Flash, I passed on a 2-bit build because it agreed with the full-precision model on the top token only about 79 percent of the time, which wasn't good enough for a coding assistant. That number belongs to a different model and a different quantizer, so it doesn't transfer to Strata. It does tell me what to measure: run your own task set against an IQ3 quant and a Q2 one, and compare both with a hosted endpoint for the same model if you can get one.

Context length. The model supports 262,144 tokens natively, and Strata's own table shows generation dropping from 93 to 73.7 tokens per second between 4K and 128K of context. Find where your workload sits on that curve, and watch system RAM while you do.

Sustained throughput. A benchmark run tells you the peak. A consumer GPU and a desktop CPU splitting every token for ten minutes straight is a different test, with heat and memory pressure in play. Run it for a while before you trust the first number.

MoE is the reason a model this size fits on hardware people already own. The active parameter count buys you speed and the quantization buys you the fit, and you pay for both somewhere: in system RAM, in bits per weight, and in a server that handles one request at a time. Strata is a good way to try it on a gaming PC. Treat 8GB as the minimum that boots, and plan on 12GB of VRAM and 64GB of RAM if you want it to feel fast.