Guide13 min read

Strata Hardware Requirements: Run 125B Qwen3.8-Flash-Next on a 12GB GPU (2026)

Strata runs the 125B-parameter Qwen3.8-Flash-Next on a 12GB NVIDIA or AMD GPU, but the real requirement is RAM: 32GB is the floor and 64GB unlocks every model size. Intel Arc and Mac aren't supported at all. Here's the pass/fail checklist, what the benchmark numbers actually mean, and the cheapest upgrade for each starting point.

C

Compute Market Team

Our Top Pick

NVIDIA GeForce RTX 5060 Ti 16GB

NVIDIA GeForce RTX 5060 Ti 16GB

$429 – $479
16GB GDDR7448 GB/s4,608

Strata v0.1.38 shipped on 3 October 2026. It's a free, MIT-licensed engine that runs Alibaba's 125-billion-parameter Qwen3.8-Flash-Next on an ordinary gaming PC. The headlines say "any gaming PC." The project's own README is more specific, and the details decide whether your PC qualifies.

Here's the short version:

Strata runs the 125B-parameter Qwen3.8-Flash-Next on a 12GB NVIDIA or AMD GPU, but RAM is the real requirement: 32GB is the floor, 64GB is the recommended target, and Intel Arc and Mac aren't supported.

The rest of this guide covers the pass/fail checklist, what the benchmark numbers do and don't tell you, who's locked out, and the cheapest upgrade from each starting point. For most readers who already own a 12–16GB card, that upgrade is RAM, not a new GPU.

The short answer: can your PC run Strata?

Three gates. Pass all three and Strata will install. Everything below is quoted from the Strata README, read on 4 October 2026.

GateRequirementWhat actually happens
GPUNVIDIA RTX 20/30/40/50, or a listed AMD Radeon card, "with 12 GB of VRAM or more"Hard gate. 8GB cards and Intel Arc fail.
RAM"32 GB or more"; "64 GB runs every size"Soft gate. More RAM unlocks larger, better model sizes.
Disk"about 80 GB free, on an SSD if you can"An SSD mainly makes the first start much faster.
OSWindows 10/11 or Linux, with a current NVIDIA or AMD driverNo macOS build.

Here's the verdict for the configurations we see most often:

  • 8GB GPU (RTX 3070, 4060, 5060), any RAM: no. Below the VRAM floor. Jump to what to buy.
  • 12GB GPU + 32GB RAM: yes, but the installer will point you to the Coder variant. That's fine for code and weaker at everything else.
  • 12–16GB GPU + 64GB RAM: the sweet spot. Every model size fits, including IQ3_S, the best-quality standard variant.
  • 24GB GPU + 96GB+ RAM: room for the largest configurations with other apps still open.
  • Intel Arc (any) or any Mac: not supported. See who's locked out.

What Strata actually does (and why 125B fits in 12GB)

Qwen3.8-Flash-Next is a mixture-of-experts model. Per CellCog's spec breakdown, it has 125B main parameters but only 6B are active per token. There are 48 layers of 512 experts each, and 10 routed experts plus 1 shared expert fire per token. On top of that sit a 51B-parameter n-gram embedding table and a 4B multi-token-prediction head.

Strata uses that sparsity to split the model across your PC. The README describes the layout: your graphics card keeps the experts that are asked for most often, your RAM holds all of them, your CPU works on the rest at the same time, and your SSD holds a big lookup table. The n-gram table, per CellCog, is "explicitly designed to sit in host memory with asynchronous prefetch rather than VRAM." The MTP head drives speculative decoding: a small helper guesses the next few tokens and the big model checks them all at once. The README puts the gain at 1.6–1.8x.

That's why the RAM line matters more than the GPU line: every expert lives in system RAM. The GPU is a cache for the hot ones. For the general theory of expert offload and how llama.cpp's --n-cpu-moe does it by hand, see our guide to running large MoE models on a small GPU.

Hardware requirements, tier by tier

Strata comes in several compressed sizes, from Q2_0 (smallest and fastest) to IQ3_S (best of the standard set), plus specialty builds. The installer recommends one based on your RAM. This table merges the README's requirements with its "Which model should I pick?" table:

TierVRAMSystem RAMRecommended sizeDisk
Minimum12GB32GBCoder (Q2_0 and IQ2_XS too if you have a 24GB card)~80GB free, SSD preferred
Step up12GB+48GBIQ2_XS, or Q2_0 for speed. Larger sizes don't fit.~80GB, SSD
Recommended12–16GB64GBIQ2_XS by default; IQ3_XXS / IQ3_S also fit~80GB, NVMe
Largest configs16–24GB96GB+IQ3_S, or Unsloth UD-Q4_K_XL (experimental)NVMe strongly preferred

The upgrade most readers need is in the RAM column. A 12GB card with 64GB of RAM runs every standard size. A 24GB card with 32GB of RAM is still stuck at the bottom of the list. If you're choosing between spending on VRAM and spending on RAM, read how much RAM you need for local AI first. For this model specifically, RAM wins.

Two practical notes from the README's troubleshooting section. First, if the disk light keeps blinking and generation crawls, you're out of free RAM, and browsers are the usual culprit. Second, the first message of a chat is read in full at about 1 minute per 30,000 tokens. Follow-up messages start in seconds.

Measured speeds: what the numbers do and don't say

The "94 tokens per second" figure in every Strata headline is real, but it comes with three qualifiers. Here are the README's measurements in full:

SizeRTX 5070 (12GB) outputRTX 5070 promptRX 9070 XT (16GB) outputRX 9070 XT prompt
Q2_094 tok/s2,650 tok/s60 tok/s1,160 tok/s
IQ2_XS79 tok/s2,090 tok/s52 tok/s1,110 tok/s
IQ3_XXS62 tok/s1,750 tok/snot reportednot reported
IQ3_S53 tok/s1,620 tok/snot reportednot reported
Coder55 tok/s2,180 tok/s44 tok/s1,420 tok/s

Source: Strata README, "How fast is it?". NVIDIA test PC: RTX 5070, Ryzen 5 7600, 64GB RAM. AMD test PC: RX 9070 XT, Ryzen 9 3900X, 47GB RAM. "Output" is generation speed in a short chat; "prompt" is how fast a 32K-token input is read. Developer-reported; not independently verified by us.

The three qualifiers:

  1. The headline is the smallest size. Q2_0 is aggressive 2-bit compression (see quantization). At IQ3_S, the best standard size, the same RTX 5070 does 53 tok/s. That's still well above reading speed, but it's 44% slower than the headline. How much quality Q2_0 gives up versus IQ3_S on this model is something we haven't measured, so treat it as needing verification.
  2. It's an RTX 5070, not a 5070 Ti. The 5070 is a 12GB card. The RTX 5070 Ti has 16GB and more compute, so it should match or beat these numbers. Nobody has published a 5070 Ti result yet, so we won't print one.
  3. The two test PCs differ in more than the GPU. The AMD rig had an older Ryzen 9 3900X and 47GB of RAM; the NVIDIA rig had a Ryzen 5 7600 and 64GB. Because Strata keeps every expert in RAM and runs some of them on the CPU, the CPU and memory are doing real work. Our reading, which is interpretation and not a measured result, is that the 12GB NVIDIA card out-generates the 16GB AMD card partly because of the platform around it and partly because CUDA builds of llama.cpp-derived engines have historically been more mature than ROCm ones. Don't read the gap as "AMD is 36% slower."

One independent data point backs up the AMD side. Linuxcompatible.org ran v0.1.38 on Debian 13 with ROCm and an RX 7900 XTX (24GB). It saw up to about 120 tok/s on favorable prompts, settling to a more realistic 75 tok/s across mixed queries. Expect that kind of spread on your own machine too: peak and typical tokens per second are different numbers.

The README also estimates that "an RTX 3090 (24 GB) should write roughly 100-140 tokens per second." That's the developer's projection, not a published measurement, and we've labeled it as such everywhere it appears below.

Who's locked out: Intel Arc, Mac, and 8GB cards

This is the part "runs on any gaming PC" skips.

Intel Arc. The Arc B580 has 12GB of VRAM, the same as the RTX 5070 in the benchmark, and it still isn't supported. Strata's GPU list names NVIDIA RTX 20–50 and specific AMD Radeon cards only. If you own a B580, run the model through llama.cpp with manual expert offload (our MoE offload guide covers the flags). Our Arc B580 local-AI review covers what else it's good for. If you're still deciding between the two budget cards, RTX 5060 Ti 16GB vs Arc B580 now has a tiebreaker: one is on Strata's list and the other isn't.

Mac. Strata is Windows and Linux only. The Mac route to this model is Ollama's qwen3.8-flash-next:125b-mlx tag, which is a 105GB download with 256K context. That realistically means a 128GB unified-memory Mac. A Mac Studio M5 Max tops out at exactly 128GB, but its "From $2,499" price is the 36GB base, and you'd need the top memory tier. Even at 128GB it's tight: 105GB of weights leaves little room for the KV cache and macOS.

No discrete GPU, or a Strix Halo box. The Radeon 8060S iGPU in Strix Halo mini PCs like the GMKtec EVO-X2 isn't on Strata's list either. A 128GB EVO-X2 is still a reasonable way to run large MoE models through llama.cpp, and our Strix Halo mini PC guide covers what to expect. It just isn't a Strata machine.

8GB cards. These are below the floor, with no workaround inside Strata. The good news is that the cheapest fix is a mid-range card, not a flagship.

What to buy: three upgrade paths

Match the path to what you already own. The first path is the cheapest and helps the most people.

Path 1: You already have a 12–16GB card. Buy RAM and an NVMe drive.

If you have an RTX 3060 12GB, 4070, 5070, 5060 Ti 16GB, RX 7800 XT, or RX 9070 XT, your GPU already passes. What's holding you back is almost certainly 16GB or 32GB of system RAM.

  • Go to 64GB of RAM. That's the README's "runs every size" line, and it moves you from Coder-only (at 32GB) to IQ3_S. Check how many DIMM slots your board has and what speed it supports before buying. DDR5 prices are still elevated in 2026, and our DRAM shortage buying guide covers when it makes sense to buy now.
  • Put the model on an NVMe SSD. You need about 80GB free. A fast drive shortens first start and matters a lot if you try the experimental 4-bit variant, which reads most of the model from the SSD while answering.

The Samsung 990 Pro 4TB ($289 – $339) is our pick for the model drive. It's PCIe 4.0 with sequential reads up to 7,450 MB/s, and 4TB holds every Strata size plus the rest of your local model library. Smaller capacities will work. Our best NVMe SSD for local AI guide covers cheaper options.

Best for: gamers with a recent 12–16GB card. This is the cheapest way to get the full model experience, and you don't need a new GPU.

Path 2: You're on an 8GB card or have no GPU. Buy a 16GB or 24GB card.

The RTX 5060 Ti 16GB ($429 – $479) is the cheapest new card that clears the floor comfortably. It's an RTX 50-series card with 16GB, 4GB above the minimum, so Strata can keep more hot experts on the GPU. We haven't seen a published Strata benchmark for it. It meets the spec, and the 12GB RTX 5070 result is the closest reference point. Pair it with 64GB of RAM.

The used RTX 3090 ($699 – $999) is the headroom pick. 24GB lets Strata cache many more experts on the GPU, and the README's own estimate for this card is 100–140 tok/s (a projection, not a published measurement). The 24GB also matters at the low end of RAM: the README notes that with a 24GB card, Q2_0 and IQ2_XS run on 32GB of RAM, not just Coder. The trade-offs are higher power draw, the risks of buying used, and no warranty in most cases. We compare them directly in used RTX 3090 vs RTX 5060 Ti for local AI.

Best for: the 5060 Ti 16GB if you want new, efficient, and cheap. The 3090 if you want the most VRAM per dollar and don't mind buying used. AMD buyers should look at the RX 9070 XT, the README's other benchmark card, in our RX 9070 XT vs RTX 5060 Ti comparison and best AMD GPU for local LLMs.

Path 3: You want the largest configs. Buy a 16GB Blackwell card and 96–128GB of RAM.

The 96GB+ tier is where IQ3_S runs with everything else open and where the experimental Unsloth 4-bit becomes an option. On the GPU side, the RTX 5070 Ti ($1,049 – $1,299) and RTX 5080 ($999 – $1,099) are both 16GB Blackwell cards. Check current prices before you buy: at our listed ranges the 5080 can come in cheaper than the 5070 Ti, because 5070 Ti street prices are well above MSRP. Our RTX 5070 Ti vs RTX 5080 comparison covers the rest. If you're comparing the 5080 against a used 24GB card, see RTX 5080 vs RTX 3090, and for the budget alternative, RTX 5080 vs RTX 5060 Ti 16GB.

The RTX 4090 ($1,599 – $1,999) and RTX 5090 ($1,999 – $2,199) will run Strata too, but they're overkill for a 6B-active model where RAM sets the quality ceiling. That money is better spent on 128GB of RAM. Both cards also currently sell well above those ranges at major retailers, so check live prices.

Best for: people who want IQ3_S or 4-bit quality and will keep the machine running long sessions. Our VRAM guide and AI GPU buying guide cover everything this model doesn't need.

Strata vs Ollama vs llama.cpp for this model

Strata is built with parts of llama.cpp/ggml, per its credits section. The choice is mostly about how much setup you want to do and what hardware you have.

ToolBest forHardwareTrade-off
StrataTurnkey GPU + RAM + SSD offload on a gaming PCNVIDIA RTX 20–50, listed AMD Radeon; Windows/LinuxOne model family. Answers one request at a time.
llama.cppManual control, unsupported GPUs (Intel Arc, Strix Halo iGPU)Nearly anythingYou tune expert offload yourself.
Ollama (MLX tag)Mac owners~128GB unified-memory Mac105GB download. Needs top-tier memory.

For coding, Strata's Coder variant removes half of the experts. Per the README, it keeps 91% of the full model's SWE-bench Verified score (as reported by its authors), fits in 32GB of RAM, and is weaker outside code, including on Chinese and other CJK text. Strata exposes an OpenAI-compatible endpoint at http://127.0.0.1:8080/v1 and an Anthropic-style /v1/messages endpoint, so coding agents can point at it directly. See our best GPU for local coding LLMs and local LLM coding setup with Cursor for the editor side.

If you want the Qwen family at a size that fits entirely in VRAM, see our Qwen3.8-27B hardware guide, plus the Qwen3.6 and Qwen3-Coder-Next guides. The local LLM hub collects the rest.

License check before you build a product on it

Strata is MIT-licensed, but the model isn't. Per CellCog, Qwen3.8-Flash-Next ships under the Qwen Community License 1.0, not Apache-2.0. Products above 100M monthly active users or $20M monthly revenue must display the model name, and "Model-as-a-Service or AI Work Assistant" businesses need a separate license from Qwen. Personal and internal use is straightforward. If you're building a product, read the license itself first. This is operational guidance, not legal advice.

The verdict

Strata runs the 125B-parameter Qwen3.8-Flash-Next on a 12GB NVIDIA or AMD GPU, but RAM is the real requirement: 32GB is the floor, 64GB is the recommended target, and Intel Arc and Mac aren't supported.

  • Have a 12–16GB card? Buy 64GB of RAM and put the model on an NVMe drive like the Samsung 990 Pro. This is the cheapest upgrade and the biggest quality jump.
  • On 8GB or no GPU? Get the RTX 5060 Ti 16GB new, or a used RTX 3090 for 24GB of headroom, plus 64GB of RAM.
  • Want the largest configs? Get a 16GB Blackwell card (RTX 5070 Ti or RTX 5080) and 96–128GB of RAM.
  • On Arc or a Mac? Strata isn't for you. Use llama.cpp on Arc, or Ollama's 105GB MLX build on a 128GB Mac.

And keep the benchmark in context: 94 tok/s is a 12GB RTX 5070 running the smallest 2-bit size. It's a strong result, but it isn't what every PC will get. If budget is the main constraint, the AI on a budget hub and our best budget GPU for AI pick up from here.

Frequently Asked Questions

Can I run Qwen3.8-Flash-Next on an RTX 3060 12GB with Strata?

Yes, on paper. The RTX 3060 12GB is an RTX 30-series card with exactly the 12GB of VRAM Strata's README asks for, so it clears the GPU gate. Your RAM decides what runs: with 32GB, Strata's installer points you to the Coder variant; with 48GB, IQ2_XS or Q2_0; with 64GB, every size fits. We haven't measured a 3060 under Strata, so expect it to be slower than the README's RTX 5070 figures (94 tok/s at Q2_0), but we can't put a number on how much slower.

Does Strata work on Intel Arc GPUs like the B580?

No. Strata's supported-GPU list covers NVIDIA GeForce RTX 20/30/40/50 and a named set of AMD Radeon cards. No Intel Arc card is on it, even though the Arc B580 has 12GB of VRAM. Arc owners who want this model should use llama.cpp with manual MoE expert offload instead.

Does Strata work on a Mac?

No. Strata supports Windows 10/11 and Linux only. Mac owners can run Qwen3.8-Flash-Next through Ollama's qwen3.8-flash-next:125b-mlx tag, but that download is 105GB, so in practice you need a 128GB unified-memory Mac.

How much RAM does Strata need?

32GB minimum. Strata keeps every expert of the model in system RAM, and the amount you have decides which compressed size you can run. At 32GB it recommends the Coder variant; at 48GB, IQ2_XS or Q2_0; at 64GB, every size fits; at 96GB or more, IQ3_S or Unsloth's experimental 4-bit with room to keep other apps open. For most readers 64GB is the right target.

Is Qwen3.8-Flash-Next free for commercial use?

Partly. The weights ship under the Qwen Community License 1.0, not Apache-2.0. Commercial use is allowed with conditions: very large products must display the model name, and Model-as-a-Service or AI Work Assistant businesses need a separate license from Qwen. Strata itself is MIT-licensed, but that doesn't change the model's license. Read the license before you ship a product on it. This is operational guidance, not legal advice.

StrataQwen3.8-Flash-NextQwenMoEexpert offloading12GB GPURAMNVMelocal AIhardware requirementsbuying guide
NVIDIA GeForce RTX 5060 Ti 16GB

NVIDIA GeForce RTX 5060 Ti 16GB

$429 – $479

Check Price

More from the blog

Stay ahead in AI hardware

Weekly deals, GPU reviews, and build guides. No spam.

Unsubscribe anytime. We respect your inbox.