AMD Strix Halo Mini PCs: The Best 128 GB Machines for Running Local AI in 2026
Strix Halo mini PCs pack 128 GB of unified memory into a sub-3-liter chassis — running 70B+ parameter models that no 16 GB discrete GPU can touch. But 128 GB now costs $3,449 and up, not the $1,499 of early 2026. Here are the real prices, the LLM benchmarks, and who should still buy one.
Compute Market Team
Our Top Pick

GMKtec EVO-X2 (Ryzen AI Max+ 395)
$2,199 – $3,649A class of hardware that arrived in early 2026 is still the most memory-rich way to run local AI on a desk: AMD Strix Halo mini PCs. These sub-3-liter machines pack up to 128 GB of unified LPDDR5X memory — with up to 96 GB directly addressable by the integrated GPU. That means you can run 70B+ parameter LLMs that no 16 GB or even 32 GB discrete GPU can touch, in a box that sits quietly on your desk and draws under 120W.
The key insight is simple: for large language model inference, memory capacity matters more than raw GPU speed. A model that doesn't fit in VRAM doesn't run — period. Strix Halo solves that problem on an integrated GPU, which nothing else at this size does.
What changed in 2026: the price. A 128 GB Strix Halo box launched around $1,499. As of 22 September 2026 the cheapest 128 GB configuration we can find is $3,449 (Framework Desktop DIY), with the GMKtec EVO-X2 128 GB at $3,499. LPDDR5X contract prices rose sharply through the year — the same memory squeeze we covered in our 2026 GPU pricing guide — and 128 GB of soldered unified memory is exactly the part that got expensive. Every price below is dated, and the old sub-$1,500 figures no longer exist.
This guide compares every Strix Halo mini PC you can buy right now, benchmarks them on real LLM workloads, and tells you exactly who should buy one — and who's better served by a discrete GPU build or Mac.
What Is AMD Strix Halo and Why Does It Matter for Local AI?
AMD's Ryzen AI Max+ 395 (codenamed "Strix Halo") is the most memory-rich consumer processor ever built. It's not a GPU. It's not a CPU. It's a monolithic APU — a single chip with everything integrated:
- 16 Zen 5 CPU cores (32 threads) — competitive with desktop Ryzen 9 chips
- 40 RDNA 3.5 compute units — roughly equivalent to a Radeon RX 7800 XT in shader count
- XDNA 2 NPU — 50 TOPS for Windows AI features (less relevant for LLM inference)
- Up to 128 GB LPDDR5X unified memory — shared between CPU and GPU, with up to 96 GB allocatable to the GPU partition
- 256-bit memory bus — delivering approximately 218 GB/s of memory bandwidth
The architecture that makes this revolutionary for local AI is unified memory. On a traditional PC, the CPU has system RAM and the GPU has its own VRAM — and these are separate pools. An RTX 5090 has 32 GB of GDDR7 VRAM. If your model is 40 GB, it doesn't fit. End of story.
Strix Halo eliminates this wall. The CPU and GPU share the same physical memory pool, and you can allocate most of it to the GPU. A 128 GB Strix Halo system with 96 GB allocated to the GPU has 3× the effective VRAM of an RTX 5090 and 4× the VRAM of an RTX 4090.
The headline is simply that no consumer GPU on the market gives you this much usable memory for inference. Background on the part itself is in Tom's Hardware's Ryzen AI Max+ 395 review.
If you're new to why VRAM matters so much, our VRAM guide breaks down the math in detail. The short version: a 70B parameter model at Q4 quantization needs roughly 40 GB of VRAM. That rules out every consumer GPU except the RTX 5090 (32 GB — still too small without heavy quantization) and the A100 80 GB ($12,000+). A 128 GB Strix Halo mini PC handles it with room to spare — at $3,449 and up.
Best Strix Halo Mini PCs You Can Buy Right Now
Prices below were read from each vendor's US store on 22 September 2026. We list only configurations we could price directly; several other vendors ship Strix Halo boxes whose current pricing we have not re-verified, so they are left out rather than quoted from memory.
| Machine | Memory | Storage | Price (22 Sep 2026) |
|---|---|---|---|
| GMKtec EVO-X2 | 64 GB | 1 TB | $2,199.99 |
| GMKtec EVO-X2 | 128 GB | 1 TB | $3,499.99 |
| GMKtec EVO-X2 | 128 GB | 2 TB | $3,649.99 |
| Framework Desktop DIY (Max 385) | 32 GB | not included | $1,269 |
| Framework Desktop DIY (Max+ 395) | 64 GB | not included | $1,959 |
| Framework Desktop DIY (Max+ 395) | 128 GB | not included | $3,449 |
Two things to read off that table. First, the jump from 64 GB to 128 GB costs more than the rest of the machine — $1,959 to $3,449 on Framework, $2,199 to $3,499 on GMKtec. That is the memory market, not a vendor markup. Second, GMKtec's figures are current promotional prices; list prices are $2,599.99, $3,999.99 and $4,039.99 respectively. Framework's DIY prices are for the system without storage or an OS, so budget another $100–$250 for an NVMe drive.
GMKtec EVO-X2 — The Turnkey Option
At $3,499.99 for the 128 GB / 1 TB configuration (or $3,649.99 with 2 TB), the GMKtec EVO-X2 is a complete machine out of the box: aluminum chassis, dual USB4, storage and Windows already fitted. If you want 128 GB of unified memory without assembling anything, this is the shortest path.
If you're coming from a Beelink SER8 ($449 – $599), the leap in capability is large — from running 7B models slowly to running 70B models comfortably — though so is the leap in price. For a unit-level teardown — thermals under sustained load, real memory-allocation limits, and per-model tok/s on this exact box — see DataHardware's full GMKtec EVO-X2 review.
Framework Desktop — The Cheapest 128 GB, and the Most Repairable
Framework's DIY Desktop is $3,449 for the 128 GB Max+ 395 board, which makes it the least expensive route to 128 GB we could price on 22 September 2026 — but read the asterisk: that figure is the system without storage or an operating system, so the delivered cost lands close to the EVO-X2 once you add an NVMe drive.
What you get for the assembly work is Framework's usual proposition: standard parts, published schematics, and user-replaceable components in a category where almost everything else is sealed. The memory itself is soldered LPDDR5X on every Strix Halo machine, Framework included — the 128 GB you buy is the 128 GB you keep — but storage, fans, and the side panel are all serviceable.
Strix Halo LLM Benchmarks — What Can You Actually Run?
Benchmarks matter more than specs. The figures below are community-reported numbers for Strix Halo systems with 128 GB of memory (96 GB allocated to the GPU), gathered from the forums and trackers named in the source column. We have not re-run them on our own hardware, and community results vary with kernel, driver and llama.cpp build — treat them as an order-of-magnitude guide. Our own sourced, dated tok/s figures live on the benchmark reference.
| Model | Quantization | VRAM Used | Tok/s (Generate) | Source |
|---|---|---|---|---|
| Llama 3.1 8B | Q4_K_M | ~6 GB | ~45 tok/s | Level1Techs Forums |
| Llama 3.3 70B | Q4_K_M | ~40 GB | ~12 tok/s | Level1Techs Forums |
| Llama 3.3 70B | Q8_0 | ~70 GB | ~8 tok/s | Framework Community |
| DeepSeek R1 (671B distill 70B) | Q4_K_M | ~40 GB | ~11 tok/s | TweakTown |
| Llama 4 Scout (109B MoE) | Q4_K_M | ~60 GB | ~9 tok/s | llm-tracker.info |
| Mistral 7B | Q4_K_M | ~5 GB | ~50 tok/s | Level1Techs Forums |
| Qwen 2.5 32B | Q4_K_M | ~20 GB | ~22 tok/s | Framework Community |
The pattern worth understanding is why Strix Halo can beat a much faster GPU on large models at all. It is not that the integrated GPU is quick — it isn't. It is that a 16 GB card like the RTX 5080 has to quantize aggressively or spill layers to system RAM once a model exceeds its VRAM, and spilling costs far more than the raw compute gap. Strix Halo keeps the whole model in GPU-addressable memory, so it never pays that penalty.
Our own read of those numbers: per-token throughput does not match an RTX 4090 on models that fit in 24 GB, and it isn't close. Above 24 GB the comparison stops being about speed and starts being about whether the model runs at all.
For context on how these models perform on our recommended products, see our DeepSeek R1 local setup guide and Llama 4 hardware guide. For a continuously updated tok/s table across quantizations and context lengths on Strix Halo silicon, DataHardware maintains a dedicated Strix Halo tokens-per-second reference.
What the Numbers Mean in Practice
- 8–12 tok/s on 70B models: Usable for interactive chat. You won't notice the speed difference vs. a cloud API for single-turn conversations. Multi-turn or long-context gets slow.
- 40–50 tok/s on 7B–8B models: Instant-feeling responses. More than fast enough for AI coding assistants, agents, and RAG pipelines.
- ~9 tok/s on Llama 4 Scout (109B MoE): Functional for interactive use. The MoE architecture means the model is smarter than 70B dense models despite similar tok/s.
Strix Halo vs Mac Studio M4 Max for Local AI
This is the comparison everyone wants. Both platforms offer 128 GB of unified memory for large model inference. But the similarities end at the memory spec.
| Spec | Strix Halo Mini PC (128 GB) | Mac Studio M4 Max (128 GB) |
|---|---|---|
| Price (22 Sep 2026) | $3,449 – $3,649 | See current Mac Studio pricing |
| GPU Compute | 40 RDNA 3.5 CUs | 40-core Apple GPU |
| Memory Bandwidth | ~218 GB/s | 546 GB/s |
| GPU-Addressable Memory | Up to 96 GB | Full 128 GB |
| CPU | 16× Zen 5 cores | 16× Apple P/E cores |
| NPU | 50 TOPS (XDNA 2) | 38 TOPS |
| OS | Linux / Windows | macOS only |
| AI Software | ROCm, Vulkan, llama.cpp, Ollama | Metal, mlx, llama.cpp, Ollama |
| Noise | Low (fan-cooled) | Silent (passive under most loads) |
| Expandability | USB4, NVMe (model-dependent) | Thunderbolt 5, NVMe |
Where Strix Halo Wins
Price — but much less than it used to. This was the headline argument for Strix Halo, and through 2026 it eroded. A 128 GB Strix Halo box has gone from roughly $1,499 to $3,449–$3,649, which puts it in the same conversation as a comparably configured Mac Studio rather than at half the price. Apple has also moved the Studio line on to M5 silicon since this comparison was first written, so check current pricing before treating either side as the cheap option. If you are buying several units the remaining gap still compounds — but it is a margin now, not a rout.
Linux native. If your AI workflow runs on Linux — and most serious production AI does — Strix Halo gives you first-class support. ROCm, Docker, CUDA translation layers, and the full Python ML ecosystem work natively. The Mac Studio requires macOS, which means Metal-only GPU access and no ROCm.
Where Mac Studio Wins
Memory bandwidth. At 546 GB/s vs. ~218 GB/s, the M4 Max has 2.5× the memory bandwidth. For LLM inference, memory bandwidth is the primary bottleneck after capacity — it directly determines tok/s. This means the Mac Studio will be noticeably faster on the same model at the same quantization level.
Software maturity. Apple's mlx framework and Metal backend for llama.cpp are well-optimized and Just Work. ROCm on Strix Halo is improving but still requires more manual configuration. For a "download Ollama, run a model, done" experience, the Mac wins.
Silence. The Mac Studio is essentially silent under all workloads. Strix Halo mini PCs are quiet but not silent — fans spin up during sustained inference.
The Verdict
Buy Strix Halo if: you're budget-conscious, you prefer Linux, you want multiple units, or you need 128 GB of memory capacity without paying the Apple tax. Buy Mac Studio if: you value silence, want the best out-of-box experience, need maximum tok/s per dollar of memory bandwidth, or your workflow is already macOS-based. See our Mac mini AI guide and Mac mini alternatives for more Apple vs. AMD comparisons.
Strix Halo vs Discrete GPU Builds for Local AI
The other big question: should you buy a Strix Halo mini PC, or just build a desktop with an RTX 4090 ($1,599 – $1,999)?
| Factor | Strix Halo Mini PC (128 GB) | RTX 4090 Desktop Build | RTX 5090 Desktop Build |
|---|---|---|---|
| Total Cost (22 Sep 2026) | $3,449 – $3,649 | ~$2,200 – $2,800 | ~$2,800 – $3,500 |
| VRAM / GPU Memory | 96 GB (unified) | 24 GB GDDR6X | 32 GB GDDR7 |
| Max Model Size (Q4) | ~150B+ params | ~30B params | ~45B params |
| 7B Model Speed | ~45 tok/s | ~62 tok/s | ~95 tok/s |
| 70B Model Speed | ~12 tok/s | Doesn't fit (offload: ~3 tok/s) | Doesn't fit (offload: ~5 tok/s) |
| Power Draw | ~80–120W system | ~450W GPU + ~150W system | ~575W GPU + ~200W system |
| Noise | Low | Moderate to loud | Loud |
| Size | ~2.5–4 liters | ~30+ liters (ATX case) | ~30+ liters (ATX case) |
| CUDA Support | No (ROCm/Vulkan) | Yes | Yes |
When Strix Halo Wins
- You need to run models larger than 32 GB. Llama 3.3 70B, DeepSeek R1, Llama 4 Scout — these require more VRAM than any consumer GPU offers. Strix Halo runs them natively.
- You want a small, quiet, low-power system. At 80–120W total system power in a 2.5-liter chassis, Strix Halo is 5× more power-efficient and 10× smaller than an RTX 4090 build.
- You're hosting AI agents or always-on services. The power savings compound — running 24/7, a Strix Halo system costs roughly $8/month in electricity vs. $40+/month for an RTX 4090 build. See our AI agent hardware guide for more on always-on deployments.
When Discrete GPUs Win
- You need maximum speed on models that fit in VRAM. An RTX 4090 runs 7B–13B models 40–50% faster than Strix Halo. An RTX 5090 is roughly 2× faster on small models.
- You need CUDA. Training, fine-tuning, and many ML frameworks still require CUDA. ROCm is catching up but isn't at parity yet. See our budget GPU guide for the best value CUDA cards.
- You're doing batch inference or training. Raw FP16/BF16 throughput on NVIDIA tensor cores vastly outperforms RDNA 3.5 compute units.
The simplest decision rule: if your model fits in 24–32 GB, buy a discrete GPU. If it doesn't, buy Strix Halo. For budget VRAM options on the NVIDIA side, an RTX 3090 ($699 – $999) gives you 24 GB of VRAM at a fraction of the RTX 4090 price.
Software Setup — Running LLMs on Strix Halo
Getting LLMs running on Strix Halo is straightforward but requires choosing the right software stack. Here's the current state of the stack:
Option 1: Ollama (Easiest)
Ollama is the fastest path to running models. Install it on Linux or Windows, and it automatically detects Strix Halo's GPU via the Vulkan backend:
# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh
# Pull and run a 70B model
ollama run llama3.3:70b-instruct-q4_K_M
# Verify GPU usage
ollama ps
Ollama handles model downloading, quantization selection, and GPU memory allocation automatically. For most users, this is all you need. See our full Ollama setup guide for detailed instructions.
Option 2: llama.cpp with Vulkan (Best Performance)
For maximum tok/s, compile llama.cpp with the Vulkan backend. This gives you direct GPU access and fine-grained control over memory allocation:
# Clone and build with Vulkan
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
cmake -B build -DGGML_VULKAN=ON
cmake --build build --config Release -j$(nproc)
# Run with GPU layers
./build/bin/llama-cli -m models/llama-3.3-70b-q4_K_M.gguf \
-ngl 99 --ctx-size 8192
The -ngl 99 flag offloads all layers to the GPU. On a 128 GB system with 96 GB allocated to the GPU, this fits any model up to ~90 GB comfortably.
Option 3: ROCm (For Advanced Users)
ROCm support for Strix Halo uses the gfx1151 GPU target, available in recent ROCm 6.x builds. You may need to set environment variables:
# Set the GPU target for ROCm
export HSA_OVERRIDE_GFX_VERSION=11.5.1
export HIP_VISIBLE_DEVICES=0
# Verify detection
rocminfo | grep gfx
ROCm gives you access to PyTorch and other ML frameworks on the GPU, but the Vulkan path in llama.cpp is generally the more stable one for pure inference. Our guidance: start on Vulkan, and move to ROCm when you actually need PyTorch.
Storage Recommendation
Large models eat disk space fast — Llama 3.3 70B at Q4 is roughly 40 GB per file. A fast NVMe drive dramatically improves model loading times. We recommend the Samsung 990 Pro 4TB ($289 – $339) for its 7,450 MB/s sequential reads, which loads a 40 GB model in under 6 seconds.
Who Should Buy a Strix Halo Mini PC?
Ideal For
- Developers running 30B–70B+ models locally. If you're building with Llama 3.3 70B, DeepSeek R1, Llama 4 Scout, or any model that exceeds 24 GB of VRAM, Strix Halo is the most affordable path. Period.
- Small businesses wanting on-premises AI. A $3,499 mini PC running a 70B model replaces API costs that can easily exceed $500/month — a longer payback than it was at $1,499, but still under a year at that burn rate. See our local AI for small business guide for the ROI math.
- AI agent hosting. Always-on AI agents need low power, small footprint, and enough memory to run capable models. Strix Halo checks every box. Our agent hardware guide covers deployment patterns in detail.
- Anyone who wants a capable general-purpose mini PC that also happens to be the best local AI machine in its price class.
Not Ideal For
- Training and fine-tuning. RDNA 3.5 compute units lack the tensor core throughput of NVIDIA GPUs. If you're training models, an RTX 4090 ($1,599 – $1,999) or RTX 5090 ($1,999 – $2,199) is still the right choice.
- Batch inference at scale. If you're serving hundreds of concurrent requests, you need the raw throughput of NVIDIA data-center GPUs, not a mini PC.
- Users who need a mature CUDA ecosystem today. ROCm and Vulkan are improving rapidly, but if your workflow depends on CUDA-only tools (certain PyTorch extensions, TensorRT, etc.), Strix Halo will cause friction.
- Budget under $1,000. If you're spending under $1,000, an RTX 5060 Ti 16GB ($429 – $479) in a budget build or a used RTX 3090 ($699 – $999) gets you into the local AI game. See our budget GPU guide.
Decision Flowchart
- Do you need to run models larger than 32 GB? → Yes: Strix Halo or Mac Studio M4 Max ($1,999 – $4,499)
- Need 128 GB and prefer Linux? → Strix Halo mini PC (Framework Desktop DIY at $3,449, or GMKtec EVO-X2 at $3,499 built)
- Want Strix Halo but not at that price? → the 64 GB tier ($1,959 Framework / $2,199 EVO-X2) still holds a 70B model at Q4 with less context headroom
- Want silence and macOS ecosystem? → Mac Studio M4 Max
- Models fit in 24 GB and need CUDA? → RTX 4090 desktop build
- Models fit in 16 GB and budget is tight? → RTX 5060 Ti 16GB ($429 – $479)
- Want the absolute cheapest 24 GB option? → Used RTX 3090 ($699 – $999)
The Bottom Line
Strix Halo is still the only way to put 128 GB of GPU-addressable memory on a desk in a 2.5-litre box that draws under 120W, and that remains genuinely useful. What is no longer true is the price story it launched with: 128 GB costs $3,449–$3,649 today, not $1,499, and the gap to Apple Silicon has narrowed to the point where the decision turns on operating system, bandwidth and serviceability instead.
The tradeoffs are real: Strix Halo is slower per-token than discrete GPUs on small models, ROCm is less mature than CUDA, and the memory bandwidth gap vs. Apple Silicon means you're leaving some performance on the table. But for the target use case — running models too large for any consumer GPU, in a small, quiet, low-power box — nothing else in the category does it.
If you're running models that fit in 24–32 GB of VRAM, you're still better served by an RTX 4090 ($1,599 – $1,999) or RTX 5090 ($1,999 – $2,199). But if you need a small, low-power machine that can hold the models that actually matter in 2026 — the 70B+ class — Strix Halo is the category, and you should go in knowing what 128 GB now costs.
For a broader look at mini PCs for AI, see our mini PC for LLM guide. For prebuilt options across all form factors, check our best prebuilt AI workstation roundup. And for a deeper dive into the software side, our guide to running LLMs locally covers everything from installation to optimization.
Frequently Asked Questions
How much VRAM does an AMD Strix Halo mini PC have for AI?
Strix Halo mini PCs with the Ryzen AI Max+ 395 support up to 128 GB of LPDDR5X unified memory, with up to 96 GB allocatable directly to the integrated GPU. This is 4× the VRAM of an RTX 4090 (24 GB) and 3× the VRAM of an RTX 5090 (32 GB), making it the most memory-rich consumer platform for running large language models locally.
Can a Strix Halo mini PC run a 70B parameter model?
Yes. A 128 GB Strix Halo mini PC with 96 GB allocated to the GPU can comfortably run Llama 3.3 70B at Q4 quantization (~40 GB), and even at Q8 (~70 GB). Community benchmarks from Level1Techs show approximately 10–14 tok/s on 70B models — slower than an RTX 4090 on small models, but the RTX 4090 can't load 70B at all without heavy quantization or offloading.
Is Strix Halo better than a Mac Studio for local AI?
Less clear-cut than it was. As of September 2026 a 128 GB Strix Halo box runs $3,449–$3,649 — Framework Desktop DIY at $3,449, GMKtec EVO-X2 at $3,499 — after LPDDR5X memory prices climbed through 2026. That is close enough to a comparably configured Mac Studio that price alone no longer settles it. Strix Halo runs Linux natively on standard, repairable parts; Apple Silicon has roughly 2.5x the memory bandwidth and a more mature toolchain (Metal, MLX). Choose on operating system and bandwidth, not on price.
What software do I need to run LLMs on Strix Halo?
The easiest path is Ollama with the Vulkan or ROCm backend — it works on both Linux and Windows. For maximum performance, use llama.cpp compiled with the Vulkan backend on Linux. ROCm support for Strix Halo uses the gfx1151 target, available in recent ROCm 6.x builds.
Should I buy a Strix Halo mini PC or build a desktop with an RTX 4090?
Buy Strix Halo if you need to run models larger than 24 GB (30B–70B+ parameters) in a compact, quiet form factor — but note that a 128 GB box now starts at $3,449, so the decision is about memory capacity, not about saving money. Build an RTX 4090 desktop ($1,599 – $1,999 for the GPU alone, plus $500+ for the rest) if you need maximum tok/s on models that fit in 24 GB, or if you need CUDA for training and fine-tuning workloads. Strix Halo is about memory capacity; discrete GPUs are about raw throughput.