Reference
LLM tokens/sec Benchmark Reference
Real, single-stream text-generation speeds for running LLMs locally — across NVIDIA GPUs, Apple Silicon, and DGX Spark. Every number links to the public source it came from. Where no verified benchmark exists, we say so instead of guessing.
Last reviewed 2026-08-23 · Next review 2026-11 · Reviewed quarterly
Llama 7B · Q4_0
| Hardware | Generation | Prompt | Backend | Confidence | Source |
|---|---|---|---|---|---|
| Mac Studio (M3 Ultra, 80-core GPU) | 92.14 tok/s | 1471.24 tok/s | llama.cpp (Metal) | Maintainer-measured | llama.cpp Discussion #4167 |
| Mac (M4 Max, 40-core GPU) | 83.06 tok/s | 885.68 tok/s | llama.cpp (Metal) | Maintainer-measured | llama.cpp Discussion #4167 |
| Strix Halo (Ryzen AI Max+ 395, 128GB)Range across 5 backends. | 50.59–55.73 tok/s | 1545.36 tok/s | llama.cpp (ROCm/Vulkan) | Community-measured | kyuz0 Strix Halo benchmark grid |
| Mac (M4 Pro, 20-core GPU) | 50.74 tok/s | 439.78 tok/s | llama.cpp (Metal) | Maintainer-measured | llama.cpp Discussion #4167 |
Llama 3 8B · Q4_K_M
| Hardware | Generation | Prompt | Backend | Confidence | Source |
|---|---|---|---|---|---|
| RTX 4090 (24GB) | 127.74 tok/s | — | llama.cpp (CUDA) | Community-measured | XiongjieDai/GPU-Benchmarks-on-LLM-Inference |
| RTX 3090 (24GB) | 111.74 tok/s | — | llama.cpp (CUDA) | Community-measured | XiongjieDai/GPU-Benchmarks-on-LLM-Inference |
| Mac Studio (M2 Ultra, 76-core GPU) | 76.28 tok/s | — | llama.cpp (Metal) | Community-measured | XiongjieDai/GPU-Benchmarks-on-LLM-Inference |
| Mac (M3 Max, 40-core GPU) | 50.74 tok/s | — | llama.cpp (Metal) | Community-measured | XiongjieDai/GPU-Benchmarks-on-LLM-Inference |
| Mac (M1 Max, 32-core GPU) | 34.49 tok/s | — | llama.cpp (Metal) | Community-measured | XiongjieDai/GPU-Benchmarks-on-LLM-Inference |
Llama 3 70B · Q4_K_M
| Hardware | Generation | Prompt | Backend | Confidence | Source |
|---|---|---|---|---|---|
| Mac Studio (M2 Ultra, 76-core GPU)70B does not fit in a 24GB consumer GPU; unified memory is the enabler here. | 12.13 tok/s | — | llama.cpp (Metal) | Community-measured | XiongjieDai/GPU-Benchmarks-on-LLM-Inference |
| Mac (M3 Max, 40-core GPU) | 7.53 tok/s | — | llama.cpp (Metal) | Community-measured | XiongjieDai/GPU-Benchmarks-on-LLM-Inference |
| Mac (M1 Max, 32-core GPU) | 4.09 tok/s | — | llama.cpp (Metal) | Community-measured | XiongjieDai/GPU-Benchmarks-on-LLM-Inference |
Qwen3 30B-A3B (MoE) · Q4_K_XL
| Hardware | Generation | Prompt | Backend | Confidence | Source |
|---|---|---|---|---|---|
| RTX 5090 (32GB)234.3 tok/s at 4K context, falling to 110.7 at 32K. | 110.65–234.3 tok/s | — | llama.cpp (CUDA) | Independent-measured | hardware-corner.net RTX 5090 LLM benchmarks |
Qwen3 8B · Q4_K_XL
| Hardware | Generation | Prompt | Backend | Confidence | Source |
|---|---|---|---|---|---|
| RTX 5090 (32GB)185.9 tok/s at 4K context, falling to 111.9 at 32K. | 111.91–185.91 tok/s | — | llama.cpp (CUDA) | Independent-measured | hardware-corner.net RTX 5090 LLM benchmarks |
Qwen3 14B · Q4_K_XL
| Hardware | Generation | Prompt | Backend | Confidence | Source |
|---|---|---|---|---|---|
| RTX 5090 (32GB)123.8 tok/s at 4K context, falling to 82.4 at 32K. | 82.35–123.79 tok/s | — | llama.cpp (CUDA) | Independent-measured | hardware-corner.net RTX 5090 LLM benchmarks |
Qwen3 32B · Q4_K_XL
| Hardware | Generation | Prompt | Backend | Confidence | Source |
|---|---|---|---|---|---|
| RTX 5090 (32GB)61.4 tok/s at 4K context, falling to 43.8 at 32K. | 43.82–61.38 tok/s | — | llama.cpp (CUDA) | Independent-measured | hardware-corner.net RTX 5090 LLM benchmarks |
gpt-oss 20B · MXFP4
| Hardware | Generation | Prompt | Backend | Confidence | Source |
|---|---|---|---|---|---|
| Strix Halo (Ryzen AI Max+ 395, 128GB)Range across 5 backends. | 72.68–79.78 tok/s | 1812.57 tok/s | llama.cpp (ROCm/Vulkan) | Community-measured | kyuz0 Strix Halo benchmark grid |
| DGX Spark (GB10) | 60.85 tok/s | 2009 tok/s | llama.cpp (CUDA) | Maintainer-measured | llama.cpp Discussion #16578 |
gpt-oss 120B · MXFP4
| Hardware | Generation | Prompt | Backend | Confidence | Source |
|---|---|---|---|---|---|
| DGX Spark (GB10)60.6 tok/s at short context, falling to 51.5 at 8K depth. | 51.54–60.57 tok/s | 1956 tok/s | llama.cpp (CUDA) | Maintainer-measured | llama.cpp Discussion #16578 |
| Strix Halo (Ryzen AI Max+ 395, 128GB)Framework Desktop; same chip as GMKtec EVO-X2 / Beelink GTR9 Pro. Range across 5 backends. | 51.02–56.61 tok/s | 719.91 tok/s | llama.cpp (ROCm/Vulkan) | Community-measured | kyuz0 Strix Halo benchmark grid |
Qwen3 Coder 30B-A3B · Q8_0
| Hardware | Generation | Prompt | Backend | Confidence | Source |
|---|---|---|---|---|---|
| DGX Spark (GB10) | 44.26 tok/s | 1654 tok/s | llama.cpp (CUDA) | Maintainer-measured | llama.cpp Discussion #16578 |
Llama 3.1 8B · Q4_K_M
| Hardware | Generation | Prompt | Backend | Confidence | Source |
|---|---|---|---|---|---|
| RTX 4090 (24GB) | 91–120 tok/s | 6697 tok/s | LocalScore (llamafile) | Community-measured | LocalScore leaderboard |
| RTX 3090 (24GB) | 95.7 tok/s | 3536 tok/s | LocalScore (llamafile) | Community-measured | LocalScore leaderboard |
| RTX 5090 (32GB)llamafile trails llama.cpp CUDA on Blackwell — compare only within this table. | 66.3 tok/s | 6297 tok/s | LocalScore (llamafile) | Community-measured | LocalScore leaderboard |
| RTX 5070 Ti (16GB) | 34.7–65.5 tok/s | — | LocalScore (llamafile) | Community-measured | LocalScore leaderboard |
| Mac Studio (M3 Ultra, 80-core GPU) | 62.7–63.3 tok/s | 1109 tok/s | LocalScore (llamafile) | Community-measured | LocalScore leaderboard |
| Mac (M4 Max, 40-core GPU) | 47.9–55.1 tok/s | 663 tok/s | LocalScore (llamafile) | Community-measured | LocalScore leaderboard |
| DGX Spark (GB10) | 33.6–34.7 tok/s | 2010 tok/s | LocalScore (llamafile) | Community-measured | LocalScore leaderboard |
| Mac mini (M4 Pro, 20-core GPU) | 32.5–32.9 tok/s | 361 tok/s | LocalScore (llamafile) | Community-measured | LocalScore leaderboard |
Known gaps — no number we'd stand behind
These segments have no verified single-stream benchmark we can cite. We list them explicitly rather than fill the cell with a guess.
| Segment | Status | Why |
|---|---|---|
| Intel Arc B580 | Data exists, not reproducible as a single figure | Community Vulkan-backend runs exist but are sparse and not yet standardized enough to quote. See the llama.cpp Vulkan thread. Source |
| A100 / H100 / MI250X (datacenter) | Only batched-throughput data — not comparable | Serving benchmarks (vLLM) measure batched throughput, not single-stream tok/s — a different metric than everything else on this page. |
| CPU-only mini PCs (Beelink, GMKtec, etc.) | No verified public benchmark | LocalScore has scattered CPU entries, but no standardized single-stream tok/s coverage for these specific SKUs. Source |
| RTX Spark / N1X (expected fall 2026) | No verified public benchmark | Unreleased. No public benchmarks exist — any number you see quoted elsewhere is a guess. |
Not sure which of these fits your workload? Use the GPU Advisor to match a model and budget to hardware, or browse the full hardware catalog.
Disclosure: Some links on this page are affiliate links. We may earn a commission if you make a purchase — at no extra cost to you. Benchmark source links above are not affiliate links.