Reference

LLM tokens/sec Benchmark Reference

Real, single-stream text-generation speeds for running LLMs locally — across NVIDIA GPUs, Apple Silicon, and DGX Spark. Every number links to the public source it came from. Where no verified benchmark exists, we say so instead of guessing.

Last reviewed 2026-08-23 · Next review 2026-11 · Reviewed quarterly

How to read this: figures are text generation (tok/s) at batch size 1 — the speed you feel when a model replies. A range (e.g. 82–124) means the source measured it across context depths; a single value is a single reported figure. Prompt processing (pp) is how fast the model ingests your input. Numbers are only comparable within the same model × quantization × backend — a CUDA number and a Metal number are different machines running different code.

Llama 7B · Q4_0

HardwareGenerationPromptBackendConfidenceSource
Mac Studio (M3 Ultra, 80-core GPU)92.14 tok/s1471.24 tok/sllama.cpp (Metal)Maintainer-measuredllama.cpp Discussion #4167
Mac (M4 Max, 40-core GPU)83.06 tok/s885.68 tok/sllama.cpp (Metal)Maintainer-measuredllama.cpp Discussion #4167
Strix Halo (Ryzen AI Max+ 395, 128GB)Range across 5 backends.50.59–55.73 tok/s1545.36 tok/sllama.cpp (ROCm/Vulkan)Community-measuredkyuz0 Strix Halo benchmark grid
Mac (M4 Pro, 20-core GPU)50.74 tok/s439.78 tok/sllama.cpp (Metal)Maintainer-measuredllama.cpp Discussion #4167

Llama 3 8B · Q4_K_M

HardwareGenerationPromptBackendConfidenceSource
RTX 4090 (24GB)127.74 tok/sllama.cpp (CUDA)Community-measuredXiongjieDai/GPU-Benchmarks-on-LLM-Inference
RTX 3090 (24GB)111.74 tok/sllama.cpp (CUDA)Community-measuredXiongjieDai/GPU-Benchmarks-on-LLM-Inference
Mac Studio (M2 Ultra, 76-core GPU)76.28 tok/sllama.cpp (Metal)Community-measuredXiongjieDai/GPU-Benchmarks-on-LLM-Inference
Mac (M3 Max, 40-core GPU)50.74 tok/sllama.cpp (Metal)Community-measuredXiongjieDai/GPU-Benchmarks-on-LLM-Inference
Mac (M1 Max, 32-core GPU)34.49 tok/sllama.cpp (Metal)Community-measuredXiongjieDai/GPU-Benchmarks-on-LLM-Inference

Llama 3 70B · Q4_K_M

HardwareGenerationPromptBackendConfidenceSource
Mac Studio (M2 Ultra, 76-core GPU)70B does not fit in a 24GB consumer GPU; unified memory is the enabler here.12.13 tok/sllama.cpp (Metal)Community-measuredXiongjieDai/GPU-Benchmarks-on-LLM-Inference
Mac (M3 Max, 40-core GPU)7.53 tok/sllama.cpp (Metal)Community-measuredXiongjieDai/GPU-Benchmarks-on-LLM-Inference
Mac (M1 Max, 32-core GPU)4.09 tok/sllama.cpp (Metal)Community-measuredXiongjieDai/GPU-Benchmarks-on-LLM-Inference

Qwen3 30B-A3B (MoE) · Q4_K_XL

HardwareGenerationPromptBackendConfidenceSource
RTX 5090 (32GB)234.3 tok/s at 4K context, falling to 110.7 at 32K.110.65–234.3 tok/sllama.cpp (CUDA)Independent-measuredhardware-corner.net RTX 5090 LLM benchmarks

Qwen3 8B · Q4_K_XL

HardwareGenerationPromptBackendConfidenceSource
RTX 5090 (32GB)185.9 tok/s at 4K context, falling to 111.9 at 32K.111.91–185.91 tok/sllama.cpp (CUDA)Independent-measuredhardware-corner.net RTX 5090 LLM benchmarks

Qwen3 14B · Q4_K_XL

HardwareGenerationPromptBackendConfidenceSource
RTX 5090 (32GB)123.8 tok/s at 4K context, falling to 82.4 at 32K.82.35–123.79 tok/sllama.cpp (CUDA)Independent-measuredhardware-corner.net RTX 5090 LLM benchmarks

Qwen3 32B · Q4_K_XL

HardwareGenerationPromptBackendConfidenceSource
RTX 5090 (32GB)61.4 tok/s at 4K context, falling to 43.8 at 32K.43.82–61.38 tok/sllama.cpp (CUDA)Independent-measuredhardware-corner.net RTX 5090 LLM benchmarks

gpt-oss 20B · MXFP4

HardwareGenerationPromptBackendConfidenceSource
Strix Halo (Ryzen AI Max+ 395, 128GB)Range across 5 backends.72.68–79.78 tok/s1812.57 tok/sllama.cpp (ROCm/Vulkan)Community-measuredkyuz0 Strix Halo benchmark grid
DGX Spark (GB10)60.85 tok/s2009 tok/sllama.cpp (CUDA)Maintainer-measuredllama.cpp Discussion #16578

gpt-oss 120B · MXFP4

HardwareGenerationPromptBackendConfidenceSource
DGX Spark (GB10)60.6 tok/s at short context, falling to 51.5 at 8K depth.51.54–60.57 tok/s1956 tok/sllama.cpp (CUDA)Maintainer-measuredllama.cpp Discussion #16578
Strix Halo (Ryzen AI Max+ 395, 128GB)Framework Desktop; same chip as GMKtec EVO-X2 / Beelink GTR9 Pro. Range across 5 backends.51.02–56.61 tok/s719.91 tok/sllama.cpp (ROCm/Vulkan)Community-measuredkyuz0 Strix Halo benchmark grid

Qwen3 Coder 30B-A3B · Q8_0

HardwareGenerationPromptBackendConfidenceSource
DGX Spark (GB10)44.26 tok/s1654 tok/sllama.cpp (CUDA)Maintainer-measuredllama.cpp Discussion #16578

Llama 3.1 8B · Q4_K_M

HardwareGenerationPromptBackendConfidenceSource
RTX 4090 (24GB)91–120 tok/s6697 tok/sLocalScore (llamafile)Community-measuredLocalScore leaderboard
RTX 3090 (24GB)95.7 tok/s3536 tok/sLocalScore (llamafile)Community-measuredLocalScore leaderboard
RTX 5090 (32GB)llamafile trails llama.cpp CUDA on Blackwell — compare only within this table.66.3 tok/s6297 tok/sLocalScore (llamafile)Community-measuredLocalScore leaderboard
RTX 5070 Ti (16GB)34.7–65.5 tok/sLocalScore (llamafile)Community-measuredLocalScore leaderboard
Mac Studio (M3 Ultra, 80-core GPU)62.7–63.3 tok/s1109 tok/sLocalScore (llamafile)Community-measuredLocalScore leaderboard
Mac (M4 Max, 40-core GPU)47.9–55.1 tok/s663 tok/sLocalScore (llamafile)Community-measuredLocalScore leaderboard
DGX Spark (GB10)33.6–34.7 tok/s2010 tok/sLocalScore (llamafile)Community-measuredLocalScore leaderboard
Mac mini (M4 Pro, 20-core GPU)32.5–32.9 tok/s361 tok/sLocalScore (llamafile)Community-measuredLocalScore leaderboard

Known gaps — no number we'd stand behind

These segments have no verified single-stream benchmark we can cite. We list them explicitly rather than fill the cell with a guess.

SegmentStatusWhy
Intel Arc B580Data exists, not reproducible as a single figureCommunity Vulkan-backend runs exist but are sparse and not yet standardized enough to quote. See the llama.cpp Vulkan thread. Source
A100 / H100 / MI250X (datacenter)Only batched-throughput data — not comparableServing benchmarks (vLLM) measure batched throughput, not single-stream tok/s — a different metric than everything else on this page.
CPU-only mini PCs (Beelink, GMKtec, etc.)No verified public benchmarkLocalScore has scattered CPU entries, but no standardized single-stream tok/s coverage for these specific SKUs. Source
RTX Spark / N1X (expected fall 2026)No verified public benchmarkUnreleased. No public benchmarks exist — any number you see quoted elsewhere is a guess.

Not sure which of these fits your workload? Use the GPU Advisor to match a model and budget to hardware, or browse the full hardware catalog.

Disclosure: Some links on this page are affiliate links. We may earn a commission if you make a purchase — at no extra cost to you. Benchmark source links above are not affiliate links.