The bench.

Local LLM inference, measured on hardware I own. Every result comes with its methodology and run log — numbers you can reproduce.

rig: ⟨ GPU · CPU · RAM — you approve the exact spec line ⟩

At a glance

sample data
2,847
Commits, past year
+18% vs prior year
23
Models benchmarked
on own hardware
41
Posts written
since 2023
9
Years shipping
and counting

Decode throughput

sample data
FP8Q4_K_M
0 125 250 Llama 3.1 8B Llama 3.1 8B · FP8 — 214 tokens/sec 214 Llama 3.1 8B · Q4_K_M — 187 tokens/sec 187 Gemma 2 9B Gemma 2 9B · FP8 — 186 tokens/sec 186 Gemma 2 9B · Q4_K_M — 165 tokens/sec 165 Mistral 7B Mistral 7B · FP8 — 231 tokens/sec 231 Mistral 7B · Q4_K_M — 202 tokens/sec 202 Qwen 2.5 14B Qwen 2.5 14B · FP8 — 142 tokens/sec 142 Qwen 2.5 14B · Q4_K_M — 121 tokens/sec 121

source: sample fixture · real runs pending approval · llama.cpp, greedy decode, batch 1

Runs

sample data
ModelFP8 (tokens/sec)Q4_K_M (tokens/sec)
Llama 3.1 8B 214187
Gemma 2 9B 186165
Mistral 7B 231202
Qwen 2.5 14B 142121

Methodology

⟨ Your methodology, in your words: build flags, context length, sampling settings, how many runs per number, thermals, what gets discarded. This section is what makes the bench credible. ⟩