The bench.
Local LLM inference, measured on hardware I own. Every result comes with its methodology and run log — numbers you can reproduce.
At a glance
sample data2,847
Commits, past year
+18% vs prior year
23
Models benchmarked
on own hardware
41
Posts written
since 2023
9
Years shipping
and counting
Decode throughput
sample data FP8Q4_K_M
source: sample fixture · real runs pending approval · llama.cpp, greedy decode, batch 1
Runs
sample data| Model | FP8 (tokens/sec) | Q4_K_M (tokens/sec) |
|---|---|---|
| Llama 3.1 8B | 214 | 187 |
| Gemma 2 9B | 186 | 165 |
| Mistral 7B | 231 | 202 |
| Qwen 2.5 14B | 142 | 121 |
Methodology
⟨ Your methodology, in your words: build flags, context length, sampling settings, how many runs per number, thermals, what gets discarded. This section is what makes the bench credible. ⟩
Useful bench? Tap the heart demo