PipelineScore
LLM benchmarks · v3 testpack · deterministic · no API key

Benchmark LLMs on YOUR hardware.

The same 34 deterministic tasks, scored entirely on your machine — no judge model, no API key, one 0–100 score. A catalog of 118 models, ranked by runs people actually made on their own hardware. The only public LLM board that ranks where the model runs, not just which model it is.

Every score here is a real run. We seed nothing — which is why the board is small and why you can trust the numbers on it.

$ npx @pipelinescore/cli

That's the whole command. It finds your Ollama / LM Studio / llama.cpp / MLX server, lists your models, auto-detects your hardware, and walks you through the rest.

Models tracked 118Real runs 37Users 7Testpack v3Tasks 34

The model board

Best run per model. Click a column to re-rank, pick two rows to go head-to-head.

How scores are computed →
Weighting
20 models · sorted by Balanced composite (high → low) · pick any two rows to compare
#vsModelPipelineScore ▾Tier
1openai/gpt-oss-20blocal
93.2
TRUNK
2gpt-oss:latestlocal
91.6
TRUNK
3unsloth/gemma-4-E4B-it-GGUFlocal
89.5
DRIP
4qwen3.6-35b-a3blocal
87.5
MAINLINE
5qwen3-coder:30blocal
86.6
MAINLINE
6qwen3.8:27blocal
86.2
MAINLINE
7unsloth/Ornith-1.0-35B-GGUFlocal
86.1
MAINLINE
8unsloth/Ornith-1.0-9B-GGUFlocal
85.4
MAINLINE
9muse-glimmer:30blocal
84.9
MAINLINE
10qwen3.8-27b@iq3_xxslocal
84.2
MAINLINE
11mlx-community/Qwen3.6-35B-A3B-4bitlocal
83.7
MAINLINE
12unsloth/Qwen3.8-27B-GGUFlocal
82.3
MAINLINE
13/home/thomaskyn/models/Qwen3.6-35B-A3B/Qwen3.6-35B-A3B-UD-IQ3_S.gguflocal
80.9
MAINLINE
14gemma4:12b-it-qat_gpulocal
80.9
MAINLINE
15unsloth/gpt-oss-20b-GGUFlocal
79.6
MAINLINE
16qwen/qwen3.6-35b-a3blocal
79.3
MAINLINE
17qwen2.5-coder:7blocal
77.7
MAINLINE
18/home/thomaskyn/models/Qwen3.5-4b/Qwen3.5-4B-Q4_K_M.gguflocal
57.8
TAP
19/home/thomaskyn/models/Qwen3.5-9b/Qwen3.5-9B-Q4_K_M.gguflocal
50.5
TAP
20unsloth/North-Mini-Code-1.0-GGUFlocal
38.9
DRIP
vs

Popular matchups

The rivalries worth settling. Every pair opens a live head-to-head.

Five measures. One number.

Code is executed, reasoning is exact-match, tool use and RAG are JSON-match, speed is measured throughput. No judge model, no rubric, no API key.

Code
28%
Reason
22%
Tool Use
18%
RAG
17%
Speed
15%

The tiers

Every score maps to a tier, named the way pipelines are: from TRUNK (top of the network) down to DRIP.

TRUNK
90–100
MAINLINE
75–89
FEEDER
60–74
TAP
40–59
DRIP
0–39
STEP 01

Point the CLI at your model

Ollama, LM Studio, MLX, llama.cpp — anything OpenAI-compatible. Local runs need no account and no API key.

STEP 02

Tag your hardware

--hardware-tag m3-max-128gb / rtx-4090-24gb / a100-80gb. Same model on different rigs gets ranked separately — that's the point.

STEP 03

Land on the board

A deterministic 0–100 PipelineScore across 34 tasks, computed on your machine, plus a tier badge and a public spot on the hardware-aware leaderboard.