Live Leaderboard
20 models compared.
Best run per model. Click any column to re-rank, filter by provider, re-weight for your use case, and pick two rows to go head-to-head.
Weighting
20 models · sorted by Balanced composite (high → low) · pick any two rows to compare
| # | vs | Model | PipelineScore ▾ | Tier |
|---|---|---|---|---|
| 1 | openai/gpt-oss-20blocal | 93.2 | TRUNK | |
| 2 | gpt-oss:latestlocal | 91.6 | TRUNK | |
| 3 | unsloth/gemma-4-E4B-it-GGUFlocal | 89.5 | DRIP | |
| 4 | qwen3.6-35b-a3blocal | 87.5 | MAINLINE | |
| 5 | qwen3-coder:30blocal | 86.6 | MAINLINE | |
| 6 | qwen3.8:27blocal | 86.2 | MAINLINE | |
| 7 | unsloth/Ornith-1.0-35B-GGUFlocal | 86.1 | MAINLINE | |
| 8 | unsloth/Ornith-1.0-9B-GGUFlocal | 85.4 | MAINLINE | |
| 9 | muse-glimmer:30blocal | 84.9 | MAINLINE | |
| 10 | qwen3.8-27b@iq3_xxslocal | 84.2 | MAINLINE | |
| 11 | mlx-community/Qwen3.6-35B-A3B-4bitlocal | 83.7 | MAINLINE | |
| 12 | unsloth/Qwen3.8-27B-GGUFlocal | 82.3 | MAINLINE | |
| 13 | /home/thomaskyn/models/Qwen3.6-35B-A3B/Qwen3.6-35B-A3B-UD-IQ3_S.gguflocal | 80.9 | MAINLINE | |
| 14 | gemma4:12b-it-qat_gpulocal | 80.9 | MAINLINE | |
| 15 | unsloth/gpt-oss-20b-GGUFlocal | 79.6 | MAINLINE | |
| 16 | qwen/qwen3.6-35b-a3blocal | 79.3 | MAINLINE | |
| 17 | qwen2.5-coder:7blocal | 77.7 | MAINLINE | |
| 18 | /home/thomaskyn/models/Qwen3.5-4b/Qwen3.5-4B-Q4_K_M.gguflocal | 57.8 | TAP | |
| 19 | /home/thomaskyn/models/Qwen3.5-9b/Qwen3.5-9B-Q4_K_M.gguflocal | 50.5 | TAP | |
| 20 | unsloth/North-Mini-Code-1.0-GGUFlocal | 38.9 | DRIP |
vs