The model matters.
So does the harness.
FairBench evaluates models with the same tasks, then changes only the harness. Compare capability and harness effect without sponsored ranks, hidden weighting, or missing receipts.
Benchmarkcore/v1
Task records24/24 complete
Catalog4 reference runs
Funding biasNone · zero ads
Live public results
Model × harness matrix
Across a row: harness effect on one model
Down a column: models under one harness
Showing the FairBench Core v1 model by harness matrix.
Swipe or scroll horizontally to compare harnesses →
| Model / harness | Minimalv1 · pinned | JSONv1 · pinned | Harness effectobserved spread |
|---|---|---|---|
| Qwen2.5 1.5B InstructGGUF Q4_K_M | 45.8%24/24 · reference | 12.5%24/24 · reference | 33.3 ptsmax − min |
| SmolLM2 1.7B InstructGGUF Q4_K_M | 51.7%24/24 · reference | 2.8%24/24 · reference | 49.0 ptsmax − min |
Macro-average across four categories · 95% CI shown in leaderboard view · malformed output and timeouts score zero.
Open full explorer →Harness effect
The wrapper can change the result.
In the founding cohort, the same model can move by more than thirty points when only the answer format and parser change.
See how harnesses are pinned →Descriptive, not causal · two pinned harnesses · 24 tasks
The fairness contract
A result is more than a score.
Model, harness, suite, and runtime are separate versioned inputs. Change one, and FairBench treats it as a new result.
- 01
Same constraints
Identical tasks, token limits, sampling settings, retry policy, and scoring rules.
- 02
Pinned harnesses
Every prompt, parser, and adapter is versioned and hashed. The wrapper is part of the result.
- 03
Receipts, not screenshots
Scores ship with item outputs, parser outcomes, latency, runtime fields, and a reproducible command template.
- 04
Uncertainty stays visible
Coverage, category breakdowns, failures, and confidence intervals live next to the headline score.
The benchmark is public
Keep the method open. Rotate the answers.
The proposal combines open methodology, rotating seasonal suites, timestamped manifest commitments, and contamination canaries.
Post-release optimization cannot disappear. FairBench can label it, preserve historical runs, and make provenance difficult to hand-wave away.
Reproduce the evaluation
Your machine. The same protocol.
The dependency-free runner works with any OpenAI-compatible local endpoint and hashes the complete result bundle.
$ npm run fairbench -- run \ --model qwen2.5:1.5b \ --harness minimal-answer/v1 \ --suite core-v1 \ --out result.jsonIntelligence deserves a fair measurement.