Independent evaluation · Protocol v0.1

The model matters.
So does the harness.

FairBench evaluates models with the same tasks, then changes only the harness. Compare capability and harness effect without sponsored ranks, hidden weighting, or missing receipts.

Benchmarkcore/v1

Task records24/24 complete

Catalog4 reference runs

Funding biasNone · zero ads

Live public results

Model × harness matrix

Across a row: harness effect on one model

Down a column: models under one harness

Results
Suite FairBench Core v1 / Tasks 24 / Attempts 1

Showing the FairBench Core v1 model by harness matrix.

Swipe or scroll horizontally to compare harnesses →

FairBench Core v1: model by harness score matrix across 24 tasks.
Model / harnessMinimalv1 · pinnedJSONv1 · pinnedHarness effectobserved spread
Qwen2.5 1.5B InstructGGUF Q4_K_M45.8%24/24 · reference12.5%24/24 · reference33.3 ptsmax − min
SmolLM2 1.7B InstructGGUF Q4_K_M51.7%24/24 · reference2.8%24/24 · reference49.0 ptsmax − min

Macro-average across four categories · 95% CI shown in leaderboard view · malformed output and timeouts score zero.

Open full explorer →

Harness effect

The wrapper can change the result.

In the founding cohort, the same model can move by more than thirty points when only the answer format and parser change.

See how harnesses are pinned →
Qwen2.5 1.5B InstructGGUF Q4_K_M
33.3 ptsobserved spread
SmolLM2 1.7B InstructGGUF Q4_K_M
49.0 ptsobserved spread

Descriptive, not causal · two pinned harnesses · 24 tasks

The fairness contract

A result is more than a score.

Model, harness, suite, and runtime are separate versioned inputs. Change one, and FairBench treats it as a new result.

  1. 01

    Same constraints

    Identical tasks, token limits, sampling settings, retry policy, and scoring rules.

  2. 02

    Pinned harnesses

    Every prompt, parser, and adapter is versioned and hashed. The wrapper is part of the result.

  3. 03

    Receipts, not screenshots

    Scores ship with item outputs, parser outcomes, latency, runtime fields, and a reproducible command template.

  4. 04

    Uncertainty stays visible

    Coverage, category breakdowns, failures, and confidence intervals live next to the headline score.

Research track · Anti-benchmaxxing v0.1

The benchmark is public

Keep the method open. Rotate the answers.

The proposal combines open methodology, rotating seasonal suites, timestamped manifest commitments, and contamination canaries.

Post-release optimization cannot disappear. FairBench can label it, preserve historical runs, and make provenance difficult to hand-wave away.

Reproduce the evaluation

Your machine. The same protocol.

The dependency-free runner works with any OpenAI-compatible local endpoint and hashes the complete result bundle.

Terminal · local evaluation
$ npm run fairbench -- run \ --model qwen2.5:1.5b \ --harness minimal-answer/v1 \ --suite core-v1 \ --out result.json
✓ 24/24 tasks recordedbundle SHA-256 written

Intelligence deserves a fair measurement.