Model × harness explorer

One matrix. Two fair questions.

The model is only one input. Keep the harness fixed to compare models—or keep the model fixed to measure the harness.

FairBench Core v1 · 24 tasks · deterministic gradingReference runs
Results
Suite FairBench Core v1 / Tasks 24 / Attempts 1

Showing FairBench Core v1 results for Minimal answer v1.

Swipe or scroll horizontally to inspect every column →

FairBench Core v1: models compared with Minimal answer v1, 24 tasks, one attempt per task.
RankModelScore95% CICoverageProvenance
1=SmolLM2 1.7B InstructGGUF Q4_K_M51.7%32.3%–71.2%24/24reference
1=Qwen2.5 1.5B InstructGGUF Q4_K_M45.8%25.0%–66.7%24/24reference