Evaluation Report
Measuring how well frontier models
actually write code.
Every candidate solution is executed in an isolated sandbox against a hidden suite of edge-case tests, and graded differentially against a trusted oracle — not hand-written expected values.
Model Leaderboard
Ranked by hidden-test pass rate across the full suite.
Coverage Matrix
Per-task pass rate for each model. Click a cell to inspect it.
Failure Analysis
Systematic, category-level weaknesses with minimal reproductions — the signal a model team can act on.
Task Explorer
Inspect a benchmark task, its hidden cases, and the oracle vs. candidate code.