VERITAS Frontier-Model Code Evaluation Harness
REPRODUCIBLE

Evaluation Report

Measuring how well frontier models
actually write code.

Every candidate solution is executed in an isolated sandbox against a hidden suite of edge-case tests, and graded differentially against a trusted oracle — not hand-written expected values.

Model Leaderboard

Ranked by hidden-test pass rate across the full suite.

Coverage Matrix

Per-task pass rate for each model. Click a cell to inspect it.

Failure Analysis

Systematic, category-level weaknesses with minimal reproductions — the signal a model team can act on.

Task Explorer

Inspect a benchmark task, its hidden cases, and the oracle vs. candidate code.