vraj

harness bench

harness bench

Same model, different harness. Three coding harnesses run the same models on the same real tasks. Every run is graded by hidden tests the agent never sees. Each dot below is one of those tests.

Loading results…

test passed test failed one row per harness and model, tasks left to right

leaderboard

who finishes the job

Resolved means every hidden correctness test passed. Rank goes by resolved tasks, then tests passed, then cost, then tokens. Ticks on each bar are single tasks. Click a column header to re-sort.

by task

every test, every run

Hover or focus a cell for time, cost and tokens.

profiles

the shape of each pairing

method

how it's graded

Hidden tests are the score
Each task ships a workspace, hidden tests and a reference solution. The agent sees only the workspace and the prompt. Correctness tests decide the score; quality and docs tests feed the side columns.
Sandboxed runs
Harnesses run under macOS sandbox-exec with the tests, references and repo out of reach. A run that touches them is flagged contaminated. None were.
Cost from real usage
Cost is recorded token usage times a fixed OpenRouter price snapshot, not the harness's own estimate. Free routes cost $0.
Same everything else
Same prompt, same task, same model route for every harness. Only the harness changes, so the gaps between rows are harness gaps.