11benchreview-backed cataloglive
Same task. Different agent. Comparable evidence.
A public, review-backed catalog of how coding agents solve the same product and design tasks. Explore reviewed quality, judging, accounting, token usage, audit state, and run provenance without flattening missing data.
$ cd v0/<benchmark>
$ npm install && npm run build
$ # facts come from reviewed cycle data
Benchmark suites
4
Registered runs
111
Eligible / judged
16
Reviewed tokens
784.6M
Structured sources
4 reviewed cycles
Evidence coverage
4 of 4 priced
Comparison surface
16 model labels
Loading catalog…
Method
A catalog over immutable reviewed cycles.
Scores, prices, ranks, and reconciliation stay owned by the benchmark artifacts. This site only reshapes reviewed fields for navigation, tables, and comparison views.