Vraj Trying out ai

score table

72 hours in, scored.

Same harness, real repos, 72 hours on the clock. Scores land here once the runs finish — cost, limits, and whether it finished the job.

fixed suite v1

The first set, ranked.

My own set, frozen. Every model gets the same tasks at each effort level; the ranking below only counts v1 runs.

No v1 scores yet — nothing baked. Check back once the runs finish.

fixed suite v2

The next set, on its own.

A new fixed set gets a new table. v2 scores never mix into the v1 ranking, and v1 never props up v2.

No v2 scores yet — nothing baked. The table fills in when real v2 runs land.