vraj

you just left the internet. this is vraj.pro

can it actually enter a repo and finish the job?

That is the only question I care about with a coding agent. I build harnesses and test the models that run in them — cheap ones against frontier ones, on real repositories.

costlimitsavailabilitycan it finish the job?

things with a pulse

Two projects I keep alive, and whatever I pushed to GitHub lately.

featuredevolving

skills

Agent skills that run software work like a factory instead of a conversation: a four-stage pipeline where no agent reviews its own work, plus a zero-dependency installer. Done is a locked command that exits 0.

role
design & build
surface
JavaScript, no dependencies
shape
four-stage review pipeline
open repository

score tablefirst run baked

72 hours in

Score table for my own model runs on real repositories. Five tasks, two effort levels, Fixed Suite v1 and v2 ranked apart.

96.9 deepseek-v4.1-flash, Fixed Suite v1

open score table

recent on GitHub …

all repositories

Syncing public repositories…

vendor numbers are still vendor numbers

Three things a leaderboard rank does not tell you.

right now

  1. building my own eval set

    Testing the things the leaderboards usually miss.

  2. cheap models against frontier ones

    Whether the cheap one still finishes the repo.

  3. tracking what the labs ship next

    Codenames, leaks, and whether they map to a real release.

About

people obsess over #1 but cost, limits and availability matter just as much once you actually use these things all day

Vraj on X, 29 Aug 2026