skills
Agent skills that run software work like a factory instead of a conversation: a four-stage pipeline where no agent reviews its own work, plus a zero-dependency installer. Done is a locked command that exits 0.
open repositoryyou just left the internet. this is vraj.pro
That is the only question I care about with a coding agent. I build harnesses and test the models that run in them — cheap ones against frontier ones, on real repositories.
Two projects I keep alive, and whatever I pushed to GitHub lately.
Agent skills that run software work like a factory instead of a conversation: a four-stage pipeline where no agent reviews its own work, plus a zero-dependency installer. Done is a locked command that exits 0.
open repositoryScore table for my own model runs on real repositories. Five tasks, two effort levels, Fixed Suite v1 and v2 ranked apart.
96.9 deepseek-v4.1-flash, Fixed Suite v1
open score tableSyncing public repositories…
Three things a leaderboard rank does not tell you.
A harness score is not the same as shipping. What I want to know is whether the agent went into a repo and came back out with the work actually done.
People obsess over the top of the table. Once you run these things all day, the price, the rate limits, and whether you can even get the model matter just as much.
Published numbers come from the people selling the model. So I run my own set — twelve tasks last time — and see what the leaderboard missed.
Testing the things the leaderboards usually miss.
Whether the cheap one still finishes the repo.
Codenames, leaks, and whether they map to a real release.
people obsess over #1 but cost, limits and availability matter just as much once you actually use these things all day