Vraj Trying out ai

Trying out ai local

can it actually enter a repo and finish the job?

That is the only question I care about with a coding agent. I build harnesses and test the models that run in them — cheap ones against frontier ones, on real repositories.

Focus
Coding agents, harnesses, and evals
Testing now
GLM, Qwen, DeepSeek, Opus, Fable
Status
Building my own eval set
Cost Limits Availability Can it finish the job?

01 — Live work

Projects with a pulse.

A few things I’m building, exploring, or keeping alive. The repository list pulls fresh data straight from GitHub.

Featured Evolving

Skills

Agent skills that run software work like a factory instead of a conversation: a four-stage pipeline where no agent reviews its own work, plus a zero-dependency installer. Done is a locked command that exits 0.

  • RoleDesign & build
  • SurfaceJavaScript, no dependencies
  • ShapeFour-stage review pipeline
Open repository

Score table 72 hours in

72 hours in

Score table for my own model runs on real repositories. Nothing baked yet — runs land here when baked.

  • WhatModel runs, scored
  • StatusNothing baked yet
  • ShapeFixed Suite v1 · v2, ranked apart
Open score table

From the source

Recent GitHub activity

All repositories

Syncing public repositories…

02 — How I judge a model

Vendor numbers are still vendor numbers.

Three things a leaderboard rank does not tell you.

03 — What’s happening

The now page.

A snapshot of the direction, not a promise to keep still.

  1. 01

    Building my own eval set

    Testing the things the leaderboards usually miss.

    In progress
  2. 02

    Cheap models against frontier ones

    Whether the cheap one still finishes the repo.

    Ongoing
  3. 03

    Tracking what the labs ship next

    Codenames, leaks, and whether they map to a real release.

    Daily

04 — A little context

The short version.

people obsess over #1 but cost, limits and availability matter just as much once you actually use these things all day

— Vraj on X, 29 Aug 2026