I measure the instruments that measure AI agents. A benchmark is an instrument; it has a noise floor, it drifts, and current reporting practice has no field that would show either. I build the small pieces that make those visible: a probe that runs every day, a format for recording what produced a number, and the analysis that reads a leaderboard against its own noise.
Endpoint sentinel
The same fixed probe, every day, against hosted inference endpoints. Determinism, output length, latency, and step changes, per provider.
liveLog
One line per change: what was measured, added, or corrected, and when.
running