I measure the instruments that measure AI agents. A benchmark is an instrument; it has a noise floor, it drifts, and current reporting practice has no field that would show either. I build the small pieces that make those visible: a probe that runs every day, a format for recording what produced a number, and the analysis that reads a leaderboard against its own noise.

Endpoint sentinel

The same fixed probe, every day, against hosted inference endpoints. Determinism, output length, latency, and step changes, per provider.

live

Log

One line per change: what was measured, added, or corrected, and when.

running