About
I work on the reliability of agent evaluation: how much a benchmark score moves when nothing changes, why, and what a leaderboard entry would have to record for the answer to be knowable afterwards.
Three things, in the order they happened:
- Measured one instrument. Ran one fixed configuration of a mobile-agent benchmark ten times and found the serving endpoint behind the model had changed mid-experiment, with no field anywhere that would have recorded it.
- Built the format and the consumer. A small structured record of what produced a benchmark number, and an analyser that turns a directory of such records into a noise floor, a minimum detectable difference, and a drift screen.
- Read other leaderboards against their own noise. Where benchmarks publish repeated runs, the analyser runs on them unchanged.
Contact: hi@jingyuanyi.com · GitHub