AI Evaluation
What a benchmark score establishes, what it only appears to, and how to tell the difference.
Every claim about an AI system's capability eventually rests on a measurement, and measurements carry assumptions that can quietly stop holding. An agent benchmark assumes the only tractable route to a high score is doing the work. An LLM judge assumes the judge understands the task well enough to grade it. Both assumptions have been tested recently, and both turn out to be conditional rather than permanent. These essays look at evaluation as a measurement problem: what the instruments assume, when those assumptions fail, and what to trust in the meantime.