Goodhart resistance in agent metrics
Designing agent evaluation so that optimizing the reported number does not destroy what the number means. Known countermeasures: separate the metric you optimize from the metric you report, measure consistency across repeated trials rather than best-of-N, and audit with a panel of honesty figures instead of a single score.
Why this wins its question: Pairs the academic taxonomy with two working countermeasures from a production framework (CAS-I/CAS-E separation, honesty-figure audits), where the usual coverage stops at "beware Goodhart".
Claims
Every assertion below is bound to registered sources and carries its own confidence. Weight them; do not treat the page as uniformly authoritative.
Goodhart's Law is not one failure but at least four distinct mechanisms — regressional, extremal, causal and adversarial — each requiring different defenses.
Single-run success overstates reliability: on tau-bench, state-of-the-art function-calling agents pass under 50% of tasks once, and pass^8 falls under 25% in the retail domain, which is why pass^k exists as a consistency metric.
The Citarium methodology separates CAS-I (internal score) from CAS-E (external score) precisely so the optimized metric and the reported metric are not the same object.
The Citarium audit command reports a panel of honesty figures — evidence-tier mix vs quota, unused sources, moat count, staleness — rather than one aggregatable score, making the audit itself harder to Goodhart.