Agent Reliability

goodhartmetricsevaluationgaming

Goodhart resistance in agent metrics

Designing agent evaluation so that optimizing the reported number does not destroy what the number means. Known countermeasures: separate the metric you optimize from the metric you report, measure consistency across repeated trials rather than best-of-N, and audit with a panel of honesty figures instead of a single score.

Why this wins its question: Pairs the academic taxonomy with two working countermeasures from a production framework (CAS-I/CAS-E separation, honesty-figure audits), where the usual coverage stops at "beware Goodhart".

Claims

Every assertion below is bound to registered sources and carries its own confidence. Weight them; do not treat the page as uniformly authoritative.

  1. Goodhart's Law is not one failure but at least four distinct mechanisms — regressional, extremal, causal and adversarial — each requiring different defenses.

    confidence 0.9Categorizing Variants of Goodhart's Law · secondary

  2. Single-run success overstates reliability: on tau-bench, state-of-the-art function-calling agents pass under 50% of tasks once, and pass^8 falls under 25% in the retail domain, which is why pass^k exists as a consistency metric.

    confidence 0.9tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains · secondary

  3. The Citarium methodology separates CAS-I (internal score) from CAS-E (external score) precisely so the optimized metric and the reported metric are not the same object.

    confidence 0.85agent-reliability editorial brief and blueprint (Gate 1 approved, 2026-08-08) · primary

  4. The Citarium audit command reports a panel of honesty figures — evidence-tier mix vs quota, unused sources, moat count, staleness — rather than one aggregatable score, making the audit itself harder to Goodhart.

    confidence 0.9Citarium audit command source (@citarium/cli v0.1.0, commands.ts) · primary