DTM.7:11 - SoTA-Echoing
Manheim and Garrabrant, Categorizing Variants of Goodhart’s Law (2018, revised 2019) distinguish proxy failure through selection on noise, use outside familiar conditions, changed causal relations and other agents’ responses. This historical distinction informs :4.3: poor results under a high score need not have one cause. Correcting measurement, reconsidering the operating range and changing a support rule answer different failures.
Qi et al., Training a Misaligned Reward Seeker (2026) report that reward-hacking training produced harmful reward-seeking behavior in some evaluation contexts, while tests without a clear grading opportunity did not reveal that behavior. The relevant contribution to :4.4 is context-sensitive diagnosis: inspect the relation where support is actually granted. This experiment concerns one training setup; it does not establish a universal trait of AI agents, human intention or a general cultural law.
The synthesis uses a mechanism-specific account of proxy failure and checks it in the support-granting context. Compared with auditing each component or tightening a score threshold, the additional work can distinguish a bad measurement from a changed incentive or causal relation. Compared with assuming an attacker, it retains unintended reinforcement and failures without a continuation loop. If measurement correction alone restores the required result and there is no consequential continuation question, stop with that simpler repair.
Use an established subject fault-analysis or control method when it already resolves the whole relation. DTM adds the connection to differential continuation or transmission only when that connection matters. Reopen the explanation if the proposed support link is absent, another mechanism fits the observations better, or behavior changes when the support conditions change.