CHK.6:6 - Bias-Annotation
People tend to invent tests that confirm their implementation. Generated criteria and generated answers can share the same blind spot. Recover a failure from the receiving work, seek an unlike case where it matters, and distinguish the data used to tune a checker from the evidence used to judge its broader use.
A benchmark’s accessible cases can underrepresent the domain or participants who bear false failures. Make that limit part of the reliance decision.