KCAE.EVAL:6 - Bias-Annotation
Authors tend to select examples their design can answer and count improvement where it is easiest to measure. Start from the receiving work and include natural failures and no-use cases. Avoid tuning on the final test, treating model agreement as independent truth, or omitting preparation and interruption costs. A smaller honest conclusion is more reusable than an inflated score.