Library / Checklist Principles Framework
Jump to passage
In this reading

Link to current text

Published source confirmed at last check

Source changed 2026-10-03 11:52:20 UTC · snapshot created 2026-10-03 11:53:41 UTC · last check 2026-10-03 12:45:07 UTC

CHK.6:11 - SoTA-Echoing

Which checking means and qualification effort fit the intended reliance? Adopt a direct comparison for an explicit observable property; adapt selective rubric or pairwise judging for a genuinely interpretive or relative question, with discriminating cases when its adequacy is uncertain. These are serious alternatives, not successive compulsory stages. In §§4.2 and 5, inspecting identifier or catalogue equality supplies the relevant grounds directly. Asking a model for a persuasive overall score adds cost without settling those properties. Pairwise judging can instead help choose between two explanations, but the preferred one still needs a separate answer to an absolute suitability question.

Selective Checklist Evaluation supplies a current rival to full-rubric/direct scoring, with effects that differ by setting and comparison mode. It warrants choosing and qualifying the mode for the receiving question, not declaring one universally superior. The study by Hong and colleagues supplies counterevidence about detail and uneven emphasis in generated criteria. Section 4.2 therefore examines consequential errors, missing answers and induced weighting alongside judgement cost. The extra qualification is worthwhile when it can defeat misplaced reliance or select better means; an applicable prior qualification remains sufficient for its unchanged scope.

The study by Bagaria and colleagues supplies bounded counterevidence about judges responding to changed answers and reversed criteria. Its generated and qualitatively inspected contrasts leave uncertainty about some expected verdicts. The adaptation in §§4.2 and 5 fixes expected answers from a simple fixture, then varies subject facts or question meaning separately. RIPD supplies a further benchmark-versus-target counterexample motivating examination after rubric edits; it does not qualify a particular deployed judge.

For single, grouped or repeated model judgements, adapt protocol comparison to the consequential uncertainty in §§4.1–4.3 and the fixed-fixture case in §5. RuVerBench v2 supplies counterevidence about prompting, grouping and voting under long research/coding contexts; its selected binary cases exclude substantial ambiguity. PReMISE v1 separates stability from preference fit and robustness under model-mediated evaluation. Neither determines this deployment’s best arrangement. Compare saved calls with consequential errors and examination effort; retain adequate prior qualification, and reopen its affected scope when a changed arrangement can alter a relied-on answer.

AutoChecklist and RLCF supply criterion refinement and checklist-feedback training alternatives. Reject their benchmark or training gains as substitutes for qualifying this checker. CHERRL and Rubric Dropout expose proxy exploitation and divergence in bounded training settings; the latter’s single-run configurations and fallible external judge limit the result. Its training intervention gives no general reason to omit operational checks. Section 4.2 instead examines actual results and new cases beyond the optimized signal.

The deliberate trade-off is the cost of obtaining discriminating grounds against the cost of a consequential false answer. None of these studies supplies a common error/cost ranking for every domain. Anthropic’s harness-design account supplies a worked calibration candidate as models and tools change; ME.11 and ME.14 guide the local trial and worth comparison when the choice remains consequential. Reopen qualification for a material criterion, tool or use change, an unexplained verdict on a known contrast, or a cheaper adequate observation.