Library / Checklist Principles Framework
Jump to passage
In this reading

Link to current text

Published source confirmed at last check

Source changed 2026-10-03 07:42:37 UTC · snapshot created 2026-10-03 07:43:27 UTC · last check 2026-10-03 07:55:10 UTC

CHK.6:4.2 - Construct a discriminating challenge

Choose cases with known or competently established differences relevant to the criterion. Include a satisfactory case and a plausible failure. Vary the property of interest while holding irrelevant differences small enough to interpret the result. For a subjective criterion, recover the judgement basis and a meaningful disagreement rather than inventing a false objective threshold.

Predict which answers those differences support before inspecting the checker output. Then compare the outputs with adequate grounds for those expected distinctions. A false pass and a false failure can have different costs; inspect both where they affect the work.

Where a rubric may drive the verdict independently of the subject, also hold the subject facts fixed and change the question. Establish what answer the new predicate supports before asking the checker. A reversed condition can call for the opposite answer; if only wording changes and the question stays the same, the expected answer stays the same. This distinction tests responsiveness to meaning without requiring every wording change to reverse a result.

When uncertainty concerns the judging arrangement, compare it on those grounded cases with a cheaper or less interfering alternative. Grouping criteria can save calls; separating them can reveal interference. Repetition and a rule for combining verdicts may reduce sampling instability, but repeated agreement cannot correct a consistent error or supply independent truth. Select the arrangement from consequential errors and total effort; repeat only when that comparison can change reliance.

If the receiving use consumes a total or percentage, examine what that number means and how the questions affect it. Splitting one easy condition into ten items can increase its influence without improving the subject. Ten successful formatting checks may conceal one failed catalogue-preservation check. Challenge the implied priorities, duplicate contributions and treatment of missing answers. Keep consequential answers visible. Where a numerical aggregate is actually needed, FPF A.19.ULSAM governs the admitted measures and lawful aggregation; a convenient average does not supply its own justification.

Start cheaply. One counterexample can defeat an overbroad claim. Use a direct executable comparison when the question and observable property admit it. A selective rubric or pairwise judgement can help with an interpretive or relative question, but choosing the better of two results does not establish that either meets an absolute condition. Qualify that judging arrangement on cases whose consequential answers have adequate independent grounds. Compare false passes, false failures and obtaining those grounds as well as the cost per judgement. Broader reliance may call for representative cases, independent judgement, executable observation or a field trial through ME.11. Choose another reader or agent when it supplies a materially different basis, expertise or observation. Mere duplication of the same unsupported inference does not establish independence.

When work is repeatedly optimized against a score, compare earlier and later actual results on consequential characteristics and suitable new cases. A rising score can reward surface changes while the intended result deteriorates. For the importer, a more persuasive success report leaves identifier loss unchanged. Inspect the proposed shortcut and obtain evidence beyond the optimized signal. A stronger second model can help investigate disagreement, but its score still needs grounds and does not become ground truth by rank.