Library / Knowledge-Corpus Access Engineering Principles Framework
Jump to passage
In this reading

Link to current text

Published source confirmed at last check

Source changed 2026-10-03 17:24:51 UTC · snapshot created 2026-10-03 17:30:20 UTC · last check 2026-10-03 18:00:10 UTC

KCAE.ASSESS:4.4 - Qualify scores for their intended decision

Ranking, calibration, classification at a threshold and agreement among logically related questions are separate properties. A probability distribution over a supplied group is conditional on that group. It can confidently prefer an inadequate candidate when none of the choices is suitable. Keep an explicit adequacy question and permit unresolved/no-useful-candidate results.

When candidates are split into groups, do not compare their locally normalized Choice probabilities as one global distribution. Screen with stable pointwise criteria or jointly reconsider finalists under the same question. For a multi-contribution result, assess the actual composition; multiplying unrelated usefulness probabilities assumes dependence facts that the model has not supplied.

Bind an adequacy result to the candidate actually returned. Suppose a ranking selects A, adequacy scores are A = 0.2 and B = 0.8, and an illustrative screening threshold is 0.3. The maximum score establishes only that some candidate passed; it cannot qualify A. Test the selected candidate’s score and source conditions, choose an independently qualified alternative, or return the shortlist for another judgement. Preserve candidate identity, criterion and source context across that selection boundary.

Choose thresholds on development cases representative of the action and evaluate them on separate cases. Retain model, prompt/question, option order/wording, preprocessing, language, source scope and decision cost with the setting. Test near misses, unknown facts, late conditions, all-inadequate pools, longer inputs and meaningful negations. Recheck the affected qualification after a model or criterion change. The same threshold need not govern cheap further reading and a costly user interruption.

Where a score is useful only for ordering inspection, keep that limited claim and avoid pretending to have calibrated adequacy. A human can read the highest-ranked candidates and make the stronger judgement from source evidence. This often provides a viable first arrangement before fully automated recommendation is justified.