Library / Knowledge-Corpus Access Engineering Principles Framework
Jump to passage
In this reading

Link to current text

Published source confirmed at last check

Source changed 2026-10-03 14:36:52 UTC · snapshot created 2026-10-03 14:38:14 UTC · last check 2026-10-03 15:35:10 UTC

KCAE.EVAL:4.5 - Qualify scores for the decision they control

For retrieval, recall over known relevant units, rank-sensitive measures and coverage can help diagnose finding. Their meaning depends on the judgement set. For screening, inspect false positives, false negatives, abstentions and unknown inputs at the proposed threshold. For claimed probabilities, test calibration on the relevant population and scoring event. For typed group choice, retain its within-group meaning; test logical relations among questions separately when the application depends on them.

Choose thresholds on development data using the costs of missed and incorrect actions. Freeze them for the held-out comparison, or account explicitly for adaptive tuning. Include near-boundary cases and no-use outcomes. A threshold that filters deliberately wrong repositories may fail on naturally difficult questions in the right repository. A calibrated score cannot create missing source evidence or domain permission.

Test the identity preserved across ranking, screening and return. Include a case where the ranking winner fails the adequacy threshold while another shortlisted candidate passes. The output must qualify its own candidate, choose a separately qualified alternative, or abstain; the maximum over the pool is not the returned candidate’s adequacy. Also test a high-scoring candidate with a contradicted necessary source condition. These cases distinguish a working score interface from a correctly connected decision.

Use executable calculation for counts, ratios, confidence intervals or cost arithmetic where needed. Preserve denominators and the treatment of unknowns. Report “not assessed” separately from failure and success. If a metric cannot distinguish the practical failure that motivated the change, add an appropriate receiving-result observation rather than decorating the same metric.