KCAE.EVAL:4 - Solution
KCAE.EVAL:4.1 - State the decision and serious alternatives
Recover the receiving use, population, corpus editions, permitted data flows and adoption consequence from KCAE.USE. Choose an incumbent that a capable practitioner would actually use. It may be manual reading, exact/lexical search, a current hybrid retriever, an established long-context workflow or a competent existing agent. A weak toy baseline makes an addition easy to justify and hard to trust.
State the proposed difference and the expected useful change. Compare complete feasible arrangements, including required extraction, source reading, assessment and delivery. Keep common components, instructions and sources equal when the question is one component’s contribution; disclose changes when equality is impossible. A provider comparison confounded by different prompts or missing tools does not isolate the provider’s effect.
Choose whether the decision requires better useful results at a fixed resource boundary, lower cost at a required quality, or a transparent tradeoff. Keep non-negotiable permission and safety constraints separate from preferences; cheap unauthorized processing is not a feasible alternative.
KCAE.EVAL:4.2 - Build episodes that can defeat the design
Use actual receiving episodes where access is permitted, or independently construct realistic cases from work conditions. Retain the original request, relevant source basis, available case facts and expected kind of useful result. Separate cases used to design queries, tune thresholds or write features from cases used to assess them. Prevent near-duplicate leakage across that boundary.
Include cases with unknown source names, cross-language wording, useful non-pattern prose, distant prerequisites, misleading vocabulary, an important exception, missing facts and no useful corpus contribution. Include naturally answerless cases from the receiving work, not only deliberately irrelevant collections. Add source changes, stale indexes, revoked access and partial availability when the intended installation must handle them. Select proportions from the expected workload or report separate strata; a hand-balanced set is not automatically representative.
Gold material may be incomplete. Pool candidates from several serious routes and obtain source-based judgements from appropriate readers; sample beyond the pool where practical. Record uncertainty about unjudged material. Do not label every unpooled passage irrelevant. When several answers can serve the work, judge sufficiency and conditions rather than exact string agreement with one preferred answer.
Synthetic questions generated from an index can be useful development probes. Keep that origin visible and do not use them as the only evidence that the same index finds unfamiliar distinctions. For an English-primary assessor serving another language, test that language and important negation/condition forms rather than assuming a good English average transfers.
For a broad synthesis, vary the underlying evidence independently of its presentation. Duplicating an old edition or adding an overlapping summary must not create another population member or inflate a count. A new current counterexample or changed qualifying note must be able to alter the claim it defeats. Inspect whether a rare consequential position survives the map, reduction and delivery, and whether unknown membership remains unknown. Compare the final claims with original-source coding on a manageable dossier; merely counting processed summaries cannot test this contribution. KCAE.SEARCH:5.2 provides a constructed example, not an empirical success rate.
KCAE.EVAL:4.3 - Observe the complete useful result and full cost
Have the intended recipient or a qualified judge assess what the arrangement actually enabled: a correct explanation, an applicable conditional proposal, a necessary missing-fact question, a justified no-use result or a completed application. Also inspect unsupported advice, omitted exceptions, false absence claims and unusable deliveries. Source relevance, recommendation worth and actual work benefit remain distinct outcomes.
Measure cost over the selected horizon: preparation, extraction repair, indexing, storage, refresh, all model/tool calls, latency distribution, main-context consumption, human reading, interruptions, follow-up questions, application and recovery. Report total model input separately from residual main-window capacity. Report medians together with tails or failures where those change adoption. Use actual observations for empirical claims; estimates remain estimates with their assumptions.
For paired episodes, compare alternatives on the same case where this does not contaminate the reader. Randomize order or use separate recipients when the first exposure teaches the answer. Keep the judge’s criteria independent of the preferred implementation. If a model judges outputs, test its agreement and failure cases against source-grounded human or otherwise competent judgement; do not assume self-evaluation is neutral.
Return distributions or uncertainty appropriate to the data and decision. A handful of designed examples can establish that a mechanism runs and expose failures; it does not estimate a population success rate reliably. A small useful trial can justify a reversible next trial without pretending to prove general superiority.
KCAE.EVAL:4.4 - Isolate the failed contribution before repairing
Use controlled substitutions after the whole comparison identifies a consequential question:
- Supply adequate source candidates manually to the same assessor and delivery path. If the result recovers, finding was a limiting factor; if not, downstream work remains.
- Hold the candidate pool and source context fixed while comparing assessors. This separates candidate availability from ranking and condition judgement.
- Hold a qualified contribution fixed while varying delivery. A failed recipient can then expose omitted context, excessive compression, unavailable tools or missing preparation.
- Hold the source snapshot and queries fixed while varying representation or route. Inspect not only aggregate retrieval scores but unique useful discoveries and unique harmful omissions.
- Replay a controlled source change through build, persistence, reload and use. Compare the affected result with unchanged controls.
- Hold inquiry content fixed while changing the encounter occasion. Observe interruption and uptake, not merely whether a notification was sent.
These substitutions diagnose an operation under controlled inputs. They do not prove that the repaired whole will produce those inputs naturally. Re-run the relevant complete use after repair. Components can interact: a larger pool can improve recall and overwhelm an assessor, while a more selective delivery can save context and remove a decisive exception. Do not add separate component gains as if they were independent benefits.
For encounter recovery, interrupt the actual implementation after claiming work, during execution and after committing the result. Present a concurrent duplicate and a later completed duplicate. Observe whether unfinished work has an actual continuation, whether a stale owner can publish, and whether an external effect can be repeated or remains unknown. A constructed state trace can expose a missing branch; it does not establish a real database’s atomicity, scheduler behavior or recipient idempotency. Keep those implementation questions separate from a classifier’s inquiry quality.
KCAE.EVAL:4.5 - Qualify scores for the decision they control
For retrieval, recall over known relevant units, rank-sensitive measures and coverage can help diagnose finding. Their meaning depends on the judgement set. For screening, inspect false positives, false negatives, abstentions and unknown inputs at the proposed threshold. For claimed probabilities, test calibration on the relevant population and scoring event. For typed group choice, retain its within-group meaning; test logical relations among questions separately when the application depends on them.
Choose thresholds on development data using the costs of missed and incorrect actions. Freeze them for the held-out comparison, or account explicitly for adaptive tuning. Include near-boundary cases and no-use outcomes. A threshold that filters deliberately wrong repositories may fail on naturally difficult questions in the right repository. A calibrated score cannot create missing source evidence or domain permission.
Test the identity preserved across ranking, screening and return. Include a case where the ranking winner fails the adequacy threshold while another shortlisted candidate passes. The output must qualify its own candidate, choose a separately qualified alternative, or abstain; the maximum over the pool is not the returned candidate’s adequacy. Also test a high-scoring candidate with a contradicted necessary source condition. These cases distinguish a working score interface from a correctly connected decision.
Use executable calculation for counts, ratios, confidence intervals or cost arithmetic where needed. Preserve denominators and the treatment of unknowns. Report “not assessed” separately from failure and success. If a metric cannot distinguish the practical failure that motivated the change, add an appropriate receiving-result observation rather than decorating the same metric.
KCAE.EVAL:4.6 - Make the adoption and refresh decision
Relate observed gains and losses to the original use. Keep the incumbent when the addition does not earn its build and maintenance burden. Adopt a narrower profile when it helps one stratum but not another. Change the source preparation, route, assessor, delivery, observer or memory indicated by the evidence. A source gap may require obtaining a source rather than improving retrieval.
Name what the evidence supports, what it leaves untested and which condition reopens the decision. A changed corpus, model, language population, authority rule, workload or privacy constraint can defeat the original comparison. Preserve an operational fallback and observation during a bounded rollout if that is the selected use. A favorable laboratory result alone is not deployment or human benefit.