KCAE.EVAL - Compare Access Arrangements by Useful Results and Full Cost
Type: Method pattern Status: Stable
KCAE.EVAL:1 - Problem frame
Use this when adopting, replacing or repairing an access arrangement requires evidence of its practical gain, or a failure needs attribution to the operation that can be repaired. A benchmark score, successful API call or convincing demonstration cannot answer every such question.
The gain is a bounded decision to adopt, retain, change or stop an arrangement on the evidence that can support it. Do not run a large study for a harmless reversible lookup whose adequacy is already inspectable. Scale the comparison to the consequence and uncertainty of the receiving choice.
KCAE.EVAL:2 - Problem
Retrieval can improve while the final answer remains wrong. A better assessor can appear worse because it receives poorer candidates. An inexpensive model can raise human review or refresh costs. A synthetic test set derived from the same summaries as the index can hide the exact distinctions the arrangement was meant to recover.
KCAE.EVAL:3 - Forces
Real episodes have greater practical relevance and less convenient labels. Whole-use comparisons establish value but can hide the cause. Controlled component comparisons explain a difference but narrow its reach. Broad coverage, independent judging and repeated trials consume the same resources the system is meant to save.
KCAE.EVAL:4 - Solution
KCAE.EVAL:4.1 - State the decision and serious alternatives
Recover the receiving use, population, corpus editions, permitted data flows and adoption consequence from KCAE.USE. Choose an incumbent that a capable practitioner would actually use. It may be manual reading, exact/lexical search, a current hybrid retriever, an established long-context workflow or a competent existing agent. A weak toy baseline makes an addition easy to justify and hard to trust.
State the proposed difference and the expected useful change. Compare complete feasible arrangements, including required extraction, source reading, assessment and delivery. Keep common components, instructions and sources equal when the question is one component’s contribution; disclose changes when equality is impossible. A provider comparison confounded by different prompts or missing tools does not isolate the provider’s effect.
Choose whether the decision requires better useful results at a fixed resource boundary, lower cost at a required quality, or a transparent tradeoff. Keep non-negotiable permission and safety constraints separate from preferences; cheap unauthorized processing is not a feasible alternative.
KCAE.EVAL:4.2 - Build episodes that can defeat the design
Use actual receiving episodes where access is permitted, or independently construct realistic cases from work conditions. Retain the original request, relevant source basis, available case facts and expected kind of useful result. Separate cases used to design queries, tune thresholds or write features from cases used to assess them. Prevent near-duplicate leakage across that boundary.
Include cases with unknown source names, cross-language wording, useful non-pattern prose, distant prerequisites, misleading vocabulary, an important exception, missing facts and no useful corpus contribution. Include naturally answerless cases from the receiving work, not only deliberately irrelevant collections. Add source changes, stale indexes, revoked access and partial availability when the intended installation must handle them. Select proportions from the expected workload or report separate strata; a hand-balanced set is not automatically representative.
Gold material may be incomplete. Pool candidates from several serious routes and obtain source-based judgements from appropriate readers; sample beyond the pool where practical. Record uncertainty about unjudged material. Do not label every unpooled passage irrelevant. When several answers can serve the work, judge sufficiency and conditions rather than exact string agreement with one preferred answer.
Synthetic questions generated from an index can be useful development probes. Keep that origin visible and do not use them as the only evidence that the same index finds unfamiliar distinctions. For an English-primary assessor serving another language, test that language and important negation/condition forms rather than assuming a good English average transfers.
For a broad synthesis, vary the underlying evidence independently of its presentation. Duplicating an old edition or adding an overlapping summary must not create another population member or inflate a count. A new current counterexample or changed qualifying note must be able to alter the claim it defeats. Inspect whether a rare consequential position survives the map, reduction and delivery, and whether unknown membership remains unknown. Compare the final claims with original-source coding on a manageable dossier; merely counting processed summaries cannot test this contribution. KCAE.SEARCH:5.2 provides a constructed example, not an empirical success rate.
KCAE.EVAL:4.3 - Observe the complete useful result and full cost
Have the intended recipient or a qualified judge assess what the arrangement actually enabled: a correct explanation, an applicable conditional proposal, a necessary missing-fact question, a justified no-use result or a completed application. Also inspect unsupported advice, omitted exceptions, false absence claims and unusable deliveries. Source relevance, recommendation worth and actual work benefit remain distinct outcomes.
Measure cost over the selected horizon: preparation, extraction repair, indexing, storage, refresh, all model/tool calls, latency distribution, main-context consumption, human reading, interruptions, follow-up questions, application and recovery. Report total model input separately from residual main-window capacity. Report medians together with tails or failures where those change adoption. Use actual observations for empirical claims; estimates remain estimates with their assumptions.
For paired episodes, compare alternatives on the same case where this does not contaminate the reader. Randomize order or use separate recipients when the first exposure teaches the answer. Keep the judge’s criteria independent of the preferred implementation. If a model judges outputs, test its agreement and failure cases against source-grounded human or otherwise competent judgement; do not assume self-evaluation is neutral.
Return distributions or uncertainty appropriate to the data and decision. A handful of designed examples can establish that a mechanism runs and expose failures; it does not estimate a population success rate reliably. A small useful trial can justify a reversible next trial without pretending to prove general superiority.
KCAE.EVAL:4.4 - Isolate the failed contribution before repairing
Use controlled substitutions after the whole comparison identifies a consequential question:
- Supply adequate source candidates manually to the same assessor and delivery path. If the result recovers, finding was a limiting factor; if not, downstream work remains.
- Hold the candidate pool and source context fixed while comparing assessors. This separates candidate availability from ranking and condition judgement.
- Hold a qualified contribution fixed while varying delivery. A failed recipient can then expose omitted context, excessive compression, unavailable tools or missing preparation.
- Hold the source snapshot and queries fixed while varying representation or route. Inspect not only aggregate retrieval scores but unique useful discoveries and unique harmful omissions.
- Replay a controlled source change through build, persistence, reload and use. Compare the affected result with unchanged controls.
- Hold inquiry content fixed while changing the encounter occasion. Observe interruption and uptake, not merely whether a notification was sent.
These substitutions diagnose an operation under controlled inputs. They do not prove that the repaired whole will produce those inputs naturally. Re-run the relevant complete use after repair. Components can interact: a larger pool can improve recall and overwhelm an assessor, while a more selective delivery can save context and remove a decisive exception. Do not add separate component gains as if they were independent benefits.
For encounter recovery, interrupt the actual implementation after claiming work, during execution and after committing the result. Present a concurrent duplicate and a later completed duplicate. Observe whether unfinished work has an actual continuation, whether a stale owner can publish, and whether an external effect can be repeated or remains unknown. A constructed state trace can expose a missing branch; it does not establish a real database’s atomicity, scheduler behavior or recipient idempotency. Keep those implementation questions separate from a classifier’s inquiry quality.
KCAE.EVAL:4.5 - Qualify scores for the decision they control
For retrieval, recall over known relevant units, rank-sensitive measures and coverage can help diagnose finding. Their meaning depends on the judgement set. For screening, inspect false positives, false negatives, abstentions and unknown inputs at the proposed threshold. For claimed probabilities, test calibration on the relevant population and scoring event. For typed group choice, retain its within-group meaning; test logical relations among questions separately when the application depends on them.
Choose thresholds on development data using the costs of missed and incorrect actions. Freeze them for the held-out comparison, or account explicitly for adaptive tuning. Include near-boundary cases and no-use outcomes. A threshold that filters deliberately wrong repositories may fail on naturally difficult questions in the right repository. A calibrated score cannot create missing source evidence or domain permission.
Test the identity preserved across ranking, screening and return. Include a case where the ranking winner fails the adequacy threshold while another shortlisted candidate passes. The output must qualify its own candidate, choose a separately qualified alternative, or abstain; the maximum over the pool is not the returned candidate’s adequacy. Also test a high-scoring candidate with a contradicted necessary source condition. These cases distinguish a working score interface from a correctly connected decision.
Use executable calculation for counts, ratios, confidence intervals or cost arithmetic where needed. Preserve denominators and the treatment of unknowns. Report “not assessed” separately from failure and success. If a metric cannot distinguish the practical failure that motivated the change, add an appropriate receiving-result observation rather than decorating the same metric.
KCAE.EVAL:4.6 - Make the adoption and refresh decision
Relate observed gains and losses to the original use. Keep the incumbent when the addition does not earn its build and maintenance burden. Adopt a narrower profile when it helps one stratum but not another. Change the source preparation, route, assessor, delivery, observer or memory indicated by the evidence. A source gap may require obtaining a source rather than improving retrieval.
Name what the evidence supports, what it leaves untested and which condition reopens the decision. A changed corpus, model, language population, authority rule, workload or privacy constraint can defeat the original comparison. Preserve an operational fallback and observation during a bounded rollout if that is the selected use. A favorable laboratory result alone is not deployment or human benefit.
KCAE.EVAL:5 - Archetypal Grounding
Suppose CedarBench compares an existing hybrid search workflow with a proposed direct-block supplement on 40 withheld support episodes. These are invented study conditions, not reported measurements. Cases include protocol ambiguity, old translations, useful table notes and naturally unsupported requests. Both alternatives use the same source snapshot, assessor, delivery and recipient rules. The proposed route adds its actual call and maintenance costs.
If a missed exception appears after manually supplying its source, the original finding route was a limitation. If the assessor still recommends resend on protocol v1, retrieval repair alone cannot close the case. If the judgement is correct but the customer receives only the upbeat first sentence, delivery remains defective. The adoption question returns only after the relevant whole is tried again.
KCAE.EVAL:6 - Bias-Annotation
Authors tend to select examples their design can answer and count improvement where it is easiest to measure. Start from the receiving work and include natural failures and no-use cases. Avoid tuning on the final test, treating model agreement as independent truth, or omitting preparation and interruption costs. A smaller honest conclusion is more reusable than an inflated score.
KCAE.EVAL:7 - Conformance Checklist
Does the comparison answer a real adoption or repair decision? Is the baseline serious? Are sources, episodes, recipients and costs comparable? Could the cases expose an unknown or absent answer? Are development and evaluation separated? Can the result distinguish finding, assessment, delivery and application? Have component repairs returned to whole use? Are evidence reach and refresh conditions explicit?
KCAE.EVAL:8 - Common Anti-Patterns and How to Avoid Them
A green API test establishes interface behavior, not useful advice. A stronger retriever with a weaker prompt is not a clean model comparison. An oracle candidate trial is a diagnosis, not end-to-end performance. A gold-free case is not necessarily a true no-answer case; inspect its source basis. A cheap query does not make an expensive maintained service cheap.
KCAE.EVAL:9 - Consequences
Engineering choices become revisable on evidence relevant to the receiving work. Some promising additions will be rejected or narrowed. Evaluation itself costs work, so reuse current matching results and inspect only changes that can alter the decision. Neither exhaustive metrics nor a universal test-set size is required.
KCAE.EVAL:10 - Architectural Rationale
Whole-use comparison determines whether the arrangement is worth having. Controlled component substitutions determine where to intervene. Separating those questions avoids both an unexplained aggregate score and a pile of excellent components that fail together.
KCAE.EVAL:11 - SoTA-Echoing
AgentRetrievalBench contributes a concrete warning about natural no-gold retrieval and limited evidence from artificial controls. Typed-decision audits expose confounding and coherence questions beyond a neat output schema. KCAE.Reference:1 preserves their scope; neither establishes support-library efficacy. Classical information-retrieval evaluation remains useful for component diagnosis. This pattern adds receiving use, source change, reader sufficiency, encounter and full lifecycle cost to the adoption question. Reopen metrics when they cease to distinguish the failure that matters.
KCAE.EVAL:12 - Relations
KCAE.USE supplies the decision and workload. Every other KCAE method supplies an operation that can be examined without making its local success the whole result. C.11.DUA supplies advice and evidence burden; ME.13 supplies qualified method-transfer questions when transfer is claimed. A publication’s own quality and admission remain governed by its authoring and evaluation methods, separate from an installation’s runtime comparison.