KCAE.INDEX:4.4 - Combine candidates without inventing a common probability
Run the selected routes against the same intended corpus/edition conditions. Deduplicate by source identity and overlap, while retaining which routes found a candidate. Keep competing editions distinct where the use needs them. Merge adjacent fragments into one reading candidate when their overlap would otherwise consume the budget repeatedly; do not merge conflicting source claims merely because their words are similar.
Raw lexical, vector and graph scores usually have different meanings. A simple candidate-combination baseline is rank fusion. Reciprocal rank fusion assigns a candidate the sum of 1/(c + rank) across lists in which it appears. It uses order rather than pretending that raw scores share a scale. The positive constant c controls how much top ranks dominate; the original 2009 study used 60, which is historical experimental selection rather than a universal setting. RRF source.
Fusion still favours material with several appearances. Preserve a tested allocation for route-unique candidates before the later shortlist: for example, take a fused core, then admit the highest-ranked unrepresented candidate from each materially different route, within the same budget. Treat query paraphrases from one model as related search attempts, not independent votes proving relevance. Compare this diversity policy with simple fusion on held-out questions; keep it only where it recovers valuable contributions at acceptable cost.
For a concrete constructed pool, suppose lexical order is A, B, C and dense order is D, A, B. With c = 60, A receives 1/61 + 1/62, B receives 1/62 + 1/63, and D receives 1/61. A and B win the top two positions because each appears in both lists. The route-unique D disappears from a two-item shortlist. Retaining D for inspection may expose a cross-language exception. Its eventual usefulness must be judged from the source; its uniqueness is a reason to inspect, not evidence of truth.
Choose candidate-window size from downstream reading capacity and error costs. A pool of 100 hits is an internal search result, not a request to stuff 100 fragments into the principal agent’s context. Assess and expand within the search service or separate reader where supported; deliver the sufficient result through KCAE.DELIVER. Too narrow a pool cannot be repaired by a perfect reranker, because the relevant source never reaches it.