KCAE.INDEX:4.2 - Construct views with different selection grounds
For an exact/lexical route, define normalization and fields deliberately. Preserve exact identifiers and error codes even if a second field lowercases, stems or expands terms. Decide how the languages in the actual corpus are tokenized. A BM25-style inverted index ranks term matches; a raw-text search remains useful for newly changed or unusual strings that the index does not yet represent. Test punctuation-sensitive identifiers, inflections and negation-bearing phrases on actual source snippets. Full-text relevance and exact equality are different operations.
For a dense route, select an embedding model whose language and domain behaviour can be examined on the workload. Embed source units, not only cards. Use the compatible query encoding and similarity operation; record model, dimensions, preprocessing, normalization and unit-generation versions. Keep independently built spaces separate until compatibility has been established. Comparing a new query embedding with vectors from an incompatible earlier model yields a numerically computable but uninterpretable score.
An approximate nearest-neighbour index trades search work and storage for approximation. On a manageable, representative subset, compare its returned neighbours with an exact search in the same vector space. This tests approximation, not semantic relevance. Separately judge whether the model ranks the needed source distinctions. A high approximation recall cannot repair an embedding that never captured the relevant condition. Token-level late interaction is another candidate when finer matching earns its additional representation and query cost; KCAE.Profiles:1 explains the ColBERT lineage without making it mandatory.
For a structural route, extract relations whose meaning is established: section containment, explicit source references, table ownership, edition succession or symbol definitions from a qualified code parser. Store relation kind, source basis and target identity. A query such as “what defines this symbol?” can use that structure directly. A query such as “what would improve this work?” needs additional interpretation; proximity in a document graph does not establish useful contribution.
For broad corpus questions, consider summaries or entity/community views at several scales. They can expose themes that local nearest-neighbour retrieval scatters. Retain source membership for each summary and permit leaf/original retrieval independently of the hierarchy. If an answer claims “the main objections across this corpus,” its sampling and coverage question differs from finding one useful objection. KCAE.SEARCH:4.6 constructs the coverage-to-answer operation; KCAE.ASSESS checks both its individual claims and the scope of its aggregate. A view supplies material to that operation, not the broad conclusion by itself.