KCAE.Profiles:1.1 - Exact, lexical and semantic units
Exact and lexical retrieval. Retain a resolver for named addresses and an inverted index for original terms. BM25 is a credible term-ranking choice: its weighting accounts for term occurrence, frequency and document length. Its tuning belongs to development data and the chosen retrieval unit. An exact technical identifier may deserve a separate lane because stemming or dense similarity can weaken its distinction. Lexical retrieval can work very well when the question and source share vocabulary; it can miss a useful paraphrase or another language. The algorithmic supplier is the historical BM25 treatment in Introduction to Information Retrieval.
Dense retrieval over original units. Encode source units with a selected document encoder, encode the question compatibly, and retrieve nearby vectors under the model’s similarity convention. Store source identity and context with each vector. An approximate nearest-neighbour index trades search work against agreement with exact vector neighbours; this agreement is not semantic recall. HNSW is an established historical algorithmic option, not a guarantee about the useful passages in a new corpus. Test multilingual queries, rare identifiers, conditions and negation. Keep a source route outside the embedding view.
Late interaction. The original ColBERT contribution represents a document and query with token-level embeddings and postpones fine-grained interaction until retrieval. It offers a richer matching alternative to a single vector per unit, with different storage and query costs. Its 2020 evaluation is historical evidence for that construction, not a current ranking of all retrieval models. Apply KCAE.SOURCE and KCAE.CHANGE to its larger representation just as to a simpler index.