Library / Knowledge-Corpus Access Engineering Principles Framework
Jump to passage
In this reading

Link to current text

Published source confirmed at last check

Source changed 2026-10-03 08:25:59 UTC · snapshot created 2026-10-03 10:17:34 UTC · last check 2026-10-03 10:35:10 UTC

KCAE.Profiles:1 - Retrieval and query construction

These profiles are replaceable implementations of the methods, not successive generations every installation must adopt. Select from the receiving workload and test the complete arrangement. A name such as RAG, agentic search or knowledge graph does not specify source authority, reading sufficiency or update semantics.

KCAE.Profiles:1.1 - Exact, lexical and semantic units

Exact and lexical retrieval. Retain a resolver for named addresses and an inverted index for original terms. BM25 is a credible term-ranking choice: its weighting accounts for term occurrence, frequency and document length. Its tuning belongs to development data and the chosen retrieval unit. An exact technical identifier may deserve a separate lane because stemming or dense similarity can weaken its distinction. Lexical retrieval can work very well when the question and source share vocabulary; it can miss a useful paraphrase or another language. The algorithmic supplier is the historical BM25 treatment in Introduction to Information Retrieval.

Dense retrieval over original units. Encode source units with a selected document encoder, encode the question compatibly, and retrieve nearby vectors under the model’s similarity convention. Store source identity and context with each vector. An approximate nearest-neighbour index trades search work against agreement with exact vector neighbours; this agreement is not semantic recall. HNSW is an established historical algorithmic option, not a guarantee about the useful passages in a new corpus. Test multilingual queries, rare identifiers, conditions and negation. Keep a source route outside the embedding view.

Late interaction. The original ColBERT contribution represents a document and query with token-level embeddings and postpones fine-grained interaction until retrieval. It offers a richer matching alternative to a single vector per unit, with different storage and query costs. Its 2020 evaluation is historical evidence for that construction, not a current ranking of all retrieval models. Apply KCAE.SOURCE and KCAE.CHANGE to its larger representation just as to a simpler index.

KCAE.Profiles:1.2 - Query variants and contextual representations

Question reformulation can expose aliases, translations and the result a method must provide. In HyDE, a generated hypothetical document supplies text to encode for retrieval. That 2023 mechanism can bridge vocabulary without making the generated document evidence. Keep the original question, reject invented case facts, and test whether expansion helps the receiving population rather than only increasing pool size.

Contextual Retrieval attaches generated document context to chunks before embedding and lexical indexing, then combines retrieval and reranking. Its 2024 account supplies a useful context-preservation construction and reported provider experiments. Here it is a historical comparator: the prefix can make an otherwise ambiguous chunk findable, but it is derived text and has context dependencies to refresh. It cannot replace the source reading or warrant transfer of the reported gains.

For fusion, Cormack, Clarke and Buettcher’s Reciprocal Rank Fusion supplies the rank-based mechanism used in KCAE.INDEX:4.4. It avoids requiring comparable raw score magnitudes. The original constant and results belong to the original experiment. Choose pool depths and fusion behavior with the actual candidate population, and retain source-return and permission checks after fusion.

KCAE.Profiles:1.3 - Structural and multiscale retrieval

Explicit section containment, definitions and references are useful low-cost structure. A generated graph or hierarchy adds inferred relations and its own loss. Use those relations to propose material, then inspect the source support appropriate to the question.

GraphRAG §2.6 selects a community level, shuffles its summaries into bounded contexts, produces intermediate answers, then combines them within a final context budget after helpfulness-based filtering. Its direct source-text map/reduce comparator is also a serious alternative. Prepared communities can amortize repeated broad questions, while preparation, updating and successive compression add burden and loss. The reported answer comparisons do not establish completeness of every summary or resulting answer.

RAPTOR recursively clusters and summarizes material. Its query construction offers both traversal with pruning at successive levels and collapsed-tree retrieval across all levels. The latter can recover a useful node without requiring a successful top-down path, but a selected cross-level set is still not the whole source population. Preserve underlying membership when parent, child or overlapping summaries contribute to one answer. These 2024 sources supply implementable alternatives; KCAE.SEARCH:4.6 supplies the population and aggregation conditions for using their output in a bounded synthesis.

LazyGraphRAG builds noun-phrase co-occurrence communities without advance LLM summaries. At query time it develops subqueries, explores communities through relevance testing, groups relevant source chunks, extracts claims and reduces selected claims to an answer. A relevance-test budget bounds exploration. This shifts preparation toward query-time work and is a serious comparator where advance summary cost is hard to amortize. Its selected claims and stopping condition still need an honest coverage account; deferred interpretation does not eliminate selection loss. The provider’s bounded experiments do not transfer their quality/cost ratios to another corpus.

KCAE.Profiles:1.4 - Direct semantic inspection and long-context reading

A direct original-block pass is useful when the question needs distinctions outside prepared selectors, or when the corpus is small enough that preparation would not pay. Enumerate blocks from the source inventory, supply the bounded state and criterion, retain plausible and unresolved blocks, and inspect their reading closures. The TypeSafe semantic-find cookbook demonstrates addressed clause selection over a supplied document with a separate adequacy question. Its Jev 1.12 example is a bounded mechanism, not evidence of exhaustive semantic search in an arbitrarily large corpus.

Long-context reading can supply another direct profile, especially for a small stable dossier or repeated cached use. The Gemini long-context documentation treats caching as a cost option and warns that multiple-item retrieval performance can vary. Compare actual model/version behavior and full lifecycle cost. A large accepted input window neither ensures every relevant relation is recovered nor makes a changed source cache current. KCAE.DELIVER remains necessary even when the whole dossier fits.