Patterns
KCAE.USE - Choose an Access Arrangement for the Work
Type: Method pattern Status: Stable
KCAE.USE:1 - Problem frame
Use this when you are constructing, buying or changing access to a documentary corpus and several technically feasible arrangements would serve different uses. The governed object is that access arrangement under a workload. The first useful result is a bounded design question with credible alternatives, sources, receiving results and decisive constraints. You need access to representative work and someone who can settle the relevant source authority and data-flow rules. If one known source already answers a one-time question, use it directly.
KCAE.USE:2 - Problem
“Improve search” does not identify the result that matters. A system can increase matching passages, reduce a single model bill and still produce worse decisions or consume more recipient effort. A design chosen from a familiar tool can also make an incidental local baseline into an unnecessary prerequisite for every corpus.
KCAE.USE:3 - Forces
Frequent similar requests reward preparation; rare novel requests reward flexible inspection. Different languages and source formats enlarge discovery but increase interpretation and extraction work. Strong currentness, privacy and latency requirements can rule out otherwise attractive routes. The engineer must compare attainable complete results rather than maximize one retrieval score.
KCAE.USE:4 - Solution
KCAE.USE:4.1 - Recover the receiving operation
Observe or reconstruct one work episode. Name who must obtain what result, what they can already do, which sources and tools are available, and which error would change the next action. Write a success condition in the vocabulary of that work: “The analyst obtains the governing retry condition and the missing destination fact before suggesting automation,” rather than “The top result is relevant.” Preserve a different use when its success condition changes, such as explaining which rule was available during an old incident.
Separate the contribution supplied by access from the later domain operation. An access system may return a governing clause and comparison evidence; a qualified domain engineer still approves an equipment change. If the user expects the system to perform that approval, the arrangement needs the relevant capability and authority, not merely a better source link. Establish which receiving results are to be supplied by people, software or a coordinated combination.
KCAE.USE:4.2 - Establish the source and authority boundary
For every source class that can change this use, obtain its membership, edition rule, applicable authority and access conditions. Ask the responsible owner what makes a source current for this decision and what happens when sources disagree. Keep a record precise enough to implement that decision: for example, “Published English administrator manual governs supported automation; the localization assists interpretation; release notes override the named clauses from their effective date; the installed destination version comes from the configuration service.” This is a supplied organizational rule in the example, not a universal hierarchy of document types.
Distinguish source unavailability, lack of permission, extraction failure and a source known to contain no relevant passage after an adequate bounded inspection. They imply different next moves. A provider’s inability to expose historical editions limits historical reproducibility even if it rejects a stale snapshot identifier. When authority is disputed, obtain the owner’s decision or return competing conditional readings; do not resolve the dispute by upload date.
Define the data boundary for each route. Public source text, a private query, case records, assessor prompts, logs, caches and recipient messages are separate flows. A remote model may be permitted for a public manual while the event description must stay local. Construct a permitted projection or use a local alternative. Removing a name does not by itself establish that sensitive meaning has been removed; have the responsible data owner settle the permissible projection at the grain the use requires.
KCAE.USE:4.3 - Turn anticipated use into a workload
Sample real requests and failures when available. Include known-name lookup, uncertain vocabulary, cross-document synthesis, historical use, no useful advice and source changes. Record approximate corpus size and growth, request rate, update rate, language/format distribution, acceptable latency, required source return and recipient preparation. Distinguish observations from forecasts. If no log exists, construct explicit scenarios and treat the workload as a design assumption to revisit after use.
Select a finite horizon and identify who bears each cost. Build and extraction costs occur before requests; maintenance follows changes; discovery and reading occur per question; interpretation and action burden fall on the recipient. Retention and notification add separate costs. Shared work is counted once, while costs displaced to a colleague remain costs. KCAE.EVAL supplies the comparison method; the workload supplies its population and use conditions.
KCAE.USE:4.4 - Compare complete feasible arrangements
Construct at least the credible existing or simple arrangement and the proposed alternative at the same receiving result. A small dossier can use full reading with a cache. A large evolving corpus can use full-text plus semantic retrieval and independent block inspection. A structured codebase can use qualified symbol navigation. Include what each needs for source preparation, assessment, delivery, updates and recovery. An unsupported capability such as history deletion makes an arrangement unavailable until that capability is supplied.
Eliminate a choice that violates a hard condition before comparing cheaper variants. Then compare quality and cost by workload segment. An untested semantic route can be a reasonable initial design hypothesis for frequent cross-language questions; it does not need an artificial history of lexical failures to be considered. Equally, a graph is not earned by its name: identify the relations its intended queries actually require and the cost of maintaining them.
Use a break-even calculation only for commensurable quantities. Suppose A costs 180 engineering minutes to prepare, 8 minutes per source change and 6 minutes of recurring human effort per request. B costs 780 minutes, 28 minutes per change and 2 minutes per request. For N requests and U changes, B uses less of this measured/estimated effort when 4N exceeds 600 + 20U. With U = 10, that is N greater than 200. These invented values omit model charges and elapsed delay, which remain separate constraints. Neither arrangement is acceptable solely because its effort total is lower; it must supply the required result.
KCAE.USE:4.5 - Return the design basis and its reopen conditions
The result identifies the chosen or proposed arrangement, first usable result, alternative retained for comparison, unknown premises and observations that would change the choice. Feed source requirements to KCAE.SOURCE, discovery requirements to KCAE.INDEX, delivery capability to KCAE.DELIVER and adoption questions to KCAE.EVAL. A single concise account can carry this basis; no mandatory per-query form follows.
KCAE.USE:5 - Archetypal Grounding
CedarBench receives repeated unfamiliar-language support requests over a large corpus. Persistent full-body semantic retrieval is selected for comparison from that workload, alongside lexical and exact routes. The installed protocol remains local; only public text and an authorized abstract query may leave the organization. In a second use, an analyst asks which of three known documents was available during an earlier incident. Direct reading of those editions is sufficient. The same source collection supports two different access choices because the receiving operation changed.
KCAE.USE:6 - Bias-Annotation
Logs favour questions users already know how to ask. Include episodes where they compensated manually or abandoned inquiry. Engineers also tend to omit the recipient’s reading effort and the maintainer’s recovery work; obtain those costs from their actual performers.
KCAE.USE:7 - Conformance Checklist
Can the named recipient obtain the stated result with the supplied preparation and tools? Is source authority obtained independently of ranking? Are private data flows explicit? Does the comparison include a credible simpler arrangement and the same required quality? Are forecasts and unknowns marked where they affect selection?
KCAE.USE:8 - Common Anti-Patterns and How to Avoid Them
Starting with “add vectors” hides the workload that could justify them; reconstruct the receiving uses first. Requiring every corpus to fail an exact-term search before considering a persistent semantic route hides foreseeable repeated need; compare feasible arrangements directly. Treating the newest file as governing confuses observation with authority; recover the owner’s applicable rule.
KCAE.USE:9 - Consequences
The design becomes explainable and testable. Acquiring representative episodes and authority information costs effort, but it prevents implementing a fast answer to the wrong question. The result can legitimately be retention of the existing arrangement or a bounded missing-capability decision.
KCAE.USE:10 - Architectural Rationale
Workload and receiving result connect all later engineering choices. A global technology ranking cannot express their different constraints. Keeping authority, access and cost separate also prevents a cheap but impermissible route from winning through a combined score.
KCAE.USE:11 - SoTA-Echoing
The practice question is whether to precompute, inspect on demand or combine them. Current provider accounts support both just-in-time exploration and long-context/cached reading; the latter still has task-dependent retrieval limits. KCAE.Profiles:1 compares those mechanisms with persistent retrieval. This pattern adopts workload-specific comparison, while rejecting universal vendor or lexical-first ordering. C.11.DUA supplies the broader burden question; the domain addition is the lifecycle workload and source/data-flow contract. Reopen the choice when demand, source volatility, recipient effort or executor capability changes.
KCAE.USE:12 - Relations
F.1 supplies question-relative source selection; C.11.DUA supplies practical burden comparison. KCAE.SOURCE realizes the chosen source boundary. KCAE.EVAL turns the design questions into evidence. All other KCAE methods can return a failed premise here without invalidating unrelated uses.
KCAE.USE:End
KCAE.SOURCE - Prepare Addressable Sources without Losing Their Context
Type: Method pattern Status: Stable
KCAE.SOURCE:1 - Problem frame
Use this when documents must become searchable or machine-readable, when a hit cannot be opened reliably, or when extracted text can lose a decisive relation. The governed object is the relation between an exact source edition, its extracted/segmented representations and obtainable reading context. The result is a source inventory and reader that can return that context with known fidelity limits. A clean, fully addressable source already sufficient for the use needs no new extraction pipeline.
KCAE.SOURCE:2 - Problem
Document bytes, extracted text and meaningful source units are different. A table flattened into lines can attach a permission to the wrong model; OCR can drop “not”; a translation can omit an exception; a heading-based address can drift. A flawless embedding of the resulting text preserves those errors.
KCAE.SOURCE:3 - Forces
Small retrieval units discriminate topics but sever context. Large reading units preserve relations but cost attention. Normalization improves matching but can erase exact identifiers or provenance. Stable identity must survive harmless movement while still revealing a meaningful edition change.
KCAE.SOURCE:4 - Solution
KCAE.SOURCE:4.1 - Preserve the source and describe its provenance
Obtain an immutable source edition or make a permitted snapshot before extracting it. Preserve publisher/collection, document identity, edition, effective interval when relevant, access class and original format. Use a content digest to detect byte changes; retain semantic identity separately so that identical text in different documents is not accidentally merged. A corpus snapshot fixes a set of source editions, not merely the name of a branch that can move.
For a mutable collection, first capture a coherent set through a source-system snapshot or an approved freeze/copy procedure. A manifest computed while files are changing can describe an impossible combination. If the source system cannot provide a coherent snapshot, state the consistency limit and narrow the usable operation; do not label a best-effort crawl as an atomic edition. KCAE.CHANGE develops live and historical reading contracts.
KCAE.SOURCE:4.2 - Extract the relationships needed for reading
Choose the extractor by the actual carrier. Markdown and HTML can expose headings, lists and links. Born-digital PDFs can still scramble columns or reading order. Scanned documents need OCR and inspection of decisive glyphs, tables and notes. Audio/video transcripts need time offsets and returns to the original segment when tone, a demonstration or a visual relation matters. Machine-readable data needs schema, units, null meaning and provenance. Store the original alongside the extraction while permission permits it.
Define a small extraction qualification set from the source’s likely failures: an exception containing negation, a multi-column page, a repeated heading, a table with merged cells, a footnote and a cross-document reference. Compare rendered source and extraction for the use-changing relationships, not only character count. Random sampling can detect wider drift, but a sample does not certify every uninspected decisive passage. Mark unqualified regions and allow direct visual or specialist reading where needed.
For a table, retain its title, row identity, column identity, units, span relations and notes. An indexed cell can be small, but its read operation expands to these relations. For example, the extracted CedarBench cell “No” is unusable. The qualified reading says “Destination protocol v1 / automatic resend / No; reconcile receipts manually; note a: receiver lacks duplicate suppression.” The source image or native table remains reachable. If OCR cannot distinguish v1 from v2, return an uncertain extraction and inspect the original before applying the condition.
Distinguish faithful extraction from an authored interpretation. A generated explanatory prefix can make an index easier to search; it remains a derived claim, with the original text separately available. Translation has the same discipline: keep source language and edition, preserve important negations, numerical units and scope, and compare the consequential reading with a competent translator or the source owner when reliance requires it. An old translation can be a search aid or historical artifact without being the current operational authority.
KCAE.SOURCE:4.3 - Separate finding units from reading units
Segment by intelligible source structure before imposing a token ceiling. Preserve parent headings, paragraph/list membership, table and figure ownership, explicit references and adjacent conditions. Split an oversized section into bounded retrieval units while leaving its reading unit intact. Overlap can reduce boundary losses, but blind overlap duplicates text and still does not recover a distant definition.
Give each unit a document/edition identity, a local locator, an exact source span where available and an expansion relation. Keep retrieval text, structural metadata and generated context in distinguishable fields. Record which parents or definitions a generated context used. That makes a later parent change capable of invalidating an apparently unchanged child representation.
Choose unit size using the retrieval question and the operation that consumes the hit. A broad research theme may benefit from section summaries; a protocol exception may need paragraph or cell retrieval. Keep an independent raw/full-body path so that summary selection is not the only way to reach the source. Test a question whose deciding text lies beyond the initial unit; the opening operation should recover it without requiring the user to guess its exact location.
KCAE.SOURCE:4.4 - Implement and test the source return
A practical exact reader accepts a source edition and a locator, then returns the requested material, its structural envelope, continuation information and the edition actually served. A pagination cursor should bind the edition and requested unit. Expose an end marker or explicit remaining extent so that truncated output cannot masquerade as complete reading. Do not rely on a line number without its edition.
When a heading moves, a semantic document/section identity can help resolve the new location, but exact historical reading still selects the old edition. When a section splits or merges, preserve explicit predecessor/successor relations and inspect them before treating them as equivalent. A fuzzy match may suggest a new location; it cannot silently establish identity or preserve an earlier judgement.
Verify build, persist, reload, find and read on a fixture that includes additions, deletions and shifted headings. Check that the returned text corresponds to the selected source, that a deleted locator yields an explicit result and that a child unit expands through the correct note. Preserve the qualification result with the extraction/version where the installation needs it. KCAE.EVAL distinguishes this mechanical correctness from useful question answering.
KCAE.SOURCE:5 - Archetypal Grounding
CedarBench’s table has two rows: protocol v1 requires receipt reconciliation; v2 permits idempotent resend after absence is confirmed. A flat extractor returns “v1 v2 No Yes” and loses the note. The engineer retains a row/column representation and source-page locator, verifies both rows against the page, and indexes each row with its heading and note. A query can now find the v1 limitation; opening it returns the complete condition. When a later PDF changes only its pagination, source content can remain semantically equal while exact page coordinates change; the new reader generation updates the locators.
KCAE.SOURCE:6 - Bias-Annotation
Clean text collections make extraction seem trivial. Include the actual poor scans, minority language and awkward tables that the workload contains. Generated context can project the author’s interpretation onto ambiguous text; mark it as derived and preserve the source ambiguity.
KCAE.SOURCE:7 - Conformance Checklist
Can every decisive hit return to its source edition and reading context? Are source and generated text distinguishable? Were characteristic extraction failures inspected against the original? Does the reader expose truncation, inaccessible regions and uncertain mappings? Does persistence/reload preserve the same return?
KCAE.SOURCE:8 - Common Anti-Patterns and How to Avoid Them
Indexing only summaries makes their omissions universal; retain searchable original units. Hash-only identity merges unrelated copies; include source provenance. Returning a table cell without row, column and note creates a false claim; expand the reading unit. Treating OCR as semantic proof ignores its loss boundary; qualify the relied-on extraction.
KCAE.SOURCE:9 - Consequences
The arrangement gains trustworthy source return and reusable extraction. It costs storage and qualification work, and some formats continue to require visual or expert reading. Explicit uncertainty prevents a damaged extraction from appearing to be a confident source claim.
KCAE.SOURCE:10 - Architectural Rationale
The source remains the basis for later questions, while retrieval units optimize discovery. Their separate identities let the system change chunking or embedding without changing what an exact read means. Preserved context dependencies make selective refresh possible.
KCAE.SOURCE:11 - SoTA-Echoing
The practice question is how to retrieve small passages without making them contextless. Anthropic’s 2024 Contextual Retrieval is a useful historical mechanism: attach context to chunks before indexing. This pattern adapts that idea by retaining generated context separately and exposing its source dependencies; it does not adopt the reported corpus-wide gains as a forecast. Full-section indexing is a simpler rival with greater per-hit reading cost. Reopen the unit design when an actual miss depends on omitted context or when expansion overwhelms the reader. NOT.5 supplies a fuller translation/loss method where notation changes; ordinary file parsing alone is not an A.6.3.RT semantic translation claim.
KCAE.SOURCE:12 - Relations
KCAE.USE supplies the source roles and allowed use. KCAE.INDEX consumes retrieval units; KCAE.DELIVER consumes reading units; KCAE.CHANGE maintains their identities and dependencies. A.6.3.RT and NOT.5 govern their respective representation-loss questions without replacing this domain extraction and source-reader construction.
KCAE.SOURCE:End
KCAE.INDEX - Build Complementary Ways to Find Unknown Material
Type: Method pattern Status: Stable
KCAE.INDEX:1 - Problem frame
Use this when repeated questions concern a large corpus and the useful source, phrase or contribution is often unknown. The governed object is the reusable retrieval arrangement: indexed units, selection grounds, query operations, result combination and independent source access. Its useful result is an implemented or implementable route design that can expose material a single representation would miss. The source reader and workload must be available. If direct reading of a bounded dossier already supplies the result at acceptable cost, it is a sufficient rival.
KCAE.INDEX:2 - Problem
A lexical system can miss another profession’s vocabulary. A dense embedding can blur negation or a rare identifier. A summary tree can exclude the only branch containing an exception. Adding a second endpoint over the same summaries preserves the same omission. Merely listing these technologies leaves the engineer without the construction that joins their strengths, pays their maintenance cost and preserves a way beyond their common blind spots.
KCAE.INDEX:3 - Forces
Prepared views reduce recurring work but cost construction and refresh. Several views increase discovery opportunities and duplicate results. Early filtering saves work and can exclude the unknown answer. Rich context improves interpretation but lowers the number of candidates that fit a fixed budget. The aim is a useful and maintainable arrangement for the workload, not an exhaustive representation of all knowledge.
KCAE.INDEX:4 - Solution
KCAE.INDEX:4.1 - Specify the retrieval operations and their source units
Start with the question families from KCAE.USE and the addressable units from KCAE.SOURCE. Define what each route returns: source edition, unit identifier, matched text or span, route and query used, ranking value with its local meaning, and a way to open the reading unit. Preserve version, permission and authority metadata independently of the score. The first interface is a candidate-finding operation; it does not claim that the material supports the user’s action.
Build an exact reader and an inventory enumerator even when semantic retrieval is the primary discovery route. The enumerator lists the permitted units in a selected snapshot without requiring a relevance match. It enables update validation and a direct original scan. A source corpus that cannot be enumerated can still be searched through its provider, but the independent-scan promise is then bounded by what that provider actually exposes.
Choose retrieval units by the distinctions the queries need. Index whole source bodies through bounded units, including examples, appendices and footnotes where they can matter. Index titles, controlled vocabulary and summaries as additional fields or views. Keep their source spans and derivation separate so that a hit from an explanatory prefix cannot be misquoted as original text. KCAE.SOURCE supplies the expansion from a small hit to its enclosing conditions.
KCAE.INDEX:4.2 - Construct views with different selection grounds
For an exact/lexical route, define normalization and fields deliberately. Preserve exact identifiers and error codes even if a second field lowercases, stems or expands terms. Decide how the languages in the actual corpus are tokenized. A BM25-style inverted index ranks term matches; a raw-text search remains useful for newly changed or unusual strings that the index does not yet represent. Test punctuation-sensitive identifiers, inflections and negation-bearing phrases on actual source snippets. Full-text relevance and exact equality are different operations.
For a dense route, select an embedding model whose language and domain behaviour can be examined on the workload. Embed source units, not only cards. Use the compatible query encoding and similarity operation; record model, dimensions, preprocessing, normalization and unit-generation versions. Keep independently built spaces separate until compatibility has been established. Comparing a new query embedding with vectors from an incompatible earlier model yields a numerically computable but uninterpretable score.
An approximate nearest-neighbour index trades search work and storage for approximation. On a manageable, representative subset, compare its returned neighbours with an exact search in the same vector space. This tests approximation, not semantic relevance. Separately judge whether the model ranks the needed source distinctions. A high approximation recall cannot repair an embedding that never captured the relevant condition. Token-level late interaction is another candidate when finer matching earns its additional representation and query cost; KCAE.Profiles:1 explains the ColBERT lineage without making it mandatory.
For a structural route, extract relations whose meaning is established: section containment, explicit source references, table ownership, edition succession or symbol definitions from a qualified code parser. Store relation kind, source basis and target identity. A query such as “what defines this symbol?” can use that structure directly. A query such as “what would improve this work?” needs additional interpretation; proximity in a document graph does not establish useful contribution.
For broad corpus questions, consider summaries or entity/community views at several scales. They can expose themes that local nearest-neighbour retrieval scatters. Retain source membership for each summary and permit leaf/original retrieval independently of the hierarchy. If an answer claims “the main objections across this corpus,” its sampling and coverage question differs from finding one useful objection. KCAE.SEARCH:4.6 constructs the coverage-to-answer operation; KCAE.ASSESS checks both its individual claims and the scope of its aggregate. A view supplies material to that operation, not the broad conclusion by itself.
KCAE.INDEX:4.3 - Make selection filters explicit
Apply access restrictions before material can reach an unauthorized component. Other filters are claims about the question. If the user asks about current operational policy and an authority rule selects the governing edition, use that rule; if the question asks about history or conflicts, retain the relevant older editions. Do not infer a subject/Suite filter from the first few words and then make every route inherit it.
Compare early filtering with later assessment. Early exclusion is appropriate when the predicate is reliable and required, such as an access boundary. A guessed topic is better used as one retrieval preference with a broad route still available. When a route supports only post-filtering of its first K hits, permission-safe candidate inspection may still suffer poor recall because eligible hits were below K. Over-fetching or a natively filtered index can help; measure the actual operation and report the boundary.
A route can fail because the useful document is absent, its extraction is incomplete, the model missed its meaning, the approximate search lost it, a filter excluded it, or the candidate window was too small. Preserve enough route information to distinguish these causes during evaluation. Otherwise every failure invites another model while the source never entered the index.
KCAE.INDEX:4.4 - Combine candidates without inventing a common probability
Run the selected routes against the same intended corpus/edition conditions. Deduplicate by source identity and overlap, while retaining which routes found a candidate. Keep competing editions distinct where the use needs them. Merge adjacent fragments into one reading candidate when their overlap would otherwise consume the budget repeatedly; do not merge conflicting source claims merely because their words are similar.
Raw lexical, vector and graph scores usually have different meanings. A simple candidate-combination baseline is rank fusion. Reciprocal rank fusion assigns a candidate the sum of 1/(c + rank) across lists in which it appears. It uses order rather than pretending that raw scores share a scale. The positive constant c controls how much top ranks dominate; the original 2009 study used 60, which is historical experimental selection rather than a universal setting. RRF source.
Fusion still favours material with several appearances. Preserve a tested allocation for route-unique candidates before the later shortlist: for example, take a fused core, then admit the highest-ranked unrepresented candidate from each materially different route, within the same budget. Treat query paraphrases from one model as related search attempts, not independent votes proving relevance. Compare this diversity policy with simple fusion on held-out questions; keep it only where it recovers valuable contributions at acceptable cost.
For a concrete constructed pool, suppose lexical order is A, B, C and dense order is D, A, B. With c = 60, A receives 1/61 + 1/62, B receives 1/62 + 1/63, and D receives 1/61. A and B win the top two positions because each appears in both lists. The route-unique D disappears from a two-item shortlist. Retaining D for inspection may expose a cross-language exception. Its eventual usefulness must be judged from the source; its uniqueness is a reason to inspect, not evidence of truth.
Choose candidate-window size from downstream reading capacity and error costs. A pool of 100 hits is an internal search result, not a request to stuff 100 fragments into the principal agent’s context. Assess and expand within the search service or separate reader where supported; deliver the sufficient result through KCAE.DELIVER. Too narrow a pool cannot be repaired by a perfect reranker, because the relevant source never reaches it.
KCAE.INDEX:4.5 - Provide an independent original-text route
Implement a route that can inspect source units without passing the same relevance filter. Enumerate permitted units from the snapshot, divide them into bounded reading batches, include their source addresses and necessary local context, and ask question-relative screening questions. Preserve the unprocessed extent and batch outcomes. The examiner may be a person, a generative model or a qualified typed assessor. The route can be used initially for a rare novel question, as a complement to indexes, or after a consequential suspected miss.
A first pass can ask whether a block contains a potentially useful condition, explanation or method; a second expands promising blocks into reading units. A later pass can search for an unresolved relationship across the retained sources. A single winner from each batch is unsafe where several contributions are needed: allow multiple candidates and an explicit no-candidate/unknown outcome. Scores normalized inside different batches are not globally comparable probabilities. Reassess finalists together or use separately qualified pointwise criteria.
A large corpus makes complete inspection expensive. Schedule by source order, strata or a declared sample that is independent of the failed selector; use affordable partial coverage when that can answer the current question. A stratified sample is a diagnostic route, not an exhaustive scan. If the route first chooses batches by the same summary classifier that failed, its independence has been lost. If the examiner reads all blocks but misses a cross-block relation, byte coverage has increased without semantic completeness. KCAE.SEARCH preserves both limits in its return.
KCAE.INDEX:4.6 - Join the routes to maintenance and evaluation
Publish each view with its source generation, covered subset and build configuration. Define readiness per operation: an exact reader can be ready before dense indexing; a graph can support containment while semantic edges remain unqualified. KCAE.CHANGE supplies coherent switching, partial readiness and deletion handling. Decide how new or changed units enter current queries while expensive views catch up.
Test complementary retrieval on cases independently authored from the working needs. Include a rare identifier, vocabulary mismatch, late exception, graph-relevant question, broad synthesis and source not represented by any summary. Measure sufficient contributions recovered within the recipient’s budget, not just union size. Remove a redundant route if its maintenance and query cost earn no useful difference. Add or replace a route when a repeated consequential gap exposes a different required distinction.
KCAE.INDEX:5 - Archetypal Grounding
CedarBench has a hypothetical snapshot of 6,000 documents and 32,000 retrieval units. Many support requests use customer vocabulary absent from the English manual. The engineer builds a lexical field preserving product identifiers, a dense field from complete source units, structural parent/reference relations and an exact reader. The query “our analyst removes double reports every morning” yields a throughput article lexically and a receipt-reconciliation note semantically. The pool retains both and opens their conditions.
A new supplier note appears before its embedding is ready. The current-source inventory and direct reader expose it; the route reports that dense coverage lags. An independent block pass can examine the new units. This is a designed integration of persistent additional retrieval and current source reading, not proof that the dense route is superior. In a three-document historical dossier, a cached full reading can replace this machinery with less total effort.
KCAE.INDEX:6 - Bias-Annotation
Queries derived from titles favour title-based systems. Test vocabulary and relationships supplied by real work. Correlated retrievers can create apparent agreement; preserve provenance and inspect route-unique gains. Majority-language averages can hide important minority-language losses.
KCAE.INDEX:7 - Conformance Checklist
Does the design explain source units, encoding, filters, candidate combination, source expansion and maintenance? Can the independent route reach material the principal selector excludes? Are scores interpreted locally? Can a known loss be attributed to ingestion, representation, filtering, approximation, pool size or later assessment? Has the chosen arrangement been compared at the actual reading budget?
KCAE.INDEX:8 - Common Anti-Patterns and How to Avoid Them
Calling several searches over one card store “independent” hides a shared bottleneck; add an original-body route. Multiplying scores turns retrieval ranks into unsupported probabilities; use a declared ranking policy and a separate assessment. A graph of document mentions is not a graph of necessary method relations; interpret the receiving relationship through KCAE.COMPOSE. A failed lexical trial is not required to justify engineering a semantic route for a forecast workload.
KCAE.INDEX:9 - Consequences
Unknown material becomes reachable through more than one selection ground, and the system can explain which parts remain unsearched. The design increases storage, update and evaluation work. It cannot guarantee future semantic completeness; it provides attainable ways to investigate consequential omissions.
KCAE.INDEX:10 - Architectural Rationale
Prepared representations and query-time inspection solve different cost problems. Their common source-address contract makes them interchangeable and combinable, while route independence protects against compression becoming an exclusive ontology. The evaluation decides how much of that plurality earns its cost.
KCAE.INDEX:11 - SoTA-Echoing
The current practice question concerns hybrid, graph, multiscale and agentic retrieval under a changing corpus. KCAE.Profiles:1 compares their constructions, including the historical ColBERT, HyDE, GraphRAG, RAPTOR and LazyGraphRAG lines with current source/view studies. This pattern adopts multiple candidate grounds and exact source return, adapts fusion to preserve useful unique findings, and rejects a technology-independent winner. Code-domain findings remain code-domain evidence. Reopen the selection after a changed query family, model, update regime or measured loss.
KCAE.INDEX:12 - Relations
KCAE.SOURCE supplies units and reading expansion. KCAE.SEARCH operates the available routes for a developing question. KCAE.ASSESS decides contribution after retrieval. KCAE.CHANGE maintains source/view correspondence, and KCAE.EVAL tests marginal and whole-use value. CMP.10 can supply the underlying data-structure and operation-cost reasoning.
KCAE.INDEX:End
KCAE.SEARCH - Continue a Search as the Question Develops
Type: Method pattern Status: Stable
KCAE.SEARCH:1 - Problem frame
Use this when a question has no sufficient known source, when a first result leaves a consequential gap, when new reading changes what should be sought, or when the answer concerns a collection rather than one passage. The governed object is one bounded inquiry through available access routes. The result is useful inspectable material, a supported synthesis within a declared corpus scope, or a precise unresolved/access limit, with enough history to continue without repeating a failed interpretation. An already sufficient source or result is a direct exit.
KCAE.SEARCH:2 - Problem
A fixed-query loop can repeatedly find the same plausible material. A free-ranging agent can spend its budget without reducing the important uncertainty. Both can return “nothing found” while having considered only one vocabulary or one subset. The next search must follow the receiving question and what the previous reading actually changed. A broad answer additionally needs a way to combine source claims under a declared coverage basis; finding several relevant passages does not perform that aggregation.
KCAE.SEARCH:3 - Forces
Preserving the original question protects its meaning, while inquiry legitimately changes it. Broader exploration improves opportunity and consumes scarce reading resources. Early useful candidates justify focused reading, but can anchor interpretation. A useful stopping decision needs the limits of the available routes, not a claim of universal absence.
KCAE.SEARCH:4 - Solution
KCAE.SEARCH:4.1 - Keep the question, facts and variants separate
Retain the user’s wording or source episode, the receiving result, known facts and current unknowns. Add a corpus-language translation, a description of the missing result, and a rival interpretation when those can expose different material. For CedarBench, “buy faster workers” remains the request; “prevent duplicate resend” is a search hypothesis supported by the morning correction episode, not an established cause.
Generate variants from relations as well as nouns. Ask what operation might produce the missing result, what condition might defeat it and which result an already found method requires. A hypothetical answer can provide retrieval vocabulary, but any invented product, causal link or configuration remains outside the case facts. When an actual reading changes the question, state the change and preserve the unresolved part of the earlier one.
KCAE.SEARCH:4.2 - Select the next route by the missing contribution
Choose a route whose selection ground can plausibly reach what is missing. Known identifiers favour exact lookup. Vocabulary mismatch favours corpus-language variants or semantic retrieval. A missing definition favours structural expansion. A global synthesis can need several regions or multiscale summaries. An index blind spot or a newly changed region can justify original-block inspection. The available arrangement may supply only some of these routes; lack of a route is a capability limit.
Keep a small search state: original/current question, inspected source units and editions, useful contributions, disputed assumptions, pending source returns, covered extent and remaining resources. Avoid retaining every verbose search result in the principal reader’s context. The state guides the next action; an authorized detailed log can remain outside that context for diagnosis.
For a question with several needed aspects, state what each contribution must establish and mark which remains missing or disputed. More documents about an already answered aspect do not supply another one. Compare new material by the contribution it adds, not only by whether its identifier is new. For example, ten resend paragraphs cannot replace the missing destination configuration; once the governing rule and configuration are sufficient, another synonym search needs a separate reason to continue.
The next action should have an explicit expected contribution: “Open the compatibility note to determine whether protocol v1 supports deduplication,” rather than “search more.” When the required fact is the customer’s installed version, use a permitted configuration source or ask the customer. No amount of searching the public manual will establish that local fact. This is a return from finding literature to obtaining case data.
KCAE.SEARCH:4.3 - Read promising material and test the interpretation
Open the source around decisive hits through KCAE.SOURCE. Expand through the condition, exception, definition or example that can alter the proposed contribution. Preserve conflicts instead of choosing the more convenient passage. A preliminary score can prioritize this reading; KCAE.ASSESS turns it into a supported or unresolved contribution.
Actively seek a discriminating countercase when two readings would lead to materially different actions. If a retry paragraph appears to allow unattended resend, inspect restrictions on receiver versions and receipt state. That is targeted challenge, not a demand to read every related publication. If the interpretation survives and the receiving result is available, stop. If it fails, use the discovered condition to formulate the next question.
KCAE.SEARCH:4.4 - Escape a consequential blind spot
When the query family is unfamiliar, the user rejects the interpretation or the necessary result remains missing, inspect the selection assumptions. Was the corpus reduced to one department? Were examples excluded? Did every query use the same translated diagnosis? Change the relevant assumption and choose a route not bound by it. The alternative need not be expensive: a glossary term, source table of contents or colleague’s exact reference may suffice.
For a direct semantic pass, use the independent enumerator from KCAE.INDEX:4.5. Partition by actual accessible source units, not by the failed relevance category. Allocate a bounded initial portion; retain source order or sampled strata, examined extent and reasons for further expansion. Inspect promising units fully and search for missing links across them. A new question can invalidate a prior “not useful” screen, so cache that result with its criterion and question rather than suppressing the unit forever.
If only part of a corpus was examined, return that extent and the practical consequence. “No adequate current automation rule was found in the inspected public manual and release notes; the supplier bulletin was inaccessible” is actionable. “There is no rule” is a stronger conclusion unsupported by that search. A complete deterministic search can establish absence of an exact string from an exact corpus; it cannot establish absence of every relevant meaning.
KCAE.SEARCH:4.5 - Spend and stop at the receiving result
Allocate resources to finding, assessment and the reading still required for correct use. Reserve enough for the final necessary source context; spending all capacity on candidates prevents application. Use actual tool/model limits, including latency and rate limits. Stop or narrow the inquiry when a required channel is unavailable, a hard budget is reached or additional accessible work is unlikely to change the next decision enough to justify its burden.
The judgement can be qualitative. Compare the next affordable action with continuing from the present result: what uncertainty could it remove, what action would change and what would the inspection cost? If a numerical value-of-information calculation is used, obtain its probabilities and costs independently; a retrieval score cannot stand in for those values. KCAE.USE supplies the decision context, while C.11.DUA supplies the general burden question.
Return the source candidates and why they matter, the inspected conditions, unresolved premises and any coverage limit that changes use. A short reply can preserve all of these where the case is simple. Distinguish “sufficient material found,” “current candidates rejected,” “missing fact,” “access constrained” and “budget stopped.” They imply different continuations and should remain different in a machine interface.
KCAE.SEARCH:4.6 - Construct a bounded corpus-wide synthesis
Use this branch when the requested contribution concerns a collection: its themes, objections, changes, contrasts or distribution of stated positions. First distinguish three results. An exemplar shows that a particular contribution occurs. A thematic overview organizes identified contributions and their differences over declared material. A frequency claim counts a defined feature in a defined unit population. Finding a good example answers the first question; repeatedly retrieving it cannot answer the other two.
Set the population and the claim before choosing the aggregation route. Specify membership, time or edition rule, permissions, and the unit about which the answer will speak. Eight current submissions, twelve stored files and eight submitting organizations can be three different populations. Decide whether earlier editions are historical evidence to compare or superseded copies to exclude from the current count. Distinguish documents that repeat one underlying report from independently produced evidence. For themes, state what “main” means in the receiving use: commonly recorded, explanatory of a contrast, or consequential to the decision. Frequency alone need not determine importance.
Obtain an inventory from the source system or KCAE.INDEX’s enumerator. Keep each included unit’s source identity and selected edition, its assigned reading portion, and whether the relevant content was read, excluded by a stated rule, inaccessible or still unexamined. If only a provider’s selected hits are available, that is the available set; do not call it the complete collection. A broad question can legitimately produce an exploratory overview of selected material, provided the answer keeps that scope.
Choose a route that can cover the required material at an affordable cost. For a small dossier, a qualified reader can read every included unit, record its question-relevant claims with source returns, and compare them directly. This avoids index and summary preparation and is a serious choice for infrequent inquiries. If the material exceeds one reader’s working capacity, partition the enumerated units into bounded reading portions. Give each portion the same question and inclusion rule. Split a long document without losing its unit identity; preserve the necessary cross-boundary context and reconnect claims that span portions. This direct partitioned map/reduce route is available without an entity graph or precomputed summaries.
For a large, repeatedly queried collection, use prepared summaries or communities where their saved query effort earns their preparation and update burden. Choose the regions or hierarchy levels to inspect and resolve their membership back to source units. A parent and its child summaries can describe the same evidence; several graph communities can reach one underlying document. Selection across levels, including RAPTOR’s collapsed-tree profile, can find useful material without inspecting the entire population. A query-guided route such as LazyGraphRAG can allocate reading to promising communities, but the retained selection boundary still limits the answer. KCAE.Profiles:1.3 compares these constructions. Use independent original inspection to investigate consequential regions or distinctions the prepared view may omit.
Map source material to claims that can be combined. For each reading portion, obtain the proposed answer-bearing claim, its source units and exact passages, the relevant entity/time/condition, and any contradiction or unresolved interpretation. Preserve which words are the source’s position and which relation the reader inferred. One portion may supply several themes, a counterexample, an exception or no relevant claim; do not require one winner or one positive answer. A “no claim” screen remains revisable when the question or coding criterion changes.
Carry lineage through intermediate summaries: a claim refers to its contributing source-unit set, not just to the name of the summary that repeated it. When an operation cannot preserve that lineage, use its summary to locate original passages before relying on the aggregate. Keep unread or excluded portions visible alongside positive results. A polished partial answer with no account of what it left out is insufficient input for a claimed whole-corpus conclusion.
Reduce by meaning and source support, not by the number of summary votes. Align claims that concern the same object, time, condition and predicate. Merge genuine paraphrases while taking the union of their underlying source-unit sets. Do not merge contrary positions or a conditional exception into the majority wording. Group the resulting claims into themes that answer the receiving question, then explain both recurring relationships and consequential differences. If an initial theme does not explain a substantial contrast, refine it and reopen the source portions whose interpretation can change; merely relabelling the final paragraph leaves the earlier coding unchanged.
For a descriptive count, define the predicate and count each eligible population unit once for that predicate. A unit may belong to several themes; disclose that those categories overlap rather than forcing their percentages to total one hundred. Retain unknown or unread units in the account of the denominator. Do not divide a relevance-selected hit count by the whole collection and call it prevalence. A claim about a wider population needs a justified sampling/estimation method and its assumptions; otherwise return the observed corpus count or a qualitative overview. The number of documents repeating a claim also does not by itself establish its truth.
Protect consequential minority evidence before compressing the answer. Maintain the supported objections, contrary cases, important qualifications and unresolved conflicts that could change the receiving decision, even when they are rare or rank poorly. Ask the relevant domain reader which differences have that consequence when ordinary interpretation cannot settle it. Reserve enough final reading and answer space for those differences, or explicitly narrow the promised answer. If intermediate claims exceed the reduction window, reduce them in further bounded groups while preserving source sets and these outstanding differences; keep the originals obtainable for the final comparison. This controls the next reading operation, not semantic loss by fiat.
Return an answer whose scope survives the final wording. State the population and edition basis, the supported themes or counts, material counterevidence and the unexamined or inaccessible remainder. Link decisive generalizations through their intermediate claims to originals. Verify that words such as “all,” “most,” “typical,” “increasing” and “consensus” have the required comparison or counting basis. A minority warning can be important without being typical; silence on a question is not agreement with another submission.
Stop when the intended bounded answer has adequate inspected support, consequential conflicts have dispositions, and the remaining affordable inquiry would not change that receiving result enough to justify its cost. A thematic overview can stop with a named uncovered region when the recipient accepts that narrower use. A complete descriptive count must instead obtain the necessary unit judgements or report the unresolved count; a representative estimate needs its own sampling basis. When the hard budget ends first, return a partial synthesis and the particular next region or ambiguity that matters. Reading every unit establishes an inspection extent, not guaranteed recognition of every possible meaning. KCAE.ASSESS checks the proposed answer against this basis before KCAE.DELIVER carries it to the recipient.
KCAE.SEARCH:5 - Archetypal Grounding
KCAE.SEARCH:5.1 - An evolving local question
The first CedarBench query finds worker-capacity instructions. The retained episode includes duplicate removal, so the second query concerns receipt reconciliation. The source then distinguishes destination protocols. The next useful operation is a permitted local configuration read. It returns v1, redirecting the answer from automatic resend to reconciliation. If the configuration channel is unavailable, the result is a conditional answer and one precise information request. The search has made useful progress without pretending that it established the missing fact.
KCAE.SEARCH:5.2 - A dossier’s objections, recurring themes and minority condition
Consider a constructed consultation on a research archive’s proposed deposit requirements. The receiving question is: “What objections do the current submissions raise, and which conditions should the archive examine before adopting the proposal?” There are eight submitting groups, twelve source files because four are superseded editions, and two overlapping generated summaries. The declared population is the eight groups’ latest submissions at the closing date. Earlier editions remain available for historical comparison; the summaries are access views. Neither adds another current respondent.
The eight short submissions fit direct reading in four portions of two, with the same question and source-role instructions. The reader opens every selected edition and returns the following claim map. These invented labels abbreviate source-addressed readings; an installed system retains the actual addresses and qualifiers.
| Current source unit | Question-relevant contribution | Place in the aggregate |
|---|---|---|
| R1 | Objects to metadata-entry effort and storage cost. | Metadata effort; storage cost. |
| R2 | Objects to metadata-entry effort. | Metadata effort. |
| R3 | Objects to metadata-entry effort. | Metadata effort. |
| R4 | Objects to metadata-entry effort and storage cost. | Metadata effort; storage cost. |
| R5 | Objects to metadata-entry effort. | Metadata effort. |
| R6 | Objects to storage cost. | Storage cost. |
| R7 | Warns that publishing precise collection locations could damage a protected site. | A consequential disclosure objection, even though raised once. |
| R8 | Says metadata effort is acceptable if the archive supplies a working template. | A conditional counter-position; inspect the template premise rather than counting this as an unconditional effort objection. |
For the descriptive count, a direct effort objection means an explicit rejection of the proposed metadata workload. A conditional acceptance such as R8 is coded separately and its unresolved condition remains in the overview. For this coding, the direct metadata-objection set is {R1, R2, R3, R4, R5}; storage cost is {R1, R4, R6}. A summary of R1–R4 and another of R3–R8 both mention metadata. They supply overlapping paths to those sets, not two more respondents or two independent confirmations. An older R5 file likewise cannot increase the count. The union operation preserves five current direct metadata objections and three storage objections out of eight, with R1 and R4 in both categories.
The resulting bounded answer is: direct objection to metadata effort is the most frequently recorded objection category in these eight current responses (five of eight), followed by storage cost (three of eight). Those counts describe the submissions, not all potential contributors or the strength of an objection. R8 identifies a template condition under which the effort concern may be resolved; R7 identifies a disclosure problem that an effort/cost majority does not answer. The archive needs to inspect those conditions before treating the aggregate as support for one uniform policy. Each theme returns to its listed sources, and the important contrasting statements return specifically to R7 and R8. No inference of consensus is made from the other groups’ silence about locations.
Now change one condition. R5 submits an authorized replacement before the closing date: after trying the supplied template, it withdraws the effort objection. Reopen R5’s coded contribution and the aggregate that used it, leaving the other inspected claims intact. The current metadata-objection set becomes {R1, R2, R3, R4}: four of eight, so “more than half object” is no longer justified. The template contrast strengthens in a stated way, but the disclosure concern remains. The old five-of-eight result still describes the earlier snapshot; keeping the old file or an unrefreshed summary cannot make it current. If the replacement’s meaning cannot be read, report four confirmed objections and R5 unresolved, rather than silently counting the previous position.