Library / Knowledge-Corpus Access Engineering Principles Framework
Jump to passage
In this reading

Link to current text

Published source confirmed at last check

Source changed 2026-10-03 16:02:47 UTC · snapshot created 2026-10-03 16:03:51 UTC · last check 2026-10-03 16:30:20 UTC

KCAE.SEARCH:4.6 - Construct a bounded corpus-wide synthesis

Use this branch when the requested contribution concerns a collection: its themes, objections, changes, contrasts or distribution of stated positions. First distinguish three results. An exemplar shows that a particular contribution occurs. A thematic overview organizes identified contributions and their differences over declared material. A frequency claim counts a defined feature in a defined unit population. Finding a good example answers the first question; repeatedly retrieving it cannot answer the other two.

Set the population and the claim before choosing the aggregation route. Specify membership, time or edition rule, permissions, and the unit about which the answer will speak. Eight current submissions, twelve stored files and eight submitting organizations can be three different populations. Decide whether earlier editions are historical evidence to compare or superseded copies to exclude from the current count. Distinguish documents that repeat one underlying report from independently produced evidence. For themes, state what “main” means in the receiving use: commonly recorded, explanatory of a contrast, or consequential to the decision. Frequency alone need not determine importance.

Obtain an inventory from the source system or KCAE.INDEX’s enumerator. Keep each included unit’s source identity and selected edition, its assigned reading portion, and whether the relevant content was read, excluded by a stated rule, inaccessible or still unexamined. If only a provider’s selected hits are available, that is the available set; do not call it the complete collection. A broad question can legitimately produce an exploratory overview of selected material, provided the answer keeps that scope.

Choose a route that can cover the required material at an affordable cost. For a small dossier, a qualified reader can read every included unit, record its question-relevant claims with source returns, and compare them directly. This avoids index and summary preparation and is a serious choice for infrequent inquiries. If the material exceeds one reader’s working capacity, partition the enumerated units into bounded reading portions. Give each portion the same question and inclusion rule. Split a long document without losing its unit identity; preserve the necessary cross-boundary context and reconnect claims that span portions. This direct partitioned map/reduce route is available without an entity graph or precomputed summaries.

For a large, repeatedly queried collection, use prepared summaries or communities where their saved query effort earns their preparation and update burden. Choose the regions or hierarchy levels to inspect and resolve their membership back to source units. A parent and its child summaries can describe the same evidence; several graph communities can reach one underlying document. Selection across levels, including RAPTOR’s collapsed-tree profile, can find useful material without inspecting the entire population. A query-guided route such as LazyGraphRAG can allocate reading to promising communities, but the retained selection boundary still limits the answer. KCAE.Profiles:1.3 compares these constructions. Use independent original inspection to investigate consequential regions or distinctions the prepared view may omit.

Map source material to claims that can be combined. For each reading portion, obtain the proposed answer-bearing claim, its source units and exact passages, the relevant entity/time/condition, and any contradiction or unresolved interpretation. Preserve which words are the source’s position and which relation the reader inferred. One portion may supply several themes, a counterexample, an exception or no relevant claim; do not require one winner or one positive answer. A “no claim” screen remains revisable when the question or coding criterion changes.

Carry lineage through intermediate summaries: a claim refers to its contributing source-unit set, not just to the name of the summary that repeated it. When an operation cannot preserve that lineage, use its summary to locate original passages before relying on the aggregate. Keep unread or excluded portions visible alongside positive results. A polished partial answer with no account of what it left out is insufficient input for a claimed whole-corpus conclusion.

Reduce by meaning and source support, not by the number of summary votes. Align claims that concern the same object, time, condition and predicate. Merge genuine paraphrases while taking the union of their underlying source-unit sets. Do not merge contrary positions or a conditional exception into the majority wording. Group the resulting claims into themes that answer the receiving question, then explain both recurring relationships and consequential differences. If an initial theme does not explain a substantial contrast, refine it and reopen the source portions whose interpretation can change; merely relabelling the final paragraph leaves the earlier coding unchanged.

For a descriptive count, define the predicate and count each eligible population unit once for that predicate. A unit may belong to several themes; disclose that those categories overlap rather than forcing their percentages to total one hundred. Retain unknown or unread units in the account of the denominator. Do not divide a relevance-selected hit count by the whole collection and call it prevalence. A claim about a wider population needs a justified sampling/estimation method and its assumptions; otherwise return the observed corpus count or a qualitative overview. The number of documents repeating a claim also does not by itself establish its truth.

Protect consequential minority evidence before compressing the answer. Maintain the supported objections, contrary cases, important qualifications and unresolved conflicts that could change the receiving decision, even when they are rare or rank poorly. Ask the relevant domain reader which differences have that consequence when ordinary interpretation cannot settle it. Reserve enough final reading and answer space for those differences, or explicitly narrow the promised answer. If intermediate claims exceed the reduction window, reduce them in further bounded groups while preserving source sets and these outstanding differences; keep the originals obtainable for the final comparison. This controls the next reading operation, not semantic loss by fiat.

Return an answer whose scope survives the final wording. State the population and edition basis, the supported themes or counts, material counterevidence and the unexamined or inaccessible remainder. Link decisive generalizations through their intermediate claims to originals. Verify that words such as “all,” “most,” “typical,” “increasing” and “consensus” have the required comparison or counting basis. A minority warning can be important without being typical; silence on a question is not agreement with another submission.

Stop when the intended bounded answer has adequate inspected support, consequential conflicts have dispositions, and the remaining affordable inquiry would not change that receiving result enough to justify its cost. A thematic overview can stop with a named uncovered region when the recipient accepts that narrower use. A complete descriptive count must instead obtain the necessary unit judgements or report the unresolved count; a representative estimate needs its own sampling basis. When the hard budget ends first, return a partial synthesis and the particular next region or ambiguity that matters. Reading every unit establishes an inspection extent, not guaranteed recognition of every possible meaning. KCAE.ASSESS checks the proposed answer against this basis before KCAE.DELIVER carries it to the recipient.