KCAE.SOURCE:4 - Solution
KCAE.SOURCE:4.1 - Preserve the source and describe its provenance
Obtain an immutable source edition or make a permitted snapshot before extracting it. Preserve publisher/collection, document identity, edition, effective interval when relevant, access class and original format. Use a content digest to detect byte changes; retain semantic identity separately so that identical text in different documents is not accidentally merged. A corpus snapshot fixes a set of source editions, not merely the name of a branch that can move.
For a mutable collection, first capture a coherent set through a source-system snapshot or an approved freeze/copy procedure. A manifest computed while files are changing can describe an impossible combination. If the source system cannot provide a coherent snapshot, state the consistency limit and narrow the usable operation; do not label a best-effort crawl as an atomic edition. KCAE.CHANGE develops live and historical reading contracts.
KCAE.SOURCE:4.2 - Extract the relationships needed for reading
Choose the extractor by the actual carrier. Markdown and HTML can expose headings, lists and links. Born-digital PDFs can still scramble columns or reading order. Scanned documents need OCR and inspection of decisive glyphs, tables and notes. Audio/video transcripts need time offsets and returns to the original segment when tone, a demonstration or a visual relation matters. Machine-readable data needs schema, units, null meaning and provenance. Store the original alongside the extraction while permission permits it.
Define a small extraction qualification set from the source’s likely failures: an exception containing negation, a multi-column page, a repeated heading, a table with merged cells, a footnote and a cross-document reference. Compare rendered source and extraction for the use-changing relationships, not only character count. Random sampling can detect wider drift, but a sample does not certify every uninspected decisive passage. Mark unqualified regions and allow direct visual or specialist reading where needed.
For a table, retain its title, row identity, column identity, units, span relations and notes. An indexed cell can be small, but its read operation expands to these relations. For example, the extracted CedarBench cell “No” is unusable. The qualified reading says “Destination protocol v1 / automatic resend / No; reconcile receipts manually; note a: receiver lacks duplicate suppression.” The source image or native table remains reachable. If OCR cannot distinguish v1 from v2, return an uncertain extraction and inspect the original before applying the condition.
Distinguish faithful extraction from an authored interpretation. A generated explanatory prefix can make an index easier to search; it remains a derived claim, with the original text separately available. Translation has the same discipline: keep source language and edition, preserve important negations, numerical units and scope, and compare the consequential reading with a competent translator or the source owner when reliance requires it. An old translation can be a search aid or historical artifact without being the current operational authority.
KCAE.SOURCE:4.3 - Separate finding units from reading units
Segment by intelligible source structure before imposing a token ceiling. Preserve parent headings, paragraph/list membership, table and figure ownership, explicit references and adjacent conditions. Split an oversized section into bounded retrieval units while leaving its reading unit intact. Overlap can reduce boundary losses, but blind overlap duplicates text and still does not recover a distant definition.
Give each unit a document/edition identity, a local locator, an exact source span where available and an expansion relation. Keep retrieval text, structural metadata and generated context in distinguishable fields. Record which parents or definitions a generated context used. That makes a later parent change capable of invalidating an apparently unchanged child representation.
Choose unit size using the retrieval question and the operation that consumes the hit. A broad research theme may benefit from section summaries; a protocol exception may need paragraph or cell retrieval. Keep an independent raw/full-body path so that summary selection is not the only way to reach the source. Test a question whose deciding text lies beyond the initial unit; the opening operation should recover it without requiring the user to guess its exact location.
KCAE.SOURCE:4.4 - Implement and test the source return
A practical exact reader accepts a source edition and a locator, then returns the requested material, its structural envelope, continuation information and the edition actually served. A pagination cursor should bind the edition and requested unit. Expose an end marker or explicit remaining extent so that truncated output cannot masquerade as complete reading. Do not rely on a line number without its edition.
When a heading moves, a semantic document/section identity can help resolve the new location, but exact historical reading still selects the old edition. When a section splits or merges, preserve explicit predecessor/successor relations and inspect them before treating them as equivalent. A fuzzy match may suggest a new location; it cannot silently establish identity or preserve an earlier judgement.
Verify build, persist, reload, find and read on a fixture that includes additions, deletions and shifted headings. Check that the returned text corresponds to the selected source, that a deleted locator yields an explicit result and that a child unit expands through the correct note. Preserve the qualification result with the extraction/version where the installation needs it. KCAE.EVAL distinguishes this mechanical correctness from useful question answering.