Library / Systems Engineering Principles Framework
Jump to passage
In this reading

Link to current text

Published source confirmed at last check

Source changed 2026-10-02 23:06:08 UTC · snapshot created 2026-10-03 01:38:24 UTC · last check 2026-10-03 03:05:10 UTC

SYSE.34 - Change Software and Persistent Data with a Recovery Boundary

Normativity: Guidance within the stated software-and-data change use; examples are illustrative.

SYSE.34:1 - Problem frame

Use this pattern when a software change can alter stored information, its interpretation, or which old and new consumers can use it. Start with the values that must survive, the permitted writers and readers, and the point at which a previous executable would cease to understand the resulting state.

The first result is a bounded change-and-recovery procedure: compatible states, the selected write/synchronization mechanism, validation, the irreversible boundary, operating limits and the actual decision holder. If a required data fact, database-specific Method or permission is missing, return that precise gap.

Do not undertake a migration merely because an executable changes. An unchanged data contract may need only a compatibility check. Conversely, a schema that looks unchanged can conceal a changed meaning. This pattern does not provide a universal database algorithm or promise zero downtime.

SYSE.34:2 - Problem

“Add fields, copy data, switch readers” leaves important work unspecified. Old writers may continue changing the original representation after a copy. Two new authorities can disagree. A retried backfill may overwrite a legitimate later write with a cached earlier value.

Returning the old executable is not enough when the new software has removed information or written values the old software cannot interpret. Even a successful database restore can omit later business events or repeat external effects. Recovery needs a real state boundary, not only a rollback button.

SYSE.34:3 - Forces

ForcePractical tension
Continuing service and simple changeCoexistence reduces coordinated interruption but adds compatibility and synchronization work.
New meaning and reversibilityA richer representation can improve use while making an exact return impossible.
Bounded migration and concurrencySmall batches control load; concurrent writes still need an explicit ordering rule.
Technical recovery and acceptable lossA restore mechanism can work while its time or information loss remains unacceptable to the people relying on the service.

SYSE.34:4 - Solution

SYSE.34:4.1 - Recover meaning, consumers and authority

Describe the old and proposed values using the application’s meaning, including null, empty, ordering, precision and identity distinctions that matter. Define a correspondence and any inverse. Identify inputs that cannot be represented or carried back without loss; do not hide them behind a convenient conversion.

Find every relevant writer and reader, including jobs, imports, replicas, recovery tools and infrequent clients. Recover the actual database/storage edition, constraints, triggers, isolation and permission conditions that the procedure will use.

Obtain the required meaning and acceptance limits from their holders. The platform can support the mechanism without deciding which information may be lost or how long the business can operate without the service.

SYSE.34:4.2 - Choose a change strategy and its states

Compare a bounded coordinated interruption with compatible coexistence. A short interruption can be simpler and safer for a small service; coexistence is useful only when its compatibility can actually be maintained. For either strategy, qualify the recovery actions available from its intermediate states, including forward repair where it is needed.

For parallel change, add the needed representation without prematurely removing the old contract, qualify coexistence, migrate consumers, and contract only after the old use has been resolved. The historical Parallel Change Method supplies that strategy, not an engine-specific synchronization mechanism.

Identify what is valid at each intermediate state: which executable can read or write, what data transformation has occurred, and what return remains possible. A failed step must have a known stop or recovery action from its actual partial state.

SYSE.34:4.3 - Construct the write and backfill mechanism

Choose one authoritative representation during coexistence when that can satisfy the use, and derive the other views through a specified transactional or otherwise qualified mechanism. If independent writers are required, obtain an explicit conflict/order resolution Method; the instruction “write both” does not supply one.

Select a backfill rule that accounts for concurrent updates. Use the database’s qualified locking, current-row or version-check behavior so that an old observation cannot silently replace a later legitimate write. Bound work and resource effects, including locks, storage growth and effects on serving traffic.

Make retry behavior explicit. Determine which transaction effects and progress observations are committed together, which may be replayed, and which external effects cannot be repeated. After a lost acknowledgement, first recover the actual state; a migration log entry alone is not proof that every dependent effect occurred.

SYSE.34:4.4 - Validate coexistence and the actual recovery claim

Exercise ordinary and boundary values, old and new writers/readers, relevant concurrent histories, partial failure and replay. Verify information correspondence as well as successful execution. A job that finishes can still populate the wrong values.

Do not activate a consumer that relies on a new representation before its required population and continuing-write invariant are established. If some data is intentionally outside the scope, ensure that the consumer can identify and handle that boundary.

Test the proposed return under the state it will actually encounter. Returning an executable, restoring information and reconciling later external events are different operations. Obtain a database-specific restore Method and suitable backup/archive material where restoration is part of the promise; measure its achieved loss and time against the authorized limits.

SYSE.34:4.5 - Execute within the boundary and separate contraction

Carry out the qualified bounded increment with the actual permission and observation needed to stop. A timeout or inconsistent state suspends dependent reliance; it does not justify blindly restarting the whole procedure.

Retain the old contract while the selected return depends on it. Removing old fields, accepting non-invertible new values or changing the write authority is a later state change with its own consumer and recovery question. Use SYSE.29 to arrange migration or continued support for remaining users, decide what must happen to retained state, and fulfil or obtain permission to end support commitments that outlive the technical transition.

Return the achieved compatible state, applicable evidence and remaining limits to deployment and release work. An engineered procedure is not evidence that it has already run, and a successful migration is not general release permission.

SYSE.34:5 - Archetypal Grounding

SYSE.34:5.1 - Define a reversible address correspondence

In a constructed ParcelWorks example, one non-partitioned PostgreSQL 18 table has an immutable row ID and a non-null Unicode text value D, stored as delivery_address. Old applications read and replace that whole value. The new form edits the first display line L separately from the remaining display text T. It does not infer postal structure.

Split D at its first line-feed character, LF. If no LF exists, set L = D and T = null. Otherwise L is the text before that LF and T is all text after it, including any further line feeds. Reconstruct D as L when T is null, or L + LF + T otherwise. L must contain no LF. Preserve every character; do not normalize whitespace.

Old DNew LNew TReconstructed meaning
12 Oak St12 Oak StnullNo line separator was present.
12 Oak St followed by LF12 Oak Stempty textThe trailing separator is preserved.
12 Oak St, LF, North, LF, Depot12 Oak StNorth, LF, DepotRemaining lines stay in the tail.
Empty textempty textnullEmpty but non-null old text remains representable.

This correspondence makes old and new text mutually recoverable within the stated domain. Postal parsing, normalization or an additional field that cannot be joined back is outside it and reopens the recovery decision.

SYSE.34:5.2 - Keep one write authority during coexistence

The selected authority remains D. Old applications write D directly. The new application’s adapter validates L/T, joins them into D and writes D; it does not independently write derived columns.

A row-level BEFORE INSERT OR UPDATE trigger computes L/T from NEW.delivery_address and returns the modified row. The exact PostgreSQL 18 trigger behavior supplies same-transaction execution and returned-row semantics. This construction excludes other address-writing triggers, replicated writers that bypass this trigger and external side effects of these row updates. If those premises are not established, this mechanism is not yet qualified.

The authorized operator briefly fences new address writes and drains in-flight writers while installing the schema and trigger. Ordinary application roles must not bypass the mechanism or independently write the derived columns. The PostgreSQL 18 privilege rules require inspection of table grants, inherited rights, ownership and superuser access; revoking a column privilege does not cancel a broad table grant. Resume old writers only after the actual configuration satisfies those conditions.

SYSE.34:5.3 - Backfill current rows rather than cached values

Backfill operates in bounded Read Committed transactions over stable ID ranges. For each named ID, it performs the equivalent of:

UPDATE address
SET delivery_address = delivery_address
WHERE id = the_named_id;

Here the_named_id is a bound value for the immutable row identity, not literal executable SQL. The trigger derives from the row being updated. The procedure never writes a D/L/T tuple cached by an earlier SELECT.

Under PostgreSQL 18 Read Committed, an updater waiting on another updater proceeds against the qualifying updated row after that transaction commits. The immutable ID keeps this case’s selection stable. This is not a general guarantee for arbitrary multi-row predicates.

Advance the batch position only after commit. An aborted transaction retains neither its row changes nor an advanced position. If commit succeeded but its acknowledgement was lost, replay the range using current values. The case’s exclusion of other update side effects is necessary for that replay to remain safe.

Constructed historyResult
Backfill derives A, then an old writer commits B.The trigger derives B with that write; both representations describe B.
Old writer B holds the row while backfill waits.Backfill acts on the committed current B and does not restore A.
Backfill commits A, acknowledgement is lost, old writer commits B, then the range is replayed.Replay derives B from current D.
New writer submits L = PO Box 7 and T = North Depot.The adapter writes their exact joined D; the trigger reconstructs the same pair.
A backfill transaction fails before commit.Its derived-field changes are rolled back and the range remains unfinished.

Enable L/T-dependent readers only after the scoped backfill and a correspondence check establish their required population. Continuing writes must preserve the same invariant. The table contains constructed histories, not a report of executed PostgreSQL qualification; deployment use requires exercising the actual schema, roles, concurrency and recovery arrangement.

SYSE.34:5.4 - State the return and its end

This release retains D, the trigger and both reader contracts. Subject to the application’s other compatibility conditions, the old executable can return while preserving legitimate address writes made through the new form. It is not necessary to undo those writes.

Contraction is separate. Before removing the old contract, fence and drain all old writers and compatibility adapters, select the new write authority and validate its consumers. Once a lossy transformation or incompatible new write has occurred, an old binary alone cannot restore the prior usable state.

Suppose a later contraction removes a trailing LF and retains neither the original D nor another copy of the absent-versus-empty tail distinction. Both 12 Oak St (T = null) and 12 Oak St followed by LF (T = empty text) then become 12 Oak St: whether a separator existed is lost, and returning the old executable cannot reconstruct it. An exact-return promise therefore requires retaining D or that lossless distinction before this change; accepting its removal requires the actual decision holder to authorize a narrower promise. After the information is gone, exact restoration needs an independently retained, qualified recovery source and is unavailable without one. A qualified forward repair must meet the actually accepted result; it is not an inverse obtainable from the transformed value alone.

A distinct fallback uses PostgreSQL 18 point-in-time recovery. It needs a suitable base backup and continuous required WAL archive; a logical dump is not a substitute for that mechanism. Restore the whole cluster into an appropriately isolated recovery arrangement, inspect its state and evaluate achieved loss/time. Choosing an earlier recovery point does not reconcile all later business events.

If an archive gap prevents that fallback, its recovery claim stops. A qualified forward repair may still be possible, and an independent non-data-changing artifact test can continue.

What changes in practice is that the team can explain which state remains usable after each change, why concurrent writes are preserved, and where “return to the old version” ceases to mean recovery.

SYSE.34:6 - Bias-Annotation

Success on a static sample can hide concurrency, inherited privileges and rare representations. Teams can also prefer zero downtime without pricing the complexity of coexistence. Inspect real writers and adverse histories, and compare a bounded interruption honestly with a more elaborate online construction.

SYSE.34:7 - Conformance Checklist

  • Old/new meaning and information-loss boundaries are explicit.
  • All relevant writers/readers and engine-specific conditions are accounted for.
  • Coexistence has a concrete authority and synchronization mechanism.
  • Backfill cannot overwrite a later legitimate write with a stale observation.
  • Partial failure, lost acknowledgement and replay have qualified behavior.
  • A new reader starts only with the population and continuing invariant it needs.
  • Executable return, data restore and later-event reconciliation are separately qualified.
  • Contraction and permission to accept loss are not inferred from successful deployment.

SYSE.34:8 - Common Anti-Patterns and How to Avoid Them

MisuseRepair
Read old values, copy them later, then prefer any non-null new field.Account for intervening writes and establish a maintained correspondence before new-reader reliance.
Make both old and new fields authoritative.Choose one authority or supply a real conflict-resolution Method.
Repeat the entire migration after a timeout.Recover committed and partial effects, then apply the qualified replay rule.
Reinstall an old binary after irreversible data loss.Use a qualified forward repair or information recovery with accepted limits.

SYSE.34:9 - Consequences

The pattern makes migration and recovery claims narrower but usable. Compatible coexistence can reduce interruption while adding temporary schema, adapter and maintenance cost. Some changes remain genuinely irreversible or require a planned outage; exposing that fact is more useful than promising a rollback that cannot preserve the information.

SYSE.34:10 - Rationale

Compatibility is a relation between actual representations and consumers over time. A transition must maintain that relation under the writers that remain active. Recovery is therefore a property of the resulting state and permitted operations, not merely the availability of an earlier executable.

SYSE.34:11 - SoTA-Echoing

For “How can application and data changes proceed without an unexamined all-at-once switch?”, adapt DORA Database change management and the historical 2014 Parallel Change strategy. Versioned migration and compatible intermediate states are useful; a synchronized outage remains a serious alternative when coexistence costs more than it gains.

Reject treating “write both” or “prefer the new non-null field” as complete concurrency reasoning. Sections 5.2–5.3 use the current PostgreSQL 18 trigger, privilege and isolation contracts to construct one bounded alternative. The cost is engine-specific qualification and a single-authority compatibility adapter, not a general bidirectional migration recipe.

Reopen the relying procedure when representation meaning, writer population, engine behavior, trigger/permission conditions, side effects or recovery limits change. The historical strategy remains useful, but its age or a current database manual cannot qualify an untested production arrangement.

SYSE.34:12 - Relations

SYSE.13 identifies software and data configurations. SYSE.27 and SYSE.29 address consumer compatibility and migration; SYSE.33 supplies controlled test conditions. SYSE.41 consumes the data-change result during deployment, SYSE.35 respects it during exposure recovery, and SYSE.38 uses it when an incident has partial data effects. SYSE.14 retains the release decision.

SYSE.34:End

Referenced in the corpus

29 literal mentions in other sections. Read their context to establish the relation.