Library / Systems Engineering Principles Framework
Jump to passage
In this reading

Link to current text

Published source confirmed at last check

Source changed 2026-10-03 05:29:54 UTC · snapshot created 2026-10-03 05:30:57 UTC · last check 2026-10-03 05:45:20 UTC

SYSE.46:4 - Solution

SYSE.46:4.1 - Select the claim and comparison

State the intended task family, actual result, protected effects, available supports and proposed reliance. Recover the actual performing arrangement. For a technical agent this includes model/settings, instructions, next-input assembly, memory, tools/editions, executor and environment. For a person it includes applicable prior performance, instructions, available aids, working material and conditions such as task order or fatigue. Identify what changed, what can be held fixed and what must instead be recorded.

Choose a baseline that represents a real alternative, including a direct deterministic operation or human-supported arrangement when appropriate. Compare the same receiving result and consequential conditions. Separate result quality, latency, cost, necessary access, restraint and side effects; combine them only when the decision supplies a justified trade-off.

Choose observations from failures that could change the reliance. A model-output test cannot establish a state-changing result. Obtain the domain’s acceptance predicates and the evidence that connects the measured signal to them.

SYSE.46:4.2 - Construct representative contrasts

Select tasks that admit useful progress and cases that require a bounded stop. Vary the location or value of a needed contribution where that distinction can reveal a failure: an adequate current premise, a genuinely missing fact, stale memory, a changed interface, unavailable support or an ambiguous effect.

Use SYSE.49 if the required task/feedback conditions must be built. Qualify fixture fidelity and reset through the appropriate domain and SYSE.33. Keep final cases and their expected outcomes outside adaptive construction or candidate selection. Synthetic consistency alone does not qualify a real-world consequence.

When the question is whether the agent chooses support well, construct matched cases with a sufficient supplied premise, a decisive fact available only through the selected support, and offered material that is irrelevant or misleading. Add verification-required and presently unanswerable cases where the receiving use needs them. Keep result, correctness basis, relevant limits and performing arrangement comparable.

For the optional contribution being tested, compare three regimes:

  • No optional support: the selected contribution is unavailable; other means and protected conditions stay as declared.
  • Supplied support: the protocol fixes which supported means or operation to use and makes it available. The agent still binds inputs, performs it, interprets and uses the actual return. This intervention supplies no oracle answer or assumed upper bound.
  • Agent-selected support: the agent chooses whether and how to obtain the contribution under the same result and protection conditions.

Record the exact access and guidance intervention. When the task requires an external world change, retain its executor, current-state observations, permission and protected controls in every regime. Removing the actuator would defeat the task, not isolate this help decision. A withheld indispensable fact likewise changes access rather than proving an inability to reason. Preserve different valid trajectories that satisfy the result.

For people, use matched tasks and an appropriate order or allocation across participants; record learning, fatigue and carryover that can change the comparison. Repeating the same question after showing its answer would not isolate support choice. HCD.12/.13 supplies any separate claim about acquired unaided capability, transfer or retention. An ordinary tool use with adequate existing evidence needs no new experiment.

SYSE.46:4.3 - Observe attempts and localize failures

Run the configured task under its stated permissions and stopping rules. Retain inputs, configuration, supplied support, selected action/call, actual return, next working material, continuation, effects and result use to the extent needed to reconstruct the comparison. These observable records need no hidden chain of reasoning. A missing observation is a measurement gap, not an unsuccessful or successful event inferred from silence.

Choose repetition from the variability and consequence of the intended reliance. Distinguish “succeeded at least once in several attempts” from “succeeded on every required repetition”; report the unit and attempt policy. Do not count a successful retry without its failed attempts and effects.

Use the support contrast and trace together. Supplied-support success with agent-selected failure can mean a skipped necessary call, malformed arguments, a return lost from the next input, or a good return ignored by the procedure. These need different repairs. Failure even with supplied support leaves support quality, actual availability, binding and use open before an intrinsic-capability conclusion.

Read the other failures at their actual location: missing write, retrieval miss, wrong target, effect uncertainty, bad integration or an unsupported environment judgement. A trace can narrow these questions without revealing a unique hidden cause. Return unresolved causation as such.

SYSE.46:4.4 - Test persistence and changed conditions when relied on

When the claim includes lasting improvement or response to change, compare initial, post-change, delayed and shifted use. Select the interval and shift from that reliance: normal restart, memory expiry, source revision, tool withdrawal or altered workload. There is no universal waiting period.

At each observation retain the task/result, configuration identity, support actually available, evaluator basis and intervening changes. Test delayed persistence under stable comparison conditions separately from changed support. An unidentifiable model-provider update may leave current system performance observable while making a component-specific persistence claim unresolved.

Keep warranted restraint, unnecessary and missed necessary access, actual obtained-and-used result, verification, sufficient-result stopping and total burden visible. A changed authoritative fact should be sought and used even after repetitive retrieval becomes cheaper. A previously useful operator should be rejected when its applicability is defeated. Removing a necessary aid changes the tested performing arrangement.

SYSE.46:4.5 - Return evidence at its qualified reach

State what was compared, what the observations support, the limits and the next condition that would reopen reliance. Preserve useful partial results when another claim is unqualified. A few favorable cases can support a bounded trial decision without proving a general capability or absence of rare failure.

Return invocation/effect failure to SYSE.42, memory failure to SYSE.43, division/join failure to SYSE.44, unsupported policy change to SYSE.45, controller failure to SYSE.47, operator failure to SYSE.48 and misleading task/feedback construction to SYSE.49. Return an assistance-selection defect to SYSE.50, distinguishing it from a lost input or ignored return. Return failed adaptive allocation to SYSE.51: compare entire trajectories, including evaluator cost, unsuccessful branches, completion reserves and warranted early stopping, against fixed/manual allocation. Compare the acted policy’s unsupported continuation, unnecessary interruption and total task burden separately from its signal calibration. Return consequential selection, summary or tool-view loss to SYSE.52, inspecting the actual request and its raw-evidence recovery path. Existing interface, domain and measurement owners retain their results. Stop when the selected qualification question is answered.