Library / Systems Engineering Principles Framework
Jump to passage
In this reading

Link to current text

Published source confirmed at last check

Source changed 2026-10-03 05:29:54 UTC · snapshot created 2026-10-03 05:30:57 UTC · last check 2026-10-03 06:25:20 UTC

SYSE.36:4 - Solution

SYSE.36:4.1 - Define the subject and the good task result

Name the service, user population, supported task class and relevant configuration. Recover the task’s acceptance meaning from the responsible product or practice owner. A platform team can implement observation without acquiring authority to redefine application behavior.

Define an eligible attempt, its beginning and end, and the condition for a good result. Include timing, correctness, freshness or other properties only as the task needs them. Separate unsupported or invalid requests under a stated rule; do not recategorize a valid but failed attempt merely to improve the denominator.

Distinguish a user attempt from internal retries or backend jobs. Several internal operations can serve one attempt, and a failed attempt may produce no backend job. Decide how cancellation, abandonment and incomplete observation affect the stated question rather than excluding them silently.

SYSE.36:4.2 - Separate measurement meaning from implementation

Write the service-level indicator, SLI, specification in terms of the user outcome before choosing a counter or query. Then identify how the available instrumentation could observe that specification and where it would miss or distort it.

For an event-ratio indicator, define both good events and eligible events over the same population and interval. For a time-based indicator, define the evaluated periods and their conditions. Do not combine request counts, user minutes and task outcomes in one ratio without a justified measurement Method.

If the important semantic result is not observable, return that gap. An easier proxy may still be reported under its own narrower meaning, but it cannot silently replace the required result. A dependency-health result provides evidence about the conditions checked on that dependency; an HTTP success status shows that the server reported success.

SYSE.36:4.3 - Construct and exercise the observation path

Place observation where material entry, execution and result failures can be detected. Connect the relevant metrics, logs or traces to the right subject and configuration using sufficient correlation. The current OpenTelemetry signals model supplies interoperable telemetry categories; it does not choose the user-success predicate.

Exercise a known good task, a meaningful failure, a failure before the backend and a loss of observation. Check delay, duplication, sampling and aggregation effects on the definition. Sampled diagnostic traces do not by themselves establish a complete event denominator.

Keep missing data visible. If eligible attempts are known but their results are not, retain that distinction. If the entry population itself is unobserved, even the denominator may be unknown. Apply an agreed conservative policy where needed, but distinguish that policy disposition from an observed service failure.

Use the authorized data-handling arrangement. Collect only what the question requires, and avoid exposing secrets or unrestricted personal task content. A correctness check may run within an authorized application boundary and emit a bounded result; exporting all input values is not inherently required.

SYSE.36:4.4 - Agree an objective and its use

Use observed task difficulty, user needs, failure consequences and feasible provision to propose a target and interval. Current performance can inform a starting proposal but does not determine what users should accept. Avoid importing another service’s percentage or assuming that a more demanding number is always a better choice.

Agree the target with the people depending on and providing the service, including the holder who can make the relevant trade-offs. Define the error budget in the same event/time basis and the actions that follow material consumption or exhaustion. Those actions need actual authority and capacity; a dashboard configuration does not create either.

A policy can prioritize repair, restrict a class of changes or require a decision before further risk. Choose its scope for the task and consequences. The budget is not permission to cause any kind of harm up to a numerical allowance, and it does not cancel independent constraints.

If agreement, meaningful timing or adequate observation is missing, return a proposal and its unresolved condition rather than claiming an operative SLO.

SYSE.36:4.5 - Interpret and improve without changing the subject

Report the observed value with its subject, population, interval and coverage limits. Distinguish measured reliability, proposed or agreed target, delivery speed, user satisfaction and improvement value.

Send platform-task evidence to the corresponding provider/practitioner decision, including SYSE.25 where a platform improvement is being chosen. Send application evidence to its application and exposure owners. SYSE.35 and SYSE.37 consume only the objective and observations that match their own named subject and population.

Compare reported reliability with actual user difficulty. Revise the definition or implementation when it misses consequential failures; retain enough continuity to explain why old and new values differ. Do not improve the apparent history by quietly changing who or what is counted.