Library / Systems Engineering Principles Framework
Jump to passage
In this reading

Link to current text

Published source confirmed at last check

Source changed 2026-10-03 02:22:15 UTC · snapshot created 2026-10-03 03:38:22 UTC · last check 2026-10-03 05:05:10 UTC

SYSE.45:4.3 - Select and execute the learning mechanism

First choose the experience transformation that can supply the intended policy target. The learning algorithm then consumes that target; the names supervised learning, reinforcement learning and distillation do not construct it.

Available experience and intended changeConstruct the signalFurther-use question and return
A qualified action sequence demonstrates the needed behaviorPair each retained decision history with the agent action it warrants. Keep tool observations as inputs, and fit only the intended agent outputs. Remove incidental target identifiers or values by varying them in applicable examples.Does the policy choose the action on a new applicable input and abstain on an unsupported one? An unqualified successful transcript returns to result/trajectory assessment.
A recorded decision is wrong or unnecessarily costlyAn experience-informed teacher proposes a corrected next action at that history. Qualify it against the task, interface and facts the student will actually have. Train the student on that history/action pair without the teacher’s extra experience.Does the correction remain warranted without hidden teacher facts? Obtain a needed fact or retain support if it does not. The old next observation follows the old action, not the proposed correction.
A search or deliberation procedure finds useful candidatesPreserve selected candidate comparisons, evaluation grounds and backtracking decisions as targets where their signals are available to the student. Distillation can transfer how candidates are generated or compared.Test unseen alternatives and defeated evaluator premises. Imitating a planner’s trace does not transfer its proof or guarantee.
Actual attempts supply outcome or process feedbackBind feedback to the required result and protected conditions. When a final reward leaves the responsible decision unclear, use a qualified intermediate state or a discriminating action contrast to localize the target. Keep an uncertain credit assignment at that evidential strength.Test the actual result and the relevant intermediate behavior separately. Reward growth with duplicated effects or lost required access fails.
Qualified comparative judgements express a preferenceName whose preference, the task/context, the compared action or answer pair and the grounds for its ordering. Give those pairs to a supported preference-training Method and implementation; return a missing judgement or implementation.Preferred behavior still needs independent factual, permission and protected-result grounds. A rater’s preference supplies no missing world effect.

Match the update to that signal. Supervised learning fits qualified actions or corrected continuations; reinforcement learning uses attempted actions and qualified reward; distillation transfers the selected teacher or search behavior. CMP.7 supplies learner construction and the distinction between obtaining data, fit and further use. The selected implementation supplies the exact update algorithm, tokenizer and trainable parameters. Bind them to these data transformations and the later input format. Keep the configuration that actually produced each comparison.

Construct tool knowledge and use separately when that is the gap. Tool-Internalized Reasoning separates learning a tool’s documented semantics, supervised preparation on tool-use trajectories, and optimization of subsequent tool reasoning. For an interval setter, the semantic target includes what operation and unit its arguments denote; a trajectory target then puts the call at the right decision with the right observations. A later reward can distinguish the selected tool and argument choice only at its qualified meaning. The source’s special tool vocabulary, document/token mapping and reward are implementation choices, not prerequisites for every adapter. A tool/argument proxy does not establish the resulting service effect. A plain converter or retained description may supply the needed operation more cheaply.

When reducing repeated procedural guidance, vary only the contribution intended to become unnecessary. Keep execution and current observations available. Test separately a stable mapping presented with all its needed inputs and a decision whose new interface fact was withheld. Their different repairs are worked below. Restore needed support rather than rewarding unsupported fluency.

Bound repeated and episode-local updates. Identify the state or parameters being changed, the qualified signal that permits an update, the experience retained, and the lifetime/reset rule. An episode-local adapter may be discarded at the end; durable shared parameters need a recoverable baseline and regression tests across later tasks. A runtime hidden state or learned memory token is not automatically either shared-parameter learning or an external text record. SYSE.52 prepares the exposed working representation, SYSE.43 maintains separately persisted episodes, and this Method governs selected parameter learning. Specialized backbone or latent-memory-module construction remains with its direct implementation.

Choose update extent and stop from applicable held-out behavior and the available means. Test older useful restraint and correct actions alongside the new target. Repeated self-generated failures are observations of failure; frequency does not turn them into positive labels. Reset, revert or return the unqualified signal when the update destroys earlier required behavior.