SYSE.45:5 - Archetypal Grounding
SYSE.45:5.1 - A converter, a learned interpretation and two failed withdrawals
A service accepts an interval in seconds. Target identity, revision and actual effect must be obtained through the current interface. For already structured input duration_ms=12000, a deterministic conversion returns seconds=12. Under the supplied conditions that small controller is adequate and cheaper to construct and qualify than an adapter. Stable recurrence alone supplies no reason to replace it.
Now change the target to interpreting recurring, varied requests such as “take three readings each minute” and “take one reading every twenty seconds.” Both call for seconds=20 under the supplied meanings. The fixed converter still works after a supported structured duration exists, but it does not obtain that interpretation. Compare a phrase-rule parser, retained specialist/guidance support, and an adapter with the same converter, current interface and executor. In this constructed case, representative source examples establish that a small phrase rule leaves many intended forms unsupported; expanding and maintaining it has greater estimated whole-horizon burden than the offered bounded adapter trial, including target qualification, tests and failed-trial fallback. C.11 therefore selects that trial. If those grounds or training access are absent, use the supported parser/guidance way or obtain a worthwhile comparison premise.
Training histories include the request, applicable tool definition, current target/revision and the observations needed for a call. Targets vary phrases, rates, durations and identifiers. Missing-unit or ambiguous requests target clarification. The learner fits the intended interpretation/call, while the executor retains permission, actual-state binding and result checks. Optional semantic preparation and trajectory warm-up address different failures; any later reward must use qualified tool/argument meaning and protected results.
Two held-out failures discriminate what was lost when guidance was reduced. Inspect the candidate’s proposed call before execution; the retained binding checks reject an unsupported call:
| Input and condition | Candidate output and observed defect | Smallest supported repair |
|---|---|---|
duration_ms=12000; the applicable definition still says the argument is seconds, and current target/revision are present | seconds=12000 instead of 12. The stable conversion is wrong despite sufficient inputs. | Retain the small converter or restore guidance; if learning that mapping remains worthwhile, repair its targets and test new values. Another current-state lookup cannot supply the missing transformation. |
The provider changed to a millisecond argument, but the new definition was omitted from the input; the old call uses seconds=12 | The call no longer matches the current interface. Restoring the fresh definition, with the same weights, yields the supported duration_ms=12000 call. | Retain the required interface observation and repair its retrieval/input path. Training on old documentation cannot supply an unobserved future contract. |
Guidance reduction therefore tests a particular stable contribution, not all support. Untouched cases also include ambiguity, unsupported units and restraint after a lost reply. A delayed comparison must record intervening adapter, provider, controller and memory changes before attributing retained benefit.
SYSE.45:5.2 - Construct a corrected continuation from a recorded history
A recorded history H contains the permitted outcome-lookup operation, the original attempt A17, an acknowledgement followed by a lost reply, and the at-most-one-effect requirement. The recorded next action was an unsafe repeat of the mutation. An experience-informed teacher, using that failure and other qualified episodes, proposes lookup_outcome(attempt=A17).
The engineer checks that H itself supplies the attempt identity, supported lookup and unresolved effect that warrant this correction. The student receives H without the teacher’s extra experience; its action target is the lookup, not the old repeated mutation or the following observation. The old next observation remains evidence about the old action. To establish what the correction does, exercise that lookup in a new qualified test occurrence or use applicable environment evidence; relabelling the old observation would fabricate its consequence.
The training example fits only the corrected agent action. If H had lost A17 or the lookup contract, restore that input or target obtaining the missing contribution instead. A final success reward alone would not identify whether safe recovery or a lucky duplicate caused the result; the supplied intermediate attempt/effect state localizes the unsafe replay decision. A contrasting acknowledgement-without-effect case tests the feedback rule.
For a preference branch, the service owner may compare two factually supported reports and prefer one that exposes unresolved status before optional explanation. Record that owner, task and report pair with the judgement. Both reports must still meet the effect and factual requirements before the pair can support that preference target.
Untouched trials compare the trained policy with the baseline on new attempt identities, known and unknown outcomes, a missing lookup operation and earlier useful no-replay behavior. An episode-local update declares when the adapter resets; a durable update tests that older behavior after the new target is learned. The result is a candidate and bounded evidence, not an inferred successful deployment.