Part VIII - Engineering Agent Work and Support
SYSE.42 - Use a Selected Tool and Apply Its Result
Type: Method pattern Status: Candidate
SYSE.42:1 - Problem frame
Use this when a person or technical agent has selected an external aid and the task depends on applying it to the intended input and using the actual result. A calculator can execute a mistyped expression correctly; software can act on the wrong target; a useful return can be ignored.
Start with the supported action, intended expression or target, and the result the receiving work needs. Bind the input, use the selected means under its actual semantics, check the obtained contribution and put it to use. The first useful result is that used contribution or the precise condition preventing it.
Choose among complete available ways through C.38/C.11 when that choice is unsettled. A.15.7 steers the current action; C.24 plans a selected technical action’s calls when needed. An already sufficient result needs no fresh call. This pattern applies an available tool or aid; SYSE.26 designs a supported service interaction, while SYSE.48 constructs a missing reusable aid.
SYSE.42:2 - Problem
A selected aid does not by itself complete the work. The performer must bind the intended input, apply the aid correctly and relate the result to the task. A calculator executes its implemented operation; paper preserves the marks while the person calculates. In an LLM realization, the model proposes symbols and the executing software acts under its interface and permissions. Correct execution on the wrong input still fails the receiving use.
SYSE.42:3 - Forces
- Automatic argument filling saves effort but can replace a missing fact with a plausible guess.
- A tool may acknowledge receipt before producing an effect. Retrying a lost reply can duplicate that effect.
- Returned text can mix useful data with instructions intended for another setting.
- Fresh evidence can invalidate the action being attempted, making another invocation less useful than reconsideration.
- Detailed checking costs time. Its extent depends on what the receiving task relies on.
SYSE.42:4 - Solution
SYSE.42:4.1 - Bind the selected means to the supported action
Recover what the task needs, which expression or target is meant, what may be done and what result will be used. Inspect the available means and the behavior that matters: entered expression, units, identifiers, valid range, current state and any effect/recovery conditions. For paper-supported arithmetic, establish the readable layout and supplied digit operations; the person performs them.
For every consequential argument, identify its source. Copy a known identifier from the current task or authoritative lookup; derive a quantity only through an applicable conversion. If the proposal omits a value or gives two plausible targets, obtain the smallest discriminating fact. A default is usable only when the contract makes it the intended value in these conditions.
Check the proposed use against the means actually available. For a calculator, inspect the entered expression before taking its display as the requested product. For software, check the actual tool set and current permissions; a generated tool name or explanation establishes neither. Place reliably enforced controls in the executor under SYSE.28. Use an explanation to locate a mismatch, while retaining the actual check.
SYSE.42:4.2 - Execute within the interface’s effect boundary
Check the expression or layout against the required operation. For a structured call, validate syntax and then its state-dependent meaning: a well-formed request can concern the wrong configuration or a superseded target. Use the current-state precondition when intervening changes matter. Perform the supported operation with the selected means.
For a state-changing software call, retain the attempt identity when recovery needs it. Distinguish rejection before effect, acknowledged progress, confirmed result and unresolved effect. Follow SYSE.26’s same-attempt recovery when an effect may have occurred. Replay is justified only by the interface’s actual retry semantics and observed state; a timeout alone does not supply that basis.
Keep the effort and waiting within the task’s remaining budget. SYSE.40 supplies combined resource and retry limits when several calls or services share them. On exhaustion, return the unfinished condition and recoverable state.
SYSE.42:4.3 - Interpret the return and test its receiving use
Check that the response concerns the requested target, operation, edition and relevant interval. Parse quantities with their units, distinguish observations from predictions, and retain partial or unknown results. A free-text success statement can be checked against a supported state observation when the claim concerns a changed target.
Treat source content as material to interpret under the current task. Embedded instructions in a retrieved page or tool response do not acquire authority to redirect the operation. Preserve provenance so that a later step can distinguish the task’s governing instruction from returned evidence.
Use the qualified return in the receiving work: update the outstanding premise, calculation, proposed continuation or user-facing result. A new observation that defeats the action’s basis returns to the domain decision or A.15.7. A technical failure returns to the interface/provider; a recurring failure to consume a good result returns to SYSE.47. Stop when the task has its sufficient result.
SYSE.42:4.4 - Qualify stronger reliance
For a consequential or repeated use, exercise valid progress alongside a wrong target, missing argument, changed state and ambiguous effect that could defeat the relied-on behavior. SYSE.46 tests the actual performing arrangement. The evidence concerns the actual interface and receiving result; correct entry or schema formation is only one part of that use.
Keep the result’s source, configuration and unresolved conditions to the extent needed by that reliance. This can be a few fields in an ordinary task record; a second invocation ledger is unnecessary when the executor already retains them.
SYSE.42:5 - Archetypal Grounding
A person uses a calculator and reports the product
Six lots contain 347 items each. A calculator has been chosen because it is already open and, under the supplied conditions, using it protects other intermediate values the person is holding. The person enters 347 × 6, inspects that expression, obtains 2082 and reports the total for the six lots. The device computes; the person binds, inspects and applies the result.
If the input is mistyped as 374 × 6, the functioning calculator returns 2244. The input check localizes the defect; correcting the expression restores 2082. Where the receiving reliance requires an arithmetic check, 350 × 6 − 3 × 6 supplies 2100 − 18. A paper alternative retains the aligned digits and carries while the person computes them. A lost carry/column relation returns to the working layout or retained record, not to a supposed calculation performed by paper.
For a search return, determine whether it is an actual computation, a recoverable cited claim or generated text before using it. A matching snippet and two interfaces to the same backend do not supply two independent checks. Stop when the required answer is used; later unaided ability is a separate development result.
A generated service call has an uncertain effect
A test service exposes a fictional operation set_interval(target, seconds, expected_revision, attempt) and a query for an attempt’s outcome. The engineer is permitted to change meter-2’s sampling interval to 10 seconds. Its current revision is 81.
The model proposes seconds=10000, having reused a millisecond value. The binding step compares the requested duration with the actual seconds field and constructs 10 from the supplied unit conversion. It obtains revision 81 from the current target rather than copying another meter’s cached revision.
The request with attempt A17 is accepted, but its reply is lost. The executor queries A17 under the supplied contract. That return identifies meter-2, revision 82 and interval 10 seconds. The next report uses those observed values and closes the change task; it does not issue another mutation.
If the attempt query instead reports an unknown outcome, the task remains unresolved. If a fresh read shows an independently changed revision before invocation, the expected-revision condition fails and the engineer reconsiders the change against that state. Neither case becomes success because the generated explanation is convincing.
For a request merely to explain a current configuration, an already applicable observation may suffice. Reissuing a state-changing call would add no needed result.
SYSE.42:6 - Bias-Annotation
Tool-friendly tasks can hide missing authorization, physical action or specialist interpretation. A simulated success can also make recovery appear easier than the real provider allows. Select the actual effect and observation conditions of the receiving work before extending the case.
SYSE.42:7 - Conformance Checklist
- The proposed invocation is bound to a supported action, actual target, current interface and sourced consequential arguments.
- Validation distinguishes shape from permission, applicability and state.
- A partial or unknown effect follows the contract’s recovery before replay.
- The return’s subject, units and evidence status support its receiving use.
- The task consumes the result or returns the precise missing condition.
- Stronger reliance includes the failure cases that could change it.
SYSE.42:8 - Common Anti-Patterns and How to Avoid Them
Schema success becomes task success. Inspect the target/result relation and the downstream use, not just parser acceptance.
Retry every exception. Recover the earlier attempt and its possible effects. Return unresolved state when the provider cannot establish replay conditions.
The response chooses the next authority. Interpret instructions found inside returned content as source material; use the current task and permitted action basis to decide what follows.
SYSE.42:9 - Consequences
The task gains an inspectable connection from a selected means and intended input to an obtained contribution and its use. Additional reads or checks can increase latency; retaining a sufficient existing result and checking only consequential arguments limits that burden. Unsupported intent and broken interfaces remain visible returns rather than fabricated completions.
SYSE.42:10 - Architectural Rationale
Execution binding sits between action selection and the receiving work. Keeping it explicit makes three repairs distinguishable: choose another action, repair the aid/interface, or correct input binding and result use. A good calculation on the wrong numbers and a correct return that the next step ignores need different repairs.
A deterministic direct call is preferable when its input mapping already suffices. Model participation is useful where interpretation is needed, while the contract remains the basis for actual effects.
SYSE.42:11 - SoTA-Echoing
The common operation connects input binding, supported application and receiving use. Kirsh, 2010 is a historical conceptual account of external representations in thinking; it helps distinguish a persistent layout from the person’s performance. The technical branch combines interface-grounded execution with result-use testing. Theory of Agent, v1, §§3–4, supplies the distinction between generating a call and using external interaction.
ToolSandbox, 2024 is a historical stateful evaluation anchor. Its contribution here is to challenge dependencies and state changes; this pattern adds the receiving task’s actual effect and recovery contract. A format-only test remains cheaper when only parsing is claimed, but cannot qualify a state-changing task. Reopen the binding when the tool contract, source authority or required result changes.
SYSE.42:12 - Relations
A.15.7 and C.24 supply action choice and fixed-action call planning. SYSE.26/.27 supply supported interaction and compatibility; SYSE.28 supplies qualified controls. SYSE.43 can supply relevant memory, while SYSE.47 constructs the procedure that invokes this operation and consumes its result. SYSE.46 tests the resulting configuration. A.15.9 and SYSE.9 obtain missing subject results without transferring responsibility for their qualified use.
SYSE.42:End
SYSE.43 - Maintain External Memory for Continuing Work
Type: Method pattern Status: Candidate
SYSE.43:1 - Problem frame
Use this when a person or technical agent resumes without a needed premise, or repeated work keeps reconstructing usable experience. A notebook can lose the place of a carry; a software archive can retain the answer while retrieval supplies the wrong episode. The next useful action is to recover what the continuation needs and locate where that relation disappeared.
The subject is the external recording and retrieval arrangement: what is retained, how it is organized and found, when it changes and how the receiving task uses it. A worksheet or notebook is external to biological memory; a software store is external to model parameters. Either may be part of an equipped performing whole. The first result is a memory-supported continuation with recoverable grounds, or a localized write, retrieval or use failure.
Start with the next action that lacks context. If its needed facts already fit reliably in the current working context, use them. Cross-task storage adds value only when later access justifies selection, maintenance and retrieval effort. Method-library discovery belongs to ME.10; constructing a new reusable operator belongs to SYSE.48.
SYSE.43:2 - Problem
Memory can preserve a sentence while losing the conditions that made it true. Retrieval can then increase confidence without supplying the fact needed now. Conversely, discarding all old context makes every interruption a new inquiry and destroys informative failures.
SYSE.43:3 - Forces
A compact summary is cheap to consume but can erase a decisive exception. Detailed episodes preserve evidence but compete for attention and storage. Current facts, preferences, attempted actions and procedural knowledge change under different conditions. Better retrieval cannot repair a fact that was never recorded, and a correct return is useless when the controller does not consume it.
SYSE.43:4 - Solution
SYSE.43:4.1 - Choose the receiving use and retained content
Name the continuation or recurring task that memory should support. Recover the facts or experience it would otherwise have to obtain again. Keep content that changes a later choice, result interpretation or recovery, rather than storing a transcript because it is available.
Distinguish authoritative observations, past episodes, tentative inferences, preferences and pointers to procedural material. Retain source, subject, relevant time or edition, and applicability where their loss could change use. A historical observation stays historical after the world changes. A remembered permission does not establish permission for another operation.
For an interrupted effect, preserve the attempt and unresolved state needed by SYSE.42/SYSE.26. For an episode, retain what was attempted, under which conditions, what was observed and why the result mattered. Avoid replacing an observed failure with the later model’s explanation of its cause.
SYSE.43:4.2 - Construct writing and retrieval together
Choose granularity from the query and action. A single measurement may need an exact record; a recurring recovery may need a short episode linked to its trace. Summarize only after checking that the summary preserves the receiving-use conditions, and keep a return to the source when compression could be consequential.
Build the retrieval request from the unresolved premise: target, needed contribution, applicable edition or interval, and the kind of evidence required. Use exact identifiers or filters when those conditions are known; use semantic retrieval where wording varies. Rank by applicability as well as similarity. Bound what the receiver must inspect and retain enough provenance to qualify use. A person’s page/tab organization and a software index implement different access paths. For a model request, fit the retrieved material to its input budget through SYSE.52.
Keep read/write scope aligned with the permitted user or task. Prevent a record from another subject or access scope from becoming evidence merely because its wording is similar. Preserve the known origin when presenting an external record to its receiver.
SYSE.43:4.3 - Reorganize experience, reconcile change and use the return
Several valid episodes may need a better organization even when none is false. Begin with a small exact record or index. Extract the action-changing features of each episode: task/configuration, input, action, effect status, actual observation and recovery. Propose a conditional summary, grouping or link only when it answers the receiving query better than that simple arrangement.
Test the proposed relation against its source episodes and a contrary episode. A shared word such as “timeout” does not establish the same recovery condition or cause. Preserve exact exceptional facts and links to raw evidence. Mark generated associations and explanations as candidate interpretations; do not rewrite an observed event to fit them.
Change the index, retrieval description or linked-note organization, then run the same receiving query. Check what is now returned and which continuation it supports. Keep the richer organization only when the gain warrants writing, retrieval, checking and maintenance burden. Retain a rare contrary episode when forgetting it changes a consequential action. An executable recovery operation belongs to SYSE.48; this step organizes the evidence used to select or perform it.
When sources disagree, first determine whether they concern the same subject, period and meaning. A newer observation can supersede an old current-state claim without invalidating the earlier event. An unresolved conflict remains visible; combining incompatible statements into one fluent summary does not resolve it.
Mark superseded material so ordinary retrieval does not present it as current. Correct an erroneous record at its source-linked claim, update dependent summaries or indexes, and retain enough history for the actual recovery need. Retire content when its permitted retention or useful applicability ends.
Have the receiver state what the retrieved contribution supplies. Incorporate it into the next calculation, action condition or supported answer. Obtain a fresh observation when current state is required. Return an unavailable premise through A.15.9 rather than inferring it from an old successful episode.
SYSE.43:4.4 - Locate the failed operation
Test a representative continuation with its expected relevant material known. Ask separately whether that material was written faithfully, whether the current retrieval returned it, and whether the task used it correctly. Supplying the needed record directly can help distinguish retrieval from utilization, provided that the comparison changes only that support.
SYSE.46 qualifies stronger reliance, including delayed and shifted use when relevant. Check an old plausible fact, a newly authoritative change, a no-match case and a distractor that could displace the needed contribution. Preserve necessary fresh access even when repeated memory calls become cheaper. A retrieval score or larger store cannot alone establish improved work.
SYSE.43:5 - Archetypal Grounding
Restore the relation between a carry and its column
A person is interrupted after calculating 7 × 6 = 42 in 347 × 6. The retained worksheet should record the units result 2 and a carry of 4 into the tens column, with the original expression. A copied note preserves “2, 4” but drops their positions. That is a recording loss. Recover the original sheet when available, restore 2 under units and 4 above tens, then continue 4 × 6 + 4 = 28 and 3 × 6 + 2 = 20 to report 2082. If the original is unavailable, recompute the bounded units operation; do not guess what the two digits meant.
If the complete sheet was retained but the current view cropped away its carry row, repair the view through SYSE.52. If the correct sheet is supplied and the person still skips adding the carry, the receiving operation needs attention; recopying the sheet does not supply that performance. Stop with a usable record and resumed calculation. Later unaided retention is a separate HCD question.
Reorganize still-valid timeout episodes
An agent’s small exact store contains three episodes under a common “timeout” keyword. In E1 the interface establishes rejection before acceptance. In E2 the request was accepted, its reply was lost, and querying the original attempt later confirmed an effect. In E3 acceptance is known but the supported query still leaves the effect unknown. The original traces remain valid.
A proposed summary “retry after timeout” fits neither E2 nor E3. The engineer extracts acceptance and effect status and builds two conditional index entries: established rejection before effect permits reconsidering a new call under the current preconditions; accepted or uncertain execution requires same-attempt outcome recovery before any replay. Both entries link to the original traces and their interface edition. E2 remains a useful contrary case even if most examples are E1.
The current interrupted attempt A19 was accepted. Retrieval by that known status now returns the recovery branch and E2/E3 evidence rather than the most common retry narrative. The controller requests A19’s outcome and uses that actual return through SYSE.42; a past successful lookup does not establish A19’s present effect. If the return remains unknown, so does this task. Dropping E3 to make the summary shorter would remove that stop and is rejected.
For three episodes, a two-entry conditional index can suffice. A larger corpus with differently worded recovery questions may justify candidate links and revised note descriptions, tested against the same queries and contrary episodes. Compare total maintenance and retrieval effort before adopting it. This changes the organization of valid experience, unlike correction of a false current-state summary below.
Preserve an earlier observation without treating it as current
An agent resumes a test-service change after interruption. Its memory contains “meter-2 used interval 10 seconds at revision 82” and a trace link. A later authoritative observation says revision 83 uses 2 seconds.
The task is to explain current sampling, so the query includes meter-2 and current configuration. The record at 82 remains evidence of the earlier change, but the answer consumes the observation at 83. A summary saying “meter-2’s interval is 10 seconds” would lose the time condition and is repaired. Its dependent retrieval entry is updated.
In a diagnostic task, the earlier 10-second episode may still matter: it explains when the behavior changed. The same stored record therefore remains useful without being promoted to a current fact.
Now suppose the correct revision-83 record is returned but the final report still says 10 seconds. Enlarging the index does not repair that failure. Inspect the actual next input through SYSE.52 and the step that consumes it through SYSE.47. Restore a field lost during assembly in the former; repair an ignored return in the latter. If revision 83 was never captured, repair writing or access instead.
SYSE.43:6 - Bias-Annotation
Conversational recall tests favor facts easy to express in text. Engineering continuation may depend on an unobserved effect, an exact configuration or a source the memory service cannot access. Let the receiving work determine what must remain external and fresh.
SYSE.43:7 - Conformance Checklist
- Retained content serves a named continuation and preserves consequential provenance and applicability.
- Retrieval uses the unresolved premise and separates relevance from currentness.
- Summaries retain their needed conditions and source return.
- Reorganization over valid episodes tests proposed links or abstractions against originals and contrary experience.
- Conflict, supersession and retirement change the affected records and direct projections; rare action-changing counterexamples remain recoverable.
- The receiver uses the returned contribution or identifies the missing one.
- Qualification separates writing, retrieval and utilization, and preserves necessary fresh observation.
SYSE.43:8 - Common Anti-Patterns and How to Avoid Them
Everything becomes durable memory. Select from later use and retirement conditions; irrelevant history can make retrieval worse.
Latest text wins. Establish the source’s authority, subject and interval before replacing a claim.
The record is correct, therefore memory works. Inspect what the receiver actually obtained and used. The failure may be in retrieval, context assembly or the next action.
SYSE.43:9 - Consequences
The system can resume with less reconstruction while retaining the difference between past experience and current evidence. The cost includes write selection, indexing, conflict handling and retirement. A smaller exact store or ordinary task record can be better than a general memory service.
SYSE.43:10 - Architectural Rationale
Writing, retrieval and use form distinct failure locations within one memory arrangement. Keeping them visible makes repair local: a missed write does not call for a larger prompt, and ignored evidence does not call for another embedding model. External storage changes available support. A retained worksheet is not evidence of acquired unaided human skill, and a software record is not a model-parameter update.
SYSE.43:11 - SoTA-Echoing
For continuing work, the selected line tests memory through its receiving use and keeps representation changes separate from changes to past observations. Kirsh, 2010 is a historical conceptual anchor for persistent external representations; it supports asking which positional relation the next human action uses. LongMemEval-V2 contributes premise-sensitive experience retrieval; AMemGym v1, §3, separates writing, reading and utilization under evolving conditions. Their bounded tasks do not establish correct engineering action.
A-Mem v11, §§3.1–3.4, supplies a historical construction alternative: retain interaction content/time, propose descriptive attributes and links to related notes, then revise affected descriptions as experience grows. Generated links remain proposed associations; its conversational evaluation does not qualify engineering action. A small conditional index can be sufficient.
Adapt these distinctions to actual configuration, provenance and fresh-state needs. Compare with a small explicit task record before adding a general memory system. Reopen the arrangement when source access, update behavior, task family or observed retrieval burden changes.
SYSE.43:12 - Relations
A.15.8 recovers the continuation state and probes consequential support loss. SYSE.42 supplies observed tool outcomes, SYSE.52 constructs the next input from retrieved material, SYSE.47 binds that input into execution, and SYSE.46 tests the configured use. ME.10 retains Method-material discovery; ME.15 retains semantic repertoire and edition distinctions. SYSE.45 changes parameters only when a separate learning intervention is selected.
SYSE.43:End
SYSE.44 - Divide and Recombine Work across Agents
Type: Method pattern Status: Stable
SYSE.44:1 - Problem frame
Use this when people or technical agents could usefully divide a needed result, but preparation, dependencies or joining errors may consume the gain. Start with the whole result and compare one-performer completion with a concrete division.
The subject is the working division, the material each participant receives and the recombination of their contributions. People may bring different source access or expertise; technical participants may use separate contexts of one model or different tools. Those arrangements need different preparation but share the requirement that their returned contributions fit the same whole. The first result is an integrated contribution that the receiver can use, or the dependency, conflict or evidence gap preventing it.
Keep dependent steps in order: dividing 347 × 6 at each digit introduces carry handovers into a calculation one person can complete. A proposed division must repay preparation, communication, duplicate work and integration. A.15.9/SYSE.9 supplies bounded requests; independently governed Systems retain SYSE.18’s governance question.
SYSE.44:2 - Problem
Splitting a brief or prompt can remove the common premise that made its parts meaningful. Even with the same supplied material, participants may interpret a quantity differently or intend different next contributions. Agents can then solve different versions of the problem, duplicate the same source error, or change shared state incompatibly. A final summarizer can conceal those differences instead of resolving them.
SYSE.44:3 - Forces
Separate contexts can protect attention and allow parallel inquiry, but every boundary needs a usable handover. Narrow context limits distraction while risking loss of a global condition. Diversity can expose mistakes, although different role names do not produce independent evidence. Integration must preserve justified disagreement without making every disagreement a reason to repeat all work.
SYSE.44:4 - Solution
SYSE.44:4.1 - Select contributions and the receiving result
Name the whole result, its protected conditions and the person or system that will integrate it. Recover which missing results could be obtained separately and which inputs they share. Distinguish a calculation, source interpretation, alternative proposal and independent check; their return conditions differ.
Compare one-performer completion, including one-context LLM work, with the proposed division on expected task quality and total effort. Include preparation, source access, repeated work, waiting and integration. Use a small representative comparison through SYSE.46 when the advantage is uncertain. More agents is a candidate arrangement, not a quality measure.
For the selected arrangement, state each contribution’s receiving use, required inputs, permissible actions, available sources, result form and stop. A.15.9 and SYSE.9 supply bounded requesting and reliance. A concise passage or a structured return is sufficient when the recipient can act from it.
SYSE.44:4.2 - Construct usable contexts and effect boundaries
Give each participant the common premise and source/configuration needed for the contribution, plus the relevant constraints of the whole. A human work brief and an LLM request must each preserve those meanings; their size and presentation need not be identical. Separate governing instructions from quoted source material. Preserve a route to the original when a summary could hide a decisive qualification.
When a message or work product leaves consequential uncertainty about how that material is understood or what the participant intends to do next, use the Reference’s collaborator application. Construct the alternatives that would change the explanation or division, then choose a small distinguishing return only when its value warrants the burden. Use the qualified result to clarify the needed meaning, revise the requested contribution or proceed with the supported arrangement. An unclear return leaves alternatives open; a supported action common to them can end the inquiry. Supplying a common brief alone establishes neither common understanding nor an accepted commitment. Keep ordinary settled collaboration lightweight.
Make result dependencies explicit. A source interpretation can begin alongside an independent calculation only if the calculation’s inputs are already settled. Otherwise obtain the prerequisite first or keep the dependent conclusion conditional.
For actions on shared executable state, establish who can write what and how incompatible effects are prevented. Isolate working copies or serialize conflicting actions under the real interface contract. Two participants cannot both rely on an old target state merely because their briefs or prompts differ. SYSE.20 and SYSE.40 supply overlap and shared-capacity conditions; SYSE.47 implements the chosen routing and limits.
SYSE.44:4.3 - Integrate claims against their actual grounds
Ask each return to preserve the result, premises, source/configuration, unsupported conditions and what it did or did not establish. Check that it answers the requested contribution for the same receiving use. A confident response to a nearby question is a missing contribution.
Reconcile disagreements by identifying the claim and its source or assumption. Compare original evidence, run a discriminating calculation or obtain a qualified specialist result when needed. Do not resolve common-source error by majority vote. Independence requires the evidence and checking relation that the claim needs; a second persona or fresh context alone does not supply it.
Combine only mutually usable returns. Test the joins: compatible units, definitions, input versions, scope and whole-result conditions. Preserve a supported partial result when another required result is missing, and state why the whole is still incomplete. The integrator retains the whole result’s required conditions.
SYSE.44:4.4 - Localize change and stop
When an input or source changes, identify the returns that used it and reopen their dependent conclusions. Keep unaffected results. Cancel or redirect obsolete work when doing so is permitted and useful; do not let its later arrival replace a result based on the new premise.
Use the integrated result in the receiving decision or action and observe whether it supplies the needed contribution. A handover delivered on time can still fail there. Return a context omission to context construction, a defective component to its supplier, a bad join to integration and an unsupported whole condition to the domain decision.
Stop when the whole result suffices or the remaining work cannot supply a worthwhile permitted continuation. A timed-out agent is an unavailable result, not evidence for whichever alternative the other agent favors.
SYSE.44:5 - Archetypal Grounding
Compare divisions for the same service-update proposal
The receiving result is a bounded service-update proposal with an applicable interruption/recovery contract and a capacity/backlog calculation. Target, source edition, demand interval, interruption duration, worker effects and allowed backlog must agree. The following are constructed comparison conditions, not measured agent timings.
One available LLM context can do both jobs. A two-context alternative gives the long contract examination to one context and the capacity calculation to another. Here the numerical input is already settled, the common brief is short, both contexts can obtain their needed source, and the integration check is a small comparison of duration, capacity and limit. A stipulated representative trial for this configuration shows that the one-context way repeatedly loses a needed contract qualification when moving between the long source and calculation; separate working material avoids that repetition without adding an equally costly join. The preference is adequate result with least total preparation, execution and repair burden.
Compare three complete ways. One-context completion retains the repeated rereading/checking cost but needs no handover. The proposed division adds the short common brief and join, while keeping each examination usable. A third proposed division omits the settled duration from the calculation brief and makes it a later return from the contract reader. That arrangement creates a serial dependency: the calculation must wait or stay conditional, so its apparent parallel gain is unavailable. C.38 constructs these ways and C.11 selects the second under the stated evidence and preference. No number of contexts by itself supplies the choice.
The contract reader returns the allowed recovery operation, duration, worker effects and source edition. The calculator consumes the settled inputs and returns the backlog trajectory with units and assumptions. Integration compares the modeled interruption with the contract and uses the joined proposal in the engineering decision. A fluent summary that omits the duration would fail this join.
Now supply a short authoritative contract excerpt and a calculation that the same one context can keep and check reliably. The repeated-rereading cost disappears while the extra brief and join remain. Reopen the same-result comparison and keep the work together. If instead duration was genuinely unsettled, obtain it first; useful source examination may continue while the capacity result stays conditional.
During a divided attempt, a new contract edition doubles the interruption duration. Reopen the dependent trajectory and the proposal that consumed it. An unrelated description of the monitoring interface remains usable. If both contexts copied an incorrect capacity from one summary, their agreement adds no independent evidence: obtain the actual capacity basis and retain the calculations as conditional if it is unavailable.
People supply the same joined contributions
A contract engineer and a capacity analyst can enact this division with source access, a written brief and an explicit receiving decision. The engineer returns the applicable duration and recovery condition; the analyst calculates from that version; the integrator checks their shared assumptions and uses the proposal. If both copy the same erroneous capacity sheet, different professional roles do not repair the source. A changed duration reopens the dependent calculation just as in the technical case.
This common operation does not require identical human and model mechanisms. Human availability, preparation and coordination cost enter the whole comparison; model input limits and repeated invocation cost enter their technical realization. For the small dependent digit operations of 347 × 6, keeping one performer and a readable carry layout can remain the cheaper adequate way.
SYSE.44:6 - Bias-Annotation
Benchmarks with easily separable questions can exaggerate the benefit of parallel agents. Shared state, source revisions and specialist authority often dominate real engineering work. Compare the actual arrangement at its required joins, not a generic multi-agent label.
SYSE.44:7 - Conformance Checklist
- The division is compared with one-performer completion for the same whole result, including preparation, dependency and joining costs.
- Each participant receives the needed common premises, sources, boundaries and receiving use; consequential uncertainty about interpretation or intended continuation is resolved enough for the selected action, or remains explicit.
- Serial dependencies and conflicting effects remain explicit.
- Returns retain their evidence and unsupported conditions.
- Integration tests compatibility and common-source failure rather than counting agreement.
- Changed premises reopen only affected returns; the final contribution is used.
SYSE.44:8 - Common Anti-Patterns and How to Avoid Them
Invent roles before finding contributions. Identify a missing result and its integration condition first.
Give every agent the entire history. Select the needed context and source return; duplicated overload can defeat the proposed division.
The integrator smooths over disagreement. Preserve the exact incompatible claim until evidence or a decision resolves it.
SYSE.44:9 - Consequences
The arrangement can obtain separable results with better attention or shorter elapsed time. It adds communication, integration and shared-resource costs, and can amplify a common error. A usable result may therefore justify fewer contexts or a narrower division after comparison.
SYSE.44:10 - Architectural Rationale
Contribution design and executable routing are different results. This pattern constructs what the participants separately supply and how their returns combine, consuming the existing whole-way comparison and choice. SYSE.47 builds the surrounding procedure. Keeping the distinction permits a useful human-integrated comparison before automating the arrangement.
SYSE.44:11 - SoTA-Echoing
The practice question is when divided work improves the receiving result after its joins and total burden are counted. For the LLM realization, Towards a Science of Scaling Agent Systems, v3 supports a task-dependent comparison rather than automatic gains from team size. Its tested configurations do not establish a universal architecture ranking.
Adopt the comparison of task structure and coordination burden, and qualify the actual joins and evidence dependencies here. A single capable context is the serious baseline. Reopen the division when shared errors, context loss or integration cost defeats its expected contribution.
SYSE.44:12 - Relations
A.15.9 and SYSE.9 supply bounded requests and qualified use. SYSE.18 addresses independent governance; SYSE.20 addresses overlapping work. SYSE.42 executes needed calls, SYSE.43 supplies selected retained context, SYSE.47 builds routing, and SYSE.46 tests the configured result. C.38/C.11.CRC compare complete alternatives when changing the overall obtaining arrangement. A.3.3.PI supplies the consequential hidden distinction and its revision; MMP.8.SD compares information-dependent continuations, with C.11.DUA judging inquiry burden. The Reference application connects those results to the participant’s contribution; A.15.7 uses the qualified observation in the next action.
SYSE.44:End
SYSE.45 - Train an LLM Policy from Qualified Interaction Experience
Type: Method pattern Status: Candidate
SYSE.45:1 - Problem frame
Use this when a recurring task warrants changing model parameters or adapters, and qualified interaction experience can teach the targeted behavior. Repeatedly supplying stable guidance may be costly; an existing policy may also keep making a consequential mistake.
The subject is the trained policy candidate and the experience, learning mechanism and retained support that make its changed behavior interpretable. The first useful result is a changed candidate with a bounded comparison, or the precise data, training or evaluation gap.
Compare training with retaining help, improving retrieval, constructing a tool or repairing the external controller. Selecting or transforming the next input belongs to SYSE.52, routing and return consumption to SYSE.47, and external recording to SYSE.43. If training is unavailable, those feasible alternatives remain available. Human practice and learning use E.23.CDI/HCD with their own mechanisms. This pattern supplies a bounded machine-policy intervention; foundation-model pretraining and the receiving domain’s correctness criterion remain external.
SYSE.45:2 - Problem
A successful trajectory can contain incidental values, unnecessary steps or a hidden external contribution. Training on it can reproduce the answer while losing the conditions for success. Removing guidance may then look like internalization even though the agent has simply stopped obtaining necessary evidence.
SYSE.45:3 - Forces
Stable knowledge may be cheaper to use through a trained policy, but changing interfaces make that knowledge costly to maintain. Successful demonstrations give a target, while informative failures reveal its limits. Feedback can reward an observable proxy rather than the required result. Training and validation share a budget, yet repeated validation can make the final comparison optimistic.
SYSE.45:4 - Solution
SYSE.45:4.1 - Fix the target and retained configuration
Name the behavior to change and its receiving result: for example, forming valid arguments, selecting a useful observation or applying stable procedural guidance. Name the base model and parameter or adapter set to update. Separate that change from context, memory, tool, controller and environment changes. When the target is assistance selection, consume SYSE.50’s decision unit, available signals, qualified need evidence and fallback; selecting training does not itself establish which help is needed. For learned effort control, consume SYSE.51’s feasible moves, intermediate decision basis, cost accounting and completion obligations; retain externally enforced hard limits.
Specify the external contributions that remain: execution, current observations, access permissions, retrieval when facts change and independent checking where the claim needs it. Compare complete later configurations through C.38/C.11.CRC and C.11. Include stability of the proposed mapping, expected recurrence, target preparation, training, independent qualification, maintenance and the fallback after a failed trial. A cheaper response that loses a necessary contribution is a different result. The Reference comparison selects an explicit rule for its supplied 100-update conditions and manual work for five; recurrence alone selects no learner.
Keep an unchanged baseline and the means to restore it. Establish the task-family boundary and what finding would make training no longer worthwhile.
SYSE.45:4.2 - Obtain experience with interpretable feedback
Collect observation/action trajectories whose inputs, source/configuration and outcomes can be recovered. Include successful alternatives, informative failures, required restraint and changed-condition cases. Use SYSE.49 when the needed experience must be constructed; qualify its feedback before using it as a training signal.
Separate the agent’s actions from tool observations and the judge’s conclusions. For each selected target, identify what counts as correct and why. A failed trace can teach a corrected action or a failure condition; blindly imitating its actions teaches the failure.
Partition construction, adaptive validation and final evaluation by the dependencies that could leak the answer, such as shared scenarios, source instances or near-duplicate trajectories. Keep the final cases outside training and repeated candidate selection. Generated labels remain provisional until their result meaning is qualified.
SYSE.45:4.3 - Select and execute the learning mechanism
First choose the experience transformation that can supply the intended policy target. The learning algorithm then consumes that target; the names supervised learning, reinforcement learning and distillation do not construct it.
| Available experience and intended change | Construct the signal | Further-use question and return |
|---|---|---|
| A qualified action sequence demonstrates the needed behavior | Pair each retained decision history with the agent action it warrants. Keep tool observations as inputs, and fit only the intended agent outputs. Remove incidental target identifiers or values by varying them in applicable examples. | Does the policy choose the action on a new applicable input and abstain on an unsupported one? An unqualified successful transcript returns to result/trajectory assessment. |
| A recorded decision is wrong or unnecessarily costly | An experience-informed teacher proposes a corrected next action at that history. Qualify it against the task, interface and facts the student will actually have. Train the student on that history/action pair without the teacher’s extra experience. | Does the correction remain warranted without hidden teacher facts? Obtain a needed fact or retain support if it does not. The old next observation follows the old action, not the proposed correction. |
| A search or deliberation procedure finds useful candidates | Preserve selected candidate comparisons, evaluation grounds and backtracking decisions as targets where their signals are available to the student. Distillation can transfer how candidates are generated or compared. | Test unseen alternatives and defeated evaluator premises. Imitating a planner’s trace does not transfer its proof or guarantee. |
| Actual attempts supply outcome or process feedback | Bind feedback to the required result and protected conditions. When a final reward leaves the responsible decision unclear, use a qualified intermediate state or a discriminating action contrast to localize the target. Keep an uncertain credit assignment at that evidential strength. | Test the actual result and the relevant intermediate behavior separately. Reward growth with duplicated effects or lost required access fails. |
| Qualified comparative judgements express a preference | Name whose preference, the task/context, the compared action or answer pair and the grounds for its ordering. Give those pairs to a supported preference-training Method and implementation; return a missing judgement or implementation. | Preferred behavior still needs independent factual, permission and protected-result grounds. A rater’s preference supplies no missing world effect. |
Match the update to that signal. Supervised learning fits qualified actions or corrected continuations; reinforcement learning uses attempted actions and qualified reward; distillation transfers the selected teacher or search behavior. CMP.7 supplies learner construction and the distinction between obtaining data, fit and further use. The selected implementation supplies the exact update algorithm, tokenizer and trainable parameters. Bind them to these data transformations and the later input format. Keep the configuration that actually produced each comparison.
Construct tool knowledge and use separately when that is the gap. Tool-Internalized Reasoning separates learning a tool’s documented semantics, supervised preparation on tool-use trajectories, and optimization of subsequent tool reasoning. For an interval setter, the semantic target includes what operation and unit its arguments denote; a trajectory target then puts the call at the right decision with the right observations. A later reward can distinguish the selected tool and argument choice only at its qualified meaning. The source’s special tool vocabulary, document/token mapping and reward are implementation choices, not prerequisites for every adapter. A tool/argument proxy does not establish the resulting service effect. A plain converter or retained description may supply the needed operation more cheaply.
When reducing repeated procedural guidance, vary only the contribution intended to become unnecessary. Keep execution and current observations available. Test separately a stable mapping presented with all its needed inputs and a decision whose new interface fact was withheld. Their different repairs are worked below. Restore needed support rather than rewarding unsupported fluency.
Bound repeated and episode-local updates. Identify the state or parameters being changed, the qualified signal that permits an update, the experience retained, and the lifetime/reset rule. An episode-local adapter may be discarded at the end; durable shared parameters need a recoverable baseline and regression tests across later tasks. A runtime hidden state or learned memory token is not automatically either shared-parameter learning or an external text record. SYSE.52 prepares the exposed working representation, SYSE.43 maintains separately persisted episodes, and this Method governs selected parameter learning. Specialized backbone or latent-memory-module construction remains with its direct implementation.
Choose update extent and stop from applicable held-out behavior and the available means. Test older useful restraint and correct actions alongside the new target. Repeated self-generated failures are observations of failure; frequency does not turn them into positive labels. Reset, revert or return the unqualified signal when the update destroys earlier required behavior.
SYSE.45:4.4 - Compare the intended later use
Through SYSE.46, compare the trained candidate with the baseline under matched task and retained-support conditions. Read actual result, necessary access/use, restraint, protected effects and total effort separately. Adaptive validation selects candidates; an untouched final comparison supports the bounded reliance claim.
Test delayed or shifted use when the claimed benefit includes persistence or adaptation. Identify intervening model, memory, controller or source changes before attributing a difference to training. A policy that uses fewer tools while missing a newly changed fact fails that receiving use.
Retain, revise or reject the candidate according to the comparison. Return bad feedback to its supplier, unsupported transfer to the training question and runtime routing defects to SYSE.47. A learned predictor or auxiliary target retains its MMP/CMP meaning; correct predictions alone do not establish a useful acting policy.
SYSE.45:5 - Archetypal Grounding
A converter, a learned interpretation and two failed withdrawals
A service accepts an interval in seconds. Target identity, revision and actual effect must be obtained through the current interface. For already structured input duration_ms=12000, a deterministic conversion returns seconds=12. Under the supplied conditions that small controller is adequate and cheaper to construct and qualify than an adapter. Stable recurrence alone supplies no reason to replace it.
Now change the target to interpreting recurring, varied requests such as “take three readings each minute” and “take one reading every twenty seconds.” Both call for seconds=20 under the supplied meanings. The fixed converter still works after a supported structured duration exists, but it does not obtain that interpretation. Compare a phrase-rule parser, retained specialist/guidance support, and an adapter with the same converter, current interface and executor. In this constructed case, representative source examples establish that a small phrase rule leaves many intended forms unsupported; expanding and maintaining it has greater estimated whole-horizon burden than the offered bounded adapter trial, including target qualification, tests and failed-trial fallback. C.11 therefore selects that trial. If those grounds or training access are absent, use the supported parser/guidance way or obtain a worthwhile comparison premise.
Training histories include the request, applicable tool definition, current target/revision and the observations needed for a call. Targets vary phrases, rates, durations and identifiers. Missing-unit or ambiguous requests target clarification. The learner fits the intended interpretation/call, while the executor retains permission, actual-state binding and result checks. Optional semantic preparation and trajectory warm-up address different failures; any later reward must use qualified tool/argument meaning and protected results.
Two held-out failures discriminate what was lost when guidance was reduced. Inspect the candidate’s proposed call before execution; the retained binding checks reject an unsupported call:
| Input and condition | Candidate output and observed defect | Smallest supported repair |
|---|---|---|
duration_ms=12000; the applicable definition still says the argument is seconds, and current target/revision are present | seconds=12000 instead of 12. The stable conversion is wrong despite sufficient inputs. | Retain the small converter or restore guidance; if learning that mapping remains worthwhile, repair its targets and test new values. Another current-state lookup cannot supply the missing transformation. |
The provider changed to a millisecond argument, but the new definition was omitted from the input; the old call uses seconds=12 | The call no longer matches the current interface. Restoring the fresh definition, with the same weights, yields the supported duration_ms=12000 call. | Retain the required interface observation and repair its retrieval/input path. Training on old documentation cannot supply an unobserved future contract. |
Guidance reduction therefore tests a particular stable contribution, not all support. Untouched cases also include ambiguity, unsupported units and restraint after a lost reply. A delayed comparison must record intervening adapter, provider, controller and memory changes before attributing retained benefit.
Construct a corrected continuation from a recorded history
A recorded history H contains the permitted outcome-lookup operation, the original attempt A17, an acknowledgement followed by a lost reply, and the at-most-one-effect requirement. The recorded next action was an unsafe repeat of the mutation. An experience-informed teacher, using that failure and other qualified episodes, proposes lookup_outcome(attempt=A17).
The engineer checks that H itself supplies the attempt identity, supported lookup and unresolved effect that warrant this correction. The student receives H without the teacher’s extra experience; its action target is the lookup, not the old repeated mutation or the following observation. The old next observation remains evidence about the old action. To establish what the correction does, exercise that lookup in a new qualified test occurrence or use applicable environment evidence; relabelling the old observation would fabricate its consequence.
The training example fits only the corrected agent action. If H had lost A17 or the lookup contract, restore that input or target obtaining the missing contribution instead. A final success reward alone would not identify whether safe recovery or a lucky duplicate caused the result; the supplied intermediate attempt/effect state localizes the unsafe replay decision. A contrasting acknowledgement-without-effect case tests the feedback rule.
For a preference branch, the service owner may compare two factually supported reports and prefer one that exposes unresolved status before optional explanation. Record that owner, task and report pair with the judgement. Both reports must still meet the effect and factual requirements before the pair can support that preference target.