3. Working combinations
Use a combination where its intermediate results are needed. Enter with a result already available, take an alternative branch when conditions require it, and stop when the question has a usable answer. The short sequences below explain possible uses; the linked pattern bodies supply the methods.
Finding FPF and DPF bodies. Open FPF-Spec.md or a linked DPF, search for the full PatternID, and read its Problem frame and Solution. A DPF’s own Table of Contents links to the bodies and practical entries. For a large file that GitHub cannot display, use View raw or Download raw file and search the downloaded text.
3.1. Construct an unfamiliar account
Recover the question → choose objects and operations → formulate relations → obtain a consequence → interpret it or revise the failed contribution.
Start with FPF B.5.FM when the mathematical question is missing. MATH.16 chooses a construction from the maps it must support. MATH.1 and MATH.5 can then construct and interpret composable expressions. MMP.10 expresses admissible subject cases; MMP.11 constructs an unknown relation inside supported conditions. Use FPF C.29 to interpret the consequence through the correspondence between the mathematical account and the working situation.
Small use. A cart log records two unit movements. Total distance is two both for forward-then-backward and forward-then-forward. If each entry records an actual signed displacement, assign +1 and -1 and use addition for composition: the final displacements are zero and two. MATH.1/.5 supplies that interpretation. If entries record commands and the cart can slip, the same sum answers a command question. Observation or a movement model must supply the relation to achieved position.
That missing physical account changes what can be concluded. PHY.4 develops physical grounds for restrictions on an unknown law. These restrictions can still leave a physical relation unresolved; use supported subject knowledge, observation or a collaborator for the needed premise. B.5.MPC connects the physical, mathematical and computational contributions. Stop with the qualified consequence, or with the particular relation still needed.
The physical connected-use example follows an unknown response through a balance, a conditional prediction and a distinguishing observation. It shows when an experiment is unnecessary and which premise to revisit after conditions change.
3.2. Understand and change a construction
Recover the operations → identify the result to preserve → compare their compositions → carry the consequence into the changed work.
FPF B.5.RC and B.5.RA recover a construction or argument. MATH.1/.5 builds and interprets operation sequences; MATH.2 tests whether identifying descriptions preserves the required operations and answer. MATH.16 treats functions as objects with evaluation maps.
Small use. Let a stored quantity initially be 1. Operation R returns its present value; W doubles the stored quantity. Performing R then W and W then R both leaves 2 stored. The returned reading is respectively 1 and 2. If the next action uses that reading, equivalence based only on final storage loses the difference the action needs.
Method Engineering, especially ME.3 and ME.7, supplies the working-method account and composition question. Use FPF C.29 to establish what the mathematical elements and operations represent in that work, and which conclusions the correspondence supports. ME.12 checks claims about the description; it does not construct the replacement method.
MATH.17 constructs transformations of rules and derives the laws their use needs. MATH.18 compares descriptions through interpretations and recoverable consequences. For example, preserving how operations compose can permit a calculation in the second description; recovering the first answer also needs a return that retains its required distinctions.
CMP.14 develops the algorithmic interaction: expose shared state and allowed observations, construct coordination, and retain separate correctness and progress arguments. CMP.12 supplies interpretation or translation when executable descriptions change.
ME.6.MC derives the consequences of proposed arrangements under a mathematical account of the work. ME.25 uses a mathematical transformation to construct a changed working procedure. The connected ME example develops local summaries for a distributed calculation, then changes those summaries when the recipient asks for another statistic. The mathematical preservation argument supplies one ground for the change; available performers, timing and practical benefit remain working questions.
For a longer construction, use MP-COMBINE-RESULTS. It follows an unreliable aggregation rule through a counterexample, compatible summary, operation on summaries and proof, then reopens retained information when the answer changes. Each contribution supplies a result the next uses.
3.3. Choose with incomplete information
Recover the available information → formulate the allowed choice → compare consequences → act, revise the requirement or obtain a useful missing indication.
MMP.8 formulates what may depend on information available at the time of choice. When a probabilistic recording question is needed, MMP.7 derives the record’s law. FPF C.16.IR addresses what an observation relation distinguishes, including bounded indications. C.11.DUA compares the value and cost of further inquiry. The MMP Readme’s observation-to-action entry connects these contributions with a usable instruction.
Small use. One resource must go to one of two requests. A report identifies the request that needs it with probability 0.8 in the first circumstance and 0.7 in the second. Following the report meets a requirement of at least 0.65 success in either circumstance, provided the report arrives before allocation. A zero-failure requirement remains unmet. If the first circumstance instead has probability 0.95 and the criterion is average success, always choosing it gives 0.95, while following the report gives 0.795.
The criterion, information and timing determine the next move. The mathematical comparison does not require further evidence when those inputs are already sufficient for the intended decision. A statistical fit is only one way to resolve an unknown. If unresolved alternatives all permit the same action, a bound may suffice. Observational agreement also leaves an intervention question open when it depends on an unestablished mechanism.
MMP.12 formulates recovery from records and makes any added restriction or penalty explicit. MMP.13 constructs an inferential conclusion and propagates its uncertainty to the requested quantity. MMP.14 constructs a comparison for a consequential discrepancy, revises the implicated component and recalculates the receiving prediction. Use these contributions where the existing law or bound leaves that work unresolved.
MMP.15 determines which intervention consequence follows from the available data and causal assumptions. MMP.16 constructs an obtainable observation that separates a consequential ambiguity and compares its benefit with its burden. MMP.8.SD retains the information and change law needed for continuing choices.
For a question about the same observed case under another action, MMP.19 connects factual inference with changed mechanisms through common underlying conditions. It constructs a joint or nested comparison, or bounds what remains ambiguous. An expected-outcome choice can still be settled by the intervention distributions alone.
The intervention-and-information example connects these methods: use the supported intervention losses, derive the possible reports, choose whether to observe, and make the later action depend on the received report. If the report no longer distinguishes the possibilities, the observation ceases to justify its cost.
3.4. Obtain a result under limits
State the consequence to retain → derive what removed detail contributes → obtain a replacement or bound → interpret the result at its supported reach.
MMP.9 develops this operation for evolution laws. FPF C.29.2 connects a question with an obtaining process; C.29.3 connects that process to physical means. Use the MMP Readme’s reduction entry when the removed detail is the difficulty.
Small use. Suppose two nonnegative populations satisfy x’=-x and y’=-2y, with x(0)+y(0)=1. Retain only z=x+y. Its derivative is -x-2y, which z alone does not determine: at z=1 it can be -1 or -2. Yet the source laws give e^(-2t) ≤ z(t) ≤ e^(-t) for t≥0. At t=3 the upper bound is below 0.05. A question asking whether the remaining total is below 0.1 is settled without recovering the initial split. At t=1, a changed threshold of 0.2 lies between the two bounds, so deciding whether the total is below that threshold requires more about the initial split.
A smaller model can therefore be sufficient for one decision while lacking a complete evolution law in the retained quantity. New interventions or a different time horizon can require the missing distinction again. FPF’s ordinary comparison and improvement methods can retain several complementary models.
MATH.20 now supplies the general bounding method: derive a comparison covering all admitted cases, carry it through the needed operation, and tighten it only if the answer needs more. MATH.21 constructs limits and justifies the operations performed on them. Physical Thinking supplies PHY.3 for limits derived from permitted physical transformations, and PHY.5 for choosing which effects must be retained. CMP.8 constructs an effective approximation and finite return condition; CMP.9 constructs sampling and estimation. Numerical procedures are one application of these algorithmic methods. MMP.17 constructs a cheaper supplier of the responses the receiving operation needs. MMP.18 combines models through compatible exchanges and joint assumptions, retaining consequential approximation and dependence.
The replacement-and-coupling example sums two distinct amounts with their error bounds. A tighter decision threshold makes one coarse response insufficient; replacing that response can settle the decision without refining the other model. Asking when the threshold was crossed requires temporal information that the final amounts do not contain.
The full MMP-SUFFICIENT-ANSWER example works this connection across MMP, MATH and FPF. Its changed-time question derives the particular initial-state distinction worth recovering.
3.5. Continue and distribute thinking
Expose the missing contribution → recover or obtain the method → use it under changed conditions → choose a worthwhile next question → retain the way of obtaining and using the result.
FPF B.5.RC/RA recovers the needed construction or argument, B.5.RR follows changed premises, and B.5.QD develops a further question. C.40.CD connects developing problems with developing ways to address them; C.36.RP addresses retaining and renewing methods. MATH.19 constructs a missing argument, MATH.22 follows a change of axioms through its consequences, and MATH.23 develops a next conjecture from a changed construction. MATH.4 constructs a witness by induction; MATH.12 extracts a construction from a proof.
Small use. Ask for a rule that selects one member of an unordered two-element set and respects renaming. Swapping the elements leaves the set unchanged but moves either possible selection. Such a rule is impossible. MATH.13/.9 exposes the obstruction. Adding a distinguished member permits a choice; returning both members changes the requested answer. A program that takes the first stored element uses an ordering that the original question did not provide.
This obstruction opens useful next questions: may the representation introduce an order, does the intended work permit that extra structure, or does it need the whole set instead? A mathematical, computational and methodological discussion can now concern the same identified difference.
Make the next contribution obtainable. First locate what prevents its use: an unavailable input, an unexplained relation, an operation the recipient cannot yet perform, or a contributor they cannot reach. C.36.RP distinguishes these repairs. DOCA.3 formulates what a development opportunity could supply; DOCA.4 constructs prospective ways to obtain it, including help, access, learning and changed work, with the support each requires. Use an adequate existing arrangement when it already supplies the contribution.
For an explanation, EXD.1 identifies the relation this recipient needs. NOT.3 supplies the reading procedure; NOT.7 can repair an expression that makes the operation difficult. When a human needs practice in using the relation, HCD.6 constructs a task from the later work and its permitted help. EXD.5 can guide the person’s explanation and a revealing retry. Use the response to decide what help is still needed. Reading a worked answer supplies no observation of that person’s learning. If later work requires retention or transfer, test performance in those conditions. Use the appropriate training or configuration method when an AI contributor needs a new capability.
Connected use: develop a selection method that a team can change. The following is a constructed design example, not an observed training result. A dispatcher must choose one of two equally eligible requests. Their identifiers are arbitrary: exchanging the identifiers must exchange which request a deterministic rule selects, while changing nothing else. The two-element obstruction above shows why those conditions cannot all hold. Choosing the first identifier introduces an ordering the requirement did not authorize. B.5.RA and MATH.13 recover the reason; EXD.1 makes that missing relation the explanation’s target.
The team now has a consequential choice. It can supply a meaningful priority, request both items, or change the requirement to equal probabilities of selection. These changes answer different questions. Suppose the receiver permits the last option. MATH.23 develops the changed claim: a uniform distribution on a finite eligible set is preserved by renaming, because renaming permutes equal probabilities. For two requests, an available unbiased random bit assigns probability one half to each. CMP.9 supplies the construction and its randomness conditions. A temporary ordering can map the two bit values to the two requests; the distribution remains uniform after renaming even though a particular bit value need not select the corresponding renamed request. The resulting guarantee concerns the distribution.
This result permits a division of work. One contributor can establish the selection law, another obtain and implement the random choice, and the receiving dispatcher decide whether probability-based fairness serves the work. People, AI and tools may provide these contributions in different combinations. ME.6.MC and ME.25 connect the mathematical account to a changed working arrangement: the selected request has to reach the performer, and the permitted source of randomness has to be available. Choose which derivations each contributor needs for their part; the receiver must understand which guarantee the arrangement supplies.
If a human dispatcher can run the routine but says that one unlucky choice disproves equal probability, the explanation has a specific target. Show the two equally likely bit outcomes and their respective requests; ask which outcomes are possible under equal probabilities and what one selection can establish. For practice, change the condition to an available bit that returns 1 three quarters of the time. Directly mapping its values to the requests no longer gives equal selection probabilities. The learner can identify the failed assumption and return the need for another random-choice construction; they need not invent a randomness extractor to make that useful return. HCD.6 and EXD.5 supply the practice and assistance choices. An actual response would support a conclusion about that attempt under its stated help, not a claim of general competence.
Changed working arrangement. Two dispatchers now act simultaneously, and each eligible request may be assigned only once. Independent fair choices select the same request with probability one half: of the four equally likely pairs AA, AB, BA and BB, two repeat a request. The earlier marginal fairness calculation remains true for each dispatcher but does not meet the new joint requirement. CMP.14 exposes the shared state and permitted histories. A single allocator can randomly permute the two requests and assign distinct entries, or coordinated dispatchers can use an available indivisible claim-and-remove operation. For the latter arrangement, a failed claim must lead to the remaining eligible request; its progress also depends on the service and communication conditions. ME then reconstructs the corresponding allocation and support arrangement. The mathematical calculation alone does not install that service or give a performer access to it.
The new problem is worth pursuing because its answer enables parallel assignment without duplicate work. C.40.CD connects that receiving need to the changed method; MATH.23 can investigate how the construction extends to larger sets. A first inquiry can compare allowed assignment histories with the two independent choices, before investing in a general implementation. If one dispatcher already meets the work’s needs, this branch can remain a future opportunity; continuing research is not a condition for using the current result.
Keep what another contributor will need to renew the method: the chosen fairness meaning, the construction, its randomness and coordination assumptions, and a case exposing the difference between separate and joint guarantees. C.36.RP also retains access to the needed help or executor. This preserves a way to obtain, explain and change the answer. The same sequence can begin with a failed physical interpretation, an unfamiliar proof or a changed notation: recover the consequential relation, obtain the missing contribution, use it in the receiving work, and let the result open a justified next question.
3.6. Construct and change an algorithm
Specify the required answer → construct its obtaining procedure → share or summarize only what the answer permits → bound cost and error → use the result → revise the changed dependency.
CMP.1 builds effective conversions and answer recovery when another solver can help. CMP.2 constructs a recurrence; CMP.3 chooses reuse and order. A mathematical construction or the formulation in MMP supplies their intended objects and answer conditions. It does not supply an algorithm merely by defining the answer.
The connected CMP example shows how a relaxed bound can finish a search for one best selection. Asking for every equally good selection changes which branches may be discarded and what a shared table must retain. CMP.4 supplies the exclusion rule, CMP.5 the bound and CMP.10 the representation comparison. The mathematical optimum can remain unchanged while its obtaining and output procedure changes.
If a weaker answer is acceptable, CMP.8 constructs an approximation and its error account. If the question concerns every possible algorithm under given access operations, CMP.11 constructs a lower bound rather than extrapolating the cost of one implementation. A changed input promise or tolerated error can reopen that limit. C.29.3 then connects the selected operations to physical execution where that contribution is needed.
These methods also support changing a way of working. A solver can move a contribution to another agent; a changed representation can make sharing possible; an interaction rule can prevent one contribution from invalidating another. C.29 and Method Engineering establish the correspondence to the actual working arrangement. They retain any physical or organizational requirement that the algorithmic argument did not address.
3.7. Carry meaning through different expressions
Recover the operation → establish references and interpretation → preserve or expose translation loss → carry the intended change → revise the affected rule.
NOT.1–.3 connect the work requirement to formation and interpretation. NOT.4 constructs a transformation under its preservation conditions; NOT.5 exposes what a translation cannot recover; NOT.6 coordinates the useful remaining forms. NOT.7 repairs the resulting reading or editing burden. A single adequate expression can be used without this whole combination.
The connected NOT example keeps a formula, operation graph and table usable through a parameter change. The table’s sampled values do not determine the formula outside those inputs. A new parameter changes derived values while an independent observation remains a record of what was observed. Replacing a supplied variable by repeated sensor reads changes the interpretation again, so an arithmetic rewrite must be reconsidered.
MATH.18 supplies mathematical interpretation and its preservation arguments. CMP.12 constructs an effective interpreter when the expressions must be executed computationally. C.29 and the physical or modeling method supply the correspondence to the measured subject. A human reading a diagram can obtain a consequence without running software; a formal meaning alone does not supply an algorithm for every consequence.
When the expression is a score or gesture, NOT.8 adds the needed pulse, frame, segmentation and reading procedure. The direction a sign represents and the capability to perform that motion are different contributions. Method Engineering helps change the working method when the newly expressed or computed result makes another way of working possible.
3.8. Change an operating flow without hiding its waiting
A manager wants work to reach its recipient sooner. A local queue becomes shorter, yet the recipient waits just as long. Begin with the recipient’s completion event and trace what a proposed change does to the work before, inside and after the measured operation. Operations Management supplies the operating subjects, policies and consequences; this Suite helps construct and interpret the model.
Connect the question to the arrangement. Use OPS.1/.3 to identify the requested result and participating work, then OPS.11 to recover consequential relations across resources and services. Distinguish a work route from the people or machines that perform it. Two stations in a diagram can need the same operator; one station can instead contain several interchangeable machines. Recover the intervals during which each resource is needed, including any unattended running time. B.5.MPC connects this account of the work with its mathematical description and computation. A changed physical arrangement can therefore change the answer even when the route diagram stays the same.
A constructed case has four orders available from time zero. Each needs one hour at A, then two hours at B. Each station has its own continuously available resource, jobs use each station in order, and there are no setups, failures or returns. Completion at B makes an order ready for its recipient. Compare transferring all four orders together, transferring each as soon as A finishes, and adding an internal limit of two unfinished orders to that second policy. For each policy, use the earliest permitted starts. When an internal place becomes free, admit the next waiting order immediately.
Construct alternatives before selecting a sequence. A.22.CGUS gives the ordinary operation: name the alternatives, their conditions and the facts that enable, block or leave them unresolved. For the unfolding work, ask whether A can start another admitted order and whether B can start a transferred order. E.18.3 gives the conditions for identifying a transformation-flow unfolding structure when the work needs that structural account. The ordinary continuation comparison is sufficient for this calculation.
MMP.10 turns the scheduling question into time variables and conditions. Let a_i be order i’s internal admission, s_Ai and f_Ai its A start and finish, and s_Bi and f_Bi its B start and finish. For the given station order:
- s_Ai is at least a_i and the preceding A finish; f_Ai = s_Ai + 1.
- s_Bi is at least f_Ai and the preceding B finish; f_Bi = s_Bi + 2.
- Batch transfer additionally requires B to wait for every A finish.
- With two internal places and single-order transfer, orders 1 and 2 can enter at zero; each later admission waits for the B completion that frees its place.
Taking the maximum of the stated lower bounds constructs the earliest starts for each fixed policy. CMP.2 and CMP.3 explain how to obtain these dependent results in order. A spreadsheet or a short program can execute the recurrence; four orders can also be calculated by hand. C.29.2 keeps the calculated schedule and the means of obtaining it distinguishable. If a resource is shared, its competing operations also need a chosen order. Construct feasible alternatives and compare their completion times; the order at each station alone may leave that resource conflict unresolved.
Compare what actually changes. The resulting times, in hours, are:
| Policy | B completions | Last completion | Mean from customer arrival | Mean from internal admission |
|---|---|---|---|---|
| All orders admitted at zero; batch transfer of four | 6, 8, 10, 12 | 12 | 9 | 9 |
| All orders admitted at zero; single-order transfer | 3, 5, 7, 9 | 9 | 6 | 6 |
| Single-order transfer; two internal places | 3, 5, 7, 9 | 9 | 6 | 4 |
The third policy admits orders at 0, 0, 3 and 5. Its internal residence times are 3, 5, 4 and 4; the excluded waits total eight order-hours. OPS.15 keeps both event pairs available, so OPS.10 compares service from the recipient’s waiting origin. Smaller transfer batches improve service in this case. The additional internal limit relocates waiting without further improving those completion times. In another operation, a limit can change interference, returns or service duration; those effects need their own account.
The count-time relation makes the boundary visible. Over the nine-hour empty-to-empty interval, internal unfinished work occupies 16 order-hours. Its mean count is 16/9, equal to the completed-order rate 4/9 times mean internal residence 4. The customer-boundary area is 24 order-hours and mean residence 6. For an observation window that cuts through unfinished orders, count only each residence interval’s overlap with that window; averaging the completed orders alone can omit the work occupying it. This finite-window reasoning follows the area construction in Sigman’s notes on Little’s Law, rather than assuming a steady regime or a delay distribution.
Return to the changed premise. Suppose one operator must now perform all A and B work, with no overlap. Add that shared occupancy to the model. The work needs twelve operator-hours, so the previous nine-hour finish is impossible. Performing A and B for each order in turn attains twelve hours under these conditions. If a decision requires completion within ten hours, this bound already settles the proposed arrangement. Choosing a remedy returns to the available ways of changing access, work or the commitment; further queue statistics cannot make twelve required hours fit into ten.
OPS.8 uses a chosen comparison to set release and protection, OPS.14 contributes the financial consequences when they matter, and ME.25 helps reconstruct a changed working method. Observation after implementation can reopen the resource occupancy, duration, transfer or completion premise. The example combines a subject account, mathematical constraints, an obtaining procedure, measurement and an operating decision; it is one application of foundational thinking. Its finite orders do not set the scope of the general methods.
3.9. Keep the vertical visible while doing the work
Contributions also meet within the same ongoing work. While constructing a model, you may interpret a notation to carry out a calculation that is part of testing a physical account. Ask what the calculation is doing in that inquiry, which constituent operations it needs and which whole conditions constrain them. B.1.5.EW gives that recovery Method.
For example, an engineer computes an average during the analysis of a response experiment. Adding samples constitutes part of computing the average, and that computation is part of the analysis currently under way. If the inquiry changes from average response to the first limit crossing, retaining only the sum and count no longer preserves the required answer. Mathematical Thinking helps identify the lost information and construction; Computational Thinking changes the calculation and its storage; Mathematical Modeling and Physical Thinking retain the observation conditions. Notational Engineering becomes relevant if the expressions hide which quantity or time each value denotes.
Knowing addition and knowing the experiment’s purpose can leave the intermediate calculation or interpretation beyond a contributor’s present capability. Obtain that contribution, explain or practise it, or change the arrangement. The result can depend on several people or AI agents, but their available contributions and communication must fit together. Successful separate operations do not establish that compatibility. Earlier observations remain earlier work; an ongoing analysis does not imply that the instrument is still observing.
A proposed faster constituent goes through B.1.5.RS: check what each relevant encompassing use needs, what survives and what adaptation is required. This vertical complements the longer result-to-use routes above. It does not prescribe a fixed number of levels or turn the five DPFs into five successive stages.
3.10. Develop a way to act from trials
Make a candidate act → obtain informative feedback → change what can produce the behavior → examine the resulting way in its intended use → retain or revisit the contribution that matters.
Use this connection when you can try a proposed way of acting but cannot yet obtain the behavior the work needs. A walking controller may succeed on its training terrain and fail elsewhere. A learned procedure may work only after inheriting earlier training. An interacting group may lose the behavior its members showed separately. The useful result is an acting construction with a justified next use, or an identified contribution still needed to obtain it. A sufficient existing controller, direct calculation or affordable finite comparison can already provide that result.
FPF C.40:4.1–4.4 and :4.7 explains the general connection between changing a way, actually applying it, examining what that application obtains and choosing further development. CMP.7 constructs a learner from the information its feedback supplies; its finite examples can suffice without neural search. MMP.8.SD helps formulate which observations can influence an action and which consequences matter later. The application below joins these contributions to neural representation, trials and change. It assumes a reader who can understand a parameterized policy and implement or obtain its execution. Use a specialist source or contributor for an unfamiliar algorithm; the particular contribution to request is identified at each return.
Obtain a controller through a complete small search
Start with the behavior and the means of trying it. In the book’s walking case, §3.2, the task is simulated locomotion over uneven terrain. The environment returns 24 observations and accepts four bounded motor commands. Specify the simulator version, terrain conditions, episode ending, observations, action interpretation and receiving criterion. The fitness used to select a candidate must come from that candidate acting under these conditions.
A fixed neural representation can make the first experiment concrete. For example, choose two hidden layers of 16 units and compute h1 = tanh(W1 o + b1), h2 = tanh(W2 h1 + b2) and a = tanh(W3 h2 + b3). Here o has 24 entries, h1 and h2 have 16, and a has four entries in the permitted interval. The matrices have shapes 16×24, 16×16 and 4×16; the biases have lengths 16, 16 and 4. Concatenate their 740 scalar parameters in a fixed declared order to obtain a candidate vector theta. Decoding must restore that same order and those same shapes. This is an explicit small architecture for the example, not a claim that these widths are optimal or the widths used in the reported book experiment.
For one trial, decode theta, reset the environment and the controller’s episode state, and obtain the first observation. Compute an action, let the environment perform it, accumulate its returned reward and continue from the new observation until the defined episode ending. The return is the sum of those rewards together with the identity of theta and the trial conditions. For repeated trials, reset afresh and compute the selected aggregate, such as mean return. The weights of this fixed controller remain unchanged during each episode. A recurrent or learning controller needs a different, explicit rule for which state or weights may change and when they reset.
An elementary search needs no hidden optimizer step. Begin with a parameter vector m, a positive perturbation scale sigma, a finite batch size and a trial budget. Sample independent vectors epsilon_i of standard normal values and form theta_i = m + sigma epsilon_i. Try the incumbent m and these proposals on the same newly sampled set of terrains; using shared trial seeds can reduce differences caused only by which terrain each received. Retain a tested vector attaining the highest mean return as the next m. Retain the incumbent when it is among those maximizers; otherwise choose the first maximizing proposal in generation order. Repeat while the remaining budget and receiving need justify another batch. This mutation-and-selection construction is a usable baseline. Its fixed scale and finite trials can miss useful changes; selecting the highest measured mean does not prove improvement on unseen terrain.
For the book’s CMA-ES alternative, replace that proposal and update operation with the complete algorithm in Hansen’s tutorial, Appendix A and Figure 6. Its generated parameter vectors enter the same trial operation. Return each vector’s evaluation to the matching distribution update, including the previous covariance, evolution paths and step-size adaptation. The book’s covariance illustration explains one contribution but is insufficient to implement the complete update. If the chosen implementation minimizes its objective, supply the negative of the return being maximized, consistently with its selection and stopping rule. The new distribution mean and the best evaluated controller are different objects: a mean that has not been tried is another candidate, not an established replacement for a tested controller.
Retain the selected evaluated vector with its decoder, input and action interpretation, any normalization and state-initialization rule. Apply that retained controller in further trials answering the receiving question. This closes the connection: numerical description → acting network → experienced trajectory → measured return → changed proposal or retained controller → further-use evidence. An optimizer checkpoint or a high score alone cannot perform the next walk.
The environment’s actual semantics can change this connection. In the Gym v0.21 implementation, reward uses change in position-and-angle shaping, a penalty linear in clipped absolute action magnitude and a fall penalty. The book’s descriptions of a squared-torque penalty, a constant alive bonus and absolute body position among the observations do not describe that implementation. Its Figure 3.1 also calls reward end-only, whereas the rollout sums step rewards. A policy cannot use a position that it never observes, and changing the reward changes which behavior the search favors.
For an implementation, obtain the actual reset, action and episode-ending interface from the chosen environment version. The Gymnasium task documentation and implementation provide that return; its interface distinguishes termination from truncation. The printed book listings also need repair of their step unpacking and inconsistent fitlist/fitness_list name. Use the operations above to construct the implementation and its checks; verify the selected software before relying on its execution.
Use failure to change the relevant contribution
In the book’s experiment, a controller selected on a single rollout can win because its terrain was easy. Testing it on 100 further trials exposes weaker average performance. Averaging 16 rollouts during selection changes the evidence that drives development; the later 100-trial examination answers a different question about the selected controller. More representative trials can reduce accidental selection, while consuming more interactions. The full cost also depends on episode lengths, candidates, updates and computing arrangement, so the repeat count alone does not establish a 16-fold ratio of total cost or a universal comparison with another learning method.
Suppose the required use includes slippery terrain that the selection trials omitted. Add or obtain representative trials of that condition and compare the controller and sufficient alternatives there. If two situations require different actions but give the controller indistinguishable observations, more trials alone cannot supply the missing distinction. Obtain a useful observation, retain relevant history through memory, or change the promised behavior. If the information is available but the selected representation cannot express the needed response, change the network or its generator. If an adequate response is expressible but the chosen perturbations rarely reach it, compare a different changing operation, initialization or learning arrangement. C.40:4.11 joins that representation choice to the actual changes it makes attainable.
A rare-failure constraint requires its own examination; a satisfactory mean does not settle it. Likewise, performance in a simulator must be related to the sensing, actuation and conditions of a physical robot before it supports that use. Chapter 6’s transfer constructions connect preparation of these conditions, training and actual transfer. Return to the implicated relation when a result fails, while retaining what the old trials still establish.
Some intended behaviors need random action at use time. Against an opponent who exploits a predictable response, a network can instead produce the parameters of a distribution and the agent actually samples its next action. For independent Gaussian action components, require positive standard deviations; if their vector is sigma, covariance is diag(sigma²), with an explicit treatment of action bounds. Evaluate the resulting stochastic policy under its own variability. Random mutations while developing a deterministic policy do not supply this behavior. Book §3.2.5 gives this branch; the receiving task decides whether it is useful.
Compare complete ways of obtaining the behavior. Lack of correct target actions does not preclude learning from delayed reward or estimating a gradient of expected return. Salimans and colleagues, §§2–3, explain an evolutionary strategy that estimates a gradient for the expected performance under parameter perturbations. Its distribution, rank transformation and parallel trials determine what signal and cost it uses. A direct controller, reinforcement learning, population search or a hybrid remains eligible under the same receiving test. Choose by obtainable information, reachable changes and complete cost, rather than by a promise that one method always escapes local optima.
Change the representation and retain useful alternatives
When changing a flat vector cannot express the needed structure, an inherited description can specify a graph or a generator of networks. The NEAT construction, book §3.3, connects historical correspondence of genes, protection of new structures, time for their weights to adjust and selection of useful further development. Stanley and Miikkulainen’s primary account supplies the actual mutation, alignment and species operations. Aligning genes permits a defined crossover; it does not guarantee that the parents’ useful behaviors survive in the offspring. The resulting network still has to act in the receiving trial.
An indirect encoding inserts another performed operation: inherited material generates a network or developing system, which then produces the behavior being judged. Chapter 4 develops spatial generators, developmental processes and input-dependent constructions. Choose the generator together with the changes it permits. A small inherited edit may alter many connections; a regular generator can make a coordinated change easy while preventing a useful local exception. Return to the generated behavior and its attainable variations when deciding whether to retain, extend or replace the encoding. Compactness alone establishes neither interpretability nor useful future change.
If one best-so-far controller destroys access to materially different continuations, retain an executable collection. In MAP-Elites, evaluate a candidate’s task quality and behavior descriptor, place it in the corresponding region of a defined archive, retain the better candidate for an occupied region, and generate further candidates from retained material. The descriptor decides which differences survive this local competition. Keep each stored controller and the conditions needed to execute it; a plotted dot is insufficient for renewed search or use. Chapter 5 explains related choices, including novelty, local competition, ensembles and reuse of acquired material. C.40:4.12 connects those differences to a consequential later use rather than treating diversity as an automatic benefit.
Uncertain measurements can change which candidate deserves a place, and changing the descriptor changes the archive’s distinction. Re-evaluate the relevant stored material when that difference matters. If the receiving question asks for one adequate controller, retaining a large repertoire may add needless work. If it asks for alternatives across energy, speed or another objective, bounded multiobjective retention is a different construction: NSGA-II selects among nondominated fronts and can discard members when capacity is exhausted. Neither an archive nor a nondominated set proves that every retained controller is suitable for deployment.
Develop the way that learns or obtains a result
A candidate can describe how to obtain a controller instead of describing only its final weights. For architecture search, construct the proposed network, initialize and train it under the declared procedure, then use its resulting behavior and cost to judge the architecture. Chapter 10 develops this whole, including reusable modules and task-specific combinations. Chapter 11 changes other obtaining contributions, such as loss, activation, training-data use and learning code. Comparing untrained architectures does not substitute for the training whose result the receiving question needs. A shortened trial or surrogate is useful only at the conclusions its relation to that full application supports.
State what one candidate may inherit. Continuing from trained weights asks whether that continuation works; initializing afresh asks whether the proposed obtaining way works from that start. For a controller that learns during a lifetime, execute that learning across the specified experiences before assessing its result. Reset, permitted feedback, memory and inherited initial state are parts of the construction. Chapter 12 develops evolutionary/RL combinations, learnability and plasticity. In the primary Baldwin-effect meta-learning construction, lifetime learning affects fitness while the evolved starting conditions are what reproduction preserves. The update or learner remains an actual operation to implement or obtain, not a label attached to a successful final network.
These alternatives can also exchange useful material. A population can provide experiences for a gradient learner; an improved learner can return a policy to the population. The evolution-guided policy-gradient construction supplies that specific exchange. Preserve how experiences are collected, which parameters are updated and how a returned policy enters further comparison. Copying a score between the two procedures would not supply the exchanged experience or behavior. C.40:4.7 keeps the candidate way, its real application and the receiving result connected across these cases.
Make interacting and continuing behavior work as a whole
When a controller is assembled from separately developed components, test the actual combinations from which each component receives credit. Chapter 7 develops cooperative neural components, teams, adapting opponents and local cellular rules. A component that works with one partner may fail with another. Construct the partner selection, shared observations, action interface and credit relation before interpreting its fitness. C.40:4.5/.6 supplies the general combination and adaptive-trial connections. For competition, retain appropriate earlier or alternative opponents when a victory over the current one could conceal lost ability; a changing opponent also changes the meaning of the comparison.
If successful joint action depends on information available only to one partner, signaling and a shared convention may be needed. When each actor already observes enough to perform the task, communication can be unnecessary. Inherited coordinated responses and a code acquired from partners require different constructions. In Li and colleagues’ learned-communication construction, retain older agents who carry the acquired code, and let newborns learn through actual rewarded interaction with them. The newborns inherit parameters governing learning: memory size, reward discount, decay and the threshold for fixing an established response. Their acquired policy maps and event memories start empty; the parents’ learned maps are not copied into them.
Perform the generation in order: newborns first interact with both parents; then the population socializes across pairs; rewards accumulated during socializing rank agents for survival. In this construction, the 25 best agents whose lives span fewer than four generations become the next seniors. They retain their acquired responses and provide parents and partners for new learners. For the actual associative update, return to the source’s Real Time Learning and Evolution: a learner unit maps input patterns to activation parameters, records input-output events and changes the corresponding parameters from discounted reward, while potentiation protects established responses from unstable newcomer feedback. Its population, trial and reproduction rules supply the conditions in which that update obtains a shared code. Keep this performed learning and the surviving carriers alongside the inherited learning parameters. If the partners, informative feedback or learning operation are unavailable, obtain that missing contribution before claiming transmission; a description of the learner cannot supply a population carrying a convention.
A local update rule offers another whole: initialize a cellular state, repeatedly apply the rule using its permitted neighborhood information, obtain a larger structure or functioning system, and assess its behavior. Vary continuation time or apply a disturbance when persistence or recovery matters. A snapshot resembling the target can be transient; successful recovery can depend on information or actuation absent after a different injury. Retain the rule and initialization needed for renewed growth, or the live state needed for continuation, according to the actual receiving task. The source’s neural cellular-automata sections specify those operations; the desired whole is not obtained by merely naming its cells.
Body changes can invalidate acquired skills. The ESP extension gives a concrete construction: change morphology together with a new body-affecting skill, re-evaluate the protected older skills, reject excessive losses, then hold the selected morphology fixed while older controllers adapt to it. Its prescribed syllabus and skill interfaces supply the dependencies; further skill composition uses the resulting body. This can obtain a new capability while retaining needed earlier ones, at the cost of those repeated trials. Chapter 9 opens related questions about reachable further development, coupled environments and solutions, and changes of organization. A richer successor or a finite recovery result supports its tested continuation; it does not establish unlimited innovation.
For continued discovery, inspect what a retained basis can actually develop into. Try a feasible further change, obtain its quality and difference, and use those results to decide which basis merits continuation. Preserving current behavior and preserving future possibilities can lead to different choices. C.40:4.11/.12 explains that general return; the neural encoding, body, challenge generator or interaction supplies the particular possibilities. If an adequate current result serves the work, continued discovery can remain a separate purpose.
Use human or generated contributions in the receiving operation
A person can change task conditions, demonstrate a behavior, choose among candidates or propose a useful alteration. Chapter 8 explains interactive development, branching and preparation of material people can meaningfully judge. Identify what their response supplies and perform the resulting change. A selected image may identify an interesting branch; it does not yet supply the working controller behind that image. Prepare informative comparisons and preserve access to the chosen material, then try the resulting candidate. C.40:4.9/.10 connects comparison and human contribution to this actual continuation.
Generated material can enter as the candidate, a changing operation, training experience or an environment. In Chapter 13, these placements lead to different constructions. For example, a language model can propose code, the code is executed under a specified test, and that result selects or informs the next proposal. The execution and test provide a way to reject fluent but ineffective output. A generated training set instead has to train a learner whose further performance judges the generator; visual plausibility is not that performance. Retain the generator or prompt, generated material, actual receiving operation and its returned result at their respective roles.