Library / Foundational Thinking DPF Suite Reference
Jump to passage
In this reading

Link to current text

Published source confirmed at last check

Source changed 2026-10-02 23:06:08 UTC · snapshot created 2026-10-03 01:38:24 UTC · last check 2026-10-03 03:10:20 UTC

3.10. Develop a way to act from trials

Make a candidate act → obtain informative feedback → change what can produce the behavior → examine the resulting way in its intended use → retain or revisit the contribution that matters.

Use this connection when you can try a proposed way of acting but cannot yet obtain the behavior the work needs. A walking controller may succeed on its training terrain and fail elsewhere. A learned procedure may work only after inheriting earlier training. An interacting group may lose the behavior its members showed separately. The useful result is an acting construction with a justified next use, or an identified contribution still needed to obtain it. A sufficient existing controller, direct calculation or affordable finite comparison can already provide that result.

FPF C.40:4.1–4.4 and :4.7 explains the general connection between changing a way, actually applying it, examining what that application obtains and choosing further development. CMP.7 constructs a learner from the information its feedback supplies; its finite examples can suffice without neural search. MMP.8.SD helps formulate which observations can influence an action and which consequences matter later. The application below joins these contributions to neural representation, trials and change. It assumes a reader who can understand a parameterized policy and implement or obtain its execution. Use a specialist source or contributor for an unfamiliar algorithm; the particular contribution to request is identified at each return.

Obtain a controller through a complete small search

Start with the behavior and the means of trying it. In the book’s walking case, §3.2, the task is simulated locomotion over uneven terrain. The environment returns 24 observations and accepts four bounded motor commands. Specify the simulator version, terrain conditions, episode ending, observations, action interpretation and receiving criterion. The fitness used to select a candidate must come from that candidate acting under these conditions.

A fixed neural representation can make the first experiment concrete. For example, choose two hidden layers of 16 units and compute h1 = tanh(W1 o + b1), h2 = tanh(W2 h1 + b2) and a = tanh(W3 h2 + b3). Here o has 24 entries, h1 and h2 have 16, and a has four entries in the permitted interval. The matrices have shapes 16×24, 16×16 and 4×16; the biases have lengths 16, 16 and 4. Concatenate their 740 scalar parameters in a fixed declared order to obtain a candidate vector theta. Decoding must restore that same order and those same shapes. This is an explicit small architecture for the example, not a claim that these widths are optimal or the widths used in the reported book experiment.

For one trial, decode theta, reset the environment and the controller’s episode state, and obtain the first observation. Compute an action, let the environment perform it, accumulate its returned reward and continue from the new observation until the defined episode ending. The return is the sum of those rewards together with the identity of theta and the trial conditions. For repeated trials, reset afresh and compute the selected aggregate, such as mean return. The weights of this fixed controller remain unchanged during each episode. A recurrent or learning controller needs a different, explicit rule for which state or weights may change and when they reset.

An elementary search needs no hidden optimizer step. Begin with a parameter vector m, a positive perturbation scale sigma, a finite batch size and a trial budget. Sample independent vectors epsilon_i of standard normal values and form theta_i = m + sigma epsilon_i. Try the incumbent m and these proposals on the same newly sampled set of terrains; using shared trial seeds can reduce differences caused only by which terrain each received. Retain a tested vector attaining the highest mean return as the next m. Retain the incumbent when it is among those maximizers; otherwise choose the first maximizing proposal in generation order. Repeat while the remaining budget and receiving need justify another batch. This mutation-and-selection construction is a usable baseline. Its fixed scale and finite trials can miss useful changes; selecting the highest measured mean does not prove improvement on unseen terrain.

For the book’s CMA-ES alternative, replace that proposal and update operation with the complete algorithm in Hansen’s tutorial, Appendix A and Figure 6. Its generated parameter vectors enter the same trial operation. Return each vector’s evaluation to the matching distribution update, including the previous covariance, evolution paths and step-size adaptation. The book’s covariance illustration explains one contribution but is insufficient to implement the complete update. If the chosen implementation minimizes its objective, supply the negative of the return being maximized, consistently with its selection and stopping rule. The new distribution mean and the best evaluated controller are different objects: a mean that has not been tried is another candidate, not an established replacement for a tested controller.

Retain the selected evaluated vector with its decoder, input and action interpretation, any normalization and state-initialization rule. Apply that retained controller in further trials answering the receiving question. This closes the connection: numerical description → acting network → experienced trajectory → measured return → changed proposal or retained controller → further-use evidence. An optimizer checkpoint or a high score alone cannot perform the next walk.

The environment’s actual semantics can change this connection. In the Gym v0.21 implementation, reward uses change in position-and-angle shaping, a penalty linear in clipped absolute action magnitude and a fall penalty. The book’s descriptions of a squared-torque penalty, a constant alive bonus and absolute body position among the observations do not describe that implementation. Its Figure 3.1 also calls reward end-only, whereas the rollout sums step rewards. A policy cannot use a position that it never observes, and changing the reward changes which behavior the search favors.

For an implementation, obtain the actual reset, action and episode-ending interface from the chosen environment version. The Gymnasium task documentation and implementation provide that return; its interface distinguishes termination from truncation. The printed book listings also need repair of their step unpacking and inconsistent fitlist/fitness_list name. Use the operations above to construct the implementation and its checks; verify the selected software before relying on its execution.

Use failure to change the relevant contribution

In the book’s experiment, a controller selected on a single rollout can win because its terrain was easy. Testing it on 100 further trials exposes weaker average performance. Averaging 16 rollouts during selection changes the evidence that drives development; the later 100-trial examination answers a different question about the selected controller. More representative trials can reduce accidental selection, while consuming more interactions. The full cost also depends on episode lengths, candidates, updates and computing arrangement, so the repeat count alone does not establish a 16-fold ratio of total cost or a universal comparison with another learning method.

Suppose the required use includes slippery terrain that the selection trials omitted. Add or obtain representative trials of that condition and compare the controller and sufficient alternatives there. If two situations require different actions but give the controller indistinguishable observations, more trials alone cannot supply the missing distinction. Obtain a useful observation, retain relevant history through memory, or change the promised behavior. If the information is available but the selected representation cannot express the needed response, change the network or its generator. If an adequate response is expressible but the chosen perturbations rarely reach it, compare a different changing operation, initialization or learning arrangement. C.40:4.11 joins that representation choice to the actual changes it makes attainable.

A rare-failure constraint requires its own examination; a satisfactory mean does not settle it. Likewise, performance in a simulator must be related to the sensing, actuation and conditions of a physical robot before it supports that use. Chapter 6’s transfer constructions connect preparation of these conditions, training and actual transfer. Return to the implicated relation when a result fails, while retaining what the old trials still establish.

Some intended behaviors need random action at use time. Against an opponent who exploits a predictable response, a network can instead produce the parameters of a distribution and the agent actually samples its next action. For independent Gaussian action components, require positive standard deviations; if their vector is sigma, covariance is diag(sigma²), with an explicit treatment of action bounds. Evaluate the resulting stochastic policy under its own variability. Random mutations while developing a deterministic policy do not supply this behavior. Book §3.2.5 gives this branch; the receiving task decides whether it is useful.

Compare complete ways of obtaining the behavior. Lack of correct target actions does not preclude learning from delayed reward or estimating a gradient of expected return. Salimans and colleagues, §§2–3, explain an evolutionary strategy that estimates a gradient for the expected performance under parameter perturbations. Its distribution, rank transformation and parallel trials determine what signal and cost it uses. A direct controller, reinforcement learning, population search or a hybrid remains eligible under the same receiving test. Choose by obtainable information, reachable changes and complete cost, rather than by a promise that one method always escapes local optima.

Change the representation and retain useful alternatives

When changing a flat vector cannot express the needed structure, an inherited description can specify a graph or a generator of networks. The NEAT construction, book §3.3, connects historical correspondence of genes, protection of new structures, time for their weights to adjust and selection of useful further development. Stanley and Miikkulainen’s primary account supplies the actual mutation, alignment and species operations. Aligning genes permits a defined crossover; it does not guarantee that the parents’ useful behaviors survive in the offspring. The resulting network still has to act in the receiving trial.

An indirect encoding inserts another performed operation: inherited material generates a network or developing system, which then produces the behavior being judged. Chapter 4 develops spatial generators, developmental processes and input-dependent constructions. Choose the generator together with the changes it permits. A small inherited edit may alter many connections; a regular generator can make a coordinated change easy while preventing a useful local exception. Return to the generated behavior and its attainable variations when deciding whether to retain, extend or replace the encoding. Compactness alone establishes neither interpretability nor useful future change.

If one best-so-far controller destroys access to materially different continuations, retain an executable collection. In MAP-Elites, evaluate a candidate’s task quality and behavior descriptor, place it in the corresponding region of a defined archive, retain the better candidate for an occupied region, and generate further candidates from retained material. The descriptor decides which differences survive this local competition. Keep each stored controller and the conditions needed to execute it; a plotted dot is insufficient for renewed search or use. Chapter 5 explains related choices, including novelty, local competition, ensembles and reuse of acquired material. C.40:4.12 connects those differences to a consequential later use rather than treating diversity as an automatic benefit.

Uncertain measurements can change which candidate deserves a place, and changing the descriptor changes the archive’s distinction. Re-evaluate the relevant stored material when that difference matters. If the receiving question asks for one adequate controller, retaining a large repertoire may add needless work. If it asks for alternatives across energy, speed or another objective, bounded multiobjective retention is a different construction: NSGA-II selects among nondominated fronts and can discard members when capacity is exhausted. Neither an archive nor a nondominated set proves that every retained controller is suitable for deployment.

Develop the way that learns or obtains a result

A candidate can describe how to obtain a controller instead of describing only its final weights. For architecture search, construct the proposed network, initialize and train it under the declared procedure, then use its resulting behavior and cost to judge the architecture. Chapter 10 develops this whole, including reusable modules and task-specific combinations. Chapter 11 changes other obtaining contributions, such as loss, activation, training-data use and learning code. Comparing untrained architectures does not substitute for the training whose result the receiving question needs. A shortened trial or surrogate is useful only at the conclusions its relation to that full application supports.

State what one candidate may inherit. Continuing from trained weights asks whether that continuation works; initializing afresh asks whether the proposed obtaining way works from that start. For a controller that learns during a lifetime, execute that learning across the specified experiences before assessing its result. Reset, permitted feedback, memory and inherited initial state are parts of the construction. Chapter 12 develops evolutionary/RL combinations, learnability and plasticity. In the primary Baldwin-effect meta-learning construction, lifetime learning affects fitness while the evolved starting conditions are what reproduction preserves. The update or learner remains an actual operation to implement or obtain, not a label attached to a successful final network.

These alternatives can also exchange useful material. A population can provide experiences for a gradient learner; an improved learner can return a policy to the population. The evolution-guided policy-gradient construction supplies that specific exchange. Preserve how experiences are collected, which parameters are updated and how a returned policy enters further comparison. Copying a score between the two procedures would not supply the exchanged experience or behavior. C.40:4.7 keeps the candidate way, its real application and the receiving result connected across these cases.

Make interacting and continuing behavior work as a whole

When a controller is assembled from separately developed components, test the actual combinations from which each component receives credit. Chapter 7 develops cooperative neural components, teams, adapting opponents and local cellular rules. A component that works with one partner may fail with another. Construct the partner selection, shared observations, action interface and credit relation before interpreting its fitness. C.40:4.5/.6 supplies the general combination and adaptive-trial connections. For competition, retain appropriate earlier or alternative opponents when a victory over the current one could conceal lost ability; a changing opponent also changes the meaning of the comparison.

If successful joint action depends on information available only to one partner, signaling and a shared convention may be needed. When each actor already observes enough to perform the task, communication can be unnecessary. Inherited coordinated responses and a code acquired from partners require different constructions. In Li and colleagues’ learned-communication construction, retain older agents who carry the acquired code, and let newborns learn through actual rewarded interaction with them. The newborns inherit parameters governing learning: memory size, reward discount, decay and the threshold for fixing an established response. Their acquired policy maps and event memories start empty; the parents’ learned maps are not copied into them.

Perform the generation in order: newborns first interact with both parents; then the population socializes across pairs; rewards accumulated during socializing rank agents for survival. In this construction, the 25 best agents whose lives span fewer than four generations become the next seniors. They retain their acquired responses and provide parents and partners for new learners. For the actual associative update, return to the source’s Real Time Learning and Evolution: a learner unit maps input patterns to activation parameters, records input-output events and changes the corresponding parameters from discounted reward, while potentiation protects established responses from unstable newcomer feedback. Its population, trial and reproduction rules supply the conditions in which that update obtains a shared code. Keep this performed learning and the surviving carriers alongside the inherited learning parameters. If the partners, informative feedback or learning operation are unavailable, obtain that missing contribution before claiming transmission; a description of the learner cannot supply a population carrying a convention.

A local update rule offers another whole: initialize a cellular state, repeatedly apply the rule using its permitted neighborhood information, obtain a larger structure or functioning system, and assess its behavior. Vary continuation time or apply a disturbance when persistence or recovery matters. A snapshot resembling the target can be transient; successful recovery can depend on information or actuation absent after a different injury. Retain the rule and initialization needed for renewed growth, or the live state needed for continuation, according to the actual receiving task. The source’s neural cellular-automata sections specify those operations; the desired whole is not obtained by merely naming its cells.

Body changes can invalidate acquired skills. The ESP extension gives a concrete construction: change morphology together with a new body-affecting skill, re-evaluate the protected older skills, reject excessive losses, then hold the selected morphology fixed while older controllers adapt to it. Its prescribed syllabus and skill interfaces supply the dependencies; further skill composition uses the resulting body. This can obtain a new capability while retaining needed earlier ones, at the cost of those repeated trials. Chapter 9 opens related questions about reachable further development, coupled environments and solutions, and changes of organization. A richer successor or a finite recovery result supports its tested continuation; it does not establish unlimited innovation.

For continued discovery, inspect what a retained basis can actually develop into. Try a feasible further change, obtain its quality and difference, and use those results to decide which basis merits continuation. Preserving current behavior and preserving future possibilities can lead to different choices. C.40:4.11/.12 explains that general return; the neural encoding, body, challenge generator or interaction supplies the particular possibilities. If an adequate current result serves the work, continued discovery can remain a separate purpose.

Use human or generated contributions in the receiving operation

A person can change task conditions, demonstrate a behavior, choose among candidates or propose a useful alteration. Chapter 8 explains interactive development, branching and preparation of material people can meaningfully judge. Identify what their response supplies and perform the resulting change. A selected image may identify an interesting branch; it does not yet supply the working controller behind that image. Prepare informative comparisons and preserve access to the chosen material, then try the resulting candidate. C.40:4.9/.10 connects comparison and human contribution to this actual continuation.

Generated material can enter as the candidate, a changing operation, training experience or an environment. In Chapter 13, these placements lead to different constructions. For example, a language model can propose code, the code is executed under a specified test, and that result selects or informs the next proposal. The execution and test provide a way to reject fluent but ineffective output. A generated training set instead has to train a learner whose further performance judges the generator; visual plausibility is not that performance. Retain the generator or prompt, generated material, actual receiving operation and its returned result at their respective roles.

A world model adds a further distinction. In Ha and Schmidhuber’s recurrent world-model construction, observations train a compact predictive representation, a controller can develop using the learned dynamics, and its behavior returns to the actual task environment for examination. Search can exploit errors in a learned model, so simulated success leaves that receiving test consequential. The Dreamer 3 construction instead combines continued real interaction, learning a world model and learning behavior through imagined trajectories. The learning operations and information flow differ; calling both a world model does not make their controllers or guarantees interchangeable. C.40:4.8 supplies the general prediction–policy–application return, while the selected source supplies the concrete learner and its update.

Distinguish useful behavior from an explanation of its origin

The same trial construction can investigate a mechanism rather than deliver a controller. In the primary hyena-mobbing study, controllers choose to approach a lion or wait. They receive five binary indicators of location, sufficient company and whether mobbing has occurred. The environment supplies movement and the rule that four hyenas inside the interaction circle can mob the lion; death remains probabilistic, with greater danger before sufficient company arrives. Random groups interact, their obtained rewards guide evolutionary selection, and the study examines the resulting populations and combinations. These supplied conditions are part of what the model can explain.

This model has no explicit emotional state or signaling channel. Its successful coordination therefore supplies neither of those mechanisms. Homogeneous copies of a successful controller and heterogeneous groups can also produce different results; preserve the grouping when interpreting individual fitness. A proposed further experiment could clamp the sufficient-company indicator to zero, preserving the five-input interface, and examine whether the acquired coordination depends on that cue. That intervention would test the implemented system. Biological use, discussed in chapter 14, additionally needs a justified correspondence between its conditions, mechanisms and observations and the population being explained. MMP.15 examines what an intervention claim follows from; MMP.16 and PHY.10 help construct an observation separating consequential alternatives. A different natural mechanism producing the same visible behavior leaves the origin question open.

A sufficient direct mechanism remains valuable here. Ijspeert and colleagues’ salamander study constructs a coupled-oscillator model from physiological hypotheses, connects it to a robot, changes drive to obtain swimming, walking and transitions, and compares its consequences with biological observations. The book’s attribution of an evolved network to this particular 2007 construction is mistaken. Direct model construction and its discriminating comparison preserve the useful research move. An engineered transition, including one found by evolutionary search, does not by itself establish the historical origin of the natural transition.

The change of receiving question determines what to retain. Engineering needs the reproducible behaving construction and evidence for its conditions of use. Learning-method development needs the obtaining operation, its starting conditions and its resulting learner. Mechanistic inquiry needs the competing account, correspondence, intervention and observation that bear on its claim. Some experiment results can contribute to more than one use, but their grounds must actually support each conclusion.

Count the whole cost and return a usable result

Include generating and decoding candidates, trials and resets, learning within a trial, optimizer updates, repeated assessment, retained alternatives, communication and memory. EvoJAX explains acceleration of the connected optimizer–policy–task computation; speeding neural inference alone can leave the simulator as the limiting contribution. Deep GA reconstructs parameters from initialization and mutation seeds, reducing stored or transmitted material while adding reconstruction work. Both are conditional implementation choices, not substitutes for informative trials or a suitable representation.

Return the result the next use needs: a tested controller and its interpretation; a learning way with its reset and inheritance rules; executable retained alternatives; or a supported mechanistic conclusion with its remaining ambiguity. Include the condition that would change the next action. If the only missing contribution is an environment adapter, request that adapter with its observation, action and episode-ending behavior. If it is an unexplained specialist operation, return to the linked source or obtain that contribution. Stop with that identified need when it cannot be supplied. The recipient should not have to reconstruct the connection from pattern names or infer a performed experiment from an article describing one.