Library / Foundational Thinking DPF Suite Reference
Jump to passage
In this reading

Link to current text

Published source confirmed at last check

Source changed 2026-10-03 02:22:15 UTC · snapshot created 2026-10-03 03:38:22 UTC · last check 2026-10-03 04:20:18 UTC

Obtain a controller through a complete small search

Start with the behavior and the means of trying it. In the book’s walking case, §3.2, the task is simulated locomotion over uneven terrain. The environment returns 24 observations and accepts four bounded motor commands. Specify the simulator version, terrain conditions, episode ending, observations, action interpretation and receiving criterion. The fitness used to select a candidate must come from that candidate acting under these conditions.

A fixed neural representation can make the first experiment concrete. For example, choose two hidden layers of 16 units and compute h1 = tanh(W1 o + b1), h2 = tanh(W2 h1 + b2) and a = tanh(W3 h2 + b3). Here o has 24 entries, h1 and h2 have 16, and a has four entries in the permitted interval. The matrices have shapes 16×24, 16×16 and 4×16; the biases have lengths 16, 16 and 4. Concatenate their 740 scalar parameters in a fixed declared order to obtain a candidate vector theta. Decoding must restore that same order and those same shapes. This is an explicit small architecture for the example, not a claim that these widths are optimal or the widths used in the reported book experiment.

For one trial, decode theta, reset the environment and the controller’s episode state, and obtain the first observation. Compute an action, let the environment perform it, accumulate its returned reward and continue from the new observation until the defined episode ending. The return is the sum of those rewards together with the identity of theta and the trial conditions. For repeated trials, reset afresh and compute the selected aggregate, such as mean return. The weights of this fixed controller remain unchanged during each episode. A recurrent or learning controller needs a different, explicit rule for which state or weights may change and when they reset.

An elementary search needs no hidden optimizer step. Begin with a parameter vector m, a positive perturbation scale sigma, a finite batch size and a trial budget. Sample independent vectors epsilon_i of standard normal values and form theta_i = m + sigma epsilon_i. Try the incumbent m and these proposals on the same newly sampled set of terrains; using shared trial seeds can reduce differences caused only by which terrain each received. Retain a tested vector attaining the highest mean return as the next m. Retain the incumbent when it is among those maximizers; otherwise choose the first maximizing proposal in generation order. Repeat while the remaining budget and receiving need justify another batch. This mutation-and-selection construction is a usable baseline. Its fixed scale and finite trials can miss useful changes; selecting the highest measured mean does not prove improvement on unseen terrain.

For the book’s CMA-ES alternative, replace that proposal and update operation with the complete algorithm in Hansen’s tutorial, Appendix A and Figure 6. Its generated parameter vectors enter the same trial operation. Return each vector’s evaluation to the matching distribution update, including the previous covariance, evolution paths and step-size adaptation. The book’s covariance illustration explains one contribution but is insufficient to implement the complete update. If the chosen implementation minimizes its objective, supply the negative of the return being maximized, consistently with its selection and stopping rule. The new distribution mean and the best evaluated controller are different objects: a mean that has not been tried is another candidate, not an established replacement for a tested controller.

Retain the selected evaluated vector with its decoder, input and action interpretation, any normalization and state-initialization rule. Apply that retained controller in further trials answering the receiving question. This closes the connection: numerical description → acting network → experienced trajectory → measured return → changed proposal or retained controller → further-use evidence. An optimizer checkpoint or a high score alone cannot perform the next walk.

The environment’s actual semantics can change this connection. In the Gym v0.21 implementation, reward uses change in position-and-angle shaping, a penalty linear in clipped absolute action magnitude and a fall penalty. The book’s descriptions of a squared-torque penalty, a constant alive bonus and absolute body position among the observations do not describe that implementation. Its Figure 3.1 also calls reward end-only, whereas the rollout sums step rewards. A policy cannot use a position that it never observes, and changing the reward changes which behavior the search favors.

For an implementation, obtain the actual reset, action and episode-ending interface from the chosen environment version. The Gymnasium task documentation and implementation provide that return; its interface distinguishes termination from truncation. The printed book listings also need repair of their step unpacking and inconsistent fitlist/fitness_list name. Use the operations above to construct the implementation and its checks; verify the selected software before relying on its execution.