Use failure to change the relevant contribution
In the book’s experiment, a controller selected on a single rollout can win because its terrain was easy. Testing it on 100 further trials exposes weaker average performance. Averaging 16 rollouts during selection changes the evidence that drives development; the later 100-trial examination answers a different question about the selected controller. More representative trials can reduce accidental selection, while consuming more interactions. The full cost also depends on episode lengths, candidates, updates and computing arrangement, so the repeat count alone does not establish a 16-fold ratio of total cost or a universal comparison with another learning method.
Suppose the required use includes slippery terrain that the selection trials omitted. Add or obtain representative trials of that condition and compare the controller and sufficient alternatives there. If two situations require different actions but give the controller indistinguishable observations, more trials alone cannot supply the missing distinction. Obtain a useful observation, retain relevant history through memory, or change the promised behavior. If the information is available but the selected representation cannot express the needed response, change the network or its generator. If an adequate response is expressible but the chosen perturbations rarely reach it, compare a different changing operation, initialization or learning arrangement. C.40:4.11 joins that representation choice to the actual changes it makes attainable.
A rare-failure constraint requires its own examination; a satisfactory mean does not settle it. Likewise, performance in a simulator must be related to the sensing, actuation and conditions of a physical robot before it supports that use. Chapter 6’s transfer constructions connect preparation of these conditions, training and actual transfer. Return to the implicated relation when a result fails, while retaining what the old trials still establish.
Some intended behaviors need random action at use time. Against an opponent who exploits a predictable response, a network can instead produce the parameters of a distribution and the agent actually samples its next action. For independent Gaussian action components, require positive standard deviations; if their vector is sigma, covariance is diag(sigma²), with an explicit treatment of action bounds. Evaluate the resulting stochastic policy under its own variability. Random mutations while developing a deterministic policy do not supply this behavior. Book §3.2.5 gives this branch; the receiving task decides whether it is useful.
Compare complete ways of obtaining the behavior. Lack of correct target actions does not preclude learning from delayed reward or estimating a gradient of expected return. Salimans and colleagues, §§2–3, explain an evolutionary strategy that estimates a gradient for the expected performance under parameter perturbations. Its distribution, rank transformation and parallel trials determine what signal and cost it uses. A direct controller, reinforcement learning, population search or a hybrid remains eligible under the same receiving test. Choose by obtainable information, reachable changes and complete cost, rather than by a promise that one method always escapes local optima.