Library / Mathematical Modeling DPF
Jump to passage
In this reading

Link to current text

Published source confirmed at last check

Source changed 2026-10-03 11:52:20 UTC · snapshot created 2026-10-03 11:53:41 UTC · last check 2026-10-03 14:35:13 UTC

MMP.8.SD:4.4 - Derive current gain plus continuation

First evaluate a policy that selects its actions from the retained state. With a sufficient state, let V_t^π(z) mean expected gain from time t onward when the policy is followed, conditional on current retained state z. For a deterministic policy, the finite model gives

V_(T+1)^π(z) = g(z)
V_t^π(z) = r_t(z, π_t(z))
           + sum_z' P_t(z' | z, π_t(z)) V_(t+1)^π(z').

Here g(z) is the conditional expected terminal contribution, r_t(z,a) the expected current gain, and P_t the next retained state’s law derived from :4.2–:4.3. A randomized rule averages the right-hand side over its action probabilities.

The equation follows by splitting the accumulated gain into the current term and the remaining terms, then conditioning on the next information state. Thus each action is compared with what can follow it, including the information then available.

For a finite state and action problem with nonempty allowed action sets, whose only policy restrictions are those local sets, the best attainable value satisfies

V_(T+1)(z) = g(z)
Q_t(z,a) = r_t(z,a) + sum_z' P_t(z' | z,a) V_(t+1)(z')
V_t(z) = max over a in A_t(z) of Q_t(z,a).

An attaining action at each reached state defines an optimal policy for this model and criterion. Finite sets make these maxima attainable. For more general spaces, determine the relevant existence, measurability and integrability conditions. When a maximum is not attained, distinguish the supremum from any obtained approximate policy.

The order of choice and averaging matters. When a later observation is available before the next action, its branch can use its own continuation. A fixed action sequence cannot use that observation. At a common decision node, selecting a different action for each still hidden state would add unavailable information.

Retain restrictions coupling choices across histories or times. A total resource limit can often be represented by remaining resource in Z; a constraint on the whole policy may instead require a constrained formulation. Independently maximizing every node can violate a coupling that the state has omitted.

These equations define the mathematical continuation problem. CMP.3 supplies sharing and scheduling of repeated subcomputations; CMP.9 supplies a sampling procedure and the error of estimated expectations; CMP.5 can supply a relaxation, a usable policy and an improvement bound when its recovery conditions hold. C.29.2 supplies the wider computational formulation, including a continuing response rather than a terminating answer.

For indefinitely continuing use, first select a finite total, discounted sum, average rate or other well-defined criterion. For example, bounded per-step gains and a discount factor γ with 0≤γ<1 make the infinite discounted sum finite. A fixed-point or limiting equation then needs conditions appropriate to that criterion. A solver’s convergence and a policy’s consequences remain separate questions.