MMP.8.SD:4.1 - Fix the decision times and the consequence being compared
Recover from MMP.8 the choices, uncontrolled circumstances and information available before each choice. Mark the order of observation, action, transition and consequence. Include delays, commitments and stopping opportunities when they change what can be done.
For decisions at times t=1,…,T, let H_t be the available history immediately before action A_t. It contains received observations and known previous actions. Include a realized gain or cost in this history only if the decision maker can know it then. A policy π_t(H_t) selects an allowed action from that history; randomization is an additional allowed operation when the problem permits it.
State the consequence criterion supplied by the receiving use. One common finite model maximizes
J(π) = E^π[R_1 + R_2 + … + R_T + G].
Here R_t is the gain at decision step t, G is a terminal contribution, and the expectation uses the trajectory law induced by policy π and the model. Gains can be negative costs. Their addition presupposes a common meaning and scale. Specify the initial information or distribution under which policies are compared.
An expected sum is one choice of criterion. For a target such as “finish with at least k points,” the consequence can instead be the terminal indicator, equal to 1 when the target is met and 0 otherwise. Its expectation is the attainment probability. For several criteria, retain their comparison or trade-offs until an applicable choice method supplies a selection rule. A scalar “reward” does not decide that rule.
Use a horizon or stopping condition appropriate to the work. Assign the consequences left at the horizon, such as unused stock or unfinished obligations. A discounted sum gives later gains smaller weights; choose that meaning deliberately. Shortening a computational planning window does not make omitted consequences disappear.