MMP.8.SD:5 - Archetypal Grounding
The following are constructed finite models. Their numbers show what follows from the stated assumptions; they are not measured effects or recommendations for a particular organization.
MMP.8.SD:5.1 - Reserve a resource for an observed later request
A portable power unit has enough energy to serve one task in either of two slots, with no recharge. Serving the current task in slot 1 yields 4 units on an already selected gain scale. In slot 2 a different task arrives with probability 0.6 and yields 10 if served. Arrival is observed before the slot-2 decision. Unused energy has terminal value zero; the current task cannot be deferred. Serving consumes the whole unit.
Let b be remaining energy, either 0 or 1, and d indicate the later arrival. At slot 2,
V_2(b,d) = 10*b*d.
Use the retained state (time, remaining energy, current request). The slot-1 comparison is
serve now: 4 + 0 = 4
reserve: 0 + 0.6*10 + 0.4*0 = 6.
The resulting policy reserves energy, then serves if the request arrives. It has larger expected gain than immediate service, although it yields zero when no request arrives. If the requirement is a gain of at least 4 in every admitted case, serving now meets it and reservation does not.
Two histories can show the same later request but differ in whether the energy was already used. Merging them would make the slot-2 service appear available after both histories. Retaining b prevents that error.
Suppose the arrival probability changes to 0.3, with the other premises unchanged. Reservation now gives 3 and immediate service gives 4, so the first action changes. If instead the only available probability statement is 0.55≤p≤0.65, reservation gives expected gain between 5.5 and 6.5 and exceeds 4 throughout. Refining p is unnecessary for that comparison.
MMP.8.SD:5.2 - Pay for information only when a later choice can use it
An unfamiliar encoded document uses format A or B, initially with equal probabilities. Two supplied decoders are available: the matching decoder gives a usable output worth 10, and the other gives 0. The model permits one final decoder application. An optional diagnostic costs 1 on the same gain scale and leaves the format unchanged.
The diagnostic returns + with probability 0.8 in format A and 0.2 in format B; the complementary probabilities give −. Its result arrives before decoder selection. The supplied observation model and A.3.3.PI’s conditioning give:
| Information before decoder choice | Probability of format A | Best decoder | Expected final gain |
|---|---|---|---|
| No diagnostic | 0.5 | A or B | 5 |
| Diagnostic + | 0.8 | A | 8 |
| Diagnostic − | 0.2 | B | 8 |
The decision state can be (stage, p), where p is the current probability of A. At the final stage,
V_2(p) = max(10*p, 10*(1-p)).
Each diagnostic result has probability 0.5 under the initial model. Thus
skip diagnostic: V_2(0.5) = 5
use diagnostic: -1 + 0.5*V_2(0.8) + 0.5*V_2(0.2) = 7.
The continuation’s choice changes with the observation. Fixing one decoder in advance gives 5 before the diagnostic cost, or 4 after it. This is why the extra information has value here.
Now the diagnostic result is delayed until after the final decoder choice. The decoder choice can no longer depend on that result. Using the diagnostic then has net expected gain 4, so skipping it gives the better value 5. This repair changes the information timing rather than the arithmetic of conditioning.
A different changed condition also reverses the comparison: with a timely but weaker symmetric diagnostic that is correct with probability 0.55, using its result gives -1+10*0.55=4.5. Skipping still gives 5 and is preferable.
The decoder and diagnostic behavior are premises of this example. A real use requires the applicable processing and observation methods to supply those consequences.
MMP.8.SD:5.3 - Retain accumulated progress when the terminal goal changes
In a two-round game, all awarded points are observed immediately. The first move is either A, giving 2 points with probability 0.5 and 0 otherwise, or B, giving 1 point certainly. In the last round, the player chooses Safe, adding 1 point certainly, or Gamble, adding 3 with probability 0.4 and 0 otherwise. The gamble’s outcome is independent of the first round.
The objective is initially to maximize the probability of finishing with at least 3 points. Let c be points already earned. The terminal contribution is g(c)=1 when c≥3 and 0 otherwise. With zero intermediate contribution, the continuation compares terminal attainment:
| Points before last round | Safe: attainment probability | Gamble: attainment probability | Best continuation |
|---|---|---|---|
| 0 | 0 | 0.4 | Gamble |
| 1 | 0 | 0.4 | Gamble |
| 2 | 1 | 0.4 | Safe |
Therefore first move A gives 0.5*1+0.5*0.4=0.7, while B gives 0.4. Choose A and condition the last move on earned points. Retain the state (round, c). The round number alone is insufficient: histories with c=0 and c=2 require different continuations.
If the objective were expected final points, Gamble’s expected increment 1.2 exceeds Safe’s 1 at every c. Both first moves have expected increment 1, so both then give 2.2 expected final points. That calculation answers a different question from attainment probability.
Now change the requested terminal target from 3 to 4 points. From c=2, Gamble attains it with probability 0.4 and Safe fails. From c=0, neither last move can attain it. Thus A gives 0.5*0.4+0.5*0=0.2. B followed by Gamble gives 0.4 and becomes preferable. The transition probabilities remain unchanged; the changed terminal criterion alters both the continuation and the first choice.
MMP.8.SD:5.4 - Preserve one unknown condition across stages
A two-stage processing route incurs costs w and 1-w, where an unknown w is fixed for the whole run and belongs to {0,1}. Choosing that route commits to both stages. A supplied alternative has total cost 1.5. The question is to minimize worst total cost, with no probability model.
The two possible cost sequences for the route are (0,1) and (1,0); both total 1. The route therefore beats the alternative. Adding the worst first-stage cost 1 to the worst second-stage cost 1 gives 2, but combines different possible runs.
For the first choice, the common total 1 already suffices. If the remaining cost is later needed and the first-stage cost is observed, that observation identifies w and determines the remaining cost.
Now the operating condition is allowed to change between stages: costs are w_1 and 1-w_2, with all four pairs (w_1,w_2) allowed. The pair (1,0) gives total cost 2. The route’s worst total is now 2, so the fixed-cost alternative 1.5 is preferable. The old conclusion fails because the admitted dependence changed.