Library / Computational Thinking DPF
Jump to passage
In this reading

Link to current text

Published source confirmed at last check

Source changed 2026-10-03 05:29:54 UTC · snapshot created 2026-10-03 05:30:57 UTC · last check 2026-10-03 06:05:20 UTC

CMP.7:4.2 - Recover the examples and feedback actually available

Identify the supplied inputs, labels, rewards, demonstrations or other observations and how they were obtained. For missing labels, delayed feedback, selection effects or dependence between examples, state what the learning procedure can actually observe. MMP.7 supplies the probability model of recorded data when such a model is needed.

Choose the feedback regime before deriving the update. In supervised learning, the supplied target response can evaluate a candidate prediction directly. In bandit feedback, only the consequence of a chosen action is observed; an update requiring all unchosen consequences is unavailable. A self-produced label remains the output of another rule and can propagate its errors.

State which conditions are fixed during the learning claim. Independent examples from one distribution, an arbitrary sequence generated by one stable target rule, and a changing target support different arguments. Data rows alone do not supply any of these assumptions.