MMP.8.SD:11 - SoTA-Echoing
The practice question is how to retain enough information for a useful sequential choice without demanding an unnecessarily complete reconstruction.
| Question and comparison | Adopted or adapted contribution and limit |
|---|---|
| Can a compact summary replace the full history tree? | Adopt the information-state line of Subramanian, Sinha, Seraj and Mahajan, Approximate Information State for Approximate Planning and Reinforcement Learning in Partially Observed Systems, JMLR 23 (2022), §§2.2–2.3 and 3.2, especially Theorems 5 and 9. Preservation of expected gain and the next summary’s law supports the reduction in :4.3–:4.4; approximate preservation needs a consequence bound. A full history tree remains preferable when small. This line saves representation and computation only when its conditions are obtainable; prediction fit alone supplies less. Reopen when a newly relevant action, history or criterion changes the equivalence between merged histories. |
| What if compact memory is useful but not sufficient? | Adapt Sinha and Mahajan, Agent-state based policies in POMDPs: Beyond belief-state MDPs, arXiv v1, 24 September 2024, §§II-C–II-D and III. Its comparison of policy classes supports :4.5: evaluate a restricted controller as such, rather than assuming an arbitrary recurrent memory admits the continuation equation in :4.4. Direct policy search is a serious alternative when sufficient-state construction is too costly. Preserve whether its result is locally optimal or best within a restricted policy class. Reopen when memory, available computation or required policy class changes. |
| Does convergence of a modern learning algorithm close the state question? | Reject that inference, using Sinha, Geist and Mahajan, Convergence of regularized agent-state-based Q-learning in POMDPs, arXiv v2, 2 September 2025, §§II-B–IV and Theorem 1. Under its learning-rate and visitation assumptions, the limit is for a regularized model that depends on the behavior policy’s limiting distribution. The result sharpens :4.5’s separation of numerical convergence from the intended policy comparison. Regularized learning remains a computational option; it brings its objective and representation conditions. Reopen when a proposed solver claims a stronger use than its result supports. |
| What if probabilities are unavailable and the criterion is worst consequence? | Adapt Dave, Venkatesh and Malikopoulos, Approximate Information States for Worst-Case Control and Learning in Uncertain Systems, arXiv v2, 6 April 2024, §§II–III. Conditional ranges and a criterion-specific continuation are a serious alternative to expected gain. The paper derives continuations for maximum instantaneous and terminal cost. For accumulated costs, :4.5 and :5.4 preserve the allowed complete trajectories. The trade-off is a worst-case comparison in place of a probabilistic average; reopen when the uncertainty set or its cross-stage dependence changes. |