SYSE.45:11 - SoTA-Echoing
The practice question is which stable contribution can be learned without losing necessary external support. Tool-Internalized Reasoning, ACL 2026, §§3.3–3.5, separates tool-semantics acquisition, trajectory warm-up and tool-reasoning optimization. This Method uses that distinction to diagnose what the learner must acquire; its specialized tokens and reward design remain optional direct implementations, and real-tool coverage and functionally equivalent tool labels limit the reported reach. Skill0 v2 supplies selective guidance withdrawal, with execution retained.
Sample-Efficient Learning from Agent Experience v1, §3, supplies the teacher-at-recorded-history construction above. Its useful target can avoid new interaction during target production, while the recorded successor still belongs to the original action. Generated corrections require qualification; task-specific consolidation and cross-task transfer remain different evidence claims.
Adapt their selective target, rather than treating all tool dependence as a defect. A fixed-model controller or retained guidance is the serious alternative at comparable whole-task cost. The exact learning algorithm and its evidence stay within the selected source edition; a changed interface or a failed untouched use reopens that reliance.