Library / First Principles Framework (FPF) - Core Conceptual Specification
Jump to passage
In this reading

Link to current text

Published source confirmed at last check

Source changed 2026-10-03 05:29:54 UTC · snapshot created 2026-10-03 05:30:57 UTC · last check 2026-10-03 05:35:10 UTC

E.23:11 - SoTA-Echoing

SoTA here means the current best-known problem-solving practice for the stated question, not the newest, most official, or most familiar source. The comparison below is current to 2026-08-19. Lineage, bounded current research, internal governing dependencies, and rejected transfers are named as such.

Practice questionExact source and statusSelected payload and limitSource-use decision, receiving locus, qualification, and reopen
What must a repeated improvement pass expose so that it produces learning rather than merely naming a cycle?Gerald Langley et al., The Improvement Guide, 2nd ed. (2009), retained Model-for-Improvement lineage; Michael Taylor et al., Systematic review of the application of the plan-do-study-act method to improve quality in healthcare, BMJ Quality & Safety 23 (2014), DOI 10.1136/bmjqs-2013-001862; Julie Reed and Alan Card, The problem with Plan-Do-Study-Act cycles, BMJ Quality & Safety 25 (2016), DOI 10.1136/bmjqs-2015-005076; D. Royce Sadler, Formative assessment and the design of instructional systems, Instructional Science 18 (1989), DOI 10.1007/BF00117714; John Hattie and Helen Timperley, The Power of Feedback, Review of Educational Research 77(1) (2007), DOI 10.3102/003465430298487. The last two are retained formative-feedback lineage.The sources contribute aim, explicit measures, prediction or proposal, tested change, comparison with the prior result, learning, and the connection among desired condition, current condition, and next move. Their healthcare and education evidence establishes neither a universal lifecycle nor FPF ontology or improvement.Adapt as lineage — reason: the common learning structure changes the E.23 action, while the named domain methods do not dominate current cross-domain practice. Receiving loci: E.23:4.1 steps 1, 3, 7–10; Affordable floor evaluation; Pattern exceptional improvement. Qualification/currentness: retained for this structure, not as present-front authority. Reopen: comparative evidence overturns the learning structure, or E.23 changes its proposal, re-evaluation, or stop action.
When does a specialized improvement-cycle family fit better than a general adaptive loop?Theory of Constraints Institute, Five Focusing Steps (https://www.tocinstitute.org/five-focusing-steps.html), living institutional explanation of POOGI; John R. Boyd, The Essence of Winning and Losing (1996 briefing), retained OODA lineage.POOGI contributes constraint selection, throughput-shaped improvement, and attention to inertia after a constraint shifts. OODA contributes orientation and feedback under changing external conditions. Neither source makes every quality problem a constraint or makes loop speed, cadence, or action volume a quality result.Adapt as lineage — reason: each branch supplies a useful selection discriminator without supplying a universal method. Receiving loci: POOGIFamily and OODAFamily in E.23:4.4. Qualification/currentness: living explanation and historical lineage, not current SoTA for all improvement. Reopen: either family is used outside its stated discriminator, or current comparative practice supplies a better branch at comparable effort.
Which repeated-agent and harness mechanisms warrant bounded use rather than one generic “agent loop”?Thoughtworks Technology Radar Vol. 34, Ralph loop (2026-04-15, Assess), current external-technique signal; Ralph CLI, Ralph loop (https://ralph-cli.dev/docs/core-concepts/ralph-loop/), and Wiggum.dev, The Loop (https://wiggum.dev/concepts/the-loop/), implementation/rationale sources; Noah Shinn et al., Reflexion (arXiv:2303.11366), Aman Madaan et al., Self-Refine (arXiv:2303.17651), Shunyu Yao et al., ReAct (arXiv:2210.03629), Andy Zhou et al., LATS (arXiv:2310.04406), and John Yang et al., SWE-agent (arXiv:2405.15793), retained stepping stones; Boyuan Wang et al., Harnesses for Inference-Time Alignment over Execution Trajectories (arXiv:2605.21516), Wenze Wang et al., A Physical Agentic Loop for Language-Guided Grasping with Execution-State Monitoring (arXiv:2604.07395), and Roxana Geambasu et al., Engineering Robustness into Personal Agents with the AI Workflow Store (arXiv:2605.10907), current 2026 preprints in distinct settings.The combined branch contributes fresh-context work against a specification, feedback memory, action–observation coupling, decomposition versus guided execution, partial-harness and over-structuring limits, bounded physical monitoring/retry/escalation, finite termination, and the flexibility–robustness trade-off of hardened workflows. It does not establish convergence by repetition, a general FPF loop kind, or that more harness is better.Adapt and combine — reason: Thoughtworks supplies a current practice signal, the 2026 papers supply distinct mechanisms and failure limits, and the older papers remain mechanism lineage. Reject as load-bearing current evidence: Ralph CLI and Wiggum documentation remain implementation/rationale only. Receiving loci: RalphLikeGeneralAdaptiveFamily, E.23:4.5, and Agent harness improvement. Qualification/currentness: primary official and arXiv sources checked 2026-08-19; evidence remains setting- and benchmark-bound. Reopen: the Radar posture or cited papers materially change, or comparative evidence changes the mechanism, cost, or stop choice.
When do an extra supervisor or accumulated search memory improve the loop rather than add apparatus?Zeda Xu, Nikolas Martelaro, and Christopher McComb, Supervising Ralph Wiggum / CRDAL (arXiv:2603.24768v2, 2026-05-07), current engineering-design research signal; Yanlong Wang et al., FactorMiner (arXiv:2602.14670v1, 2026-02-16), current financial-alpha-search research signal.CRDAL contributes metacognitive co-regulation when fixation, underexploration, or expensive design mistakes are live. FactorMiner contributes retrieve/generate/evaluate/distill, modular skills, experience memory, and reduced redundant search in a large comparable search space. Neither supports a universal supervisor, financial-alpha objectives, or automatic transfer to every object.Adapt as two domain-bounded branches — reason: the operation cues are useful only when E.23 can name the corresponding risk or comparable search. Receiving loci: E.23:4.5 operation-family selection, E.23:4.6 cost/risk discipline, and Agent harness improvement. Qualification/currentness: current preprints for engineering design and financial alpha, not general improvement evidence. Reopen: either paper changes materially, a comparative source dominates its cue, or E.23 begins using the domain objective as a general quality value.
When may a fixed performer improve by changing one external method-description object?Yifan Yang et al., SkillOpt: Executive Strategy for Self-Evolving Agent Skills (arXiv:2605.23904v2, 2026-05-25), current preprint.SkillOpt keeps the target model fixed while a separate optimizer makes bounded add/delete/replace edits to one external skill document, accepts only held-out improvement, and keeps rejected-edit and optimizer memory separate. Its benchmark results do not transfer automatically to arbitrary Methods, physical systems, or work-facing system-role kinds and assignments.Adapt — reason: the fixed-performer, mutable-object, bounded-edit, held-out-acceptance split directly sharpens an E.23 family without importing the optimizer as FPF method law. Receiving loci: FixedPerformerObjectVersionUnderImprovementOptimizationFamily, BoundedObjectChangeBudget, held-out re-evaluation, and Agent harness improvement. Qualification/currentness: current preprint, limited to its models, benchmarks, skill documents, and harnesses. Reopen: the paper changes materially or stronger comparable evidence changes the held-out acceptance or bounded-edit rule.
How should several quality coordinates and OEE/NQD alternatives be compared without turning one score or algorithm into authority?Xi Lin et al., Quality-Diversity Optimization as Multi-Objective Optimization (arXiv:2602.00478, 2026), current preprint; Haoxiang Qin et al., A survey on Quality-Diversity optimization: Approaches, applications, and challenges, Swarm and Evolutionary Computation 100:102240 (2026), DOI 10.1016/j.swevo.2025.102240, current survey; Rick Kazman, Mark Klein, and Paul Clements, ATAM (CMU/SEI-2000-TR-004, 2000), retained software-architecture lineage.QD/MOO contributes set-valued multi-coordinate comparison and explicit behavior or descriptor spaces. ATAM contributes scenario-based exposure of quality-attribute trade-offs. These sources do not supply one universal scalar, an FPF archive/front policy, or authority over candidate generation, retention, selection, parity, refresh, or publication.Adapt — reason: set-valued and scenario-visible trade-offs support non-dominated choice while direct neighbours retain their own decisions. Receiving loci: E.23:4.1 steps 8–9, Physical prototype improvement, Three proposals, and NQD quality-side improvement; C.17–C.19/G.5/G.9/G.11 remain the exits. Qualification/currentness: QD sources are current for QD/MOO; ATAM is domain lineage. Reopen: a current QD overview dominates these payloads, or E.23 starts defining a neighbour-owned archive, front, pool, selection, parity, or refresh rule.
When does optimization of a visible measure cease to improve the intended value?Jacek Karwowski et al., Goodhart’s Law in Reinforcement Learning (ICLR 2024), and Thomas Kwa, Drake Thomas, and Adrià Garriga-Alonso, Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification (NeurIPS 2024), current narrow proxy-risk anchors; Charles Goodhart, Problems of Monetary Management: The U.K. Experience (1975); Donald T. Campbell, Assessing the Impact of Planned Social Change, Occasional Paper 8 (1976); David Manheim and Scott Garrabrant, Categorizing Variants of Goodhart’s Law (arXiv:1803.04585, 2018); and Jongwoon Choi, Gary Hecht, and William Tayler, Lost in Translation: The Effects of Incentive Compensation on Strategy Surrogation, The Accounting Review 87(4) (2012), retained monetary, social-indicator, taxonomy, and strategy-surrogation lineage or evidence.The sources distinguish proxy misspecification, optimization pressure, behavior change, heavy-tailed reward error, and strategy surrogation. They do not forbid measurement, predict every failure, or make RL/RLHF results a general project-value model.Adapt and combine — reason: the 2024 papers provide current domain-bounded mechanisms while the older sources preserve distinct cross-domain failure explanations. Receiving loci: protected trade-offs, E.23:4.3 stop, Goodharted pass, and the E.13 exit. Qualification/currentness: the current anchors are narrow RL/RLHF results; older sources remain lineage or bounded experimental evidence. Reopen: stronger current proxy-risk evidence changes a used mechanism, or E.13’s intended-value boundary changes.
When can the evaluation change without silently moving the improvement target?Iacob et al., The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators (2026-06-29), bounded current research; Kim et al., Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization (2025), reward-evaluation research.RQGM separates fixed within-epoch criteria from changes between epochs; Kim et al. shows that optimizing an evaluator proxy need not optimize downstream performance.Adapt: §4.1a keeps one comparison stable and consumes A.19.ECS:4.5 for justified evaluator change. Qualification: source scopes checked 2026-09-30; coding, review and mathematical benchmarks do not prove general convergence, practical framework quality or an FPF cost saving. Reopen: changed task populations, leakage, harmful repairs or cost invalidate the adoption basis.
How should a source-composed improvement claim expose what each source contributes?Joanne McKenzie and Sue Brennan, Cochrane Handbook for Systematic Reviews of Interventions, version 6.5 (2024), Chapter 12, current evidence-synthesis reference; FPF G.2 and G.11, current internal governing dependencies for source-use decisions and source currentness.Chapter 12 contributes disclosure of the selected synthesis method and limitations. FPF supplies the cross-domain decision and currentness rules. Cochrane’s healthcare evidence hierarchy and statistical methods are not imported, and the broad synthesis of cycle, agent, trade-off, proxy, and OEE/NQD lines remains rationale rather than independent current evidence.Adapt Chapter 12 for transparent contribution and limitation reporting; adopt as governing dependencies G.2/G.11; reject healthcare-specific apparatus and citation-count front claims. Receiving loci: SourceComposedResultClaim, E.23:4.7, and the source-use/currentness exits. Qualification/currentness: current reference plus internal rules, not external cross-domain SoTA. Reopen: Chapter 12 materially changes the used reporting rule, G.2/G.11 changes, or a composed claim no longer names its admitted sources and limits.