Library / First Principles Framework (FPF) - Core Conceptual Specification
Jump to passage
In this reading

Link to current text

Published source confirmed at last check

Source changed 2026-10-03 08:25:59 UTC · snapshot created 2026-10-03 08:26:43 UTC · last check 2026-10-03 08:35:10 UTC

E.22:11.2 - Conditions that change the evaluation question

The following contributions address additional conditions: actionable feedback, protected trade-offs and the reliability of an automated judge. They constrain the question when that condition is present; they do not require every ordinary requester to run those branches.

Practice questionExact source and statusSelected payload and domain limitSource-use decision, changed E.22 locus, qualification, and reopen
A rubric-level evaluation needs its own reliability check rather than trust in one aggregate judge verdict.Tianjun Pan et al., RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following, arXiv:2603.25133 (2026), and Hongli Zhou et al., Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization, arXiv:2603.08091 (2026), are current preprints for automated LLM judging.Pan et al. show that fine-grained rubric judging can remain inaccurate and variable; Zhou et al. test a taxonomy of twelve bias types across generative and discriminative judges. These works concern LLM judges and instruction-following benchmarks; they do not validate an FPF evaluation or generalize their numeric results to physical, medical, or organizational evaluation.Adapt — reason: use rubric-level variability and the twelve-bias taxonomy only to require reliability evidence when an automated LLM judge is selected; neither payload supports a general judge verdict. Changed loci: QualityEvaluationUseDeclaration, ExpectedEvaluationEvidenceBasis@Context, and the Floor evaluation and Exceptional improvement slices. Qualification/currentness as of 2026-08-19: the exact cited 2026 preprints apply only to their automated-judge and instruction-following settings. Reopen: a later benchmark or replication changes either payload, or an E.22 use adds, removes, or materially changes its LLM-judge branch.
Actionable formative feedback distinguishes the desired condition, current performance, and a move that can close the gap.D. Royce Sadler, Formative assessment and the design of instructional systems, Instructional Science 18, 119-144 (1989), DOI 10.1007/BF00117714; John Hattie and Helen Timperley, The Power of Feedback, Review of Educational Research 77(1), 81-112 (2007), DOI 10.3102/003465430298487. Both are historical formative-feedback sources, used here for their standard/current-work/action distinction.Sadler supplies the comparison between a quality standard and current work plus action by the learner; Hattie and Timperley synthesize goal, current progress, and next-step feedback questions. Their classroom evidence does not establish FPF kinds, project authority, or the quality of a proposed repair.Adapt — reason: use the standard/current-gap/action and goal/progress/next-step structures to keep aim, present result, and possible repair distinct; classroom evidence does not validate FPF evaluation. Changed loci: floor and aim bindings in the question frame, proposal/no-proposal result boundaries, and the Absorption slice. Qualification/currentness: these historical sources support the stated feedback structure; they do not establish its superiority to a sufficient direct evaluation or validate it across domains. Reopen: current formative-feedback evidence overturns that structure, or E.22 stops using it to change proposal, no-proposal, or absorption action.
Multi-coordinate improvement needs set-valued alternatives and explicit trade-offs rather than one scalar winner.Xi Lin et al., Quality-Diversity Optimization as Multi-Objective Optimization, arXiv:2602.00478 (2026), current preprint; Haoxiang Qin et al., A survey on Quality-Diversity optimization: Approaches, applications, and challenges, Swarm and Evolutionary Computation 100:102240 (2026), DOI 10.1016/j.swevo.2025.102240, current survey.Lin et al. reformulate QD as a large multi-objective problem and use set-based scalarization; Qin et al. survey high-performing collections over descriptor spaces. These algorithmic results do not assign FPF archive, front, publication, or selection authority.Adapt — reason: use set-valued alternatives and explicit descriptor/coordinate trade-offs, not the QD algorithms or any implied authority, because one scalar winner can hide protected-quality loss. Changed loci: paretoTradeoffEvaluation, TradeoffProtectionSet@Context, CandidateImprovementProposalPortfolio@Context, and the Proposal portfolio and Physical-system proposal slices. Qualification/currentness as of 2026-08-19: the exact cited 2026 preprint and survey apply to QD/MOO optimization. Reopen: current QD/MOO evidence changes the case for set-valued trade-offs, or E.22 begins asserting archive, front, publication, or selection authority.
Optimizing a measure can damage the intended value through several different mechanisms.Charles Goodhart, Problems of Monetary Management: The U.K. Experience (1975), historical monetary-control failure evidence; Donald T. Campbell, Assessing the Impact of Planned Social Change, Occasional Paper 8 (1976), historical social-indicator failure evidence; David Manheim and Scott Garrabrant, Categorizing Variants of Goodhart’s Law, arXiv:1803.04585 (2018), later taxonomy; Jongwoon Choi, Gary Hecht, and William Tayler, Lost in Translation: The Effects of Incentive Compensation on Strategy Surrogation, The Accounting Review 87(4), 1135-1164 (2012), peer-reviewed experimental evidence.Goodhart concerns control that changes an observed regularity; Campbell concerns corruption pressure on social indicators; Manheim and Garrabrant distinguish several overoptimization mechanisms; Choi et al. show managers treating a measure as the strategic construct. None says that every metric is invalid or supplies the intended value automatically.Adapt — reason: use the distinct proxy-failure mechanisms and observed strategy surrogation to ask what worsened and protect the intended value; reject the inference that every metric is invalid. Changed loci: CC-E22-7, CC-E22-8a, the Goodharted improvement repair, and the E.13 relation. Qualification/currentness: the historical accounts, taxonomy and experiment support the named proxy-failure mechanisms, not a universal anti-measure rule. Reopen: evidence overturns a mechanism used here, or E.13’s intended-value and protected-quality test changes.
Automated-judge mitigation is model-dependent and can itself require a declared guarantee or evidence profile.Sadman Kabir Soumik, Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines, arXiv:2604.23178 (2026), current preprint; Benjamin Feuer, Lucas Rosenblatt, and Oussama Elachqar, Towards Provably Unbiased LLM Judges via Bias-Bounded Evaluation, arXiv:2603.05485 (2026), current preprint.Soumik compares nine mitigations and reports model-dependent effects across four bias types; Feuer et al. define average bias-boundedness for specified judge settings. These results are benchmark- and model-bound and do not make any LLM judge generally unbiased.Adapt — reason: use model dependence and the bounded-guarantee form to require a declared reliability evidence profile and qualification, not to call an LLM judge generally unbiased. Changed loci: ExpectedEvaluationEvidenceBasis@Context, EvaluationQualificationWindow, CC-E22-9, and the Exceptional improvement slice. Qualification/currentness as of 2026-08-19: the exact cited 2026 preprints remain benchmark-, model-, and judge-setting-bound. Reopen: later evaluation establishes materially different mitigation transfer or guarantee conditions, or E.22 changes the evidence profile required for an automated judge.
OEE and NQD can use proposal-shaped quality pressure without collapsing proposal, candidate retention, and selection.Xi Lin et al., Quality-Diversity Optimization as Multi-Objective Optimization, arXiv:2602.00478 (2026), current preprint; Haoxiang Qin et al., A survey on Quality-Diversity optimization: Approaches, applications, and challenges, Swarm and Evolutionary Computation 100:102240 (2026), DOI 10.1016/j.swevo.2025.102240, current survey.The shared comparison question is how to preserve several high-performing alternatives across declared coordinates or descriptors. The sources do not say that an evaluation proposal is already a generated candidate, archive insertion, front update, or selected result.Adapt — reason: use the collection-over-coordinates payload only to keep an E.22 proposal portfolio distinct from candidate generation, retention, and selection; it supplies no OEE/NQD authority. Changed loci: CandidateImprovementProposalRow@Context, E.22:4.6, the Proposal portfolio slice, and the C.17–C.19/G.5 relation boundary. Qualification/currentness as of 2026-08-19: the same exact 2026 QD sources apply here only as a bounded OEE/NQD proposal-framing comparison. Reopen: current QD evidence changes the collection/selection distinction, or the direct C.17–C.19 or G.5 consumer boundary changes.