E.22:11 - SoTA-Echoing
E.22:11.1 - Choosing the lightest sufficient question construction
Practice question. How should a requester obtain an evaluation question that will support the recipient’s actual next use, without adding another description when direct evaluation already suffices? The selected line is conditional: retain an adequate direct question; otherwise recover its missing purpose, object, criterion, grounds or use. Express exact E.22 bindings when they are needed to preserve that question across evaluators and consuming uses.
Two serious alternatives already provide purpose-first construction. Basili et al., Linking Software Development and Business Strategy Through Measurement (2010), especially “Background” and “GQM+Strategies”, links goals, derived measurement questions and interpretation to context and assumptions. It is a historical method still relevant to this comparison; its organizational alignment model is useful when that larger linkage is the question. Kidder et al., CDC Program Evaluation Framework, 2024, Steps 3–5, develops evaluation purpose, intended users and use, questions, design, credible evidence and supported conclusions. It expressly accommodates available resources and revision when needed evidence is infeasible. Their software-organization and program-evaluation settings supply no FPF characteristic values.
Comparison on the same receiving question. Keep the complete M1, reader, studio use, available premises and E.21 criteria from section 5 fixed. The following is a worked adaptation of the methods to that question, not a measurement of practitioner time or error rates.
| Adequate approach | Execution on this case | Sufficient result and burden |
|---|---|---|
| Direct E.21 evaluation of the filled request | The evaluator receives M1, the studio use, nineteen-coordinate floor, expected grounds and complete result form; evaluation starts immediately. | It answers the stated quality question. A separate E.22 frame supplies no additional judgement and adds a description to maintain. Retain this direct execution. |
| Goal–question–measurement construction using the GQM contribution | Starting with “review the studio method”, distinguish the goal of obtaining usable development guidance from the goal of qualifying a recording. Derive questions about that guidance and connect the chosen E.21 observations to the consuming decision. | It can construct the same sufficient question. The work is clarification, selection and interpretation; a full organizational goal/strategy model is warranted only when organizational alignment is also required. |
| Purpose-and-use construction using the CDC contribution | Identify the studio practitioner as the user, the decision about using M1 as the purpose, and the complete pattern as the object. Choose an attainable examination of its guidance and premises, then require conclusions supported by those grounds. | It can also construct the same sufficient question. Use its program-design, engagement and data-collection work when the receiving evaluation requires those contributions; they are not prerequisites of this one pattern-quality question. |
| E.22 construction and, where needed, exact declaration | Resolve the same ambiguity in section 4.0. Keep the question about M1 separate from the conditional preparation, the resulting evaluation and a later recording decision. Bind the declared criterion, scope, expected grounds and return form through sections 4.1–4.3 when those links must be exchanged precisely. | The ordinary request can be the same one. Exact bindings make a changed object, criterion or consuming use inspectable, at the cost of constituting and maintaining those references. They do not supply another evaluation result. |
Choice and accepted trade-off. Adopt the purpose-to-question and use-to-evidence direction shared by the alternatives; adapt it to the FPF object, criterion and result distinctions. Sections 1 and 4.0 put that construction before schemas, section 4.3 retains the direct-evaluation exit and section 5 supplies the filled request and its changed-use boundary. The local addition is an explicit account of those bindings, useful when exchanging a question otherwise risks evaluating another object or carrying a result beyond its scope. Pay its description and maintenance cost only for that needed precision. When an adequate direct request or the rival’s question already preserves the same distinctions, reuse it. Neither a newer explanation nor the presence of a frame establishes better detection, lower cost or a reason to replace a sufficient evaluation.
The comparison selects a bounded FPF adaptation, not a ranking of whole GQM, CDC or FPF frameworks. Reopen when a competing construction preserves the needed question at lower total burden, when an intended reader cannot recover the first request, when a changed source or use defeats these transfers, or when exact binding introduces more maintenance than its receiving use warrants.
E.22:11.2 - Conditions that change the evaluation question
The following contributions address additional conditions: actionable feedback, protected trade-offs and the reliability of an automated judge. They constrain the question when that condition is present; they do not require every ordinary requester to run those branches.
| Practice question | Exact source and status | Selected payload and domain limit | Source-use decision, changed E.22 locus, qualification, and reopen |
|---|---|---|---|
| A rubric-level evaluation needs its own reliability check rather than trust in one aggregate judge verdict. | Tianjun Pan et al., RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following, arXiv:2603.25133 (2026), and Hongli Zhou et al., Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization, arXiv:2603.08091 (2026), are current preprints for automated LLM judging. | Pan et al. show that fine-grained rubric judging can remain inaccurate and variable; Zhou et al. test a taxonomy of twelve bias types across generative and discriminative judges. These works concern LLM judges and instruction-following benchmarks; they do not validate an FPF evaluation or generalize their numeric results to physical, medical, or organizational evaluation. | Adapt — reason: use rubric-level variability and the twelve-bias taxonomy only to require reliability evidence when an automated LLM judge is selected; neither payload supports a general judge verdict. Changed loci: QualityEvaluationUseDeclaration, ExpectedEvaluationEvidenceBasis@Context, and the Floor evaluation and Exceptional improvement slices. Qualification/currentness as of 2026-08-19: the exact cited 2026 preprints apply only to their automated-judge and instruction-following settings. Reopen: a later benchmark or replication changes either payload, or an E.22 use adds, removes, or materially changes its LLM-judge branch. |
| Actionable formative feedback distinguishes the desired condition, current performance, and a move that can close the gap. | D. Royce Sadler, Formative assessment and the design of instructional systems, Instructional Science 18, 119-144 (1989), DOI 10.1007/BF00117714; John Hattie and Helen Timperley, The Power of Feedback, Review of Educational Research 77(1), 81-112 (2007), DOI 10.3102/003465430298487. Both are historical formative-feedback sources, used here for their standard/current-work/action distinction. | Sadler supplies the comparison between a quality standard and current work plus action by the learner; Hattie and Timperley synthesize goal, current progress, and next-step feedback questions. Their classroom evidence does not establish FPF kinds, project authority, or the quality of a proposed repair. | Adapt — reason: use the standard/current-gap/action and goal/progress/next-step structures to keep aim, present result, and possible repair distinct; classroom evidence does not validate FPF evaluation. Changed loci: floor and aim bindings in the question frame, proposal/no-proposal result boundaries, and the Absorption slice. Qualification/currentness: these historical sources support the stated feedback structure; they do not establish its superiority to a sufficient direct evaluation or validate it across domains. Reopen: current formative-feedback evidence overturns that structure, or E.22 stops using it to change proposal, no-proposal, or absorption action. |
| Multi-coordinate improvement needs set-valued alternatives and explicit trade-offs rather than one scalar winner. | Xi Lin et al., Quality-Diversity Optimization as Multi-Objective Optimization, arXiv:2602.00478 (2026), current preprint; Haoxiang Qin et al., A survey on Quality-Diversity optimization: Approaches, applications, and challenges, Swarm and Evolutionary Computation 100:102240 (2026), DOI 10.1016/j.swevo.2025.102240, current survey. | Lin et al. reformulate QD as a large multi-objective problem and use set-based scalarization; Qin et al. survey high-performing collections over descriptor spaces. These algorithmic results do not assign FPF archive, front, publication, or selection authority. | Adapt — reason: use set-valued alternatives and explicit descriptor/coordinate trade-offs, not the QD algorithms or any implied authority, because one scalar winner can hide protected-quality loss. Changed loci: paretoTradeoffEvaluation, TradeoffProtectionSet@Context, CandidateImprovementProposalPortfolio@Context, and the Proposal portfolio and Physical-system proposal slices. Qualification/currentness as of 2026-08-19: the exact cited 2026 preprint and survey apply to QD/MOO optimization. Reopen: current QD/MOO evidence changes the case for set-valued trade-offs, or E.22 begins asserting archive, front, publication, or selection authority. |
| Optimizing a measure can damage the intended value through several different mechanisms. | Charles Goodhart, Problems of Monetary Management: The U.K. Experience (1975), historical monetary-control failure evidence; Donald T. Campbell, Assessing the Impact of Planned Social Change, Occasional Paper 8 (1976), historical social-indicator failure evidence; David Manheim and Scott Garrabrant, Categorizing Variants of Goodhart’s Law, arXiv:1803.04585 (2018), later taxonomy; Jongwoon Choi, Gary Hecht, and William Tayler, Lost in Translation: The Effects of Incentive Compensation on Strategy Surrogation, The Accounting Review 87(4), 1135-1164 (2012), peer-reviewed experimental evidence. | Goodhart concerns control that changes an observed regularity; Campbell concerns corruption pressure on social indicators; Manheim and Garrabrant distinguish several overoptimization mechanisms; Choi et al. show managers treating a measure as the strategic construct. None says that every metric is invalid or supplies the intended value automatically. | Adapt — reason: use the distinct proxy-failure mechanisms and observed strategy surrogation to ask what worsened and protect the intended value; reject the inference that every metric is invalid. Changed loci: CC-E22-7, CC-E22-8a, the Goodharted improvement repair, and the E.13 relation. Qualification/currentness: the historical accounts, taxonomy and experiment support the named proxy-failure mechanisms, not a universal anti-measure rule. Reopen: evidence overturns a mechanism used here, or E.13’s intended-value and protected-quality test changes. |
| Automated-judge mitigation is model-dependent and can itself require a declared guarantee or evidence profile. | Sadman Kabir Soumik, Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines, arXiv:2604.23178 (2026), current preprint; Benjamin Feuer, Lucas Rosenblatt, and Oussama Elachqar, Towards Provably Unbiased LLM Judges via Bias-Bounded Evaluation, arXiv:2603.05485 (2026), current preprint. | Soumik compares nine mitigations and reports model-dependent effects across four bias types; Feuer et al. define average bias-boundedness for specified judge settings. These results are benchmark- and model-bound and do not make any LLM judge generally unbiased. | Adapt — reason: use model dependence and the bounded-guarantee form to require a declared reliability evidence profile and qualification, not to call an LLM judge generally unbiased. Changed loci: ExpectedEvaluationEvidenceBasis@Context, EvaluationQualificationWindow, CC-E22-9, and the Exceptional improvement slice. Qualification/currentness as of 2026-08-19: the exact cited 2026 preprints remain benchmark-, model-, and judge-setting-bound. Reopen: later evaluation establishes materially different mitigation transfer or guarantee conditions, or E.22 changes the evidence profile required for an automated judge. |
| OEE and NQD can use proposal-shaped quality pressure without collapsing proposal, candidate retention, and selection. | Xi Lin et al., Quality-Diversity Optimization as Multi-Objective Optimization, arXiv:2602.00478 (2026), current preprint; Haoxiang Qin et al., A survey on Quality-Diversity optimization: Approaches, applications, and challenges, Swarm and Evolutionary Computation 100:102240 (2026), DOI 10.1016/j.swevo.2025.102240, current survey. | The shared comparison question is how to preserve several high-performing alternatives across declared coordinates or descriptors. The sources do not say that an evaluation proposal is already a generated candidate, archive insertion, front update, or selected result. | Adapt — reason: use the collection-over-coordinates payload only to keep an E.22 proposal portfolio distinct from candidate generation, retention, and selection; it supplies no OEE/NQD authority. Changed loci: CandidateImprovementProposalRow@Context, E.22:4.6, the Proposal portfolio slice, and the C.17–C.19/G.5 relation boundary. Qualification/currentness as of 2026-08-19: the same exact 2026 QD sources apply here only as a bounded OEE/NQD proposal-framing comparison. Reopen: current QD evidence changes the collection/selection distinction, or the direct C.17–C.19 or G.5 consumer boundary changes. |