HCD.18:11 - SoTA-Echoing
The working question is how to obtain an evaluation that discriminates useful material support at reasonable effort. The selected line combines use-bound validity, examples that distinguish rival rules, functional feedback, and calibrated interpretation of reader evidence. Its advantage over a heading checklist or an unqualified correct-answer count is that a detected difference selects a different repair or evidence claim. The deliberate cost is constructing meaningful contrasts.
| Practice choice | Adopt, adapt, or reject; effect on this pattern | Source contribution, limit, and reopen condition |
|---|---|---|
| Tie a judgement to the intended interpretation and use | Adapt use-bound validity in sections 4.1–4.2 and S1. A single context-free material score cannot answer an assistance-dependent use question. | The AERA/APA/NCME Standards (2014) supply the assessment-validity anchor, not a book-quality scale. Reopen when the audience, task, or inference changes. |
| Make an example distinguish consequentially different rules | Adapt contrastive example design in sections 4.2 and 4.6. The uninterrupted example in M0 cannot discriminate the two time rules; an interruption can at similar reading effort. | Wesenberg et al. (2025) supply failure evidence about ambiguous worked examples in brief units and two topics. The local time example is a constructed application, not their experiment. Reopen when ambiguity is deliberate preparation and later instruction actually resolves the alternatives. |
| Evaluate what feedback permits next | Adapt a separate feedback-and-continuation property in section 4.3 and S1. A key can mark an error without enabling correction. | Wisniewski et al. (2020) synthesize heterogeneous feedback effects; presence or quantity is an inadequate default. They establish no local passing threshold. Reopen when a different support arrangement changes the next action or required evidence. |
| Calibrate an automated judgement for the property being judged | Adopt bounded calibration in section 4.6 and reject simulated novice success as human evidence. This costs a discriminating probe but can expose expert rescue of defective material. | Bavaresco et al. (2025) find task-dependent LLM/human agreement across NLP evaluations. This is counterevidence to unrestricted substitution, not validation of a learning-material rubric. Reopen for a new evaluator, property, or audience-dependent inference. |