HCD.18 - Select or Construct a Characterization and Evaluation Specification for Instructional Material
Type: Human Capability Development practitioner evaluation-specification construction pattern Status: Stable
HCD.18:0 - Use This When
Use this pattern when you must decide what to repair, select, or rely on in instructional material, but the available evaluation cannot distinguish the material’s useful contribution from an attractive presentation, a correct answer supplied by an expert, or help supplied by a teacher.
An editor preparing a worksheet, a teacher choosing a worked explanation, and a designer evaluating a connected course route face the same first question: what work should this material help this audience do, with what support? Write that promise and one task that would expose a consequential failure. Then check whether an existing specification already answers it.
The first useful result is either a compatible specification ready for evaluation, or a new specification that tells an evaluator what to inspect, how to interpret observations, and what decision each difference can change. The specification concerns the instructional material in its declared use. Wider human capability development remains outside this construction task.
If a compatible specification is already available, use it directly through HCD.19. If the question is one person’s performance, transfer, or retention, use HCD.11, HCD.12, or HCD.13. A professional reference used to make a decision is evaluated for that reference use; it acquires a learning-task requirement only when it promises instruction.
HCD.18:0.1 - Working Terms
| Term | Meaning in this pattern |
|---|---|
| Instructional material | The explanations, examples, tasks, representations, and usable returns offered to support specified learning work. A worksheet, book, or guided interface may carry this content. |
| Declared use | The audience, tasks, prerequisites, ordinary assistance, and decision for which the evaluation is intended. |
| Task denominator | The work and consequential variations the material promises to support, recovered from that work and its subject requirements rather than counted from the material’s headings. |
| Characteristic | A property whose possible values can change use, repair, evidence collection, or protection of another property. |
| Evaluation specification | The defined object and use, selected characteristics and scales, evidence requirements, result meanings, and conditions under which an evaluator may draw a conclusion. |
| Material value | A qualified judgement about a selected property of the material. The evidence supporting the judgement and the material’s fit to the specification remain separately interpretable. |
These are working terms, not new artifact types. A short written agreement can carry the specification when it supplies the required meaning.
HCD.18:1 - Problem Frame
A time-accounting handout tells a learner to count from a task’s first start to handoff. Its example contains no interruption, so the intended active effort and the elapsed turnaround have the same value. The example looks correct. A prepared reader later separates the quantities without help, while another reader may follow the handout’s wrong general rule.
Counting correct answers alone would miss the defective explanation. Counting headings would find an explanation, an example, and a task and miss it again. A useful evaluation must make the distinction matter: introduce an interruption, ask for both quantities, inspect the explanation that supports them, and preserve what the reader supplied from prior knowledge.
The evaluator also needs to know what the material promises. A handout used with a teacher’s corrective question has a different support arrangement from a self-study unit. The question can legitimately enable correction without proving that the first explanation enabled independent recognition.
A specification fixes these distinctions before observations are interpreted. It keeps subject truth, audience support, the task’s validity, and the strength of the resulting inference connected without turning them into one score.
HCD.18:2 - Problem
How can a designer construct a small but complete evaluation specification that distinguishes action-changing material defects, supports defensible values and repairs, and exposes the evidence still needed for the declared use?
HCD.18:3 - Forces
- A reusable profile saves effort, but a fixed quality list can miss a task absent from the material itself.
- Shared observations can support several properties, while those properties require different repairs.
- A correct first answer can come from the material, prior expertise, or an available helper; the evaluation must distinguish those contributions when they change reliance.
- More detail can clarify a distinction or obstruct the next action. The property concerns support for the use, not volume.
- A usable ordinal judgement can guide repair without pretending to measure equal intervals or predict learning.
- A demanding task can reveal useful learning work or a missing prerequisite. Immediate ease cannot decide between them.
HCD.18:4 - Solution
Construct the specification from the intended work to the evidence and resulting decision. A.19.ECS supplies the general construction requirements; this pattern supplies their instructional-material application. The instructional-material evaluation profile offers selectable property families and an optional repair/use scale.
HCD.18:4.1 - Recover the promise and stop if an existing specification fits
Name the material edition and the portion to be judged. Distinguish its content from a particular export when font size, navigation, image resolution, or interaction changes access. Identify the intended audience and prerequisites at the level that changes the tasks: “can add time intervals” is useful in the worked case; an invented demographic persona is not.
State the ordinary material, tools, source returns, teacher help, and peer or AI assistance. Separate what is promised from what is available and what later proves to be used. For an intended future audience, these are design conditions, not observations about already identified people.
Name the decision: for example, select between two explanations for a supervised exercise, repair a course’s planning task, or authorize a specified self-study use. Say what a wrong positive conclusion would cost. This selects the evaluation grain and evidence burden.
Compare an existing specification by value: object kind, audience, task denominator, support arrangement, relied-on subject state, properties, value meanings, evidence, and consuming decision. A matching title or the same 0–5 labels do not establish compatibility. If the existing specification fits, stop construction and use it. If only one requirement changed, retain the compatible parts and repair the affected specification content.
HCD.18:4.2 - Derive the tasks and contrast cases outside the present table of contents
Start from the work the audience should be able to perform or practise. Recover its required distinctions, prerequisite operations, acceptable variants, and consequential errors from the problem and a qualified subject basis. A material’s claim of coverage does not define its own denominator.
For each promised task, ask what a successful contribution would contain and which plausible wrong contribution an apparently adequate explanation might also produce. Choose a case on which the alternatives differ. In the time example, an uninterrupted interval cannot distinguish active effort from elapsed turnaround; an interruption can.
Compare this task-derived need with the actual explanations, examples, practice, and accessible returns. A precise external return can supply a contribution if it is usable by this audience at this point. A title that promises the contribution cannot.
Prepare three kinds of contrast:
- an in-scope material that supplies the needed contribution;
- an in-scope material with a consequential deficiency;
- a nearby object or use for which this specification would answer the wrong question.
These contrasts test fit as well as quality. A personal capability judgement is outside a material specification even when the same learner task appears in both.
Distinguish the subject basis from manufacturing history. An existing directly authored book can be evaluated without knowing how its author wrote it. If fidelity to a transcript, prescribed edition, or earlier rendering was promised, evaluate that source-to-material relation on its own evidence.
HCD.18:4.3 - Select properties by the decisions their values change
Use the profile’s property families as a search aid. Keep a characteristic when its values can change the use decision, identify a distinct repair, call for different evidence, or protect a contribution that another repair might damage.
For each proposed characteristic, try two possible material conditions. If both would receive the same practical response and evidence requirement, refine the distinction or remove the redundant coordinate. Split a compound coordinate when its parts can fail independently with different repairs. Subject accuracy and currentness may share a passage but require different evidence; explanatory sufficiency and criterion validity can diverge when a correct explanation is followed by a key that rejects a valid answer.
Declare the characteristic’s bearer and grain. A false arrow in one diagram, a missing operation across an explanation–exercise chain, and a whole-route omission have different reach. A local defect does not inherit the whole book’s grain merely because the evaluator opened the whole book.
When the decision depends on how much selected structure a reader can extract, select C.2.8 U.ExtractableStructuralInformation. Identify the material episteme, its expressing publication form and the reader or observer; fix the selected relations and correctness basis, usable preparation, operations, help, access and budget only as precisely as this comparison needs. These conditions qualify the three bearers. A material reference may identify the account and expression without new paperwork. Explanatory sufficiency and overall material adequacy remain different questions.
Make the selected set complete for the declared decision. A short diagnostic can legitimately borrow two profile questions, but then its conclusion is a bounded diagnosis, not a complete material evaluation. Do not call an omitted required property “not applicable” merely because it is inconvenient to observe.
HCD.18:4.4 - Bind values, evidence, and decisions without collapsing them
For each selected characteristic, define its value domain, scale type, polarity, anchors, and conditions of applicability. A numerical scale needs an interpretation: a count, a measured duration, and an ordinal repair judgement support different operations.
For structural amount, a qualitative account of additional and missing relations can suffice. A count or fraction needs fixed semantic units and a nonempty selected set; another scale needs the interpretation required by C.2.8. Keep the observed recovery, conditional design estimate or mapped formal estimate identifiable. A missing denominator or criterion is a gap, not zero. Define any required-relation use condition separately from amount; more unrelated structure cannot compensate for a missing necessary relation.
The profile’s optional 0–5 rendering is ordinal. Its anchors distinguish absence of usable support, limited support, connected central repair, bounded repair, adequacy, and additional useful resilience. They describe material use and repair, not a structural amount, and do not license averages, equal-step differences, or a universal floor. Choose a floor only when the consuming decision needs one, and explain why that decision can rely on it. State any non-compensable condition at its actual grain.
Make adjacent values distinguishable in the domain. For explanatory sufficiency, distinguish rebuilding the central explanation–task chain from repairing one missing transition. Distinguish adequacy for the selected cases from an evidenced advantage across consequential variations. A value of 5 requires the latter advantage over 4, not more pages or a history of having been repaired.
Attach evidence requirements to each value. Source inspection can establish a false unit relation. It cannot establish how an unobserved learner experienced the passage. Where adequacy for the audience requires reader work, say what contribution must be observed and under which assistance conditions. Keep a missing observation explicit; lack of evidence is not the material value 0.
Define what a result returns for each selected coordinate: value or lawful missingness, evidence basis, limiting case, grounds against adjacent alternatives, implication for use, and smallest worthwhile repair. Shared evidence may be cited by several rows without being counted as independent observations.
Resolve differences at their cause. If two evaluators disagree between 2 and 3, trace whether the repair crosses a central connected chain or remains bounded. Do not average the judgements. If the disputed distinction repeatedly fails to change action, repair the scale rather than demand stronger agreement with an arbitrary label.
HCD.18:4.5 - Make the specification usable by another evaluator
The following content completes an A.19.ECS specification. A table is optional; the values are not.
| Specification content | What the evaluator must be able to recover |
|---|---|
| Object, use, and qualification | Exact material or material family, evaluation grain, audience prerequisites, tasks, assistance, subject/source state, consuming decision, and window in which the basis remains valid. |
| Fit and contrasts | In-scope adequate and deficient examples, outside-kind or outside-use example, entry decision, and the disposition of a forced invocation on an incompatible object. |
| Characteristics and scales | Selected properties, their applicability, value domains and scale types, polarity, anchors, any floor or exceptional-value condition, and protected trade-offs. |
| Evidence and calibration | Coordinate-specific evidence minima, qualified subject criteria, actual tasks and competing answers, admissible inference, missingness, known contrasts, and the needed reader or evaluator calibration. |
| Result and evidence payload | Every selected coordinate with value or missingness; material locus, task/observation, available and actual help, adjacent-value grounds, limiting example, repair, and claim limits. Preserve the task, criterion, first answer, and later assistance when they carry the inference. |
| Decisions and continuation | Supported-use, repair-needed, insufficient-evidence, disputed-criterion, and fit-defect meanings as needed; stop, reopen, and next-evidence conditions; the starting result for E.23 if improvement is needed. |
| Neighbouring claims and comparisons | Exact exits for human performance, transfer, retention, source fidelity, or production history. If selection among alternatives is live, the common comparison basis and E.22 use, including incomparable cases. |
A complete specification need not already have every observation that its later use requires. It must state the missing observations and prevent a stronger result from being claimed without them. Conversely, a calibrated illustration does not make every future application adequate.
An incompatible object normally stops before invocation. If an invocation is nevertheless required, return an explicit fit defect and the blocked conclusion. If a selected coordinate becomes unevaluable during an in-scope application, return its evidence or criterion defect in the complete result. Neither case authorizes silently dropping that coordinate.
HCD.18:4.6 - Try the specification on contrasts and establish its continuation
Have an evaluator use the same specification on its adequate, deficient, and outside-use contrasts. Examine the resulting decisions and repairs, not merely agreement on labels. Check whether a known central defect can pass because an example is ambiguous, whether a correct alternative is rejected, and whether a missing observation is mistakenly converted to a low material score.
Where an AI reader or evaluator is used, test the contribution it is actually able to make. Subject-grounded counterexamples can calibrate defect diagnosis; simulated novice behaviour cannot establish human learning. A fresh reader is useful when private knowledge of the intended defect or answer would defeat the selected inference.
Stop construction when another evaluator can identify the object and tasks, collect the specified evidence, return all selected coordinates with their limits, and decide the declared next move. An unresolved criterion or discriminating-evidence gap is a legitimate returned result.
Name the specification version when results will be compared later. A changed value meaning, task denominator, or support arrangement may invalidate numeric comparison with earlier results. Preserve earlier results under their original meanings; establish a justified comparison of meanings or decline equivalence. Reopen only the affected specification content when a new task, audience, source premise, assistance change, or recurrent calibration disagreement changes it.
HCD.18:5 - Archetypal Grounding
HCD.18:5.1 - A complete small specification for active effort and elapsed turnaround
A teacher is preparing a short introduction for readers who can read clock times, add intervals, and use a calculator. The promised work is to separate task effort from turnaround and compare consumed person-time with available person-time. The material stays available. A teacher may ask a corrective question after the first answer. Independent first recognition and supported correction are recorded separately.
Call the following specification Time-use S1. Its decision is whether the material supports this bounded introduction with the declared teacher contribution, and what to repair before relying on it. The object is the handout’s explanation, examples, task, and usable correction support, not a learner’s durable capability or the visual quality of an uninspected export. Qualification holds while these tasks, audience prerequisites, support, and time-accounting meanings remain unchanged.
The subject criteria are stipulated before an attempt: active effort sums the named person’s active intervals; elapsed turnaround spans first start to handoff; exclusive active intervals are not double-counted; simultaneous work by different people is added as person-time; a resource comparison uses the same population and unit on both sides.
| Task derived from the promise | Required contribution and discriminating error |
|---|---|
| Interrupted report | One person works on a report 09:00–09:30 and 11:30–12:00, with incident work 09:30–11:30. Return report effort 60 person-minutes, incident effort 120 person-minutes, report turnaround 180 clock-minutes, total effort 180 person-minutes, and incident share 120/180. Calling report effort 180 exposes the rival rule. |
| Concurrent people | A works on X 09:00–10:00; B works on Y 09:00–11:00; both are available 09:00–11:00. Return consumed effort 3 person-hours, available effort 4 person-hours, and elapsed window 2 hours. Compare 3 with 4, not 3 with 2. |
| Correction and continuation | Preserve the first report answer, then ask which intervals contain report work and which belong to the incident. Ask the reader to correct or defend the answer and apply the distinction to the concurrent-people task. A repeated answer without a reason does not show that feedback enabled correction. |
These tasks are not inferred from the handout’s headings. The interruption distinguishes two time concepts; the concurrent case distinguishes person-time from clock duration. Both are necessary to the stated promise.
Time-use S1 selects six characteristics. It uses the profile’s ordinal repair/use anchors, with higher values better for this use. For the supported-use conclusion, every selected property must be at least 4 on the evidence required below. This is S1’s conservative local decision rule, not an HCD floor. No averaging compensates a failed required property. Missing evidence returns “insufficient evidence for supported use”; it does not prohibit a controlled probe designed to obtain that evidence.
| Selected characteristic | Evidence required for 4 in S1 | What distinguishes the neighbouring repair or resilience judgement |
|---|---|---|
| Subject validity | Inspect the explanation, example, task, key, and correction against the stipulated meanings; challenge the interruption and population denominators. | One bounded incorrect unit label with otherwise consistent uses calls for 3; a wrong central rule propagated through the explanation and use calls for 2. For 5, demonstrate an additional useful boundary treatment, such as separating waiting from other-task work without corrupting either quantity. |
| Task and dependency coverage | Trace both task families and their prerequisites to actual material or accessible declared support; inspect that each needed operation is supplied. | A missing usable return for one supplied operation is bounded repair; absence of the concurrent-person operation needs a connected addition. For 5, show useful coverage of a consequential further case, such as unequal availability, while preserving the original promise. |
| Explanatory sufficiency | Inspect the grounds linking intervals, quantities, and comparison; obtain reader work under the declared prerequisite and support conditions that reveals how the reader recovers those grounds. | One missing transition differs from rebuilding the active/elapsed explanation chain. For 5, the reader can use the supplied explanation on an additional consequential variation, with its source contribution and assistance visible. |
| Example discrimination | Apply the correct and rival rules to the supplied examples. At least one example must distinguish interruption and one must distinguish aggregate person-time from elapsed duration, in material or the declared accessible support. | An isolated misleading example amid a sufficient discriminating pair differs from an example set compatible throughout with the rival rule. For 5, an additional boundary contrast prevents a further identified misapplication. |
| Task and criterion validity | Independently solve the tasks; test the key against a correct alternative and the named rival errors. Equivalent units and justified exact or rounded shares are accepted. | A key’s bounded omission differs from a task–criterion chain that elicits or rewards the wrong quantity. For 5, the criterion demonstrably discriminates a further consequential error without rejecting a valid alternative. |
| Feedback and continuation support | Inspect the available feedback and observe a first answer followed by the declared question, an explained correction or justified retention, and the changed task. Record the teacher’s actual contribution. | A missing next-task return differs from a support arrangement that cannot expose or correct the central discrepancy. For 5, an additional relevant difficulty can be handled through the declared support without supplying the next answer wholesale. |
The remaining profile families do not become automatic extra scores. In this bounded case, semantic representations and language are inspected where they carry the selected explanations and criteria; no separate alternative export, long narrative route, or changing subject edition is promised. If one of those conditions changes, revisit the selection.
For each coordinate, S1 returns the value or named missingness, the inspected material passage and task, the criterion and actual response, source inspection versus reader evidence, available and actual help, adjacent-value grounds, and repair reach. A 4/5 judgement without its required observation is blocked. An observed subject error can still support a lower value without waiting for a learner to reproduce it.
HCD.18:5.2 - Contrasts that test what the specification can conclude
A deficient handout, M0, says that time spent doing a task runs from first touch to handoff. Its only worked example is uninterrupted. It offers some useful time-accounting content, but its central rule and example do not support the interruption distinction. S1 can locate a connected explanation–example repair even if a prepared reader independently returns all correct quantities.
A revised explanation, M1, distinguishes active intervals from start-to-handoff turnaround, separates the columns, gives the interrupted report result, and states how to sum person-time across people. This repairs the named subject relation. To become an adequate contrast for all of S1, the material-and-support arrangement must also supply the concurrent example, a criterion that accepts valid equivalents, and the specified reader/feedback evidence. “Better explanation” alone is not a complete six-coordinate result.
An adequate constructed contrast combines M1 with the concurrent worked comparison, the S1 criteria, and an available teacher question and retry. It shows what adequate content and support would contain. Until an appropriate reader actually uses that arrangement, explanatory and feedback adequacy remain evidence requirements, not observed successes.
The outside-use contrast is “this reader will retain the distinction independently next month.” It concerns a human retention claim. S1 returns outside use and points to HCD.13; it does not award the handout 0.
The practical change is visible before collecting new learner data: M0 can no longer pass on heading coverage and expert accuracy alone; the repaired material cannot acquire a learning claim merely by containing the right definitions; and the remaining observation is specific enough to undertake.
HCD.18:6 - Bias-Annotation
The selected task denominator can inherit the designer’s narrow experience. Compare it with consequential work variants and a qualified domain account, especially when the present material omits them. Declared prerequisites can also conceal missing instruction by demanding knowledge the intended audience was never promised to have.
Prepared readers and language models may repair an explanation without noticing how much prior knowledge they used. Their diagnosis can establish a source defect, but their success does not automatically qualify material for a less-prepared audience. Preserve actual human difficulty reports when their task or assistance differs.
An ordinal scale can make evaluators seek a number before understanding repair reach. Read the limiting case and evidence first; retain an unresolved boundary when its meaning is disputed.
HCD.18:7 - Conformance Checklist
This is assurance guidance for a constructed specification; the working method is in sections 4 and 5.
- HCD18-CC1. The specification SHALL identify the material, grain, audience, tasks, prerequisites, assistance, subject basis, consuming decision, and qualification conditions.
- HCD18-CC2. The constructor SHALL check compatibility with an existing specification and derive the task denominator from the promised work rather than the present headings alone.
- HCD18-CC3. The specification SHALL provide adequate, deficient, and outside-use contrasts and preserve the distinction between fit, material value, and evidence.
- HCD18-CC4. Each selected characteristic SHALL have an action-changing purpose, applicability, scale binding, anchors, evidence requirements, and an explicit result or missingness disposition.
- HCD18-CC5. An ordinal numeric rendering SHALL NOT be treated as equal intervals or silently averaged. Any floor, exceptional value, comparison, or non-compensable condition SHALL be justified for the consuming use.
- HCD18-CC6. The result requirements SHALL preserve the evidence and adjacent-value grounds needed to recover the judgement, including first response and actual assistance when they change the inference.
- HCD18-CC7. Calibration SHALL test discriminating contrasts and preserve criterion or evidence disagreements. Constructed cases SHALL be distinguished from observed reader work.
- HCD18-CC8. The specification SHALL state supported conclusions, stops, reopens, neighbouring-claim exits, and the E.23 starting result when repair is needed.
HCD.18:8 - Common Anti-Patterns and How to Avoid Them
| Misuse | Why it changes the wrong decision | Repair |
|---|---|---|
| Count every present heading as coverage | The material defines its own denominator and hides omitted work. | Recover the required task and compare its dependencies with the content. |
| Treat every profile family as a mandatory score | Inert coordinates add effort while an unlisted live property can still be missed. | Select by changed action, evidence, or protection; complete the set for the actual use. |
| Award 0 for no learner observation | Ignorance becomes a demonstrated material failure. | Return missing evidence and the conclusion it blocks. |
| Infer material adequacy from an expert’s correct answer | The expert may supply the missing or corrected relation. | Inspect the material and the actual source of the response; obtain audience-compatible evidence if needed. |
| Copy another programme’s 0–5 values | Identical numbers can encode different fit, repair, and floor meanings. | Retain the original scale and establish only a justified comparison of meanings. |
| Require construction records for every old book | A current product question is replaced by a historical claim. | Evaluate current support; open source-fidelity or history only when promised or consequential. |
HCD.18:9 - Consequences
A constructed specification makes the next evaluation executable: the evaluator knows what work matters, what a defect would change, what evidence to obtain, and which stronger conclusion remains unavailable. Repairs can target a criterion, example, explanation, return, or assistance arrangement instead of expanding the material indiscriminately.
The cost is up-front task and contrast design. That cost is worthwhile when an intuitive score could drive consequential selection or repair. A compatible existing specification avoids repeating it. New audiences or purposes can require a changed specification; preserving the original result meanings makes that change inspectable.
HCD.18:10 - Architectural Rationale
Material evaluation sits inside human capability development because the material is one contribution to learning work. Keeping its evaluation distinct allows the same explanation to be judged within a teacher-supported activity, a self-study unit, or a narrower preparatory use without attributing every outcome to the text.
A.19.ECS already supplies the general specification-construction discipline. The additional domain contribution is the task-derived denominator, discrimination of plausible learner rules, assistance-sensitive evidence, and material-specific repair reach. HCD.19 consumes the result to inspect material and actual reader work. Their relationship is a direct result dependency when construction is needed, not a mandatory two-step lifecycle.
The common profile supplies reusable property meanings and comparison cautions. It does not replace task-specific selection or turn a framework publication into an admitted MethodDescription. An independently live Method or MethodDescription claim retains its own admission requirements.
HCD.18:11 - SoTA-Echoing
The working question is how to obtain an evaluation that discriminates useful material support at reasonable effort. The selected line combines use-bound validity, examples that distinguish rival rules, functional feedback, and calibrated interpretation of reader evidence. Its advantage over a heading checklist or an unqualified correct-answer count is that a detected difference selects a different repair or evidence claim. The deliberate cost is constructing meaningful contrasts.
| Practice choice | Adopt, adapt, or reject; effect on this pattern | Source contribution, limit, and reopen condition |
|---|---|---|
| Tie a judgement to the intended interpretation and use | Adapt use-bound validity in sections 4.1–4.2 and S1. A single context-free material score cannot answer an assistance-dependent use question. | The AERA/APA/NCME Standards (2014) supply the assessment-validity anchor, not a book-quality scale. Reopen when the audience, task, or inference changes. |
| Make an example distinguish consequentially different rules | Adapt contrastive example design in sections 4.2 and 4.6. The uninterrupted example in M0 cannot discriminate the two time rules; an interruption can at similar reading effort. | Wesenberg et al. (2025) supply failure evidence about ambiguous worked examples in brief units and two topics. The local time example is a constructed application, not their experiment. Reopen when ambiguity is deliberate preparation and later instruction actually resolves the alternatives. |
| Evaluate what feedback permits next | Adapt a separate feedback-and-continuation property in section 4.3 and S1. A key can mark an error without enabling correction. | Wisniewski et al. (2020) synthesize heterogeneous feedback effects; presence or quantity is an inadequate default. They establish no local passing threshold. Reopen when a different support arrangement changes the next action or required evidence. |
| Calibrate an automated judgement for the property being judged | Adopt bounded calibration in section 4.6 and reject simulated novice success as human evidence. This costs a discriminating probe but can expose expert rescue of defective material. | Bavaresco et al. (2025) find task-dependent LLM/human agreement across NLP evaluations. This is counterevidence to unrestricted substitution, not validation of a learning-material rubric. Reopen for a new evaluator, property, or audience-dependent inference. |
HCD.18:12 - Relations
- A.19.ECS defines the general content of an evaluation specification and its fit, evidence, comparison, and continuation requirements. A.17, A.18, and C.16 retain the general characteristic and scale meanings used by its bindings.
- HCD.19 applies a compatible specification to instructional material and representative reader work. E.22 supplies comparison when selection is live; E.23 supplies improvement from the resulting defect and protected trade-offs.
- HCD.3 and HCD.6 can supply the target contribution and representative task design. HCD.18 can also use qualified task and audience inputs obtained elsewhere; those patterns are not mandatory preliminary stages.
- HCD.11, HCD.12, and HCD.13 govern evidence about human performance, transfer, and retention. A material judgement contributes to those questions without replacing their evidence.
- NSTD.6 and NSTD.8 contribute current narrative-product adequacy, attachment, and progressive reconstruction when the selected material’s narrative use makes those properties live.
- E.4.DPF.DA governs the adequacy of a DPF package used as a framework. An instructional-material profile applies additionally only for an instructional promise; it does not turn every reference framework into a course.
HCD.18:End