HCD.8:5.1 - Calibrate Assessors on Three Contrasting Answers
HCD.7 has exposed a missing contribution: assessors cannot yet apply the engineering-case criteria consistently. Their provider target task is to judge a short answer, identify the exact requirement and evidence, preserve supported parts, distinguish a correct bounded use from a resource error or justified request, and resolve disagreement within the final-week assessment window.
Preparation uses three contrasting answers:
| Case answer | Required assessor judgement |
|---|---|
| The report matches HW3/FW8/SCHED2, mandatory checks are confirmed, and 24 test hours are available. The learner proposes ten changes in three homogeneous batches, requiring 23 hours, and states that two of the twelve requested changes remain in the queue. | Supported bounded proposal; evidence applicability, resource limit, and queue consequence are preserved. |
| The same data are supplied, but the learner promises twelve changes because batching is faster. | Unsupported: twelve changes require 27 hours, exceeding the 24-hour limit. |
| The changed report does not identify the tested firmware, and the learner asks which version was tested before relying on it for FW8. | Justified missing-information return. The same refusal would be incomplete when the first row’s full compatible report is present. |
Assessors judge independently, then locate each disagreement in the disputed requirement, answer fragment, and supplied evidence. A calibrated specialist result or answer key resolves the domain question; smooth writing or a general impression of the learner does not. The assessors then receive another representative case, such as the same ten-change proposal with only 22 hours. They must reject the now-unsupported promise, preserve unaffected criteria, and invoke the disagreement route within the promised time.
The three examples, key, and reserved calibration hours are preparation. Consistent and timely judgement on the changed case supports the bounded assessor contribution. It does not prove general assessor capability, operation at the whole final-week load, or any learner effect; those stronger uses need their own operating evidence.