HCD.11 — Assess Human Performance in Representative Work
Type: Human Capability Development practitioner method pattern Status: Stable
HCD.11:1 - Problem Frame
Use this when a person’s document, course grade, helped answer or successful task is about to support a claim about what that person can contribute in later Work. Start by writing the contribution to be judged, the conditions in which it is needed and the decision that will use the judgement. A compatible HCD.1 demand account can supply this frame; an equivalent qualified account is sufficient.
The first useful result is a bounded capability-evidence account: what the person demonstrably did, with which means, against which criteria, and what that supports for the receiving use. The result may instead identify an absent observation or an invalid assessment. This lets a practitioner choose the next useful practice or observation without promoting a good product into an unsupported claim about its maker.
Assess the human contribution within equipped work. If the work requires a current source, calculator, accessibility support or specialist return, competent use of that support may be part of the target. An internally available contribution is a different target and needs its own task.
Do not start an assessment when the real question is still what contribution is needed, what counts as domain-correct work, or whether the person had a workable assignment and access. Obtain that input first. Familiar performance does not answer unfamiliar transfer or retention after delay; those are HCD.12 and HCD.13 questions. Release, credentialing and employment decisions may require stronger evidence and retain their own responsible decision makers.
HCD.11:2 - Problem
A visible output often mixes contributions from the person, a template, an AI assistant, a colleague and feedback. A correct answer after a cue can conceal the very recognition failure the assessment was meant to expose. Conversely, removing a required source can make competent source-dependent work look like a personal deficit.
A single score compounds the problem when correct calculation offsets an unsupported compatibility claim, missing observations are treated as errors, or a representative sample is replaced by whichever exercise is easiest to mark. The assessment then directs practice toward the wrong problem and gives later users more confidence than the observations warrant.
HCD.11:3 - Forces
| Force | Working tension |
|---|---|
| Authentic work and attribution | Preserve the means used in later work while making the person’s decision-bearing contribution observable. |
| Coverage and cost | Expose consequential variation without making every small practice decision require a full examination programme. |
| Feedback and observation | Feedback can help learning while changing what a later answer demonstrates. |
| Comparable judgement and domain specificity | Use explicit criteria and consistent scoring without replacing domain correctness with a generic rubric. |
| Useful conclusions and uncertainty | Return the claim already supported while retaining untested conditions and unresolved evidence gaps. |
HCD.11:4 - Solution
HCD.11:4.1 - Frame one assessable contribution
Name the person, representative work family, relevant configuration, allowed means, evidence window and receiving use. Describe an observable contribution such as “detect an unsupported configuration claim and propose a justified continuation,” not “understand engineering.”
Separate the contributions when their evidence can differ. Calculation, source selection, independent challenge, coordination and explanation may all matter, but success in one does not establish the others. State which contributions must be available before a hint and which may properly use sources or other people.
If the demand is only prospective, an assessment may explore the proposed contribution; its result remains evidence for that declared task, not confirmation that the demand is current.
HCD.11:4.2 - Connect the claim to a task and an inference
Before seeing the answer, make this short chain recoverable:
| Question | Minimum answer needed |
|---|---|
| What is claimed? | The person’s contribution under the stated conditions. |
| What task could elicit it? | A representative task containing the decision-bearing difficulty, with the required access and information. |
| What will count as evidence? | The action, explanation, work product, error or omission that would distinguish successful contribution from a plausible rival. |
| What may be inferred? | The bounded conclusion that those observations can support, and the conclusion they cannot support. |
Obtain domain criteria from a current qualified Method, result or competent specialist. Specify material errors, acceptable variants, and any condition that cannot be compensated by success elsewhere. Do not construct a scoring rule that can award overall success while such a condition fails.
Select tasks for consequential variation: for example a matching and a mismatching source, a feasible and an infeasible resource condition, or an ordinary and an exceptional coordination demand. Explain why the selected sample is sufficient for the receiving decision. There is no universal number of tasks. If a rare but consequential condition is untested, exclude it from the claim or obtain the relevant evidence before stronger reliance.
Return a missing criterion, unsuitable task or unavailable competent assessor before judging the person. These are assessment defects, not performance failures.
HCD.11:4.3 - Observe the attempt without hiding help
Explain the task conditions and permitted means. Observe enough of the attempt to attribute the contribution: relevant source use, choice, adaptation, action, result and any required recognition before a cue. Keep material errors, help, corrections and unfinished work visible.
A product alone is enough only when its origin and means of production cannot change the intended inference. Otherwise, use the least intrusive discriminating observation: a short explanation of one consequential choice, a source trace, a direct observation or a fresh task. Do not add an oral examination by default.
Keep the first attempt separate from correction after feedback. A teacher’s question about the missing configuration can improve the answer; it also removes independent detection from that later observation. Record what the cue supplied rather than merely marking the answer “corrected.”
HCD.11:4.4 - Judge contributions and repair only the affected inference
Compare each observed contribution with its criterion. Distinguish a demonstrated error from an absent observation, an inaccessible task and a disputed judgement. Do not give a missing observation the value “failed” unless the task conditions made the required non-response itself informative.
For a material scoring disagreement, return to the same observation and criterion. Clarify an ambiguous criterion or obtain another qualified assessor when needed. If the original task did not discriminate the relevant alternatives, withhold the dependent judgement. Select a new task when resolving that distinction for the receiving use warrants the obtainable observation and its whole burden. Averaging incompatible judgements does not settle what happened.
When feedback or a task repair changes the conditions, preserve the first result and report the later result under the changed conditions. A later success can support a new claim; it does not rewrite an earlier failure.
HCD.11:4.5 - Return a result proportionate to its use
A short account is sufficient when it preserves the person and contribution, task and conditions, observations including help, criterion-based judgement, bounded inference and relevant uncertainty. No particular form is compulsory. Include a further observation only for a selected receiving use that needs it, when its obtainable contribution warrants the full learner, assessor, preparation, delay and displaced-work burden. A current limited result can finish while a stronger claim remains unsupported.
Use the following distinctions in the conclusion:
| Finding | Useful return |
|---|---|
| The required contribution was observed | State the supported performance and the narrower capability claim, with tested conditions and sample limits. |
| A material criterion failed | State which contribution failed in which attempt; preserve successful contributions. |
| Observations cannot distinguish the claim from a live rival | State the unresolved claim and current limit; retain independently supported conclusions. Name a discriminating observation when the receiving use needs that next question. |
| Task, access or criterion was defective | Return the defect to its owner and identify which judgement must wait or be replaced. |
Stop when the intended reversible decision has enough evidence or when a named missing input prevents it. Escalate assurance only when the receiving use needs broader coverage, independent judgement, scoring consistency, accessibility checks or a defensible validity argument. The assessor’s account supplies evidence; it does not grant release or professional authority.
Reopen only the affected claim when a new observation, changed configuration, source, tool, criterion or receiving use changes what the evidence supports.
HCD.11:5 - Archetypal Grounding
HCD.11:5.1 - Correct arithmetic, failed compatibility, later independent performance
Learner-L17 is a fictional person. The attempts below are constructed observations for learning how to assess; they are not observations of Engineer-E27, Engineer-P14 or programme participants.
The task is to recommend a defensible continuation for an engineering change service. HW3/FW8 is required. The supplied financial assumptions give approximate NPVs of 48.9, 51.6 and 44.6 thousand for alternatives A, B and C. B’s highest NPV does not supply compatibility: its available report concerns FW7. A needs a further shared-power check. C retains HW3/FW8 and uses SCHED2, with two hours of required checks per change and one hour of preparation per homogeneous batch of at most four.
With 24 test hours available, ten changes need 20 + 3 = 23 hours; twelve need 24 + 3 = 27. C can therefore support a proposal for ten, not twelve, with two new changes per week left in the queue. The release manager must agree to that limit and retains the release decision. These engineering criteria are stipulated case facts, not general engineering or safety guidance.
The assessed contribution is to inspect configuration evidence, preserve required checks, and explain a supported continuation and its queue consequence. L17 has the reports, task tables, calculator and ordinary AI assistance; no teacher hint is permitted during an independent attempt. The assessor observes source and assistant traces and asks for an explanation only where needed to attribute the conclusion. In the two fresh cases below, the trace places L17’s recognition before any teacher or AI cue names the mismatch; ordinary tool access remains available.
| Attempt and conditions | Constructed observation | Judgement and inference |
|---|---|---|
| First attempt, no teacher cue | L17 calculates the financial ordering correctly and recommends B without noticing that the report is for FW7. | Calculation is supported. The compatibility criterion fails; the first recommendation is unsupported. |
| Same task after “Which configuration was tested?” | L17 finds FW7, withdraws B and explains A’s remaining check and C’s bounded ten-change proposal. | The correction is supported with a configuration cue. Independent detection has not been shown in this attempt. |
| Fresh representative case, ordinary means, no cue | L17 personally finds a mismatching report and explains why the higher financial value cannot settle compatibility; L17 proposes the supported C limit and queue return. | Independent configuration checking and bounded continuation are supported in this case. |
| Second fresh representative case, ordinary means, no cue | L17 again detects the mismatch, traces the relevant evidence and explains the 23-hour limit and remaining queue. | The specified contribution is supported across the two fresh cases, with a small-sample boundary. |
The result does not average the four attempts. It preserves the original miss, the helped correction and the later independent observations. For local reversible practice planning in these tested conditions, the practitioner can preserve the supported calculation and configuration contribution and finish that decision. A proposed material resource change raises a separate question: select its observation when it is obtainable and worthwhile for the next use, rather than making it a condition of retaining the contribution already demonstrated. These observations do not establish performance across every engineering configuration or authorize a release.
HCD.11:5.2 - An inaccessible assessment is not a demonstrated deficit
A task asks a person to use a current technical source, but the assessment removes the source and an accessibility aid normally required for reading it. Failure to retrieve its detail cannot support the intended source-use judgement. Restore the target conditions or explicitly select a different, justified independent-core question; do not prescribe remedial learning from the invalid comparison.
HCD.11:6 - Bias-Annotation
Visible fluency, confidence, speed and polished AI output can draw attention away from the contribution that matters. Inspect the criterion-bearing action, including a justified refusal or a justified positive proposal.
A familiar task can favour prior exposure; an unnecessarily unfamiliar interface can measure access difficulty. Keep familiarity and access conditions visible. High-consequence uses need the relevant accessibility and fairness evidence; a convenient local task does not supply it.
HCD.11:7 - Conformance Checklist
- The same person, contribution, conditions and receiving use are recoverable.
- The task elicits the claimed contribution and covers the consequential variation being claimed.
- Domain criteria, material errors and non-compensable conditions were explicit before scoring.
- The observation distinguishes the person’s contribution from assistance where that distinction matters.
- First attempts, feedback, corrections and unfinished work remain separately interpretable.
- Missing evidence, assessment defects, scoring disagreements and performance errors receive different returns.
- The conclusion preserves uncertainty, untested conditions and an ordinary stop. A further observation is selected for the stronger use actually being pursued and its obtainable, worthwhile contribution; the current supported result requires no future-study field.
HCD.11:8 - Common Anti-Patterns and How to Avoid Them
| Anti-pattern | Better move |
|---|---|
| “The submitted answer is good, therefore the person can do the work.” | Inspect the person’s contribution where the producer or means could change the inference. |
| “The corrected answer counts as independent success.” | Retain the cue and test the contribution in a fresh uncued attempt. |
| “The total score passes despite the compatibility error.” | Keep the non-compensable criterion decisive and preserve the correct calculation separately. |
| “Removing every tool reveals real capability.” | Assess the required equipped contribution; specify an independent core only when the work needs it. |
| “One success proves the whole capability.” | State the tested envelope and finish the use it supports. Keep consequential untested variants outside the claim; obtain their evidence only for a selected stronger use when the observation is feasible and worthwhile. Missing evidence still withholds that stronger reliance. |
HCD.11:9 - Consequences
The assessment returns less impressive but more useful claims: a practitioner can see what to preserve, what to practise, and what cannot yet be judged. Targeted observations can cost less than a larger examination that misses the decisive contribution.
The cost is explicit task selection and observation of help and material errors. Richer scoring and independent assessment are justified by the receiving use, not imposed on every small practice decision. A higher score that hides a failed constraint is a worse result for that use.
HCD.11:10 - Rationale
Capability is inferred from performance under conditions; it is not the document, grade or observation itself. Making the claim-to-observation link explicit exposes both overclaiming and needless retesting. Separating contributions also makes repair local: a failed compatibility check need not erase a valid calculation, and an invalid task need not become a diagnosis of its participant.
HCD.11:11 - SoTA-Echoing
Working question: how can an assessment support a claim about a person’s contribution when tasks, assistance and later use differ?
Selected line — adopt and adapt. The 2014 Testing Standards, Chapter 1 tie evidence to an intended interpretation and use. Expanded evidence-centred design connects tasks and observations to inferences and makes learning during assessment relevant. HCD.11 adapts that line into the short chain in §4.2 and the separate attempts in §5.1.
Serious alternative: score the final work product with one rubric. That is cheaper when the product’s origin is settled and the inference concerns only that product. Where help or a non-compensable criterion changes the human claim, the same rubric can give a confident wrong answer. Keep its useful criterion-based scoring, but reject the substitution of total score for the intended contribution.
The AI tutoring experiment by Bastani and colleagues adds a bounded counterexample: assisted practice gains in high-school mathematics did not guarantee later unassisted performance. It supports keeping support conditions visible, not an adult-workplace effect estimate or a ban on AI.
The selected line changes §4.3’s observation of help and §4.4’s separate judgements, at the cost of only the discriminating evidence needed for the use. These sources guide design; they do not validate this local assessment. Reopen the task or inference when local evidence shows that it rewards a proxy, misses a material error or attributes a tool’s contribution to the person.
HCD.11:12 - Relations
HCD.1 can supply the human-demand and later-Work frame; equivalent qualified input permits direct entry. HCD.3 diagnoses a limiting human target when the evidence needs a causal differential rather than another score. HCD.4 can use compatible new contribution evidence in a profile comparison.
HCD.12 tests material unfamiliar variation; HCD.13 tests delay and support dependence; HCD.14 can revise an affected development assumption from this result alone. E.23.CAE supplies a bounded differential among applicability, access, adaptation, enactment and capability change when that distinction changes the next action. Target-domain Methods and specialist results continue to supply correctness and criticality.
Use C.16 for an actual measurement claim, including its scale, method and comparability basis, and B.3 for a stronger assurance claim when required. This Method introduces neither a universal mastery score nor a new decision authority.
HCD.11:End