C.40:5.5 - Transfer a parsing repair and reconsider which task is informative
A publisher develops a checker for a miniature link format. A link is a filename, one literal #, and an exact anchor. The rule is to split at the first literal separator, percent-decode the filename once, then look up the decoded filename and anchor in an independently supplied complete catalogue. The catalogue contains guide.md/intro, Guide One.md/intro, C# notes.md/usage and C%23 notes.md/archive. Directories, external URLs and encoded anchors are outside this model.
Prepare four challenges. Each contains the valid link below and a second link with the same path and anchor missing. Correctly accepting the first and rejecting the second gives two correct decisions. The catalogue, rather than the candidate’s output, determines the answers.
| Challenge | Valid link | A: split, no decoding | B: decode whole link once, then split | C: split, decode filename once | D: split, decode filename twice |
|---|---|---|---|---|---|
| Plain | guide.md#intro | 2 | 2 | 2 | 2 |
| Space | Guide%20One.md#intro | 1 | 2 | 2 | 2 |
| Hash | C%23%20notes.md#usage | 1 | 1 | 2 | 2 |
| Percent | C%2523%20notes.md#archive | 1 | 2 | 2 | 1 |
These are exact results of the stipulated parsing operations on the eight inputs, not measurements on a production manual. Rejecting every link scores only one on each challenge; the valid/invalid pair prevents that shortcut from looking successful.
Suppose Space is being developed from A while Hash is being developed from B. Both incumbents score one. Inspecting Hash explains B’s failure: decoding first creates a # inside the filename, and splitting there changes the target. The programmer constructs C by moving the split before the single decoding operation. Hash now scores two. Test C directly on Space; it also scores two and can replace A there without another independent repair. That target test, not the Hash result, supports the replacement.
A proposed new Plain challenge is already solved and adds no needed distinction, so keep it only if its regression use is worthwhile. Percent is also solved by C. That success can remove it from active repair, but deleting the case altogether would lose a useful distinction: the plausible repeated-normalization variation D fails on it. Retain Percent when such a variation is a live possibility.
The response profiles also explain how the comparison basis changes. Under A, B and C, Space and Percent both have profile (1, 2, 2). That profile alone cannot distinguish them. Including D changes their profiles to (1, 2, 2, 2) and (1, 2, 2, 1). The tasks did not change; the set of tried operations did. Recompute that comparison and retain the difference where it affects further development. Surface novelty such as another filename would not necessarily reveal this distinction.
Now suppose one shared checker C must handle all four challenges, and a proposed change produces D. Retaining C as a specialist elsewhere does not make the changed common checker adequate. Rerun the common checker on the retained tasks: Percent falls from two correct decisions to one while the other three stay at two. That exact counterexample establishes a regression without a noisy-rate estimate. Return Percent to the repair inputs, restore single decoding and rerun the four challenges before replacing the common checker. If restoring it prevents a required new use, identify that conflict and construct another operation instead of alternating two inadequate releases. This demonstrates the shared-candidate return for a software repair; it does not demonstrate learning dynamics or an effective training mixture.
For this fully specified small grammar, directly implementing C and retaining the four discriminating challenges is sufficient. A population of parsers and an automatic task generator would add burden without supplying a missing answer. The construction nevertheless explains what a larger search would have to connect: independently assessable tasks, informative variation, local repair, target-tested reuse and reconsideration of retained challenges. If the catalogue becomes incomplete, a larger search over these same binary answers supplies no missing target fact; return that information condition through C.40.CD. If the grammar changes, reconsider the affected rules and challenges while preserving results within the original scope.