Library / First Principles Framework (FPF) - Core Conceptual Specification
Jump to passage
In this reading

Link to current text

Published source confirmed at last check

Source changed 2026-10-03 05:29:54 UTC · snapshot created 2026-10-03 05:30:57 UTC · last check 2026-10-03 07:00:10 UTC

A.6.B:8.3 - Show #2: ML evaluation protocol boundary (reproducibility discipline)

A published “evaluation protocol” boundary (common in modern ML governance) benefits from strict classification:

  • L: metric definitions and invariants (e.g., what counts as AUROC; data partition invariants).
  • A: admissibility gates (dataset usage-term constraints; pinned environment constraints; seed policy).
  • D: checker and author duties (publish required faces; use declared dataset version; retention duties for run evidence carriers).
  • E: admitted system EvaluationRunner-A performed MLProtocolEvaluation-T1 : U.Work over Model-M1, Dataset-D7, the pinned environment and seed policy, and evaluation window T; the exact AUROC metric application declared by the L claim returned AUROCResult-T1 with measuredAUROC=0.91. When a checker, gate, or audit decision relies on that result, an A.10 evidence-provenance path links it to exact carriers RunLog-T1, DatasetHash-D7, EvaluationReport-R1, and TraceSet-T1.

The square keeps “must use dataset vX” (D) separate from “evaluation is admissible iff dataset usage terms match” (A), and both separate from “MLProtocolEvaluation-T1 returned AUROCResult-T1(measuredAUROC=0.91)” (E). The report and log carriers may support reliance on that result; producing a carrier is not the measured result.