| 1 |
BEN-01 |
Freeze intended use, operating domain, claims and prohibited
interpretations. |
| 2 |
BEN-02 |
Select immutable dataset, ontology, guideline, split and
reference-policy versions. |
| 3 |
BEN-03 |
Register benchmark profile, primary tasks, metrics, tolerances and
decision gates. |
| 4 |
BEN-04 |
Audit consent, integrity, synchronization, calibration and scorable
reference tiers. |
| 5 |
BEN-05 |
Verify grouped splits, duplicate hashes, scene leakage and test
quarantine. |
| 6 |
BEN-06 |
Build deterministic evaluation container and record all artifact
hashes. |
| 7 |
BEN-07 |
Run constant, rule, modality and oracle baselines. |
| 8 |
BEN-08 |
Evaluate primitive pose, scene and tracking tasks. |
| 9 |
BEN-09 |
Evaluate contact, support and force tasks by reference tier. |
| 10 |
BEN-10 |
Evaluate phase boundaries and multilanes. |
| 11 |
BEN-11 |
Evaluate kinematic trajectories, derivatives and summaries. |
| 12 |
BEN-12 |
Evaluate movement/failure classification and 28 contrast sets. |
| 13 |
BEN-13 |
Evaluate calibration, prediction sets, abstention and open-set
behavior. |
| 14 |
BEN-14 |
Run full attempt reconstruction with real upstream predictions. |
| 15 |
BEN-15 |
Run distribution-shift views and worst-group analysis. |
| 16 |
BEN-16 |
Run controlled corruptions, sensor dropout and calibration
drift. |
| 17 |
BEN-17 |
Execute prespecified ablations on paired units. |
| 18 |
BEN-18 |
Measure latency, throughput, memory, energy and operational failures
at quality floor. |
| 19 |
BEN-19 |
Compute cluster-aware confidence intervals and paired
comparisons. |
| 20 |
BEN-20 |
Sample evidence ledgers and attribute end-to-end errors
upstream. |
| 21 |
BEN-21 |
Re-run audit subset from immutable artifacts to test
reproducibility. |
| 22 |
BEN-22 |
For deployment claims, execute prospective frozen field pilot. |
| 23 |
BEN-23 |
For coach claims, execute separate blinded/controlled human-factors
study. |
| 24 |
BEN-24 |
Publish result bundle, limitations, unsupported classes and gate
decisions. |