15.230 14.8.20 Decision gates
| Gate | Ámbito | Condición | Acción al fallar |
|---|---|---|---|
DG-01 |
all | dataset release and split hashes match profile | BLOCK |
DG-02 |
all | no unresolved leakage or test quarantine incident | INVALIDATE_RUN |
DG-03 |
all | all primary metrics computable or explicitly NotEstimable | BLOCK_CLAIM |
DG-04 |
pose | pose quality floor met on required visibility strata | BLOCK_DOWNSTREAM_CLAIM |
DG-05 |
scene | scene identity and geometry floor met | BLOCK_CONTACT_CLAIM |
DG-06 |
contact | contact/support floor met on required reference tier | BLOCK_PHASE_AND_CLASS_CLAIM |
DG-07 |
phase | segment and boundary floor met | BLOCK_TEMPORAL_CLAIM |
DG-08 |
kinematics | native-unit bias/error and coverage floor met | BLOCK_KINEMATIC_PUBLICATION |
DG-09 |
classification | macro/hierarchical and contrast floors met | BLOCK_CLASSIFIER_CLAIM |
DG-10 |
open-set | risk-coverage and unknown floors met | REQUIRE_HUMAN_REVIEW_MODE |
DG-11 |
calibration | calibration evaluated globally and by required groups | NO_PROBABILITY_CLAIM |
DG-12 |
robustness | degradation within prespecified envelope or safe abstention | RESTRICT_OPERATING_DOMAIN |
DG-13 |
sensor dropout | missing modalities cause declared degradation, not silent confidence | BLOCK_DEPLOYMENT |
DG-14 |
reproducibility | audited rerun within profile tolerance | BLOCK_RELEASE |
DG-15 |
system | operational success and resource limits met at quality floor | BLOCK_TARGET_PROFILE |
DG-16 |
privacy | consent, withdrawal and access controls valid | BLOCK_DATA_AND_RUN |
DG-17 |
evidence | explanations trace to actual inputs and intermediate states | NO_EXPLANATION_CLAIM |
DG-18 |
field | prospective frozen pilot completed for deployment claims | RESEARCH_ONLY |
DG-19 |
human | human study shows no unacceptable harmful reliance | NO_COACH_DECISION_SUPPORT_CLAIM |
DG-20 |
release | limitations, unsupported labels and operating domain published | BLOCK_RELEASE |
Los thresholds numéricos se almacenan en cada benchmark profile. Esta especificación define qué debe medirse y qué bloquea un claim, pero no inventa valores universales antes de un piloto representativo.