MET-CLS-01 |
Macro F1 |
all labels and each hierarchy level |
MET-CLS-02 |
Micro F1 |
multi-label aggregate |
MET-CLS-03 |
Mean Average Precision |
310 movements and 34 failures |
MET-CLS-04 |
Per-class precision/recall/F1 |
with support and confidence intervals |
MET-CLS-05 |
Exact match ratio |
episode label sets |
MET-CLS-06 |
Top-k recall |
candidate assistance only |
MET-CLS-07 |
Hierarchy-aware precision/recall |
penalize errors by ontology distance |
MET-CLS-08 |
Segmental F1@10/25/50 |
temporal episodes |
MET-CLS-09 |
Edit score |
sequence over-segmentation/order |
MET-CLS-10 |
Boundary F1 and tolerance error |
onset/offset intervals |
MET-CLS-11 |
Contrast-set accuracy and margin |
each high-priority diagnostic group |
MET-CLS-12 |
Expected Calibration Error |
global and classwise with bin sensitivity |
MET-CLS-13 |
Brier score |
multi-label probability quality |
MET-CLS-14 |
Negative log-likelihood |
calibrated probabilistic output |
MET-CLS-15 |
Reliability diagrams |
overall, class, athlete and domain groups |
MET-CLS-16 |
AUROC unknown detection |
known versus held-out unknown |
MET-CLS-17 |
AUPR unknown detection |
unknown-positive and known-positive views |
MET-CLS-18 |
FPR@95TPR |
open-set rejection |
MET-CLS-19 |
OSCR curve |
open-set classification rate |
MET-CLS-20 |
Risk–coverage curve |
selective prediction |
MET-CLS-21 |
Conformal coverage and mean set size |
marginal plus class/group diagnostics |
MET-CLS-22 |
Cross-athlete/route/gym drop |
domain generalization |
MET-CLS-23 |
Shortcut sensitivity |
background, route, athlete and camera perturbations |
MET-CLS-24 |
Human review yield and override rate |
deployment utility and error discovery |