P48-R1 |
National Institute of Standards and Technology |
AI Test, Evaluation, Validation and Verification (TEVV) |
https://www.nist.gov/ai-test-evaluation-validation-and-verification-tevv |
P48-R2 |
Luiten, Osep, Dendorfer et al. |
HOTA: A Higher Order Metric for Evaluating Multi-Object
Tracking |
https://arxiv.org/abs/2009.07736 |
P48-R3 |
Kim, Woo, Lee and Kweon |
Video Panoptic Segmentation |
https://openaccess.thecvf.com/content_CVPR_2020/html/Kim_Video_Panoptic_Segmentation_CVPR_2020_paper.html |
P48-R4 |
Koh, Sagawa, Marklund et al. |
WILDS: A Benchmark of in-the-Wild Distribution Shifts |
https://proceedings.mlr.press/v139/koh21a.html |
P48-R5 |
Hendrycks and Dietterich |
Benchmarking Neural Network Robustness to Common Corruptions and
Perturbations |
https://openreview.net/forum?id=HJz6tiCqYm |
P48-R6 |
Zhou, Van Landeghem, Popordanoska and Blaschko |
A Novel Characterization of the Population Area Under the Risk
Coverage Curve |
https://proceedings.mlr.press/v267/zhou25y.html |
P48-R7 |
MLCommons |
MLPerf Inference Benchmarks Documentation |
https://docs.mlcommons.org/inference/ |
P48-R8 |
Demšar |
Statistical Comparisons of Classifiers over Multiple Data Sets |
https://jmlr.org/papers/v7/demsar06a.html |
P48-R9 |
Efron and Tibshirani |
An Introduction to the Bootstrap |
https://doi.org/10.1201/9780429246593 |
P48-R10 |
Dietterich |
Approximate Statistical Tests for Comparing Supervised
Classification Learning Algorithms |
https://doi.org/10.1162/089976698300017197 |
P48-R11 |
Purkrabek and Matas |
ProbPose: A Probabilistic Approach to 2D Human Pose Estimation |
https://openaccess.thecvf.com/content/CVPR2025/html/Purkrabek_ProbPose_A_Probabilistic_Approach_to_2D_Human_Pose_Estimation_CVPR_2025_paper.html |
P48-R12 |
de Geus, Meletis, Lu, Wen and Dubbelman |
Part-Aware Panoptic Segmentation |
https://openaccess.thecvf.com/content/CVPR2021/html/de_Geus_Part-Aware_Panoptic_Segmentation_CVPR_2021_paper.html |
P48-R13 |
Ovadia, Fertig, Ren et al. |
Can You Trust Your Model’s Uncertainty? Evaluating Predictive
Uncertainty Under Dataset Shift |
https://proceedings.neurips.cc/paper/2019/hash/8558cb408c1d76621371888657d2eb1d-Abstract.html |
P48-R14 |
Nixon, Dusenberry, Zhang, Jerfel and Tran |
Measuring Calibration in Deep Learning |
https://arxiv.org/abs/1904.01685 |
P48-R15 |
Angelopoulos and Bates |
A Gentle Introduction to Conformal Prediction and Distribution-Free
Uncertainty Quantification |
https://arxiv.org/abs/2107.07511 |
P48-R16 |
Saito and Rehmsmeier |
The Precision-Recall Plot Is More Informative than the ROC Plot When
Evaluating Binary Classifiers on Imbalanced Datasets |
https://doi.org/10.1371/journal.pone.0118432 |
P48-R17 |
Brier |
Verification of Forecasts Expressed in Terms of Probability |
https://doi.org/10.1175/1520-0493(1950)078%3C0001:VOFEIT%3E2.0.CO;2 |
P48-R18 |
Pineau, Vincent-Lamarre, Sinha et al. |
Improving Reproducibility in Machine Learning Research |
https://jmlr.org/papers/v22/20-303.html |
P41-R13 |
COCO Consortium |
COCO Dataset — person keypoints and evaluation context |
https://cocodataset.org/ |
P42-R2 |
Kirillov, He, Girshick, Rother and Dollár |
Panoptic Segmentation |
https://arxiv.org/abs/1801.00868 |
P44-R5 |
Abu Farha and Gall |
MS-TCN: Multi-Stage Temporal Convolutional Network for Action
Segmentation |
https://openaccess.thecvf.com/content_CVPR_2019/html/Abu_Farha_MS-TCN_Multi-Stage_Temporal_Convolutional_Network_for_Action_Segmentation_CVPR_2019_paper.html |
P46-R10 |
Bendale and Boult |
Towards Open Set Deep Networks |
https://openaccess.thecvf.com/content_cvpr_2016/html/Bendale_Towards_Open_Set_CVPR_2016_paper.html |
P46-R14 |
Guo, Pleiss, Sun and Weinberger |
On Calibration of Modern Neural Networks |
https://proceedings.mlr.press/v70/guo17a.html |
P46-R15 |
Dabah and Tirer |
On Temperature Scaling and Conformal Prediction of Deep
Classifiers |
https://proceedings.mlr.press/v267/dabah25a.html |
P46-R21 |
Yan et al. |
CIMI4D: A Large Multimodal Climbing Motion Dataset Under Human-Scene
Interactions |
https://openaccess.thecvf.com/content/CVPR2023/html/Yan_CIMI4D_A_Large_Multimodal_Climbing_Motion_Dataset_Under_Human-Scene_Interactions_CVPR_2023_paper.html |
P46-R22 |
Yan et al. |
ClimbingCap: Multi-Modal Dataset and Method for Rock Climbing in
World Coordinate |
https://openaccess.thecvf.com/content/CVPR2025/html/Yan_ClimbingCap_Multi-Modal_Dataset_and_Method_for_Rock_Climbing_in_World_CVPR_2025_paper.html |
P46-R23 |
Maschek and Schedl |
The Way Up: A Dataset for Hold Usage Detection in Sport
Climbing |
https://openaccess.thecvf.com/content/CVPR2025W/CVSPORTS/html/Maschek_The_Way_Up_A_Dataset_for_Hold_Usage_Detection_in_CVPRW_2025_paper.html |
P46-R25 |
Geifman and El-Yaniv |
SelectiveNet: A Deep Neural Network with an Integrated Reject
Option |
https://proceedings.mlr.press/v97/geifman19a.html |
P46-R26 |
Mitchell et al. |
Model Cards for Model Reporting |
https://doi.org/10.1145/3287560.3287596 |
P46-R27 |
Gebru et al. |
Datasheets for Datasets |
https://doi.org/10.1145/3458723 |
P47-R16 |
Heilbron, Escorcia, Ghanem and Niebles |
ActivityNet: A Large-Scale Video Benchmark for Human Activity
Understanding |
https://openaccess.thecvf.com/content_cvpr_2015/html/Heilbron_ActivityNet_A_Large-Scale_2015_CVPR_paper.html |
P47-R20 |
National Institute of Standards and Technology |
Artificial Intelligence Risk Management Framework (AI RMF 1.0) |
https://doi.org/10.6028/NIST.AI.100-1 |