Comment from Anonymous
Anonymous AnonymousSupportOther
Summary: The commenter supports the inclusion of Rear Automatic Braking (RAB) with pedestrian detection in the New Car Assessment Program but argues that the current single-trial testing protocol lacks statistical rigor. They recommend a multi-trial sampling design and the inclusion of specificity controls to better measure system reliability and account for false positives.
While the inclusion of RAB with pedestrian detection is a vital initiative to mitigate backover crashes, the currently proposed evaluation protocol—specifically the single-trial per test condition requirement presents statistical vulnerabilities. From a statistical quality control and measurement theory perspective, the proposed design risks high Type I (false negative) error rates and Type II (false positive) for high-performing systems and fails to measure the probabilistic nature of Advanced Driver Assistance Systems (ADAS). Recommendations are provided below to enhance the statistical rigor of the protocol. If an RAB system suffers from frequent "phantom braking," drivers will become frustrated and manually disable the system. Therefore a good measurement of the RAB system must account for both false positives and false negatives.
To accurately measure system performance, NHTSA should consider a single-trial testing with a multi-trial sampling design for each condition. Using a multi-trial approach dramatically reduces sensitivity to random outliers and provides a statistically better estimate of a system's true reliability parameter.
Sample Representativeness and Test Matrix Covariates
A robust experimental design must ensure that the 20 test conditions accurately represent the parameter space of real-world backover crashes.
1. A number of other comments have pointed out and we concur that the addition of illumination variance in low-light environment can better covariate for different light levels. If camera-based perception models suffer degraded performance in low-light conditions, the inclusion of low-light trials would improve the validity of the dataset.
2. Early computer vision algorithms trained on homogeneous datasets demonstrated measurable disparities in detecting darker-skinned pedestrians in low-light conditions.
Recommendation: NHTSA should integrate Specificity Controls (Nuisance/False-Positive Testing) into the test matrix. To receive NCAP credit, a vehicle must demonstrate both high Sensitivity (> X% pedestrian collision avoidance) and high Specificity (> Y% false activation rate on non-hazardous targets).