Test case sampling optimization for safety validation of automated driving systems

Qian, Chen; Xu, Jingbin; Xing, Xin; Guo, Feng · 2026 · Crossref

DOI: 10.1038/s41467-026-69675-8

archive: archived pipeline: cataloged verified

Get this paper ↗ (DOI — opens at the source; we link to it, we don't host it)

Summary

This paper addresses the challenge of selecting efficient and representative test cases for the safety validation of automated driving systems (ADS). Current validation methods often rely on heuristic rules or simplified surrogate models that fail to capture the complexity of real-world driving, particularly rare, safety-critical "corner cases." The authors propose a Kernel Test Case Sampling (KTCS) method designed to simultaneously optimize for two criteria: representativeness, ensuring the selected cases align with real-world driving distributions, and coverage, ensuring the inclusion of high-risk, low-probability scenarios. The goal is to create a standardized, scalable framework that enables robust accident-rate estimation and fair comparisons between ADS performance and human driving benchmarks. The study utilizes data from the Second Strategic Highway Research Program (SHRP2) Naturalistic Driving Study, the largest-scale naturalistic driving dataset available. The test case pool consists of approximately 300,000 normal driving segments and 90 safety-critical crash events, each represented as a 15-second driving segment with 48 extracted features covering driving dynamics, environmental factors, and interaction behaviors. The KTCS method employs a dual optimization approach: it minimizes Information Potential (IP) via stochastic sampling to ensure coverage of the testable space and minimizes Maximum Mean Discrepancy (MMD) using an attention mechanism to ensure representativeness. The method was evaluated against five state-of-the-art sampling techniques, including Minimum Energy Design, MMD-Critic, Support Point, SPARTAN, and Uniform Sampling, using metrics such as Empirical Hellinger distance, Absolute Distribution Difference, and pairwise diversity distances. The results demonstrate that KTCS outperforms existing methods in balancing representativeness and coverage. While methods like Uniform Sampling and SPARTAN failed to capture high-risk corner cases, and Minimum Energy Design overemphasized rare events at the expense of distributional realism, KTCS effectively approximated the density distribution of the full dataset across all 48 features. Quantitative analysis showed that KTCS achieved competitive scores on both IP and MMD metrics, with lower values indicating superior performance compared to competitors. The selected subset of 118 cases closely mirrored the statistical properties of the entire pool, including long-tailed scenarios critical for safety validation. Additionally, the authors introduced a "Scaling Risk" metric to quantify ADS safety relative to human drivers, demonstrating the framework's utility in interpreting test results. The significance of this work lies in providing a principled, data-driven approach to ADS safety validation that mitigates biases inherent in static testing procedures. By leveraging large-scale naturalistic data and optimizing for both realism and risk coverage, the proposed method supports more credible safety evaluations. This framework facilitates accelerated development and deployment of automated driving systems by enabling efficient, standardized testing that builds public trust and regulatory confidence. The ability to reliably estimate accident rates and compare ADS performance against human benchmarks addresses critical gaps in current validation practices, offering a robust solution to the curse of dimensionality and rarity in safety testing.

Provenance

The full processing record for this entry. Every stage of this paper's journey through the pipeline is logged — what ran, with which tool and model, how many attempts it took, and when it last completed.

StageOutcomeToolModelPromptAttemptsCompleted
discover success Crossref 1 2026-08-09
archive success canonical_url 1 2026-08-09
extract success pdftotext 4 2026-08-10
clean success clean 2 2026-08-10
chunk success chunk 2 2026-08-10
embed success embed Qwen/Qwen3-Embedding-8B 2 2026-08-10
promote success 1 2026-08-09
summarize success llm qwen3.6-27b-nvidia summ-v5 2 2026-08-10
tag success vector_similarity 17 2026-08-11
verify success 2 2026-08-10

Summary generated by qwen3.6-27b-nvidia on 2026-08-10; verification: verified.

Topics

Ranked by relevance to this paper. Hover a topic for its definition.

Information type

What kind of knowledge this paper contributes, grouped by family — independent of topic (what it is about) and method (how it was studied).