Identifying and managing data quality requirements: a design science study in the field of automated driving
DOI: 10.1007/s11219-023-09622-8
archive: archived pipeline: cataloged verified
Get this paper ↗ (DOI — opens at the source; we link to it, we don't host it)
Summary
This study addresses the critical lack of systematic processes for identifying and managing data quality requirements in safety-critical, data-driven systems, specifically within the context of automated driving. As Advanced Driver Assistance Systems (ADAS) increasingly rely on deep learning for perception tasks, the quality of training, validation, and runtime data directly impacts system safety. Poor data quality can lead to fatal faults, yet current practices often rely on undocumented expert knowledge rather than structured frameworks. The authors aim to identify relevant data quality challenges and develop a candidate framework to help stakeholders systematically assess and maintain data quality. The research employs a Design Science Research (DSR) methodology conducted in three iterative cycles: problem identification, solution design, and evaluation. The study is grounded in a case study involving a Swedish Tier 1 automotive supplier. Data was collected through a literature review, semi-structured interviews with eight industry and research experts, and two surveys. In the first cycle, 27 data quality challenges were identified and ranked using a calculated "Challenge Score" metric derived from survey responses. The second cycle focused on designing the artifact components, verified through additional interviews and thematic coding. The third cycle evaluated the proposed framework via a focus group with five participants and a survey of ten anonymous respondents from a deep learning research project. The primary contribution is the Candidate Framework for Data Quality Assessment and Maintenance (CaFDaQAM), which consists of four integrated components. First, a Data Quality Workflow provides a six-step process for identifying challenges, collecting attributes, associating them, defining metrics, identifying solutions, and presenting findings to stakeholders. Second, a List of Data Quality Challenges offers a template to document challenges, including their type, severity score, and impact on AI functions. Third, a List of Data Quality Attributes defines qualitative concepts and their quantitative metrics. Fourth, Solution Candidates provide templates for mitigating identified challenges. The framework establishes a many-to-many relationship between challenges and attributes, allowing for targeted mitigation strategies. The significance of this work lies in providing a structured, extendable blueprint for managing data quality in deep learning applications, particularly in safety-critical domains. By consolidating challenges, attributes, and solutions into a unified framework, CaFDaQAM bridges the gap between data science practices and requirements engineering. The study validates that such a framework can help practitioners move beyond ad-hoc data handling toward systematic quality assurance, thereby enhancing the reliability and safety of automated driving systems. The authors position CaFDaQAM as a stepping stone toward a comprehensive standard for data quality management in data-centric developments.
Provenance
The full processing record for this entry. Every stage of this paper's journey through the pipeline is logged — what ran, with which tool and model, how many attempts it took, and when it last completed.
| Stage | Outcome | Tool | Model | Prompt | Attempts | Completed |
|---|---|---|---|---|---|---|
| discover | success | Crossref | — | — | 1 | 2026-08-09 |
| archive | success | canonical_url | — | — | 1 | 2026-08-09 |
| extract | success | pdftotext | — | — | 4 | 2026-08-10 |
| clean | success | clean | — | — | 2 | 2026-08-10 |
| chunk | success | chunk | — | — | 2 | 2026-08-10 |
| embed | success | embed | Qwen/Qwen3-Embedding-8B | — | 2 | 2026-08-10 |
| promote | success | — | — | — | 1 | 2026-08-09 |
| summarize | success | llm | qwen3.6-27b-nvidia | summ-v5 | 2 | 2026-08-10 |
| tag | success | vector_similarity | — | — | 17 | 2026-08-11 |
| verify | success | — | — | — | 2 | 2026-08-10 |
Summary generated by qwen3.6-27b-nvidia on 2026-08-10; verification: verified.
Topics
Ranked by relevance to this paper. Hover a topic for its definition.
Information type
What kind of knowledge this paper contributes, grouped by family — independent of topic (what it is about) and method (how it was studied).
- Methodological Resource: dataset resource, validation psychometrics, tool software