Scaling-Aware Data Selection for End-to-End Autonomous Driving Systems

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dimlioglu, Tolga, Chang, Nadine, Shen, Maying, Mahmood, Rafid, Alvarez, Jose M.
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911578636419072
author Dimlioglu, Tolga
Chang, Nadine
Shen, Maying
Mahmood, Rafid
Alvarez, Jose M.
author_facet Dimlioglu, Tolga
Chang, Nadine
Shen, Maying
Mahmood, Rafid
Alvarez, Jose M.
contents Large-scale deep learning models for physical AI applications depend on diverse training data collection efforts. These models and correspondingly, the training data, must address different evaluation criteria necessary for the models to be deployable in real-world environments. Data selection policies can guide the development of the training set, but current frameworks do not account for the ambiguity in how data points affect different metrics. In this work, we propose Mixture Optimization via Scaling-Aware Iterative Collection (MOSAIC), a general data selection framework that operates by: (i) partitioning the dataset into domains; (ii) fitting neural scaling laws from each data domain to the evaluation metrics; and (iii) optimizing a data mixture by iteratively adding data from domains that maximize the change in metrics. We apply MOSAIC to autonomous driving (AD), where an End-to-End (E2E) planner model is evaluated on the Extended Predictive Driver Model Score (EPDMS), an aggregate of driving rule compliance metrics. Here, MOSAIC outperforms a diverse set of baselines on EPDMS with up to 80\% less data.
format Preprint
id arxiv_https___arxiv_org_abs_2604_08366
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Scaling-Aware Data Selection for End-to-End Autonomous Driving Systems
Dimlioglu, Tolga
Chang, Nadine
Shen, Maying
Mahmood, Rafid
Alvarez, Jose M.
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Large-scale deep learning models for physical AI applications depend on diverse training data collection efforts. These models and correspondingly, the training data, must address different evaluation criteria necessary for the models to be deployable in real-world environments. Data selection policies can guide the development of the training set, but current frameworks do not account for the ambiguity in how data points affect different metrics. In this work, we propose Mixture Optimization via Scaling-Aware Iterative Collection (MOSAIC), a general data selection framework that operates by: (i) partitioning the dataset into domains; (ii) fitting neural scaling laws from each data domain to the evaluation metrics; and (iii) optimizing a data mixture by iteratively adding data from domains that maximize the change in metrics. We apply MOSAIC to autonomous driving (AD), where an End-to-End (E2E) planner model is evaluated on the Extended Predictive Driver Model Score (EPDMS), an aggregate of driving rule compliance metrics. Here, MOSAIC outperforms a diverse set of baselines on EPDMS with up to 80\% less data.
title Scaling-Aware Data Selection for End-to-End Autonomous Driving Systems
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.08366