ETA: Efficiency through Thinking Ahead, A Dual Approach to Self-Driving with Large Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hamdan, Shadi, Sima, Chonghao, Yang, Zetong, Li, Hongyang, Güney, Fatma
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916785354178560
author Hamdan, Shadi
Sima, Chonghao
Yang, Zetong
Li, Hongyang
Güney, Fatma
author_facet Hamdan, Shadi
Sima, Chonghao
Yang, Zetong
Li, Hongyang
Güney, Fatma
contents How can we benefit from large models without sacrificing inference speed, a common dilemma in self-driving systems? A prevalent solution is a dual-system architecture, employing a small model for rapid, reactive decisions and a larger model for slower but more informative analyses. Existing dual-system designs often implement parallel architectures where inference is either directly conducted using the large model at each current frame or retrieved from previously stored inference results. However, these works still struggle to enable large models for a timely response to every online frame. Our key insight is to shift intensive computations of the current frame to previous time steps and perform a batch inference of multiple time steps to make large models respond promptly to each time step. To achieve the shifting, we introduce Efficiency through Thinking Ahead (ETA), an asynchronous system designed to: (1) propagate informative features from the past to the current frame using future predictions from the large model, (2) extract current frame features using a small model for real-time responsiveness, and (3) integrate these dual features via an action mask mechanism that emphasizes action-critical image regions. Evaluated on the Bench2Drive CARLA Leaderboard-v2 benchmark, ETA advances state-of-the-art performance by 8% with a driving score of 69.53 while maintaining a near-real-time inference speed at 50 ms.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07725
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ETA: Efficiency through Thinking Ahead, A Dual Approach to Self-Driving with Large Models
Hamdan, Shadi
Sima, Chonghao
Yang, Zetong
Li, Hongyang
Güney, Fatma
Computer Vision and Pattern Recognition
Artificial Intelligence
How can we benefit from large models without sacrificing inference speed, a common dilemma in self-driving systems? A prevalent solution is a dual-system architecture, employing a small model for rapid, reactive decisions and a larger model for slower but more informative analyses. Existing dual-system designs often implement parallel architectures where inference is either directly conducted using the large model at each current frame or retrieved from previously stored inference results. However, these works still struggle to enable large models for a timely response to every online frame. Our key insight is to shift intensive computations of the current frame to previous time steps and perform a batch inference of multiple time steps to make large models respond promptly to each time step. To achieve the shifting, we introduce Efficiency through Thinking Ahead (ETA), an asynchronous system designed to: (1) propagate informative features from the past to the current frame using future predictions from the large model, (2) extract current frame features using a small model for real-time responsiveness, and (3) integrate these dual features via an action mask mechanism that emphasizes action-critical image regions. Evaluated on the Bench2Drive CARLA Leaderboard-v2 benchmark, ETA advances state-of-the-art performance by 8% with a driving score of 69.53 while maintaining a near-real-time inference speed at 50 ms.
title ETA: Efficiency through Thinking Ahead, A Dual Approach to Self-Driving with Large Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.07725