On the Feasibility and Opportunity of Autoregressive 3D Object Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Zanming, Yoo, Jinsu, Jeon, Sooyoung, Liu, Zhenzhen, Campbell, Mark, Weinberger, Kilian Q, Hariharan, Bharath, Chao, Wei-Lun, Luo, Katie Z
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918379812552704
author Huang, Zanming
Yoo, Jinsu
Jeon, Sooyoung
Liu, Zhenzhen
Campbell, Mark
Weinberger, Kilian Q
Hariharan, Bharath
Chao, Wei-Lun
Luo, Katie Z
author_facet Huang, Zanming
Yoo, Jinsu
Jeon, Sooyoung
Liu, Zhenzhen
Campbell, Mark
Weinberger, Kilian Q
Hariharan, Bharath
Chao, Wei-Lun
Luo, Katie Z
contents LiDAR-based 3D object detectors typically rely on proposal heads with hand-crafted components like anchor assignment and non-maximum suppression (NMS), complicating training and limiting extensibility. We present AutoReg3D, an autoregressive 3D detector that casts detection as sequence generation. Given point-cloud features, AutoReg3D emits objects in a range-causal (near-to-far) order and encodes each object as a short, discrete-token sequence consisting of its center, size, orientation, velocity, and class. This near-to-far ordering mirrors LiDAR geometry--near objects occlude far ones but not vice versa--enabling straightforward teacher forcing during training and autoregressive decoding at test time. AutoReg3D is compatible across diverse point-cloud or backbones and attains competitive nuScenes performance without anchors or NMS. Beyond parity, the sequential formulation unlocks language-model advances for 3D perception, including GRPO-style reinforcement learning for task-aligned objectives. These results position autoregressive decoding as a viable, flexible alternative for LiDAR-based detection and open a path to importing modern sequence-modeling tools into 3D perception.
format Preprint
id arxiv_https___arxiv_org_abs_2603_07985
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle On the Feasibility and Opportunity of Autoregressive 3D Object Detection
Huang, Zanming
Yoo, Jinsu
Jeon, Sooyoung
Liu, Zhenzhen
Campbell, Mark
Weinberger, Kilian Q
Hariharan, Bharath
Chao, Wei-Lun
Luo, Katie Z
Computer Vision and Pattern Recognition
LiDAR-based 3D object detectors typically rely on proposal heads with hand-crafted components like anchor assignment and non-maximum suppression (NMS), complicating training and limiting extensibility. We present AutoReg3D, an autoregressive 3D detector that casts detection as sequence generation. Given point-cloud features, AutoReg3D emits objects in a range-causal (near-to-far) order and encodes each object as a short, discrete-token sequence consisting of its center, size, orientation, velocity, and class. This near-to-far ordering mirrors LiDAR geometry--near objects occlude far ones but not vice versa--enabling straightforward teacher forcing during training and autoregressive decoding at test time. AutoReg3D is compatible across diverse point-cloud or backbones and attains competitive nuScenes performance without anchors or NMS. Beyond parity, the sequential formulation unlocks language-model advances for 3D perception, including GRPO-style reinforcement learning for task-aligned objectives. These results position autoregressive decoding as a viable, flexible alternative for LiDAR-based detection and open a path to importing modern sequence-modeling tools into 3D perception.
title On the Feasibility and Opportunity of Autoregressive 3D Object Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.07985