On the Feasibility and Opportunity of Autoregressive 3D Object Detection
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918379812552704 |
|---|---|
| author | Huang, Zanming Yoo, Jinsu Jeon, Sooyoung Liu, Zhenzhen Campbell, Mark Weinberger, Kilian Q Hariharan, Bharath Chao, Wei-Lun Luo, Katie Z |
| author_facet | Huang, Zanming Yoo, Jinsu Jeon, Sooyoung Liu, Zhenzhen Campbell, Mark Weinberger, Kilian Q Hariharan, Bharath Chao, Wei-Lun Luo, Katie Z |
| contents | LiDAR-based 3D object detectors typically rely on proposal heads with hand-crafted components like anchor assignment and non-maximum suppression (NMS), complicating training and limiting extensibility. We present AutoReg3D, an autoregressive 3D detector that casts detection as sequence generation. Given point-cloud features, AutoReg3D emits objects in a range-causal (near-to-far) order and encodes each object as a short, discrete-token sequence consisting of its center, size, orientation, velocity, and class. This near-to-far ordering mirrors LiDAR geometry--near objects occlude far ones but not vice versa--enabling straightforward teacher forcing during training and autoregressive decoding at test time. AutoReg3D is compatible across diverse point-cloud or backbones and attains competitive nuScenes performance without anchors or NMS. Beyond parity, the sequential formulation unlocks language-model advances for 3D perception, including GRPO-style reinforcement learning for task-aligned objectives. These results position autoregressive decoding as a viable, flexible alternative for LiDAR-based detection and open a path to importing modern sequence-modeling tools into 3D perception. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_07985 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | On the Feasibility and Opportunity of Autoregressive 3D Object Detection Huang, Zanming Yoo, Jinsu Jeon, Sooyoung Liu, Zhenzhen Campbell, Mark Weinberger, Kilian Q Hariharan, Bharath Chao, Wei-Lun Luo, Katie Z Computer Vision and Pattern Recognition LiDAR-based 3D object detectors typically rely on proposal heads with hand-crafted components like anchor assignment and non-maximum suppression (NMS), complicating training and limiting extensibility. We present AutoReg3D, an autoregressive 3D detector that casts detection as sequence generation. Given point-cloud features, AutoReg3D emits objects in a range-causal (near-to-far) order and encodes each object as a short, discrete-token sequence consisting of its center, size, orientation, velocity, and class. This near-to-far ordering mirrors LiDAR geometry--near objects occlude far ones but not vice versa--enabling straightforward teacher forcing during training and autoregressive decoding at test time. AutoReg3D is compatible across diverse point-cloud or backbones and attains competitive nuScenes performance without anchors or NMS. Beyond parity, the sequential formulation unlocks language-model advances for 3D perception, including GRPO-style reinforcement learning for task-aligned objectives. These results position autoregressive decoding as a viable, flexible alternative for LiDAR-based detection and open a path to importing modern sequence-modeling tools into 3D perception. |
| title | On the Feasibility and Opportunity of Autoregressive 3D Object Detection |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2603.07985 |