LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
Fuente:
arXiv
Saved in:
| Main Authors: | Man, Yunze, Wang, Shihao, Zhang, Guowen, Bjorck, Johan, Li, Zhiqi, Gui, Liang-Yan, Fan, Jim, Kautz, Jan, Wang, Yu-Xiong, Yu, Zhiding |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
by: Wang, Shihao, et al.
Published: (2026)
by: Wang, Shihao, et al.
Published: (2026)
Situational Awareness Matters in 3D Vision Language Reasoning
by: Man, Yunze, et al.
Published: (2024)
by: Man, Yunze, et al.
Published: (2024)
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
by: Man, Yunze, et al.
Published: (2025)
by: Man, Yunze, et al.
Published: (2025)
Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding
by: Man, Yunze, et al.
Published: (2024)
by: Man, Yunze, et al.
Published: (2024)
DualCross: Cross-Modality Cross-Domain Adaptation for Monocular BEV Perception
by: Man, Yunze, et al.
Published: (2023)
by: Man, Yunze, et al.
Published: (2023)
PaintScene4D: Consistent 4D Scene Generation from Text Prompts
by: Gupta, Vinayak, et al.
Published: (2024)
by: Gupta, Vinayak, et al.
Published: (2024)
Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning
by: Huang, Chi-Pin, et al.
Published: (2026)
by: Huang, Chi-Pin, et al.
Published: (2026)
SceneCraft: Layout-Guided 3D Scene Generation
by: Yang, Xiuyu, et al.
Published: (2024)
by: Yang, Xiuyu, et al.
Published: (2024)
PhyCritic: Multimodal Critic Models for Physical AI
by: Xiong, Tianyi, et al.
Published: (2026)
by: Xiong, Tianyi, et al.
Published: (2026)
Floating No More: Object-Ground Reconstruction from a Single Image
by: Man, Yunze, et al.
Published: (2024)
by: Man, Yunze, et al.
Published: (2024)
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
by: Wang, Shihao, et al.
Published: (2025)
by: Wang, Shihao, et al.
Published: (2025)
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
by: Wang, Shihao, et al.
Published: (2024)
by: Wang, Shihao, et al.
Published: (2024)
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
by: Huang, De-An, et al.
Published: (2025)
by: Huang, De-An, et al.
Published: (2025)
InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object Interaction
by: Xu, Sirui, et al.
Published: (2024)
by: Xu, Sirui, et al.
Published: (2024)
Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?
by: Li, Zhiqi, et al.
Published: (2023)
by: Li, Zhiqi, et al.
Published: (2023)
StreamChat: Chatting with Streaming Video
by: Liu, Jihao, et al.
Published: (2024)
by: Liu, Jihao, et al.
Published: (2024)
Capturing Visual Environment Structure Correlates with Control Performance
by: Dong, Jiahua, et al.
Published: (2026)
by: Dong, Jiahua, et al.
Published: (2026)
Frozen Transformers in Language Models Are Effective Visual Encoder Layers
by: Pang, Ziqi, et al.
Published: (2023)
by: Pang, Ziqi, et al.
Published: (2023)
PPTArena: A Benchmark for Agentic PowerPoint Editing
by: Ofengenden, Michael, et al.
Published: (2025)
by: Ofengenden, Michael, et al.
Published: (2025)
Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation
by: Li, Zhenxin, et al.
Published: (2024)
by: Li, Zhenxin, et al.
Published: (2024)
Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation
by: Lee, Junha, et al.
Published: (2025)
by: Lee, Junha, et al.
Published: (2025)
Improving Distant 3D Object Detection Using 2D Box Supervision
by: Yang, Zetong, et al.
Published: (2024)
by: Yang, Zetong, et al.
Published: (2024)
MVPaint: Synchronized Multi-View Diffusion for Painting Anything 3D
by: Cheng, Wei, et al.
Published: (2024)
by: Cheng, Wei, et al.
Published: (2024)
SAM 3D: 3Dfy Anything in Images
by: SAM 3D Team, et al.
Published: (2025)
by: SAM 3D Team, et al.
Published: (2025)
Gaussian Grouping: Segment and Edit Anything in 3D Scenes
by: Ye, Mingqiao, et al.
Published: (2023)
by: Ye, Mingqiao, et al.
Published: (2023)
ElectricSight: 3D Hazard Monitoring for Power Lines Using Low-Cost Sensors
by: Li, Xingchen, et al.
Published: (2025)
by: Li, Xingchen, et al.
Published: (2025)
Omni-3DEdit: Generalized Versatile 3D Editing in One-Pass
by: Liyi, Chen, et al.
Published: (2026)
by: Liyi, Chen, et al.
Published: (2026)
3DGS-Drag: Dragging Gaussians for Intuitive Point-Based 3D Editing
by: Dong, Jiahua, et al.
Published: (2026)
by: Dong, Jiahua, et al.
Published: (2026)
COLMAP-Free 3D Gaussian Splatting
by: Fu, Yang, et al.
Published: (2023)
by: Fu, Yang, et al.
Published: (2023)
SAGD: Boundary-Enhanced Segment Anything in 3D Gaussian via Gaussian Decomposition
by: Hu, Xu, et al.
Published: (2024)
by: Hu, Xu, et al.
Published: (2024)
Detect Anything 3D in the Wild
by: Zhang, Hanxue, et al.
Published: (2025)
by: Zhang, Hanxue, et al.
Published: (2025)
STP4D: Spatio-Temporal-Prompt Consistent Modeling for Text-to-4D Gaussian Splatting
by: Deng, Yunze, et al.
Published: (2025)
by: Deng, Yunze, et al.
Published: (2025)
Fast Multi-view Consistent 3D Editing with Video Priors
by: Chen, Liyi, et al.
Published: (2025)
by: Chen, Liyi, et al.
Published: (2025)
Instruct 4D-to-4D: Editing 4D Scenes as Pseudo-3D Scenes Using 2D Diffusion
by: Mou, Linzhan, et al.
Published: (2024)
by: Mou, Linzhan, et al.
Published: (2024)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
by: Wang, Shihao, et al.
Published: (2025)
by: Wang, Shihao, et al.
Published: (2025)
Hiding in Plain Sight: Finding MAHA on Reddit
by: Ahmed, Sabit, et al.
Published: (2026)
by: Ahmed, Sabit, et al.
Published: (2026)
3D Aware Region Prompted Vision Language Model
by: Cheng, An-Chieh, et al.
Published: (2025)
by: Cheng, An-Chieh, et al.
Published: (2025)
Lift3D: Zero-Shot Lifting of Any 2D Vision Model to 3D
by: T, Mukund Varma, et al.
Published: (2024)
by: T, Mukund Varma, et al.
Published: (2024)
SAM3D: Zero-Shot 3D Object Detection via Segment Anything Model
by: Zhang, Dingyuan, et al.
Published: (2023)
by: Zhang, Dingyuan, et al.
Published: (2023)
Material Anything: Generating Materials for Any 3D Object via Diffusion
by: Huang, Xin, et al.
Published: (2024)
by: Huang, Xin, et al.
Published: (2024)
Similar Items
-
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
by: Wang, Shihao, et al.
Published: (2026) -
Situational Awareness Matters in 3D Vision Language Reasoning
by: Man, Yunze, et al.
Published: (2024) -
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
by: Man, Yunze, et al.
Published: (2025) -
Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding
by: Man, Yunze, et al.
Published: (2024) -
DualCross: Cross-Modality Cross-Domain Adaptation for Monocular BEV Perception
by: Man, Yunze, et al.
Published: (2023)