A Unified Framework for 3D Scene Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Wei, Shi, Chunsheng, Tu, Sifan, Zhou, Xin, Liang, Dingkang, Bai, Xiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
von: Zhou, Xin, et al.
Veröffentlicht: (2026)
von: Zhou, Xin, et al.
Veröffentlicht: (2026)
PointTPA: Dynamic Network Parameter Adaptation for 3D Scene Understanding
von: Liu, Siyuan, et al.
Veröffentlicht: (2026)
von: Liu, Siyuan, et al.
Veröffentlicht: (2026)
A Unified Image-Dense Annotation Generation Model for Underwater Scenes
von: Lin, Hongkai, et al.
Veröffentlicht: (2025)
von: Lin, Hongkai, et al.
Veröffentlicht: (2025)
The Role of World Models in Shaping Autonomous Driving: A Comprehensive Survey
von: Tu, Sifan, et al.
Veröffentlicht: (2025)
von: Tu, Sifan, et al.
Veröffentlicht: (2025)
NAUTILUS: A Large Multimodal Model for Underwater Scene Understanding
von: Xu, Wei, et al.
Veröffentlicht: (2025)
von: Xu, Wei, et al.
Veröffentlicht: (2025)
SOOD++: Leveraging Unlabeled Data to Boost Oriented Object Detection
von: Liang, Dingkang, et al.
Veröffentlicht: (2024)
von: Liang, Dingkang, et al.
Veröffentlicht: (2024)
UniFuture: A 4D Driving World Model for Future Generation and Perception
von: Liang, Dingkang, et al.
Veröffentlicht: (2025)
von: Liang, Dingkang, et al.
Veröffentlicht: (2025)
More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models
von: Lin, Hongkai, et al.
Veröffentlicht: (2025)
von: Lin, Hongkai, et al.
Veröffentlicht: (2025)
Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding
von: Wu, Xianjin, et al.
Veröffentlicht: (2026)
von: Wu, Xianjin, et al.
Veröffentlicht: (2026)
AV-Unified: A Unified Framework for Audio-visual Scene Understanding
von: Li, Guangyao, et al.
Veröffentlicht: (2026)
von: Li, Guangyao, et al.
Veröffentlicht: (2026)
AVS-Net: Point Sampling with Adaptive Voxel Size for 3D Scene Understanding
von: Yang, Hongcheng, et al.
Veröffentlicht: (2024)
von: Yang, Hongcheng, et al.
Veröffentlicht: (2024)
Dynamic Adapter Meets Prompt Tuning: Parameter-Efficient Transfer Learning for Point Cloud Analysis
von: Zhou, Xin, et al.
Veröffentlicht: (2024)
von: Zhou, Xin, et al.
Veröffentlicht: (2024)
ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding
von: Guan, Yiran, et al.
Veröffentlicht: (2026)
von: Guan, Yiran, et al.
Veröffentlicht: (2026)
PointMamba: A Simple State Space Model for Point Cloud Analysis
von: Liang, Dingkang, et al.
Veröffentlicht: (2024)
von: Liang, Dingkang, et al.
Veröffentlicht: (2024)
DriveEditor: A Unified 3D Information-Guided Framework for Controllable Object Editing in Driving Scenes
von: Liang, Yiyuan, et al.
Veröffentlicht: (2024)
von: Liang, Yiyuan, et al.
Veröffentlicht: (2024)
Unified Semantic Transformer for 3D Scene Understanding
von: Koch, Sebastian, et al.
Veröffentlicht: (2025)
von: Koch, Sebastian, et al.
Veröffentlicht: (2025)
Parameter-Efficient Fine-Tuning in Spectral Domain for Point Cloud Learning
von: Liang, Dingkang, et al.
Veröffentlicht: (2024)
von: Liang, Dingkang, et al.
Veröffentlicht: (2024)
MINIMA: Modality Invariant Image Matching
von: Ren, Jiangwei, et al.
Veröffentlicht: (2024)
von: Ren, Jiangwei, et al.
Veröffentlicht: (2024)
Unified 3D Scene Understanding Through Physical World Modeling
von: Lee, Wanhee, et al.
Veröffentlicht: (2026)
von: Lee, Wanhee, et al.
Veröffentlicht: (2026)
PnP-U3D: Plug-and-Play 3D Framework Bridging Autoregression and Diffusion for Unified Understanding and Generation
von: Chen, Yongwei, et al.
Veröffentlicht: (2026)
von: Chen, Yongwei, et al.
Veröffentlicht: (2026)
Make Your ViT-based Multi-view 3D Detectors Faster via Token Compression
von: Zhang, Dingyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Dingyuan, et al.
Veröffentlicht: (2024)
When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
von: Sun, Zhengyang, et al.
Veröffentlicht: (2026)
von: Sun, Zhengyang, et al.
Veröffentlicht: (2026)
HUGS: Holistic Urban 3D Scene Understanding via Gaussian Splatting
von: Zhou, Hongyu, et al.
Veröffentlicht: (2024)
von: Zhou, Hongyu, et al.
Veröffentlicht: (2024)
3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
von: Huang, Ting, et al.
Veröffentlicht: (2025)
von: Huang, Ting, et al.
Veröffentlicht: (2025)
TB-HSU: Hierarchical 3D Scene Understanding with Contextual Affordances
von: Xu, Wenting, et al.
Veröffentlicht: (2024)
von: Xu, Wenting, et al.
Veröffentlicht: (2024)
You Only Look Bottom-Up for Monocular 3D Object Detection
von: Xiong, Kaixin, et al.
Veröffentlicht: (2024)
von: Xiong, Kaixin, et al.
Veröffentlicht: (2024)
Scene Reconstruction as Mapping Priors for 3D Detection
von: Fu, Yang, et al.
Veröffentlicht: (2026)
von: Fu, Yang, et al.
Veröffentlicht: (2026)
Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
von: Huang, Mingxin, et al.
Veröffentlicht: (2024)
Anomaly Detection by Adapting a pre-trained Vision Language Model
von: Cai, Yuxuan, et al.
Veröffentlicht: (2024)
von: Cai, Yuxuan, et al.
Veröffentlicht: (2024)
Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution
von: Liang, Dingkang, et al.
Veröffentlicht: (2025)
von: Liang, Dingkang, et al.
Veröffentlicht: (2025)
3D Question Answering for City Scene Understanding
von: Sun, Penglei, et al.
Veröffentlicht: (2024)
von: Sun, Penglei, et al.
Veröffentlicht: (2024)
Reg3D: Reconstructive Geometry Instruction Tuning for 3D Scene Understanding
von: Zheng, Hongpei, et al.
Veröffentlicht: (2025)
von: Zheng, Hongpei, et al.
Veröffentlicht: (2025)
Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy
von: Jingyu, Gong, et al.
Veröffentlicht: (2025)
von: Jingyu, Gong, et al.
Veröffentlicht: (2025)
GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation
von: Deng, Tianchen, et al.
Veröffentlicht: (2025)
von: Deng, Tianchen, et al.
Veröffentlicht: (2025)
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
von: Zhao, Zongchuang, et al.
Veröffentlicht: (2025)
von: Zhao, Zongchuang, et al.
Veröffentlicht: (2025)
Calib3D: Calibrating Model Preferences for Reliable 3D Scene Understanding
von: Kong, Lingdong, et al.
Veröffentlicht: (2024)
von: Kong, Lingdong, et al.
Veröffentlicht: (2024)
CAGS: Open-Vocabulary 3D Scene Understanding with Context-Aware Gaussian Splatting
von: Sun, Wei, et al.
Veröffentlicht: (2025)
von: Sun, Wei, et al.
Veröffentlicht: (2025)
Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving
von: Kong, Lingdong, et al.
Veröffentlicht: (2024)
von: Kong, Lingdong, et al.
Veröffentlicht: (2024)
SAM3D: Zero-Shot 3D Object Detection via Segment Anything Model
von: Zhang, Dingyuan, et al.
Veröffentlicht: (2023)
von: Zhang, Dingyuan, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation
von: Zhou, Xin, et al.
Veröffentlicht: (2025) -
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
von: Zhou, Xin, et al.
Veröffentlicht: (2026) -
PointTPA: Dynamic Network Parameter Adaptation for 3D Scene Understanding
von: Liu, Siyuan, et al.
Veröffentlicht: (2026) -
A Unified Image-Dense Annotation Generation Model for Underwater Scenes
von: Lin, Hongkai, et al.
Veröffentlicht: (2025) -
The Role of World Models in Shaping Autonomous Driving: A Comprehensive Survey
von: Tu, Sifan, et al.
Veröffentlicht: (2025)