Empowering Large Language Models with 3D Situation Awareness
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Zhihao, Peng, Yibo, Ren, Jinke, Liao, Yinghong, Han, Yatong, Feng, Chun-Mei, Zhao, Hengshuang, Li, Guanbin, Cui, Shuguang, Li, Zhen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual Programming for Zero-shot Open-Vocabulary 3D Visual Grounding
by: Yuan, Zhihao, et al.
Published: (2023)
by: Yuan, Zhihao, et al.
Published: (2023)
Instance-free Text to Point Cloud Localization with Relative Position Awareness
by: Wang, Lichao, et al.
Published: (2024)
by: Wang, Lichao, et al.
Published: (2024)
Adaptive Pruning for Large Language Models with Structural Importance Awareness
by: Zheng, Haotian, et al.
Published: (2024)
by: Zheng, Haotian, et al.
Published: (2024)
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations
by: Yuan, Zhihao, et al.
Published: (2025)
by: Yuan, Zhihao, et al.
Published: (2025)
STMA: A Spatio-Temporal Memory Agent for Long-Horizon Embodied Task Planning
by: Lei, Mingcong, et al.
Published: (2025)
by: Lei, Mingcong, et al.
Published: (2025)
CLEA: Closed-Loop Embodied Agent for Enhancing Task Execution in Dynamic Environments
by: Lei, Mingcong, et al.
Published: (2025)
by: Lei, Mingcong, et al.
Published: (2025)
Let Video Teaches You More: Video-to-Image Knowledge Distillation using DEtection TRansformer for Medical Video Lesion Detection
by: Jiang, Yuncheng, et al.
Published: (2024)
by: Jiang, Yuncheng, et al.
Published: (2024)
RoboSeek: You Need to Interact with Your Objects
by: Peng, Yibo, et al.
Published: (2025)
by: Peng, Yibo, et al.
Published: (2025)
Generative Semantic Communication for Text-to-Speech Synthesis
by: Zheng, Jiahao, et al.
Published: (2024)
by: Zheng, Jiahao, et al.
Published: (2024)
Movable-Antenna Empowered AAV-Enabled Data Collection over Low-Altitude Wireless Networks
by: Zhang, Xuhui, et al.
Published: (2025)
by: Zhang, Xuhui, et al.
Published: (2025)
Latency-Aware Resource Allocation for Mobile Edge Generation and Computing via Deep Reinforcement Learning
by: Wu, Yinyu, et al.
Published: (2024)
by: Wu, Yinyu, et al.
Published: (2024)
PhyScene3D: Physically Consistent Interactive 3D Tabletop Scene Generation
by: Chen, Weixing, et al.
Published: (2026)
by: Chen, Weixing, et al.
Published: (2026)
DV-3DLane: End-to-end Multi-modal 3D Lane Detection with Dual-view Representation
by: Luo, Yueru, et al.
Published: (2024)
by: Luo, Yueru, et al.
Published: (2024)
SA-GS: Semantic-Aware Gaussian Splatting for Large Scene Reconstruction with Geometry Constrain
by: Xiong, Butian, et al.
Published: (2024)
by: Xiong, Butian, et al.
Published: (2024)
Towards Flexible 3D Perception: Object-Centric Occupancy Completion Augments 3D Object Detection
by: Zheng, Chaoda, et al.
Published: (2024)
by: Zheng, Chaoda, et al.
Published: (2024)
Generative Semantic Communication for Joint Image Transmission and Segmentation
by: Yuan, Weiwen, et al.
Published: (2024)
by: Yuan, Weiwen, et al.
Published: (2024)
UAV-Enabled Wireless Networks with Movable-Antenna Array: Flexible Beamforming and Trajectory Design
by: Liu, Wenchao, et al.
Published: (2024)
by: Liu, Wenchao, et al.
Published: (2024)
PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Models
by: Guo, Zilu, et al.
Published: (2025)
by: Guo, Zilu, et al.
Published: (2025)
Empowering Small Language Models with Factual Hallucination-Aware Reasoning for Financial Classification
by: Yuan, Han, et al.
Published: (2026)
by: Yuan, Han, et al.
Published: (2026)
ECC-PolypDet: Enhanced CenterNet with Contrastive Learning for Automatic Polyp Detection
by: Jiang, Yuncheng, et al.
Published: (2024)
by: Jiang, Yuncheng, et al.
Published: (2024)
Joint Signal Detection and Automatic Modulation Classification via Deep Learning
by: Xing, Huijun, et al.
Published: (2024)
by: Xing, Huijun, et al.
Published: (2024)
Knowledge Base Enabled Semantic Communication: A Generative Perspective
by: Ren, Jinke, et al.
Published: (2023)
by: Ren, Jinke, et al.
Published: (2023)
Latency Minimization for UAV-Enabled Federated Learning: Trajectory Design and Resource Allocation
by: Zhang, Xuhui, et al.
Published: (2024)
by: Zhang, Xuhui, et al.
Published: (2024)
Scalable Federated Unlearning via Isolated and Coded Sharding
by: Lin, Yijing, et al.
Published: (2024)
by: Lin, Yijing, et al.
Published: (2024)
RoboMemory: A Brain-inspired Multi-memory Agentic Framework for Interactive Environmental Learning in Physical Embodied Systems
by: Lei, Mingcong, et al.
Published: (2025)
by: Lei, Mingcong, et al.
Published: (2025)
Fully Test-Time Adaptation for Monocular 3D Object Detection
by: Lin, Hongbin, et al.
Published: (2024)
by: Lin, Hongbin, et al.
Published: (2024)
SDesc3D: Towards Layout-Aware 3D Indoor Scene Generation from Short Descriptions
by: Feng, Jie, et al.
Published: (2026)
by: Feng, Jie, et al.
Published: (2026)
EconAgent: Large Language Model-Empowered Agents for Simulating Macroeconomic Activities
by: Li, Nian, et al.
Published: (2023)
by: Li, Nian, et al.
Published: (2023)
MVHumanNet++: A Large-scale Dataset of Multi-view Daily Dressing Human Captures with Richer Annotations for 3D Human Digitization
by: Li, Chenghong, et al.
Published: (2025)
by: Li, Chenghong, et al.
Published: (2025)
PlaneCycle: Training-Free 2D-to-3D Lifting of Foundation Models Without Adapters
by: Yu, Yinghong, et al.
Published: (2026)
by: Yu, Yinghong, et al.
Published: (2026)
Using Large Language Models to Categorize Strategic Situations and Decipher Motivations Behind Human Behaviors
by: Xie, Yutong, et al.
Published: (2025)
by: Xie, Yutong, et al.
Published: (2025)
Survey-Free Radio Map Construction via HMM-Based Coarse-to-Fine Inference
by: Xing, Zheng, et al.
Published: (2026)
by: Xing, Zheng, et al.
Published: (2026)
Situat3DChange: Situated 3D Change Understanding Dataset for Multimodal Large Language Model
by: Liu, Ruiping, et al.
Published: (2025)
by: Liu, Ruiping, et al.
Published: (2025)
SVLL: Staged Vision-Language Learning for Physically Grounded Embodied Task Planning
by: Yang, Yuyuan, et al.
Published: (2026)
by: Yang, Yuyuan, et al.
Published: (2026)
Sequential Treatment of Iliopsoas Tendon Cysts Combined With Medial Hip Snapping by Hip Arthroscopy
by: Yanlin Li, et al.
Published: (2024)
by: Yanlin Li, et al.
Published: (2024)
Mobile-Agent-RAG: Driving Smart Multi-Agent Coordination with Contextual Knowledge Empowerment for Long-Horizon Mobile Automation
by: Zhou, Yuxiang, et al.
Published: (2025)
by: Zhou, Yuxiang, et al.
Published: (2025)
MAC: Masked Agent Collaboration Boosts Large Language Model Medical Decision-Making
by: Peng, Zhihao, et al.
Published: (2025)
by: Peng, Zhihao, et al.
Published: (2025)
Situational Awareness Matters in 3D Vision Language Reasoning
by: Man, Yunze, et al.
Published: (2024)
by: Man, Yunze, et al.
Published: (2024)
DAIAN: Deep Adaptive Intent-Aware Network for CTR Prediction in Trigger-Induced Recommendation
by: Lv, Zhihao, et al.
Published: (2026)
by: Lv, Zhihao, et al.
Published: (2026)
Traj-LLM: A New Exploration for Empowering Trajectory Prediction with Pre-trained Large Language Models
by: Lan, Zhengxing, et al.
Published: (2024)
by: Lan, Zhengxing, et al.
Published: (2024)
Similar Items
-
Visual Programming for Zero-shot Open-Vocabulary 3D Visual Grounding
by: Yuan, Zhihao, et al.
Published: (2023) -
Instance-free Text to Point Cloud Localization with Relative Position Awareness
by: Wang, Lichao, et al.
Published: (2024) -
Adaptive Pruning for Large Language Models with Structural Importance Awareness
by: Zheng, Haotian, et al.
Published: (2024) -
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations
by: Yuan, Zhihao, et al.
Published: (2025) -
STMA: A Spatio-Temporal Memory Agent for Long-Horizon Embodied Task Planning
by: Lei, Mingcong, et al.
Published: (2025)