Do Vision--Language Models Understand 3D Scenes or Just Catalogue Objects?
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Maheshwari, Animesh, Sahu, Divyansh, Verma, Nishit |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
VLG-CBM: Training Concept Bottleneck Models with Vision-Language Guidance
par: Srivastava, Divyansh, et autres
Publié: (2024)
par: Srivastava, Divyansh, et autres
Publié: (2024)
Dynamic Scene Understanding from Vision-Language Representations
par: Pruss, Shahaf, et autres
Publié: (2025)
par: Pruss, Shahaf, et autres
Publié: (2025)
Lay-Your-Scene: Natural Scene Layout Generation with Diffusion Transformers
par: Srivastava, Divyansh, et autres
Publié: (2025)
par: Srivastava, Divyansh, et autres
Publié: (2025)
Contrastive Language-Colored Pointmap Pretraining for Unified 3D Scene Understanding
par: Mao, Ye, et autres
Publié: (2026)
par: Mao, Ye, et autres
Publié: (2026)
SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
par: Jia, Baoxiong, et autres
Publié: (2024)
par: Jia, Baoxiong, et autres
Publié: (2024)
T-3DGS: Removing Transient Objects for 3D Scene Reconstruction
par: Markin, Alexander, et autres
Publié: (2024)
par: Markin, Alexander, et autres
Publié: (2024)
ESGNN: Towards Equivariant Scene Graph Neural Network for 3D Scene Understanding
par: Pham, Quang P. M., et autres
Publié: (2024)
par: Pham, Quang P. M., et autres
Publié: (2024)
Calib3D: Calibrating Model Preferences for Reliable 3D Scene Understanding
par: Kong, Lingdong, et autres
Publié: (2024)
par: Kong, Lingdong, et autres
Publié: (2024)
Understanding Task Transfer in Vision-Language Models
par: Sachdeva, Bhuvan, et autres
Publié: (2025)
par: Sachdeva, Bhuvan, et autres
Publié: (2025)
Bootstraping Clustering of Gaussians for View-consistent 3D Scene Understanding
par: Zhang, Wenbo, et autres
Publié: (2024)
par: Zhang, Wenbo, et autres
Publié: (2024)
Dream2Real: Zero-Shot 3D Object Rearrangement with Vision-Language Models
par: Kapelyukh, Ivan, et autres
Publié: (2023)
par: Kapelyukh, Ivan, et autres
Publié: (2023)
TESGNN: Temporal Equivariant Scene Graph Neural Networks for Efficient and Robust Multi-View 3D Scene Understanding
par: Pham, Quang P. M., et autres
Publié: (2024)
par: Pham, Quang P. M., et autres
Publié: (2024)
EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding
par: Wu, Yuqi, et autres
Publié: (2024)
par: Wu, Yuqi, et autres
Publié: (2024)
Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving
par: Gao, Haoxiang, et autres
Publié: (2025)
par: Gao, Haoxiang, et autres
Publié: (2025)
Do Language Models Understand Time?
par: Ding, Xi, et autres
Publié: (2024)
par: Ding, Xi, et autres
Publié: (2024)
HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
par: Bao, Chen, et autres
Publié: (2024)
par: Bao, Chen, et autres
Publié: (2024)
DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding
par: Zhuo, Dong, et autres
Publié: (2026)
par: Zhuo, Dong, et autres
Publié: (2026)
Towards Understanding How Knowledge Evolves in Large Vision-Language Models
par: Wang, Sudong, et autres
Publié: (2025)
par: Wang, Sudong, et autres
Publié: (2025)
Exploring the Limits of Zero Shot Vision Language Models for Hate Meme Detection: The Vulnerabilities and their Interpretations
par: Rizwan, Naquee, et autres
Publié: (2024)
par: Rizwan, Naquee, et autres
Publié: (2024)
HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models
par: Zohrabi, Reihaneh, et autres
Publié: (2026)
par: Zohrabi, Reihaneh, et autres
Publié: (2026)
Adapt3R: Adaptive 3D Scene Representation for Domain Transfer in Imitation Learning
par: Wilcox, Albert, et autres
Publié: (2025)
par: Wilcox, Albert, et autres
Publié: (2025)
Beyond Human Vision: The Role of Large Vision Language Models in Microscope Image Analysis
par: Verma, Prateek, et autres
Publié: (2024)
par: Verma, Prateek, et autres
Publié: (2024)
Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving
par: Kong, Lingdong, et autres
Publié: (2024)
par: Kong, Lingdong, et autres
Publié: (2024)
VoteSplat: Hough Voting Gaussian Splatting for 3D Scene Understanding
par: Jiang, Minchao, et autres
Publié: (2025)
par: Jiang, Minchao, et autres
Publié: (2025)
U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound Understanding
par: Le, Anjie, et autres
Publié: (2025)
par: Le, Anjie, et autres
Publié: (2025)
Enhancing Large Vision Model in Street Scene Semantic Understanding through Leveraging Posterior Optimization Trajectory
par: Kou, Wei-Bin, et autres
Publié: (2025)
par: Kou, Wei-Bin, et autres
Publié: (2025)
InseRF: Text-Driven Generative Object Insertion in Neural 3D Scenes
par: Shahbazi, Mohamad, et autres
Publié: (2024)
par: Shahbazi, Mohamad, et autres
Publié: (2024)
What Makes "Good" Distractors for Object Hallucination Evaluation in Large Vision-Language Models?
par: Xie, Ming-Kun, et autres
Publié: (2025)
par: Xie, Ming-Kun, et autres
Publié: (2025)
ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
par: Li, Zhaoyang, et autres
Publié: (2025)
par: Li, Zhaoyang, et autres
Publié: (2025)
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
par: Zhou, Yiyang, et autres
Publié: (2023)
par: Zhou, Yiyang, et autres
Publié: (2023)
Uncertainty-Aware Vision-Language Segmentation for Medical Imaging
par: Das, Aryan, et autres
Publié: (2026)
par: Das, Aryan, et autres
Publié: (2026)
Learning 3D Scene Analogies with Neural Contextual Scene Maps
par: Kim, Junho, et autres
Publié: (2025)
par: Kim, Junho, et autres
Publié: (2025)
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
par: Sui, Elaine, et autres
Publié: (2024)
par: Sui, Elaine, et autres
Publié: (2024)
Vision-Based Natural Language Scene Understanding for Autonomous Driving: An Extended Dataset and a New Model for Traffic Scene Description Generation
par: Zadeh, Danial Sadrian, et autres
Publié: (2026)
par: Zadeh, Danial Sadrian, et autres
Publié: (2026)
Sampling 3D Gaussian Scenes in Seconds with Latent Diffusion Models
par: Henderson, Paul, et autres
Publié: (2024)
par: Henderson, Paul, et autres
Publié: (2024)
Toward Autonomous Laboratory Safety Monitoring with Vision Language Models: Learning to See Hazards Through Scene Structure
par: Chakraborty, Trishna, et autres
Publié: (2026)
par: Chakraborty, Trishna, et autres
Publié: (2026)
Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Models
par: Wu, Ziyi, et autres
Publié: (2024)
par: Wu, Ziyi, et autres
Publié: (2024)
SIP: Site in Pieces- A Dataset of Disaggregated Construction-Phase 3D Scans for Semantic Segmentation and Scene Understanding
par: Kim, Seongyong, et autres
Publié: (2025)
par: Kim, Seongyong, et autres
Publié: (2025)
Large Language Models Implicitly Learn to See and Hear Just By Reading
par: Verma, Prateek, et autres
Publié: (2025)
par: Verma, Prateek, et autres
Publié: (2025)
Sources of Uncertainty in 3D Scene Reconstruction
par: Klasson, Marcus, et autres
Publié: (2024)
par: Klasson, Marcus, et autres
Publié: (2024)
Documents similaires
-
VLG-CBM: Training Concept Bottleneck Models with Vision-Language Guidance
par: Srivastava, Divyansh, et autres
Publié: (2024) -
Dynamic Scene Understanding from Vision-Language Representations
par: Pruss, Shahaf, et autres
Publié: (2025) -
Lay-Your-Scene: Natural Scene Layout Generation with Diffusion Transformers
par: Srivastava, Divyansh, et autres
Publié: (2025) -
Contrastive Language-Colored Pointmap Pretraining for Unified 3D Scene Understanding
par: Mao, Ye, et autres
Publié: (2026) -
SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
par: Jia, Baoxiong, et autres
Publié: (2024)