Enregistré dans:
| Auteurs principaux: | Rosenberg, Gillian, Stadhard, Skylar, Hansen, Bruce C., Greene, Michelle R. |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2603.26589 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Embodied Scene Understanding for Vision Language Models via MetaVQA
par: Wang, Weizhen, et autres
Publié: (2025)
par: Wang, Weizhen, et autres
Publié: (2025)
Places in the Wild: A Large, High-Resolution RAW Photograph Dataset for Ecologically Valid Vision Research
par: Greene, Michelle R.
Publié: (2026)
par: Greene, Michelle R.
Publié: (2026)
Environmental Understanding Vision-Language Model for Embodied Agent
par: Bang, Jinsik, et autres
Publié: (2026)
par: Bang, Jinsik, et autres
Publié: (2026)
JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
par: Maeda, Koki, et autres
Publié: (2026)
par: Maeda, Koki, et autres
Publié: (2026)
SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining
par: Li, Yue, et autres
Publié: (2025)
par: Li, Yue, et autres
Publié: (2025)
EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding
par: Wu, Yuqi, et autres
Publié: (2024)
par: Wu, Yuqi, et autres
Publié: (2024)
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
par: Qi, Zhangyang, et autres
Publié: (2025)
par: Qi, Zhangyang, et autres
Publié: (2025)
RLM: A Vision-Language Model Approach for Radar Scene Understanding
par: Mishra, Pushkal, et autres
Publié: (2025)
par: Mishra, Pushkal, et autres
Publié: (2025)
MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning
par: Ju, Yuanchen, et autres
Publié: (2025)
par: Ju, Yuanchen, et autres
Publié: (2025)
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
par: Shu, Yan, et autres
Publié: (2025)
par: Shu, Yan, et autres
Publié: (2025)
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding
par: Fan, Yue, et autres
Publié: (2024)
par: Fan, Yue, et autres
Publié: (2024)
VEOcc: Voxel-Centric Online Semantic Occupancy Prediction For Embodied Scene Understanding
par: Wang, Ruoyu, et autres
Publié: (2026)
par: Wang, Ruoyu, et autres
Publié: (2026)
Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
par: Ma, Jingtian, et autres
Publié: (2025)
par: Ma, Jingtian, et autres
Publié: (2025)
VL-Reader: Vision and Language Reconstructor is an Effective Scene Text Recognizer
par: Zhong, Humen, et autres
Publié: (2024)
par: Zhong, Humen, et autres
Publié: (2024)
RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
par: Liu, Hanqing, et autres
Publié: (2026)
par: Liu, Hanqing, et autres
Publié: (2026)
From Scan to Action: Leveraging Realistic Scans for Embodied Scene Understanding
par: Halacheva, Anna-Maria, et autres
Publié: (2025)
par: Halacheva, Anna-Maria, et autres
Publié: (2025)
Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
par: Yang, Ganlin, et autres
Publié: (2025)
par: Yang, Ganlin, et autres
Publié: (2025)
Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models
par: Ranasinghe, Yasiru, et autres
Publié: (2025)
par: Ranasinghe, Yasiru, et autres
Publié: (2025)
Dynamic Scene Understanding from Vision-Language Representations
par: Pruss, Shahaf, et autres
Publié: (2025)
par: Pruss, Shahaf, et autres
Publié: (2025)
GaussianVLM: Scene-centric 3D Vision-Language Models using Language-aligned Gaussian Splats for Embodied Reasoning and Beyond
par: Halacheva, Anna-Maria, et autres
Publié: (2025)
par: Halacheva, Anna-Maria, et autres
Publié: (2025)
Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models
par: Li, Yueyan, et autres
Publié: (2025)
par: Li, Yueyan, et autres
Publié: (2025)
SceneGPT: A Language Model for 3D Scene Understanding
par: Chandhok, Shivam
Publié: (2024)
par: Chandhok, Shivam
Publié: (2024)
VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?
par: Liu, Qing'an, et autres
Publié: (2026)
par: Liu, Qing'an, et autres
Publié: (2026)
Descrip3D: Enhancing Large Language Model-based 3D Scene Understanding with Object-Level Text Descriptions
par: Xue, Jintang, et autres
Publié: (2025)
par: Xue, Jintang, et autres
Publié: (2025)
Vision-Language Integration for Zero-Shot Scene Understanding in Real-World Environments
par: Rajiv, Manjunath Prasad Holenarasipura, et autres
Publié: (2025)
par: Rajiv, Manjunath Prasad Holenarasipura, et autres
Publié: (2025)
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes
par: Rivera, Antonio Carlos, et autres
Publié: (2024)
par: Rivera, Antonio Carlos, et autres
Publié: (2024)
Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
par: Li, Haoyuan, et autres
Publié: (2025)
par: Li, Haoyuan, et autres
Publié: (2025)
Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models
par: Mohamud, Safaa Abdullahi Moallim, et autres
Publié: (2025)
par: Mohamud, Safaa Abdullahi Moallim, et autres
Publié: (2025)
Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding
par: Sharma, Shivam, et autres
Publié: (2026)
par: Sharma, Shivam, et autres
Publié: (2026)
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
par: Liu, Parker, et autres
Publié: (2025)
par: Liu, Parker, et autres
Publié: (2025)
Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models
par: Zhang, Jiyao, et autres
Publié: (2026)
par: Zhang, Jiyao, et autres
Publié: (2026)
Bharat Scene Text: A Novel Comprehensive Dataset and Benchmark for Indian Language Scene Text Understanding
par: De, Anik, et autres
Publié: (2025)
par: De, Anik, et autres
Publié: (2025)
Enhancing Vision-Language Models with Scene Graphs for Traffic Accident Understanding
par: Lohner, Aaron, et autres
Publié: (2024)
par: Lohner, Aaron, et autres
Publié: (2024)
Scene Change Detection with Vision-Language Representation Learning
par: Sheng, Diwei, et autres
Publié: (2026)
par: Sheng, Diwei, et autres
Publié: (2026)
Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering
par: Li, Zechuan, et autres
Publié: (2025)
par: Li, Zechuan, et autres
Publié: (2025)
LET-US: Long Event-Text Understanding of Scenes
par: Chen, Rui, et autres
Publié: (2025)
par: Chen, Rui, et autres
Publié: (2025)
Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models
par: Berman, Nimrod, et autres
Publié: (2025)
par: Berman, Nimrod, et autres
Publié: (2025)
Vision-Language Models for Autonomous Driving: CLIP-Based Dynamic Scene Understanding
par: Elhenawy, Mohammed, et autres
Publié: (2025)
par: Elhenawy, Mohammed, et autres
Publié: (2025)
CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model
par: Zhao, Shuai, et autres
Publié: (2023)
par: Zhao, Shuai, et autres
Publié: (2023)
HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model
par: Li, Chen, et autres
Publié: (2025)
par: Li, Chen, et autres
Publié: (2025)
Documents similaires
-
Embodied Scene Understanding for Vision Language Models via MetaVQA
par: Wang, Weizhen, et autres
Publié: (2025) -
Places in the Wild: A Large, High-Resolution RAW Photograph Dataset for Ecologically Valid Vision Research
par: Greene, Michelle R.
Publié: (2026) -
Environmental Understanding Vision-Language Model for Embodied Agent
par: Bang, Jinsik, et autres
Publié: (2026) -
JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
par: Maeda, Koki, et autres
Publié: (2026) -
SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining
par: Li, Yue, et autres
Publié: (2025)