Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sharma, Shivam, Nagaonkar, Sankalp, Choithani, Ashish, Trivedi, Ashutosh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments
von: Nagaonkar, Sankalp, et al.
Veröffentlicht: (2025)
von: Nagaonkar, Sankalp, et al.
Veröffentlicht: (2025)
BadScan: An Architectural Backdoor Attack on Visual State Space Models
von: Deshmukh, Om Suhas, et al.
Veröffentlicht: (2024)
von: Deshmukh, Om Suhas, et al.
Veröffentlicht: (2024)
SceneGPT: A Language Model for 3D Scene Understanding
von: Chandhok, Shivam
Veröffentlicht: (2024)
von: Chandhok, Shivam
Veröffentlicht: (2024)
TM-PATHVQA:90000+ Textless Multilingual Questions for Medical Visual Question Answering
von: Rajkhowa, Tonmoy, et al.
Veröffentlicht: (2024)
von: Rajkhowa, Tonmoy, et al.
Veröffentlicht: (2024)
Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Models
von: Zhang, Jialiang, et al.
Veröffentlicht: (2026)
von: Zhang, Jialiang, et al.
Veröffentlicht: (2026)
Open-ended Hierarchical Streaming Video Understanding with Vision Language Models
von: Kang, Hyolim, et al.
Veröffentlicht: (2025)
von: Kang, Hyolim, et al.
Veröffentlicht: (2025)
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
von: Qi, Zhangyang, et al.
Veröffentlicht: (2025)
von: Qi, Zhangyang, et al.
Veröffentlicht: (2025)
Do Vision-Language Foundational models show Robust Visual Perception?
von: Chandhok, Shivam, et al.
Veröffentlicht: (2024)
von: Chandhok, Shivam, et al.
Veröffentlicht: (2024)
Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark
von: Chen, Seng Nam, et al.
Veröffentlicht: (2026)
von: Chen, Seng Nam, et al.
Veröffentlicht: (2026)
Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models
von: Lee, Jonathan, et al.
Veröffentlicht: (2025)
von: Lee, Jonathan, et al.
Veröffentlicht: (2025)
Streaming Long Video Understanding with Large Language Models
von: Qian, Rui, et al.
Veröffentlicht: (2024)
von: Qian, Rui, et al.
Veröffentlicht: (2024)
StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA
von: Hu, Yuhang, et al.
Veröffentlicht: (2025)
von: Hu, Yuhang, et al.
Veröffentlicht: (2025)
Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models
von: Berman, Nimrod, et al.
Veröffentlicht: (2025)
von: Berman, Nimrod, et al.
Veröffentlicht: (2025)
Do Vision--Language Models Understand 3D Scenes or Just Catalogue Objects?
von: Maheshwari, Animesh, et al.
Veröffentlicht: (2026)
von: Maheshwari, Animesh, et al.
Veröffentlicht: (2026)
Vision-Language Models Do Not Understand Negation
von: Alhamoud, Kumail, et al.
Veröffentlicht: (2025)
von: Alhamoud, Kumail, et al.
Veröffentlicht: (2025)
Where Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization
von: Bharadwaj, Siddhant, et al.
Veröffentlicht: (2026)
von: Bharadwaj, Siddhant, et al.
Veröffentlicht: (2026)
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
von: Liu, Parker, et al.
Veröffentlicht: (2025)
von: Liu, Parker, et al.
Veröffentlicht: (2025)
ProGAL-VLA: Grounded Alignment through Prospective Reasoning in Vision-Language-Action Models
von: Darabi, Nastaran, et al.
Veröffentlicht: (2026)
von: Darabi, Nastaran, et al.
Veröffentlicht: (2026)
Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
von: Pham, Tan-Hanh, et al.
Veröffentlicht: (2025)
von: Pham, Tan-Hanh, et al.
Veröffentlicht: (2025)
RLM: A Vision-Language Model Approach for Radar Scene Understanding
von: Mishra, Pushkal, et al.
Veröffentlicht: (2025)
von: Mishra, Pushkal, et al.
Veröffentlicht: (2025)
Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning
von: Liang, Dayong, et al.
Veröffentlicht: (2025)
von: Liang, Dayong, et al.
Veröffentlicht: (2025)
SceneGraphVLM: Dynamic Scene Graph Generation from Video with Vision-Language Models
von: Makarov, Vladislav, et al.
Veröffentlicht: (2026)
von: Makarov, Vladislav, et al.
Veröffentlicht: (2026)
Do Vision-Language Models Really Understand Visual Language?
von: Hou, Yifan, et al.
Veröffentlicht: (2024)
von: Hou, Yifan, et al.
Veröffentlicht: (2024)
Do Vision-Language Models Understand Compound Nouns?
von: Kumar, Sonal, et al.
Veröffentlicht: (2024)
von: Kumar, Sonal, et al.
Veröffentlicht: (2024)
Do Vision-Language Models Understand Visual Persuasiveness?
von: Park, Gyuwon
Veröffentlicht: (2025)
von: Park, Gyuwon
Veröffentlicht: (2025)
SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining
von: Li, Yue, et al.
Veröffentlicht: (2025)
von: Li, Yue, et al.
Veröffentlicht: (2025)
Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding
von: Ma, Ke, et al.
Veröffentlicht: (2026)
von: Ma, Ke, et al.
Veröffentlicht: (2026)
Leveraging Vision-Language Models for Improving Domain Generalization in Image Classification
von: Addepalli, Sravanti, et al.
Veröffentlicht: (2023)
von: Addepalli, Sravanti, et al.
Veröffentlicht: (2023)
Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
von: Gholami, Mohsen, et al.
Veröffentlicht: (2025)
von: Gholami, Mohsen, et al.
Veröffentlicht: (2025)
Embodied Scene Understanding for Vision Language Models via MetaVQA
von: Wang, Weizhen, et al.
Veröffentlicht: (2025)
von: Wang, Weizhen, et al.
Veröffentlicht: (2025)
VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding
von: Kim, Younggun, et al.
Veröffentlicht: (2025)
von: Kim, Younggun, et al.
Veröffentlicht: (2025)
On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning
von: Kim, Geewook, et al.
Veröffentlicht: (2024)
von: Kim, Geewook, et al.
Veröffentlicht: (2024)
Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models
von: Wang, Yuqing, et al.
Veröffentlicht: (2023)
von: Wang, Yuqing, et al.
Veröffentlicht: (2023)
Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language Understanding
von: Jeon, Yerim, et al.
Veröffentlicht: (2025)
von: Jeon, Yerim, et al.
Veröffentlicht: (2025)
CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models
von: Kim, Joowon, et al.
Veröffentlicht: (2026)
von: Kim, Joowon, et al.
Veröffentlicht: (2026)
The Limits of Learning from Pictures and Text: Vision-Language Models and Embodied Scene Understanding
von: Rosenberg, Gillian, et al.
Veröffentlicht: (2026)
von: Rosenberg, Gillian, et al.
Veröffentlicht: (2026)
Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models
von: Chen, Yangyi, et al.
Veröffentlicht: (2023)
von: Chen, Yangyi, et al.
Veröffentlicht: (2023)
FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning
von: Chen, Haodong, et al.
Veröffentlicht: (2025)
von: Chen, Haodong, et al.
Veröffentlicht: (2025)
StreamingEval: A Unified Evaluation Protocol towards Realistic Streaming Video Understanding
von: Tang, Guowei, et al.
Veröffentlicht: (2026)
von: Tang, Guowei, et al.
Veröffentlicht: (2026)
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
von: Feng, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Feng, Zhiyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments
von: Nagaonkar, Sankalp, et al.
Veröffentlicht: (2025) -
BadScan: An Architectural Backdoor Attack on Visual State Space Models
von: Deshmukh, Om Suhas, et al.
Veröffentlicht: (2024) -
SceneGPT: A Language Model for 3D Scene Understanding
von: Chandhok, Shivam
Veröffentlicht: (2024) -
TM-PATHVQA:90000+ Textless Multilingual Questions for Medical Visual Question Answering
von: Rajkhowa, Tonmoy, et al.
Veröffentlicht: (2024) -
Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Models
von: Zhang, Jialiang, et al.
Veröffentlicht: (2026)