Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Loginova, Olga, Loguinova, Sofía Ortega |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CAST: Cross-modal Alignment Similarity Test for Vision Language Models
von: Dagan, Gautier, et al.
Veröffentlicht: (2024)
von: Dagan, Gautier, et al.
Veröffentlicht: (2024)
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models
von: Liang, Yiming, et al.
Veröffentlicht: (2026)
von: Liang, Yiming, et al.
Veröffentlicht: (2026)
Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging
von: Fu, Zihang, et al.
Veröffentlicht: (2026)
von: Fu, Zihang, et al.
Veröffentlicht: (2026)
TemMed-Bench: Evaluating Temporal Medical Image Reasoning in Vision-Language Models
von: Zhang, Junyi, et al.
Veröffentlicht: (2025)
von: Zhang, Junyi, et al.
Veröffentlicht: (2025)
Temporal Reasoning Transfer from Text to Video
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
How to Correctly Make Mistakes: A Framework for Constructing and Benchmarking Mistake Aware Egocentric Procedural Videos
von: Loginova, Olga, et al.
Veröffentlicht: (2026)
von: Loginova, Olga, et al.
Veröffentlicht: (2026)
ProGAL-VLA: Grounded Alignment through Prospective Reasoning in Vision-Language-Action Models
von: Darabi, Nastaran, et al.
Veröffentlicht: (2026)
von: Darabi, Nastaran, et al.
Veröffentlicht: (2026)
LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation
von: Irawan, Patrick Amadeus, et al.
Veröffentlicht: (2026)
von: Irawan, Patrick Amadeus, et al.
Veröffentlicht: (2026)
CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models
von: Poppi, Tobia, et al.
Veröffentlicht: (2026)
von: Poppi, Tobia, et al.
Veröffentlicht: (2026)
Demographic and Linguistic Bias Evaluation in Omnimodal Language Models
von: Elobaid, Alaa
Veröffentlicht: (2026)
von: Elobaid, Alaa
Veröffentlicht: (2026)
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
von: Li, Yun, et al.
Veröffentlicht: (2025)
von: Li, Yun, et al.
Veröffentlicht: (2025)
Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
von: Wei, Yana, et al.
Veröffentlicht: (2025)
von: Wei, Yana, et al.
Veröffentlicht: (2025)
STAIR: Spatial-Temporal Reasoning with Auditable Intermediate Results for Video Question Answering
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
Tracing the Arrow of Time: Diagnosing Temporal Information Flow in Video-LLMs
von: Han, Peitao, et al.
Veröffentlicht: (2026)
von: Han, Peitao, et al.
Veröffentlicht: (2026)
Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
von: Fei, Hao, et al.
Veröffentlicht: (2024)
von: Fei, Hao, et al.
Veröffentlicht: (2024)
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
von: Hannan, Tanveer, et al.
Veröffentlicht: (2024)
von: Hannan, Tanveer, et al.
Veröffentlicht: (2024)
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
von: Wang, Ye, et al.
Veröffentlicht: (2025)
von: Wang, Ye, et al.
Veröffentlicht: (2025)
LLM Post-Training: A Deep Dive into Reasoning Large Language Models
von: Kumar, Komal, et al.
Veröffentlicht: (2025)
von: Kumar, Komal, et al.
Veröffentlicht: (2025)
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
von: Zhu, Wenxin, et al.
Veröffentlicht: (2025)
von: Zhu, Wenxin, et al.
Veröffentlicht: (2025)
The Visual Iconicity Challenge: Evaluating Vision-Language Models on Sign Language Form-Meaning Mapping
von: Keleş, Onur, et al.
Veröffentlicht: (2025)
von: Keleş, Onur, et al.
Veröffentlicht: (2025)
Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map Alignment
von: Rassin, Royi, et al.
Veröffentlicht: (2023)
von: Rassin, Royi, et al.
Veröffentlicht: (2023)
On the Cultural Anachronism and Temporal Reasoning in Vision Language Models
von: Ranjan, Mukul, et al.
Veröffentlicht: (2026)
von: Ranjan, Mukul, et al.
Veröffentlicht: (2026)
Are Vision Language Models Cross-Cultural Theory of Mind Reasoners?
von: Nazi, Zabir Al, et al.
Veröffentlicht: (2025)
von: Nazi, Zabir Al, et al.
Veröffentlicht: (2025)
Evaluating Reasoning Faithfulness in Medical Vision-Language Models using Multimodal Perturbations
von: Moll, Johannes, et al.
Veröffentlicht: (2025)
von: Moll, Johannes, et al.
Veröffentlicht: (2025)
SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing
von: Biyyala, Varun, et al.
Veröffentlicht: (2025)
von: Biyyala, Varun, et al.
Veröffentlicht: (2025)
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
von: Zhu, Kejian, et al.
Veröffentlicht: (2025)
von: Zhu, Kejian, et al.
Veröffentlicht: (2025)
RiskCueBench: Benchmarking Anticipatory Reasoning from Early Risk Cues in Video-Language Models
von: Luo, Sha, et al.
Veröffentlicht: (2026)
von: Luo, Sha, et al.
Veröffentlicht: (2026)
Cerberus: Real-Time Video Anomaly Detection via Cascaded Vision-Language Models
von: Zheng, Yue, et al.
Veröffentlicht: (2025)
von: Zheng, Yue, et al.
Veröffentlicht: (2025)
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
von: Cheng, Zesen, et al.
Veröffentlicht: (2024)
von: Cheng, Zesen, et al.
Veröffentlicht: (2024)
VidCoM: Fast Video Comprehension through Large Language Models with Multimodal Tools
von: Qi, Ji, et al.
Veröffentlicht: (2023)
von: Qi, Ji, et al.
Veröffentlicht: (2023)
TimeRefine: Temporal Grounding with Time Refining Video LLM
von: Wang, Xizi, et al.
Veröffentlicht: (2024)
von: Wang, Xizi, et al.
Veröffentlicht: (2024)
VL-GLUE: A Suite of Fundamental yet Challenging Visuo-Linguistic Reasoning Tasks
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
Arbitration Failure, Not Perceptual Blindness: How Vision-Language Models Resolve Visual-Linguistic Conflicts
von: Nooralahzadeh, Farhad, et al.
Veröffentlicht: (2026)
von: Nooralahzadeh, Farhad, et al.
Veröffentlicht: (2026)
GenieBlue: Integrating both Linguistic and Multimodal Capabilities for Large Language Models on Mobile Devices
von: Lu, Xudong, et al.
Veröffentlicht: (2025)
von: Lu, Xudong, et al.
Veröffentlicht: (2025)
Visual Merit or Linguistic Crutch? A Close Look at DeepSeek-OCR
von: Liang, Yunhao, et al.
Veröffentlicht: (2026)
von: Liang, Yunhao, et al.
Veröffentlicht: (2026)
Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning
von: Li, Kailing, et al.
Veröffentlicht: (2025)
von: Li, Kailing, et al.
Veröffentlicht: (2025)
Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning
von: Dogan, Mustafa, et al.
Veröffentlicht: (2024)
von: Dogan, Mustafa, et al.
Veröffentlicht: (2024)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks
von: Ashraf, Tajamul, et al.
Veröffentlicht: (2025)
von: Ashraf, Tajamul, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CAST: Cross-modal Alignment Similarity Test for Vision Language Models
von: Dagan, Gautier, et al.
Veröffentlicht: (2024) -
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models
von: Liang, Yiming, et al.
Veröffentlicht: (2026) -
Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging
von: Fu, Zihang, et al.
Veröffentlicht: (2026) -
TemMed-Bench: Evaluating Temporal Medical Image Reasoning in Vision-Language Models
von: Zhang, Junyi, et al.
Veröffentlicht: (2025) -
Temporal Reasoning Transfer from Text to Video
von: Li, Lei, et al.
Veröffentlicht: (2024)