Reasoning over Video: Evaluating How MLLMs Extract, Integrate, and Reconstruct Spatiotemporal Evidence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bang, Seunghwan, Song, Hwanjun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence
von: Qiang, Chenhui, et al.
Veröffentlicht: (2025)
von: Qiang, Chenhui, et al.
Veröffentlicht: (2025)
Video-R1: Reinforcing Video Reasoning in MLLMs
von: Feng, Kaituo, et al.
Veröffentlicht: (2025)
von: Feng, Kaituo, et al.
Veröffentlicht: (2025)
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
von: Liu, Yuanxin, et al.
Veröffentlicht: (2025)
von: Liu, Yuanxin, et al.
Veröffentlicht: (2025)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
von: Ouyang, Kun, et al.
Veröffentlicht: (2025)
von: Ouyang, Kun, et al.
Veröffentlicht: (2025)
UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?
von: Liu, Yuanxin, et al.
Veröffentlicht: (2025)
von: Liu, Yuanxin, et al.
Veröffentlicht: (2025)
Vivid4D: Improving 4D Reconstruction from Monocular Video by Video Inpainting
von: Huang, Jiaxin, et al.
Veröffentlicht: (2025)
von: Huang, Jiaxin, et al.
Veröffentlicht: (2025)
VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning
von: Wang, Qi, et al.
Veröffentlicht: (2025)
von: Wang, Qi, et al.
Veröffentlicht: (2025)
Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs
von: Zhu, Rui, et al.
Veröffentlicht: (2026)
von: Zhu, Rui, et al.
Veröffentlicht: (2026)
Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
von: Han, Su Ho, et al.
Veröffentlicht: (2025)
von: Han, Su Ho, et al.
Veröffentlicht: (2025)
Can MLLMs Reason About Visual Persuasion? Evaluating the Efficacy and Faithfulness of Reasoning
von: Lee, Naeun, et al.
Veröffentlicht: (2026)
von: Lee, Naeun, et al.
Veröffentlicht: (2026)
Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs
von: Zhao, Zijia, et al.
Veröffentlicht: (2024)
von: Zhao, Zijia, et al.
Veröffentlicht: (2024)
ST-SimDiff: Balancing Spatiotemporal Similarity and Difference for Efficient Video Understanding with MLLMs
von: Luo, Bingjun, et al.
Veröffentlicht: (2026)
von: Luo, Bingjun, et al.
Veröffentlicht: (2026)
VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification
von: Meng, Jiahao, et al.
Veröffentlicht: (2026)
von: Meng, Jiahao, et al.
Veröffentlicht: (2026)
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
von: Fang, Bo, et al.
Veröffentlicht: (2025)
von: Fang, Bo, et al.
Veröffentlicht: (2025)
Training-Free Reasoning and Reflection in MLLMs
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
Think 360°: Evaluating the Width-centric Reasoning Capability of MLLMs Beyond Depth
von: Chen, Mingrui, et al.
Veröffentlicht: (2026)
von: Chen, Mingrui, et al.
Veröffentlicht: (2026)
Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos
von: Tang, Yuqi, et al.
Veröffentlicht: (2026)
von: Tang, Yuqi, et al.
Veröffentlicht: (2026)
SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark
von: Wang, Gui, et al.
Veröffentlicht: (2026)
von: Wang, Gui, et al.
Veröffentlicht: (2026)
Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs
von: Wu, Jingze, et al.
Veröffentlicht: (2026)
von: Wu, Jingze, et al.
Veröffentlicht: (2026)
LLM-based User Profile Management for Recommender System
von: Bang, Seunghwan, et al.
Veröffentlicht: (2025)
von: Bang, Seunghwan, et al.
Veröffentlicht: (2025)
Q-HyViT: Post-Training Quantization of Hybrid Vision Transformers with Bridge Block Reconstruction for IoT Systems
von: Lee, Jemin, et al.
Veröffentlicht: (2023)
von: Lee, Jemin, et al.
Veröffentlicht: (2023)
How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024)
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024)
Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task
von: Fan, Sunqi, et al.
Veröffentlicht: (2025)
von: Fan, Sunqi, et al.
Veröffentlicht: (2025)
Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
von: Tong, Jintao, et al.
Veröffentlicht: (2025)
von: Tong, Jintao, et al.
Veröffentlicht: (2025)
Explore How to Inject Beneficial Noise in MLLMs
von: Zhu, Ruishu, et al.
Veröffentlicht: (2025)
von: Zhu, Ruishu, et al.
Veröffentlicht: (2025)
SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs
von: Gu, Bo, et al.
Veröffentlicht: (2026)
von: Gu, Bo, et al.
Veröffentlicht: (2026)
Reinforcing Consistency in Video MLLMs with Structured Rewards
von: Quan, Yihao, et al.
Veröffentlicht: (2026)
von: Quan, Yihao, et al.
Veröffentlicht: (2026)
Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding
von: Huang, Yanxiang, et al.
Veröffentlicht: (2026)
von: Huang, Yanxiang, et al.
Veröffentlicht: (2026)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
von: Yilmaz, Nilay, et al.
Veröffentlicht: (2025)
von: Yilmaz, Nilay, et al.
Veröffentlicht: (2025)
Co-speech Gesture Video Generation via Motion-Based Graph Retrieval
von: Song, Yafei, et al.
Veröffentlicht: (2025)
von: Song, Yafei, et al.
Veröffentlicht: (2025)
Touch-R1: Reinforcing Touch Reasoning in MLLMs
von: Lai, Yingxin, et al.
Veröffentlicht: (2026)
von: Lai, Yingxin, et al.
Veröffentlicht: (2026)
IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams
von: Li, Jinzhao, et al.
Veröffentlicht: (2026)
von: Li, Jinzhao, et al.
Veröffentlicht: (2026)
EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy
von: Li, Jinzhao, et al.
Veröffentlicht: (2026)
von: Li, Jinzhao, et al.
Veröffentlicht: (2026)
Exploring Spatiotemporal Feature Propagation for Video-Level Compressive Spectral Reconstruction: Dataset, Model and Benchmark
von: Cai, Lijing, et al.
Veröffentlicht: (2026)
von: Cai, Lijing, et al.
Veröffentlicht: (2026)
VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
Unhackable Temporal Rewarding for Scalable Video MLLMs
von: Yu, En, et al.
Veröffentlicht: (2025)
von: Yu, En, et al.
Veröffentlicht: (2025)
AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMs
von: Chang, Boyu, et al.
Veröffentlicht: (2026)
von: Chang, Boyu, et al.
Veröffentlicht: (2026)
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs
von: Wang, Haicheng, et al.
Veröffentlicht: (2026)
von: Wang, Haicheng, et al.
Veröffentlicht: (2026)
Incentivizing Cardiologist-Like Reasoning in MLLMs for Interpretable Echocardiographic Diagnosis
von: Qin, Yi, et al.
Veröffentlicht: (2026)
von: Qin, Yi, et al.
Veröffentlicht: (2026)
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
von: Xiao, Kelaiti, et al.
Veröffentlicht: (2025)
von: Xiao, Kelaiti, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence
von: Qiang, Chenhui, et al.
Veröffentlicht: (2025) -
Video-R1: Reinforcing Video Reasoning in MLLMs
von: Feng, Kaituo, et al.
Veröffentlicht: (2025) -
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
von: Liu, Yuanxin, et al.
Veröffentlicht: (2025) -
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
von: Ouyang, Kun, et al.
Veröffentlicht: (2025) -
UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?
von: Liu, Yuanxin, et al.
Veröffentlicht: (2025)