On the Consistency of Video Large Language Models in Temporal Comprehension
Fuente:
arXiv
Saved in:
| Main Authors: | Jung, Minjoon, Xiao, Junbin, Zhang, Byoung-Tak, Yao, Angela |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
by: Jung, Minjoon, et al.
Published: (2025)
by: Jung, Minjoon, et al.
Published: (2025)
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding
by: Jung, Minjoon, et al.
Published: (2026)
by: Jung, Minjoon, et al.
Published: (2026)
Exploring Ordinal Bias in Action Recognition for Instructional Videos
by: Kim, Joochan, et al.
Published: (2025)
by: Kim, Joochan, et al.
Published: (2025)
Background-aware Moment Detection for Video Moment Retrieval
by: Jung, Minjoon, et al.
Published: (2023)
by: Jung, Minjoon, et al.
Published: (2023)
Towards Spatially Consistent Image Generation: On Incorporating Intrinsic Scene Properties into Diffusion Models
by: Lee, Hyundo, et al.
Published: (2025)
by: Lee, Hyundo, et al.
Published: (2025)
Question-Answering Dense Video Events
by: Qin, Hangyu, et al.
Published: (2024)
by: Qin, Hangyu, et al.
Published: (2024)
PGA: Personalizing Grasping Agents with Single Human-Robot Interaction
by: Kim, Junghyun, et al.
Published: (2023)
by: Kim, Junghyun, et al.
Published: (2023)
Ego-Grounding for Personalized Question-Answering in Egocentric Videos
by: Xiao, Junbin, et al.
Published: (2026)
by: Xiao, Junbin, et al.
Published: (2026)
State-Space Hierarchical Compression with Gated Attention and Learnable Sampling for Hour-Long Video Understanding in Large Multimodal Models
by: Kim, Geewook, et al.
Published: (2025)
by: Kim, Geewook, et al.
Published: (2025)
TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models
by: Chen, Junzhe, et al.
Published: (2026)
by: Chen, Junzhe, et al.
Published: (2026)
Can I Trust Your Answer? Visually Grounded Video Question Answering
by: Xiao, Junbin, et al.
Published: (2023)
by: Xiao, Junbin, et al.
Published: (2023)
DBMovi-GS: Dynamic View Synthesis from Blurry Monocular Video via Sparse-Controlled Gaussian Splatting
by: Song, Yeon-Ji, et al.
Published: (2025)
by: Song, Yeon-Ji, et al.
Published: (2025)
OV-MAP : Open-Vocabulary Zero-Shot 3D Instance Segmentation Map for Robots
by: Kim, Juno, et al.
Published: (2025)
by: Kim, Juno, et al.
Published: (2025)
Edit Temporal-Consistent Videos with Image Diffusion Model
by: Wang, Yuanzhi, et al.
Published: (2023)
by: Wang, Yuanzhi, et al.
Published: (2023)
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
by: Xiao, Junbin, et al.
Published: (2026)
by: Xiao, Junbin, et al.
Published: (2026)
EXOT: Exit-aware Object Tracker for Safe Robotic Manipulation of Moving Object
by: Kim, Hyunseo, et al.
Published: (2023)
by: Kim, Hyunseo, et al.
Published: (2023)
Video Summarization with Large Language Models
by: Lee, Min Jung, et al.
Published: (2025)
by: Lee, Min Jung, et al.
Published: (2025)
V-CORE: Temporally Consistent Video Understanding for Video-LLM
by: Kang, Zhengjian, et al.
Published: (2026)
by: Kang, Zhengjian, et al.
Published: (2026)
ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language Models
by: Qu, Mengxue, et al.
Published: (2024)
by: Qu, Mengxue, et al.
Published: (2024)
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
by: Li, Yun, et al.
Published: (2025)
by: Li, Yun, et al.
Published: (2025)
Zero-Shot Vision-and-Language Navigation with Collision Mitigation in Continuous Environment
by: Jeong, Seongjun, et al.
Published: (2024)
by: Jeong, Seongjun, et al.
Published: (2024)
PiTe: Pixel-Temporal Alignment for Large Video-Language Model
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models
by: Souza, Rafael, et al.
Published: (2024)
by: Souza, Rafael, et al.
Published: (2024)
On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning
by: Kim, Geewook, et al.
Published: (2024)
by: Kim, Geewook, et al.
Published: (2024)
Unveiling the Tapestry of Consistency in Large Vision-Language Models
by: Zhang, Yuan, et al.
Published: (2024)
by: Zhang, Yuan, et al.
Published: (2024)
Locality-aware Concept Bottleneck Model
by: Jeon, Sujin, et al.
Published: (2025)
by: Jeon, Sujin, et al.
Published: (2025)
Temporal-Consistent Video Restoration with Pre-trained Diffusion Models
by: Wang, Hengkang, et al.
Published: (2025)
by: Wang, Hengkang, et al.
Published: (2025)
VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
by: Zhao, Fufangchen, et al.
Published: (2025)
by: Zhao, Fufangchen, et al.
Published: (2025)
OCK: Unsupervised Dynamic Video Prediction with Object-Centric Kinematics
by: Song, Yeon-Ji, et al.
Published: (2024)
by: Song, Yeon-Ji, et al.
Published: (2024)
Spatial Degradation-Aware and Temporal Consistent Diffusion Model for Compressed Video Super-Resolution
by: An, Hongyu, et al.
Published: (2025)
by: An, Hongyu, et al.
Published: (2025)
GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding
by: Fan, Rong, et al.
Published: (2026)
by: Fan, Rong, et al.
Published: (2026)
SC-Tune: Unleashing Self-Consistent Referential Comprehension in Large Vision Language Models
by: Yue, Tongtian, et al.
Published: (2024)
by: Yue, Tongtian, et al.
Published: (2024)
ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
by: Wang, Yueqian, et al.
Published: (2025)
by: Wang, Yueqian, et al.
Published: (2025)
Scene-Text Grounding for Text-Based Video Question Answering
by: Zhou, Sheng, et al.
Published: (2024)
by: Zhou, Sheng, et al.
Published: (2024)
Learning Temporally Consistent Video Depth from Video Diffusion Priors
by: Shao, Jiahao, et al.
Published: (2024)
by: Shao, Jiahao, et al.
Published: (2024)
Continual Vision-and-Language Navigation
by: Jeong, Seongjun, et al.
Published: (2024)
by: Jeong, Seongjun, et al.
Published: (2024)
CoDeF: Content Deformation Fields for Temporally Consistent Video Processing
by: Ouyang, Hao, et al.
Published: (2023)
by: Ouyang, Hao, et al.
Published: (2023)
Collaborative Temporal Consistency Learning for Point-supervised Natural Language Video Localization
by: Tao, Zhuo, et al.
Published: (2025)
by: Tao, Zhuo, et al.
Published: (2025)
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
by: Shi, Jiapeng, et al.
Published: (2026)
by: Shi, Jiapeng, et al.
Published: (2026)
UVCG: Leveraging Temporal Consistency for Universal Video Protection
by: Li, KaiZhou, et al.
Published: (2024)
by: Li, KaiZhou, et al.
Published: (2024)
Similar Items
-
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
by: Jung, Minjoon, et al.
Published: (2025) -
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding
by: Jung, Minjoon, et al.
Published: (2026) -
Exploring Ordinal Bias in Action Recognition for Instructional Videos
by: Kim, Joochan, et al.
Published: (2025) -
Background-aware Moment Detection for Video Moment Retrieval
by: Jung, Minjoon, et al.
Published: (2023) -
Towards Spatially Consistent Image Generation: On Incorporating Intrinsic Scene Properties into Diffusion Models
by: Lee, Hyundo, et al.
Published: (2025)