An Empirical Study on How Video-LLMs Answer Video Questions
Fuente:
arXiv
Saved in:
| Main Authors: | Gou, Chenhui, Ma, Ziyu, Duan, Zicheng, He, Haoyu, Chen, Feng, Liu, Akide, Zhuang, Bohan, Cai, Jianfei, Rezatofighi, Hamid |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DrVideo: Document Retrieval Based Long Video Understanding
by: Ma, Ziyu, et al.
Published: (2024)
by: Ma, Ziyu, et al.
Published: (2024)
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
by: Shaker, Abdelrahman, et al.
Published: (2025)
by: Shaker, Abdelrahman, et al.
Published: (2025)
How Well Can Vision Language Models See Image Details?
by: Gou, Chenhui, et al.
Published: (2024)
by: Gou, Chenhui, et al.
Published: (2024)
VQ-VA World: Towards High-Quality Visual Question-Visual Answering
by: Gou, Chenhui, et al.
Published: (2025)
by: Gou, Chenhui, et al.
Published: (2025)
JRDB-PanoTrack: An Open-world Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments
by: Le, Duy-Tho, et al.
Published: (2024)
by: Le, Duy-Tho, et al.
Published: (2024)
Marginalized Generalized IoU (MGIoU): A Unified Objective Function for Optimizing Any Convex Parametric Shapes
by: Le, Duy-Tho, et al.
Published: (2025)
by: Le, Duy-Tho, et al.
Published: (2025)
DifFUSER: Diffusion Model for Robust Multi-Sensor Fusion in 3D Object Detection and BEV Segmentation
by: Le, Duy-Tho, et al.
Published: (2024)
by: Le, Duy-Tho, et al.
Published: (2024)
LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models
by: Duan, Zicheng, et al.
Published: (2026)
by: Duan, Zicheng, et al.
Published: (2026)
ASAP-Textured Gaussians: Enhancing Textured Gaussians with Adaptive Sampling and Anisotropic Parameterization
by: Wei, Meng, et al.
Published: (2025)
by: Wei, Meng, et al.
Published: (2025)
Normal-GS: 3D Gaussian Splatting with Normal-Involved Rendering
by: Wei, Meng, et al.
Published: (2024)
by: Wei, Meng, et al.
Published: (2024)
Efficient Stitchable Task Adaptation
by: He, Haoyu, et al.
Published: (2023)
by: He, Haoyu, et al.
Published: (2023)
LongVLM: Efficient Long Video Understanding via Large Language Models
by: Weng, Yuetian, et al.
Published: (2024)
by: Weng, Yuetian, et al.
Published: (2024)
OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
by: Chen, Feng, et al.
Published: (2025)
by: Chen, Feng, et al.
Published: (2025)
Streaming Video Diffusion: Online Video Editing with Diffusion Models
by: Chen, Feng, et al.
Published: (2024)
by: Chen, Feng, et al.
Published: (2024)
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
by: Chen, Zhuokun, et al.
Published: (2026)
by: Chen, Zhuokun, et al.
Published: (2026)
CoV: Chain-of-View Prompting for Spatial Reasoning
by: Zhao, Haoyu, et al.
Published: (2026)
by: Zhao, Haoyu, et al.
Published: (2026)
GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering
by: Ma, Ziyu, et al.
Published: (2024)
by: Ma, Ziyu, et al.
Published: (2024)
ReCA: Multi-Shot Long Video Extrapolation via Recursive Context Allocation
by: Liu, Akide, et al.
Published: (2026)
by: Liu, Akide, et al.
Published: (2026)
Pruning Self-attentions into Convolutional Layers in Single Path
by: He, Haoyu, et al.
Published: (2021)
by: He, Haoyu, et al.
Published: (2021)
Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-Answering
by: Liao, Zhaohe, et al.
Published: (2024)
by: Liao, Zhaohe, et al.
Published: (2024)
Video Detective: Seek Critical Clues Recurrently to Answer Question from Long Videos
by: Du, Henghui, et al.
Published: (2025)
by: Du, Henghui, et al.
Published: (2025)
Admitting Ignorance Helps the Video Question Answering Models to Answer
by: Li, Haopeng, et al.
Published: (2025)
by: Li, Haopeng, et al.
Published: (2025)
EZIGen: Enhancing zero-shot personalized image generation with precise subject encoding and decoupled guidance
by: Duan, Zicheng, et al.
Published: (2024)
by: Duan, Zicheng, et al.
Published: (2024)
Motion Mamba: Efficient and Long Sequence Motion Generation
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
JRDB-Social: A Multifaceted Robotic Dataset for Understanding of Context and Dynamics of Human Interactions Within Social Groups
by: Jahangard, Simindokht, et al.
Published: (2024)
by: Jahangard, Simindokht, et al.
Published: (2024)
VideoQA in the Era of LLMs: An Empirical Study
by: Xiao, Junbin, et al.
Published: (2024)
by: Xiao, Junbin, et al.
Published: (2024)
FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion
by: Liu, Akide, et al.
Published: (2025)
by: Liu, Akide, et al.
Published: (2025)
ZPressor: Bottleneck-Aware Compression for Scalable Feed-Forward 3DGS
by: Wang, Weijie, et al.
Published: (2025)
by: Wang, Weijie, et al.
Published: (2025)
Physics-Grounded Motion Forecasting via Equation Discovery for Trajectory-Guided Image-to-Video Generation
by: Feng, Tao, et al.
Published: (2025)
by: Feng, Tao, et al.
Published: (2025)
InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
Social-MAE: Social Masked Autoencoder for Multi-person Motion Representation Learning
by: Ehsanpour, Mahsa, et al.
Published: (2024)
by: Ehsanpour, Mahsa, et al.
Published: (2024)
How Important are Videos for Training Video LLMs?
by: Lydakis, George, et al.
Published: (2025)
by: Lydakis, George, et al.
Published: (2025)
Less Detail, Better Answers: Degradation-Driven Prompting for VQA
by: Han, Haoxuan, et al.
Published: (2026)
by: Han, Haoxuan, et al.
Published: (2026)
From Priors to Perception: Grounding Video-LLMs in Physical Reality
by: Zhao, Zicheng, et al.
Published: (2026)
by: Zhao, Zicheng, et al.
Published: (2026)
Point-Cache: Test-time Dynamic and Hierarchical Cache for Robust and Generalizable Point Cloud Analysis
by: Sun, Hongyu, et al.
Published: (2025)
by: Sun, Hongyu, et al.
Published: (2025)
Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models
by: Yoon, Eunseop, et al.
Published: (2025)
by: Yoon, Eunseop, et al.
Published: (2025)
RISE-Video: Can Video Generators Decode Implicit World Rules?
by: Liu, Mingxin, et al.
Published: (2026)
by: Liu, Mingxin, et al.
Published: (2026)
Steering and Rectifying Latent Representation Manifolds in Frozen Multi-modal LLMs for Video Anomaly Detection
by: Cai, Zhaolin, et al.
Published: (2026)
by: Cai, Zhaolin, et al.
Published: (2026)
JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
by: Jahangard, Simindokht, et al.
Published: (2025)
by: Jahangard, Simindokht, et al.
Published: (2025)
dinov3.seg: Open-Vocabulary Semantic Segmentation with DINOv3
by: Dutta, Saikat, et al.
Published: (2026)
by: Dutta, Saikat, et al.
Published: (2026)
Similar Items
-
DrVideo: Document Retrieval Based Long Video Understanding
by: Ma, Ziyu, et al.
Published: (2024) -
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
by: Shaker, Abdelrahman, et al.
Published: (2025) -
How Well Can Vision Language Models See Image Details?
by: Gou, Chenhui, et al.
Published: (2024) -
VQ-VA World: Towards High-Quality Visual Question-Visual Answering
by: Gou, Chenhui, et al.
Published: (2025) -
JRDB-PanoTrack: An Open-world Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments
by: Le, Duy-Tho, et al.
Published: (2024)