Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yiheng, Zhu, Lichen, Lin, Yueqian, Liu, Yudong, Zhang, Jingyang, Li, Hai "Helen", Chen, Yiran |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SD-NAE: Generating Natural Adversarial Examples with Stable Diffusion
by: Lin, Yueqian, et al.
Published: (2023)
by: Lin, Yueqian, et al.
Published: (2023)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
by: Wang, Shaoguang, et al.
Published: (2026)
by: Wang, Shaoguang, et al.
Published: (2026)
Adaptive Keyframe Sampling for Long Video Understanding
by: Tang, Xi, et al.
Published: (2025)
by: Tang, Xi, et al.
Published: (2025)
MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models
by: Wang, Zhongxi, et al.
Published: (2026)
by: Wang, Zhongxi, et al.
Published: (2026)
MSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video Understanding
by: Tan, Wenhui, et al.
Published: (2026)
by: Tan, Wenhui, et al.
Published: (2026)
Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval
by: Shlapentokh-Rothman, Michal, et al.
Published: (2026)
by: Shlapentokh-Rothman, Michal, et al.
Published: (2026)
Test-Time Temporal Sampling for Efficient MLLM Video Understanding
by: Wang, Kaibin, et al.
Published: (2025)
by: Wang, Kaibin, et al.
Published: (2025)
FOCUS: Efficient Keyframe Selection for Long Video Understanding
by: Zhu, Zirui, et al.
Published: (2025)
by: Zhu, Zirui, et al.
Published: (2025)
Text-Conditioned Resampler For Long Form Video Understanding
by: Korbar, Bruno, et al.
Published: (2023)
by: Korbar, Bruno, et al.
Published: (2023)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
by: Yin, Yufei, et al.
Published: (2026)
by: Yin, Yufei, et al.
Published: (2026)
KFFocus: Highlighting Keyframes for Enhanced Video Understanding
by: Nie, Ming, et al.
Published: (2025)
by: Nie, Ming, et al.
Published: (2025)
From Captions to Keyframes: KeyScore for Multimodal Frame Scoring and Video-Language Understanding
by: Lin, Shih-Yao, et al.
Published: (2025)
by: Lin, Shih-Yao, et al.
Published: (2025)
REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding
by: Li, Jiaze, et al.
Published: (2025)
by: Li, Jiaze, et al.
Published: (2025)
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
by: Chen, Boyu, et al.
Published: (2025)
by: Chen, Boyu, et al.
Published: (2025)
AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
by: Zhang, Xian, et al.
Published: (2025)
by: Zhang, Xian, et al.
Published: (2025)
Keyframe-oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-Form Video Processing
by: Liu, Yudong, et al.
Published: (2025)
by: Liu, Yudong, et al.
Published: (2025)
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
by: Fang, Bo, et al.
Published: (2025)
by: Fang, Bo, et al.
Published: (2025)
LongViTU: Instruction Tuning for Long-Form Video Understanding
by: Wu, Rujie, et al.
Published: (2025)
by: Wu, Rujie, et al.
Published: (2025)
Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing
by: Yang, Hao, et al.
Published: (2026)
by: Yang, Hao, et al.
Published: (2026)
Zero-Shot Long-Form Video Understanding through Screenplay
by: Wu, Yongliang, et al.
Published: (2024)
by: Wu, Yongliang, et al.
Published: (2024)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
by: Fang, Xinyu, et al.
Published: (2024)
by: Fang, Xinyu, et al.
Published: (2024)
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding
by: Lin, Kuanwei, et al.
Published: (2026)
by: Lin, Kuanwei, et al.
Published: (2026)
DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes
by: Song, Zhende, et al.
Published: (2024)
by: Song, Zhende, et al.
Published: (2024)
VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
by: He, Jianxiang, et al.
Published: (2025)
by: He, Jianxiang, et al.
Published: (2025)
FlashFPS: Efficient Farthest Point Sampling for Large-Scale Point Clouds via Pruning and Caching
by: Fu, Yuzhe, et al.
Published: (2026)
by: Fu, Yuzhe, et al.
Published: (2026)
Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
by: Li, Jialuo, et al.
Published: (2025)
by: Li, Jialuo, et al.
Published: (2025)
Agentic Keyframe Search for Video Question Answering
by: Fan, Sunqi, et al.
Published: (2025)
by: Fan, Sunqi, et al.
Published: (2025)
T*: Re-thinking Temporal Search for Long-Form Video Understanding
by: Ye, Jinhui, et al.
Published: (2025)
by: Ye, Jinhui, et al.
Published: (2025)
Unleashing Hour-Scale Video Training for Long Video-Language Understanding
by: Lin, Jingyang, et al.
Published: (2025)
by: Lin, Jingyang, et al.
Published: (2025)
Long-VMNet: Accelerating Long-Form Video Understanding via Fixed Memory
by: Gurukar, Saket, et al.
Published: (2025)
by: Gurukar, Saket, et al.
Published: (2025)
DrawVideo: Generating Long Video from Storyboard Keyframe Sketches
by: Xu, Chuanzhi, et al.
Published: (2026)
by: Xu, Chuanzhi, et al.
Published: (2026)
Large Model based Sequential Keyframe Extraction for Video Summarization
by: Tan, Kailong, et al.
Published: (2024)
by: Tan, Kailong, et al.
Published: (2024)
Removing Averaging: Personalized Lip-Sync Driven Characters Based on Identity Adapter
by: Zhu, Yanyu, et al.
Published: (2025)
by: Zhu, Yanyu, et al.
Published: (2025)
AdaFlow: Efficient Long Video Editing via Adaptive Attention Slimming And Keyframe Selection
by: Zhang, Shuheng, et al.
Published: (2025)
by: Zhang, Shuheng, et al.
Published: (2025)
SmartDirector: Keyframe-Conditioned Cinematic Video Generation with Narrative Pacing Control
by: Zhang, Zhida, et al.
Published: (2026)
by: Zhang, Zhida, et al.
Published: (2026)
AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding
by: Qi, Haozhe, et al.
Published: (2026)
by: Qi, Haozhe, et al.
Published: (2026)
Occlusion-Aware Physics-Semantic Keyframe Selection for Robust Video Editing
by: Liu, Lin, et al.
Published: (2026)
by: Liu, Lin, et al.
Published: (2026)
RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
by: Xun, Shuhang, et al.
Published: (2025)
by: Xun, Shuhang, et al.
Published: (2025)
CAD-MLLM: Unifying Multimodality-Conditioned CAD Generation With MLLM
by: Xu, Jingwei, et al.
Published: (2024)
by: Xu, Jingwei, et al.
Published: (2024)
Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play
by: Wang, Qinsi, et al.
Published: (2025)
by: Wang, Qinsi, et al.
Published: (2025)
Similar Items
-
SD-NAE: Generating Natural Adversarial Examples with Stable Diffusion
by: Lin, Yueqian, et al.
Published: (2023) -
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
by: Wang, Shaoguang, et al.
Published: (2026) -
Adaptive Keyframe Sampling for Long Video Understanding
by: Tang, Xi, et al.
Published: (2025) -
MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models
by: Wang, Zhongxi, et al.
Published: (2026) -
MSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video Understanding
by: Tan, Wenhui, et al.
Published: (2026)