LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917483539070976 |
|---|---|
| author | Chen, Jingfeng Qian, Jiawen Deng, Wendi Guo, Yinuo Yu, Jiaqi Leng, Sicong Thirukovalluru, Raghuveer Dhingra, Bhuwan |
| author_facet | Chen, Jingfeng Qian, Jiawen Deng, Wendi Guo, Yinuo Yu, Jiaqi Leng, Sicong Thirukovalluru, Raghuveer Dhingra, Bhuwan |
| contents | Video understanding in multimodal large language models requires selecting informative frames from long, redundant videos under limited visual-token budgets. Existing methods often rely on uniform sampling, point-wise relevance scoring, chunk-wise selection, or agentic exploration, which either miss global dependencies or introduce substantial overhead. We propose LDDR (Linear DPP-Based Dynamic Resolution), a training-free, plug-and-play, and budget-aware video frame sampling framework. LDDR performs query-aware Determinantal Point Process (DPP) frame selection in a task-conditioned feature space, achieving a 3x runtime speedup over standard DPP baselines. It further introduces a Group DPP importance metric to guide frame retention and dynamic resolution allocation, assigning more tokens to informative, non-redundant frames while downscaling or pruning less useful ones. Across four video benchmarks spanning short-, medium-, and long-range videos, LDDR consistently outperforms the next-best baselines, achieving gains of 2.5 points under budget-constrained settings and 1.6 points in high-budget scenarios. These improvements are consistently observed across multiple MLLM backbones, including both open- and closed-source models. Qualitative analysis confirms that relevant frames are selected and allocated a higher budget, facilitating improved video understanding. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_11477 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs Chen, Jingfeng Qian, Jiawen Deng, Wendi Guo, Yinuo Yu, Jiaqi Leng, Sicong Thirukovalluru, Raghuveer Dhingra, Bhuwan Computer Vision and Pattern Recognition Video understanding in multimodal large language models requires selecting informative frames from long, redundant videos under limited visual-token budgets. Existing methods often rely on uniform sampling, point-wise relevance scoring, chunk-wise selection, or agentic exploration, which either miss global dependencies or introduce substantial overhead. We propose LDDR (Linear DPP-Based Dynamic Resolution), a training-free, plug-and-play, and budget-aware video frame sampling framework. LDDR performs query-aware Determinantal Point Process (DPP) frame selection in a task-conditioned feature space, achieving a 3x runtime speedup over standard DPP baselines. It further introduces a Group DPP importance metric to guide frame retention and dynamic resolution allocation, assigning more tokens to informative, non-redundant frames while downscaling or pruning less useful ones. Across four video benchmarks spanning short-, medium-, and long-range videos, LDDR consistently outperforms the next-best baselines, achieving gains of 2.5 points under budget-constrained settings and 1.6 points in high-budget scenarios. These improvements are consistently observed across multiple MLLM backbones, including both open- and closed-source models. Qualitative analysis confirms that relevant frames are selected and allocated a higher budget, facilitating improved video understanding. |
| title | LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2605.11477 |