Adaptive Keyframe Sampling for Long Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tang, Xi, Qiu, Jihao, Xie, Lingxi, Tian, Yunjie, Jiao, Jianbin, Ye, Qixiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fast-iTPN: Integrally Pre-Trained Transformer Pyramid Network with Token Migration
von: Tian, Yunjie, et al.
Veröffentlicht: (2022)
von: Tian, Yunjie, et al.
Veröffentlicht: (2022)
Artemis: Towards Referential Understanding in Complex Videos
von: Qiu, Jihao, et al.
Veröffentlicht: (2024)
von: Qiu, Jihao, et al.
Veröffentlicht: (2024)
LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding
von: Qiu, Jihao, et al.
Veröffentlicht: (2026)
von: Qiu, Jihao, et al.
Veröffentlicht: (2026)
Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
von: Wang, Yiheng, et al.
Veröffentlicht: (2026)
von: Wang, Yiheng, et al.
Veröffentlicht: (2026)
FOCUS: Efficient Keyframe Selection for Long Video Understanding
von: Zhu, Zirui, et al.
Veröffentlicht: (2025)
von: Zhu, Zirui, et al.
Veröffentlicht: (2025)
ChatterBox: Multi-round Multimodal Referring and Grounding
von: Tian, Yunjie, et al.
Veröffentlicht: (2024)
von: Tian, Yunjie, et al.
Veröffentlicht: (2024)
YOLOv12: Attention-Centric Real-Time Object Detectors
von: Tian, Yunjie, et al.
Veröffentlicht: (2025)
von: Tian, Yunjie, et al.
Veröffentlicht: (2025)
Neptune: The Long Orbit to Benchmarking Long Video Understanding
von: Nagrani, Arsha, et al.
Veröffentlicht: (2024)
von: Nagrani, Arsha, et al.
Veröffentlicht: (2024)
Motion meets Attention: Video Motion Prompts
von: Chen, Qixiang, et al.
Veröffentlicht: (2024)
von: Chen, Qixiang, et al.
Veröffentlicht: (2024)
Depth-guided Texture Diffusion for Image Semantic Segmentation
von: Sun, Wei, et al.
Veröffentlicht: (2024)
von: Sun, Wei, et al.
Veröffentlicht: (2024)
Dual-Signal Adaptive KV-Cache Optimization for Long-Form Video Understanding in Vision-Language Models
von: Sai, Vishnu, et al.
Veröffentlicht: (2026)
von: Sai, Vishnu, et al.
Veröffentlicht: (2026)
PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation
von: Li, Xiaolong, et al.
Veröffentlicht: (2025)
von: Li, Xiaolong, et al.
Veröffentlicht: (2025)
VideoNSA: Native Sparse Attention Scales Video Understanding
von: Song, Enxin, et al.
Veröffentlicht: (2025)
von: Song, Enxin, et al.
Veröffentlicht: (2025)
Tuning-Free Multi-Event Long Video Generation via Synchronized Coupled Sampling
von: Kim, Subin, et al.
Veröffentlicht: (2025)
von: Kim, Subin, et al.
Veröffentlicht: (2025)
Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
von: Li, Jialuo, et al.
Veröffentlicht: (2025)
von: Li, Jialuo, et al.
Veröffentlicht: (2025)
Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior
von: Li, Yulin, et al.
Veröffentlicht: (2025)
von: Li, Yulin, et al.
Veröffentlicht: (2025)
ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension
von: Ma, Tianren, et al.
Veröffentlicht: (2024)
von: Ma, Tianren, et al.
Veröffentlicht: (2024)
Uncertainty-guided Optimal Transport in Depth Supervised Sparse-View 3D Gaussian
von: Sun, Wei, et al.
Veröffentlicht: (2024)
von: Sun, Wei, et al.
Veröffentlicht: (2024)
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing
von: Lee, Hosu, et al.
Veröffentlicht: (2024)
von: Lee, Hosu, et al.
Veröffentlicht: (2024)
Do Language Models Understand Time?
von: Ding, Xi, et al.
Veröffentlicht: (2024)
von: Ding, Xi, et al.
Veröffentlicht: (2024)
HourVideo: 1-Hour Video-Language Understanding
von: Chandrasegaran, Keshigeyan, et al.
Veröffentlicht: (2024)
von: Chandrasegaran, Keshigeyan, et al.
Veröffentlicht: (2024)
LongVideoAgent: Multi-Agent Reasoning with Long Videos
von: Liu, Runtao, et al.
Veröffentlicht: (2025)
von: Liu, Runtao, et al.
Veröffentlicht: (2025)
VMamba: Visual State Space Model
von: Liu, Yue, et al.
Veröffentlicht: (2024)
von: Liu, Yue, et al.
Veröffentlicht: (2024)
Compressing Model with Few Class-Imbalance Samples: An Out-of-Distribution Expedition
von: Wu, Tian-Shuang, et al.
Veröffentlicht: (2025)
von: Wu, Tian-Shuang, et al.
Veröffentlicht: (2025)
FedGSCA: Medical Federated Learning with Global Sample Selector and Client Adaptive Adjuster under Label Noise
von: Ye, Mengwen, et al.
Veröffentlicht: (2025)
von: Ye, Mengwen, et al.
Veröffentlicht: (2025)
VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
von: He, Jianxiang, et al.
Veröffentlicht: (2025)
von: He, Jianxiang, et al.
Veröffentlicht: (2025)
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement
von: Rao, Zhefan, et al.
Veröffentlicht: (2024)
von: Rao, Zhefan, et al.
Veröffentlicht: (2024)
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
von: Yuan, Yuqian, et al.
Veröffentlicht: (2024)
von: Yuan, Yuqian, et al.
Veröffentlicht: (2024)
Correspondence-Guided SfM-Free 3D Gaussian Splatting for NVS
von: Sun, Wei, et al.
Veröffentlicht: (2024)
von: Sun, Wei, et al.
Veröffentlicht: (2024)
Small Vision-Language Models are Smart Compressors for Long Video Understanding
von: Fei, Junjie, et al.
Veröffentlicht: (2026)
von: Fei, Junjie, et al.
Veröffentlicht: (2026)
Radial Attention: $O(n\log n)$ Sparse Attention with Energy Decay for Long Video Generation
von: Li, Xingyang, et al.
Veröffentlicht: (2025)
von: Li, Xingyang, et al.
Veröffentlicht: (2025)
Temporal Preference Optimization for Long-Form Video Understanding
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
Enhancing Long Video Generation Consistency without Tuning
von: Li, Xingyao, et al.
Veröffentlicht: (2024)
von: Li, Xingyao, et al.
Veröffentlicht: (2024)
Video Understanding by Design: How Datasets Shape Architectures and Insights
von: Wang, Lei, et al.
Veröffentlicht: (2025)
von: Wang, Lei, et al.
Veröffentlicht: (2025)
VIDEOP2R: Video Understanding from Perception to Reasoning
von: Jiang, Yifan, et al.
Veröffentlicht: (2025)
von: Jiang, Yifan, et al.
Veröffentlicht: (2025)
Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion Models
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2023)
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2023)
Controllable Video Generation with Provable Disentanglement
von: Shen, Yifan, et al.
Veröffentlicht: (2025)
von: Shen, Yifan, et al.
Veröffentlicht: (2025)
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
von: Woo, Byeongju, et al.
Veröffentlicht: (2026)
von: Woo, Byeongju, et al.
Veröffentlicht: (2026)
Understanding-Enhanced Model Collaboration for Long-Tailed Egocentric Mistake Detection
von: Han, Boyu, et al.
Veröffentlicht: (2026)
von: Han, Boyu, et al.
Veröffentlicht: (2026)
Cert-SSBD: Certified Backdoor Defense with Sample-Specific Smoothing Noises
von: Qiao, Ting, et al.
Veröffentlicht: (2025)
von: Qiao, Ting, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Fast-iTPN: Integrally Pre-Trained Transformer Pyramid Network with Token Migration
von: Tian, Yunjie, et al.
Veröffentlicht: (2022) -
Artemis: Towards Referential Understanding in Complex Videos
von: Qiu, Jihao, et al.
Veröffentlicht: (2024) -
LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding
von: Qiu, Jihao, et al.
Veröffentlicht: (2026) -
Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
von: Wang, Yiheng, et al.
Veröffentlicht: (2026) -
FOCUS: Efficient Keyframe Selection for Long Video Understanding
von: Zhu, Zirui, et al.
Veröffentlicht: (2025)