One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Zheyu, Pang, Ziqi, Chen, Shixing, Hao, Xiang, Bhat, Vimal, Wang, Yu-Xiong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MR. Video: "MapReduce" is the Principle for Long Video Understanding
von: Pang, Ziqi, et al.
Veröffentlicht: (2025)
von: Pang, Ziqi, et al.
Veröffentlicht: (2025)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
RefTok: Reference-Based Tokenization for Video Generation
von: Fan, Xiang, et al.
Veröffentlicht: (2025)
von: Fan, Xiang, et al.
Veröffentlicht: (2025)
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
von: Zhang, Yunzhu, et al.
Veröffentlicht: (2025)
von: Zhang, Yunzhu, et al.
Veröffentlicht: (2025)
Event-Anchored Frame Selection for Effective Long-Video Understanding
von: Chen, Wang, et al.
Veröffentlicht: (2026)
von: Chen, Wang, et al.
Veröffentlicht: (2026)
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
von: Huang, De-An, et al.
Veröffentlicht: (2025)
von: Huang, De-An, et al.
Veröffentlicht: (2025)
Generative Frame Sampler for Long Video Understanding
von: Yao, Linli, et al.
Veröffentlicht: (2025)
von: Yao, Linli, et al.
Veröffentlicht: (2025)
MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
DiffSign: AI-Assisted Generation of Customizable Sign Language Videos With Enhanced Realism
von: Krishnamurthy, Sudha, et al.
Veröffentlicht: (2024)
von: Krishnamurthy, Sudha, et al.
Veröffentlicht: (2024)
Dynamic Token Compression for Efficient Video Understanding through Reinforcement Learning
von: Wang, Shida, et al.
Veröffentlicht: (2026)
von: Wang, Shida, et al.
Veröffentlicht: (2026)
Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding
von: Chen, Wang, et al.
Veröffentlicht: (2026)
von: Chen, Wang, et al.
Veröffentlicht: (2026)
METok: Multi-Stage Event-based Token Compression for Efficient Long Video Understanding
von: Wang, Mengyue, et al.
Veröffentlicht: (2025)
von: Wang, Mengyue, et al.
Veröffentlicht: (2025)
Adaptive Greedy Frame Selection for Long Video Understanding
von: Huang, Yuning, et al.
Veröffentlicht: (2026)
von: Huang, Yuning, et al.
Veröffentlicht: (2026)
Video Token Merging for Long-form Video Understanding
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
GLUS: Global-Local Reasoning Unified into A Single Large Language Model for Video Segmentation
von: Lin, Lang, et al.
Veröffentlicht: (2025)
von: Lin, Lang, et al.
Veröffentlicht: (2025)
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2024)
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2024)
AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding
von: Qi, Haozhe, et al.
Veröffentlicht: (2026)
von: Qi, Haozhe, et al.
Veröffentlicht: (2026)
RMem: Restricted Memory Banks Improve Video Object Segmentation
von: Zhou, Junbao, et al.
Veröffentlicht: (2024)
von: Zhou, Junbao, et al.
Veröffentlicht: (2024)
From Frames to Clips: Training-free Adaptive Key Clip Selection for Long-Form Video Understanding
von: Sun, Guangyu, et al.
Veröffentlicht: (2025)
von: Sun, Guangyu, et al.
Veröffentlicht: (2025)
STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
von: Jiang, Jindong, et al.
Veröffentlicht: (2025)
von: Jiang, Jindong, et al.
Veröffentlicht: (2025)
EarlyTom: Early Token Compression Completes Fast Video Understanding
von: Wang, Hesong, et al.
Veröffentlicht: (2026)
von: Wang, Hesong, et al.
Veröffentlicht: (2026)
A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
von: Kerssies, Tommie, et al.
Veröffentlicht: (2026)
von: Kerssies, Tommie, et al.
Veröffentlicht: (2026)
X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding
von: Zhou, Wenqi, et al.
Veröffentlicht: (2025)
von: Zhou, Wenqi, et al.
Veröffentlicht: (2025)
LVBench: An Extreme Long Video Understanding Benchmark
von: Wang, Weihan, et al.
Veröffentlicht: (2024)
von: Wang, Weihan, et al.
Veröffentlicht: (2024)
FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
von: Cho, Janghoon, et al.
Veröffentlicht: (2025)
von: Cho, Janghoon, et al.
Veröffentlicht: (2025)
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning
von: Wang, Yicheng, et al.
Veröffentlicht: (2024)
von: Wang, Yicheng, et al.
Veröffentlicht: (2024)
Text-Guided Video Masked Autoencoder
von: Fan, David, et al.
Veröffentlicht: (2024)
von: Fan, David, et al.
Veröffentlicht: (2024)
Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity
von: Zhang, Huaxin, et al.
Veröffentlicht: (2024)
von: Zhang, Huaxin, et al.
Veröffentlicht: (2024)
ALLVB: All-in-One Long Video Understanding Benchmark
von: Tan, Xichen, et al.
Veröffentlicht: (2025)
von: Tan, Xichen, et al.
Veröffentlicht: (2025)
M-LLM Based Video Frame Selection for Efficient Video Understanding
von: Hu, Kai, et al.
Veröffentlicht: (2025)
von: Hu, Kai, et al.
Veröffentlicht: (2025)
StreamingTOM: Streaming Token Compression for Efficient Video Understanding
von: Chen, Xueyi, et al.
Veröffentlicht: (2025)
von: Chen, Xueyi, et al.
Veröffentlicht: (2025)
Aligning Generative Denoising with Discriminative Objectives Unleashes Diffusion for Visual Perception
von: Pang, Ziqi, et al.
Veröffentlicht: (2025)
von: Pang, Ziqi, et al.
Veröffentlicht: (2025)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
von: Wang, Shaoguang, et al.
Veröffentlicht: (2026)
von: Wang, Shaoguang, et al.
Veröffentlicht: (2026)
VeRVE: Versatile Retrieval for Videos via Unified Embeddings
von: Halbe, Shaunak, et al.
Veröffentlicht: (2026)
von: Halbe, Shaunak, et al.
Veröffentlicht: (2026)
AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark
von: Gauba, Aruna, et al.
Veröffentlicht: (2025)
von: Gauba, Aruna, et al.
Veröffentlicht: (2025)
Frame by Familiar Frame: Understanding Replication in Video Diffusion Models
von: Rahman, Aimon, et al.
Veröffentlicht: (2024)
von: Rahman, Aimon, et al.
Veröffentlicht: (2024)
Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
von: Li, Jialuo, et al.
Veröffentlicht: (2025)
von: Li, Jialuo, et al.
Veröffentlicht: (2025)
Towards Understanding Unsafe Video Generation
von: Pang, Yan, et al.
Veröffentlicht: (2024)
von: Pang, Yan, et al.
Veröffentlicht: (2024)
What Happens Next? Next Scene Prediction with a Unified Video Model
von: Li, Xinjie, et al.
Veröffentlicht: (2025)
von: Li, Xinjie, et al.
Veröffentlicht: (2025)
Principles of Visual Tokens for Efficient Video Understanding
von: Hao, Xinyue, et al.
Veröffentlicht: (2024)
von: Hao, Xinyue, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MR. Video: "MapReduce" is the Principle for Long Video Understanding
von: Pang, Ziqi, et al.
Veröffentlicht: (2025) -
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025) -
RefTok: Reference-Based Tokenization for Video Generation
von: Fan, Xiang, et al.
Veröffentlicht: (2025) -
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
von: Zhang, Yunzhu, et al.
Veröffentlicht: (2025) -
Event-Anchored Frame Selection for Effective Long-Video Understanding
von: Chen, Wang, et al.
Veröffentlicht: (2026)