FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yunzhu, Lu, Yu, Wang, Tianyi, Rao, Fengyun, Yang, Yi, Zhu, Linchao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-reward Alignment
by: Suo, Yucheng, et al.
Published: (2025)
by: Suo, Yucheng, et al.
Published: (2025)
GPD: Guided Progressive Distillation for Fast and High-Quality Video Generation
by: Liang, Xiao, et al.
Published: (2026)
by: Liang, Xiao, et al.
Published: (2026)
HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronization
by: Zhou, Zitang, et al.
Published: (2025)
by: Zhou, Zitang, et al.
Published: (2025)
FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal Attention
by: Lu, Yu, et al.
Published: (2024)
by: Lu, Yu, et al.
Published: (2024)
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
by: Zhao, Ruixiang, et al.
Published: (2026)
by: Zhao, Ruixiang, et al.
Published: (2026)
One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding
by: Zhang, Zheyu, et al.
Published: (2026)
by: Zhang, Zheyu, et al.
Published: (2026)
Combating Label Noise With A General Surrogate Model For Sample Selection
by: Liang, Chao, et al.
Published: (2023)
by: Liang, Chao, et al.
Published: (2023)
VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization
by: Atanov, Andrei, et al.
Published: (2026)
by: Atanov, Andrei, et al.
Published: (2026)
FOCUS: Efficient Keyframe Selection for Long Video Understanding
by: Zhu, Zirui, et al.
Published: (2025)
by: Zhu, Zirui, et al.
Published: (2025)
AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding
by: Qi, Haozhe, et al.
Published: (2026)
by: Qi, Haozhe, et al.
Published: (2026)
STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
by: Jiang, Jindong, et al.
Published: (2025)
by: Jiang, Jindong, et al.
Published: (2025)
DGL: Dynamic Global-Local Prompt Tuning for Text-Video Retrieval
by: Yang, Xiangpeng, et al.
Published: (2024)
by: Yang, Xiangpeng, et al.
Published: (2024)
VideoGrain: Modulating Space-Time Attention for Multi-grained Video Editing
by: Yang, Xiangpeng, et al.
Published: (2025)
by: Yang, Xiangpeng, et al.
Published: (2025)
Video Token Merging for Long-form Video Understanding
by: Lee, Seon-Ho, et al.
Published: (2024)
by: Lee, Seon-Ho, et al.
Published: (2024)
EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing
by: Yang, Xiangpeng, et al.
Published: (2024)
by: Yang, Xiangpeng, et al.
Published: (2024)
FlexTraj: Image-to-Video Generation with Flexible Point Trajectory Control
by: Zhang, Zhiyuan, et al.
Published: (2025)
by: Zhang, Zhiyuan, et al.
Published: (2025)
Context-Aware Token Selection and Packing for Enhanced Vision Transformer
by: Zhang, Tianyi, et al.
Published: (2024)
by: Zhang, Tianyi, et al.
Published: (2024)
KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs
by: Song, Baiyang, et al.
Published: (2026)
by: Song, Baiyang, et al.
Published: (2026)
D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning
by: Tang, Changli, et al.
Published: (2026)
by: Tang, Changli, et al.
Published: (2026)
Adaptive Greedy Frame Selection for Long Video Understanding
by: Huang, Yuning, et al.
Published: (2026)
by: Huang, Yuning, et al.
Published: (2026)
MVP: Multiple View Prediction Improves GUI Grounding
by: Zhang, Yunzhu, et al.
Published: (2025)
by: Zhang, Yunzhu, et al.
Published: (2025)
Event-Anchored Frame Selection for Effective Long-Video Understanding
by: Chen, Wang, et al.
Published: (2026)
by: Chen, Wang, et al.
Published: (2026)
TokenMotion: Motion-Guided Vision Transformer for Video Camouflaged Object Detection Via Learnable Token Selection
by: Yu, Zifan, et al.
Published: (2023)
by: Yu, Zifan, et al.
Published: (2023)
TR-PTS: Task-Relevant Parameter and Token Selection for Efficient Tuning
by: Luo, Siqi, et al.
Published: (2025)
by: Luo, Siqi, et al.
Published: (2025)
Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding
by: Chen, Wang, et al.
Published: (2026)
by: Chen, Wang, et al.
Published: (2026)
METok: Multi-Stage Event-based Token Compression for Efficient Long Video Understanding
by: Wang, Mengyue, et al.
Published: (2025)
by: Wang, Mengyue, et al.
Published: (2025)
AudioScenic: Audio-Driven Video Scene Editing
by: Shen, Kaixin, et al.
Published: (2024)
by: Shen, Kaixin, et al.
Published: (2024)
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
by: Song, Enxin, et al.
Published: (2023)
by: Song, Enxin, et al.
Published: (2023)
FlexPara: Flexible Neural Surface Parameterization
by: Zhao, Yuming, et al.
Published: (2025)
by: Zhao, Yuming, et al.
Published: (2025)
MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs
by: Xu, Yunqiu, et al.
Published: (2024)
by: Xu, Yunqiu, et al.
Published: (2024)
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
by: Huang, De-An, et al.
Published: (2025)
by: Huang, De-An, et al.
Published: (2025)
FlexEControl: Flexible and Efficient Multimodal Control for Text-to-Image Generation
by: He, Xuehai, et al.
Published: (2024)
by: He, Xuehai, et al.
Published: (2024)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
by: Liu, Xiangrui, et al.
Published: (2025)
by: Liu, Xiangrui, et al.
Published: (2025)
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
by: Chen, Yuxiao, et al.
Published: (2026)
by: Chen, Yuxiao, et al.
Published: (2026)
$\text{PKS}^4$:Parallel Kinematic Selective State Space Scanners for Efficient Video Understanding
by: Zeng, Lingjie, et al.
Published: (2026)
by: Zeng, Lingjie, et al.
Published: (2026)
FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion
by: Lu, Yu, et al.
Published: (2025)
by: Lu, Yu, et al.
Published: (2025)
SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
by: Xu, Mingze, et al.
Published: (2025)
by: Xu, Mingze, et al.
Published: (2025)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
by: Hu, Pengfei, et al.
Published: (2025)
by: Hu, Pengfei, et al.
Published: (2025)
Principles of Visual Tokens for Efficient Video Understanding
by: Hao, Xinyue, et al.
Published: (2024)
by: Hao, Xinyue, et al.
Published: (2024)
Collaborative Group: Composed Image Retrieval via Consensus Learning from Noisy Annotations
by: Zhang, Xu, et al.
Published: (2023)
by: Zhang, Xu, et al.
Published: (2023)
Similar Items
-
From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-reward Alignment
by: Suo, Yucheng, et al.
Published: (2025) -
GPD: Guided Progressive Distillation for Fast and High-Quality Video Generation
by: Liang, Xiao, et al.
Published: (2026) -
HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronization
by: Zhou, Zitang, et al.
Published: (2025) -
FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal Attention
by: Lu, Yu, et al.
Published: (2024) -
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
by: Zhao, Ruixiang, et al.
Published: (2026)