Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qin, Minghao, Liu, Xiangrui, Liang, Zhengyang, Shu, Yan, Yuan, Huaying, Zhou, Juenjie, Xiao, Shitao, Zhao, Bo, Liu, Zheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Task-Aware KV Compression For Cost-Effective Long Video Understanding
von: Qin, Minghao, et al.
Veröffentlicht: (2025)
von: Qin, Minghao, et al.
Veröffentlicht: (2025)
Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
von: Shu, Yan, et al.
Veröffentlicht: (2024)
von: Shu, Yan, et al.
Veröffentlicht: (2024)
TimeScope: Towards Task-Oriented Temporal Grounding In Long Videos
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
Video-Browser: Towards Agentic Open-web Video Browsing
von: Liang, Zhengyang, et al.
Veröffentlicht: (2025)
von: Liang, Zhengyang, et al.
Veröffentlicht: (2025)
MLVU: Benchmarking Multi-task Long Video Understanding
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
Agentic Very Long Video Understanding
von: Rege, Aniket, et al.
Veröffentlicht: (2026)
von: Rege, Aniket, et al.
Veröffentlicht: (2026)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
Large Language Models as Foundations for Next-Gen Dense Retrieval: A Comprehensive Empirical Assessment
von: Luo, Kun, et al.
Veröffentlicht: (2024)
von: Luo, Kun, et al.
Veröffentlicht: (2024)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding
von: Xie, Yuan, et al.
Veröffentlicht: (2025)
von: Xie, Yuan, et al.
Veröffentlicht: (2025)
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
USV: Unified Sparsification for Accelerating Video Diffusion Models
von: Wu, Xinjian, et al.
Veröffentlicht: (2025)
von: Wu, Xinjian, et al.
Veröffentlicht: (2025)
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2025)
von: Yin, Yufei, et al.
Veröffentlicht: (2025)
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
von: Ji, Yicheng, et al.
Veröffentlicht: (2026)
von: Ji, Yicheng, et al.
Veröffentlicht: (2026)
Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models
von: Sun, Haoqin, et al.
Veröffentlicht: (2026)
von: Sun, Haoqin, et al.
Veröffentlicht: (2026)
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
von: Fu, Honghao, et al.
Veröffentlicht: (2026)
von: Fu, Honghao, et al.
Veröffentlicht: (2026)
BGE Landmark Embedding: A Chunking-Free Embedding Method For Retrieval Augmented Long-Context Large Language Models
von: Luo, Kun, et al.
Veröffentlicht: (2024)
von: Luo, Kun, et al.
Veröffentlicht: (2024)
VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks
von: Jang, Lawrence, et al.
Veröffentlicht: (2024)
von: Jang, Lawrence, et al.
Veröffentlicht: (2024)
AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding
von: Li, Handong, et al.
Veröffentlicht: (2026)
von: Li, Handong, et al.
Veröffentlicht: (2026)
Flexibly Scaling Large Language Models Contexts Through Extensible Tokenization
von: Shao, Ninglu, et al.
Veröffentlicht: (2024)
von: Shao, Ninglu, et al.
Veröffentlicht: (2024)
Relax Forcing: Relaxed KV-Memory for Consistent Long Video Generation
von: Zhao, Zengqun, et al.
Veröffentlicht: (2026)
von: Zhao, Zengqun, et al.
Veröffentlicht: (2026)
VideoLucy: Deep Memory Backtracking for Long Video Understanding
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation
von: Chen, Jiayu, et al.
Veröffentlicht: (2026)
von: Chen, Jiayu, et al.
Veröffentlicht: (2026)
EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining
von: Xu, Boshen, et al.
Veröffentlicht: (2025)
von: Xu, Boshen, et al.
Veröffentlicht: (2025)
Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion
von: Cai, Peiliang, et al.
Veröffentlicht: (2026)
von: Cai, Peiliang, et al.
Veröffentlicht: (2026)
Towards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining
von: Peng, Bo, et al.
Veröffentlicht: (2026)
von: Peng, Bo, et al.
Veröffentlicht: (2026)
Towards Event-oriented Long Video Understanding
von: Du, Yifan, et al.
Veröffentlicht: (2024)
von: Du, Yifan, et al.
Veröffentlicht: (2024)
MedHorizon: Towards Long-context Medical Video Understanding in the Wild
von: Du, Bodong, et al.
Veröffentlicht: (2026)
von: Du, Bodong, et al.
Veröffentlicht: (2026)
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2025)
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2025)
MTV-Inpaint: Multi-Task Long Video Inpainting
von: Yang, Shiyuan, et al.
Veröffentlicht: (2025)
von: Yang, Shiyuan, et al.
Veröffentlicht: (2025)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
Plenoptic Video Generation
von: Fu, Xiao, et al.
Veröffentlicht: (2026)
von: Fu, Xiao, et al.
Veröffentlicht: (2026)
Video Panels for Long Video Understanding
von: Doorenbos, Lars, et al.
Veröffentlicht: (2025)
von: Doorenbos, Lars, et al.
Veröffentlicht: (2025)
Towards Long Video Understanding via Fine-detailed Video Story Generation
von: You, Zeng, et al.
Veröffentlicht: (2024)
von: You, Zeng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Task-Aware KV Compression For Cost-Effective Long Video Understanding
von: Qin, Minghao, et al.
Veröffentlicht: (2025) -
Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
von: Shu, Yan, et al.
Veröffentlicht: (2024) -
TimeScope: Towards Task-Oriented Temporal Grounding In Long Videos
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025) -
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025) -
Video-Browser: Towards Agentic Open-web Video Browsing
von: Liang, Zhengyang, et al.
Veröffentlicht: (2025)