Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shu, Yan, Liu, Zheng, Zhang, Peitian, Qin, Minghao, Zhou, Junjie, Liang, Zhengyang, Huang, Tiejun, Zhao, Bo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
von: Qin, Minghao, et al.
Veröffentlicht: (2025)
von: Qin, Minghao, et al.
Veröffentlicht: (2025)
MLVU: Benchmarking Multi-task Long Video Understanding
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
Task-Aware KV Compression For Cost-Effective Long Video Understanding
von: Qin, Minghao, et al.
Veröffentlicht: (2025)
von: Qin, Minghao, et al.
Veröffentlicht: (2025)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
TimeScope: Towards Task-Oriented Temporal Grounding In Long Videos
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
Video-Browser: Towards Agentic Open-web Video Browsing
von: Liang, Zhengyang, et al.
Veröffentlicht: (2025)
von: Liang, Zhengyang, et al.
Veröffentlicht: (2025)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
Unleashing Hour-Scale Video Training for Long Video-Language Understanding
von: Lin, Jingyang, et al.
Veröffentlicht: (2025)
von: Lin, Jingyang, et al.
Veröffentlicht: (2025)
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
Goldfish: Vision-Language Understanding of Arbitrarily Long Videos
von: Ataallah, Kirolos, et al.
Veröffentlicht: (2024)
von: Ataallah, Kirolos, et al.
Veröffentlicht: (2024)
Long Video Understanding with Learnable Retrieval in Video-Language Models
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models
von: Ma, Martin Q., et al.
Veröffentlicht: (2026)
von: Ma, Martin Q., et al.
Veröffentlicht: (2026)
Linear Scaling Video VLMs for Long Video Understanding
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2026)
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2026)
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2025)
von: Yin, Yufei, et al.
Veröffentlicht: (2025)
Small Vision-Language Models are Smart Compressors for Long Video Understanding
von: Fei, Junjie, et al.
Veröffentlicht: (2026)
von: Fei, Junjie, et al.
Veröffentlicht: (2026)
BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding
von: Liu, Shuming, et al.
Veröffentlicht: (2025)
von: Liu, Shuming, et al.
Veröffentlicht: (2025)
UVLM: Benchmarking Video Language Model for Underwater World Understanding
von: Xue, Xizhe, et al.
Veröffentlicht: (2025)
von: Xue, Xizhe, et al.
Veröffentlicht: (2025)
Language Repository for Long Video Understanding
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
HourVideo: 1-Hour Video-Language Understanding
von: Chandrasegaran, Keshigeyan, et al.
Veröffentlicht: (2024)
von: Chandrasegaran, Keshigeyan, et al.
Veröffentlicht: (2024)
VideoDistill: Language-aware Vision Distillation for Video Question Answering
von: Zou, Bo, et al.
Veröffentlicht: (2024)
von: Zou, Bo, et al.
Veröffentlicht: (2024)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
von: Ma, David, et al.
Veröffentlicht: (2025)
von: Ma, David, et al.
Veröffentlicht: (2025)
Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding
von: Xie, Yuan, et al.
Veröffentlicht: (2025)
von: Xie, Yuan, et al.
Veröffentlicht: (2025)
Understanding Long Videos with Multimodal Language Models
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
von: Zheng, Naishan, et al.
Veröffentlicht: (2025)
von: Zheng, Naishan, et al.
Veröffentlicht: (2025)
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
von: Guo, Yanan, et al.
Veröffentlicht: (2025)
von: Guo, Yanan, et al.
Veröffentlicht: (2025)
SPKLIP: Aligning Spike Video Streams with Natural Language
von: Gao, Yongchang, et al.
Veröffentlicht: (2025)
von: Gao, Yongchang, et al.
Veröffentlicht: (2025)
Self-ReS: Self-Reflection in Large Vision-Language Models for Long Video Understanding
von: Pereira, Joao, et al.
Veröffentlicht: (2025)
von: Pereira, Joao, et al.
Veröffentlicht: (2025)
Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers
von: Ren, Weiming, et al.
Veröffentlicht: (2025)
von: Ren, Weiming, et al.
Veröffentlicht: (2025)
VideoLucy: Deep Memory Backtracking for Long Video Understanding
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
von: Wang, Yi, et al.
Veröffentlicht: (2024)
von: Wang, Yi, et al.
Veröffentlicht: (2024)
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2024)
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2024)
AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding
von: Li, Handong, et al.
Veröffentlicht: (2026)
von: Li, Handong, et al.
Veröffentlicht: (2026)
v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound
von: Shi, Zhengpeng, et al.
Veröffentlicht: (2025)
von: Shi, Zhengpeng, et al.
Veröffentlicht: (2025)
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
von: Hannan, Tanveer, et al.
Veröffentlicht: (2024)
von: Hannan, Tanveer, et al.
Veröffentlicht: (2024)
Streaming Long Video Understanding with Large Language Models
von: Qian, Rui, et al.
Veröffentlicht: (2024)
von: Qian, Rui, et al.
Veröffentlicht: (2024)
Parameter-free Video Segmentation for Vision and Language Understanding
von: Mahon, Louis, et al.
Veröffentlicht: (2025)
von: Mahon, Louis, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
von: Qin, Minghao, et al.
Veröffentlicht: (2025) -
MLVU: Benchmarking Multi-task Long Video Understanding
von: Zhou, Junjie, et al.
Veröffentlicht: (2024) -
Task-Aware KV Compression For Cost-Effective Long Video Understanding
von: Qin, Minghao, et al.
Veröffentlicht: (2025) -
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025) -
TimeScope: Towards Task-Oriented Temporal Grounding In Long Videos
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)