LongVLM: Efficient Long Video Understanding via Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Weng, Yuetian, Han, Mingfei, He, Haoyu, Chang, Xiaojun, Zhuang, Bohan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos
von: Han, Mingfei, et al.
Veröffentlicht: (2023)
von: Han, Mingfei, et al.
Veröffentlicht: (2023)
SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
von: Xu, Mingze, et al.
Veröffentlicht: (2025)
von: Xu, Mingze, et al.
Veröffentlicht: (2025)
BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
Streaming Long Video Understanding with Large Language Models
von: Qian, Rui, et al.
Veröffentlicht: (2024)
von: Qian, Rui, et al.
Veröffentlicht: (2024)
VirtueBench: Evaluating Trustworthiness under Uncertainty in Long Video Understanding
von: Yu, Xueqing, et al.
Veröffentlicht: (2026)
von: Yu, Xueqing, et al.
Veröffentlicht: (2026)
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2024)
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2024)
Understanding Long Videos with Multimodal Language Models
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism
von: Chen, Tao, et al.
Veröffentlicht: (2026)
von: Chen, Tao, et al.
Veröffentlicht: (2026)
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
von: Chen, Yuxiao, et al.
Veröffentlicht: (2026)
von: Chen, Yuxiao, et al.
Veröffentlicht: (2026)
Long Video Understanding with Learnable Retrieval in Video-Language Models
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos
von: Fateh, Fawad Javed, et al.
Veröffentlicht: (2024)
von: Fateh, Fawad Javed, et al.
Veröffentlicht: (2024)
Language Repository for Long Video Understanding
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
Motion Mamba: Efficient and Long Sequence Motion Generation
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
von: Chen, Zhuokun, et al.
Veröffentlicht: (2026)
von: Chen, Zhuokun, et al.
Veröffentlicht: (2026)
Self-ReS: Self-Reflection in Large Vision-Language Models for Long Video Understanding
von: Pereira, Joao, et al.
Veröffentlicht: (2025)
von: Pereira, Joao, et al.
Veröffentlicht: (2025)
Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos
von: Cao, Meng, et al.
Veröffentlicht: (2026)
von: Cao, Meng, et al.
Veröffentlicht: (2026)
Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval
von: Chen, Tao, et al.
Veröffentlicht: (2025)
von: Chen, Tao, et al.
Veröffentlicht: (2025)
STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
von: Jiang, Jindong, et al.
Veröffentlicht: (2025)
von: Jiang, Jindong, et al.
Veröffentlicht: (2025)
BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding
von: Liu, Shuming, et al.
Veröffentlicht: (2025)
von: Liu, Shuming, et al.
Veröffentlicht: (2025)
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
von: He, Bo, et al.
Veröffentlicht: (2024)
von: He, Bo, et al.
Veröffentlicht: (2024)
TriAttention: Efficient Long Reasoning with Trigonometric KV Compression
von: Mao, Weian, et al.
Veröffentlicht: (2026)
von: Mao, Weian, et al.
Veröffentlicht: (2026)
CogVLM2: Visual Language Models for Image and Video Understanding
von: Hong, Wenyi, et al.
Veröffentlicht: (2024)
von: Hong, Wenyi, et al.
Veröffentlicht: (2024)
WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
von: Liu, Zhiheng, et al.
Veröffentlicht: (2025)
von: Liu, Zhiheng, et al.
Veröffentlicht: (2025)
LongVILA: Scaling Long-Context Visual Language Models for Long Videos
von: Chen, Yukang, et al.
Veröffentlicht: (2024)
von: Chen, Yukang, et al.
Veröffentlicht: (2024)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
Long-VMNet: Accelerating Long-Form Video Understanding via Fixed Memory
von: Gurukar, Saket, et al.
Veröffentlicht: (2025)
von: Gurukar, Saket, et al.
Veröffentlicht: (2025)
Goldfish: Vision-Language Understanding of Arbitrarily Long Videos
von: Ataallah, Kirolos, et al.
Veröffentlicht: (2024)
von: Ataallah, Kirolos, et al.
Veröffentlicht: (2024)
Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
von: Shu, Yan, et al.
Veröffentlicht: (2024)
von: Shu, Yan, et al.
Veröffentlicht: (2024)
OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
von: Chen, Feng, et al.
Veröffentlicht: (2025)
von: Chen, Feng, et al.
Veröffentlicht: (2025)
AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding
von: Li, Handong, et al.
Veröffentlicht: (2026)
von: Li, Handong, et al.
Veröffentlicht: (2026)
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
von: Gao, Zhe, et al.
Veröffentlicht: (2026)
von: Gao, Zhe, et al.
Veröffentlicht: (2026)
Mitigating Data Redundancy to Revitalize Transformer-based Long-Term Time Series Forecasting System
von: Li, Mingjie, et al.
Veröffentlicht: (2022)
von: Li, Mingjie, et al.
Veröffentlicht: (2022)
Beyond Dense Futures: World Models as Structured Planners for Robotic Manipulation
von: Jin, Minghao, et al.
Veröffentlicht: (2026)
von: Jin, Minghao, et al.
Veröffentlicht: (2026)
PersonaVLM: Long-Term Personalized Multimodal LLMs
von: Nie, Chang, et al.
Veröffentlicht: (2026)
von: Nie, Chang, et al.
Veröffentlicht: (2026)
Towards Event-oriented Long Video Understanding
von: Du, Yifan, et al.
Veröffentlicht: (2024)
von: Du, Yifan, et al.
Veröffentlicht: (2024)
TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
von: Tang, Canhui, et al.
Veröffentlicht: (2025)
von: Tang, Canhui, et al.
Veröffentlicht: (2025)
Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions
von: Zhang, Kecheng, et al.
Veröffentlicht: (2026)
von: Zhang, Kecheng, et al.
Veröffentlicht: (2026)
ReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video Understanding
von: Yashima, Daichi, et al.
Veröffentlicht: (2026)
von: Yashima, Daichi, et al.
Veröffentlicht: (2026)
LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
von: Ning, Zhenyu, et al.
Veröffentlicht: (2025)
von: Ning, Zhenyu, et al.
Veröffentlicht: (2025)
VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding
von: Yang, Ruoliu, et al.
Veröffentlicht: (2026)
von: Yang, Ruoliu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos
von: Han, Mingfei, et al.
Veröffentlicht: (2023) -
SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
von: Xu, Mingze, et al.
Veröffentlicht: (2025) -
BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025) -
Streaming Long Video Understanding with Large Language Models
von: Qian, Rui, et al.
Veröffentlicht: (2024) -
VirtueBench: Evaluating Trustworthiness under Uncertainty in Long Video Understanding
von: Yu, Xueqing, et al.
Veröffentlicht: (2026)