Visual Context Window Extension: A New Perspective for Long Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wei, Hongchen, Chen, Zhenzhong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Training-Free Reasoning and Reflection in MLLMs
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs
von: Gao, Yiling, et al.
Veröffentlicht: (2026)
von: Gao, Yiling, et al.
Veröffentlicht: (2026)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
LOP: Learning Optimal Pruning for Efficient On-Demand MLLMs Scaling
von: Zhang, Zhihan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhihan, et al.
Veröffentlicht: (2025)
RSFAKE-1M: A Large-Scale Dataset for Detecting Diffusion-Generated Remote Sensing Forgeries
von: Tan, Zhihong, et al.
Veröffentlicht: (2025)
von: Tan, Zhihong, et al.
Veröffentlicht: (2025)
SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context
von: Li, Jungang, et al.
Veröffentlicht: (2024)
von: Li, Jungang, et al.
Veröffentlicht: (2024)
Exploring Long- and Short-Range Temporal Information for Learned Video Compression
von: Wang, Huairui, et al.
Veröffentlicht: (2022)
von: Wang, Huairui, et al.
Veröffentlicht: (2022)
Memory Consolidation Enables Long-Context Video Understanding
von: Balažević, Ivana, et al.
Veröffentlicht: (2024)
von: Balažević, Ivana, et al.
Veröffentlicht: (2024)
Intention-driven Ego-to-Exo Video Generation
von: Luo, Hongchen, et al.
Veröffentlicht: (2024)
von: Luo, Hongchen, et al.
Veröffentlicht: (2024)
VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding
von: He, Haichen, et al.
Veröffentlicht: (2026)
von: He, Haichen, et al.
Veröffentlicht: (2026)
Context Forcing: Consistent Autoregressive Video Generation with Long Context
von: Chen, Shuo, et al.
Veröffentlicht: (2026)
von: Chen, Shuo, et al.
Veröffentlicht: (2026)
LongVILA: Scaling Long-Context Visual Language Models for Long Videos
von: Chen, Yukang, et al.
Veröffentlicht: (2024)
von: Chen, Yukang, et al.
Veröffentlicht: (2024)
DiffVC-RT: Towards Practical Real-Time Diffusion-based Perceptual Neural Video Compression
von: Ma, Wenzhuo, et al.
Veröffentlicht: (2026)
von: Ma, Wenzhuo, et al.
Veröffentlicht: (2026)
Leverage Task Context for Object Affordance Ranking
von: Huang, Haojie, et al.
Veröffentlicht: (2024)
von: Huang, Haojie, et al.
Veröffentlicht: (2024)
LaMD: Latent Motion Diffusion for Image-Conditional Video Generation
von: Hu, Yaosi, et al.
Veröffentlicht: (2023)
von: Hu, Yaosi, et al.
Veröffentlicht: (2023)
VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation
von: Ren, Weiming, et al.
Veröffentlicht: (2024)
von: Ren, Weiming, et al.
Veröffentlicht: (2024)
Visual-Geometric Collaborative Guidance for Affordance Learning
von: Luo, Hongchen, et al.
Veröffentlicht: (2024)
von: Luo, Hongchen, et al.
Veröffentlicht: (2024)
Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers
von: Ren, Weiming, et al.
Veröffentlicht: (2025)
von: Ren, Weiming, et al.
Veröffentlicht: (2025)
VMAD: Visual-enhanced Multimodal Large Language Model for Zero-Shot Anomaly Detection
von: Deng, Huilin, et al.
Veröffentlicht: (2024)
von: Deng, Huilin, et al.
Veröffentlicht: (2024)
Question-guided Visual Compression with Memory Feedback for Long-Term Video Understanding
von: Yamao, Sosuke, et al.
Veröffentlicht: (2026)
von: Yamao, Sosuke, et al.
Veröffentlicht: (2026)
Long Video Understanding with Learnable Retrieval in Video-Language Models
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
Internalized Reasoning for Long-Context Visual Document Understanding
von: Veselka, Austin
Veröffentlicht: (2026)
von: Veselka, Austin
Veröffentlicht: (2026)
CoS: Chain-of-Shot Prompting for Long Video Understanding
von: Hu, Jian, et al.
Veröffentlicht: (2025)
von: Hu, Jian, et al.
Veröffentlicht: (2025)
VERA: Identifying and Leveraging Visual Evidence Retrieval Heads in Long-Context Understanding
von: Pei, Rongcan, et al.
Veröffentlicht: (2026)
von: Pei, Rongcan, et al.
Veröffentlicht: (2026)
Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding
von: Xie, Yuan, et al.
Veröffentlicht: (2025)
von: Xie, Yuan, et al.
Veröffentlicht: (2025)
Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understanding
von: Wang, Zheng, et al.
Veröffentlicht: (2026)
von: Wang, Zheng, et al.
Veröffentlicht: (2026)
SeViCES: Unifying Semantic-Visual Evidence Consensus for Long Video Understanding
von: Sheng, Yuan, et al.
Veröffentlicht: (2025)
von: Sheng, Yuan, et al.
Veröffentlicht: (2025)
APVR: Hour-Level Long Video Understanding with Adaptive Pivot Visual Information Retrieval
von: Gao, Hong, et al.
Veröffentlicht: (2025)
von: Gao, Hong, et al.
Veröffentlicht: (2025)
Long Context Tuning for Video Generation
von: Guo, Yuwei, et al.
Veröffentlicht: (2025)
von: Guo, Yuwei, et al.
Veröffentlicht: (2025)
VideoPro: Adaptive Program Reasoning for Long Video Understanding
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
von: Li, Chenglin, et al.
Veröffentlicht: (2025)
Train Short, Inference Long: Training-free Horizon Extension for Autoregressive Video Generation
von: Li, Jia, et al.
Veröffentlicht: (2026)
von: Li, Jia, et al.
Veröffentlicht: (2026)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
von: Ma, Wentao, et al.
Veröffentlicht: (2025)
von: Ma, Wentao, et al.
Veröffentlicht: (2025)
Long-Context State-Space Video World Models
von: Po, Ryan, et al.
Veröffentlicht: (2025)
von: Po, Ryan, et al.
Veröffentlicht: (2025)
Video Token Merging for Long-form Video Understanding
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
von: Lee, Seon-Ho, et al.
Veröffentlicht: (2024)
Linear Scaling Video VLMs for Long Video Understanding
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2026)
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2026)
Language Repository for Long Video Understanding
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
EEA: Exploration-Exploitation Agent for Long Video Understanding
von: Yang, Te, et al.
Veröffentlicht: (2025)
von: Yang, Te, et al.
Veröffentlicht: (2025)
From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-reward Alignment
von: Suo, Yucheng, et al.
Veröffentlicht: (2025)
von: Suo, Yucheng, et al.
Veröffentlicht: (2025)
OmChat: A Recipe to Train Multimodal Language Models with Strong Long Context and Video Understanding
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2024)
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Training-Free Reasoning and Reflection in MLLMs
von: Wei, Hongchen, et al.
Veröffentlicht: (2025) -
ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs
von: Gao, Yiling, et al.
Veröffentlicht: (2026) -
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
von: Wei, Hongchen, et al.
Veröffentlicht: (2025) -
LOP: Learning Optimal Pruning for Efficient On-Demand MLLMs Scaling
von: Zhang, Zhihan, et al.
Veröffentlicht: (2025) -
RSFAKE-1M: A Large-Scale Dataset for Detecting Diffusion-Generated Remote Sensing Forgeries
von: Tan, Zhihong, et al.
Veröffentlicht: (2025)