VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ren, Weiming, Yang, Huan, Min, Jie, Wei, Cong, Chen, Wenhu |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers
par: Ren, Weiming, et autres
Publié: (2025)
par: Ren, Weiming, et autres
Publié: (2025)
ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation
par: Ren, Weiming, et autres
Publié: (2024)
par: Ren, Weiming, et autres
Publié: (2024)
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
par: Ma, Wentao, et autres
Publié: (2025)
par: Ma, Wentao, et autres
Publié: (2025)
AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
par: Ku, Max, et autres
Publié: (2024)
par: Ku, Max, et autres
Publié: (2024)
UniVideo: Unified Understanding, Generation, and Editing for Videos
par: Wei, Cong, et autres
Publié: (2025)
par: Wei, Cong, et autres
Publié: (2025)
Context Forcing: Consistent Autoregressive Video Generation with Long Context
par: Chen, Shuo, et autres
Publié: (2026)
par: Chen, Shuo, et autres
Publié: (2026)
QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design
par: Schneider, Benjamin, et autres
Publié: (2025)
par: Schneider, Benjamin, et autres
Publié: (2025)
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
par: Shen, Xiaoqian, et autres
Publié: (2024)
par: Shen, Xiaoqian, et autres
Publié: (2024)
VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding
par: Zhao, Henghao, et autres
Publié: (2025)
par: Zhao, Henghao, et autres
Publié: (2025)
HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
par: Qiu, Haonan, et autres
Publié: (2025)
par: Qiu, Haonan, et autres
Publié: (2025)
Token Merging via Spatiotemporal Information Mining for Surgical Video Understanding
par: Jiang, Xixi, et autres
Publié: (2025)
par: Jiang, Xixi, et autres
Publié: (2025)
Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling
par: Hyung, Junha, et autres
Publié: (2024)
par: Hyung, Junha, et autres
Publié: (2024)
E-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation
par: Xu, Zeyu, et autres
Publié: (2025)
par: Xu, Zeyu, et autres
Publié: (2025)
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
par: Jin, Hongbo, et autres
Publié: (2025)
par: Jin, Hongbo, et autres
Publié: (2025)
Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understanding
par: Wang, Zheng, et autres
Publié: (2026)
par: Wang, Zheng, et autres
Publié: (2026)
OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision
par: Wei, Cong, et autres
Publié: (2024)
par: Wei, Cong, et autres
Publié: (2024)
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
par: Shen, Xiaoqian, et autres
Publié: (2025)
par: Shen, Xiaoqian, et autres
Publié: (2025)
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering
par: Cheng, Zheng, et autres
Publié: (2024)
par: Cheng, Zheng, et autres
Publié: (2024)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
par: Yuan, Huaying, et autres
Publié: (2025)
par: Yuan, Huaying, et autres
Publié: (2025)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
par: Yin, Yufei, et autres
Publié: (2026)
par: Yin, Yufei, et autres
Publié: (2026)
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
par: Xue, Zhucun, et autres
Publié: (2025)
par: Xue, Zhucun, et autres
Publié: (2025)
Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task
par: Fan, Sunqi, et autres
Publié: (2025)
par: Fan, Sunqi, et autres
Publié: (2025)
STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution
par: Xie, Rui, et autres
Publié: (2025)
par: Xie, Rui, et autres
Publié: (2025)
Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation
par: Wang, Wenjing, et autres
Publié: (2023)
par: Wang, Wenjing, et autres
Publié: (2023)
Enhancing Video Understanding: Deep Neural Networks for Spatiotemporal Analysis
par: Fadaei, Amir Hosein, et autres
Publié: (2025)
par: Fadaei, Amir Hosein, et autres
Publié: (2025)
Long Video Understanding with Learnable Retrieval in Video-Language Models
par: Xu, Jiaqi, et autres
Publié: (2023)
par: Xu, Jiaqi, et autres
Publié: (2023)
VCA: Video Curious Agent for Long Video Understanding
par: Yang, Zeyuan, et autres
Publié: (2024)
par: Yang, Zeyuan, et autres
Publié: (2024)
VideoLucy: Deep Memory Backtracking for Long Video Understanding
par: Zuo, Jialong, et autres
Publié: (2025)
par: Zuo, Jialong, et autres
Publié: (2025)
Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought
par: Zhang, Shuyi, et autres
Publié: (2025)
par: Zhang, Shuyi, et autres
Publié: (2025)
Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding
par: Xie, Yuan, et autres
Publié: (2025)
par: Xie, Yuan, et autres
Publié: (2025)
VideoPro: Adaptive Program Reasoning for Long Video Understanding
par: Li, Chenglin, et autres
Publié: (2025)
par: Li, Chenglin, et autres
Publié: (2025)
DiffST: Spatiotemporal-Aware Diffusion for Real-World Space-Time Video Super-Resolution
par: Chen, Zheng, et autres
Publié: (2026)
par: Chen, Zheng, et autres
Publié: (2026)
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
par: Guo, Yanan, et autres
Publié: (2025)
par: Guo, Yanan, et autres
Publié: (2025)
Video Token Merging for Long-form Video Understanding
par: Lee, Seon-Ho, et autres
Publié: (2024)
par: Lee, Seon-Ho, et autres
Publié: (2024)
EEA: Exploration-Exploitation Agent for Long Video Understanding
par: Yang, Te, et autres
Publié: (2025)
par: Yang, Te, et autres
Publié: (2025)
Linear Scaling Video VLMs for Long Video Understanding
par: Eyzaguirre, Cristobal, et autres
Publié: (2026)
par: Eyzaguirre, Cristobal, et autres
Publié: (2026)
Visual Context Window Extension: A New Perspective for Long Video Understanding
par: Wei, Hongchen, et autres
Publié: (2024)
par: Wei, Hongchen, et autres
Publié: (2024)
VISTA: Video Interaction Spatio-Temporal Analysis Benchmark
par: Aparcedo, Alejandro, et autres
Publié: (2026)
par: Aparcedo, Alejandro, et autres
Publié: (2026)
Video Diffusion Models: A Survey
par: Melnik, Andrew, et autres
Publié: (2024)
par: Melnik, Andrew, et autres
Publié: (2024)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
par: Liu, Xiangrui, et autres
Publié: (2025)
par: Liu, Xiangrui, et autres
Publié: (2025)
Documents similaires
-
Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers
par: Ren, Weiming, et autres
Publié: (2025) -
ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation
par: Ren, Weiming, et autres
Publié: (2024) -
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
par: Ma, Wentao, et autres
Publié: (2025) -
AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
par: Ku, Max, et autres
Publié: (2024) -
UniVideo: Unified Understanding, Generation, and Editing for Videos
par: Wei, Cong, et autres
Publié: (2025)