VideoPrism: A Foundational Visual Encoder for Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Long, Gundavarapu, Nitesh B., Yuan, Liangzhe, Zhou, Hao, Yan, Shen, Sun, Jennifer J., Friedman, Luke, Qian, Rui, Weyand, Tobias, Zhao, Yue, Hornung, Rachel, Schroff, Florian, Yang, Ming-Hsuan, Ross, David A., Wang, Huisheng, Adam, Hartwig, Sirotenko, Mikhail, Liu, Ting, Gong, Boqing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VideoGLUE: Video General Understanding Evaluation of Foundation Models
by: Yuan, Liangzhe, et al.
Published: (2023)
by: Yuan, Liangzhe, et al.
Published: (2023)
Extending Video Masked Autoencoders to 128 frames
by: Gundavarapu, Nitesh Bharadwaj, et al.
Published: (2024)
by: Gundavarapu, Nitesh Bharadwaj, et al.
Published: (2024)
Neptune: The Long Orbit to Benchmarking Long Video Understanding
by: Nagrani, Arsha, et al.
Published: (2024)
by: Nagrani, Arsha, et al.
Published: (2024)
Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding
by: Xiong, Yuanhao, et al.
Published: (2023)
by: Xiong, Yuanhao, et al.
Published: (2023)
Distilling Vision-Language Models on Millions of Videos
by: Zhao, Yue, et al.
Published: (2024)
by: Zhao, Yue, et al.
Published: (2024)
MINERVA: Evaluating Complex Video Reasoning
by: Nagrani, Arsha, et al.
Published: (2025)
by: Nagrani, Arsha, et al.
Published: (2025)
LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs
by: Kim, Jihwan, et al.
Published: (2026)
by: Kim, Jihwan, et al.
Published: (2026)
VideoPoet: A Large Language Model for Zero-Shot Video Generation
by: Kondratyuk, Dan, et al.
Published: (2023)
by: Kondratyuk, Dan, et al.
Published: (2023)
Video Creation by Demonstration
by: Sun, Yihong, et al.
Published: (2024)
by: Sun, Yihong, et al.
Published: (2024)
Lifting Data-Tracing Machine Unlearning to Knowledge-Tracing for Foundation Models
by: Tan, Yuwen, et al.
Published: (2025)
by: Tan, Yuwen, et al.
Published: (2025)
Moiré Video Authentication: A Physical Signature Against AI Video Generation
by: Qing, Yuan, et al.
Published: (2026)
by: Qing, Yuan, et al.
Published: (2026)
VideoAds for Fast-Paced Video Understanding
by: Zhang, Zheyuan, et al.
Published: (2025)
by: Zhang, Zheyuan, et al.
Published: (2025)
CAViAR: Critic-Augmented Video Agentic Reasoning
by: Menon, Sachit, et al.
Published: (2025)
by: Menon, Sachit, et al.
Published: (2025)
Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding
by: Nagrani, Arsha, et al.
Published: (2026)
by: Nagrani, Arsha, et al.
Published: (2026)
Image Diffusion Preview with Consistency Solver
by: Wang, Fu-Yun, et al.
Published: (2025)
by: Wang, Fu-Yun, et al.
Published: (2025)
Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
by: Yu, Lijun, et al.
Published: (2023)
by: Yu, Lijun, et al.
Published: (2023)
VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer
by: Lin, Rui, et al.
Published: (2026)
by: Lin, Rui, et al.
Published: (2026)
SANPO: A Scene Understanding, Accessibility and Human Navigation Dataset
by: Waghmare, Sagar M., et al.
Published: (2023)
by: Waghmare, Sagar M., et al.
Published: (2023)
Sobre la realidad de la vida cotidiana de los jóvenes en poblaciones en el nuevo orden democrático: «ni tan protagonista ni tan víctima»
by: Michaela Weyand
Published: (1993)
by: Michaela Weyand
Published: (1993)
Evolution of Video Generative Foundations
by: Hu, Teng, et al.
Published: (2026)
by: Hu, Teng, et al.
Published: (2026)
MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning
by: Singh, Darshan, et al.
Published: (2026)
by: Singh, Darshan, et al.
Published: (2026)
Think over Trajectories: Leveraging Video Generation to Reconstruct GPS Trajectories from Cellular Signaling
by: Zhang, Ruixing, et al.
Published: (2026)
by: Zhang, Ruixing, et al.
Published: (2026)
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
by: Maaz, Muhammad, et al.
Published: (2024)
by: Maaz, Muhammad, et al.
Published: (2024)
PrismVAU: Prompt-Refined Inference System for Multimodal Video Anomaly Understanding
by: Erregue, Iñaki, et al.
Published: (2026)
by: Erregue, Iñaki, et al.
Published: (2026)
Optimal Investment with Herd Behaviour Using Rational Decision Decomposition
by: Wang, Huisheng, et al.
Published: (2024)
by: Wang, Huisheng, et al.
Published: (2024)
Optimal Investment under Mutual Strategy Influence among Agents
by: Wang, Huisheng, et al.
Published: (2025)
by: Wang, Huisheng, et al.
Published: (2025)
Optimal Investment under the Influence of Decision-changing Imitation
by: Wang, Huisheng, et al.
Published: (2024)
by: Wang, Huisheng, et al.
Published: (2024)
Analyzing the Crowding-Out Effect of Investment Herding on Consumption: An Optimal Control Theory Approach
by: Wang, Huisheng, et al.
Published: (2025)
by: Wang, Huisheng, et al.
Published: (2025)
Mechanism Design for Investment Regulation under Herding
by: Wang, Huisheng, et al.
Published: (2026)
by: Wang, Huisheng, et al.
Published: (2026)
PrismAudio: Decomposed Chain-of-Thoughts and Multi-dimensional Rewards for Video-to-Audio Generation
by: Liu, Huadai, et al.
Published: (2025)
by: Liu, Huadai, et al.
Published: (2025)
MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer
by: Karim, Rezaul, et al.
Published: (2023)
by: Karim, Rezaul, et al.
Published: (2023)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
Attention to Neural Plagiarism: Diffusion Models Can Plagiarize Your Copyrighted Images!
by: Zou, Zihang, et al.
Published: (2026)
by: Zou, Zihang, et al.
Published: (2026)
Culture in Action: Evaluating Text-to-Image Models through Social Activities
by: Malakouti, Sina, et al.
Published: (2025)
by: Malakouti, Sina, et al.
Published: (2025)
The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition
by: Tan, Yuwen, et al.
Published: (2025)
by: Tan, Yuwen, et al.
Published: (2025)
Epsilon-VAE: Denoising as Visual Decoding
by: Zhao, Long, et al.
Published: (2024)
by: Zhao, Long, et al.
Published: (2024)
Deranged Perfect Matchings on complete graph and balanced complete r-partite graph
by: Deng, Boqing
Published: (2025)
by: Deng, Boqing
Published: (2025)
Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models
by: Yi, Jinhui, et al.
Published: (2024)
by: Yi, Jinhui, et al.
Published: (2024)
SparseTem: Boosting the Efficiency of CNN-Based Video Encoders by Exploiting Temporal Continuity
by: Wang, Kunyun, et al.
Published: (2024)
by: Wang, Kunyun, et al.
Published: (2024)
Video Prediction Models as General Visual Encoders
by: Maier, James, et al.
Published: (2024)
by: Maier, James, et al.
Published: (2024)
Similar Items
-
VideoGLUE: Video General Understanding Evaluation of Foundation Models
by: Yuan, Liangzhe, et al.
Published: (2023) -
Extending Video Masked Autoencoders to 128 frames
by: Gundavarapu, Nitesh Bharadwaj, et al.
Published: (2024) -
Neptune: The Long Orbit to Benchmarking Long Video Understanding
by: Nagrani, Arsha, et al.
Published: (2024) -
Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding
by: Xiong, Yuanhao, et al.
Published: (2023) -
Distilling Vision-Language Models on Millions of Videos
by: Zhao, Yue, et al.
Published: (2024)