VideoGLUE: Video General Understanding Evaluation of Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Liangzhe, Gundavarapu, Nitesh Bharadwaj, Zhao, Long, Zhou, Hao, Cui, Yin, Jiang, Lu, Yang, Xuan, Jia, Menglin, Weyand, Tobias, Friedman, Luke, Sirotenko, Mikhail, Wang, Huisheng, Schroff, Florian, Adam, Hartwig, Yang, Ming-Hsuan, Liu, Ting, Gong, Boqing |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VideoPrism: A Foundational Visual Encoder for Video Understanding
by: Zhao, Long, et al.
Published: (2024)
by: Zhao, Long, et al.
Published: (2024)
Extending Video Masked Autoencoders to 128 frames
by: Gundavarapu, Nitesh Bharadwaj, et al.
Published: (2024)
by: Gundavarapu, Nitesh Bharadwaj, et al.
Published: (2024)
Neptune: The Long Orbit to Benchmarking Long Video Understanding
by: Nagrani, Arsha, et al.
Published: (2024)
by: Nagrani, Arsha, et al.
Published: (2024)
Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding
by: Xiong, Yuanhao, et al.
Published: (2023)
by: Xiong, Yuanhao, et al.
Published: (2023)
Distilling Vision-Language Models on Millions of Videos
by: Zhao, Yue, et al.
Published: (2024)
by: Zhao, Yue, et al.
Published: (2024)
MINERVA: Evaluating Complex Video Reasoning
by: Nagrani, Arsha, et al.
Published: (2025)
by: Nagrani, Arsha, et al.
Published: (2025)
Video Creation by Demonstration
by: Sun, Yihong, et al.
Published: (2024)
by: Sun, Yihong, et al.
Published: (2024)
Lifting Data-Tracing Machine Unlearning to Knowledge-Tracing for Foundation Models
by: Tan, Yuwen, et al.
Published: (2025)
by: Tan, Yuwen, et al.
Published: (2025)
LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs
by: Kim, Jihwan, et al.
Published: (2026)
by: Kim, Jihwan, et al.
Published: (2026)
Moiré Video Authentication: A Physical Signature Against AI Video Generation
by: Qing, Yuan, et al.
Published: (2026)
by: Qing, Yuan, et al.
Published: (2026)
VideoAds for Fast-Paced Video Understanding
by: Zhang, Zheyuan, et al.
Published: (2025)
by: Zhang, Zheyuan, et al.
Published: (2025)
SANPO: A Scene Understanding, Accessibility and Human Navigation Dataset
by: Waghmare, Sagar M., et al.
Published: (2023)
by: Waghmare, Sagar M., et al.
Published: (2023)
VideoPoet: A Large Language Model for Zero-Shot Video Generation
by: Kondratyuk, Dan, et al.
Published: (2023)
by: Kondratyuk, Dan, et al.
Published: (2023)
Image Diffusion Preview with Consistency Solver
by: Wang, Fu-Yun, et al.
Published: (2025)
by: Wang, Fu-Yun, et al.
Published: (2025)
CAViAR: Critic-Augmented Video Agentic Reasoning
by: Menon, Sachit, et al.
Published: (2025)
by: Menon, Sachit, et al.
Published: (2025)
Evolution of Video Generative Foundations
by: Hu, Teng, et al.
Published: (2026)
by: Hu, Teng, et al.
Published: (2026)
Sobre la realidad de la vida cotidiana de los jóvenes en poblaciones en el nuevo orden democrático: «ni tan protagonista ni tan víctima»
by: Michaela Weyand
Published: (1993)
by: Michaela Weyand
Published: (1993)
Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
by: Yu, Lijun, et al.
Published: (2023)
by: Yu, Lijun, et al.
Published: (2023)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding
by: Nagrani, Arsha, et al.
Published: (2026)
by: Nagrani, Arsha, et al.
Published: (2026)
MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning
by: Singh, Darshan, et al.
Published: (2026)
by: Singh, Darshan, et al.
Published: (2026)
Think over Trajectories: Leveraging Video Generation to Reconstruct GPS Trajectories from Cellular Signaling
by: Zhang, Ruixing, et al.
Published: (2026)
by: Zhang, Ruixing, et al.
Published: (2026)
GLUE: Gradient-free Learning to Unify Experts
by: Park, Jong-Ik, et al.
Published: (2025)
by: Park, Jong-Ik, et al.
Published: (2025)
On Discrete Prompt Optimization for Diffusion Models
by: Wang, Ruochen, et al.
Published: (2024)
by: Wang, Ruochen, et al.
Published: (2024)
HyperCore: The Core Framework for Building Hyperbolic Foundation Models with Comprehensive Modules
by: He, Neil, et al.
Published: (2025)
by: He, Neil, et al.
Published: (2025)
Cryptocurrency forensics: Forensic analysis of the Electrum wallet to uncover artifacts
by: Saloni Jain, et al.
Published: (2026)
by: Saloni Jain, et al.
Published: (2026)
Omni-Video: Democratizing Unified Video Understanding and Generation
by: Tan, Zhiyu, et al.
Published: (2025)
by: Tan, Zhiyu, et al.
Published: (2025)
GLUE: Global-Local Unified Encoding for Imitation Learning via Key-Patch Tracking
by: Chen, Ye, et al.
Published: (2025)
by: Chen, Ye, et al.
Published: (2025)
ltzGLUE: Luxembourgish General Language Understanding Evaluation
by: Plum, Alistair, et al.
Published: (2026)
by: Plum, Alistair, et al.
Published: (2026)
Unified Dense Prediction of Video Diffusion
by: Yang, Lehan, et al.
Published: (2025)
by: Yang, Lehan, et al.
Published: (2025)
VL-GLUE: A Suite of Fundamental yet Challenging Visuo-Linguistic Reasoning Tasks
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
by: Sampat, Shailaja Keyur, et al.
Published: (2024)
Improving Subject-Driven Image Synthesis with Subject-Agnostic Guidance
by: Chan, Kelvin C. K., et al.
Published: (2024)
by: Chan, Kelvin C. K., et al.
Published: (2024)
Attention to Neural Plagiarism: Diffusion Models Can Plagiarize Your Copyrighted Images!
by: Zou, Zihang, et al.
Published: (2026)
by: Zou, Zihang, et al.
Published: (2026)
Culture in Action: Evaluating Text-to-Image Models through Social Activities
by: Malakouti, Sina, et al.
Published: (2025)
by: Malakouti, Sina, et al.
Published: (2025)
The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition
by: Tan, Yuwen, et al.
Published: (2025)
by: Tan, Yuwen, et al.
Published: (2025)
InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
by: Wang, Chenting, et al.
Published: (2025)
by: Wang, Chenting, et al.
Published: (2025)
Scaling Video Pretraining for Surgical Foundation Models
by: Lu, Sicheng, et al.
Published: (2026)
by: Lu, Sicheng, et al.
Published: (2026)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
by: Wasim, Syed Talal, et al.
Published: (2023)
by: Wasim, Syed Talal, et al.
Published: (2023)
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
by: Lin, Yuanze, et al.
Published: (2025)
by: Lin, Yuanze, et al.
Published: (2025)
Video-CoM: Interactive Video Reasoning via Chain of Manipulations
by: Rasheed, Hanoona, et al.
Published: (2025)
by: Rasheed, Hanoona, et al.
Published: (2025)
Similar Items
-
VideoPrism: A Foundational Visual Encoder for Video Understanding
by: Zhao, Long, et al.
Published: (2024) -
Extending Video Masked Autoencoders to 128 frames
by: Gundavarapu, Nitesh Bharadwaj, et al.
Published: (2024) -
Neptune: The Long Orbit to Benchmarking Long Video Understanding
by: Nagrani, Arsha, et al.
Published: (2024) -
Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding
by: Xiong, Yuanhao, et al.
Published: (2023) -
Distilling Vision-Language Models on Millions of Videos
by: Zhao, Yue, et al.
Published: (2024)