ViMU: Benchmarking Video Metaphorical Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Qi, Wang, Xinchao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MetaphorStar: Image Metaphor Understanding and Reasoning with End-to-End Visual Reinforcement Learning
by: Zhang, Chenhao, et al.
Published: (2026)
by: Zhang, Chenhao, et al.
Published: (2026)
MetaphorVU: Towards Metaphorical Video Understanding
by: Li, Zhuoqun, et al.
Published: (2026)
by: Li, Zhuoqun, et al.
Published: (2026)
ViFeEdit: A Video-Free Tuner of Your Video Diffusion Transformer
by: Yu, Ruonan, et al.
Published: (2026)
by: Yu, Ruonan, et al.
Published: (2026)
Vid-SME: Membership Inference Attacks against Large Video Understanding Models
by: Li, Qi, et al.
Published: (2025)
by: Li, Qi, et al.
Published: (2025)
LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding
by: Wang, Xiaodong, et al.
Published: (2026)
by: Wang, Xiaodong, et al.
Published: (2026)
FairViT: Fair Vision Transformer via Adaptive Masking
by: Tian, Bowei, et al.
Published: (2024)
by: Tian, Bowei, et al.
Published: (2024)
BLEnD-Vis: Benchmarking Multimodal Cultural Understanding in Vision Language Models
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2025)
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2025)
Sponge Tool Attack: Stealthy Denial-of-Efficiency against Tool-Augmented Agentic Reasoning
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
Compositional Video Generation as Flow Equalization
by: Yang, Xingyi, et al.
Published: (2024)
by: Yang, Xingyi, et al.
Published: (2024)
VideoNorms: Benchmarking Cultural Awareness of Video Language Models
by: Varimalla, Nikhil Reddy, et al.
Published: (2025)
by: Varimalla, Nikhil Reddy, et al.
Published: (2025)
FakingRecipe: Detecting Fake News on Short Video Platforms from the Perspective of Creative Process
by: Bu, Yuyan, et al.
Published: (2024)
by: Bu, Yuyan, et al.
Published: (2024)
ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation
by: Jha, Akshita, et al.
Published: (2024)
by: Jha, Akshita, et al.
Published: (2024)
Understanding the Representation of Older Adults in Motion Capture Locomotion Datasets
by: Yu, Yunkai, et al.
Published: (2025)
by: Yu, Yunkai, et al.
Published: (2025)
TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
by: Chen, Harold Haodong, et al.
Published: (2025)
by: Chen, Harold Haodong, et al.
Published: (2025)
On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes
by: Modi, Rajat, et al.
Published: (2024)
by: Modi, Rajat, et al.
Published: (2024)
SeViCES: Unifying Semantic-Visual Evidence Consensus for Long Video Understanding
by: Sheng, Yuan, et al.
Published: (2025)
by: Sheng, Yuan, et al.
Published: (2025)
Towards Understanding Unsafe Video Generation
by: Pang, Yan, et al.
Published: (2024)
by: Pang, Yan, et al.
Published: (2024)
AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic Images
by: Zhang, Bo, et al.
Published: (2026)
by: Zhang, Bo, et al.
Published: (2026)
Encapsulating Knowledge in One Prompt
by: Li, Qi, et al.
Published: (2024)
by: Li, Qi, et al.
Published: (2024)
REBUS: A Robust Evaluation Benchmark of Understanding Symbols
by: Gritsevskiy, Andrew, et al.
Published: (2024)
by: Gritsevskiy, Andrew, et al.
Published: (2024)
Can Multimodal LLMs See Science Instruction? Benchmarking Pedagogical Reasoning in K-12 Classroom Videos
by: Shen, Yixuan, et al.
Published: (2026)
by: Shen, Yixuan, et al.
Published: (2026)
EmoAssist: Emotional Assistant for Visual Impairment Community
by: Qi, Xingyu, et al.
Published: (2025)
by: Qi, Xingyu, et al.
Published: (2025)
Multiple Object Detection and Tracking in Panoramic Videos for Cycling Safety Analysis
by: Guo, Jingwei, et al.
Published: (2024)
by: Guo, Jingwei, et al.
Published: (2024)
ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting
by: Lee, Yeonkyung, et al.
Published: (2026)
by: Lee, Yeonkyung, et al.
Published: (2026)
ViLLa: Video Reasoning Segmentation with Large Language Model
by: Zheng, Rongkun, et al.
Published: (2024)
by: Zheng, Rongkun, et al.
Published: (2024)
EduPersona: Benchmarking Subjective Ability Boundaries of Virtual Student Agents
by: Zhu, Buyuan, et al.
Published: (2025)
by: Zhu, Buyuan, et al.
Published: (2025)
LongViTU: Instruction Tuning for Long-Form Video Understanding
by: Wu, Rujie, et al.
Published: (2025)
by: Wu, Rujie, et al.
Published: (2025)
Detecting Cultural Differences in News Video Thumbnails via Computational Aesthetics
by: Limpijankit, Marvin, et al.
Published: (2025)
by: Limpijankit, Marvin, et al.
Published: (2025)
Video-based Pedestrian and Vehicle Traffic Analysis During Football Games
by: Fleischer, Jacques P., et al.
Published: (2024)
by: Fleischer, Jacques P., et al.
Published: (2024)
Multi-modal News Understanding with Professionally Labelled Videos (ReutersViLNews)
by: Chou, Shih-Han, et al.
Published: (2024)
by: Chou, Shih-Han, et al.
Published: (2024)
A Machine Learning Model for Crowd Density Classification in Hajj Video Frames
by: Shah, Afnan A.
Published: (2025)
by: Shah, Afnan A.
Published: (2025)
ChaosBench: A Multi-Channel, Physics-Based Benchmark for Subseasonal-to-Seasonal Climate Prediction
by: Nathaniel, Juan, et al.
Published: (2024)
by: Nathaniel, Juan, et al.
Published: (2024)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
by: Li, Kunchang, et al.
Published: (2023)
by: Li, Kunchang, et al.
Published: (2023)
ViViD: Video Virtual Try-on using Diffusion Models
by: Fang, Zixun, et al.
Published: (2024)
by: Fang, Zixun, et al.
Published: (2024)
MM-Soc: Benchmarking Multimodal Large Language Models in Social Media Platforms
by: Jin, Yiqiao, et al.
Published: (2024)
by: Jin, Yiqiao, et al.
Published: (2024)
Anatomy of a Lie: A Multi-Stage Diagnostic Framework for Tracing Hallucinations in Vision-Language Models
by: Xiong, Lexiang, et al.
Published: (2026)
by: Xiong, Lexiang, et al.
Published: (2026)
Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
by: Jin, Peng, et al.
Published: (2023)
by: Jin, Peng, et al.
Published: (2023)
ViDiC: Video Difference Captioning
by: Wu, Jiangtao, et al.
Published: (2025)
by: Wu, Jiangtao, et al.
Published: (2025)
Breaking the Global North Stereotype: A Global South-centric Benchmark Dataset for Auditing and Mitigating Biases in Facial Recognition Systems
by: Jaiswal, Siddharth D, et al.
Published: (2024)
by: Jaiswal, Siddharth D, et al.
Published: (2024)
Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
by: Wang, Youze, et al.
Published: (2025)
by: Wang, Youze, et al.
Published: (2025)
Similar Items
-
MetaphorStar: Image Metaphor Understanding and Reasoning with End-to-End Visual Reinforcement Learning
by: Zhang, Chenhao, et al.
Published: (2026) -
MetaphorVU: Towards Metaphorical Video Understanding
by: Li, Zhuoqun, et al.
Published: (2026) -
ViFeEdit: A Video-Free Tuner of Your Video Diffusion Transformer
by: Yu, Ruonan, et al.
Published: (2026) -
Vid-SME: Membership Inference Attacks against Large Video Understanding Models
by: Li, Qi, et al.
Published: (2025) -
LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding
by: Wang, Xiaodong, et al.
Published: (2026)