V2P-Bench: Evaluating Video-Language Understanding with Visual Prompts for Better Human-Model Interaction
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Yiming, Zeng, Yu, Qi, Yukun, Liu, YaoYang, Bao, Xikun, Chen, Lin, Chen, Zehui, Miao, Qing, Liu, Chenxi, Zhao, Jie, Zhao, Feng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning
by: Qi, Yukun, et al.
Published: (2025)
by: Qi, Yukun, et al.
Published: (2025)
Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models
by: Zeng, Yu, et al.
Published: (2025)
by: Zeng, Yu, et al.
Published: (2025)
VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation
by: Zhao, Yiming, et al.
Published: (2026)
by: Zhao, Yiming, et al.
Published: (2026)
SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation
by: Liu, YaoYang, et al.
Published: (2026)
by: Liu, YaoYang, et al.
Published: (2026)
Hybrid 3D Human Pose Estimation with Monocular Video and Sparse IMUs
by: Bao, Yiming, et al.
Published: (2024)
by: Bao, Yiming, et al.
Published: (2024)
ConflictBench: Evaluating Human-AI Conflict via Interactive and Visually Grounded Environments
by: Zhao, Weixiang, et al.
Published: (2026)
by: Zhao, Weixiang, et al.
Published: (2026)
ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
by: Chen, Lin, et al.
Published: (2024)
by: Chen, Lin, et al.
Published: (2024)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
by: Han, ZhaoYang, et al.
Published: (2025)
by: Han, ZhaoYang, et al.
Published: (2025)
Multi-Prompting Decoder Helps Better Language Understanding
by: Cheng, Zifeng, et al.
Published: (2024)
by: Cheng, Zifeng, et al.
Published: (2024)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
by: Zheng, Naishan, et al.
Published: (2025)
by: Zheng, Naishan, et al.
Published: (2025)
LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding
by: Wang, Xiaodong, et al.
Published: (2026)
by: Wang, Xiaodong, et al.
Published: (2026)
STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
by: Li, Yun, et al.
Published: (2025)
by: Li, Yun, et al.
Published: (2025)
VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning
by: Wang, Qiuchen, et al.
Published: (2025)
by: Wang, Qiuchen, et al.
Published: (2025)
P-Flow: Prompting Visual Effects Generation
by: Zhao, Rui, et al.
Published: (2026)
by: Zhao, Rui, et al.
Published: (2026)
RSB-Pose: Robust Short-Baseline Binocular 3D Human Pose Estimation with Occlusion Handling
by: Wan, Xiaoyue, et al.
Published: (2023)
by: Wan, Xiaoyue, et al.
Published: (2023)
Panda or not Panda? Understanding Adversarial Attacks with Interactive Visualization
by: You, Yuzhe, et al.
Published: (2023)
by: You, Yuzhe, et al.
Published: (2023)
MindSearch: Mimicking Human Minds Elicits Deep AI Searcher
by: Chen, Zehui, et al.
Published: (2024)
by: Chen, Zehui, et al.
Published: (2024)
StressPrompt: Does Stress Impact Large Language Models and Human Performance Similarly?
by: Shen, Guobin, et al.
Published: (2024)
by: Shen, Guobin, et al.
Published: (2024)
SAMITE: Position Prompted SAM2 with Calibrated Memory for Visual Object Tracking
by: Xu, Qianxiong, et al.
Published: (2025)
by: Xu, Qianxiong, et al.
Published: (2025)
T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models
by: Guo, Xuyang, et al.
Published: (2025)
by: Guo, Xuyang, et al.
Published: (2025)
VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects
by: Gao, Xiangbo, et al.
Published: (2026)
by: Gao, Xiangbo, et al.
Published: (2026)
TASTE-Rob: Advancing Video Generation of Task-Oriented Hand-Object Interaction for Generalizable Robotic Manipulation
by: Zhao, Hongxiang, et al.
Published: (2025)
by: Zhao, Hongxiang, et al.
Published: (2025)
NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts
by: Zhang, Shudan, et al.
Published: (2024)
by: Zhang, Shudan, et al.
Published: (2024)
ColorConceptBench: A Benchmark for Probabilistic Color-Concept Understanding in Text-to-Image Models
by: Ruan, Chenxi, et al.
Published: (2026)
by: Ruan, Chenxi, et al.
Published: (2026)
Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench
by: Hu, Lanxiang, et al.
Published: (2025)
by: Hu, Lanxiang, et al.
Published: (2025)
REI-Bench: Can Embodied Agents Understand Vague Human Instructions in Task Planning?
by: Jiang, Chenxi, et al.
Published: (2025)
by: Jiang, Chenxi, et al.
Published: (2025)
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos
by: Feng, Hengyi, et al.
Published: (2026)
by: Feng, Hengyi, et al.
Published: (2026)
ActTraitBench: Quantifying the Knowledge-Decision Gap in Large Language Models via Human-Grounded Behavioral Validation
by: Yang, Yutong, et al.
Published: (2026)
by: Yang, Yutong, et al.
Published: (2026)
Personalized Video Summarization by Multimodal Video Understanding
by: Chen, Brian, et al.
Published: (2024)
by: Chen, Brian, et al.
Published: (2024)
Mixup Helps Understanding Multimodal Video Better
by: Ma, Xiaoyu, et al.
Published: (2025)
by: Ma, Xiaoyu, et al.
Published: (2025)
In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
by: Peng, Taiying, et al.
Published: (2025)
by: Peng, Taiying, et al.
Published: (2025)
V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models
by: Luo, Yang, et al.
Published: (2025)
by: Luo, Yang, et al.
Published: (2025)
VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans?
by: Wang, Jiaqi, et al.
Published: (2025)
by: Wang, Jiaqi, et al.
Published: (2025)
Efficient Motion Prompt Learning for Robust Visual Tracking
by: Zhao, Jie, et al.
Published: (2025)
by: Zhao, Jie, et al.
Published: (2025)
SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning
by: Kong, Fanqi, et al.
Published: (2025)
by: Kong, Fanqi, et al.
Published: (2025)
A Vanilla Multi-Task Framework for Dense Visual Prediction Solution to 1st VCL Challenge -- Multi-Task Robustness Track
by: Chen, Zehui, et al.
Published: (2024)
by: Chen, Zehui, et al.
Published: (2024)
The Knowledge Microscope: Features as Better Analytical Lenses than Neurons
by: Chen, Yuheng, et al.
Published: (2025)
by: Chen, Yuheng, et al.
Published: (2025)
Towards Fine-grained Large Object Segmentation 1st Place Solution to 3D AI Challenge 2020 -- Instance Segmentation Track
by: Chen, Zehui, et al.
Published: (2020)
by: Chen, Zehui, et al.
Published: (2020)
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs
by: Zhang, Zicheng, et al.
Published: (2024)
by: Zhang, Zicheng, et al.
Published: (2024)
JailbreakHunter: A Visual Analytics Approach for Jailbreak Prompts Discovery from Large-Scale Human-LLM Conversational Datasets
by: Jin, Zhihua, et al.
Published: (2024)
by: Jin, Zhihua, et al.
Published: (2024)
Similar Items
-
VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning
by: Qi, Yukun, et al.
Published: (2025) -
Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models
by: Zeng, Yu, et al.
Published: (2025) -
VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation
by: Zhao, Yiming, et al.
Published: (2026) -
SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation
by: Liu, YaoYang, et al.
Published: (2026) -
Hybrid 3D Human Pose Estimation with Monocular Video and Sparse IMUs
by: Bao, Yiming, et al.
Published: (2024)