PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ouyang, Kun, Liu, Yuanxin, Li, Shicheng, Liu, Yi, Zhou, Hao, Meng, Fandong, Zhou, Jie, Sun, Xu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
von: Ouyang, Kun, et al.
Veröffentlicht: (2025)
von: Ouyang, Kun, et al.
Veröffentlicht: (2025)
Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence
von: Ouyang, Kun, et al.
Veröffentlicht: (2025)
von: Ouyang, Kun, et al.
Veröffentlicht: (2025)
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
von: Liu, Yuanxin, et al.
Veröffentlicht: (2025)
von: Liu, Yuanxin, et al.
Veröffentlicht: (2025)
Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters
von: Xu, Zhiyu, et al.
Veröffentlicht: (2026)
von: Xu, Zhiyu, et al.
Veröffentlicht: (2026)
Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models
von: Wei, Yuancheng, et al.
Veröffentlicht: (2026)
von: Wei, Yuancheng, et al.
Veröffentlicht: (2026)
TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT Alignment
von: Li, Shicheng, et al.
Veröffentlicht: (2025)
von: Li, Shicheng, et al.
Veröffentlicht: (2025)
EventBench: Towards Comprehensive Benchmarking of Event-based MLLMs
von: Liu, Shaoyu, et al.
Veröffentlicht: (2025)
von: Liu, Shaoyu, et al.
Veröffentlicht: (2025)
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
von: Gao, Tianyi, et al.
Veröffentlicht: (2025)
von: Gao, Tianyi, et al.
Veröffentlicht: (2025)
ST-BiBench: Benchmarking Multi-Stream Multimodal Coordination in Bimanual Embodied Tasks for MLLMs
von: Wu, Xin, et al.
Veröffentlicht: (2026)
von: Wu, Xin, et al.
Veröffentlicht: (2026)
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models
von: Liu, Juntao, et al.
Veröffentlicht: (2025)
von: Liu, Juntao, et al.
Veröffentlicht: (2025)
AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity
von: Lan, Zhibin, et al.
Veröffentlicht: (2024)
von: Lan, Zhibin, et al.
Veröffentlicht: (2024)
VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models
von: Li, Shicheng, et al.
Veröffentlicht: (2023)
von: Li, Shicheng, et al.
Veröffentlicht: (2023)
AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models
von: Zhou, Yutong, et al.
Veröffentlicht: (2024)
von: Zhou, Yutong, et al.
Veröffentlicht: (2024)
UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios
von: Zhou, Baichuan, et al.
Veröffentlicht: (2024)
von: Zhou, Baichuan, et al.
Veröffentlicht: (2024)
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
von: Qiu, Yansheng, et al.
Veröffentlicht: (2025)
von: Qiu, Yansheng, et al.
Veröffentlicht: (2025)
360° Image Perception with MLLMs: A Comprehensive Benchmark and a Training-Free Method
von: Tran, Huyen T. T., et al.
Veröffentlicht: (2026)
von: Tran, Huyen T. T., et al.
Veröffentlicht: (2026)
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
von: Lin, Junming, et al.
Veröffentlicht: (2024)
von: Lin, Junming, et al.
Veröffentlicht: (2024)
Punching Bag vs. Punching Person: Motion Transferability in Videos
von: Abdullah, Raiyaan, et al.
Veröffentlicht: (2025)
von: Abdullah, Raiyaan, et al.
Veröffentlicht: (2025)
MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs
von: Yuan, Jiakang, et al.
Veröffentlicht: (2025)
von: Yuan, Jiakang, et al.
Veröffentlicht: (2025)
Continuous Visual Autoregressive Generation via Score Maximization
von: Shao, Chenze, et al.
Veröffentlicht: (2025)
von: Shao, Chenze, et al.
Veröffentlicht: (2025)
Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
von: Anand, Dhruv, et al.
Veröffentlicht: (2025)
von: Anand, Dhruv, et al.
Veröffentlicht: (2025)
VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents
von: Zhang, Zhengbo, et al.
Veröffentlicht: (2026)
von: Zhang, Zhengbo, et al.
Veröffentlicht: (2026)
MileBench: Benchmarking MLLMs in Long Context
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
von: Song, Dingjie, et al.
Veröffentlicht: (2024)
MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs
von: Xu, Yunqiu, et al.
Veröffentlicht: (2024)
von: Xu, Yunqiu, et al.
Veröffentlicht: (2024)
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
von: Cheng, Ziming, et al.
Veröffentlicht: (2025)
von: Cheng, Ziming, et al.
Veröffentlicht: (2025)
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
von: Ouyang, Linke, et al.
Veröffentlicht: (2024)
von: Ouyang, Linke, et al.
Veröffentlicht: (2024)
Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
von: Mao, Jiawei, et al.
Veröffentlicht: (2025)
von: Mao, Jiawei, et al.
Veröffentlicht: (2025)
MER-Bench: A Comprehensive Benchmark for Multimodal Meme Reappraisal
von: Nie, Yiqi, et al.
Veröffentlicht: (2026)
von: Nie, Yiqi, et al.
Veröffentlicht: (2026)
OmniBrainBench: A Comprehensive Multimodal Benchmark for Brain Imaging Analysis Across Multi-stage Clinical Tasks
von: Peng, Zhihao, et al.
Veröffentlicht: (2025)
von: Peng, Zhihao, et al.
Veröffentlicht: (2025)
EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs
von: Dai, Yang, et al.
Veröffentlicht: (2026)
von: Dai, Yang, et al.
Veröffentlicht: (2026)
HM-Bench: A Comprehensive Benchmark for Multimodal Large Language Models in Hyperspectral Remote Sensing
von: Zhang, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2026)
MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification
von: Xu, Zhaopan, et al.
Veröffentlicht: (2025)
von: Xu, Zhaopan, et al.
Veröffentlicht: (2025)
MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
UVE: Are MLLMs Unified Evaluators for AI-Generated Videos?
von: Liu, Yuanxin, et al.
Veröffentlicht: (2025)
von: Liu, Yuanxin, et al.
Veröffentlicht: (2025)
VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning
von: Jia, Zi-Yi, et al.
Veröffentlicht: (2026)
von: Jia, Zi-Yi, et al.
Veröffentlicht: (2026)
ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models
von: Xue, Kaiwen, et al.
Veröffentlicht: (2026)
von: Xue, Kaiwen, et al.
Veröffentlicht: (2026)
EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios
von: Qiu, Lu, et al.
Veröffentlicht: (2024)
von: Qiu, Lu, et al.
Veröffentlicht: (2024)
MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection
von: Jiang, Xi, et al.
Veröffentlicht: (2024)
von: Jiang, Xi, et al.
Veröffentlicht: (2024)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
von: Xu, Zelin, et al.
Veröffentlicht: (2026)
von: Xu, Zelin, et al.
Veröffentlicht: (2026)
IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
von: Ouyang, Kun, et al.
Veröffentlicht: (2025) -
Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence
von: Ouyang, Kun, et al.
Veröffentlicht: (2025) -
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
von: Liu, Yuanxin, et al.
Veröffentlicht: (2025) -
Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters
von: Xu, Zhiyu, et al.
Veröffentlicht: (2026) -
Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models
von: Wei, Yuancheng, et al.
Veröffentlicht: (2026)