MANBench: Is Your Multimodal Model Smarter than Human?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Han, Xu, Qitong, Dong, Yiheng, Yang, Xin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reasoning with Sampling: Your Base Model is Smarter Than You Think
von: Karan, Aayush, et al.
Veröffentlicht: (2025)
von: Karan, Aayush, et al.
Veröffentlicht: (2025)
Are Large Language Models Truly Smarter Than Humans?
von: M, Eshwar Reddy, et al.
Veröffentlicht: (2026)
von: M, Eshwar Reddy, et al.
Veröffentlicht: (2026)
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges
von: Fu, Rao, et al.
Veröffentlicht: (2024)
von: Fu, Rao, et al.
Veröffentlicht: (2024)
Tougher Text, Smarter Models: Raising the Bar for Adversarial Defence Benchmarks
von: Wang, Yang, et al.
Veröffentlicht: (2025)
von: Wang, Yang, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Multimodal Model for Fake News Detection
von: Li, Yiheng, et al.
Veröffentlicht: (2026)
von: Li, Yiheng, et al.
Veröffentlicht: (2026)
We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
von: Qiao, Runqi, et al.
Veröffentlicht: (2024)
von: Qiao, Runqi, et al.
Veröffentlicht: (2024)
Informed Routing in LLMs: Smarter Token-Level Computation for Faster Inference
von: Han, Chao, et al.
Veröffentlicht: (2025)
von: Han, Chao, et al.
Veröffentlicht: (2025)
Are Large Language Models More Empathetic than Humans?
von: Welivita, Anuradha, et al.
Veröffentlicht: (2024)
von: Welivita, Anuradha, et al.
Veröffentlicht: (2024)
CamemBERT 2.0: A Smarter French Language Model Aged to Perfection
von: Antoun, Wissam, et al.
Veröffentlicht: (2024)
von: Antoun, Wissam, et al.
Veröffentlicht: (2024)
Multi-Sense Embeddings for Language Models and Knowledge Distillation
von: Wang, Qitong, et al.
Veröffentlicht: (2025)
von: Wang, Qitong, et al.
Veröffentlicht: (2025)
Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models
von: Xu, Xin, et al.
Veröffentlicht: (2026)
von: Xu, Xin, et al.
Veröffentlicht: (2026)
SUGAR: Leveraging Contextual Confidence for Smarter Retrieval
von: Zubkova, Hanna, et al.
Veröffentlicht: (2025)
von: Zubkova, Hanna, et al.
Veröffentlicht: (2025)
VRM: Teaching Reward Models to Understand Authentic Human Preferences
von: Liu, Biao, et al.
Veröffentlicht: (2026)
von: Liu, Biao, et al.
Veröffentlicht: (2026)
Experiential Semantic Information and Brain Alignment: Are Multimodal Models Better than Language Models?
von: Bavaresco, Anna, et al.
Veröffentlicht: (2025)
von: Bavaresco, Anna, et al.
Veröffentlicht: (2025)
Benchmarks Saturate When The Model Gets Smarter Than The Judge
von: Ballon, Marthe, et al.
Veröffentlicht: (2026)
von: Ballon, Marthe, et al.
Veröffentlicht: (2026)
Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology
von: Hu, Jucheng, et al.
Veröffentlicht: (2026)
von: Hu, Jucheng, et al.
Veröffentlicht: (2026)
MOSSBench: Is Your Multimodal Language Model Oversensitive to Safe Queries?
von: Li, Xirui, et al.
Veröffentlicht: (2024)
von: Li, Xirui, et al.
Veröffentlicht: (2024)
Your Multimodal Speech Model Says I Have a Face for Radio
von: Nachesa, Maya K., et al.
Veröffentlicht: (2026)
von: Nachesa, Maya K., et al.
Veröffentlicht: (2026)
How to Train Your Fact Verifier: Knowledge Transfer with Multimodal Open Models
von: Lee, Jaeyoung, et al.
Veröffentlicht: (2024)
von: Lee, Jaeyoung, et al.
Veröffentlicht: (2024)
Smarter, not Bigger: Fine-Tuned RAG-Enhanced LLMs for Automotive HIL Testing
von: Feng, Chao, et al.
Veröffentlicht: (2025)
von: Feng, Chao, et al.
Veröffentlicht: (2025)
De-biased Multimodal Electrocardiogram Analysis
von: Li, Haitao, et al.
Veröffentlicht: (2024)
von: Li, Haitao, et al.
Veröffentlicht: (2024)
Don't Throw Away Your Pretrained Model
von: Feng, Shangbin, et al.
Veröffentlicht: (2025)
von: Feng, Shangbin, et al.
Veröffentlicht: (2025)
How to Merge Your Multimodal Models Over Time?
von: Dziadzio, Sebastian, et al.
Veröffentlicht: (2024)
von: Dziadzio, Sebastian, et al.
Veröffentlicht: (2024)
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision
von: Pala, Tej Deep, et al.
Veröffentlicht: (2025)
von: Pala, Tej Deep, et al.
Veröffentlicht: (2025)
The Better You Learn, The Smarter You Prune: Towards Efficient Vision-language-action Models via Differentiable Token Pruning
von: Jiang, Titong, et al.
Veröffentlicht: (2025)
von: Jiang, Titong, et al.
Veröffentlicht: (2025)
TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation
von: Jeong, Yeil, et al.
Veröffentlicht: (2026)
von: Jeong, Yeil, et al.
Veröffentlicht: (2026)
Data Distribution Bottlenecks in Grounding Language Models to Knowledge Bases
von: Shu, Yiheng, et al.
Veröffentlicht: (2023)
von: Shu, Yiheng, et al.
Veröffentlicht: (2023)
TablePilot: Recommending Human-Preferred Tabular Data Analysis with Large Language Models
von: Yi, Deyin, et al.
Veröffentlicht: (2025)
von: Yi, Deyin, et al.
Veröffentlicht: (2025)
How Easy is It to Fool Your Multimodal LLMs? An Empirical Analysis on Deceptive Prompts
von: Qian, Yusu, et al.
Veröffentlicht: (2024)
von: Qian, Yusu, et al.
Veröffentlicht: (2024)
CoCA: Regaining Safety-awareness of Multimodal Large Language Models with Constitutional Calibration
von: Gao, Jiahui, et al.
Veröffentlicht: (2024)
von: Gao, Jiahui, et al.
Veröffentlicht: (2024)
Locas: Your Models are Principled Initializers of Locally-Supported Parametric Memories
von: Lu, Sidi, et al.
Veröffentlicht: (2026)
von: Lu, Sidi, et al.
Veröffentlicht: (2026)
LargePiG: Your Large Language Model is Secretly a Pointer Generator
von: Sun, Zhongxiang, et al.
Veröffentlicht: (2024)
von: Sun, Zhongxiang, et al.
Veröffentlicht: (2024)
Generating Effective CoT Traces for Mitigating Causal Hallucination
von: Zhao, Yiheng, et al.
Veröffentlicht: (2026)
von: Zhao, Yiheng, et al.
Veröffentlicht: (2026)
MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
von: Chen, Zhaorun, et al.
Veröffentlicht: (2024)
von: Chen, Zhaorun, et al.
Veröffentlicht: (2024)
Is ChatGPT More Empathetic than Humans?
von: Welivita, Anuradha, et al.
Veröffentlicht: (2024)
von: Welivita, Anuradha, et al.
Veröffentlicht: (2024)
Can Large Multimodal Models Uncover Deep Semantics Behind Images?
von: Yang, Yixin, et al.
Veröffentlicht: (2024)
von: Yang, Yixin, et al.
Veröffentlicht: (2024)
Are Smarter LLMs Safer? Exploring Safety-Reasoning Trade-offs in Prompting and Fine-Tuning
von: Li, Ang, et al.
Veröffentlicht: (2025)
von: Li, Ang, et al.
Veröffentlicht: (2025)
SpeechAgents: Human-Communication Simulation with Multi-Modal Multi-Agent Systems
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine
von: Xiao, Hanguang, et al.
Veröffentlicht: (2024)
von: Xiao, Hanguang, et al.
Veröffentlicht: (2024)
Don't Forget Your Reward Values: Language Model Alignment via Value-based Calibration
von: Mao, Xin, et al.
Veröffentlicht: (2024)
von: Mao, Xin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Reasoning with Sampling: Your Base Model is Smarter Than You Think
von: Karan, Aayush, et al.
Veröffentlicht: (2025) -
Are Large Language Models Truly Smarter Than Humans?
von: M, Eshwar Reddy, et al.
Veröffentlicht: (2026) -
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges
von: Fu, Rao, et al.
Veröffentlicht: (2024) -
Tougher Text, Smarter Models: Raising the Bar for Adversarial Defence Benchmarks
von: Wang, Yang, et al.
Veröffentlicht: (2025) -
Retrieval-Augmented Multimodal Model for Fake News Detection
von: Li, Yiheng, et al.
Veröffentlicht: (2026)