MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Shengkang, Lin, Hongzhan, Luo, Ziyang, Ye, Zhen, Chen, Guang, Ma, Jing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges
von: Fu, Rao, et al.
Veröffentlicht: (2024)
von: Fu, Rao, et al.
Veröffentlicht: (2024)
Towards Comprehensive Stage-wise Benchmarking of Large Language Models in Fact-Checking
von: Lin, Hongzhan, et al.
Veröffentlicht: (2026)
von: Lin, Hongzhan, et al.
Veröffentlicht: (2026)
SAP-Bench: Benchmarking Multimodal Large Language Models in Surgical Action Planning
von: Xu, Mengya, et al.
Veröffentlicht: (2025)
von: Xu, Mengya, et al.
Veröffentlicht: (2025)
MMCode: Benchmarking Multimodal Large Language Models for Code Generation with Visually Rich Programming Problems
von: Li, Kaixin, et al.
Veröffentlicht: (2024)
von: Li, Kaixin, et al.
Veröffentlicht: (2024)
Probabilistic Concept Graph Reasoning for Multimodal Misinformation Detection
von: Yang, Ruichao, et al.
Veröffentlicht: (2026)
von: Yang, Ruichao, et al.
Veröffentlicht: (2026)
AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception
von: Huang, Yipo, et al.
Veröffentlicht: (2024)
von: Huang, Yipo, et al.
Veröffentlicht: (2024)
VisualSimpleQA: A Benchmark for Decoupled Evaluation of Large Vision-Language Models in Fact-Seeking Question Answering
von: Wang, Yanling, et al.
Veröffentlicht: (2025)
von: Wang, Yanling, et al.
Veröffentlicht: (2025)
Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
von: Li, Chenxu, et al.
Veröffentlicht: (2025)
von: Li, Chenxu, et al.
Veröffentlicht: (2025)
Can Large Vision-Language Models Understand Multimodal Sarcasm?
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language Models
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
von: Chen, Yi, et al.
Veröffentlicht: (2023)
von: Chen, Yi, et al.
Veröffentlicht: (2023)
Heron-Bench: A Benchmark for Evaluating Vision Language Models in Japanese
von: Inoue, Yuichi, et al.
Veröffentlicht: (2024)
von: Inoue, Yuichi, et al.
Veröffentlicht: (2024)
PM4Bench: Benchmarking Large Vision-Language Models with Parallel Multilingual Multi-Modal Multi-task Corpus
von: Gao, Junyuan, et al.
Veröffentlicht: (2025)
von: Gao, Junyuan, et al.
Veröffentlicht: (2025)
VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict Resolution
von: Tian, Baoliang, et al.
Veröffentlicht: (2025)
von: Tian, Baoliang, et al.
Veröffentlicht: (2025)
MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
von: Wang, Zhaowei, et al.
Veröffentlicht: (2025)
von: Wang, Zhaowei, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations in Large Vision-Language Models with Internal Fact-based Contrastive Decoding
von: Wang, Chao, et al.
Veröffentlicht: (2025)
von: Wang, Chao, et al.
Veröffentlicht: (2025)
ChartCheck: Explainable Fact-Checking over Real-World Chart Images
von: Akhtar, Mubashara, et al.
Veröffentlicht: (2023)
von: Akhtar, Mubashara, et al.
Veröffentlicht: (2023)
CODIS: Benchmarking Context-Dependent Visual Comprehension for Multimodal Large Language Models
von: Luo, Fuwen, et al.
Veröffentlicht: (2024)
von: Luo, Fuwen, et al.
Veröffentlicht: (2024)
Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models
von: Ding, Meidan, et al.
Veröffentlicht: (2025)
von: Ding, Meidan, et al.
Veröffentlicht: (2025)
LVLM-Compress-Bench: Benchmarking the Broader Impact of Large Vision-Language Model Compression
von: Kundu, Souvik, et al.
Veröffentlicht: (2025)
von: Kundu, Souvik, et al.
Veröffentlicht: (2025)
NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
von: Li, Baiqi, et al.
Veröffentlicht: (2024)
von: Li, Baiqi, et al.
Veröffentlicht: (2024)
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
VULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understanding
von: Yu, Haorui, et al.
Veröffentlicht: (2026)
von: Yu, Haorui, et al.
Veröffentlicht: (2026)
MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
von: Xia, Peng, et al.
Veröffentlicht: (2024)
von: Xia, Peng, et al.
Veröffentlicht: (2024)
Multimodal Fact-Level Attribution for Verifiable Reasoning
von: Wan, David, et al.
Veröffentlicht: (2026)
von: Wan, David, et al.
Veröffentlicht: (2026)
CDH-Bench: A Commonsense-Driven Hallucination Benchmark for Evaluating Visual Fidelity in Vision-Language Models
von: Chen, Kesheng, et al.
Veröffentlicht: (2026)
von: Chen, Kesheng, et al.
Veröffentlicht: (2026)
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
von: Yang, Rui, et al.
Veröffentlicht: (2025)
von: Yang, Rui, et al.
Veröffentlicht: (2025)
MathScape: Benchmarking Multimodal Large Language Models in Real-World Mathematical Contexts
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
von: Jiang, Lei, et al.
Veröffentlicht: (2025)
von: Jiang, Lei, et al.
Veröffentlicht: (2025)
HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection
von: Gao, Siyuan, et al.
Veröffentlicht: (2025)
von: Gao, Siyuan, et al.
Veröffentlicht: (2025)
VLind-Bench: Measuring Language Priors in Large Vision-Language Models
von: Lee, Kang-il, et al.
Veröffentlicht: (2024)
von: Lee, Kang-il, et al.
Veröffentlicht: (2024)
A Survey on Benchmarks of Multimodal Large Language Models
von: Li, Jian, et al.
Veröffentlicht: (2024)
von: Li, Jian, et al.
Veröffentlicht: (2024)
Using Prompts to Guide Large Language Models in Imitating a Real Person's Language Style
von: Chen, Ziyang, et al.
Veröffentlicht: (2024)
von: Chen, Ziyang, et al.
Veröffentlicht: (2024)
HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
von: Guan, Tianrui, et al.
Veröffentlicht: (2023)
von: Guan, Tianrui, et al.
Veröffentlicht: (2023)
A Vision Check-up for Language Models
von: Sharma, Pratyusha, et al.
Veröffentlicht: (2024)
von: Sharma, Pratyusha, et al.
Veröffentlicht: (2024)
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation
von: Luo, Ziyang, et al.
Veröffentlicht: (2024)
von: Luo, Ziyang, et al.
Veröffentlicht: (2024)
Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models
von: Kim, Dain, et al.
Veröffentlicht: (2026)
von: Kim, Dain, et al.
Veröffentlicht: (2026)
TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions
von: He, Xingwei, et al.
Veröffentlicht: (2024)
von: He, Xingwei, et al.
Veröffentlicht: (2024)
NeuroABench: A Multimodal Evaluation Benchmark for Neurosurgical Anatomy Identification
von: Song, Ziyang, et al.
Veröffentlicht: (2025)
von: Song, Ziyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges
von: Fu, Rao, et al.
Veröffentlicht: (2024) -
Towards Comprehensive Stage-wise Benchmarking of Large Language Models in Fact-Checking
von: Lin, Hongzhan, et al.
Veröffentlicht: (2026) -
SAP-Bench: Benchmarking Multimodal Large Language Models in Surgical Action Planning
von: Xu, Mengya, et al.
Veröffentlicht: (2025) -
MMCode: Benchmarking Multimodal Large Language Models for Code Generation with Visually Rich Programming Problems
von: Li, Kaixin, et al.
Veröffentlicht: (2024) -
Probabilistic Concept Graph Reasoning for Multimodal Misinformation Detection
von: Yang, Ruichao, et al.
Veröffentlicht: (2026)