LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Raval, Ananya, Narayanan, Aravind, Khazaie, Vahid Reza, Raza, Shaina |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bias in the Picture: Benchmarking VLMs with Social-Cue News Images and LLM-as-Judge Assessment
von: Narayanan, Aravind, et al.
Veröffentlicht: (2025)
von: Narayanan, Aravind, et al.
Veröffentlicht: (2025)
HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation
von: Raza, Shaina, et al.
Veröffentlicht: (2025)
von: Raza, Shaina, et al.
Veröffentlicht: (2025)
MBIAS: Mitigating Bias in Large Language Models While Retaining Context
von: Raza, Shaina, et al.
Veröffentlicht: (2024)
von: Raza, Shaina, et al.
Veröffentlicht: (2024)
VLDBench Evaluating Multimodal Disinformation with Regulatory Alignment
von: Raza, Shaina, et al.
Veröffentlicht: (2025)
von: Raza, Shaina, et al.
Veröffentlicht: (2025)
Can Generative Models Improve Self-Supervised Representation Learning?
von: Ayromlou, Sana, et al.
Veröffentlicht: (2024)
von: Ayromlou, Sana, et al.
Veröffentlicht: (2024)
SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding
von: Radwan, Ahmed Y., et al.
Veröffentlicht: (2026)
von: Radwan, Ahmed Y., et al.
Veröffentlicht: (2026)
Do Composed Image Retrieval Benchmarks Require Multimodal Composition?
von: Attimonelli, Matteo, et al.
Veröffentlicht: (2026)
von: Attimonelli, Matteo, et al.
Veröffentlicht: (2026)
Enhancing Anomaly Detection Generalization through Knowledge Exposure: The Dual Effects of Augmentation
von: Anvari, Mohammad Akhavan, et al.
Veröffentlicht: (2024)
von: Anvari, Mohammad Akhavan, et al.
Veröffentlicht: (2024)
Multimodal Large Language Models for Image, Text, and Speech Data Augmentation: A Survey
von: Sapkota, Ranjan, et al.
Veröffentlicht: (2025)
von: Sapkota, Ranjan, et al.
Veröffentlicht: (2025)
Speak While Watching: Unleashing TRUE Real-Time Video Understanding Capability of Multimodal Large Language Models
von: Lin, Junyan, et al.
Veröffentlicht: (2026)
von: Lin, Junyan, et al.
Veröffentlicht: (2026)
Do Multimodal Large Language Models Understand Welding?
von: Khvatskii, Grigorii, et al.
Veröffentlicht: (2025)
von: Khvatskii, Grigorii, et al.
Veröffentlicht: (2025)
EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models
von: Das, Rocktim Jyoti, et al.
Veröffentlicht: (2024)
von: Das, Rocktim Jyoti, et al.
Veröffentlicht: (2024)
KazakhOCR: A Synthetic Benchmark for Evaluating Multimodal Models in Low-Resource Kazakh Script OCR
von: Gagnier, Henry, et al.
Veröffentlicht: (2026)
von: Gagnier, Henry, et al.
Veröffentlicht: (2026)
You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction
von: Lawrence, Logan, et al.
Veröffentlicht: (2025)
von: Lawrence, Logan, et al.
Veröffentlicht: (2025)
KBE-DME: Dynamic Multimodal Evaluation via Knowledge Enhanced Benchmark Evolution
von: Zhang, Junzhe, et al.
Veröffentlicht: (2025)
von: Zhang, Junzhe, et al.
Veröffentlicht: (2025)
Beyond Words: Multimodal LLM Knows When to Speak
von: Liao, Zikai, et al.
Veröffentlicht: (2025)
von: Liao, Zikai, et al.
Veröffentlicht: (2025)
Look & Mark: Leveraging Radiologist Eye Fixations and Bounding boxes in Multimodal Large Language Models for Chest X-ray Report Generation
von: Kim, Yunsoo, et al.
Veröffentlicht: (2025)
von: Kim, Yunsoo, et al.
Veröffentlicht: (2025)
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
von: Rostamkhani, Mohammadmostafa, et al.
Veröffentlicht: (2024)
von: Rostamkhani, Mohammadmostafa, et al.
Veröffentlicht: (2024)
A Culturally-diverse Multilingual Multimodal Video Benchmark & Model
von: Shafique, Bhuiyan Sanjid, et al.
Veröffentlicht: (2025)
von: Shafique, Bhuiyan Sanjid, et al.
Veröffentlicht: (2025)
Benchmarking Vision-Language Contrastive Methods for Medical Representation Learning
von: Roy, Shuvendu, et al.
Veröffentlicht: (2024)
von: Roy, Shuvendu, et al.
Veröffentlicht: (2024)
From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models
von: Yang, Cheng, et al.
Veröffentlicht: (2026)
von: Yang, Cheng, et al.
Veröffentlicht: (2026)
EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
von: Chen, Kai, et al.
Veröffentlicht: (2024)
von: Chen, Kai, et al.
Veröffentlicht: (2024)
Evaluating Fairness in Large Vision-Language Models Across Diverse Demographic Attributes and Prompts
von: Wu, Xuyang, et al.
Veröffentlicht: (2024)
von: Wu, Xuyang, et al.
Veröffentlicht: (2024)
UniSAFE: A Comprehensive Benchmark for Safety Evaluation of Unified Multimodal Models
von: Lee, Segyu, et al.
Veröffentlicht: (2026)
von: Lee, Segyu, et al.
Veröffentlicht: (2026)
GAOKAO-MM: A Chinese Human-Level Benchmark for Multimodal Models Evaluation
von: Zong, Yi, et al.
Veröffentlicht: (2024)
von: Zong, Yi, et al.
Veröffentlicht: (2024)
SAP-Bench: Benchmarking Multimodal Large Language Models in Surgical Action Planning
von: Xu, Mengya, et al.
Veröffentlicht: (2025)
von: Xu, Mengya, et al.
Veröffentlicht: (2025)
SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models
von: Xia, Haotian, et al.
Veröffentlicht: (2024)
von: Xia, Haotian, et al.
Veröffentlicht: (2024)
MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
CODIS: Benchmarking Context-Dependent Visual Comprehension for Multimodal Large Language Models
von: Luo, Fuwen, et al.
Veröffentlicht: (2024)
von: Luo, Fuwen, et al.
Veröffentlicht: (2024)
LLaVA-Critic: Learning to Evaluate Multimodal Models
von: Xiong, Tianyi, et al.
Veröffentlicht: (2024)
von: Xiong, Tianyi, et al.
Veröffentlicht: (2024)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
When Relations Break: Analyzing Relation Hallucination in Vision-Language Model Under Rotation and Noise
von: Shin, Philip Wootaek, et al.
Veröffentlicht: (2026)
von: Shin, Philip Wootaek, et al.
Veröffentlicht: (2026)
UEval: A Benchmark for Unified Multimodal Generation
von: Li, Bo, et al.
Veröffentlicht: (2026)
von: Li, Bo, et al.
Veröffentlicht: (2026)
UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents
von: Ji, Yifan, et al.
Veröffentlicht: (2026)
von: Ji, Yifan, et al.
Veröffentlicht: (2026)
MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models
von: Ruan, Jiacheng, et al.
Veröffentlicht: (2025)
von: Ruan, Jiacheng, et al.
Veröffentlicht: (2025)
Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
von: Li, Chenxu, et al.
Veröffentlicht: (2025)
von: Li, Chenxu, et al.
Veröffentlicht: (2025)
MathScape: Benchmarking Multimodal Large Language Models in Real-World Mathematical Contexts
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception
von: Huang, Yipo, et al.
Veröffentlicht: (2024)
von: Huang, Yipo, et al.
Veröffentlicht: (2024)
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
von: Ye, Maoyuan, et al.
Veröffentlicht: (2025)
von: Ye, Maoyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Bias in the Picture: Benchmarking VLMs with Social-Cue News Images and LLM-as-Judge Assessment
von: Narayanan, Aravind, et al.
Veröffentlicht: (2025) -
HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation
von: Raza, Shaina, et al.
Veröffentlicht: (2025) -
MBIAS: Mitigating Bias in Large Language Models While Retaining Context
von: Raza, Shaina, et al.
Veröffentlicht: (2024) -
VLDBench Evaluating Multimodal Disinformation with Regulatory Alignment
von: Raza, Shaina, et al.
Veröffentlicht: (2025) -
Can Generative Models Improve Self-Supervised Representation Learning?
von: Ayromlou, Sana, et al.
Veröffentlicht: (2024)