HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Raza, Shaina, Narayanan, Aravind, Khazaie, Vahid Reza, Vayani, Ashmal, Radwan, Ahmed Y., Chettiar, Mukund S., Singh, Amandeep, Shah, Mubarak, Pandya, Deval |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bias in the Picture: Benchmarking VLMs with Social-Cue News Images and LLM-as-Judge Assessment
by: Narayanan, Aravind, et al.
Published: (2025)
by: Narayanan, Aravind, et al.
Published: (2025)
VLDBench Evaluating Multimodal Disinformation with Regulatory Alignment
by: Raza, Shaina, et al.
Published: (2025)
by: Raza, Shaina, et al.
Published: (2025)
LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation
by: Raval, Ananya, et al.
Published: (2025)
by: Raval, Ananya, et al.
Published: (2025)
SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding
by: Radwan, Ahmed Y., et al.
Published: (2026)
by: Radwan, Ahmed Y., et al.
Published: (2026)
BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models
by: Narnaware, Vishal, et al.
Published: (2025)
by: Narnaware, Vishal, et al.
Published: (2025)
FairSense-AI: Responsible AI Meets Sustainability
by: Raza, Shaina, et al.
Published: (2025)
by: Raza, Shaina, et al.
Published: (2025)
Beyond Content: How Grammatical Gender Shapes Visual Representation in Text-to-Image Models
by: Saeed, Muhammed, et al.
Published: (2025)
by: Saeed, Muhammed, et al.
Published: (2025)
MedRoute: RL-Based Dynamic Specialist Routing in Multi-Agent Medical Diagnosis
by: Vayani, Ashmal, et al.
Published: (2026)
by: Vayani, Ashmal, et al.
Published: (2026)
Learning to Share: Selective Memory for Efficient Parallel Agentic Systems
by: Fioresi, Joseph, et al.
Published: (2026)
by: Fioresi, Joseph, et al.
Published: (2026)
Position: Sustainable Open-Source AI Requires Tracking the Cumulative Footprint of Derivatives
by: Raza, Shaina, et al.
Published: (2026)
by: Raza, Shaina, et al.
Published: (2026)
FAIR Enough: How Can We Develop and Assess a FAIR-Compliant Dataset for Large Language Models' Training?
by: Raza, Shaina, et al.
Published: (2024)
by: Raza, Shaina, et al.
Published: (2024)
Practical Guide for Causal Pathways and Sub-group Disparity Analysis
by: Kohankhaki, Farnaz, et al.
Published: (2024)
by: Kohankhaki, Farnaz, et al.
Published: (2024)
Just as Humans Need Vaccines, So Do Models: Model Immunization to Combat Falsehoods
by: Raza, Shaina, et al.
Published: (2025)
by: Raza, Shaina, et al.
Published: (2025)
GAEA: A Geolocation Aware Conversational Assistant
by: Campos, Ron, et al.
Published: (2025)
by: Campos, Ron, et al.
Published: (2025)
From Features to Actions: Explainability in Traditional and Agentic AI Systems
by: Chaduvula, Sindhuja, et al.
Published: (2026)
by: Chaduvula, Sindhuja, et al.
Published: (2026)
Reducing Hallucinations in LLMs via Factuality-Aware Preference Learning
by: Chaduvula, Sindhuja, et al.
Published: (2026)
by: Chaduvula, Sindhuja, et al.
Published: (2026)
The Deepfakes We Missed: We Built Detectors for a Threat That Didn't Arrive
by: Raza, Shaina
Published: (2026)
by: Raza, Shaina
Published: (2026)
Ius Humani
Published: (2013)
Published: (2013)
VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding
by: Mahmood, Ahmad, et al.
Published: (2024)
by: Mahmood, Ahmad, et al.
Published: (2024)
ViLBias: Detecting and Reasoning about Bias in Multimodal Content
by: Raza, Shaina, et al.
Published: (2024)
by: Raza, Shaina, et al.
Published: (2024)
Can Generative Models Improve Self-Supervised Representation Learning?
by: Ayromlou, Sana, et al.
Published: (2024)
by: Ayromlou, Sana, et al.
Published: (2024)
Jurnal Sosiologi Pendidikan Humanis
Published: (2019)
Published: (2019)
Can LLMs Capture Human Preferences?
by: Goli, Ali, et al.
Published: (2023)
by: Goli, Ali, et al.
Published: (2023)
Enhancing Anomaly Detection Generalization through Knowledge Exposure: The Dual Effects of Augmentation
by: Anvari, Mohammad Akhavan, et al.
Published: (2024)
by: Anvari, Mohammad Akhavan, et al.
Published: (2024)
A Narrative Review of Identity, Data, and Location Privacy Techniques in Edge Computing and Mobile Crowdsourcing
by: Bashir, Syed Raza, et al.
Published: (2024)
by: Bashir, Syed Raza, et al.
Published: (2024)
Desert Camels and Oil Sheikhs: Arab-Centric Red Teaming of Frontier LLMs
by: Saeed, Muhammed, et al.
Published: (2024)
by: Saeed, Muhammed, et al.
Published: (2024)
FROM LEADERSHIP TO ADVOCACY: HOW INTERNAL MARKETING AND HRM FOSTER EMPLOYEE–BRAND AMBASSADORS
by: Silvia Ahmed Khattak,Raza Mubarak,Ghayyur Qadir
Published: (2025)
by: Silvia Ahmed Khattak,Raza Mubarak,Ghayyur Qadir
Published: (2025)
Influencia del Índice de Masa Corporal y la actividad física en el comportamiento alimentario de los consumidores españoles
by: Amr Radwan Ahmed Radwan
Published: (2017)
by: Amr Radwan Ahmed Radwan
Published: (2017)
Comprehensive Analysis of Transparency and Accessibility of ChatGPT, DeepSeek, And other SoTA Large Language Models
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Fact or Fiction? Can LLMs be Reliable Annotators for Political Truths?
by: Chatrath, Veronica, et al.
Published: (2024)
by: Chatrath, Veronica, et al.
Published: (2024)
MBIAS: Mitigating Bias in Large Language Models While Retaining Context
by: Raza, Shaina, et al.
Published: (2024)
by: Raza, Shaina, et al.
Published: (2024)
Detecting Deception, Not Deepfakes: Why Media Forensics Needs Social Theories
by: Ho, Jessee, et al.
Published: (2026)
by: Ho, Jessee, et al.
Published: (2026)
Analyzing the Impact of Fake News on the Anticipated Outcome of the 2024 Election Ahead of Time
by: Raza, Shaina, et al.
Published: (2023)
by: Raza, Shaina, et al.
Published: (2023)
Multimodal Large Language Models for Image, Text, and Speech Data Augmentation: A Survey
by: Sapkota, Ranjan, et al.
Published: (2025)
by: Sapkota, Ranjan, et al.
Published: (2025)
Promoting wellness through pet‐friendly academic work environments
by: Ahmed Radwan
Published: (2026)
by: Ahmed Radwan
Published: (2026)
Promoting wellness through pet‐friendly work environments
by: Ahmed Radwan
Published: (2026)
by: Ahmed Radwan
Published: (2026)
Promoting wellness through pet‐friendly academic work environments
by: Ahmed Radwan
Published: (2026)
by: Ahmed Radwan
Published: (2026)
Promoting Wellness Through Pet‐Friendly Academic Work Environments
by: Ahmed Radwan
Published: (2026)
by: Ahmed Radwan
Published: (2026)
Who is Responsible? The Data, Models, Users or Regulations? A Comprehensive Survey on Responsible Generative AI for a Sustainable Future
by: Raza, Shaina, et al.
Published: (2025)
by: Raza, Shaina, et al.
Published: (2025)
Thinking Beyond Tokens: From Brain-Inspired Intelligence to Cognitive Foundations for Artificial General Intelligence and its Societal Impact
by: Qureshi, Rizwan, et al.
Published: (2025)
by: Qureshi, Rizwan, et al.
Published: (2025)
Similar Items
-
Bias in the Picture: Benchmarking VLMs with Social-Cue News Images and LLM-as-Judge Assessment
by: Narayanan, Aravind, et al.
Published: (2025) -
VLDBench Evaluating Multimodal Disinformation with Regulatory Alignment
by: Raza, Shaina, et al.
Published: (2025) -
LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation
by: Raval, Ananya, et al.
Published: (2025) -
SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding
by: Radwan, Ahmed Y., et al.
Published: (2026) -
BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models
by: Narnaware, Vishal, et al.
Published: (2025)