WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics
Fuente:
arXiv
Saved in:
| Main Authors: | Maurya, Sneha, Saboo, Pragya, Kumar, Girish |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Are LLMs Court-Ready? Evaluating Frontier Models on Indian Legal Reasoning
by: Juvekar, Kush, et al.
Published: (2025)
by: Juvekar, Kush, et al.
Published: (2025)
Topic-aware Large Language Models for Summarizing the Lived Healthcare Experiences Described in Health Stories
by: Bilalpur, Maneesh, et al.
Published: (2025)
by: Bilalpur, Maneesh, et al.
Published: (2025)
Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation
by: Gringras, David, et al.
Published: (2026)
by: Gringras, David, et al.
Published: (2026)
AI Governance and Accountability: An Analysis of Anthropic's Claude
by: Priyanshu, Aman, et al.
Published: (2024)
by: Priyanshu, Aman, et al.
Published: (2024)
Towards Safe Multilingual Frontier AI
by: Kanepajs, Artūrs, et al.
Published: (2024)
by: Kanepajs, Artūrs, et al.
Published: (2024)
White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs
by: Wan, Yixin, et al.
Published: (2024)
by: Wan, Yixin, et al.
Published: (2024)
On the Credibility of Evaluating LLMs using Survey Questions
by: Libovický, Jindřich
Published: (2026)
by: Libovický, Jindřich
Published: (2026)
MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
by: Ghosh, Rajarshi, et al.
Published: (2025)
by: Ghosh, Rajarshi, et al.
Published: (2025)
Findings of the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-powered Tutors
by: Kochmar, Ekaterina, et al.
Published: (2025)
by: Kochmar, Ekaterina, et al.
Published: (2025)
Lived Experience Not Found: LLMs Struggle to Align with Experts on Addressing Adverse Drug Reactions from Psychiatric Medication Use
by: Chandra, Mohit, et al.
Published: (2024)
by: Chandra, Mohit, et al.
Published: (2024)
Evaluating the Capabilities of LLMs for Supporting Anticipatory Impact Assessment
by: Allaham, Mowafak, et al.
Published: (2024)
by: Allaham, Mowafak, et al.
Published: (2024)
Evaluating Cultural Awareness of LLMs for Yoruba, Malayalam, and English
by: Dawson, Fiifi, et al.
Published: (2024)
by: Dawson, Fiifi, et al.
Published: (2024)
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
by: Kabir, Mohsinul, et al.
Published: (2025)
by: Kabir, Mohsinul, et al.
Published: (2025)
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
by: Haider, Batool, et al.
Published: (2025)
by: Haider, Batool, et al.
Published: (2025)
Evaluating GPT-3.5's Awareness and Summarization Abilities for European Constitutional Texts with Shared Topics
by: Greco, Candida M., et al.
Published: (2024)
by: Greco, Candida M., et al.
Published: (2024)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
by: Kabir, Mohsinul, et al.
Published: (2026)
by: Kabir, Mohsinul, et al.
Published: (2026)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
by: Rystrøm, Jonathan, et al.
Published: (2025)
by: Rystrøm, Jonathan, et al.
Published: (2025)
RedTopic: Toward Topic-Diverse Red Teaming of Large Language Models
by: Ding, Jiale, et al.
Published: (2025)
by: Ding, Jiale, et al.
Published: (2025)
PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
by: Shi, Yuzhen, et al.
Published: (2026)
by: Shi, Yuzhen, et al.
Published: (2026)
ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases
by: Li, Yuchong, et al.
Published: (2025)
by: Li, Yuchong, et al.
Published: (2025)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
Do Large Language Models Get Caught in Hofstadter-Mobius Loops?
by: Hryszko, Jaroslaw
Published: (2026)
by: Hryszko, Jaroslaw
Published: (2026)
Reducing Large Language Model Safety Risks in Women's Health using Semantic Entropy
by: Penny-Dimri, Jahan C., et al.
Published: (2025)
by: Penny-Dimri, Jahan C., et al.
Published: (2025)
Patterns vs. Patients: Evaluating LLMs against Mental Health Professionals on Personality Disorder Diagnosis through First-Person Narratives
by: Drożdż, Karolina, et al.
Published: (2025)
by: Drożdż, Karolina, et al.
Published: (2025)
Evaluation of LLMs for Process Model Analysis and Optimization
by: Kumar, Akhil, et al.
Published: (2025)
by: Kumar, Akhil, et al.
Published: (2025)
Assessing the Performance of Human-Capable LLMs -- Are LLMs Coming for Your Job?
by: Mavi, John, et al.
Published: (2024)
by: Mavi, John, et al.
Published: (2024)
Culturally Adaptive Explainable LLM Assessment for Multilingual Information Disorder: A Human-in-the-Loop Approach
by: Jouneghani, Maziar Kianimoghadam
Published: (2026)
by: Jouneghani, Maziar Kianimoghadam
Published: (2026)
The simulation of judgment in LLMs
by: Loru, Edoardo, et al.
Published: (2025)
by: Loru, Edoardo, et al.
Published: (2025)
Measuring Teaching with LLMs
by: Hardy, Michael
Published: (2025)
by: Hardy, Michael
Published: (2025)
The Political Preferences of LLMs
by: Rozado, David
Published: (2024)
by: Rozado, David
Published: (2024)
Open Source Language Models Can Provide Feedback: Evaluating LLMs' Ability to Help Students Using GPT-4-As-A-Judge
by: Koutcheme, Charles, et al.
Published: (2024)
by: Koutcheme, Charles, et al.
Published: (2024)
Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs
by: de Landa, Joseba Fernandez, et al.
Published: (2026)
by: de Landa, Joseba Fernandez, et al.
Published: (2026)
Readers Prefer Outputs of AI Trained on Copyrighted Books over Expert Human Writers
by: Chakrabarty, Tuhin, et al.
Published: (2025)
by: Chakrabarty, Tuhin, et al.
Published: (2025)
Moral Mazes in the Era of LLMs
by: Nguyen, Dang, et al.
Published: (2026)
by: Nguyen, Dang, et al.
Published: (2026)
Evaluating Patient Safety Risks in Generative AI: Development and Validation of a FMECA Framework for Generated Clinical Content
by: Bednarczyk, Lydie, et al.
Published: (2026)
by: Bednarczyk, Lydie, et al.
Published: (2026)
neuralFOMO: Can LLMs Handle Being Second Best? Measuring Envy-Like Preferences in Multi-Agent Settings
by: Ramamoorthy, Arnav, et al.
Published: (2025)
by: Ramamoorthy, Arnav, et al.
Published: (2025)
Toward Inclusive Educational AI: Auditing Frontier LLMs through a Multiplexity Lens
by: Mushtaq, Abdullah, et al.
Published: (2025)
by: Mushtaq, Abdullah, et al.
Published: (2025)
Human vs. Machine: Behavioral Differences Between Expert Humans and Language Models in Wargame Simulations
by: Lamparth, Max, et al.
Published: (2024)
by: Lamparth, Max, et al.
Published: (2024)
Are LLMs (Really) Ideological? An IRT-based Analysis and Alignment Tool for Perceived Socio-Economic Bias in LLMs
by: Wachter, Jasmin, et al.
Published: (2025)
by: Wachter, Jasmin, et al.
Published: (2025)
Interpretability Framework for LLMs in Undergraduate Calculus
by: Dakshit, Sagnik, et al.
Published: (2025)
by: Dakshit, Sagnik, et al.
Published: (2025)
Similar Items
-
Are LLMs Court-Ready? Evaluating Frontier Models on Indian Legal Reasoning
by: Juvekar, Kush, et al.
Published: (2025) -
Topic-aware Large Language Models for Summarizing the Lived Healthcare Experiences Described in Health Stories
by: Bilalpur, Maneesh, et al.
Published: (2025) -
Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation
by: Gringras, David, et al.
Published: (2026) -
AI Governance and Accountability: An Analysis of Anthropic's Claude
by: Priyanshu, Aman, et al.
Published: (2024) -
Towards Safe Multilingual Frontier AI
by: Kanepajs, Artūrs, et al.
Published: (2024)