Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hamna, Hamna, Bhat, Gayatri, Mukherjee, Sourabrata, Lalani, Faisal, Hadfield, Evan, Siddarth, Divya, Bali, Kalika, Sitaram, Sunayana |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DOSA: A Dataset of Social Artifacts from Different Indian Geographical Subcultures
von: Seth, Agrima, et al.
Veröffentlicht: (2024)
von: Seth, Agrima, et al.
Veröffentlicht: (2024)
How Deep Is Representational Bias in LLMs? The Cases of Caste and Religion
von: Seth, Agrima, et al.
Veröffentlicht: (2025)
von: Seth, Agrima, et al.
Veröffentlicht: (2025)
METAL: Towards Multilingual Meta-Evaluation
von: Hada, Rishav, et al.
Veröffentlicht: (2024)
von: Hada, Rishav, et al.
Veröffentlicht: (2024)
Alzheimer’s Disease and Healthy Aging Indicators Cognitive Decline
von: Hamna Khuld
Veröffentlicht: (2024)
von: Hamna Khuld
Veröffentlicht: (2024)
Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic Prompting
von: Mukherjee, Sagnik, et al.
Veröffentlicht: (2024)
von: Mukherjee, Sagnik, et al.
Veröffentlicht: (2024)
Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation?
von: Hada, Rishav, et al.
Veröffentlicht: (2023)
von: Hada, Rishav, et al.
Veröffentlicht: (2023)
Bridging the Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
von: Kumar, Somnath, et al.
Veröffentlicht: (2024)
von: Kumar, Somnath, et al.
Veröffentlicht: (2024)
Bridging the Language Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
von: Kumar, Somnath, et al.
Veröffentlicht: (2023)
von: Kumar, Somnath, et al.
Veröffentlicht: (2023)
HEALTH-PARIKSHA: Assessing RAG Models for Health Chatbots in Real-World Multilingual Settings
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
Beyond Metrics: Evaluating LLMs' Effectiveness in Culturally Nuanced, Low-Resource Real-World Scenarios
von: Ochieng, Millicent, et al.
Veröffentlicht: (2024)
von: Ochieng, Millicent, et al.
Veröffentlicht: (2024)
KAHANI: Culturally-Nuanced Visual Storytelling Tool for Non-Western Cultures
von: Hamna, et al.
Veröffentlicht: (2024)
von: Hamna, et al.
Veröffentlicht: (2024)
A hybrid variational quantum circuit approach for stabilizer states classifiers
von: Aslam, Hamna, et al.
Veröffentlicht: (2025)
von: Aslam, Hamna, et al.
Veröffentlicht: (2025)
Methodological Considerations in Interpreting Mortality Trends in Aortic Dissection Among Hypertensive Adults
von: Hamna Bibi, et al.
Veröffentlicht: (2026)
von: Hamna Bibi, et al.
Veröffentlicht: (2026)
Methodological Considerations and Evidence Interpretation in the Phase III Evaluation of the Absnow Biodegradable ASD Occluder
von: Saad Arif, et al.
Veröffentlicht: (2026)
von: Saad Arif, et al.
Veröffentlicht: (2026)
M5 -- A Diverse Benchmark to Assess the Performance of Large Multimodal Models Across Multilingual and Multicultural Vision-Language Tasks
von: Schneider, Florian, et al.
Veröffentlicht: (2024)
von: Schneider, Florian, et al.
Veröffentlicht: (2024)
TRUCE: Private Benchmarking to Prevent Contamination and Improve Comparative Evaluation of LLMs
von: Rajore, Tanmay, et al.
Veröffentlicht: (2024)
von: Rajore, Tanmay, et al.
Veröffentlicht: (2024)
MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks
von: Ahuja, Sanchit, et al.
Veröffentlicht: (2023)
von: Ahuja, Sanchit, et al.
Veröffentlicht: (2023)
Contamination Report for Multilingual Benchmarks
von: Ahuja, Sanchit, et al.
Veröffentlicht: (2024)
von: Ahuja, Sanchit, et al.
Veröffentlicht: (2024)
Improving Self Consistency in LLMs through Probabilistic Tokenization
von: Sathe, Ashutosh, et al.
Veröffentlicht: (2024)
von: Sathe, Ashutosh, et al.
Veröffentlicht: (2024)
Re: Review of Progress in Early Weight‐Bearing After Distal Femur Fracture Fixation
von: Faisal A. Shaikh, et al.
Veröffentlicht: (2025)
von: Faisal A. Shaikh, et al.
Veröffentlicht: (2025)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
von: Agarwal, Dhruv, et al.
Veröffentlicht: (2025)
von: Agarwal, Dhruv, et al.
Veröffentlicht: (2025)
Text Style Transfer: An Introductory Overview
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2024)
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2024)
IMPACT OF GAIT TRAINING ON LOWER EXTREMITY MOTOR FUNCTION AND BALANCE PERFORMANCE IN STROKE SURVIVORS
von: Hamna Sarfraz,Hafiz Muhammad Waseem Javaid,Nida Razzak
Veröffentlicht: (2026)
von: Hamna Sarfraz,Hafiz Muhammad Waseem Javaid,Nida Razzak
Veröffentlicht: (2026)
Fifth‐Time Recurrence of Dermatofibrosarcoma Protuberans at Distinct Sites: A Rare Case Report
von: Hamna Tariq, et al.
Veröffentlicht: (2025)
von: Hamna Tariq, et al.
Veröffentlicht: (2025)
Locating Edge Domination Number of Some Classes of Claw-Free Cubic Graphs
von: Muhammad Shoaib Sardar, et al.
Veröffentlicht: (2024)
von: Muhammad Shoaib Sardar, et al.
Veröffentlicht: (2024)
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models
von: Sathe, Ashutosh, et al.
Veröffentlicht: (2024)
von: Sathe, Ashutosh, et al.
Veröffentlicht: (2024)
Exploring Pretraining via Active Forgetting for Improving Cross Lingual Transfer for Decoder Language Models
von: Aggarwal, Divyanshu, et al.
Veröffentlicht: (2024)
von: Aggarwal, Divyanshu, et al.
Veröffentlicht: (2024)
Towards Inducing Long-Context Abilities in Multilingual Neural Machine Translation Models
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
Risk, AGAIN - Empirical Examination of the Impact of Age, Gender, Affiliation to a Political Party, Insider Ownership, and Narcissism of Top Executives on Firm Risk Taking
von: Natt, Zainib Ehsan, et al.
Veröffentlicht: (2025)
von: Natt, Zainib Ehsan, et al.
Veröffentlicht: (2025)
STUDY OF BLOOD DEFERRAL CAUSES PRE AND POST DONATION IN A TERTIARY CARE HOSPITAL
von: Hamna Arif,Minza Arif,Sadia Hameed,Fatima Hassan,Sana Saqib
Veröffentlicht: (2025)
von: Hamna Arif,Minza Arif,Sadia Hameed,Fatima Hassan,Sana Saqib
Veröffentlicht: (2025)
A Survey on Self-supervised Contrastive Learning for Multimodal Text-Image Analysis
von: Khan, Asifullah, et al.
Veröffentlicht: (2025)
von: Khan, Asifullah, et al.
Veröffentlicht: (2025)
Traversable Wormholes in f(R) Gravity: Influence of Global Monopole Charge and Energy Conditions
von: Hamna Asad, et al.
Veröffentlicht: (2025)
von: Hamna Asad, et al.
Veröffentlicht: (2025)
Speech Representation Learning Revisited: The Necessity of Separate Learnable Parameters and Robust Data Augmentation
von: Yadav, Hemant, et al.
Veröffentlicht: (2024)
von: Yadav, Hemant, et al.
Veröffentlicht: (2024)
Analysing the Masked predictive coding training criterion for pre-training a Speech Representation Model
von: Yadav, Hemant, et al.
Veröffentlicht: (2023)
von: Yadav, Hemant, et al.
Veröffentlicht: (2023)
MS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked Language Modelling methods for learning Speech Representations
von: Yadav, Hemant, et al.
Veröffentlicht: (2024)
von: Yadav, Hemant, et al.
Veröffentlicht: (2024)
JOOCI: a Framework for Learning Comprehensive Speech Representations
von: Yadav, Hemant, et al.
Veröffentlicht: (2024)
von: Yadav, Hemant, et al.
Veröffentlicht: (2024)
DEPART: DEcomposing PARiTy across Multilingual LLMs
von: Uppadhyay, Manan, et al.
Veröffentlicht: (2026)
von: Uppadhyay, Manan, et al.
Veröffentlicht: (2026)
ARTIFICIAL INTELLIGENCE IN PUBLIC GOVERNANCE: EVALUATING OPPORTUNITIES, RISKS, AND POLICY FRAMEWORKS IN PAKISTAN
von: Momna Khan, Muhammad Imran Farooq, *Usama Salis, Rashid Qutub, Hamna Anis
Veröffentlicht: (2026)
von: Momna Khan, Muhammad Imran Farooq, *Usama Salis, Rashid Qutub, Hamna Anis
Veröffentlicht: (2026)
MAPLE: Multilingual Evaluation of Parameter Efficient Finetuning of Large Language Models
von: Aggarwal, Divyanshu, et al.
Veröffentlicht: (2024)
von: Aggarwal, Divyanshu, et al.
Veröffentlicht: (2024)
Women, Infamous, and Exotic Beings: A Comparative Study of Honorific Usages in Wikipedia and LLMs for Bengali and Hindi
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2025)
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DOSA: A Dataset of Social Artifacts from Different Indian Geographical Subcultures
von: Seth, Agrima, et al.
Veröffentlicht: (2024) -
How Deep Is Representational Bias in LLMs? The Cases of Caste and Religion
von: Seth, Agrima, et al.
Veröffentlicht: (2025) -
METAL: Towards Multilingual Meta-Evaluation
von: Hada, Rishav, et al.
Veröffentlicht: (2024) -
Alzheimer’s Disease and Healthy Aging Indicators Cognitive Decline
von: Hamna Khuld
Veröffentlicht: (2024) -
Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic Prompting
von: Mukherjee, Sagnik, et al.
Veröffentlicht: (2024)