HEALTH-PARIKSHA: Assessing RAG Models for Health Chatbots in Real-World Multilingual Settings
Fuente:
arXiv
Saved in:
| Main Authors: | Gumma, Varun, Raghunath, Ananditha, Jain, Mohit, Sitaram, Sunayana |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural Data
by: Watts, Ishaan, et al.
Published: (2024)
by: Watts, Ishaan, et al.
Published: (2024)
Contamination Report for Multilingual Benchmarks
by: Ahuja, Sanchit, et al.
Published: (2024)
by: Ahuja, Sanchit, et al.
Published: (2024)
METAL: Towards Multilingual Meta-Evaluation
by: Hada, Rishav, et al.
Published: (2024)
by: Hada, Rishav, et al.
Published: (2024)
MAFIA: Multi-Adapter Fused Inclusive LanguAge Models
by: Jain, Prachi, et al.
Published: (2024)
by: Jain, Prachi, et al.
Published: (2024)
Designing Medical Chatbots where Accuracy and Acceptability are in Conflict: An Exploratory, Vignette-based Study in Urban India
by: Raghunath, Ananditha, et al.
Published: (2026)
by: Raghunath, Ananditha, et al.
Published: (2026)
M5 -- A Diverse Benchmark to Assess the Performance of Large Multimodal Models Across Multilingual and Multicultural Vision-Language Tasks
by: Schneider, Florian, et al.
Published: (2024)
by: Schneider, Florian, et al.
Published: (2024)
Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation?
by: Hada, Rishav, et al.
Published: (2023)
by: Hada, Rishav, et al.
Published: (2023)
Beyond Metrics: Evaluating LLMs' Effectiveness in Culturally Nuanced, Low-Resource Real-World Scenarios
by: Ochieng, Millicent, et al.
Published: (2024)
by: Ochieng, Millicent, et al.
Published: (2024)
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models
by: Sathe, Ashutosh, et al.
Published: (2024)
by: Sathe, Ashutosh, et al.
Published: (2024)
UPDESH: Synthesizing Grounded Instruction Tuning Data for 13 Indic Languages
by: Chitale, Pranjal A., et al.
Published: (2025)
by: Chitale, Pranjal A., et al.
Published: (2025)
MAPLE: Multilingual Evaluation of Parameter Efficient Finetuning of Large Language Models
by: Aggarwal, Divyanshu, et al.
Published: (2024)
by: Aggarwal, Divyanshu, et al.
Published: (2024)
Towards Inducing Long-Context Abilities in Multilingual Neural Machine Translation Models
by: Gumma, Varun, et al.
Published: (2024)
by: Gumma, Varun, et al.
Published: (2024)
MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks
by: Ahuja, Sanchit, et al.
Published: (2023)
by: Ahuja, Sanchit, et al.
Published: (2023)
Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
by: Hamna, Hamna, et al.
Published: (2025)
by: Hamna, Hamna, et al.
Published: (2025)
Exploring Pretraining via Active Forgetting for Improving Cross Lingual Transfer for Decoder Language Models
by: Aggarwal, Divyanshu, et al.
Published: (2024)
by: Aggarwal, Divyanshu, et al.
Published: (2024)
MS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked Language Modelling methods for learning Speech Representations
by: Yadav, Hemant, et al.
Published: (2024)
by: Yadav, Hemant, et al.
Published: (2024)
DEPART: DEcomposing PARiTy across Multilingual LLMs
by: Uppadhyay, Manan, et al.
Published: (2026)
by: Uppadhyay, Manan, et al.
Published: (2026)
Improving Self Consistency in LLMs through Probabilistic Tokenization
by: Sathe, Ashutosh, et al.
Published: (2024)
by: Sathe, Ashutosh, et al.
Published: (2024)
Speech Representation Learning Revisited: The Necessity of Separate Learnable Parameters and Robust Data Augmentation
by: Yadav, Hemant, et al.
Published: (2024)
by: Yadav, Hemant, et al.
Published: (2024)
JOOCI: a Framework for Learning Comprehensive Speech Representations
by: Yadav, Hemant, et al.
Published: (2024)
by: Yadav, Hemant, et al.
Published: (2024)
A Multilingual, Culture-First Approach to Addressing Misgendering in LLM Applications
by: Sitaram, Sunayana, et al.
Published: (2025)
by: Sitaram, Sunayana, et al.
Published: (2025)
DOSA: A Dataset of Social Artifacts from Different Indian Geographical Subcultures
by: Seth, Agrima, et al.
Published: (2024)
by: Seth, Agrima, et al.
Published: (2024)
Exploring Continual Fine-Tuning for Enhancing Language Ability in Large Language Model
by: Aggarwal, Divyanshu, et al.
Published: (2024)
by: Aggarwal, Divyanshu, et al.
Published: (2024)
Teaching LLMs to Abstain across Languages via Multilingual Feedback
by: Feng, Shangbin, et al.
Published: (2024)
by: Feng, Shangbin, et al.
Published: (2024)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
by: Agarwal, Dhruv, et al.
Published: (2025)
by: Agarwal, Dhruv, et al.
Published: (2025)
Bridging the Language Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
by: Kumar, Somnath, et al.
Published: (2023)
by: Kumar, Somnath, et al.
Published: (2023)
CultureLLM: Incorporating Cultural Differences into Large Language Models
by: Li, Cheng, et al.
Published: (2024)
by: Li, Cheng, et al.
Published: (2024)
Bridging the Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
by: Kumar, Somnath, et al.
Published: (2024)
by: Kumar, Somnath, et al.
Published: (2024)
How Deep Is Representational Bias in LLMs? The Cases of Caste and Religion
by: Seth, Agrima, et al.
Published: (2025)
by: Seth, Agrima, et al.
Published: (2025)
Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic Prompting
by: Mukherjee, Sagnik, et al.
Published: (2024)
by: Mukherjee, Sagnik, et al.
Published: (2024)
sPhinX: Sample Efficient Multilingual Instruction Fine-Tuning Through N-shot Guided Prompting
by: Ahuja, Sanchit, et al.
Published: (2024)
by: Ahuja, Sanchit, et al.
Published: (2024)
Reinforcement Learning for Optimizing RAG for Domain Chatbots
by: Kulkarni, Mandar, et al.
Published: (2024)
by: Kulkarni, Mandar, et al.
Published: (2024)
TRUCE: Private Benchmarking to Prevent Contamination and Improve Comparative Evaluation of LLMs
by: Rajore, Tanmay, et al.
Published: (2024)
by: Rajore, Tanmay, et al.
Published: (2024)
Addressing Hallucinations with RAG and NMISS in Italian Healthcare LLM Chatbots
by: Priola, Maria Paola
Published: (2024)
by: Priola, Maria Paola
Published: (2024)
Diversity Enhances an LLM's Performance in RAG and Long-context Task
by: Wang, Zhichao, et al.
Published: (2025)
by: Wang, Zhichao, et al.
Published: (2025)
Exploring Robustness of Multilingual LLMs on Real-World Noisy Data
by: Aliakbarzadeh, Amirhossein, et al.
Published: (2025)
by: Aliakbarzadeh, Amirhossein, et al.
Published: (2025)
Investigating Language Preference of Multilingual RAG Systems
by: Park, Jeonghyun, et al.
Published: (2025)
by: Park, Jeonghyun, et al.
Published: (2025)
Designing with Culture: How Social Norms Shape Trust and Preference in Health Chatbots
by: Wadhwa, Arpita, et al.
Published: (2025)
by: Wadhwa, Arpita, et al.
Published: (2025)
MunTTS: A Text-to-Speech System for Mundari
by: Gumma, Varun, et al.
Published: (2024)
by: Gumma, Varun, et al.
Published: (2024)
Which Institutional Frameworks Do Chatbots Assume? Auditing Jurisdictional Defaults in Multilingual LLMs
by: Wang, Zhizhi, et al.
Published: (2026)
by: Wang, Zhizhi, et al.
Published: (2026)
Similar Items
-
PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural Data
by: Watts, Ishaan, et al.
Published: (2024) -
Contamination Report for Multilingual Benchmarks
by: Ahuja, Sanchit, et al.
Published: (2024) -
METAL: Towards Multilingual Meta-Evaluation
by: Hada, Rishav, et al.
Published: (2024) -
MAFIA: Multi-Adapter Fused Inclusive LanguAge Models
by: Jain, Prachi, et al.
Published: (2024) -
Designing Medical Chatbots where Accuracy and Acceptability are in Conflict: An Exploratory, Vignette-based Study in Urban India
by: Raghunath, Ananditha, et al.
Published: (2026)