PeruMedQA: Benchmarking Large Language Models (LLMs) on Peruvian Medical Exams -- Dataset Construction and Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Carrillo-Larco, Rodrigo M., Melgarejo, Jesus Lovón, Castillo-Cara, Manuel, Bravo-Rocca, Gusseppe |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Feature Engineering for Agents: An Adaptive Cognitive Architecture for Interpretable ML Monitoring
by: Bravo-Rocca, Gusseppe, et al.
Published: (2025)
by: Bravo-Rocca, Gusseppe, et al.
Published: (2025)
LLMs for energy and macronutrients estimation using only text data from 24-hour dietary recalls: a parameter-efficient fine-tuning experiment using a 10-shot prompt
by: Carrillo-Larco, Rodrigo M
Published: (2025)
by: Carrillo-Larco, Rodrigo M
Published: (2025)
Autoevaluación de habilidades investigativas e intención de dedicarse a la investigación en estudiantes de primer año de medicina de una universidad privada en Lima, Perú
by: Rodrigo M. Carrillo-Larco
Published: (2013)
by: Rodrigo M. Carrillo-Larco
Published: (2013)
MedExpQA: Multilingual Benchmarking of Large Language Models for Medical Question Answering
by: Alonso, Iñigo, et al.
Published: (2024)
by: Alonso, Iñigo, et al.
Published: (2024)
MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs
by: Wei, Jianhui, et al.
Published: (2025)
by: Wei, Jianhui, et al.
Published: (2025)
PRIORIDADES NACIONALES DE INVESTIGACIÓN EN SALUD COMO CATEGORÍAS, EN EL CONGRESO CIENTÍFICO DE ESTUDIANTES DE MEDICINA 2012
by: Rodrigo M. Carrillo-Larco
Published: (2012)
by: Rodrigo M. Carrillo-Larco
Published: (2012)
OPORTUNIDADES DEL CÓDIGO QR PARA DISEMINAR INFORMACIÓN EN SALUD
by: Rodrigo M. Carrillo-Larco
Published: (2013)
by: Rodrigo M. Carrillo-Larco
Published: (2013)
APLICACIÓN ACADÉMICA DE MENSAJES DE TEXTO EN UN CURSO DE PRIMEROS AUXILIOS: ESTUDIO PILOTO EN UNA UNIVERSIDAD PRIVADA DE LIMA, PERÚ
by: Rodrigo M. Carrillo-Larco
Published: (2015)
by: Rodrigo M. Carrillo-Larco
Published: (2015)
EVALUACIÓN DE LA CALIDAD DE INFORMACIÓN SOBRE EL EMBARAZO EN PÁGINAS WEB SEGÚN LAS GUÍAS PERUANAS
by: Rodrigo M. Carrillo-Larco
Published: (2012)
by: Rodrigo M. Carrillo-Larco
Published: (2012)
PRESUPUESTO PARTICIPATIVO, ¿LAS REGIONES MÁS VULNERABLES LO INVIERTEN EN SALUD?
by: Rodrigo M. Carrillo-Larco
Published: (2012)
by: Rodrigo M. Carrillo-Larco
Published: (2012)
Nivel de conocimiento y frecuencia de autoexamen de mama en alumnos de los primeros años de la carrera de Medicina
by: Rodrigo M. Carrillo-Larco
Published: (2015)
by: Rodrigo M. Carrillo-Larco
Published: (2015)
MedConceptsQA: Open Source Medical Concepts QA Benchmark
by: Shoham, Ofir Ben, et al.
Published: (2024)
by: Shoham, Ofir Ben, et al.
Published: (2024)
FairMedQA: Benchmarking Bias in Large Language Models for Medical Question Answering
by: Xiao, Ying, et al.
Published: (2025)
by: Xiao, Ying, et al.
Published: (2025)
AfriMed-QA: A Pan-African, Multi-Specialty, Medical Question-Answering Benchmark Dataset
by: Olatunji, Tobi, et al.
Published: (2024)
by: Olatunji, Tobi, et al.
Published: (2024)
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
by: Zuo, Yuxin, et al.
Published: (2025)
by: Zuo, Yuxin, et al.
Published: (2025)
MedExQA: Medical Question Answering Benchmark with Multiple Explanations
by: Kim, Yunsoo, et al.
Published: (2024)
by: Kim, Yunsoo, et al.
Published: (2024)
EuropeMedQA Study Protocol: A Multilingual, Multimodal Medical Examination Dataset for Language Model Evaluation
by: Causio, Francesco Andrea, et al.
Published: (2026)
by: Causio, Francesco Andrea, et al.
Published: (2026)
JBE-QA: Japanese Bar Exam QA Dataset for Assessing Legal Domain Knowledge
by: Cao, Zhihan, et al.
Published: (2025)
by: Cao, Zhihan, et al.
Published: (2025)
MedFrameQA: A Multi-Image Medical VQA Benchmark for Clinical Reasoning
by: Yu, Suhao, et al.
Published: (2025)
by: Yu, Suhao, et al.
Published: (2025)
ReToP: Learning to Rewrite Electronic Health Records for Clinical Prediction
by: Lovon-Melgarejo, Jesus, et al.
Published: (2026)
by: Lovon-Melgarejo, Jesus, et al.
Published: (2026)
MedAraBench: Large-Scale Arabic Medical Question Answering Dataset and Benchmark
by: Abu-Daoud, Mouath, et al.
Published: (2026)
by: Abu-Daoud, Mouath, et al.
Published: (2026)
MedVision: Dataset and Benchmark for Quantitative Medical Image Analysis
by: Yao, Yongcheng, et al.
Published: (2025)
by: Yao, Yongcheng, et al.
Published: (2025)
MedHal: An Evaluation Dataset for Medical Hallucination Detection
by: Mehenni, Gaya, et al.
Published: (2025)
by: Mehenni, Gaya, et al.
Published: (2025)
LiveMedBench: A Contamination-Free Medical Benchmark for LLMs with Automated Rubric Evaluation
by: Yan, Zhiling, et al.
Published: (2026)
by: Yan, Zhiling, et al.
Published: (2026)
DiversityMedQA: Assessing Demographic Biases in Medical Diagnosis using Large Language Models
by: Rawat, Rajat, et al.
Published: (2024)
by: Rawat, Rajat, et al.
Published: (2024)
Nueva centralidad en interfase urbano-rural (I-UR) Caso: sector Umapalca, zona sur de Arequipa Metropolitana
by: David Jesús Lovon-Caso
Published: (2020)
by: David Jesús Lovon-Caso
Published: (2020)
ThReadMed-QA: A Multi-Turn Medical Dialogue Benchmark from Real Patient Questions
by: Munnangi, Monica, et al.
Published: (2026)
by: Munnangi, Monica, et al.
Published: (2026)
Compound-QA: A Benchmark for Evaluating LLMs on Compound Questions
by: Hou, Yutao, et al.
Published: (2024)
by: Hou, Yutao, et al.
Published: (2024)
¿Quién tiene derecho a opinar sobre política lingüística en Perú? Un análisis crítico del discurso
by: Marco Antonio Lovón-Cueva
Published: (2020)
by: Marco Antonio Lovón-Cueva
Published: (2020)
LLM-MedQA: Enhancing Medical Question Answering through Case Studies in Large Language Models
by: Yang, Hang, et al.
Published: (2024)
by: Yang, Hang, et al.
Published: (2024)
HypoTermQA: Hypothetical Terms Dataset for Benchmarking Hallucination Tendency of LLMs
by: Uluoglakci, Cem, et al.
Published: (2024)
by: Uluoglakci, Cem, et al.
Published: (2024)
OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLM
by: Hu, Yutao, et al.
Published: (2024)
by: Hu, Yutao, et al.
Published: (2024)
ECG-Expert-QA: A Benchmark for Evaluating Medical Large Language Models in Heart Disease Diagnosis
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
La netnografía en la investigación lingüística
by: Marco Lovón
Published: (2025)
by: Marco Lovón
Published: (2025)
La necesidad de decolonialismo lingüístico sobre el subtitulaje en inglés
by: Marco Lovón
Published: (2023)
by: Marco Lovón
Published: (2023)
Análisis lingüístico de voces estereotipadas y raciales en Simón Bolívar: una aproximación1*
by: Marco Lovón
Published: (2023)
by: Marco Lovón
Published: (2023)
MedBioRAG: Semantic Search and Retrieval-Augmented Generation with Large Language Models for Medical and Biological QA
by: Kim, Seonok
Published: (2025)
by: Kim, Seonok
Published: (2025)
MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations
by: Lazzaroni, Ruggero Marino, et al.
Published: (2025)
by: Lazzaroni, Ruggero Marino, et al.
Published: (2025)
VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain
by: Le-Duc, Khai
Published: (2024)
by: Le-Duc, Khai
Published: (2024)
ViMedCSS: A Vietnamese Medical Code-Switching Speech Dataset & Benchmark
by: Nguyen, Tung X., et al.
Published: (2026)
by: Nguyen, Tung X., et al.
Published: (2026)
Similar Items
-
Feature Engineering for Agents: An Adaptive Cognitive Architecture for Interpretable ML Monitoring
by: Bravo-Rocca, Gusseppe, et al.
Published: (2025) -
LLMs for energy and macronutrients estimation using only text data from 24-hour dietary recalls: a parameter-efficient fine-tuning experiment using a 10-shot prompt
by: Carrillo-Larco, Rodrigo M
Published: (2025) -
Autoevaluación de habilidades investigativas e intención de dedicarse a la investigación en estudiantes de primer año de medicina de una universidad privada en Lima, Perú
by: Rodrigo M. Carrillo-Larco
Published: (2013) -
MedExpQA: Multilingual Benchmarking of Large Language Models for Medical Question Answering
by: Alonso, Iñigo, et al.
Published: (2024) -
MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs
by: Wei, Jianhui, et al.
Published: (2025)