MedArabiQ: Benchmarking Large Language Models on Arabic Medical Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Daoud, Mouath Abu, Abouzahir, Chaimae, Kharouf, Leen, Al-Eisawi, Walid, Habash, Nizar, Shamout, Farah E. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AraHealthQA 2025: The First Shared Task on Arabic Health Question Answering
by: Alhuzali, Hassan, et al.
Published: (2025)
by: Alhuzali, Hassan, et al.
Published: (2025)
Cross-Lingual Empirical Evaluation of Large Language Models for Arabic Medical Tasks
by: Abouzahir, Chaimae, et al.
Published: (2026)
by: Abouzahir, Chaimae, et al.
Published: (2026)
MedAraBench: Large-Scale Arabic Medical Question Answering Dataset and Benchmark
by: Abu-Daoud, Mouath, et al.
Published: (2026)
by: Abu-Daoud, Mouath, et al.
Published: (2026)
Desk2Desk: Optimization-based Mixed Reality Workspace Integration for Remote Side-by-side Collaboration
by: Sidenmark, Ludwig, et al.
Published: (2024)
by: Sidenmark, Ludwig, et al.
Published: (2024)
Measuring Large Language Models Dependency: Validating the Arabic Version of the LLM-D12 Scale
by: AlShakhsi, Sameha, et al.
Published: (2025)
by: AlShakhsi, Sameha, et al.
Published: (2025)
K-QA: A Real-World Medical Q&A Benchmark
by: Manes, Itay, et al.
Published: (2024)
by: Manes, Itay, et al.
Published: (2024)
Making It Work Is the Work: Engineering Maturity as Epistemic Work
by: Leen, Danny, et al.
Published: (2026)
by: Leen, Danny, et al.
Published: (2026)
Design of a visual environment for programming by direct data manipulation
by: Adam, Michel, et al.
Published: (2025)
by: Adam, Michel, et al.
Published: (2025)
Arabic Little STT: Arabic Children Speech Recognition Dataset
by: Alkadri, Mouhand, et al.
Published: (2025)
by: Alkadri, Mouhand, et al.
Published: (2025)
Developing and Validating the Arabic Version of the Attitudes Toward Large Language Models Scale
by: Barajeeh, Basad, et al.
Published: (2025)
by: Barajeeh, Basad, et al.
Published: (2025)
MedBike: A Cardiac Patient Monitoring System Enhanced through Gamification
by: Hossain, Tahmim, et al.
Published: (2024)
by: Hossain, Tahmim, et al.
Published: (2024)
Combining Automation and Expertise: A Semi-automated Approach to Correcting Eye Tracking Data in Reading Tasks
by: Madi, Naser Al, et al.
Published: (2025)
by: Madi, Naser Al, et al.
Published: (2025)
SyriSign: A Parallel Corpus for Arabic Text to Syrian Arabic Sign Language Translation
by: Khalil, Mohammad Amer, et al.
Published: (2026)
by: Khalil, Mohammad Amer, et al.
Published: (2026)
MedFoundationHub: A Lightweight and Secure Toolkit for Deploying Medical Vision Language Foundation Models
by: Li, Xiao, et al.
Published: (2025)
by: Li, Xiao, et al.
Published: (2025)
ArEEG_Chars: Dataset for Envisioned Speech Recognition using EEG for Arabic Characters
by: Darwish, Hazem, et al.
Published: (2024)
by: Darwish, Hazem, et al.
Published: (2024)
ArEEG_Words: Dataset for Envisioned Speech Recognition using EEG for Arabic Words
by: Darwish, Hazem, et al.
Published: (2024)
by: Darwish, Hazem, et al.
Published: (2024)
Adjusting B‐Tree for Better Usability: Nodes as Files Instead of Disk Blocks
by: Majed AbuSafiya
Published: (2026)
by: Majed AbuSafiya
Published: (2026)
MedDialogRubrics: A Comprehensive Benchmark and Evaluation Framework for Multi-turn Medical Consultations in Large Language Models
by: Gong, Lecheng, et al.
Published: (2026)
by: Gong, Lecheng, et al.
Published: (2026)
Experimental Interface for Multimodal and Large Language Model Based Explanations of Educational Recommender Systems
by: Abu-Rasheed, Hasan, et al.
Published: (2024)
by: Abu-Rasheed, Hasan, et al.
Published: (2024)
Bias Beneath the Tone: Empirical Characterisation of Tone Bias in LLM-Driven UX Systems
by: Bodara, Heet, et al.
Published: (2025)
by: Bodara, Heet, et al.
Published: (2025)
Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education
by: Snoubara, Abdul Aziz, et al.
Published: (2026)
by: Snoubara, Abdul Aziz, et al.
Published: (2026)
Investigating the Task Load of Investigating the Task Load in Visualization Studies
by: Pahr, Daniel, et al.
Published: (2025)
by: Pahr, Daniel, et al.
Published: (2025)
Lemmatization as a Classification Task: Results from Arabic across Multiple Genres
by: Saeed, Mostafa, et al.
Published: (2025)
by: Saeed, Mostafa, et al.
Published: (2025)
Prompt2Task: Automating UI Tasks on Smartphones from Textual Prompts
by: Huang, Tian, et al.
Published: (2024)
by: Huang, Tian, et al.
Published: (2024)
Task Mode: Dynamic Filtering for Task-Specific Web Navigation using LLMs
by: Mohanbabu, Ananya Gubbi, et al.
Published: (2025)
by: Mohanbabu, Ananya Gubbi, et al.
Published: (2025)
Does GenAI Make Usability Testing Obsolete?
by: Pourasad, Ali Ebrahimi, et al.
Published: (2024)
by: Pourasad, Ali Ebrahimi, et al.
Published: (2024)
RFM-HRI : A Multimodal Dataset of Medical Robot Failure, User Reaction and Recovery Preferences for Item Retrieval Tasks
by: Batra, Yashika, et al.
Published: (2026)
by: Batra, Yashika, et al.
Published: (2026)
TaskLens: Generating Task-Conditioned Scaffolded Interfaces for Learning Professional Creative Software
by: Liu, Yimeng, et al.
Published: (2025)
by: Liu, Yimeng, et al.
Published: (2025)
CPS-TaskForge: Generating Collaborative Problem Solving Environments for Diverse Communication Tasks
by: Haduong, Nikita, et al.
Published: (2024)
by: Haduong, Nikita, et al.
Published: (2024)
Exploring diversity perceptions in a community through a Q&A chatbot
by: Kun, Peter, et al.
Published: (2024)
by: Kun, Peter, et al.
Published: (2024)
VisionTasker: Mobile Task Automation Using Vision Based UI Understanding and LLM Task Planning
by: Song, Yunpeng, et al.
Published: (2023)
by: Song, Yunpeng, et al.
Published: (2023)
From Following to Understanding: Investigating the Role of Reflective Prompts in AR-Guided Tasks to Promote Task Understanding
by: Zhang, Nandi, et al.
Published: (2025)
by: Zhang, Nandi, et al.
Published: (2025)
Expert-Generated Privacy Q&A Dataset for Conversational AI and User Study Insights
by: Leschanowsky, Anna, et al.
Published: (2025)
by: Leschanowsky, Anna, et al.
Published: (2025)
Management and Visualization Tools for Emergency Medical Services
by: Guigues, Vincent, et al.
Published: (2024)
by: Guigues, Vincent, et al.
Published: (2024)
PyZoBot: A Platform for Conversational Information Extraction and Synthesis from Curated Zotero Reference Libraries through Advanced Retrieval-Augmented Generation
by: Alshammari, Suad, et al.
Published: (2024)
by: Alshammari, Suad, et al.
Published: (2024)
ErgoGlide: A Wearable Trackball Device for Ergonomic Text Entry in Virtual Reality
by: Bakar, Muhammad Abu, et al.
Published: (2026)
by: Bakar, Muhammad Abu, et al.
Published: (2026)
ContextQ: Generated Questions to Support Meaningful Parent-Child Dialogue While Co-Reading
by: Smith, Griffin Dietz, et al.
Published: (2024)
by: Smith, Griffin Dietz, et al.
Published: (2024)
Task-Aware Delegation Cues for LLM Agents
by: Gu, Xingrui
Published: (2026)
by: Gu, Xingrui
Published: (2026)
A Typology of Decision-Making Tasks for Visualization
by: Brumar, Camelia D., et al.
Published: (2024)
by: Brumar, Camelia D., et al.
Published: (2024)
Haptic VR Simulation for Surgery Procedures in Medical Training
by: Jie, Lim Zheng, et al.
Published: (2024)
by: Jie, Lim Zheng, et al.
Published: (2024)
Similar Items
-
AraHealthQA 2025: The First Shared Task on Arabic Health Question Answering
by: Alhuzali, Hassan, et al.
Published: (2025) -
Cross-Lingual Empirical Evaluation of Large Language Models for Arabic Medical Tasks
by: Abouzahir, Chaimae, et al.
Published: (2026) -
MedAraBench: Large-Scale Arabic Medical Question Answering Dataset and Benchmark
by: Abu-Daoud, Mouath, et al.
Published: (2026) -
Desk2Desk: Optimization-based Mixed Reality Workspace Integration for Remote Side-by-side Collaboration
by: Sidenmark, Ludwig, et al.
Published: (2024) -
Measuring Large Language Models Dependency: Validating the Arabic Version of the LLM-D12 Scale
by: AlShakhsi, Sameha, et al.
Published: (2025)