Beyond Accuracy: Investigating Error Types in GPT-4 Responses to USMLE Questions
Fuente:
arXiv
Saved in:
| Main Authors: | Roy, Soumyadeep, Khatua, Aparup, Ghoochani, Fatemeh, Hadler, Uwe, Nejdl, Wolfgang, Ganguly, Niloy |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unlocking Efficiency: Adaptive Masking for Gene Transformer Models
by: Roy, Soumyadeep, et al.
Published: (2024)
by: Roy, Soumyadeep, et al.
Published: (2024)
TIGQA:An Expert Annotated Question Answering Dataset in Tigrinya
by: Teklehaymanot, Hailay, et al.
Published: (2024)
by: Teklehaymanot, Hailay, et al.
Published: (2024)
MEDVOC: Vocabulary Adaptation for Fine-tuning Pre-trained Language Models on Medical Text Summarization
by: Balde, Gunjan, et al.
Published: (2024)
by: Balde, Gunjan, et al.
Published: (2024)
Adaptive BPE Tokenization for Enhanced Vocabulary Adaptation in Finetuning Pretrained Language Models
by: Balde, Gunjan, et al.
Published: (2024)
by: Balde, Gunjan, et al.
Published: (2024)
Evaluation of LLMs in Medical Text Summarization: The Role of Vocabulary Adaptation in High OOV Settings
by: Balde, Gunjan, et al.
Published: (2025)
by: Balde, Gunjan, et al.
Published: (2025)
Learning Faster with Better Tokens: Parameter-Efficient Vocabulary Adaptation for Specialized Text Summarization
by: Balde, Gunjan, et al.
Published: (2026)
by: Balde, Gunjan, et al.
Published: (2026)
GPT-4's assessment of its performance in a USMLE-based case study
by: Dhakal, Uttam, et al.
Published: (2024)
by: Dhakal, Uttam, et al.
Published: (2024)
"Where does it hurt?" -- Dataset and Study on Physician Intent Trajectories in Doctor Patient Dialogues
by: Röhr, Tom, et al.
Published: (2025)
by: Röhr, Tom, et al.
Published: (2025)
Digital Diasporas: How Origin Characteristics and Host-Native Distance Shape Immigrants' Online Cultural Retention
by: Khatua, Aparup, et al.
Published: (2025)
by: Khatua, Aparup, et al.
Published: (2025)
LLM Probe: Evaluating LLMs for Low-Resource Languages
by: Teklehaymanot, Hailay Kidu, et al.
Published: (2026)
by: Teklehaymanot, Hailay Kidu, et al.
Published: (2026)
Cost-Performance Optimization for Processing Low-Resource Language Tasks Using Commercial LLMs
by: Nag, Arijit, et al.
Published: (2024)
by: Nag, Arijit, et al.
Published: (2024)
Order-Based Pre-training Strategies for Procedural Text Understanding
by: Nandy, Abhilash, et al.
Published: (2024)
by: Nandy, Abhilash, et al.
Published: (2024)
LGSE: Lexically Grounded Subword Embedding Initialization for Low-Resource Language Adaptation
by: Teklehaymanot, Hailay, et al.
Published: (2026)
by: Teklehaymanot, Hailay, et al.
Published: (2026)
Towards Transparent Stance Detection: A Zero-Shot Approach Using Implicit and Explicit Interpretability
by: Upadhyaya, Apoorva, et al.
Published: (2025)
by: Upadhyaya, Apoorva, et al.
Published: (2025)
Navigating Nuance: In Quest for Political Truth
by: Sar, Soumyadeep, et al.
Published: (2025)
by: Sar, Soumyadeep, et al.
Published: (2025)
Training-Induced Bias Toward LLM-Generated Content in Dense Retrieval
by: Xion, William, et al.
Published: (2026)
by: Xion, William, et al.
Published: (2026)
Cross-Modal Rationale Transfer for Explainable Humanitarian Classification on Social Media
by: Nguyen, Thi Huyen, et al.
Published: (2026)
by: Nguyen, Thi Huyen, et al.
Published: (2026)
A Trio Neural Model for Dynamic Entity Relatedness Ranking
by: Nguyen, Tu, et al.
Published: (2018)
by: Nguyen, Tu, et al.
Published: (2018)
Efficient Continual Pre-training of LLMs for Low-resource Languages
by: Nag, Arijit, et al.
Published: (2024)
by: Nag, Arijit, et al.
Published: (2024)
Effects of Different Prompts on the Quality of GPT-4 Responses to Dementia Care Questions
by: Li, Zhuochun, et al.
Published: (2024)
by: Li, Zhuochun, et al.
Published: (2024)
Program of Thoughts for Financial Reasoning: Leveraging Dynamic In-Context Examples and Generative Retrieval
by: Khatuya, Subhendu, et al.
Published: (2025)
by: Khatuya, Subhendu, et al.
Published: (2025)
FlairGPT: Repurposing LLMs for Interior Designs
by: Littlefair, Gabrielle, et al.
Published: (2025)
by: Littlefair, Gabrielle, et al.
Published: (2025)
How Robust are the Tabular QA Models for Scientific Tables? A Study using Customized Dataset
by: Ghosh, Akash, et al.
Published: (2024)
by: Ghosh, Akash, et al.
Published: (2024)
IDALC: A Semi-Supervised Framework for Intent Detection and Active Learning based Correction
by: Mullick, Ankan, et al.
Published: (2025)
by: Mullick, Ankan, et al.
Published: (2025)
Towards Sustainable NLP: Insights from Benchmarking Inference Energy in Large Language Models
by: Poddar, Soham, et al.
Published: (2025)
by: Poddar, Soham, et al.
Published: (2025)
IVP-VAE: Modeling EHR Time Series with Initial Value Problem Solvers
by: Xiao, Jingge, et al.
Published: (2023)
by: Xiao, Jingge, et al.
Published: (2023)
Tokenization Disparities as Infrastructure Bias: How Subword Systems Create Inequities in LLM Access and Efficiency
by: Teklehaymanot, Hailay Kidu, et al.
Published: (2025)
by: Teklehaymanot, Hailay Kidu, et al.
Published: (2025)
Label-semantics Aware Generative Approach for Domain-Agnostic Multilabel Classification
by: Khatuya, Subhendu, et al.
Published: (2025)
by: Khatuya, Subhendu, et al.
Published: (2025)
Is GPT-4 Less Politically Biased than GPT-3.5? A Renewed Investigation of ChatGPT's Political Biases
by: Weber, Erik, et al.
Published: (2024)
by: Weber, Erik, et al.
Published: (2024)
Understanding the Role of Temperature in Diverse Question Generation by GPT-4
by: Agarwal, Arav, et al.
Published: (2024)
by: Agarwal, Arav, et al.
Published: (2024)
Adversarial Question Answering Robustness: A Multi-Level Error Analysis and Mitigation Study
by: Choudhury, Agniv Roy, et al.
Published: (2026)
by: Choudhury, Agniv Roy, et al.
Published: (2026)
GPT-5 Model Corrected GPT-4V's Chart Reading Errors, Not Prompting
by: Yang, Kaichun, et al.
Published: (2025)
by: Yang, Kaichun, et al.
Published: (2025)
GPT-Fathom: Benchmarking Large Language Models to Decipher the Evolutionary Path towards GPT-4 and Beyond
by: Zheng, Shen, et al.
Published: (2023)
by: Zheng, Shen, et al.
Published: (2023)
Brevity is the soul of sustainability: Characterizing LLM response lengths
by: Poddar, Soham, et al.
Published: (2025)
by: Poddar, Soham, et al.
Published: (2025)
Instruction-Guided Bullet Point Summarization of Long Financial Earnings Call Transcripts
by: Khatuya, Subhendu, et al.
Published: (2024)
by: Khatuya, Subhendu, et al.
Published: (2024)
Knowledge-Augmented Question Error Correction for Chinese Question Answer System with QuestionRAG
by: Qiu, Longpeng, et al.
Published: (2025)
by: Qiu, Longpeng, et al.
Published: (2025)
A Comprehensive Evaluation of GPT-4V on Knowledge-Intensive Visual Question Answering
by: Li, Yunxin, et al.
Published: (2023)
by: Li, Yunxin, et al.
Published: (2023)
ChatGPT as a Math Questioner? Evaluating ChatGPT on Generating Pre-university Math Questions
by: Van Long, Phuoc Pham, et al.
Published: (2023)
by: Van Long, Phuoc Pham, et al.
Published: (2023)
Beyond Accuracy: Rethinking Hallucination and Regulatory Response in Generative AI
by: Li, Zihao, et al.
Published: (2025)
by: Li, Zihao, et al.
Published: (2025)
Framing Analysis of Health-Related Narratives: Conspiracy versus Mainstream Media
by: Reiter-Haas, Markus, et al.
Published: (2024)
by: Reiter-Haas, Markus, et al.
Published: (2024)
Similar Items
-
Unlocking Efficiency: Adaptive Masking for Gene Transformer Models
by: Roy, Soumyadeep, et al.
Published: (2024) -
TIGQA:An Expert Annotated Question Answering Dataset in Tigrinya
by: Teklehaymanot, Hailay, et al.
Published: (2024) -
MEDVOC: Vocabulary Adaptation for Fine-tuning Pre-trained Language Models on Medical Text Summarization
by: Balde, Gunjan, et al.
Published: (2024) -
Adaptive BPE Tokenization for Enhanced Vocabulary Adaptation in Finetuning Pretrained Language Models
by: Balde, Gunjan, et al.
Published: (2024) -
Evaluation of LLMs in Medical Text Summarization: The Role of Vocabulary Adaptation in High OOV Settings
by: Balde, Gunjan, et al.
Published: (2025)