When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Faisal, Faizan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LEGAL-UQA: A Low-Resource Urdu-English Dataset for Legal Question Answering
von: Faisal, Faizan, et al.
Veröffentlicht: (2024)
von: Faisal, Faizan, et al.
Veröffentlicht: (2024)
Frontier LLMs Still Struggle with Simple Reasoning Tasks
von: Malek, Alan, et al.
Veröffentlicht: (2025)
von: Malek, Alan, et al.
Veröffentlicht: (2025)
Skin-SOAP: A Weakly Supervised Framework for Generating Structured SOAP Notes
von: Kamal, Sadia, et al.
Veröffentlicht: (2025)
von: Kamal, Sadia, et al.
Veröffentlicht: (2025)
Towards Adapting Open-Source Large Language Models for Expert-Level Clinical Note Generation
von: Wang, Hanyin, et al.
Veröffentlicht: (2024)
von: Wang, Hanyin, et al.
Veröffentlicht: (2024)
CNSight: Evaluation of Clinical Note Segmentation Tools
von: Surana, Risha, et al.
Veröffentlicht: (2025)
von: Surana, Risha, et al.
Veröffentlicht: (2025)
Reasoning's Razor: Reasoning Improves Accuracy but Can Hurt Recall at Critical Operating Points in Safety and Hallucination Detection
von: Chegini, Atoosa, et al.
Veröffentlicht: (2025)
von: Chegini, Atoosa, et al.
Veröffentlicht: (2025)
Learning to Reason at the Frontier of Learnability
von: Foster, Thomas, et al.
Veröffentlicht: (2025)
von: Foster, Thomas, et al.
Veröffentlicht: (2025)
When Shared Knowledge Hurts: Spectral Over-Accumulation in Model Merging
von: Li, Yayuan, et al.
Veröffentlicht: (2026)
von: Li, Yayuan, et al.
Veröffentlicht: (2026)
When Models Know More Than They Say: Probing Analogical Reasoning in LLMs
von: McGovern, Hope, et al.
Veröffentlicht: (2026)
von: McGovern, Hope, et al.
Veröffentlicht: (2026)
When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning
von: Singhi, Nishad, et al.
Veröffentlicht: (2025)
von: Singhi, Nishad, et al.
Veröffentlicht: (2025)
When Do LLMs Reason? A Dynamical Systems View via Entropy Phase Transitions
von: Xia, Wei, et al.
Veröffentlicht: (2026)
von: Xia, Wei, et al.
Veröffentlicht: (2026)
Synthetic Patient-Physician Dialogue Generation from Clinical Notes Using LLM
von: Das, Trisha, et al.
Veröffentlicht: (2024)
von: Das, Trisha, et al.
Veröffentlicht: (2024)
Rewarding Graph Reasoning Process makes LLMs more Generalized Reasoners
von: Peng, Miao, et al.
Veröffentlicht: (2025)
von: Peng, Miao, et al.
Veröffentlicht: (2025)
Multitask Mayhem: Unveiling and Mitigating Safety Gaps in LLMs Fine-tuning
von: Jan, Essa, et al.
Veröffentlicht: (2024)
von: Jan, Essa, et al.
Veröffentlicht: (2024)
Dementia-R1: Reinforced Pretraining and Reasoning from Unstructured Clinical Notes for Real-World Dementia Prognosis
von: Kim, Choonghan, et al.
Veröffentlicht: (2026)
von: Kim, Choonghan, et al.
Veröffentlicht: (2026)
Where Norms and References Collide: Evaluating LLMs on Normative Reasoning
von: Abrams, Mitchell, et al.
Veröffentlicht: (2026)
von: Abrams, Mitchell, et al.
Veröffentlicht: (2026)
ECG-Reasoning-Benchmark: A Benchmark for Evaluating Clinical Reasoning Capabilities in ECG Interpretation
von: Oh, Jungwoo, et al.
Veröffentlicht: (2026)
von: Oh, Jungwoo, et al.
Veröffentlicht: (2026)
The Growing Pains of Frontier Models: When Leaderboards Stop Separating and What to Measure Next
von: Amin, Adil
Veröffentlicht: (2026)
von: Amin, Adil
Veröffentlicht: (2026)
When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure
von: Xiao, Boyu, et al.
Veröffentlicht: (2026)
von: Xiao, Boyu, et al.
Veröffentlicht: (2026)
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
von: Park, Jungsoo, et al.
Veröffentlicht: (2025)
von: Park, Jungsoo, et al.
Veröffentlicht: (2025)
Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents
von: Turk, Matt
Veröffentlicht: (2026)
von: Turk, Matt
Veröffentlicht: (2026)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
von: Fan, Run-Ze, et al.
Veröffentlicht: (2025)
von: Fan, Run-Ze, et al.
Veröffentlicht: (2025)
Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs
von: Wang, Yinong Oliver, et al.
Veröffentlicht: (2025)
von: Wang, Yinong Oliver, et al.
Veröffentlicht: (2025)
Towards Scalable SOAP Note Generation: A Weakly Supervised Multimodal Framework
von: Kamal, Sadia, et al.
Veröffentlicht: (2025)
von: Kamal, Sadia, et al.
Veröffentlicht: (2025)
GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations
von: Duan, Jinhao, et al.
Veröffentlicht: (2024)
von: Duan, Jinhao, et al.
Veröffentlicht: (2024)
Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning
von: Singh, Joykirat, et al.
Veröffentlicht: (2024)
von: Singh, Joykirat, et al.
Veröffentlicht: (2024)
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
When Should LLMs Be Less Specific? Selective Abstraction for Reliable Long-Form Text Generation
von: Goren, Shani, et al.
Veröffentlicht: (2026)
von: Goren, Shani, et al.
Veröffentlicht: (2026)
Generation and De-Identification of Indian Clinical Discharge Summaries using LLMs
von: Singh, Sanjeet, et al.
Veröffentlicht: (2024)
von: Singh, Sanjeet, et al.
Veröffentlicht: (2024)
Reasoning-Driven Synthetic Data Generation and Evaluation
von: Davidson, Tim R., et al.
Veröffentlicht: (2026)
von: Davidson, Tim R., et al.
Veröffentlicht: (2026)
SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions
von: Suvarna, Ashima, et al.
Veröffentlicht: (2026)
von: Suvarna, Ashima, et al.
Veröffentlicht: (2026)
Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning
von: Yu, Erxin, et al.
Veröffentlicht: (2025)
von: Yu, Erxin, et al.
Veröffentlicht: (2025)
Evaluating Social Bias in RAG Systems: When External Context Helps and Reasoning Hurts
von: Parihar, Shweta, et al.
Veröffentlicht: (2026)
von: Parihar, Shweta, et al.
Veröffentlicht: (2026)
What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
von: Lv, Keyu, et al.
Veröffentlicht: (2026)
von: Lv, Keyu, et al.
Veröffentlicht: (2026)
Self-Training Meets Consistency: Improving LLMs' Reasoning with Consistency-Driven Rationale Evaluation
von: Lee, Jaehyeok, et al.
Veröffentlicht: (2024)
von: Lee, Jaehyeok, et al.
Veröffentlicht: (2024)
Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking
von: Xu, Qinwu, et al.
Veröffentlicht: (2026)
von: Xu, Qinwu, et al.
Veröffentlicht: (2026)
CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics
von: Nagar, Aishik, et al.
Veröffentlicht: (2026)
von: Nagar, Aishik, et al.
Veröffentlicht: (2026)
When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
von: Wang, Kai, et al.
Veröffentlicht: (2025)
von: Wang, Kai, et al.
Veröffentlicht: (2025)
GEAR: A General Evaluation Framework for Abductive Reasoning
von: He, Kaiyu, et al.
Veröffentlicht: (2025)
von: He, Kaiyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LEGAL-UQA: A Low-Resource Urdu-English Dataset for Legal Question Answering
von: Faisal, Faizan, et al.
Veröffentlicht: (2024) -
Frontier LLMs Still Struggle with Simple Reasoning Tasks
von: Malek, Alan, et al.
Veröffentlicht: (2025) -
Skin-SOAP: A Weakly Supervised Framework for Generating Structured SOAP Notes
von: Kamal, Sadia, et al.
Veröffentlicht: (2025) -
Towards Adapting Open-Source Large Language Models for Expert-Level Clinical Note Generation
von: Wang, Hanyin, et al.
Veröffentlicht: (2024) -
CNSight: Evaluation of Clinical Note Segmentation Tools
von: Surana, Risha, et al.
Veröffentlicht: (2025)