Reinforcement Learning for Out-of-Distribution Reasoning in LLMs: An Empirical Study on Diagnosis-Related Group Coding
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Hanyin, Wu, Zhenbang, Kolar, Gururaj, Korsapati, Hariprasad, Bartlett, Brian, Hull, Bryan, Sun, Jimeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning
di: Lin, Jiacheng, et al.
Pubblicazione: (2025)
di: Lin, Jiacheng, et al.
Pubblicazione: (2025)
Process-Supervised Reward Models for Verifying Clinical Note Generation: A Scalable Approach Guided by Domain Expertise
di: Wang, Hanyin, et al.
Pubblicazione: (2024)
di: Wang, Hanyin, et al.
Pubblicazione: (2024)
Towards Adapting Open-Source Large Language Models for Expert-Level Clinical Note Generation
di: Wang, Hanyin, et al.
Pubblicazione: (2024)
di: Wang, Hanyin, et al.
Pubblicazione: (2024)
Bridging the Reproducibility Divide: Open Source Software's Role in Standardizing Healthcare AI
di: Wu, John, et al.
Pubblicazione: (2026)
di: Wu, John, et al.
Pubblicazione: (2026)
Beyond Label Attention: Transparency in Language Models for Automated Medical Coding via Dictionary Learning
di: Wu, John, et al.
Pubblicazione: (2024)
di: Wu, John, et al.
Pubblicazione: (2024)
An Empirical Study of Reasoning Steps in Thinking Code LLMs
di: Xue, Haoran, et al.
Pubblicazione: (2025)
di: Xue, Haoran, et al.
Pubblicazione: (2025)
Social Determinants of Health Prediction for ICD-9 Code with Reasoning Models
di: Khan, Sharim, et al.
Pubblicazione: (2025)
di: Khan, Sharim, et al.
Pubblicazione: (2025)
DILA: Dictionary Label Attention for Mechanistic Interpretability in High-dimensional Multi-label Medical Coding Prediction
di: Wu, John, et al.
Pubblicazione: (2024)
di: Wu, John, et al.
Pubblicazione: (2024)
Multi-Step Deductive Reasoning Over Natural Language: An Empirical Study on Out-of-Distribution Generalisation
di: Bao, Qiming, et al.
Pubblicazione: (2022)
di: Bao, Qiming, et al.
Pubblicazione: (2022)
MANet: A deep learning for object detection
di: Wu, Zhenbang
Pubblicazione: (2026)
di: Wu, Zhenbang
Pubblicazione: (2026)
Convex Holder bound and its applications
di: M, Hariprasad
Pubblicazione: (2025)
di: M, Hariprasad
Pubblicazione: (2025)
Recursive eigen extrusion: Expanding eigenbasis conjecture
di: Hariprasad, M
Pubblicazione: (2019)
di: Hariprasad, M
Pubblicazione: (2019)
Prompt Sensitivity and Answer Consistency of Small Open-Source Language Models for Clinical Question Answering in Low-Resource Healthcare
di: Hariprasad, Shravani
Pubblicazione: (2026)
di: Hariprasad, Shravani
Pubblicazione: (2026)
Sparse Regression Codes for Secret Key Agreement: Achieving Strong Secrecy and Near-Optimal Rates for Gaussian Sources
di: Athanasakos, Emmanouil M., et al.
Pubblicazione: (2025)
di: Athanasakos, Emmanouil M., et al.
Pubblicazione: (2025)
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
di: Chen, Yang, et al.
Pubblicazione: (2025)
di: Chen, Yang, et al.
Pubblicazione: (2025)
On Formally Undecidable Propositions of Nondeterministic Complexity and Related Classes
di: Kolář, Martin
Pubblicazione: (2026)
di: Kolář, Martin
Pubblicazione: (2026)
On the Empirical Complexity of Reasoning and Planning in LLMs
di: Kang, Liwei, et al.
Pubblicazione: (2024)
di: Kang, Liwei, et al.
Pubblicazione: (2024)
On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks
di: Gupta, Aarav, et al.
Pubblicazione: (2026)
di: Gupta, Aarav, et al.
Pubblicazione: (2026)
Test-Time Adaptation for EEG Foundation Models: A Systematic Study under Real-World Distribution Shifts
di: Lee, Gabriel Jason, et al.
Pubblicazione: (2026)
di: Lee, Gabriel Jason, et al.
Pubblicazione: (2026)
Towards Physiologically Sensible Predictions via the Rule-based Reinforcement Learning Layer
di: Zhu, Lingwei, et al.
Pubblicazione: (2025)
di: Zhu, Lingwei, et al.
Pubblicazione: (2025)
The conjugacy problem in Out(Fm) when the polynomial restrictions are non-growing
di: Bartlett, Gabriel
Pubblicazione: (2025)
di: Bartlett, Gabriel
Pubblicazione: (2025)
Utilizing Training Data to Improve LLM Reasoning for Tabular Understanding
di: Gao, Chufan, et al.
Pubblicazione: (2025)
di: Gao, Chufan, et al.
Pubblicazione: (2025)
TTM-RE: Memory-Augmented Document-Level Relation Extraction
di: Gao, Chufan, et al.
Pubblicazione: (2024)
di: Gao, Chufan, et al.
Pubblicazione: (2024)
LLMs are Bug Replicators: An Empirical Study on LLMs' Capability in Completing Bug-prone Code
di: Guo, Liwei, et al.
Pubblicazione: (2025)
di: Guo, Liwei, et al.
Pubblicazione: (2025)
Scouring Parrondo's Paradox in Discrete-Time Quantum Walks
di: Kadiri, Gururaj
Pubblicazione: (2024)
di: Kadiri, Gururaj
Pubblicazione: (2024)
Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning
di: Panaganti, Kishan, et al.
Pubblicazione: (2026)
di: Panaganti, Kishan, et al.
Pubblicazione: (2026)
An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
di: Jin, Bowen, et al.
Pubblicazione: (2025)
di: Jin, Bowen, et al.
Pubblicazione: (2025)
Open the Oyster: Empirical Evaluation and Improvement of Code Reasoning Confidence in LLMs
di: Wang, Shufan, et al.
Pubblicazione: (2025)
di: Wang, Shufan, et al.
Pubblicazione: (2025)
Enhancing the Code Reasoning Capabilities of LLMs via Consistency-based Reinforcement Learning
di: Qin, Zhanyue, et al.
Pubblicazione: (2026)
di: Qin, Zhanyue, et al.
Pubblicazione: (2026)
Effective Learning for Small Reasoning Models: An Empirical Study on 0.5B Reasoning LLMs
di: Zhuang, Xialie, et al.
Pubblicazione: (2025)
di: Zhuang, Xialie, et al.
Pubblicazione: (2025)
CodeReasoner: Enhancing the Code Reasoning Ability with Reinforcement Learning
di: Tang, Lingxiao, et al.
Pubblicazione: (2025)
di: Tang, Lingxiao, et al.
Pubblicazione: (2025)
How Execution Features Relate to Failures: An Empirical Study and Diagnosis Approach
di: Smytzek, Marius, et al.
Pubblicazione: (2025)
di: Smytzek, Marius, et al.
Pubblicazione: (2025)
Molecular De Novo Design through Transformer-based Reinforcement Learning
di: Xu, Pengcheng, et al.
Pubblicazione: (2023)
di: Xu, Pengcheng, et al.
Pubblicazione: (2023)
R1-Code-Interpreter: LLMs Reason with Code via Supervised and Multi-stage Reinforcement Learning
di: Chen, Yongchao, et al.
Pubblicazione: (2025)
di: Chen, Yongchao, et al.
Pubblicazione: (2025)
On Code-Induced Reasoning in LLMs
di: Waheed, Abdul, et al.
Pubblicazione: (2025)
di: Waheed, Abdul, et al.
Pubblicazione: (2025)
How Good Are LLMs at Out-of-Distribution Detection?
di: Liu, Bo, et al.
Pubblicazione: (2023)
di: Liu, Bo, et al.
Pubblicazione: (2023)
Accurate, Efficient, and Explainable Deep Learning Approaches for Environmental Science Problems
di: Shi, Jimeng
Pubblicazione: (2026)
di: Shi, Jimeng
Pubblicazione: (2026)
On the Out-of-Distribution Generalization of Reasoning in Multimodal LLMs for Simple Visual Planning Tasks
di: Neuhaus, Yannic, et al.
Pubblicazione: (2026)
di: Neuhaus, Yannic, et al.
Pubblicazione: (2026)
Early Period of Training Impacts Adaptation for Out-of-Distribution Generalization: An Empirical Study
di: Liu, Chen Cecilia, et al.
Pubblicazione: (2024)
di: Liu, Chen Cecilia, et al.
Pubblicazione: (2024)
An Empirical Study of Pre-trained Model Selection for Out-of-Distribution Generalization and Calibration
di: Naganuma, Hiroki, et al.
Pubblicazione: (2023)
di: Naganuma, Hiroki, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning
di: Lin, Jiacheng, et al.
Pubblicazione: (2025) -
Process-Supervised Reward Models for Verifying Clinical Note Generation: A Scalable Approach Guided by Domain Expertise
di: Wang, Hanyin, et al.
Pubblicazione: (2024) -
Towards Adapting Open-Source Large Language Models for Expert-Level Clinical Note Generation
di: Wang, Hanyin, et al.
Pubblicazione: (2024) -
Bridging the Reproducibility Divide: Open Source Software's Role in Standardizing Healthcare AI
di: Wu, John, et al.
Pubblicazione: (2026) -
Beyond Label Attention: Transparency in Language Models for Automated Medical Coding via Dictionary Learning
di: Wu, John, et al.
Pubblicazione: (2024)