Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization
Fuente:
arXiv
Guardado en:
| Autores principales: | Kawakami, Wataru, Suzuki, Keita, Iwasawa, Junichiro |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MedRECT: A Medical Reasoning Benchmark for Error Correction in Clinical Texts
por: Iwase, Naoto, et al.
Publicado: (2025)
por: Iwase, Naoto, et al.
Publicado: (2025)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
por: Lai, Xin, et al.
Publicado: (2024)
por: Lai, Xin, et al.
Publicado: (2024)
HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
por: Chen, Junying, et al.
Publicado: (2024)
por: Chen, Junying, et al.
Publicado: (2024)
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
por: Potamitis, Nearchos, et al.
Publicado: (2025)
por: Potamitis, Nearchos, et al.
Publicado: (2025)
Self-Evolved Preference Optimization for Enhancing Mathematical Reasoning in Small Language Models
por: Singh, Joykirat, et al.
Publicado: (2025)
por: Singh, Joykirat, et al.
Publicado: (2025)
Advancing LLM Reasoning Generalists with Preference Trees
por: Yuan, Lifan, et al.
Publicado: (2024)
por: Yuan, Lifan, et al.
Publicado: (2024)
MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMs
por: Zhao, Guojiang, et al.
Publicado: (2025)
por: Zhao, Guojiang, et al.
Publicado: (2025)
From Noisy Traces to Stable Gradients: Bias-Variance Optimized Preference Optimization for Aligning Large Reasoning Models
por: Zhu, Mingkang, et al.
Publicado: (2025)
por: Zhu, Mingkang, et al.
Publicado: (2025)
Rewarding Graph Reasoning Process makes LLMs more Generalized Reasoners
por: Peng, Miao, et al.
Publicado: (2025)
por: Peng, Miao, et al.
Publicado: (2025)
Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs
por: Kim, Jaemin, et al.
Publicado: (2025)
por: Kim, Jaemin, et al.
Publicado: (2025)
MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining
por: Xiaomi, LLM-Core, et al.
Publicado: (2025)
por: Xiaomi, LLM-Core, et al.
Publicado: (2025)
DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search
por: Yue, Murong, et al.
Publicado: (2024)
por: Yue, Murong, et al.
Publicado: (2024)
AdapThink: Adaptive Thinking Preferences for Reasoning Language Model
por: Wan, Xu, et al.
Publicado: (2025)
por: Wan, Xu, et al.
Publicado: (2025)
Active Preference Inference using Language Models and Probabilistic Reasoning
por: Piriyakulkij, Wasu Top, et al.
Publicado: (2023)
por: Piriyakulkij, Wasu Top, et al.
Publicado: (2023)
CHASE-SQL: Multi-Path Reasoning and Preference Optimized Candidate Selection in Text-to-SQL
por: Pourreza, Mohammadreza, et al.
Publicado: (2024)
por: Pourreza, Mohammadreza, et al.
Publicado: (2024)
The Mouth is Not the Brain: Bridging Energy-Based World Models and Language Generation
por: Niimi, Junichiro
Publicado: (2026)
por: Niimi, Junichiro
Publicado: (2026)
ixi-GEN: Efficient Industrial sLLMs through Domain Adaptive Continual Pretraining
por: Kim, Seonwu, et al.
Publicado: (2025)
por: Kim, Seonwu, et al.
Publicado: (2025)
Prompt Repetition Improves Non-Reasoning LLMs
por: Leviathan, Yaniv, et al.
Publicado: (2025)
por: Leviathan, Yaniv, et al.
Publicado: (2025)
Reverse Thinking Makes LLMs Stronger Reasoners
por: Chen, Justin Chih-Yao, et al.
Publicado: (2024)
por: Chen, Justin Chih-Yao, et al.
Publicado: (2024)
Reasoning Paths Optimization: Learning to Reason and Explore From Diverse Paths
por: Chia, Yew Ken, et al.
Publicado: (2024)
por: Chia, Yew Ken, et al.
Publicado: (2024)
Memorization vs. Reasoning: Updating LLMs with New Knowledge
por: Li, Aochong Oliver, et al.
Publicado: (2025)
por: Li, Aochong Oliver, et al.
Publicado: (2025)
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
por: Lin, Bill Yuchen, et al.
Publicado: (2025)
por: Lin, Bill Yuchen, et al.
Publicado: (2025)
Frontier LLMs Still Struggle with Simple Reasoning Tasks
por: Malek, Alan, et al.
Publicado: (2025)
por: Malek, Alan, et al.
Publicado: (2025)
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
por: Wang, Qibin, et al.
Publicado: (2025)
por: Wang, Qibin, et al.
Publicado: (2025)
Learning to Correct for QA Reasoning with Black-box LLMs
por: Kim, Jaehyung, et al.
Publicado: (2024)
por: Kim, Jaehyung, et al.
Publicado: (2024)
A Decomposition Perspective to Long-context Reasoning for LLMs
por: Xiao, Yanling, et al.
Publicado: (2026)
por: Xiao, Yanling, et al.
Publicado: (2026)
Do LLMs Encode Functional Importance of Reasoning Tokens?
por: Singh, Janvijay, et al.
Publicado: (2026)
por: Singh, Janvijay, et al.
Publicado: (2026)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
por: Lin, Zicheng, et al.
Publicado: (2024)
por: Lin, Zicheng, et al.
Publicado: (2024)
Can Post-Training Transform LLMs into Causal Reasoners?
por: Chen, Junqi, et al.
Publicado: (2026)
por: Chen, Junqi, et al.
Publicado: (2026)
LLMs are not Zero-Shot Reasoners for Biomedical Information Extraction
por: Nagar, Aishik, et al.
Publicado: (2024)
por: Nagar, Aishik, et al.
Publicado: (2024)
Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models
por: Gu, Xiaojie, et al.
Publicado: (2026)
por: Gu, Xiaojie, et al.
Publicado: (2026)
Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization
por: Su, Zhenpeng, et al.
Publicado: (2025)
por: Su, Zhenpeng, et al.
Publicado: (2025)
Reasoning LLMs are Wandering Solution Explorers
por: Lu, Jiahao, et al.
Publicado: (2025)
por: Lu, Jiahao, et al.
Publicado: (2025)
LLMs Can Evolve Continually on Modality for X-Modal Reasoning
por: Yu, Jiazuo, et al.
Publicado: (2024)
por: Yu, Jiazuo, et al.
Publicado: (2024)
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks
por: Xu, Yifei, et al.
Publicado: (2025)
por: Xu, Yifei, et al.
Publicado: (2025)
Time-R1: Towards Comprehensive Temporal Reasoning in LLMs
por: Liu, Zijia, et al.
Publicado: (2025)
por: Liu, Zijia, et al.
Publicado: (2025)
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
por: Qu, Yuxiao, et al.
Publicado: (2025)
por: Qu, Yuxiao, et al.
Publicado: (2025)
Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
por: Deng, Wenhao, et al.
Publicado: (2025)
por: Deng, Wenhao, et al.
Publicado: (2025)
Leveraging Parameter Space Symmetries for Reasoning Skill Transfer in LLMs
por: Horoi, Stefan, et al.
Publicado: (2025)
por: Horoi, Stefan, et al.
Publicado: (2025)
How Likely Do LLMs with CoT Mimic Human Reasoning?
por: Bao, Guangsheng, et al.
Publicado: (2024)
por: Bao, Guangsheng, et al.
Publicado: (2024)
Ejemplares similares
-
MedRECT: A Medical Reasoning Benchmark for Error Correction in Clinical Texts
por: Iwase, Naoto, et al.
Publicado: (2025) -
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
por: Lai, Xin, et al.
Publicado: (2024) -
HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
por: Chen, Junying, et al.
Publicado: (2024) -
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
por: Potamitis, Nearchos, et al.
Publicado: (2025) -
Self-Evolved Preference Optimization for Enhancing Mathematical Reasoning in Small Language Models
por: Singh, Joykirat, et al.
Publicado: (2025)