DISPO: Enhancing Training Efficiency and Stability in Reinforcement Learning for Large Language Model Mathematical Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Karaman, Batuhan K., Rawal, Aditya, Shakiah, Suhaila, Ghavamzadeh, Mohammad, Hong, Mingyi, Biswas, Arijit, Zhou, Ruida |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GROUNDHOG: Grounding Large Language Models to Holistic Segmentation
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
von: Zhang, Yichi, et al.
Veröffentlicht: (2024)
Directional-Clamp PPO
von: Karpel, Gilad, et al.
Veröffentlicht: (2025)
von: Karpel, Gilad, et al.
Veröffentlicht: (2025)
VaPR -- Vision-language Preference alignment for Reasoning
von: Wadhawan, Rohan, et al.
Veröffentlicht: (2025)
von: Wadhawan, Rohan, et al.
Veröffentlicht: (2025)
HiPER: Hierarchical Reinforcement Learning with Explicit Credit Assignment for Large Language Model Agents
von: Peng, Jiangweizhi, et al.
Veröffentlicht: (2026)
von: Peng, Jiangweizhi, et al.
Veröffentlicht: (2026)
Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback
von: Afsharrad, Amirhossein, et al.
Veröffentlicht: (2026)
von: Afsharrad, Amirhossein, et al.
Veröffentlicht: (2026)
Assessing the significance of longitudinal data in Alzheimer's Disease forecasting
von: Karaman, Batuhan K., et al.
Veröffentlicht: (2024)
von: Karaman, Batuhan K., et al.
Veröffentlicht: (2024)
Direct Preference Optimization with Rating Information: Practical Algorithms and Provable Gains
von: Viano, Luca, et al.
Veröffentlicht: (2026)
von: Viano, Luca, et al.
Veröffentlicht: (2026)
Sequence-level Large Language Model Training with Contrastive Preference Optimization
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
Malware Dataset
von: Syed Suhaila
Veröffentlicht: (2025)
von: Syed Suhaila
Veröffentlicht: (2025)
POROver: Improving Safety and Reducing Overrefusal in Large Language Models with Overgeneration and Preference Optimization
von: Karaman, Batuhan K., et al.
Veröffentlicht: (2024)
von: Karaman, Batuhan K., et al.
Veröffentlicht: (2024)
Maximum Entropy Semi-Supervised Inverse Reinforcement Learning
von: Audiffren, Julien, et al.
Veröffentlicht: (2026)
von: Audiffren, Julien, et al.
Veröffentlicht: (2026)
MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning
von: Das, Debrup, et al.
Veröffentlicht: (2024)
von: Das, Debrup, et al.
Veröffentlicht: (2024)
Reasoning Relay: Evaluating Stability and Interchangeability of Large Language Models in Mathematical Reasoning
von: Lu, Leo, et al.
Veröffentlicht: (2025)
von: Lu, Leo, et al.
Veröffentlicht: (2025)
On the Statistical Efficiency of Mean-Field Reinforcement Learning with General Function Approximation
von: Huang, Jiawei, et al.
Veröffentlicht: (2023)
von: Huang, Jiawei, et al.
Veröffentlicht: (2023)
Mastering Robot Manipulation with Multimodal Prompts through Pretraining and Multi-task Fine-tuning
von: Li, Jiachen, et al.
Veröffentlicht: (2023)
von: Li, Jiachen, et al.
Veröffentlicht: (2023)
Reframing Data Value for Large Language Models Through the Lens of Plausibility
von: Rammal, Mohamad Rida, et al.
Veröffentlicht: (2024)
von: Rammal, Mohamad Rida, et al.
Veröffentlicht: (2024)
Bayesian Regret Minimization in Offline Bandits
von: Petrik, Marek, et al.
Veröffentlicht: (2023)
von: Petrik, Marek, et al.
Veröffentlicht: (2023)
Bayesian policy gradient and actor-critic algorithms
von: Ghavamzadeh, Mohammad, et al.
Veröffentlicht: (2026)
von: Ghavamzadeh, Mohammad, et al.
Veröffentlicht: (2026)
Conservative Contextual Bandits: Beyond Linear Representations
von: Deb, Rohan, et al.
Veröffentlicht: (2024)
von: Deb, Rohan, et al.
Veröffentlicht: (2024)
Contextual Bandits with Stage-wise Constraints
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2024)
von: Pacchiano, Aldo, et al.
Veröffentlicht: (2024)
BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning
von: Zhong, Han, et al.
Veröffentlicht: (2025)
von: Zhong, Han, et al.
Veröffentlicht: (2025)
LLM Reasoning Engine: Specialized Training for Enhanced Mathematical Reasoning
von: Chen, Shuguang, et al.
Veröffentlicht: (2024)
von: Chen, Shuguang, et al.
Veröffentlicht: (2024)
OWL: Geometry-Aware Spatial Reasoning for Audio Large Language Models
von: Biswas, Subrata, et al.
Veröffentlicht: (2025)
von: Biswas, Subrata, et al.
Veröffentlicht: (2025)
Automated rice leaf disease detection using artificial intelligence deep learning
von: Suhaila, M. P., et al.
Veröffentlicht: (2025)
von: Suhaila, M. P., et al.
Veröffentlicht: (2025)
Comparison of fetal Doppler indices and growth in pregnancies with anterior or posterior placental position
von: Suhaila Fadhil Al-Shaikh
Veröffentlicht: (2020)
von: Suhaila Fadhil Al-Shaikh
Veröffentlicht: (2020)
Dual Instruction Tuning with Large Language Models for Mathematical Reasoning
von: Zhou, Yongwei, et al.
Veröffentlicht: (2024)
von: Zhou, Yongwei, et al.
Veröffentlicht: (2024)
Longitudinal Mammogram Risk Prediction
von: Karaman, Batuhan K., et al.
Veröffentlicht: (2024)
von: Karaman, Batuhan K., et al.
Veröffentlicht: (2024)
Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning
von: Xi, Zhiheng, et al.
Veröffentlicht: (2024)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2024)
CAMA: Enhancing Mathematical Reasoning in Large Language Models with Causal Knowledge
von: Zan, Lei, et al.
Veröffentlicht: (2025)
von: Zan, Lei, et al.
Veröffentlicht: (2025)
DEM: Distribution Edited Model for Training with Mixed Data Distributions
von: Ram, Dhananjay, et al.
Veröffentlicht: (2024)
von: Ram, Dhananjay, et al.
Veröffentlicht: (2024)
FANS -- Formal Answer Selection for Natural Language Math Reasoning Using Lean4
von: Yao, Jiarui, et al.
Veröffentlicht: (2025)
von: Yao, Jiarui, et al.
Veröffentlicht: (2025)
Gap-Filling Prompting Enhances Code-Assisted Mathematical Reasoning
von: Mohammadkhani, Mohammad Ghiasvand
Veröffentlicht: (2024)
von: Mohammadkhani, Mohammad Ghiasvand
Veröffentlicht: (2024)
Let's Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM's Math Capability
von: Wang, Ruida, et al.
Veröffentlicht: (2025)
von: Wang, Ruida, et al.
Veröffentlicht: (2025)
Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
von: Tan, Zelin, et al.
Veröffentlicht: (2025)
von: Tan, Zelin, et al.
Veröffentlicht: (2025)
d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning
von: Zhao, Siyan, et al.
Veröffentlicht: (2025)
von: Zhao, Siyan, et al.
Veröffentlicht: (2025)
Attentive AV-FusionNet: Audio-Visual Quality Prediction with Hybrid Attention
von: Salaj, Ina, et al.
Veröffentlicht: (2025)
von: Salaj, Ina, et al.
Veröffentlicht: (2025)
RF-GML: Reference-Free Generative Machine Listener
von: Biswas, Arijit, et al.
Veröffentlicht: (2024)
von: Biswas, Arijit, et al.
Veröffentlicht: (2024)
Towards Evaluating Generative Audio: Insights from Neural Audio Codec Embedding Distances
von: Biswas, Arijit, et al.
Veröffentlicht: (2025)
von: Biswas, Arijit, et al.
Veröffentlicht: (2025)
Enhancing Mathematical Problem Solving in LLMs through Execution-Driven Reasoning Augmentation
von: Basarkar, Aditya, et al.
Veröffentlicht: (2026)
von: Basarkar, Aditya, et al.
Veröffentlicht: (2026)
Adaptive Optimization for Enhanced Efficiency in Large-Scale Language Model Training
von: Chen, Jiajing, et al.
Veröffentlicht: (2024)
von: Chen, Jiajing, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
GROUNDHOG: Grounding Large Language Models to Holistic Segmentation
von: Zhang, Yichi, et al.
Veröffentlicht: (2024) -
Directional-Clamp PPO
von: Karpel, Gilad, et al.
Veröffentlicht: (2025) -
VaPR -- Vision-language Preference alignment for Reasoning
von: Wadhawan, Rohan, et al.
Veröffentlicht: (2025) -
HiPER: Hierarchical Reinforcement Learning with Explicit Credit Assignment for Large Language Model Agents
von: Peng, Jiangweizhi, et al.
Veröffentlicht: (2026) -
Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback
von: Afsharrad, Amirhossein, et al.
Veröffentlicht: (2026)