Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zimmer, Matthieu, Ji, Xiaotong, Nguyen, Tu, Ammar, Haitham Bou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling
von: Nguyen, Tu, et al.
Veröffentlicht: (2026)
von: Nguyen, Tu, et al.
Veröffentlicht: (2026)
Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026)
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026)
Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026)
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026)
The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus
von: Roy, Amartya, et al.
Veröffentlicht: (2026)
von: Roy, Amartya, et al.
Veröffentlicht: (2026)
Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2025)
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2025)
Mixture of Attentions For Speculative Decoding
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2026)
von: Ramesh, Shyam Sundhar, et al.
Veröffentlicht: (2026)
Risk-Controlled Lean-as-Judge for Natural-Language Mathematical Reasoning
von: Bourigault, Pauline, et al.
Veröffentlicht: (2026)
von: Bourigault, Pauline, et al.
Veröffentlicht: (2026)
On Almost Surely Safe Alignment of Large Language Models at Inference-Time
von: Ji, Xiaotong, et al.
Veröffentlicht: (2025)
von: Ji, Xiaotong, et al.
Veröffentlicht: (2025)
Tree-OPO: Off-policy Monte Carlo Tree-Guided Advantage Optimization for Multistep Reasoning
von: Huang, Bingning, et al.
Veröffentlicht: (2025)
von: Huang, Bingning, et al.
Veröffentlicht: (2025)
Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised Reasoning
von: Oomerjee, Adnan, et al.
Veröffentlicht: (2025)
von: Oomerjee, Adnan, et al.
Veröffentlicht: (2025)
OCMDP: Observation-Constrained Markov Decision Process
von: Wang, Taiyi, et al.
Veröffentlicht: (2024)
von: Wang, Taiyi, et al.
Veröffentlicht: (2024)
Rethinking the Role of Temperature in Large Language Model Distillation
von: Luong, Hoang-Chau, et al.
Veröffentlicht: (2026)
von: Luong, Hoang-Chau, et al.
Veröffentlicht: (2026)
SparsePO: Controlling Preference Alignment of LLMs via Sparse Token Masks
von: Christopoulou, Fenia, et al.
Veröffentlicht: (2024)
von: Christopoulou, Fenia, et al.
Veröffentlicht: (2024)
Why the Brain Consolidates: Predictive Forgetting for Optimal Generalisation
von: Fountas, Zafeirios, et al.
Veröffentlicht: (2026)
von: Fountas, Zafeirios, et al.
Veröffentlicht: (2026)
Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy Updates
von: Diwan, Anish, et al.
Veröffentlicht: (2026)
von: Diwan, Anish, et al.
Veröffentlicht: (2026)
Untangling Component Imbalance in Hybrid Linear Attention Conversion Methods
von: Benfeghoul, Martin, et al.
Veröffentlicht: (2025)
von: Benfeghoul, Martin, et al.
Veröffentlicht: (2025)
Subjective Depth and Timescale Transformers: Learning Where and When to Compute
von: Wieser, Frederico, et al.
Veröffentlicht: (2025)
von: Wieser, Frederico, et al.
Veröffentlicht: (2025)
Optimal Decision Tree Policies for Markov Decision Processes
von: Vos, Daniël, et al.
Veröffentlicht: (2023)
von: Vos, Daniël, et al.
Veröffentlicht: (2023)
Markov Decision Processes under External Temporal Processes
von: Ayyagari, Ranga Shaarad, et al.
Veröffentlicht: (2023)
von: Ayyagari, Ranga Shaarad, et al.
Veröffentlicht: (2023)
SuRe: Surprise-Driven Prioritised Replay for Continual LLM Learning
von: Hazard, Hugo, et al.
Veröffentlicht: (2025)
von: Hazard, Hugo, et al.
Veröffentlicht: (2025)
Policy Gradient for Robust Markov Decision Processes
von: Wang, Qiuhao, et al.
Veröffentlicht: (2024)
von: Wang, Qiuhao, et al.
Veröffentlicht: (2024)
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
von: Li, Yaxuan, et al.
Veröffentlicht: (2026)
von: Li, Yaxuan, et al.
Veröffentlicht: (2026)
Robust Lagrangian and Adversarial Policy Gradient for Robust Constrained Markov Decision Processes
von: Bossens, David M.
Veröffentlicht: (2023)
von: Bossens, David M.
Veröffentlicht: (2023)
Rethinking Data Mixing from the Perspective of Large Language Models
von: Xu, Yuanjian, et al.
Veröffentlicht: (2026)
von: Xu, Yuanjian, et al.
Veröffentlicht: (2026)
A Unified Theory of Compositionality, Modularity, and Interpretability in Markov Decision Processes
von: Ringstrom, Thomas J., et al.
Veröffentlicht: (2025)
von: Ringstrom, Thomas J., et al.
Veröffentlicht: (2025)
SPOT: Scalable Policy Optimization with Trees for Markov Decision Processes
von: Xiong, Xuyuan, et al.
Veröffentlicht: (2025)
von: Xiong, Xuyuan, et al.
Veröffentlicht: (2025)
Act as You Learn: Adaptive Decision-Making in Non-Stationary Markov Decision Processes
von: Luo, Baiting, et al.
Veröffentlicht: (2024)
von: Luo, Baiting, et al.
Veröffentlicht: (2024)
Solving Robust Markov Decision Processes: Generic, Reliable, Efficient
von: Meggendorfer, Tobias, et al.
Veröffentlicht: (2024)
von: Meggendorfer, Tobias, et al.
Veröffentlicht: (2024)
Hierarchical Average-Reward Linearly-solvable Markov Decision Processes
von: Infante, Guillermo, et al.
Veröffentlicht: (2024)
von: Infante, Guillermo, et al.
Veröffentlicht: (2024)
Al-Khwarizmi: Discovering Physical Laws with Foundation Models
von: Mower, Christopher E., et al.
Veröffentlicht: (2025)
von: Mower, Christopher E., et al.
Veröffentlicht: (2025)
Linear Mixture Distributionally Robust Markov Decision Processes
von: Liu, Zhishuai, et al.
Veröffentlicht: (2025)
von: Liu, Zhishuai, et al.
Veröffentlicht: (2025)
Diffusion-Augmented Markov Decision Processes for Maximum Entropy Reinforcement Learning
von: Sanokowski, Sebastian, et al.
Veröffentlicht: (2025)
von: Sanokowski, Sebastian, et al.
Veröffentlicht: (2025)
Homomorphic Mappings for Value-Preserving State Aggregation in Markov Decision Processes
von: Zhao, Shuo, et al.
Veröffentlicht: (2025)
von: Zhao, Shuo, et al.
Veröffentlicht: (2025)
Dual Formulation for Non-Rectangular Lp Robust Markov Decision Processes
von: Kumar, Navdeep, et al.
Veröffentlicht: (2025)
von: Kumar, Navdeep, et al.
Veröffentlicht: (2025)
Efficient and Sharp Off-Policy Evaluation in Robust Markov Decision Processes
von: Bennett, Andrew, et al.
Veröffentlicht: (2024)
von: Bennett, Andrew, et al.
Veröffentlicht: (2024)
Optimistic Regret Bounds for Online Learning in Adversarial Markov Decision Processes
von: Moon, Sang Bin, et al.
Veröffentlicht: (2024)
von: Moon, Sang Bin, et al.
Veröffentlicht: (2024)
Constrained Sampling for Language Models Should Be Easy: An MCMC Perspective
von: Gonzalez, Emmanuel Anaya, et al.
Veröffentlicht: (2025)
von: Gonzalez, Emmanuel Anaya, et al.
Veröffentlicht: (2025)
Probing the Decision Boundaries of In-context Learning in Large Language Models
von: Zhao, Siyan, et al.
Veröffentlicht: (2024)
von: Zhao, Siyan, et al.
Veröffentlicht: (2024)
A Cantor-Kantorovich Metric Between Markov Decision Processes with Application to Transfer Learning
von: Banse, Adrien, et al.
Veröffentlicht: (2024)
von: Banse, Adrien, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling
von: Nguyen, Tu, et al.
Veröffentlicht: (2026) -
Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026) -
Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026) -
The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus
von: Roy, Amartya, et al.
Veröffentlicht: (2026) -
Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2025)