ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Singh, Aasheesh, Vaddina, Vishal, Birru, Dagnachew |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PROTEUS: SLA-Aware Routing via Lagrangian RL for Multi-LLM Serving Systems
von: Bhatti, Amit Singh, et al.
Veröffentlicht: (2026)
von: Bhatti, Amit Singh, et al.
Veröffentlicht: (2026)
ORPO: Monolithic Preference Optimization without Reference Model
von: Hong, Jiwoo, et al.
Veröffentlicht: (2024)
von: Hong, Jiwoo, et al.
Veröffentlicht: (2024)
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
von: Zhang, Songming, et al.
Veröffentlicht: (2025)
von: Zhang, Songming, et al.
Veröffentlicht: (2025)
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
von: Zhou, Yongchao, et al.
Veröffentlicht: (2023)
von: Zhou, Yongchao, et al.
Veröffentlicht: (2023)
Automatic Prompt Optimization with Prompt Distillation
von: Dyagin, Ernest A., et al.
Veröffentlicht: (2025)
von: Dyagin, Ernest A., et al.
Veröffentlicht: (2025)
LLM Pruning and Distillation in Practice: The Minitron Approach
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2024)
von: Sreenivas, Sharath Turuvekere, et al.
Veröffentlicht: (2024)
Scalable LLM Reasoning Acceleration with Low-rank Distillation
von: Dong, Harry, et al.
Veröffentlicht: (2025)
von: Dong, Harry, et al.
Veröffentlicht: (2025)
Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
von: Fu, Yuqian, et al.
Veröffentlicht: (2026)
von: Fu, Yuqian, et al.
Veröffentlicht: (2026)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
von: Yang, Wenkai, et al.
Veröffentlicht: (2026)
von: Yang, Wenkai, et al.
Veröffentlicht: (2026)
KL for a KL: On-Policy Distillation with Control Variate Baseline
von: Oh, Minjae, et al.
Veröffentlicht: (2026)
von: Oh, Minjae, et al.
Veröffentlicht: (2026)
DP-OPD: Differentially Private On-Policy Distillation for Language Models
von: Khadem, Fatemeh, et al.
Veröffentlicht: (2026)
von: Khadem, Fatemeh, et al.
Veröffentlicht: (2026)
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
von: Agarwal, Rishabh, et al.
Veröffentlicht: (2023)
von: Agarwal, Rishabh, et al.
Veröffentlicht: (2023)
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
von: Li, Yaxuan, et al.
Veröffentlicht: (2026)
von: Li, Yaxuan, et al.
Veröffentlicht: (2026)
DistiLLM: Towards Streamlined Distillation for Large Language Models
von: Ko, Jongwoo, et al.
Veröffentlicht: (2024)
von: Ko, Jongwoo, et al.
Veröffentlicht: (2024)
Distillation Contrastive Decoding: Improving LLMs Reasoning with Contrastive Decoding and Distillation
von: Phan, Phuc, et al.
Veröffentlicht: (2024)
von: Phan, Phuc, et al.
Veröffentlicht: (2024)
MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate
von: Wang, Jianze, et al.
Veröffentlicht: (2026)
von: Wang, Jianze, et al.
Veröffentlicht: (2026)
SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
von: Zheng, Binbin, et al.
Veröffentlicht: (2026)
von: Zheng, Binbin, et al.
Veröffentlicht: (2026)
DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs
von: Ko, Jongwoo, et al.
Veröffentlicht: (2025)
von: Ko, Jongwoo, et al.
Veröffentlicht: (2025)
DOGe: Defensive Output Generation for LLM Protection Against Knowledge Distillation
von: Li, Pingzhi, et al.
Veröffentlicht: (2025)
von: Li, Pingzhi, et al.
Veröffentlicht: (2025)
LLM-NEO: Parameter Efficient Knowledge Distillation for Large Language Models
von: Yang, Runming, et al.
Veröffentlicht: (2024)
von: Yang, Runming, et al.
Veröffentlicht: (2024)
Merge-of-Thought Distillation
von: Shen, Zhanming, et al.
Veröffentlicht: (2025)
von: Shen, Zhanming, et al.
Veröffentlicht: (2025)
Distillation Scaling Laws
von: Busbridge, Dan, et al.
Veröffentlicht: (2025)
von: Busbridge, Dan, et al.
Veröffentlicht: (2025)
TabDistill: Distilling Transformers into Neural Nets for Few-Shot Tabular Classification
von: Dissanayake, Pasan, et al.
Veröffentlicht: (2025)
von: Dissanayake, Pasan, et al.
Veröffentlicht: (2025)
GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation
von: Li, Sijia, et al.
Veröffentlicht: (2026)
von: Li, Sijia, et al.
Veröffentlicht: (2026)
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
von: Chen, Yingfa, et al.
Veröffentlicht: (2026)
von: Chen, Yingfa, et al.
Veröffentlicht: (2026)
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
von: Yang, Zhuolin, et al.
Veröffentlicht: (2026)
von: Yang, Zhuolin, et al.
Veröffentlicht: (2026)
Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment
von: Wang, Jialu, et al.
Veröffentlicht: (2026)
von: Wang, Jialu, et al.
Veröffentlicht: (2026)
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
von: Limozin, Alexis, et al.
Veröffentlicht: (2026)
von: Limozin, Alexis, et al.
Veröffentlicht: (2026)
Accelerated Gradient-based Design Optimization Via Differentiable Physics-Informed Neural Operator: A Composites Autoclave Processing Case Study
von: Patel, Janak M., et al.
Veröffentlicht: (2025)
von: Patel, Janak M., et al.
Veröffentlicht: (2025)
FBI-LLM: Scaling Up Fully Binarized LLMs from Scratch via Autoregressive Distillation
von: Ma, Liqun, et al.
Veröffentlicht: (2024)
von: Ma, Liqun, et al.
Veröffentlicht: (2024)
On Teacher Hacking in Language Model Distillation
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2025)
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2025)
Efficiently Distilling LLMs for Edge Applications
von: Kundu, Achintya, et al.
Veröffentlicht: (2024)
von: Kundu, Achintya, et al.
Veröffentlicht: (2024)
Self-Distilled Agentic Reinforcement Learning
von: Lu, Zhengxi, et al.
Veröffentlicht: (2026)
von: Lu, Zhengxi, et al.
Veröffentlicht: (2026)
Ensemble Distillation for Unsupervised Constituency Parsing
von: Shayegh, Behzad, et al.
Veröffentlicht: (2023)
von: Shayegh, Behzad, et al.
Veröffentlicht: (2023)
KVSculpt: KV Cache Compression as Distillation
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
von: Jiang, Bo, et al.
Veröffentlicht: (2026)
Self-Distillation for Multi-Token Prediction
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026)
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026)
Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning
von: Fan, Chongyu, et al.
Veröffentlicht: (2024)
von: Fan, Chongyu, et al.
Veröffentlicht: (2024)
Local Prompt Optimization
von: Jain, Yash, et al.
Veröffentlicht: (2025)
von: Jain, Yash, et al.
Veröffentlicht: (2025)
DRAG: Distilling RAG for SLMs from LLMs to Transfer Knowledge and Mitigate Hallucination via Evidence and Graph-based Distillation
von: Chen, Jennifer, et al.
Veröffentlicht: (2025)
von: Chen, Jennifer, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PROTEUS: SLA-Aware Routing via Lagrangian RL for Multi-LLM Serving Systems
von: Bhatti, Amit Singh, et al.
Veröffentlicht: (2026) -
ORPO: Monolithic Preference Optimization without Reference Model
von: Hong, Jiwoo, et al.
Veröffentlicht: (2024) -
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
von: Zhang, Songming, et al.
Veröffentlicht: (2025) -
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
von: Zhou, Yongchao, et al.
Veröffentlicht: (2023) -
Automatic Prompt Optimization with Prompt Distillation
von: Dyagin, Ernest A., et al.
Veröffentlicht: (2025)