DP-OPD: Differentially Private On-Policy Distillation for Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Khadem, Fatemeh, Mousavi, Sajad, Fang, Yi, Liu, Yuhong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DP-Fusion: Token-Level Differentially Private Inference for Large Language Models
di: Thareja, Rushil, et al.
Pubblicazione: (2025)
di: Thareja, Rushil, et al.
Pubblicazione: (2025)
MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate
di: Wang, Jianze, et al.
Pubblicazione: (2026)
di: Wang, Jianze, et al.
Pubblicazione: (2026)
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
di: Yuan, Qianhao, et al.
Pubblicazione: (2026)
di: Yuan, Qianhao, et al.
Pubblicazione: (2026)
Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation
di: Zhao, Anhao, et al.
Pubblicazione: (2026)
di: Zhao, Anhao, et al.
Pubblicazione: (2026)
Differentially Private Zeroth-Order Methods for Scalable Large Language Model Finetuning
di: Liu, Z, et al.
Pubblicazione: (2024)
di: Liu, Z, et al.
Pubblicazione: (2024)
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
di: Zhang, Songming, et al.
Pubblicazione: (2025)
di: Zhang, Songming, et al.
Pubblicazione: (2025)
Differentially Private Learning Needs Better Model Initialization and Self-Distillation
di: Ngong, Ivoline C., et al.
Pubblicazione: (2024)
di: Ngong, Ivoline C., et al.
Pubblicazione: (2024)
DP-MemArc: Differential Privacy Transfer Learning for Memory Efficient Language Models
di: Liu, Yanming, et al.
Pubblicazione: (2024)
di: Liu, Yanming, et al.
Pubblicazione: (2024)
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
di: Li, Yaxuan, et al.
Pubblicazione: (2026)
di: Li, Yaxuan, et al.
Pubblicazione: (2026)
OPD+: Rethinking the Advantage Design for On-Policy Distillation
di: Zhao, Hanyang, et al.
Pubblicazione: (2026)
di: Zhao, Hanyang, et al.
Pubblicazione: (2026)
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
di: Agarwal, Rishabh, et al.
Pubblicazione: (2023)
di: Agarwal, Rishabh, et al.
Pubblicazione: (2023)
A Multi-Faceted Evaluation Framework for Assessing Synthetic Data Generated by Large Language Models
di: Yuan, Yefeng, et al.
Pubblicazione: (2024)
di: Yuan, Yefeng, et al.
Pubblicazione: (2024)
Learning to Diagnose Privately: DP-Powered LLMs for Radiology Report Classification
di: Bhattacharjee, Payel, et al.
Pubblicazione: (2025)
di: Bhattacharjee, Payel, et al.
Pubblicazione: (2025)
Evolutionary Contrastive Distillation for Language Model Alignment
di: Katz-Samuels, Julian, et al.
Pubblicazione: (2024)
di: Katz-Samuels, Julian, et al.
Pubblicazione: (2024)
Structured Agent Distillation for Large Language Model
di: Liu, Jun, et al.
Pubblicazione: (2025)
di: Liu, Jun, et al.
Pubblicazione: (2025)
OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification
di: Zhou, Yuhang, et al.
Pubblicazione: (2026)
di: Zhou, Yuhang, et al.
Pubblicazione: (2026)
On Teacher Hacking in Language Model Distillation
di: Tiapkin, Daniil, et al.
Pubblicazione: (2025)
di: Tiapkin, Daniil, et al.
Pubblicazione: (2025)
Private Language Models via Truncated Laplacian Mechanism
di: Huang, Tianhao, et al.
Pubblicazione: (2024)
di: Huang, Tianhao, et al.
Pubblicazione: (2024)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
di: Yang, Wenkai, et al.
Pubblicazione: (2026)
di: Yang, Wenkai, et al.
Pubblicazione: (2026)
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
di: Yang, Zhicheng, et al.
Pubblicazione: (2026)
di: Yang, Zhicheng, et al.
Pubblicazione: (2026)
Coordinated Robustness Evaluation Framework for Vision-Language Models
di: Babu, Ashwin Ramesh, et al.
Pubblicazione: (2025)
di: Babu, Ashwin Ramesh, et al.
Pubblicazione: (2025)
ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM Distillation
di: Singh, Aasheesh, et al.
Pubblicazione: (2025)
di: Singh, Aasheesh, et al.
Pubblicazione: (2025)
Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
di: Fu, Yuqian, et al.
Pubblicazione: (2026)
di: Fu, Yuqian, et al.
Pubblicazione: (2026)
Differentially Private Language Generation and Identification in the Limit
di: Mehrotra, Anay, et al.
Pubblicazione: (2026)
di: Mehrotra, Anay, et al.
Pubblicazione: (2026)
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
di: Wu, Yecheng, et al.
Pubblicazione: (2026)
di: Wu, Yecheng, et al.
Pubblicazione: (2026)
Large Language Models Explore by Latent Distilling
di: Zeng, Yuanhao, et al.
Pubblicazione: (2026)
di: Zeng, Yuanhao, et al.
Pubblicazione: (2026)
Delta Knowledge Distillation for Large Language Models
di: Cao, Yihan, et al.
Pubblicazione: (2025)
di: Cao, Yihan, et al.
Pubblicazione: (2025)
Kakugo: Distillation of Low-Resource Languages into Small Language Models
di: Devine, Peter, et al.
Pubblicazione: (2026)
di: Devine, Peter, et al.
Pubblicazione: (2026)
Efficient Differentially Private Fine-Tuning of LLMs via Reinforcement Learning
di: Khadangi, Afshin, et al.
Pubblicazione: (2025)
di: Khadangi, Afshin, et al.
Pubblicazione: (2025)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
di: Kim, Junhyuck, et al.
Pubblicazione: (2026)
di: Kim, Junhyuck, et al.
Pubblicazione: (2026)
Distilling LLMs' Decomposition Abilities into Compact Language Models
di: Tarasov, Denis, et al.
Pubblicazione: (2024)
di: Tarasov, Denis, et al.
Pubblicazione: (2024)
Compact Language Models via Pruning and Knowledge Distillation
di: Muralidharan, Saurav, et al.
Pubblicazione: (2024)
di: Muralidharan, Saurav, et al.
Pubblicazione: (2024)
Towards the Law of Capacity Gap in Distilling Language Models
di: Zhang, Chen, et al.
Pubblicazione: (2023)
di: Zhang, Chen, et al.
Pubblicazione: (2023)
DistiLLM: Towards Streamlined Distillation for Large Language Models
di: Ko, Jongwoo, et al.
Pubblicazione: (2024)
di: Ko, Jongwoo, et al.
Pubblicazione: (2024)
Leveraging Zero-Shot Prompting for Efficient Language Model Distillation
di: Vöge, Lukas, et al.
Pubblicazione: (2024)
di: Vöge, Lukas, et al.
Pubblicazione: (2024)
LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning
di: Huang, Wei, et al.
Pubblicazione: (2026)
di: Huang, Wei, et al.
Pubblicazione: (2026)
KL for a KL: On-Policy Distillation with Control Variate Baseline
di: Oh, Minjae, et al.
Pubblicazione: (2026)
di: Oh, Minjae, et al.
Pubblicazione: (2026)
Dist2ill: Distributional Distillation for One-Pass Uncertainty Estimation in Large Language Models
di: Zhao, Yicong, et al.
Pubblicazione: (2025)
di: Zhao, Yicong, et al.
Pubblicazione: (2025)
DP-FedAdamW: An Efficient Optimizer for Differentially Private Federated Large Models
di: Liu, Jin, et al.
Pubblicazione: (2026)
di: Liu, Jin, et al.
Pubblicazione: (2026)
Being Strong Progressively! Enhancing Knowledge Distillation of Large Language Models through a Curriculum Learning Framework
di: Liu, Lingyuan, et al.
Pubblicazione: (2025)
di: Liu, Lingyuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
DP-Fusion: Token-Level Differentially Private Inference for Large Language Models
di: Thareja, Rushil, et al.
Pubblicazione: (2025) -
MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate
di: Wang, Jianze, et al.
Pubblicazione: (2026) -
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
di: Yuan, Qianhao, et al.
Pubblicazione: (2026) -
Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation
di: Zhao, Anhao, et al.
Pubblicazione: (2026) -
Differentially Private Zeroth-Order Methods for Scalable Large Language Model Finetuning
di: Liu, Z, et al.
Pubblicazione: (2024)