AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Songming, Zhang, Xue, Zhang, Tong, Hu, Bojie, Chen, Yufeng, Xu, Jinan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
KDFlow: A User-Friendly and Efficient Knowledge Distillation Framework for Large Language Models
di: Zhang, Songming, et al.
Pubblicazione: (2026)
di: Zhang, Songming, et al.
Pubblicazione: (2026)
A Dual-Space Framework for General Knowledge Distillation of Large Language Models
di: Zhang, Xue, et al.
Pubblicazione: (2025)
di: Zhang, Xue, et al.
Pubblicazione: (2025)
Dual-Space Knowledge Distillation for Large Language Models
di: Zhang, Songming, et al.
Pubblicazione: (2024)
di: Zhang, Songming, et al.
Pubblicazione: (2024)
Towards Understanding and Improving Knowledge Distillation for Neural Machine Translation
di: Zhang, Songming, et al.
Pubblicazione: (2023)
di: Zhang, Songming, et al.
Pubblicazione: (2023)
Evolutionary Contrastive Distillation for Language Model Alignment
di: Katz-Samuels, Julian, et al.
Pubblicazione: (2024)
di: Katz-Samuels, Julian, et al.
Pubblicazione: (2024)
Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models
di: Li, Chengao, et al.
Pubblicazione: (2025)
di: Li, Chengao, et al.
Pubblicazione: (2025)
Towards the Law of Capacity Gap in Distilling Language Models
di: Zhang, Chen, et al.
Pubblicazione: (2023)
di: Zhang, Chen, et al.
Pubblicazione: (2023)
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level
di: Jia, Nan, et al.
Pubblicazione: (2026)
di: Jia, Nan, et al.
Pubblicazione: (2026)
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
di: Li, Yaxuan, et al.
Pubblicazione: (2026)
di: Li, Yaxuan, et al.
Pubblicazione: (2026)
Structured Agent Distillation for Large Language Model
di: Liu, Jun, et al.
Pubblicazione: (2025)
di: Liu, Jun, et al.
Pubblicazione: (2025)
Large Language Models Explore by Latent Distilling
di: Zeng, Yuanhao, et al.
Pubblicazione: (2026)
di: Zeng, Yuanhao, et al.
Pubblicazione: (2026)
TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment
di: Wang, Jiaxuan, et al.
Pubblicazione: (2026)
di: Wang, Jiaxuan, et al.
Pubblicazione: (2026)
DP-OPD: Differentially Private On-Policy Distillation for Language Models
di: Khadem, Fatemeh, et al.
Pubblicazione: (2026)
di: Khadem, Fatemeh, et al.
Pubblicazione: (2026)
Self-Distillation for Multi-Token Prediction
di: Zhao, Guoliang, et al.
Pubblicazione: (2026)
di: Zhao, Guoliang, et al.
Pubblicazione: (2026)
MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate
di: Wang, Jianze, et al.
Pubblicazione: (2026)
di: Wang, Jianze, et al.
Pubblicazione: (2026)
CM-Align: Consistency-based Multilingual Alignment for Large Language Models
di: Zhang, Xue, et al.
Pubblicazione: (2025)
di: Zhang, Xue, et al.
Pubblicazione: (2025)
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
di: Agarwal, Rishabh, et al.
Pubblicazione: (2023)
di: Agarwal, Rishabh, et al.
Pubblicazione: (2023)
TARo: Token-level Adaptive Routing for LLM Test-time Alignment
di: Rai, Arushi, et al.
Pubblicazione: (2026)
di: Rai, Arushi, et al.
Pubblicazione: (2026)
Distill and Align Decomposition for Enhanced Claim Verification
di: Magomere, Jabez, et al.
Pubblicazione: (2026)
di: Magomere, Jabez, et al.
Pubblicazione: (2026)
BOND: Aligning LLMs with Best-of-N Distillation
di: Sessa, Pier Giuseppe, et al.
Pubblicazione: (2024)
di: Sessa, Pier Giuseppe, et al.
Pubblicazione: (2024)
AlignBench: Benchmarking Chinese Alignment of Large Language Models
di: Liu, Xiao, et al.
Pubblicazione: (2023)
di: Liu, Xiao, et al.
Pubblicazione: (2023)
Probabilistic Token Alignment for Large Language Model Fusion
di: Zeng, Runjia, et al.
Pubblicazione: (2025)
di: Zeng, Runjia, et al.
Pubblicazione: (2025)
TIP: Token Importance in On-Policy Distillation
di: Xu, Yuanda, et al.
Pubblicazione: (2026)
di: Xu, Yuanda, et al.
Pubblicazione: (2026)
TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
di: Zhang, Tunyu, et al.
Pubblicazione: (2025)
di: Zhang, Tunyu, et al.
Pubblicazione: (2025)
TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models
di: Shing, Makoto, et al.
Pubblicazione: (2025)
di: Shing, Makoto, et al.
Pubblicazione: (2025)
SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
di: Zheng, Binbin, et al.
Pubblicazione: (2026)
di: Zheng, Binbin, et al.
Pubblicazione: (2026)
ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM Distillation
di: Singh, Aasheesh, et al.
Pubblicazione: (2025)
di: Singh, Aasheesh, et al.
Pubblicazione: (2025)
On Teacher Hacking in Language Model Distillation
di: Tiapkin, Daniil, et al.
Pubblicazione: (2025)
di: Tiapkin, Daniil, et al.
Pubblicazione: (2025)
TLPO: Token-Level Policy Optimization for Mitigating Language Confusion in Large Language Models
di: Choo, Jinho, et al.
Pubblicazione: (2026)
di: Choo, Jinho, et al.
Pubblicazione: (2026)
Multilingual Safety Alignment via Self-Distillation
di: Qin, Ruiyang, et al.
Pubblicazione: (2026)
di: Qin, Ruiyang, et al.
Pubblicazione: (2026)
GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation
di: Li, Sijia, et al.
Pubblicazione: (2026)
di: Li, Sijia, et al.
Pubblicazione: (2026)
LLM-Oriented Token-Adaptive Knowledge Distillation
di: Xie, Xurong, et al.
Pubblicazione: (2025)
di: Xie, Xurong, et al.
Pubblicazione: (2025)
Being Strong Progressively! Enhancing Knowledge Distillation of Large Language Models through a Curriculum Learning Framework
di: Liu, Lingyuan, et al.
Pubblicazione: (2025)
di: Liu, Lingyuan, et al.
Pubblicazione: (2025)
Merge-of-Thought Distillation
di: Shen, Zhanming, et al.
Pubblicazione: (2025)
di: Shen, Zhanming, et al.
Pubblicazione: (2025)
Delta Knowledge Distillation for Large Language Models
di: Cao, Yihan, et al.
Pubblicazione: (2025)
di: Cao, Yihan, et al.
Pubblicazione: (2025)
Reasoning Distillation and Structural Alignment for Improved Code Generation
di: Jalilifard, Amir, et al.
Pubblicazione: (2025)
di: Jalilifard, Amir, et al.
Pubblicazione: (2025)
DistiLLM: Towards Streamlined Distillation for Large Language Models
di: Ko, Jongwoo, et al.
Pubblicazione: (2024)
di: Ko, Jongwoo, et al.
Pubblicazione: (2024)
LLM-NEO: Parameter Efficient Knowledge Distillation for Large Language Models
di: Yang, Runming, et al.
Pubblicazione: (2024)
di: Yang, Runming, et al.
Pubblicazione: (2024)
Kakugo: Distillation of Low-Resource Languages into Small Language Models
di: Devine, Peter, et al.
Pubblicazione: (2026)
di: Devine, Peter, et al.
Pubblicazione: (2026)
Feature Structure Distillation with Centered Kernel Alignment in BERT Transferring
di: Jung, Hee-Jun, et al.
Pubblicazione: (2022)
di: Jung, Hee-Jun, et al.
Pubblicazione: (2022)
Documenti analoghi
-
KDFlow: A User-Friendly and Efficient Knowledge Distillation Framework for Large Language Models
di: Zhang, Songming, et al.
Pubblicazione: (2026) -
A Dual-Space Framework for General Knowledge Distillation of Large Language Models
di: Zhang, Xue, et al.
Pubblicazione: (2025) -
Dual-Space Knowledge Distillation for Large Language Models
di: Zhang, Songming, et al.
Pubblicazione: (2024) -
Towards Understanding and Improving Knowledge Distillation for Neural Machine Translation
di: Zhang, Songming, et al.
Pubblicazione: (2023) -
Evolutionary Contrastive Distillation for Language Model Alignment
di: Katz-Samuels, Julian, et al.
Pubblicazione: (2024)