Self-Distillation for Multi-Token Prediction
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhao, Guoliang, Xie, Ruobing, Wang, An, Li, Shuaipeng, Xie, Huaibing, Sun, Xingwu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
More Expressive Attention with Negative Weights
di: Lv, Ang, et al.
Pubblicazione: (2024)
di: Lv, Ang, et al.
Pubblicazione: (2024)
Towards a Comprehensive Scaling Law of Mixture-of-Experts
di: Zhao, Guoliang, et al.
Pubblicazione: (2025)
di: Zhao, Guoliang, et al.
Pubblicazione: (2025)
Language Models "Grok" to Copy
di: Lv, Ang, et al.
Pubblicazione: (2024)
di: Lv, Ang, et al.
Pubblicazione: (2024)
Proximal Supervised Fine-Tuning
di: Zhu, Wenhong, et al.
Pubblicazione: (2025)
di: Zhu, Wenhong, et al.
Pubblicazione: (2025)
LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
di: Yang, Wenkai, et al.
Pubblicazione: (2025)
di: Yang, Wenkai, et al.
Pubblicazione: (2025)
Autonomy-of-Experts Models
di: Lv, Ang, et al.
Pubblicazione: (2025)
di: Lv, Ang, et al.
Pubblicazione: (2025)
A Decomposition Perspective to Long-context Reasoning for LLMs
di: Xiao, Yanling, et al.
Pubblicazione: (2026)
di: Xiao, Yanling, et al.
Pubblicazione: (2026)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
di: Yang, Wenkai, et al.
Pubblicazione: (2026)
di: Yang, Wenkai, et al.
Pubblicazione: (2026)
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
di: Zhang, Songming, et al.
Pubblicazione: (2025)
di: Zhang, Songming, et al.
Pubblicazione: (2025)
MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate
di: Wang, Jianze, et al.
Pubblicazione: (2026)
di: Wang, Jianze, et al.
Pubblicazione: (2026)
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
di: Wang, Hao, et al.
Pubblicazione: (2026)
di: Wang, Hao, et al.
Pubblicazione: (2026)
Dynamics of Spontaneous Topic Changes in Next Token Prediction with Self-Attention
di: Jia, Mumin, et al.
Pubblicazione: (2025)
di: Jia, Mumin, et al.
Pubblicazione: (2025)
TokenButler: Token Importance is Predictable
di: Akhauri, Yash, et al.
Pubblicazione: (2025)
di: Akhauri, Yash, et al.
Pubblicazione: (2025)
Mechanics of Next Token Prediction with Self-Attention
di: Li, Yingcong, et al.
Pubblicazione: (2024)
di: Li, Yingcong, et al.
Pubblicazione: (2024)
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
di: Shen, Guobin, et al.
Pubblicazione: (2026)
di: Shen, Guobin, et al.
Pubblicazione: (2026)
Multilingual Safety Alignment via Self-Distillation
di: Qin, Ruiyang, et al.
Pubblicazione: (2026)
di: Qin, Ruiyang, et al.
Pubblicazione: (2026)
GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation
di: Li, Sijia, et al.
Pubblicazione: (2026)
di: Li, Sijia, et al.
Pubblicazione: (2026)
Data Agent: A Holistic Architecture for Orchestrating Data+AI Ecosystems
di: Sun, Zhaoyan, et al.
Pubblicazione: (2025)
di: Sun, Zhaoyan, et al.
Pubblicazione: (2025)
Cautious Next Token Prediction
di: Wang, Yizhou, et al.
Pubblicazione: (2025)
di: Wang, Yizhou, et al.
Pubblicazione: (2025)
Self-Distilled Agentic Reinforcement Learning
di: Lu, Zhengxi, et al.
Pubblicazione: (2026)
di: Lu, Zhengxi, et al.
Pubblicazione: (2026)
UltraFeedback: Boosting Language Models with Scaled AI Feedback
di: Cui, Ganqu, et al.
Pubblicazione: (2023)
di: Cui, Ganqu, et al.
Pubblicazione: (2023)
Toward Consistent World Models with Multi-Token Prediction and Latent Semantic Enhancement
di: Zhong, Qimin, et al.
Pubblicazione: (2026)
di: Zhong, Qimin, et al.
Pubblicazione: (2026)
Hybrid Policy Distillation for LLMs
di: Zhu, Wenhong, et al.
Pubblicazione: (2026)
di: Zhu, Wenhong, et al.
Pubblicazione: (2026)
M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
di: Chen, Jianlv, et al.
Pubblicazione: (2024)
di: Chen, Jianlv, et al.
Pubblicazione: (2024)
MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?
di: Che, Xinyu, et al.
Pubblicazione: (2026)
di: Che, Xinyu, et al.
Pubblicazione: (2026)
HMoE: Heterogeneous Mixture of Experts for Language Modeling
di: Wang, An, et al.
Pubblicazione: (2024)
di: Wang, An, et al.
Pubblicazione: (2024)
LLM-Enhanced Data Management
di: Zhou, Xuanhe, et al.
Pubblicazione: (2024)
di: Zhou, Xuanhe, et al.
Pubblicazione: (2024)
Lossless Token Sequence Compression via Meta-Tokens
di: Harvill, John, et al.
Pubblicazione: (2025)
di: Harvill, John, et al.
Pubblicazione: (2025)
Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization
di: Fu, Yuhan, et al.
Pubblicazione: (2024)
di: Fu, Yuhan, et al.
Pubblicazione: (2024)
Selective Preference Optimization via Token-Level Reward Function Estimation
di: Yang, Kailai, et al.
Pubblicazione: (2024)
di: Yang, Kailai, et al.
Pubblicazione: (2024)
On the Robustness of Transformers against Context Hijacking for Linear Classification
di: Li, Tianle, et al.
Pubblicazione: (2025)
di: Li, Tianle, et al.
Pubblicazione: (2025)
Reinforce LLM Reasoning through Multi-Agent Reflection
di: Yuan, Yurun, et al.
Pubblicazione: (2025)
di: Yuan, Yurun, et al.
Pubblicazione: (2025)
Post-Trained MoE Can Skip Half Experts via Self-Distillation
di: Lv, Xingtai, et al.
Pubblicazione: (2026)
di: Lv, Xingtai, et al.
Pubblicazione: (2026)
Self-Play with Adversarial Critic: Provable and Scalable Offline Alignment for Language Models
di: Ji, Xiang, et al.
Pubblicazione: (2024)
di: Ji, Xiang, et al.
Pubblicazione: (2024)
Confidence-aware Self-Semantic Distillation on Knowledge Graph Embedding
di: Liu, Yichen, et al.
Pubblicazione: (2022)
di: Liu, Yichen, et al.
Pubblicazione: (2022)
Multi-Token Prediction via Self-Distillation
di: Kirchenbauer, John, et al.
Pubblicazione: (2026)
di: Kirchenbauer, John, et al.
Pubblicazione: (2026)
Efficient Joint Prediction of Multiple Future Tokens
di: Ahn, Kwangjun, et al.
Pubblicazione: (2025)
di: Ahn, Kwangjun, et al.
Pubblicazione: (2025)
HMI: Hierarchical Knowledge Management for Efficient Multi-Tenant Inference in Pretrained Language Models
di: Zhang, Jun, et al.
Pubblicazione: (2025)
di: Zhang, Jun, et al.
Pubblicazione: (2025)
Reveal and Release: Iterative LLM Unlearning with Self-generated Data
di: Xie, Linxi, et al.
Pubblicazione: (2025)
di: Xie, Linxi, et al.
Pubblicazione: (2025)
Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge Distillation
di: Kim, Minsang, et al.
Pubblicazione: (2026)
di: Kim, Minsang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
More Expressive Attention with Negative Weights
di: Lv, Ang, et al.
Pubblicazione: (2024) -
Towards a Comprehensive Scaling Law of Mixture-of-Experts
di: Zhao, Guoliang, et al.
Pubblicazione: (2025) -
Language Models "Grok" to Copy
di: Lv, Ang, et al.
Pubblicazione: (2024) -
Proximal Supervised Fine-Tuning
di: Zhu, Wenhong, et al.
Pubblicazione: (2025) -
LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
di: Yang, Wenkai, et al.
Pubblicazione: (2025)