Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
Fuente:
arXiv
Guardado en:
| Autores principales: | Armandpour, Mohammadreza, Ilhan, Fatih, Harrison, David, Jaiswal, Ajay, Hoang, Duc N. M, Faghri, Fartash, Zhang, Yizhe, Cho, Minsik, Farajtabar, Mehrdad |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TIDE: Every Layer Knows the Token Beneath the Context
por: Jaiswal, Ajay, et al.
Publicado: (2026)
por: Jaiswal, Ajay, et al.
Publicado: (2026)
MemoryLLM: Plug-n-Play Interpretable Feed-Forward Memory for Transformers
por: Jaiswal, Ajay, et al.
Publicado: (2026)
por: Jaiswal, Ajay, et al.
Publicado: (2026)
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
por: Samragh, Mohammad, et al.
Publicado: (2024)
por: Samragh, Mohammad, et al.
Publicado: (2024)
TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining
por: Li, Jeffrey, et al.
Publicado: (2025)
por: Li, Jeffrey, et al.
Publicado: (2025)
SpecMD: A Comprehensive Study On Speculative Expert Prefetching
por: Hoang, Duc, et al.
Publicado: (2026)
por: Hoang, Duc, et al.
Publicado: (2026)
Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context
por: Alizadeh, Keivan, et al.
Publicado: (2026)
por: Alizadeh, Keivan, et al.
Publicado: (2026)
Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific Models
por: Vemulapalli, Raviteja, et al.
Publicado: (2023)
por: Vemulapalli, Raviteja, et al.
Publicado: (2023)
Computational Bottlenecks of Training Small-scale Large Language Models
por: Ashkboos, Saleh, et al.
Publicado: (2024)
por: Ashkboos, Saleh, et al.
Publicado: (2024)
Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential
por: Samragh, Mohammad, et al.
Publicado: (2025)
por: Samragh, Mohammad, et al.
Publicado: (2025)
TiC-CLIP: Continual Training of CLIP Models
por: Garg, Saurabh, et al.
Publicado: (2023)
por: Garg, Saurabh, et al.
Publicado: (2023)
Barriers for Learning in an Evolving World: Mathematical Understanding of Loss of Plasticity
por: Joudaki, Amir, et al.
Publicado: (2025)
por: Joudaki, Amir, et al.
Publicado: (2025)
CatLIP: CLIP-level Visual Recognition Accuracy with 2.7x Faster Pre-training on Web-scale Image-Text Data
por: Mehta, Sachin, et al.
Publicado: (2024)
por: Mehta, Sachin, et al.
Publicado: (2024)
Proxy-FDA: Proxy-based Feature Distribution Alignment for Fine-tuning Vision Foundation Models without Forgetting
por: Huang, Chen, et al.
Publicado: (2025)
por: Huang, Chen, et al.
Publicado: (2025)
SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding
por: Wang, Haoxiang, et al.
Publicado: (2023)
por: Wang, Haoxiang, et al.
Publicado: (2023)
Trigger Where It Hurts: Unveiling Hidden Backdoors through Sensitivity with Sensitron
por: Zhao, Gejian, et al.
Publicado: (2025)
por: Zhao, Gejian, et al.
Publicado: (2025)
Where Does It Hurt? Identifying the Real Concerns in the Ethics of Reference Service.
por: Rothstein, Samuel
Publicado: (1989)
por: Rothstein, Samuel
Publicado: (1989)
MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
por: Hannah, Lauren. A, et al.
Publicado: (2025)
por: Hannah, Lauren. A, et al.
Publicado: (2025)
Where-to-Unmask: Ground-Truth-Guided Unmasking Order Learning for Masked Diffusion Language Models
por: Asano, Hikaru, et al.
Publicado: (2026)
por: Asano, Hikaru, et al.
Publicado: (2026)
Why Self-Training Helps and Hurts: Denoising vs. Signal Forgetting
por: Wu, Mingqi, et al.
Publicado: (2026)
por: Wu, Mingqi, et al.
Publicado: (2026)
Do Compressed LLMs Forget Knowledge? An Experimental Study with Practical Implications
por: Hoang, Duc N. M, et al.
Publicado: (2023)
por: Hoang, Duc N. M, et al.
Publicado: (2023)
Reasoning's Razor: Reasoning Improves Accuracy but Can Hurt Recall at Critical Operating Points in Safety and Hallucination Detection
por: Chegini, Atoosa, et al.
Publicado: (2025)
por: Chegini, Atoosa, et al.
Publicado: (2025)
CLIP with Quality Captions: A Strong Pretraining for Vision Tasks
por: Vasu, Pavan Kumar Anasosalu, et al.
Publicado: (2024)
por: Vasu, Pavan Kumar Anasosalu, et al.
Publicado: (2024)
When, Where and Why to Average Weights?
por: Ajroldi, Niccolò, et al.
Publicado: (2025)
por: Ajroldi, Niccolò, et al.
Publicado: (2025)
Shrinking Cities in Portugal – Where and Why
por: Maria Helena Guimarães
Publicado: (2015)
por: Maria Helena Guimarães
Publicado: (2015)
From Dense to Dynamic: Token-Difficulty Driven MoEfication of Pre-Trained LLMs
por: Nishu, Kumari, et al.
Publicado: (2025)
por: Nishu, Kumari, et al.
Publicado: (2025)
LLM in a flash: Efficient Large Language Model Inference with Limited Memory
por: Alizadeh, Keivan, et al.
Publicado: (2023)
por: Alizadeh, Keivan, et al.
Publicado: (2023)
SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
por: Kim, Han-Byul, et al.
Publicado: (2025)
por: Kim, Han-Byul, et al.
Publicado: (2025)
Duo-LLM: A Framework for Studying Adaptive Computation in Large Language Models
por: Alizadeh, Keivan, et al.
Publicado: (2024)
por: Alizadeh, Keivan, et al.
Publicado: (2024)
TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment
por: Wang, Jiaxuan, et al.
Publicado: (2026)
por: Wang, Jiaxuan, et al.
Publicado: (2026)
AI Where It Matters: Where, Why, and How Developers Want AI Support in Daily Work
por: Choudhuri, Rudrajit, et al.
Publicado: (2025)
por: Choudhuri, Rudrajit, et al.
Publicado: (2025)
Causal Explanations for Disparate Trends: Where and Why?
por: Blau, Tal, et al.
Publicado: (2025)
por: Blau, Tal, et al.
Publicado: (2025)
Advances in Machine Learning: Where Can Quantum Techniques Help?
por: Kashyap, Samarth, et al.
Publicado: (2025)
por: Kashyap, Samarth, et al.
Publicado: (2025)
Theory! Where From and Where to?
por: Inga‐Britt Krause
Publicado: (2026)
por: Inga‐Britt Krause
Publicado: (2026)
It's the Information Age, so Where's the Information? Why Our Students Can't Find It and What We Can Do to Help
por: Jenson, Jill D.
Publicado: (2004)
por: Jenson, Jill D.
Publicado: (2004)
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
por: Vasu, Pavan Kumar Anasosalu, et al.
Publicado: (2023)
por: Vasu, Pavan Kumar Anasosalu, et al.
Publicado: (2023)
Bibliographical Instruction: Where It's At and Where It's Going.
por: McNallie, Bruce
Publicado: (1982)
por: McNallie, Bruce
Publicado: (1982)
A Study of Student Dependency on Artificial Intelligence Applications in their Education: With Reference to Indore City
por: Ajay Jaiswal
Publicado: (2025)
por: Ajay Jaiswal
Publicado: (2025)
Neuron-Guided Interpretation of Code LLMs: Where, Why, and How?
por: Yin, Zhe, et al.
Publicado: (2025)
por: Yin, Zhe, et al.
Publicado: (2025)
Where, What, Why: Towards Explainable Driver Attention Prediction
por: Zhou, Yuchen, et al.
Publicado: (2025)
por: Zhou, Yuchen, et al.
Publicado: (2025)
M.L.S. Librarians in Public Libraries: Where They Are and Why It Matters.
por: Lynch, Mary Jo, et al.
Publicado: (1993)
por: Lynch, Mary Jo, et al.
Publicado: (1993)
Ejemplares similares
-
TIDE: Every Layer Knows the Token Beneath the Context
por: Jaiswal, Ajay, et al.
Publicado: (2026) -
MemoryLLM: Plug-n-Play Interpretable Feed-Forward Memory for Transformers
por: Jaiswal, Ajay, et al.
Publicado: (2026) -
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
por: Samragh, Mohammad, et al.
Publicado: (2024) -
TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining
por: Li, Jeffrey, et al.
Publicado: (2025) -
SpecMD: A Comprehensive Study On Speculative Expert Prefetching
por: Hoang, Duc, et al.
Publicado: (2026)