On student-teacher deviations in distillation: does it pay to disobey?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nagarajan, Vaishnavh, Menon, Aditya Krishna, Bhojanapalli, Srinadh, Mobahi, Hossein, Kumar, Sanjiv |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Think before you speak: Training Language Models With Pause Tokens
von: Goyal, Sachin, et al.
Veröffentlicht: (2023)
von: Goyal, Sachin, et al.
Veröffentlicht: (2023)
Deep sequence models tend to memorize geometrically; it is unclear why
von: Noroozizadeh, Shahriar, et al.
Veröffentlicht: (2025)
von: Noroozizadeh, Shahriar, et al.
Veröffentlicht: (2025)
HiRE: High Recall Approximate Top-$k$ Estimation for Efficient LLM Inference
von: L, Yashas Samaga B, et al.
Veröffentlicht: (2024)
von: L, Yashas Samaga B, et al.
Veröffentlicht: (2024)
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
The pitfalls of next-token prediction
von: Bachmann, Gregor, et al.
Veröffentlicht: (2024)
von: Bachmann, Gregor, et al.
Veröffentlicht: (2024)
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2025)
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2025)
Data-Aware Random Feature Kernel for Transformers
von: Farzam, Amirhossein, et al.
Veröffentlicht: (2026)
von: Farzam, Amirhossein, et al.
Veröffentlicht: (2026)
Language Model Cascades: Token-level uncertainty and beyond
von: Gupta, Neha, et al.
Veröffentlicht: (2024)
von: Gupta, Neha, et al.
Veröffentlicht: (2024)
Mimetic Initialization Helps State Space Models Learn to Recall
von: Trockman, Asher, et al.
Veröffentlicht: (2024)
von: Trockman, Asher, et al.
Veröffentlicht: (2024)
Faster Cascades via Speculative Decoding
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
Scalable In-context Ranking with Generative Models
von: Gupta, Nilesh, et al.
Veröffentlicht: (2025)
von: Gupta, Nilesh, et al.
Veröffentlicht: (2025)
Autoregressive Ranking: Bridging the Gap Between Dual and Cross Encoders
von: Rozonoyer, Benjamin, et al.
Veröffentlicht: (2026)
von: Rozonoyer, Benjamin, et al.
Veröffentlicht: (2026)
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
von: Zhou, Yongchao, et al.
Veröffentlicht: (2023)
von: Zhou, Yongchao, et al.
Veröffentlicht: (2023)
Trust the uncertain teacher: distilling dark knowledge via calibrated uncertainty
von: Kim, Jeonghyun, et al.
Veröffentlicht: (2026)
von: Kim, Jeonghyun, et al.
Veröffentlicht: (2026)
Efficient Document Ranking with Learnable Late Interactions
von: Ji, Ziwei, et al.
Veröffentlicht: (2024)
von: Ji, Ziwei, et al.
Veröffentlicht: (2024)
Bipartite Ranking From Multiple Labels: On Loss Versus Label Aggregation
von: Lukasik, Michal, et al.
Veröffentlicht: (2025)
von: Lukasik, Michal, et al.
Veröffentlicht: (2025)
LAuReL: Learned Augmented Residual Layer
von: Menghani, Gaurav, et al.
Veröffentlicht: (2024)
von: Menghani, Gaurav, et al.
Veröffentlicht: (2024)
IDLM: Inverse-distilled Diffusion Language Models
von: Li, David, et al.
Veröffentlicht: (2026)
von: Li, David, et al.
Veröffentlicht: (2026)
Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?
von: Gatmiry, Khashayar, et al.
Veröffentlicht: (2024)
von: Gatmiry, Khashayar, et al.
Veröffentlicht: (2024)
Regression-aware Inference with LLMs
von: Lukasik, Michal, et al.
Veröffentlicht: (2024)
von: Lukasik, Michal, et al.
Veröffentlicht: (2024)
Knowledge distillation through geometry-aware representational alignment
von: Bhattarai, Prajjwal, et al.
Veröffentlicht: (2025)
von: Bhattarai, Prajjwal, et al.
Veröffentlicht: (2025)
Sharpness-Aware Minimization Enhances Feature Quality via Balanced Learning
von: Springer, Jacob Mitchell, et al.
Veröffentlicht: (2024)
von: Springer, Jacob Mitchell, et al.
Veröffentlicht: (2024)
Functional Interpolation for Relative Positions Improves Long Context Transformers
von: Li, Shanda, et al.
Veröffentlicht: (2023)
von: Li, Shanda, et al.
Veröffentlicht: (2023)
Supervised learning pays attention
von: Craig, Erin, et al.
Veröffentlicht: (2025)
von: Craig, Erin, et al.
Veröffentlicht: (2025)
Understanding the Failure Modes of Out-of-Distribution Generalization
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2020)
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2020)
Dual-Encoders for Extreme Multi-Label Classification
von: Gupta, Nilesh, et al.
Veröffentlicht: (2023)
von: Gupta, Nilesh, et al.
Veröffentlicht: (2023)
Exploring the potential of prototype-based soft-labels data distillation for imbalanced data classification
von: Rosu, Radu-Andrei, et al.
Veröffentlicht: (2024)
von: Rosu, Radu-Andrei, et al.
Veröffentlicht: (2024)
sDREAMER: Self-distilled Mixture-of-Modality-Experts Transformer for Automatic Sleep Staging
von: Chen, Jingyuan, et al.
Veröffentlicht: (2025)
von: Chen, Jingyuan, et al.
Veröffentlicht: (2025)
LASPA: Language Agnostic Speaker Disentanglement with Prefix-Tuned Cross-Attention
von: Menon, Aditya Srinivas, et al.
Veröffentlicht: (2025)
von: Menon, Aditya Srinivas, et al.
Veröffentlicht: (2025)
Deep Double Q-learning
von: Nagarajan, Prabhat, et al.
Veröffentlicht: (2025)
von: Nagarajan, Prabhat, et al.
Veröffentlicht: (2025)
Reasoning with Latent Thoughts: On the Power of Looped Transformers
von: Saunshi, Nikunj, et al.
Veröffentlicht: (2025)
von: Saunshi, Nikunj, et al.
Veröffentlicht: (2025)
Inference Time Context Sparsity: Illusion or Opportunity?
von: Joshi, Sahil, et al.
Veröffentlicht: (2026)
von: Joshi, Sahil, et al.
Veröffentlicht: (2026)
Towards a theory of model distillation
von: Boix-Adsera, Enric
Veröffentlicht: (2024)
von: Boix-Adsera, Enric
Veröffentlicht: (2024)
Learning temporal embeddings from electronic health records of chronic kidney disease patients
von: Kumar, Aditya, et al.
Veröffentlicht: (2026)
von: Kumar, Aditya, et al.
Veröffentlicht: (2026)
When is Offline Policy Selection Sample Efficient for Reinforcement Learning?
von: Liu, Vincent, et al.
Veröffentlicht: (2023)
von: Liu, Vincent, et al.
Veröffentlicht: (2023)
Learning for Interval Prediction of Electricity Demand: A Cluster-based Bootstrapping Approach
von: Dube, Rohit, et al.
Veröffentlicht: (2023)
von: Dube, Rohit, et al.
Veröffentlicht: (2023)
NoFunEval: Funny How Code LMs Falter on Requirements Beyond Functional Correctness
von: Singhal, Manav, et al.
Veröffentlicht: (2024)
von: Singhal, Manav, et al.
Veröffentlicht: (2024)
Graph-Attentive MAPPO for Dynamic Retail Pricing
von: Amma, Krishna Kumar Neelakanta Pillai Santha Kumari
Veröffentlicht: (2025)
von: Amma, Krishna Kumar Neelakanta Pillai Santha Kumari
Veröffentlicht: (2025)
Multi-Agent Reinforcement Learning for Dynamic Pricing: Balancing Profitability,Stability and Fairness
von: Amma, Krishna Kumar Neelakanta Pillai Santha Kumari
Veröffentlicht: (2026)
von: Amma, Krishna Kumar Neelakanta Pillai Santha Kumari
Veröffentlicht: (2026)
Ähnliche Einträge
-
Think before you speak: Training Language Models With Pause Tokens
von: Goyal, Sachin, et al.
Veröffentlicht: (2023) -
Deep sequence models tend to memorize geometrically; it is unclear why
von: Noroozizadeh, Shahriar, et al.
Veröffentlicht: (2025) -
HiRE: High Recall Approximate Top-$k$ Estimation for Efficient LLM Inference
von: L, Yashas Samaga B, et al.
Veröffentlicht: (2024) -
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
von: Cho, Hanseul, et al.
Veröffentlicht: (2024) -
The pitfalls of next-token prediction
von: Bachmann, Gregor, et al.
Veröffentlicht: (2024)