Language Model Distillation: A Temporal Difference Imitation Learning Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Zishun, Li, Shangzhe, Zhang, Xinhua |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Coupled Distributional Random Expert Distillation for World Model Online Imitation Learning
by: Li, Shangzhe, et al.
Published: (2025)
by: Li, Shangzhe, et al.
Published: (2025)
Beyond Imitation: Learning Key Reasoning Steps from Dual Chain-of-Thoughts in Reasoning Distillation
by: Dai, Chengwei, et al.
Published: (2024)
by: Dai, Chengwei, et al.
Published: (2024)
BiLD: Bi-directional Logits Difference Loss for Large Language Model Distillation
by: Li, Minchong, et al.
Published: (2024)
by: Li, Minchong, et al.
Published: (2024)
Same Evidence, Different Answers: Canonical-Context On-Policy Distillation for Multi-Turn Language Models
by: Lin, Zizhuo, et al.
Published: (2026)
by: Lin, Zizhuo, et al.
Published: (2026)
Does Knowledge Localization Hold True? Surprising Differences Between Entity and Relation Perspectives in Language Models
by: Wei, Yifan, et al.
Published: (2024)
by: Wei, Yifan, et al.
Published: (2024)
Chunk-Distilled Language Modeling
by: Li, Yanhong, et al.
Published: (2024)
by: Li, Yanhong, et al.
Published: (2024)
The Same But Different: Structural Similarities and Differences in Multilingual Language Modeling
by: Zhang, Ruochen, et al.
Published: (2024)
by: Zhang, Ruochen, et al.
Published: (2024)
Timo: Towards Better Temporal Reasoning for Language Models
by: Su, Zhaochen, et al.
Published: (2024)
by: Su, Zhaochen, et al.
Published: (2024)
Mixed Distillation Helps Smaller Language Model Better Reasoning
by: Li, Chenglin, et al.
Published: (2023)
by: Li, Chenglin, et al.
Published: (2023)
Pedagogically-Inspired Data Synthesis for Language Model Knowledge Distillation
by: He, Bowei, et al.
Published: (2026)
by: He, Bowei, et al.
Published: (2026)
Dual-Space Knowledge Distillation for Large Language Models
by: Zhang, Songming, et al.
Published: (2024)
by: Zhang, Songming, et al.
Published: (2024)
Pre-training Distillation for Large Language Models: A Design Space Exploration
by: Peng, Hao, et al.
Published: (2024)
by: Peng, Hao, et al.
Published: (2024)
MiniLLM: On-Policy Distillation of Large Language Models
by: Gu, Yuxian, et al.
Published: (2023)
by: Gu, Yuxian, et al.
Published: (2023)
PLPP: Prompt Learning with Perplexity Is Self-Distillation for Vision-Language Models
by: Liu, Biao, et al.
Published: (2024)
by: Liu, Biao, et al.
Published: (2024)
RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment
by: Yang, Kevin, et al.
Published: (2023)
by: Yang, Kevin, et al.
Published: (2023)
CoT is Not True Reasoning, It Is Just a Tight Constraint to Imitate: A Theory Perspective
by: Shao, Jintian, et al.
Published: (2025)
by: Shao, Jintian, et al.
Published: (2025)
A Distributional Perspective on Word Learning in Neural Language Models
by: Ficarra, Filippo, et al.
Published: (2025)
by: Ficarra, Filippo, et al.
Published: (2025)
Distilling Mathematical Reasoning Capabilities into Small Language Models
by: Zhu, Xunyu, et al.
Published: (2024)
by: Zhu, Xunyu, et al.
Published: (2024)
EasyDistill: A Comprehensive Toolkit for Effective Knowledge Distillation of Large Language Models
by: Wang, Chengyu, et al.
Published: (2025)
by: Wang, Chengyu, et al.
Published: (2025)
TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models
by: Shing, Makoto, et al.
Published: (2025)
by: Shing, Makoto, et al.
Published: (2025)
Black-Box On-Policy Distillation of Large Language Models
by: Ye, Tianzhu, et al.
Published: (2025)
by: Ye, Tianzhu, et al.
Published: (2025)
Large Language Models Explore by Latent Distilling
by: Zeng, Yuanhao, et al.
Published: (2026)
by: Zeng, Yuanhao, et al.
Published: (2026)
Cross-Modal Knowledge Distillation for Speech Large Language Models
by: Wang, Enzhi, et al.
Published: (2025)
by: Wang, Enzhi, et al.
Published: (2025)
Unveiling Imitation Learning: Exploring the Impact of Data Falsity to Large Language Model
by: Cho, Hyunsoo
Published: (2024)
by: Cho, Hyunsoo
Published: (2024)
ELAD: Explanation-Guided Large Language Models Active Distillation
by: Zhang, Yifei, et al.
Published: (2024)
by: Zhang, Yifei, et al.
Published: (2024)
Imitating Language via Scalable Inverse Reinforcement Learning
by: Wulfmeier, Markus, et al.
Published: (2024)
by: Wulfmeier, Markus, et al.
Published: (2024)
Capturing Nuanced Preferences: Preference-Aligned Distillation for Small Language Models
by: Gu, Yanggan, et al.
Published: (2025)
by: Gu, Yanggan, et al.
Published: (2025)
LLMR: Knowledge Distillation with a Large Language Model-Induced Reward
by: Li, Dongheng, et al.
Published: (2024)
by: Li, Dongheng, et al.
Published: (2024)
OPSDL: On-Policy Self-Distillation for Long-Context Language Models
by: Zhang, Xinsen, et al.
Published: (2026)
by: Zhang, Xinsen, et al.
Published: (2026)
RL from Teacher-Model Refinement: Gradual Imitation Learning for Machine Translation
by: Lee, Dongyub Jude, et al.
Published: (2025)
by: Lee, Dongyub Jude, et al.
Published: (2025)
Differences in Text Generated by Diffusion and Autoregressive Language Models
by: Zhang, Zeyang, et al.
Published: (2026)
by: Zhang, Zeyang, et al.
Published: (2026)
Protecting Language Models Against Unauthorized Distillation through Trace Rewriting
by: Ma, Xinhang, et al.
Published: (2026)
by: Ma, Xinhang, et al.
Published: (2026)
Key-Point-Driven Mathematical Reasoning Distillation of Large Language Model
by: Zhu, Xunyu, et al.
Published: (2024)
by: Zhu, Xunyu, et al.
Published: (2024)
Contextualization Distillation from Large Language Model for Knowledge Graph Completion
by: Li, Dawei, et al.
Published: (2024)
by: Li, Dawei, et al.
Published: (2024)
QCRD: Quality-guided Contrastive Rationale Distillation for Large Language Models
by: Wang, Wei, et al.
Published: (2024)
by: Wang, Wei, et al.
Published: (2024)
EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation
by: Long, Yunbo, et al.
Published: (2026)
by: Long, Yunbo, et al.
Published: (2026)
AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation
by: Zhang, Songming, et al.
Published: (2025)
by: Zhang, Songming, et al.
Published: (2025)
MedKGent: A Large Language Model Agent Framework for Constructing Temporally Evolving Medical Knowledge Graph
by: Zhang, Duzhen, et al.
Published: (2025)
by: Zhang, Duzhen, et al.
Published: (2025)
Less is More: Selective Reflection for Compatible and Efficient Knowledge Distillation in Large Language Models
by: Liu, Lingyuan, et al.
Published: (2025)
by: Liu, Lingyuan, et al.
Published: (2025)
Visuospatial Perspective Taking in Multimodal Language Models
by: Prunty, Jonathan, et al.
Published: (2026)
by: Prunty, Jonathan, et al.
Published: (2026)
Similar Items
-
Coupled Distributional Random Expert Distillation for World Model Online Imitation Learning
by: Li, Shangzhe, et al.
Published: (2025) -
Beyond Imitation: Learning Key Reasoning Steps from Dual Chain-of-Thoughts in Reasoning Distillation
by: Dai, Chengwei, et al.
Published: (2024) -
BiLD: Bi-directional Logits Difference Loss for Large Language Model Distillation
by: Li, Minchong, et al.
Published: (2024) -
Same Evidence, Different Answers: Canonical-Context On-Policy Distillation for Multi-Turn Language Models
by: Lin, Zizhuo, et al.
Published: (2026) -
Does Knowledge Localization Hold True? Surprising Differences Between Entity and Relation Perspectives in Language Models
by: Wei, Yifan, et al.
Published: (2024)