LoCa: Logit Calibration for Knowledge Distillation
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Runming, Wu, Taiqiang, Yang, Yujiu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLM-NEO: Parameter Efficient Knowledge Distillation for Large Language Models
di: Yang, Runming, et al.
Pubblicazione: (2024)
di: Yang, Runming, et al.
Pubblicazione: (2024)
Weight-Inherited Distillation for Task-Agnostic BERT Compression
di: Wu, Taiqiang, et al.
Pubblicazione: (2023)
di: Wu, Taiqiang, et al.
Pubblicazione: (2023)
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
di: Yang, Xuewei, et al.
Pubblicazione: (2026)
di: Yang, Xuewei, et al.
Pubblicazione: (2026)
Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-Contrast
di: Shi, Chufan, et al.
Pubblicazione: (2024)
di: Shi, Chufan, et al.
Pubblicazione: (2024)
ProFit: Leveraging High-Value Signals in SFT via Probability-Guided Token Selection
di: Liu, Tao, et al.
Pubblicazione: (2026)
di: Liu, Tao, et al.
Pubblicazione: (2026)
Rethinking Kullback-Leibler Divergence in Knowledge Distillation for Large Language Models
di: Wu, Taiqiang, et al.
Pubblicazione: (2024)
di: Wu, Taiqiang, et al.
Pubblicazione: (2024)
OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification
di: Zhou, Yuhang, et al.
Pubblicazione: (2026)
di: Zhou, Yuhang, et al.
Pubblicazione: (2026)
Shadow-FT: Tuning Instruct Model via Training on Paired Base Model
di: Wu, Taiqiang, et al.
Pubblicazione: (2025)
di: Wu, Taiqiang, et al.
Pubblicazione: (2025)
Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMs
di: Anshumann, et al.
Pubblicazione: (2025)
di: Anshumann, et al.
Pubblicazione: (2025)
Progressive Knowledge Graph Completion
di: Li, Jiayi, et al.
Pubblicazione: (2024)
di: Li, Jiayi, et al.
Pubblicazione: (2024)
DALD: Improving Logits-based Detector without Logits from Black-box LLMs
di: Zeng, Cong, et al.
Pubblicazione: (2024)
di: Zeng, Cong, et al.
Pubblicazione: (2024)
Exploiting LLMs for Automatic Hypothesis Assessment via a Logit-Based Calibrated Prior
di: Gong, Yue, et al.
Pubblicazione: (2025)
di: Gong, Yue, et al.
Pubblicazione: (2025)
Stabilizing Policy Optimization via Logits Convexity
di: Chen, Hongzhan, et al.
Pubblicazione: (2026)
di: Chen, Hongzhan, et al.
Pubblicazione: (2026)
ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems
di: Zhou, Wenyong, et al.
Pubblicazione: (2026)
di: Zhou, Wenyong, et al.
Pubblicazione: (2026)
Towards Token-Level Text Anomaly Detection
di: Cao, Yang, et al.
Pubblicazione: (2026)
di: Cao, Yang, et al.
Pubblicazione: (2026)
Edge-free but Structure-aware: Prototype-Guided Knowledge Distillation from GNNs to MLPs
di: Wu, Taiqiang, et al.
Pubblicazione: (2023)
di: Wu, Taiqiang, et al.
Pubblicazione: (2023)
KD-LoRA: A Hybrid Approach to Efficient Fine-Tuning with LoRA and Knowledge Distillation
di: Azimi, Rambod, et al.
Pubblicazione: (2024)
di: Azimi, Rambod, et al.
Pubblicazione: (2024)
MLKD-BERT: Multi-level Knowledge Distillation for Pre-trained Language Models
di: Zhang, Ying, et al.
Pubblicazione: (2024)
di: Zhang, Ying, et al.
Pubblicazione: (2024)
Sinkhorn Distance Minimization for Knowledge Distillation
di: Cui, Xiao, et al.
Pubblicazione: (2024)
di: Cui, Xiao, et al.
Pubblicazione: (2024)
X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillation
di: Sreenivas, Sharath Turuvekere, et al.
Pubblicazione: (2026)
di: Sreenivas, Sharath Turuvekere, et al.
Pubblicazione: (2026)
Relational Knowledge Distillation Using Fine-tuned Function Vectors
di: Kang, Andrea, et al.
Pubblicazione: (2026)
di: Kang, Andrea, et al.
Pubblicazione: (2026)
Logit Reweighting for Topic-Focused Summarization
di: Braun, Joschka, et al.
Pubblicazione: (2025)
di: Braun, Joschka, et al.
Pubblicazione: (2025)
Knowledge Distillation and Dataset Distillation of Large Language Models: Emerging Trends, Challenges, and Future Directions
di: Fang, Luyang, et al.
Pubblicazione: (2025)
di: Fang, Luyang, et al.
Pubblicazione: (2025)
Knowledge Distillation with Training Wheels
di: Liu, Guanlin, et al.
Pubblicazione: (2025)
di: Liu, Guanlin, et al.
Pubblicazione: (2025)
Revisiting Model Interpolation for Efficient Reasoning
di: Wu, Taiqiang, et al.
Pubblicazione: (2025)
di: Wu, Taiqiang, et al.
Pubblicazione: (2025)
Self-Calibrating Language Models via Test-Time Discriminative Distillation
di: Hedna, Mohamed Rissal, et al.
Pubblicazione: (2026)
di: Hedna, Mohamed Rissal, et al.
Pubblicazione: (2026)
Feature Alignment and Representation Transfer in Knowledge Distillation for Large Language Models
di: Yang, Junjie, et al.
Pubblicazione: (2025)
di: Yang, Junjie, et al.
Pubblicazione: (2025)
Dual-Head Knowledge Distillation: Enhancing Logits Utilization with an Auxiliary Head
di: Yang, Penghui, et al.
Pubblicazione: (2024)
di: Yang, Penghui, et al.
Pubblicazione: (2024)
MiniDisc: Minimal Distillation Schedule for Language Model Compression
di: Zhang, Chen, et al.
Pubblicazione: (2022)
di: Zhang, Chen, et al.
Pubblicazione: (2022)
Logits are All We Need to Adapt Closed Models
di: Hiranandani, Gaurush, et al.
Pubblicazione: (2025)
di: Hiranandani, Gaurush, et al.
Pubblicazione: (2025)
Confidence Preservation Property in Knowledge Distillation Abstractions
di: Vengertsev, Dmitry, et al.
Pubblicazione: (2024)
di: Vengertsev, Dmitry, et al.
Pubblicazione: (2024)
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
di: Li, Jin, et al.
Pubblicazione: (2025)
di: Li, Jin, et al.
Pubblicazione: (2025)
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
di: Kim, Gyeongman, et al.
Pubblicazione: (2025)
di: Kim, Gyeongman, et al.
Pubblicazione: (2025)
Efficient Knowledge Injection in LLMs via Self-Distillation
di: Kujanpää, Kalle, et al.
Pubblicazione: (2024)
di: Kujanpää, Kalle, et al.
Pubblicazione: (2024)
In Good GRACEs: Principled Teacher Selection for Knowledge Distillation
di: Panigrahi, Abhishek, et al.
Pubblicazione: (2025)
di: Panigrahi, Abhishek, et al.
Pubblicazione: (2025)
Brewing Knowledge in Context: Distillation Perspectives on In-Context Learning
di: Li, Chengye, et al.
Pubblicazione: (2025)
di: Li, Chengye, et al.
Pubblicazione: (2025)
On the Calibration of Multilingual Question Answering LLMs
di: Yang, Yahan, et al.
Pubblicazione: (2023)
di: Yang, Yahan, et al.
Pubblicazione: (2023)
PromptKD: Distilling Student-Friendly Knowledge for Generative Language Models via Prompt Tuning
di: Kim, Gyeongman, et al.
Pubblicazione: (2024)
di: Kim, Gyeongman, et al.
Pubblicazione: (2024)
LoRA-drop: Efficient LoRA Parameter Pruning based on Output Evaluation
di: Zhou, Hongyun, et al.
Pubblicazione: (2024)
di: Zhou, Hongyun, et al.
Pubblicazione: (2024)
A Survey on Symbolic Knowledge Distillation of Large Language Models
di: Acharya, Kamal, et al.
Pubblicazione: (2024)
di: Acharya, Kamal, et al.
Pubblicazione: (2024)
Documenti analoghi
-
LLM-NEO: Parameter Efficient Knowledge Distillation for Large Language Models
di: Yang, Runming, et al.
Pubblicazione: (2024) -
Weight-Inherited Distillation for Task-Agnostic BERT Compression
di: Wu, Taiqiang, et al.
Pubblicazione: (2023) -
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
di: Yang, Xuewei, et al.
Pubblicazione: (2026) -
Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-Contrast
di: Shi, Chufan, et al.
Pubblicazione: (2024) -
ProFit: Leveraging High-Value Signals in SFT via Probability-Guided Token Selection
di: Liu, Tao, et al.
Pubblicazione: (2026)