LCD: Advancing Extreme Low-Bit Clustering for Large Language Models via Knowledge Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Fangxin, Yang, Ning, Zhao, Junping, Yang, Tao, Guan, Haibing, Jiang, Li |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2025)
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2025)
QUARK: Quantization-Enabled Circuit Sharing for Transformer Acceleration by Exploiting Common Patterns in Nonlinear Operations
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2025)
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2025)
PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2025)
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2025)
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
von: Zhang, Tianao, et al.
Veröffentlicht: (2025)
von: Zhang, Tianao, et al.
Veröffentlicht: (2025)
Delta Knowledge Distillation for Large Language Models
von: Cao, Yihan, et al.
Veröffentlicht: (2025)
von: Cao, Yihan, et al.
Veröffentlicht: (2025)
LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning Accelerator
von: Li, Guoyu, et al.
Veröffentlicht: (2025)
von: Li, Guoyu, et al.
Veröffentlicht: (2025)
AdaGMLP: AdaBoosting GNN-to-MLP Knowledge Distillation
von: Lu, Weigang, et al.
Veröffentlicht: (2024)
von: Lu, Weigang, et al.
Veröffentlicht: (2024)
Extreme Low-Bit Inference in Reasoning Models: Failure Modes and Targeted Recovery
von: Alimaskina, Ekaterina, et al.
Veröffentlicht: (2026)
von: Alimaskina, Ekaterina, et al.
Veröffentlicht: (2026)
STRCMP: Integrating Graph Structural Priors with Language Models for Combinatorial Optimization
von: Li, Xijun, et al.
Veröffentlicht: (2025)
von: Li, Xijun, et al.
Veröffentlicht: (2025)
LLM-NEO: Parameter Efficient Knowledge Distillation for Large Language Models
von: Yang, Runming, et al.
Veröffentlicht: (2024)
von: Yang, Runming, et al.
Veröffentlicht: (2024)
HESTIA: A Hessian-Guided Differentiable Quantization-Aware Training Framework for Extremely Low-Bit LLMs
von: Wang, Guoan, et al.
Veröffentlicht: (2026)
von: Wang, Guoan, et al.
Veröffentlicht: (2026)
Efficient Epistemic Uncertainty Estimation for Large Language Models via Knowledge Distillation
von: Park, Seonghyeon, et al.
Veröffentlicht: (2026)
von: Park, Seonghyeon, et al.
Veröffentlicht: (2026)
Lillama: Large Language Models Compression via Low-Rank Feature Distillation
von: Sy, Yaya, et al.
Veröffentlicht: (2024)
von: Sy, Yaya, et al.
Veröffentlicht: (2024)
Large Language Model Guided Knowledge Distillation for Time Series Anomaly Detection
von: Liu, Chen, et al.
Veröffentlicht: (2024)
von: Liu, Chen, et al.
Veröffentlicht: (2024)
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
von: Hu, Xing, et al.
Veröffentlicht: (2024)
von: Hu, Xing, et al.
Veröffentlicht: (2024)
Event-CausNet: Unlocking Causal Knowledge from Text with Large Language Models for Reliable Spatio-Temporal Forecasting
von: Niu, Luyao, et al.
Veröffentlicht: (2025)
von: Niu, Luyao, et al.
Veröffentlicht: (2025)
LBLLM: Lightweight Binarization of Large Language Models via Three-Stage Distillation
von: Song, Siqing, et al.
Veröffentlicht: (2026)
von: Song, Siqing, et al.
Veröffentlicht: (2026)
Low-Dimensional Federated Knowledge Graph Embedding via Knowledge Distillation
von: Zhang, Xiaoxiong, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaoxiong, et al.
Veröffentlicht: (2024)
HyperOffload: Graph-Driven Hierarchical Memory Management for Large Language Models on SuperNode Architectures
von: Liu, Fangxin, et al.
Veröffentlicht: (2026)
von: Liu, Fangxin, et al.
Veröffentlicht: (2026)
PPC-GPT: Federated Task-Specific Compression of Large Language Models via Pruning and Chain-of-Thought Distillation
von: Fan, Tao, et al.
Veröffentlicht: (2025)
von: Fan, Tao, et al.
Veröffentlicht: (2025)
PromptKD: Distilling Student-Friendly Knowledge for Generative Language Models via Prompt Tuning
von: Kim, Gyeongman, et al.
Veröffentlicht: (2024)
von: Kim, Gyeongman, et al.
Veröffentlicht: (2024)
KnowMol: Advancing Molecular Large Language Models with Multi-Level Chemical Knowledge
von: Yang, Zaifei, et al.
Veröffentlicht: (2025)
von: Yang, Zaifei, et al.
Veröffentlicht: (2025)
Structured Agent Distillation for Large Language Model
von: Liu, Jun, et al.
Veröffentlicht: (2025)
von: Liu, Jun, et al.
Veröffentlicht: (2025)
Rethinking Kullback-Leibler Divergence in Knowledge Distillation for Large Language Models
von: Wu, Taiqiang, et al.
Veröffentlicht: (2024)
von: Wu, Taiqiang, et al.
Veröffentlicht: (2024)
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
von: Li, Yaxuan, et al.
Veröffentlicht: (2026)
von: Li, Yaxuan, et al.
Veröffentlicht: (2026)
Enhancing Graph Neural Networks with Limited Labeled Data by Actively Distilling Knowledge from Large Language Models
von: Li, Quan, et al.
Veröffentlicht: (2024)
von: Li, Quan, et al.
Veröffentlicht: (2024)
Compact Language Models via Pruning and Knowledge Distillation
von: Muralidharan, Saurav, et al.
Veröffentlicht: (2024)
von: Muralidharan, Saurav, et al.
Veröffentlicht: (2024)
A&B BNN: Add&Bit-Operation-Only Hardware-Friendly Binary Neural Network
von: Ma, Ruichen, et al.
Veröffentlicht: (2024)
von: Ma, Ruichen, et al.
Veröffentlicht: (2024)
FedMomentum: Preserving LoRA Training Momentum in Federated Fine-Tuning
von: Yan, Peishen, et al.
Veröffentlicht: (2026)
von: Yan, Peishen, et al.
Veröffentlicht: (2026)
Distillation of Large Language Models via Concrete Score Matching
von: Kim, Yeongmin, et al.
Veröffentlicht: (2025)
von: Kim, Yeongmin, et al.
Veröffentlicht: (2025)
Large Language Model Distilling Medication Recommendation Model
von: Liu, Qidong, et al.
Veröffentlicht: (2024)
von: Liu, Qidong, et al.
Veröffentlicht: (2024)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
von: Lee, Banseok, et al.
Veröffentlicht: (2025)
von: Lee, Banseok, et al.
Veröffentlicht: (2025)
OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation
von: Yang, Haochen, et al.
Veröffentlicht: (2026)
von: Yang, Haochen, et al.
Veröffentlicht: (2026)
Bridging the Semantic Gap for Categorical Data Clustering via Large Language Models
von: Yang, Zihua, et al.
Veröffentlicht: (2026)
von: Yang, Zihua, et al.
Veröffentlicht: (2026)
DASH: Input-Aware Dynamic Layer Skipping for Efficient LLM Inference with Markov Decision Policies
von: Yang, Ning, et al.
Veröffentlicht: (2025)
von: Yang, Ning, et al.
Veröffentlicht: (2025)
Advancing Real-time Pandemic Forecasting Using Large Language Models: A COVID-19 Case Study
von: Du, Hongru, et al.
Veröffentlicht: (2024)
von: Du, Hongru, et al.
Veröffentlicht: (2024)
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
von: Kim, Gyeongman, et al.
Veröffentlicht: (2025)
von: Kim, Gyeongman, et al.
Veröffentlicht: (2025)
DARK: Diagonal-Anchored Repulsive Knowledge Distillation for Vision-Language Models under Extreme Compression
von: Saeed, Numan, et al.
Veröffentlicht: (2026)
von: Saeed, Numan, et al.
Veröffentlicht: (2026)
FactorLLM: Factorizing Knowledge via Mixture of Experts for Large Language Models
von: Zhao, Zhongyu, et al.
Veröffentlicht: (2024)
von: Zhao, Zhongyu, et al.
Veröffentlicht: (2024)
Quantization Meets Reasoning: Exploring and Mitigating Degradation of Low-Bit LLMs in Mathematical Reasoning
von: Li, Zhen, et al.
Veröffentlicht: (2025)
von: Li, Zhen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2025) -
QUARK: Quantization-Enabled Circuit Sharing for Transformer Acceleration by Exploiting Common Patterns in Nonlinear Operations
von: Zhao, Zhixiong, et al.
Veröffentlicht: (2025) -
PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2025) -
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
von: Zhang, Tianao, et al.
Veröffentlicht: (2025) -
Delta Knowledge Distillation for Large Language Models
von: Cao, Yihan, et al.
Veröffentlicht: (2025)