Multi-Sense Embeddings for Language Models and Knowledge Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Qitong, Zaki, Mohammed J., Kollias, Georgios, Kalantzis, Vasileios |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Decoder-based Sense Knowledge Distillation
by: Wang, Qitong, et al.
Published: (2026)
by: Wang, Qitong, et al.
Published: (2026)
KERL: Knowledge-Enhanced Personalized Recipe Recommendation using Large Language Models
by: Mohbat, Fnu, et al.
Published: (2025)
by: Mohbat, Fnu, et al.
Published: (2025)
LLaVA-Chef: A Multi-modal Generative Model for Food Recipes
by: Mohbat, Fnu, et al.
Published: (2024)
by: Mohbat, Fnu, et al.
Published: (2024)
Needle in the Haystack for Memory Based Large Language Models
by: Nelson, Elliot, et al.
Published: (2024)
by: Nelson, Elliot, et al.
Published: (2024)
Generation Constraint Scaling Can Mitigate Hallucination
by: Kollias, Georgios, et al.
Published: (2024)
by: Kollias, Georgios, et al.
Published: (2024)
Knowledge Distillation for Temporal Knowledge Graph Reasoning with Large Language Models
by: Xing, Wang, et al.
Published: (2026)
by: Xing, Wang, et al.
Published: (2026)
Multi-Aspect Knowledge Distillation for Language Model with Low-rank Factorization
by: Liu, Zihe, et al.
Published: (2026)
by: Liu, Zihe, et al.
Published: (2026)
Can Memory-Augmented Language Models Generalize on Reasoning-in-a-Haystack Tasks?
by: Das, Payel, et al.
Published: (2025)
by: Das, Payel, et al.
Published: (2025)
Direct Preference Knowledge Distillation for Large Language Models
by: Li, Yixing, et al.
Published: (2024)
by: Li, Yixing, et al.
Published: (2024)
Memorization Dynamics in Knowledge Distillation for Language Models
by: Borkar, Jaydeep, et al.
Published: (2026)
by: Borkar, Jaydeep, et al.
Published: (2026)
Revisiting Knowledge Distillation for Autoregressive Language Models
by: Zhong, Qihuang, et al.
Published: (2024)
by: Zhong, Qihuang, et al.
Published: (2024)
NeuroPrune: A Neuro-inspired Topological Sparse Training Algorithm for Large Language Models
by: Dhurandhar, Amit, et al.
Published: (2024)
by: Dhurandhar, Amit, et al.
Published: (2024)
Multi-Level Optimal Transport for Universal Cross-Tokenizer Knowledge Distillation on Language Models
by: Cui, Xiao, et al.
Published: (2024)
by: Cui, Xiao, et al.
Published: (2024)
DDK: Distilling Domain Knowledge for Efficient Large Language Models
by: Liu, Jiaheng, et al.
Published: (2024)
by: Liu, Jiaheng, et al.
Published: (2024)
Knowledge Distillation of Black-Box Large Language Models
by: Chen, Hongzhan, et al.
Published: (2024)
by: Chen, Hongzhan, et al.
Published: (2024)
Distilling Rule-based Knowledge into Large Language Models
by: Yang, Wenkai, et al.
Published: (2023)
by: Yang, Wenkai, et al.
Published: (2023)
A Survey on Knowledge Distillation of Large Language Models
by: Xu, Xiaohan, et al.
Published: (2024)
by: Xu, Xiaohan, et al.
Published: (2024)
Enhancing Knowledge Distillation of Large Language Models through Efficient Multi-Modal Distribution Alignment
by: Peng, Tianyu, et al.
Published: (2024)
by: Peng, Tianyu, et al.
Published: (2024)
M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
by: Chen, Jianlv, et al.
Published: (2024)
by: Chen, Jianlv, et al.
Published: (2024)
Feature Alignment-Based Knowledge Distillation for Efficient Compression of Large Language Models
by: Wang, Shuo, et al.
Published: (2024)
by: Wang, Shuo, et al.
Published: (2024)
MLKD-BERT: Multi-level Knowledge Distillation for Pre-trained Language Models
by: Zhang, Ying, et al.
Published: (2024)
by: Zhang, Ying, et al.
Published: (2024)
EasyDistill: A Comprehensive Toolkit for Effective Knowledge Distillation of Large Language Models
by: Wang, Chengyu, et al.
Published: (2025)
by: Wang, Chengyu, et al.
Published: (2025)
SWITCH: Studying with Teacher for Knowledge Distillation of Large Language Models
by: Koo, Jahyun, et al.
Published: (2024)
by: Koo, Jahyun, et al.
Published: (2024)
Evolving Knowledge Distillation with Large Language Models and Active Learning
by: Liu, Chengyuan, et al.
Published: (2024)
by: Liu, Chengyuan, et al.
Published: (2024)
MiniPLM: Knowledge Distillation for Pre-Training Language Models
by: Gu, Yuxian, et al.
Published: (2024)
by: Gu, Yuxian, et al.
Published: (2024)
Cross-Modal Knowledge Distillation for Speech Large Language Models
by: Wang, Enzhi, et al.
Published: (2025)
by: Wang, Enzhi, et al.
Published: (2025)
Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs
by: Jin, Ruihan, et al.
Published: (2026)
by: Jin, Ruihan, et al.
Published: (2026)
Efficient Intent-Based Filtering for Multi-Party Conversations Using Knowledge Distillation from LLMs
by: Gody, Reem, et al.
Published: (2025)
by: Gody, Reem, et al.
Published: (2025)
Efficient Knowledge Distillation: Empowering Small Language Models with Teacher Model Insights
by: Ballout, Mohamad, et al.
Published: (2024)
by: Ballout, Mohamad, et al.
Published: (2024)
LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations
by: Vujanic, Robin, et al.
Published: (2025)
by: Vujanic, Robin, et al.
Published: (2025)
Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data
by: Pham, Bao, et al.
Published: (2026)
by: Pham, Bao, et al.
Published: (2026)
Survey on Knowledge Distillation for Large Language Models: Methods, Evaluation, and Application
by: Yang, Chuanpeng, et al.
Published: (2024)
by: Yang, Chuanpeng, et al.
Published: (2024)
The Valley of Code Reasoning: Scaling Knowledge Distillation of Large Language Models
by: He, Muyu, et al.
Published: (2025)
by: He, Muyu, et al.
Published: (2025)
Exploring and Enhancing the Transfer of Distribution in Knowledge Distillation for Autoregressive Language Models
by: Rao, Jun, et al.
Published: (2024)
by: Rao, Jun, et al.
Published: (2024)
Dual-Space Knowledge Distillation for Large Language Models
by: Zhang, Songming, et al.
Published: (2024)
by: Zhang, Songming, et al.
Published: (2024)
Multi-Granularity Semantic Revision for Large Language Model Distillation
by: Liu, Xiaoyu, et al.
Published: (2024)
by: Liu, Xiaoyu, et al.
Published: (2024)
Confidence-aware Self-Semantic Distillation on Knowledge Graph Embedding
by: Liu, Yichen, et al.
Published: (2022)
by: Liu, Yichen, et al.
Published: (2022)
MMIDR: Teaching Large Language Model to Interpret Multimodal Misinformation via Knowledge Distillation
by: Wang, Longzheng, et al.
Published: (2024)
by: Wang, Longzheng, et al.
Published: (2024)
Knowledge Distillation for Large Language Models
by: La Torre, Alejandro Paredes, et al.
Published: (2026)
by: La Torre, Alejandro Paredes, et al.
Published: (2026)
Rethinking Kullback-Leibler Divergence in Knowledge Distillation for Large Language Models
by: Wu, Taiqiang, et al.
Published: (2024)
by: Wu, Taiqiang, et al.
Published: (2024)
Similar Items
-
Decoder-based Sense Knowledge Distillation
by: Wang, Qitong, et al.
Published: (2026) -
KERL: Knowledge-Enhanced Personalized Recipe Recommendation using Large Language Models
by: Mohbat, Fnu, et al.
Published: (2025) -
LLaVA-Chef: A Multi-modal Generative Model for Food Recipes
by: Mohbat, Fnu, et al.
Published: (2024) -
Needle in the Haystack for Memory Based Large Language Models
by: Nelson, Elliot, et al.
Published: (2024) -
Generation Constraint Scaling Can Mitigate Hallucination
by: Kollias, Georgios, et al.
Published: (2024)