LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Vujanic, Robin, Rueckstiess, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Neural Topic Models with Wasserstein Knowledge Distillation
by: Adhya, Suman, et al.
Published: (2023)
by: Adhya, Suman, et al.
Published: (2023)
FaMTEB: Massive Text Embedding Benchmark in Persian Language
by: Zinvandi, Erfan, et al.
Published: (2025)
by: Zinvandi, Erfan, et al.
Published: (2025)
Optimal Embedding Guided Negative Sample Generation for Knowledge Graph Link Prediction
by: Takamoto, Makoto, et al.
Published: (2025)
by: Takamoto, Makoto, et al.
Published: (2025)
Advancing Academic Knowledge Retrieval via LLM-enhanced Representation Similarity Fusion
by: Dai, Wei, et al.
Published: (2024)
by: Dai, Wei, et al.
Published: (2024)
Tethering Broken Themes: Aligning Neural Topic Models with Labels and Authors
by: Nagda, Mayank, et al.
Published: (2024)
by: Nagda, Mayank, et al.
Published: (2024)
SLMRec: Distilling Large Language Models into Small for Sequential Recommendation
by: Xu, Wujiang, et al.
Published: (2024)
by: Xu, Wujiang, et al.
Published: (2024)
Federated Learning for ICD Classification with Lightweight Models and Pretrained Embeddings
by: Xu, Binbin, et al.
Published: (2025)
by: Xu, Binbin, et al.
Published: (2025)
UniGLM: Training One Unified Language Model for Text-Attributed Graph Embedding
by: Fang, Yi, et al.
Published: (2024)
by: Fang, Yi, et al.
Published: (2024)
Test-Time Compute for Frozen Embedding Models through Agentic Program Search
by: Xiao, Han
Published: (2026)
by: Xiao, Han
Published: (2026)
KGGen: Extracting Knowledge Graphs from Plain Text with Language Models
by: Mo, Belinda, et al.
Published: (2025)
by: Mo, Belinda, et al.
Published: (2025)
Optimizing Multi-Stage Language Models for Effective Text Retrieval
by: Trung, Quang Hoang, et al.
Published: (2024)
by: Trung, Quang Hoang, et al.
Published: (2024)
Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting
by: Qin, Zhen, et al.
Published: (2023)
by: Qin, Zhen, et al.
Published: (2023)
On the Theoretical Limitations of Embedding-Based Retrieval
by: Weller, Orion, et al.
Published: (2025)
by: Weller, Orion, et al.
Published: (2025)
Diffusion-Pretrained Dense and Contextual Embeddings
by: Eslami, Sedigheh, et al.
Published: (2026)
by: Eslami, Sedigheh, et al.
Published: (2026)
Mapping Transformer Leveraged Embeddings for Cross-Lingual Document Representation
by: Tashu, Tsegaye Misikir, et al.
Published: (2024)
by: Tashu, Tsegaye Misikir, et al.
Published: (2024)
Optimizing Small Transformer-Based Language Models for Multi-Label Sentiment Analysis in Short Texts
by: Neumann, Julius, et al.
Published: (2025)
by: Neumann, Julius, et al.
Published: (2025)
Description-Based Text Similarity
by: Ravfogel, Shauli, et al.
Published: (2023)
by: Ravfogel, Shauli, et al.
Published: (2023)
AfroXLMR-Comet: Multilingual Knowledge Distillation with Attention Matching for Low-Resource languages
by: Raju, Joshua Sakthivel, et al.
Published: (2025)
by: Raju, Joshua Sakthivel, et al.
Published: (2025)
Question-Answer Extraction from Scientific Articles Using Knowledge Graphs and Large Language Models
by: Azarbonyad, Hosein, et al.
Published: (2025)
by: Azarbonyad, Hosein, et al.
Published: (2025)
Enhancing Q&A Text Retrieval with Ranking Models: Benchmarking, fine-tuning and deploying Rerankers for RAG
by: Moreira, Gabriel de Souza P., et al.
Published: (2024)
by: Moreira, Gabriel de Souza P., et al.
Published: (2024)
From Topology to Retrieval: Decoding Embedding Spaces with Unified Signatures
by: Rottach, Florian, et al.
Published: (2025)
by: Rottach, Florian, et al.
Published: (2025)
MTEB-French: Resources for French Sentence Embedding Evaluation and Analysis
by: Ciancone, Mathieu, et al.
Published: (2024)
by: Ciancone, Mathieu, et al.
Published: (2024)
Arctic-Embed 2.0: Multilingual Retrieval Without Compromise
by: Yu, Puxuan, et al.
Published: (2024)
by: Yu, Puxuan, et al.
Published: (2024)
Efficient Rationale-based Retrieval: On-policy Distillation from Generative Rerankers based on JEPA
by: Chen, Teng, et al.
Published: (2026)
by: Chen, Teng, et al.
Published: (2026)
Methods for Generating Drift in Text Streams
by: Garcia, Cristiano Mesquita, et al.
Published: (2024)
by: Garcia, Cristiano Mesquita, et al.
Published: (2024)
ConvKGYarn: Spinning Configurable and Scalable Conversational Knowledge Graph QA datasets with Large Language Models
by: Pradeep, Ronak, et al.
Published: (2024)
by: Pradeep, Ronak, et al.
Published: (2024)
Simple Is Effective: The Roles of Graphs and Large Language Models in Knowledge-Graph-Based Retrieval-Augmented Generation
by: Li, Mufei, et al.
Published: (2024)
by: Li, Mufei, et al.
Published: (2024)
SetCSE: Set Operations using Contrastive Learning of Sentence Embeddings
by: Liu, Kang
Published: (2024)
by: Liu, Kang
Published: (2024)
Task-Adaptive Embedding Refinement via Test-time LLM Guidance
by: Gera, Ariel, et al.
Published: (2026)
by: Gera, Ariel, et al.
Published: (2026)
NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
by: Lee, Chankyu, et al.
Published: (2024)
by: Lee, Chankyu, et al.
Published: (2024)
MixLM: High-Throughput and Effective LLM Ranking via Text-Embedding Mix-Interaction
by: Li, Guoyao, et al.
Published: (2025)
by: Li, Guoyao, et al.
Published: (2025)
SoftQE: Learned Representations of Queries Expanded by LLMs
by: Pimpalkhute, Varad, et al.
Published: (2024)
by: Pimpalkhute, Varad, et al.
Published: (2024)
SemCSE: Semantic Contrastive Sentence Embeddings Using LLM-Generated Summaries For Scientific Abstracts
by: Brinner, Marc, et al.
Published: (2025)
by: Brinner, Marc, et al.
Published: (2025)
A Dynamic Framework for Semantic Grouping of Common Data Elements (CDE) Using Embeddings and Clustering
by: Krishnamurthy, Madan, et al.
Published: (2025)
by: Krishnamurthy, Madan, et al.
Published: (2025)
Text Mining Analysis of Symptom Patterns in Medical Chatbot Conversations
by: Razavi, Hamed
Published: (2025)
by: Razavi, Hamed
Published: (2025)
The Structure-Content Trade-off in Knowledge Graph Retrieval
by: Six, Valentin, et al.
Published: (2025)
by: Six, Valentin, et al.
Published: (2025)
Twitter Sentiment Analysis using Distributed Word and Sentence Representation
by: Reddy, Dwarampudi Mahidhar, et al.
Published: (2019)
by: Reddy, Dwarampudi Mahidhar, et al.
Published: (2019)
Adaptive Two-Phase Finetuning LLMs for Japanese Legal Text Retrieval
by: Trung, Quang Hoang, et al.
Published: (2024)
by: Trung, Quang Hoang, et al.
Published: (2024)
Exploring Contrastive Learning for Long-Tailed Multi-Label Text Classification
by: Audibert, Alexandre, et al.
Published: (2024)
by: Audibert, Alexandre, et al.
Published: (2024)
One Word is Enough: Minimal Adversarial Perturbations for Neural Text Ranking
by: Karmakar, Tanmay, et al.
Published: (2026)
by: Karmakar, Tanmay, et al.
Published: (2026)
Similar Items
-
Improving Neural Topic Models with Wasserstein Knowledge Distillation
by: Adhya, Suman, et al.
Published: (2023) -
FaMTEB: Massive Text Embedding Benchmark in Persian Language
by: Zinvandi, Erfan, et al.
Published: (2025) -
Optimal Embedding Guided Negative Sample Generation for Knowledge Graph Link Prediction
by: Takamoto, Makoto, et al.
Published: (2025) -
Advancing Academic Knowledge Retrieval via LLM-enhanced Representation Similarity Fusion
by: Dai, Wei, et al.
Published: (2024) -
Tethering Broken Themes: Aligning Neural Topic Models with Labels and Authors
by: Nagda, Mayank, et al.
Published: (2024)