Is Modularity Transferable? A Case Study through the Lens of Knowledge Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Klimaszewski, Mateusz, Andruszkiewicz, Piotr, Birch, Alexandra |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
No Train but Gain: Language Arithmetic for training-free Language Adapters enhancement
von: Klimaszewski, Mateusz, et al.
Veröffentlicht: (2024)
von: Klimaszewski, Mateusz, et al.
Veröffentlicht: (2024)
Is a Document Educational or Just Wikipedia-Style? -- Pitfalls of Classifier-Based Quality Filtering
von: Klimaszewski, Mateusz, et al.
Veröffentlicht: (2026)
von: Klimaszewski, Mateusz, et al.
Veröffentlicht: (2026)
EuroGEST: Investigating gender stereotypes in multilingual language models
von: Rowe, Jacqueline, et al.
Veröffentlicht: (2025)
von: Rowe, Jacqueline, et al.
Veröffentlicht: (2025)
DistillLens: Symmetric Knowledge Distillation Through Logit Lens
von: Dhakal, Manish, et al.
Veröffentlicht: (2026)
von: Dhakal, Manish, et al.
Veröffentlicht: (2026)
ExpertSteer: Intervening in LLMs through Expert Knowledge
von: Wang, Weixuan, et al.
Veröffentlicht: (2025)
von: Wang, Weixuan, et al.
Veröffentlicht: (2025)
Multilingual Retrieval-Augmented Generation for Knowledge-Intensive Task
von: Ranaldi, Leonardo, et al.
Veröffentlicht: (2025)
von: Ranaldi, Leonardo, et al.
Veröffentlicht: (2025)
Liaozhai through the Looking-Glass: On Paratextual Explicitation of Culture-Bound Terms in Machine Translation
von: Shen, Sherrie, et al.
Veröffentlicht: (2025)
von: Shen, Sherrie, et al.
Veröffentlicht: (2025)
Cache & Distil: Optimising API Calls to Large Language Models
von: Ramírez, Guillem, et al.
Veröffentlicht: (2023)
von: Ramírez, Guillem, et al.
Veröffentlicht: (2023)
Enhancing Low-Resource NMT with a Multilingual Encoder and Knowledge Distillation: A Case Study
von: Roy, Aniruddha, et al.
Veröffentlicht: (2024)
von: Roy, Aniruddha, et al.
Veröffentlicht: (2024)
Exploring and Enhancing the Transfer of Distribution in Knowledge Distillation for Autoregressive Language Models
von: Rao, Jun, et al.
Veröffentlicht: (2024)
von: Rao, Jun, et al.
Veröffentlicht: (2024)
EGAD: Entropy-Guided Adaptive Distillation for Token-Level Knowledge Transfer
von: Zhang, Hao, et al.
Veröffentlicht: (2026)
von: Zhang, Hao, et al.
Veröffentlicht: (2026)
PANDA: Prompt Transfer Meets Knowledge Distillation for Efficient Model Adaptation
von: Zhong, Qihuang, et al.
Veröffentlicht: (2022)
von: Zhong, Qihuang, et al.
Veröffentlicht: (2022)
Improving Multilingual Retrieval-Augmented Language Models through Dialectic Reasoning Argumentations
von: Ranaldi, Leonardo, et al.
Veröffentlicht: (2025)
von: Ranaldi, Leonardo, et al.
Veröffentlicht: (2025)
Optimising Calls to Large Language Models with Uncertainty-Based Two-Tier Selection
von: Ramírez, Guillem, et al.
Veröffentlicht: (2024)
von: Ramírez, Guillem, et al.
Veröffentlicht: (2024)
Evaluating Long-Term Memory for Long-Context Question Answering
von: Terranova, Alessandra, et al.
Veröffentlicht: (2025)
von: Terranova, Alessandra, et al.
Veröffentlicht: (2025)
Feature Alignment and Representation Transfer in Knowledge Distillation for Large Language Models
von: Yang, Junjie, et al.
Veröffentlicht: (2025)
von: Yang, Junjie, et al.
Veröffentlicht: (2025)
Sentence-Level or Token-Level? A Comprehensive Study on Knowledge Distillation
von: Wei, Jingxuan, et al.
Veröffentlicht: (2024)
von: Wei, Jingxuan, et al.
Veröffentlicht: (2024)
Cultural Adaptation of Menus: A Fine-Grained Approach
von: Zhang, Zhonghe, et al.
Veröffentlicht: (2024)
von: Zhang, Zhonghe, et al.
Veröffentlicht: (2024)
Can We Use Probing to Better Understand Fine-tuning and Knowledge Distillation of the BERT NLU?
von: Hościłowicz, Jakub, et al.
Veröffentlicht: (2023)
von: Hościłowicz, Jakub, et al.
Veröffentlicht: (2023)
SWITCH: Studying with Teacher for Knowledge Distillation of Large Language Models
von: Koo, Jahyun, et al.
Veröffentlicht: (2024)
von: Koo, Jahyun, et al.
Veröffentlicht: (2024)
A Multifaceted Analysis of Negative Bias in Large Language Models through the Lens of Parametric Knowledge
von: Song, Jongyoon, et al.
Veröffentlicht: (2025)
von: Song, Jongyoon, et al.
Veröffentlicht: (2025)
Explanation Regularisation through the Lens of Attributions
von: Ferreira, Pedro, et al.
Veröffentlicht: (2024)
von: Ferreira, Pedro, et al.
Veröffentlicht: (2024)
A Survey of Text Style Transfer: Applications and Ethical Implications
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2024)
von: Mukherjee, Sourabrata, et al.
Veröffentlicht: (2024)
Controlling What You Share: Assessing Language Model Adherence to Privacy Preferences
von: Ramírez, Guillem, et al.
Veröffentlicht: (2025)
von: Ramírez, Guillem, et al.
Veröffentlicht: (2025)
Transferring BERT Capabilities from High-Resource to Low-Resource Languages Using Vocabulary Matching
von: Rybak, Piotr
Veröffentlicht: (2024)
von: Rybak, Piotr
Veröffentlicht: (2024)
Hermit Kingdom Through the Lens of Multiple Perspectives: A Case Study of LLM Hallucination on North Korea
von: Cho, Eunjung, et al.
Veröffentlicht: (2025)
von: Cho, Eunjung, et al.
Veröffentlicht: (2025)
Bridging the Language Gaps in Large Language Models with Inference-Time Cross-Lingual Intervention
von: Wang, Weixuan, et al.
Veröffentlicht: (2024)
von: Wang, Weixuan, et al.
Veröffentlicht: (2024)
Demystifying Multilingual Chain-of-Thought in Process Reward Modeling
von: Wang, Weixuan, et al.
Veröffentlicht: (2025)
von: Wang, Weixuan, et al.
Veröffentlicht: (2025)
When Does Monolingual Data Help Multilingual Translation: The Role of Domain and Model Scale
von: Baziotis, Christos, et al.
Veröffentlicht: (2023)
von: Baziotis, Christos, et al.
Veröffentlicht: (2023)
HBO: Hierarchical Balancing Optimization for Fine-Tuning Large Language Models
von: Wang, Weixuan, et al.
Veröffentlicht: (2025)
von: Wang, Weixuan, et al.
Veröffentlicht: (2025)
Learning to Summarize by Learning to Quiz: Adversarial Agentic Collaboration for Long Document Summarization
von: Wang, Weixuan, et al.
Veröffentlicht: (2025)
von: Wang, Weixuan, et al.
Veröffentlicht: (2025)
MGen: Millions of Naturally Occurring Generics in Context
von: Cilleruelo, Gustavo, et al.
Veröffentlicht: (2025)
von: Cilleruelo, Gustavo, et al.
Veröffentlicht: (2025)
The Ups and Downs of Large Language Model Inference with Vocabulary Trimming by Language Heuristics
von: Bogoychev, Nikolay, et al.
Veröffentlicht: (2023)
von: Bogoychev, Nikolay, et al.
Veröffentlicht: (2023)
TinyThinker: Distilling Reasoning through Coarse-to-Fine Knowledge Internalization with Self-Reflection
von: Piao, Shengmin, et al.
Veröffentlicht: (2024)
von: Piao, Shengmin, et al.
Veröffentlicht: (2024)
ReasoningRank: Teaching Student Models to Rank through Reasoning-Based Knowledge Distillation
von: Ji, Yuelyu, et al.
Veröffentlicht: (2024)
von: Ji, Yuelyu, et al.
Veröffentlicht: (2024)
Rethinking Selective Knowledge Distillation
von: Tavor, Almog, et al.
Veröffentlicht: (2026)
von: Tavor, Almog, et al.
Veröffentlicht: (2026)
Machine Translation Meta Evaluation through Translation Accuracy Challenge Sets
von: Moghe, Nikita, et al.
Veröffentlicht: (2024)
von: Moghe, Nikita, et al.
Veröffentlicht: (2024)
Efficient Interleaved Speech Modeling through Knowledge Distillation
von: Nouriborji, Mohammadmahdi, et al.
Veröffentlicht: (2025)
von: Nouriborji, Mohammadmahdi, et al.
Veröffentlicht: (2025)
Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation
von: Mansurov, Jonibek, et al.
Veröffentlicht: (2024)
von: Mansurov, Jonibek, et al.
Veröffentlicht: (2024)
Scaling Knowledge Graph Construction through Synthetic Data Generation and Distillation
von: Choubey, Prafulla Kumar, et al.
Veröffentlicht: (2024)
von: Choubey, Prafulla Kumar, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
No Train but Gain: Language Arithmetic for training-free Language Adapters enhancement
von: Klimaszewski, Mateusz, et al.
Veröffentlicht: (2024) -
Is a Document Educational or Just Wikipedia-Style? -- Pitfalls of Classifier-Based Quality Filtering
von: Klimaszewski, Mateusz, et al.
Veröffentlicht: (2026) -
EuroGEST: Investigating gender stereotypes in multilingual language models
von: Rowe, Jacqueline, et al.
Veröffentlicht: (2025) -
DistillLens: Symmetric Knowledge Distillation Through Logit Lens
von: Dhakal, Manish, et al.
Veröffentlicht: (2026) -
ExpertSteer: Intervening in LLMs through Expert Knowledge
von: Wang, Weixuan, et al.
Veröffentlicht: (2025)