SWITCH: Studying with Teacher for Knowledge Distillation of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Koo, Jahyun, Hwang, Yerin, Kim, Yongil, Kang, Taegwan, Bae, Hyunkyung, Jung, Kyomin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMs can be easily Confused by Instructional Distractions
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
MP2D: An Automated Topic Shift Dialogue Generation Framework Leveraging Knowledge Graphs
by: Hwang, Yerin, et al.
Published: (2024)
by: Hwang, Yerin, et al.
Published: (2024)
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
by: Moon, Jiwon, et al.
Published: (2025)
by: Moon, Jiwon, et al.
Published: (2025)
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
by: Lee, Dongryeol, et al.
Published: (2026)
by: Lee, Dongryeol, et al.
Published: (2026)
When Wording Steers the Evaluation: Framing Bias in LLM judges
by: Hwang, Yerin, et al.
Published: (2026)
by: Hwang, Yerin, et al.
Published: (2026)
Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the effect of Epistemic Markers on LLM-based Evaluation
by: Lee, Dongryeol, et al.
Published: (2024)
by: Lee, Dongryeol, et al.
Published: (2024)
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
Black-Box Hallucination Detection via Consistency Under the Uncertain Expression
by: Joo, Seongho, et al.
Published: (2025)
by: Joo, Seongho, et al.
Published: (2025)
Mitigating Biases for Instruction-following Language Models via Bias Neurons Elimination
by: Yang, Nakyeong, et al.
Published: (2023)
by: Yang, Nakyeong, et al.
Published: (2023)
LifeTox: Unveiling Implicit Toxicity in Life Advice
by: Kim, Minbeom, et al.
Published: (2023)
by: Kim, Minbeom, et al.
Published: (2023)
Program Synthesis via Test-Time Transduction
by: Lee, Kang-il, et al.
Published: (2025)
by: Lee, Kang-il, et al.
Published: (2025)
OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language Models
by: Koo, Jahyun, et al.
Published: (2024)
by: Koo, Jahyun, et al.
Published: (2024)
Knowledge Beyond Language: Bridging the Gap in Multilingual Machine Unlearning Evaluation
by: Hwang, Kyomin, et al.
Published: (2026)
by: Hwang, Kyomin, et al.
Published: (2026)
Fine-grained Gender Control in Machine Translation with Large Language Models
by: Lee, Minwoo, et al.
Published: (2024)
by: Lee, Minwoo, et al.
Published: (2024)
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
by: Min, Kyungmin, et al.
Published: (2024)
by: Min, Kyungmin, et al.
Published: (2024)
FaithUn: Toward Faithful Forgetting in Language Models by Investigating the Interconnectedness of Knowledge
by: Yang, Nakyeong, et al.
Published: (2025)
by: Yang, Nakyeong, et al.
Published: (2025)
Guaranteed Generation from Large Language Models
by: Kim, Minbeom, et al.
Published: (2024)
by: Kim, Minbeom, et al.
Published: (2024)
Reducing Peak Memory Usage for Modern Multimodal Large Language Model Pipelines
by: Kim, Junwan, et al.
Published: (2026)
by: Kim, Junwan, et al.
Published: (2026)
VLind-Bench: Measuring Language Priors in Large Vision-Language Models
by: Lee, Kang-il, et al.
Published: (2024)
by: Lee, Kang-il, et al.
Published: (2024)
How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models
by: Kim, Minsung, et al.
Published: (2025)
by: Kim, Minsung, et al.
Published: (2025)
Retrieval-Augmented Generation Based Nurse Observation Extraction
by: Hwang, Kyomin, et al.
Published: (2026)
by: Hwang, Kyomin, et al.
Published: (2026)
Unplug and Play Language Models: Decomposing Experts in Language Models at Inference Time
by: Yang, Nakyeong, et al.
Published: (2024)
by: Yang, Nakyeong, et al.
Published: (2024)
Knowledge Integration Decay in Search-Augmented Reasoning of Large Language Models
by: Yu, Sangwon, et al.
Published: (2026)
by: Yu, Sangwon, et al.
Published: (2026)
Persona is a Double-edged Sword: Mitigating the Negative Impact of Role-playing Prompts in Zero-shot Reasoning Tasks
by: Kim, Junseok, et al.
Published: (2024)
by: Kim, Junseok, et al.
Published: (2024)
A Character-Centric Creative Story Generation via Imagination
by: Park, Kyeongman, et al.
Published: (2024)
by: Park, Kyeongman, et al.
Published: (2024)
Persona Switch: Mixing Distinct Perspectives in Decoding Time
by: Kim, Junseok, et al.
Published: (2026)
by: Kim, Junseok, et al.
Published: (2026)
Delta Knowledge Distillation for Large Language Models
by: Cao, Yihan, et al.
Published: (2025)
by: Cao, Yihan, et al.
Published: (2025)
Conditional [MASK] Discrete Diffusion Language Model
by: Koh, Hyukhun, et al.
Published: (2024)
by: Koh, Hyukhun, et al.
Published: (2024)
Evolving Knowledge Distillation with Large Language Models and Active Learning
by: Liu, Chengyuan, et al.
Published: (2024)
by: Liu, Chengyuan, et al.
Published: (2024)
SAFE-SQL: Self-Augmented In-Context Learning with Fine-grained Example Selection for Text-to-SQL
by: Lee, Jimin, et al.
Published: (2025)
by: Lee, Jimin, et al.
Published: (2025)
Efficient Knowledge Distillation: Empowering Small Language Models with Teacher Model Insights
by: Ballout, Mohamad, et al.
Published: (2024)
by: Ballout, Mohamad, et al.
Published: (2024)
Carpe Diem: On the Evaluation of World Knowledge in Lifelong Language Models
by: Kim, Yujin, et al.
Published: (2023)
by: Kim, Yujin, et al.
Published: (2023)
Confidence-Guided Stepwise Model Routing for Cost-Efficient Reasoning
by: Lee, Sangmook, et al.
Published: (2025)
by: Lee, Sangmook, et al.
Published: (2025)
Knowledge Distillation of Black-Box Large Language Models
by: Chen, Hongzhan, et al.
Published: (2024)
by: Chen, Hongzhan, et al.
Published: (2024)
Direct Preference Knowledge Distillation for Large Language Models
by: Li, Yixing, et al.
Published: (2024)
by: Li, Yixing, et al.
Published: (2024)
A Survey on Knowledge Distillation of Large Language Models
by: Xu, Xiaohan, et al.
Published: (2024)
by: Xu, Xiaohan, et al.
Published: (2024)
Distilling Rule-based Knowledge into Large Language Models
by: Yang, Wenkai, et al.
Published: (2023)
by: Yang, Wenkai, et al.
Published: (2023)
Casual as an Anchor: Resolving Supervision Misalignment in Formality Transfer Dataset
by: Yu, Hyojeong, et al.
Published: (2026)
by: Yu, Hyojeong, et al.
Published: (2026)
Knowledge Distillation for Temporal Knowledge Graph Reasoning with Large Language Models
by: Xing, Wang, et al.
Published: (2026)
by: Xing, Wang, et al.
Published: (2026)
Similar Items
-
LLMs can be easily Confused by Instructional Distractions
by: Hwang, Yerin, et al.
Published: (2025) -
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
by: Hwang, Yerin, et al.
Published: (2025) -
MP2D: An Automated Topic Shift Dialogue Generation Framework Leveraging Knowledge Graphs
by: Hwang, Yerin, et al.
Published: (2024) -
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
by: Moon, Jiwon, et al.
Published: (2025) -
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
by: Lee, Dongryeol, et al.
Published: (2026)