An Empirical Investigation into the Effect of Parameter Choices in Knowledge Distillation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Sultan, Md Arafat, Trivedi, Aashka, Awasthy, Parul, Sil, Avirup |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Optimal Policy Minimum Bayesian Risk
par: Astudillo, Ramón Fernandez, et autres
Publié: (2025)
par: Astudillo, Ramón Fernandez, et autres
Publié: (2025)
Prompts as Auto-Optimized Training Hyperparameters: Training Best-in-Class IR Models from Scratch with 10 Gold Labels
par: Xian, Jasper, et autres
Publié: (2024)
par: Xian, Jasper, et autres
Publié: (2024)
Multistage Collaborative Knowledge Distillation from a Large Language Model for Semi-Supervised Sequence Generation
par: Zhao, Jiachen, et autres
Publié: (2023)
par: Zhao, Jiachen, et autres
Publié: (2023)
SeaView: Software Engineering Agent Visual Interface for Enhanced Workflow
par: Bula, Timothy, et autres
Publié: (2025)
par: Bula, Timothy, et autres
Publié: (2025)
LMK > CLS: Landmark Pooling for Dense Embeddings
par: Doshi, Meet, et autres
Publié: (2026)
par: Doshi, Meet, et autres
Publié: (2026)
RMT-KD: Random Matrix Theoretic Causal Knowledge Distillation
par: Ettori, Davide, et autres
Publié: (2025)
par: Ettori, Davide, et autres
Publié: (2025)
A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning
par: Liu, Licheng, et autres
Publié: (2025)
par: Liu, Licheng, et autres
Publié: (2025)
PLD: A Choice-Theoretic List-Wise Knowledge Distillation
par: Bassam, Ejafa, et autres
Publié: (2025)
par: Bassam, Ejafa, et autres
Publié: (2025)
ReFIT: Relevance Feedback from a Reranker during Inference
par: Reddy, Revanth Gangi, et autres
Publié: (2023)
par: Reddy, Revanth Gangi, et autres
Publié: (2023)
Granite Embedding Models
par: Awasthy, Parul, et autres
Publié: (2025)
par: Awasthy, Parul, et autres
Publié: (2025)
Content-Aware Attack Detection in LLM Agent Tool-Call Traffic: An Empirical Study of Features, Architectures, and Evaluation Protocols
par: Zavrak, Sultan
Publié: (2026)
par: Zavrak, Sultan
Publié: (2026)
Citation-Grounded Code Comprehension: Preventing LLM Hallucination Through Hybrid Retrieval and Graph-Augmented Context
par: Arafat, Jahidul
Publié: (2025)
par: Arafat, Jahidul
Publié: (2025)
FIRST: Faster Improved Listwise Reranking with Single Token Decoding
par: Reddy, Revanth Gangi, et autres
Publié: (2024)
par: Reddy, Revanth Gangi, et autres
Publié: (2024)
LLM-NEO: Parameter Efficient Knowledge Distillation for Large Language Models
par: Yang, Runming, et autres
Publié: (2024)
par: Yang, Runming, et autres
Publié: (2024)
Importance Analysis for Dynamic Control of Balancing Parameter in a Simple Knowledge Distillation Setting
par: Kim, Seongmin, et autres
Publié: (2025)
par: Kim, Seongmin, et autres
Publié: (2025)
Distill to Delete: Unlearning in Graph Networks with Knowledge Distillation
par: Sinha, Yash, et autres
Publié: (2023)
par: Sinha, Yash, et autres
Publié: (2023)
SMART: Self-Aware Agent for Tool Overuse Mitigation
par: Qian, Cheng, et autres
Publié: (2025)
par: Qian, Cheng, et autres
Publié: (2025)
Parameter Choice and Neuro-Symbolic Approaches for Deep Domain-Invariant Learning
par: Dinu, Marius-Constantin
Publié: (2024)
par: Dinu, Marius-Constantin
Publié: (2024)
Short Data, Long Context: Distilling Positional Knowledge in Transformers
par: Huber, Patrick, et autres
Publié: (2026)
par: Huber, Patrick, et autres
Publié: (2026)
Truncated Matrix Completion - An Empirical Study
par: Naik, Rishhabh, et autres
Publié: (2025)
par: Naik, Rishhabh, et autres
Publié: (2025)
Parameter Efficient Diverse Paraphrase Generation Using Sequence-Level Knowledge Distillation
par: Jayawardena, Lasal, et autres
Publié: (2024)
par: Jayawardena, Lasal, et autres
Publié: (2024)
Feature Distillation is the Better Choice for Model-Heterogeneous Federated Learning
par: Li, Yichen, et autres
Publié: (2025)
par: Li, Yichen, et autres
Publié: (2025)
BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment
par: Sajid, Md. Ashiq Ul Islam, et autres
Publié: (2026)
par: Sajid, Md. Ashiq Ul Islam, et autres
Publié: (2026)
ModelingAgent: Bridging LLMs and Mathematical Modeling for Real-World Challenges
par: Qian, Cheng, et autres
Publié: (2025)
par: Qian, Cheng, et autres
Publié: (2025)
Distilled Circuits: A Mechanistic Study of Internal Restructuring in Knowledge Distillation
par: Haskins, Reilly, et autres
Publié: (2025)
par: Haskins, Reilly, et autres
Publié: (2025)
Online Knowledge Distillation with Reward Guidance
par: Jia, Chen
Publié: (2025)
par: Jia, Chen
Publié: (2025)
How Is Uncertainty Propagated in Knowledge Distillation?
par: Cui, Ziyao, et autres
Publié: (2026)
par: Cui, Ziyao, et autres
Publié: (2026)
Generalizing Scaling Laws for Dense and Sparse Large Language Models
par: Hossain, Md Arafat, et autres
Publié: (2025)
par: Hossain, Md Arafat, et autres
Publié: (2025)
Knowledge Distillation with Adapted Weight
par: Wu, Sirong, et autres
Publié: (2025)
par: Wu, Sirong, et autres
Publié: (2025)
Boosting the Cross-Architecture Generalization of Dataset Distillation through an Empirical Study
par: Zhao, Lirui, et autres
Publié: (2023)
par: Zhao, Lirui, et autres
Publié: (2023)
Semantic-Constrained Federated Aggregation: Convergence Theory and Privacy-Utility Bounds for Knowledge-Enhanced Distributed Learning
par: Arafat, Jahidul
Publié: (2025)
par: Arafat, Jahidul
Publié: (2025)
Self-Refinement of Language Models from External Proxy Metrics Feedback
par: Ramji, Keshav, et autres
Publié: (2024)
par: Ramji, Keshav, et autres
Publié: (2024)
A Teacher-Free Graph Knowledge Distillation Framework with Dual Self-Distillation
par: Wu, Lirong, et autres
Publié: (2024)
par: Wu, Lirong, et autres
Publié: (2024)
A Metamorphic Testing Perspective on Knowledge Distillation for Language Models of Code: Does the Student Deeply Mimic the Teacher?
par: Awal, Md. Abdul, et autres
Publié: (2025)
par: Awal, Md. Abdul, et autres
Publié: (2025)
Relational Linear Properties in Language Models: An Empirical Investigation
par: Valer, Giovanni, et autres
Publié: (2026)
par: Valer, Giovanni, et autres
Publié: (2026)
Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory
par: Xiao, Jiancong, et autres
Publié: (2025)
par: Xiao, Jiancong, et autres
Publié: (2025)
Adaptive Group Robust Ensemble Knowledge Distillation
par: Kenfack, Patrik, et autres
Publié: (2024)
par: Kenfack, Patrik, et autres
Publié: (2024)
BicKD: Bilateral Contrastive Knowledge Distillation
par: Zhu, Jiangnan, et autres
Publié: (2026)
par: Zhu, Jiangnan, et autres
Publié: (2026)
Knowledge Distillation Detection for Open-weights Models
par: Shi, Qin, et autres
Publié: (2025)
par: Shi, Qin, et autres
Publié: (2025)
BadImplant: Injection-based Multi-Targeted Graph Backdoor Attack
par: Khan, Md Nabi Newaz, et autres
Publié: (2026)
par: Khan, Md Nabi Newaz, et autres
Publié: (2026)
Documents similaires
-
Optimal Policy Minimum Bayesian Risk
par: Astudillo, Ramón Fernandez, et autres
Publié: (2025) -
Prompts as Auto-Optimized Training Hyperparameters: Training Best-in-Class IR Models from Scratch with 10 Gold Labels
par: Xian, Jasper, et autres
Publié: (2024) -
Multistage Collaborative Knowledge Distillation from a Large Language Model for Semi-Supervised Sequence Generation
par: Zhao, Jiachen, et autres
Publié: (2023) -
SeaView: Software Engineering Agent Visual Interface for Enhanced Workflow
par: Bula, Timothy, et autres
Publié: (2025) -
LMK > CLS: Landmark Pooling for Dense Embeddings
par: Doshi, Meet, et autres
Publié: (2026)