An Empirical Investigation into the Effect of Parameter Choices in Knowledge Distillation
Fuente:
arXiv
Salvato in:
| Autori principali: | Sultan, Md Arafat, Trivedi, Aashka, Awasthy, Parul, Sil, Avirup |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Optimal Policy Minimum Bayesian Risk
di: Astudillo, Ramón Fernandez, et al.
Pubblicazione: (2025)
di: Astudillo, Ramón Fernandez, et al.
Pubblicazione: (2025)
Prompts as Auto-Optimized Training Hyperparameters: Training Best-in-Class IR Models from Scratch with 10 Gold Labels
di: Xian, Jasper, et al.
Pubblicazione: (2024)
di: Xian, Jasper, et al.
Pubblicazione: (2024)
Multistage Collaborative Knowledge Distillation from a Large Language Model for Semi-Supervised Sequence Generation
di: Zhao, Jiachen, et al.
Pubblicazione: (2023)
di: Zhao, Jiachen, et al.
Pubblicazione: (2023)
SeaView: Software Engineering Agent Visual Interface for Enhanced Workflow
di: Bula, Timothy, et al.
Pubblicazione: (2025)
di: Bula, Timothy, et al.
Pubblicazione: (2025)
LMK > CLS: Landmark Pooling for Dense Embeddings
di: Doshi, Meet, et al.
Pubblicazione: (2026)
di: Doshi, Meet, et al.
Pubblicazione: (2026)
RMT-KD: Random Matrix Theoretic Causal Knowledge Distillation
di: Ettori, Davide, et al.
Pubblicazione: (2025)
di: Ettori, Davide, et al.
Pubblicazione: (2025)
A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning
di: Liu, Licheng, et al.
Pubblicazione: (2025)
di: Liu, Licheng, et al.
Pubblicazione: (2025)
PLD: A Choice-Theoretic List-Wise Knowledge Distillation
di: Bassam, Ejafa, et al.
Pubblicazione: (2025)
di: Bassam, Ejafa, et al.
Pubblicazione: (2025)
ReFIT: Relevance Feedback from a Reranker during Inference
di: Reddy, Revanth Gangi, et al.
Pubblicazione: (2023)
di: Reddy, Revanth Gangi, et al.
Pubblicazione: (2023)
Granite Embedding Models
di: Awasthy, Parul, et al.
Pubblicazione: (2025)
di: Awasthy, Parul, et al.
Pubblicazione: (2025)
Content-Aware Attack Detection in LLM Agent Tool-Call Traffic: An Empirical Study of Features, Architectures, and Evaluation Protocols
di: Zavrak, Sultan
Pubblicazione: (2026)
di: Zavrak, Sultan
Pubblicazione: (2026)
Citation-Grounded Code Comprehension: Preventing LLM Hallucination Through Hybrid Retrieval and Graph-Augmented Context
di: Arafat, Jahidul
Pubblicazione: (2025)
di: Arafat, Jahidul
Pubblicazione: (2025)
FIRST: Faster Improved Listwise Reranking with Single Token Decoding
di: Reddy, Revanth Gangi, et al.
Pubblicazione: (2024)
di: Reddy, Revanth Gangi, et al.
Pubblicazione: (2024)
LLM-NEO: Parameter Efficient Knowledge Distillation for Large Language Models
di: Yang, Runming, et al.
Pubblicazione: (2024)
di: Yang, Runming, et al.
Pubblicazione: (2024)
Importance Analysis for Dynamic Control of Balancing Parameter in a Simple Knowledge Distillation Setting
di: Kim, Seongmin, et al.
Pubblicazione: (2025)
di: Kim, Seongmin, et al.
Pubblicazione: (2025)
Distill to Delete: Unlearning in Graph Networks with Knowledge Distillation
di: Sinha, Yash, et al.
Pubblicazione: (2023)
di: Sinha, Yash, et al.
Pubblicazione: (2023)
SMART: Self-Aware Agent for Tool Overuse Mitigation
di: Qian, Cheng, et al.
Pubblicazione: (2025)
di: Qian, Cheng, et al.
Pubblicazione: (2025)
Parameter Choice and Neuro-Symbolic Approaches for Deep Domain-Invariant Learning
di: Dinu, Marius-Constantin
Pubblicazione: (2024)
di: Dinu, Marius-Constantin
Pubblicazione: (2024)
Short Data, Long Context: Distilling Positional Knowledge in Transformers
di: Huber, Patrick, et al.
Pubblicazione: (2026)
di: Huber, Patrick, et al.
Pubblicazione: (2026)
Truncated Matrix Completion - An Empirical Study
di: Naik, Rishhabh, et al.
Pubblicazione: (2025)
di: Naik, Rishhabh, et al.
Pubblicazione: (2025)
Parameter Efficient Diverse Paraphrase Generation Using Sequence-Level Knowledge Distillation
di: Jayawardena, Lasal, et al.
Pubblicazione: (2024)
di: Jayawardena, Lasal, et al.
Pubblicazione: (2024)
Feature Distillation is the Better Choice for Model-Heterogeneous Federated Learning
di: Li, Yichen, et al.
Pubblicazione: (2025)
di: Li, Yichen, et al.
Pubblicazione: (2025)
BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment
di: Sajid, Md. Ashiq Ul Islam, et al.
Pubblicazione: (2026)
di: Sajid, Md. Ashiq Ul Islam, et al.
Pubblicazione: (2026)
ModelingAgent: Bridging LLMs and Mathematical Modeling for Real-World Challenges
di: Qian, Cheng, et al.
Pubblicazione: (2025)
di: Qian, Cheng, et al.
Pubblicazione: (2025)
Distilled Circuits: A Mechanistic Study of Internal Restructuring in Knowledge Distillation
di: Haskins, Reilly, et al.
Pubblicazione: (2025)
di: Haskins, Reilly, et al.
Pubblicazione: (2025)
Online Knowledge Distillation with Reward Guidance
di: Jia, Chen
Pubblicazione: (2025)
di: Jia, Chen
Pubblicazione: (2025)
How Is Uncertainty Propagated in Knowledge Distillation?
di: Cui, Ziyao, et al.
Pubblicazione: (2026)
di: Cui, Ziyao, et al.
Pubblicazione: (2026)
Generalizing Scaling Laws for Dense and Sparse Large Language Models
di: Hossain, Md Arafat, et al.
Pubblicazione: (2025)
di: Hossain, Md Arafat, et al.
Pubblicazione: (2025)
Knowledge Distillation with Adapted Weight
di: Wu, Sirong, et al.
Pubblicazione: (2025)
di: Wu, Sirong, et al.
Pubblicazione: (2025)
Boosting the Cross-Architecture Generalization of Dataset Distillation through an Empirical Study
di: Zhao, Lirui, et al.
Pubblicazione: (2023)
di: Zhao, Lirui, et al.
Pubblicazione: (2023)
Semantic-Constrained Federated Aggregation: Convergence Theory and Privacy-Utility Bounds for Knowledge-Enhanced Distributed Learning
di: Arafat, Jahidul
Pubblicazione: (2025)
di: Arafat, Jahidul
Pubblicazione: (2025)
Self-Refinement of Language Models from External Proxy Metrics Feedback
di: Ramji, Keshav, et al.
Pubblicazione: (2024)
di: Ramji, Keshav, et al.
Pubblicazione: (2024)
A Teacher-Free Graph Knowledge Distillation Framework with Dual Self-Distillation
di: Wu, Lirong, et al.
Pubblicazione: (2024)
di: Wu, Lirong, et al.
Pubblicazione: (2024)
A Metamorphic Testing Perspective on Knowledge Distillation for Language Models of Code: Does the Student Deeply Mimic the Teacher?
di: Awal, Md. Abdul, et al.
Pubblicazione: (2025)
di: Awal, Md. Abdul, et al.
Pubblicazione: (2025)
Relational Linear Properties in Language Models: An Empirical Investigation
di: Valer, Giovanni, et al.
Pubblicazione: (2026)
di: Valer, Giovanni, et al.
Pubblicazione: (2026)
Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory
di: Xiao, Jiancong, et al.
Pubblicazione: (2025)
di: Xiao, Jiancong, et al.
Pubblicazione: (2025)
Adaptive Group Robust Ensemble Knowledge Distillation
di: Kenfack, Patrik, et al.
Pubblicazione: (2024)
di: Kenfack, Patrik, et al.
Pubblicazione: (2024)
BicKD: Bilateral Contrastive Knowledge Distillation
di: Zhu, Jiangnan, et al.
Pubblicazione: (2026)
di: Zhu, Jiangnan, et al.
Pubblicazione: (2026)
Knowledge Distillation Detection for Open-weights Models
di: Shi, Qin, et al.
Pubblicazione: (2025)
di: Shi, Qin, et al.
Pubblicazione: (2025)
BadImplant: Injection-based Multi-Targeted Graph Backdoor Attack
di: Khan, Md Nabi Newaz, et al.
Pubblicazione: (2026)
di: Khan, Md Nabi Newaz, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Optimal Policy Minimum Bayesian Risk
di: Astudillo, Ramón Fernandez, et al.
Pubblicazione: (2025) -
Prompts as Auto-Optimized Training Hyperparameters: Training Best-in-Class IR Models from Scratch with 10 Gold Labels
di: Xian, Jasper, et al.
Pubblicazione: (2024) -
Multistage Collaborative Knowledge Distillation from a Large Language Model for Semi-Supervised Sequence Generation
di: Zhao, Jiachen, et al.
Pubblicazione: (2023) -
SeaView: Software Engineering Agent Visual Interface for Enhanced Workflow
di: Bula, Timothy, et al.
Pubblicazione: (2025) -
LMK > CLS: Landmark Pooling for Dense Embeddings
di: Doshi, Meet, et al.
Pubblicazione: (2026)