Sparse Knowledge Distillation: A Mathematical Framework for Probability-Domain Temperature Scaling and Multi-Stage Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Flouro, Aaron R., Chadwick, Shawn P. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Teacher Ensemble Distillation: A Mathematical Framework for Probability-Domain Knowledge Aggregation
by: Flouro, Aaron R., et al.
Published: (2026)
by: Flouro, Aaron R., et al.
Published: (2026)
Recursive Meta-Distillation: An Axiomatic Framework for Iterative Knowledge Refinement
by: Flouro, Aaron R., et al.
Published: (2026)
by: Flouro, Aaron R., et al.
Published: (2026)
Hallucinations Live in Variance
by: Flouro, Aaron R., et al.
Published: (2026)
by: Flouro, Aaron R., et al.
Published: (2026)
Adaptive Weighting in Knowledge Distillation: An Axiomatic Framework for Multi-Scale Teacher Ensemble Optimization
by: Flouro, Aaron R., et al.
Published: (2026)
by: Flouro, Aaron R., et al.
Published: (2026)
Closed-Form Beta Distribution Estimation from Sparse Statistics with Random Forest Implicit Regularization
by: Landers, Jonathan R.
Published: (2025)
by: Landers, Jonathan R.
Published: (2025)
Post-Training Probability Manifold Correction via Structured SVD Pruning and Self-Referential Distillation
by: Flouro, Aaron R., et al.
Published: (2026)
by: Flouro, Aaron R., et al.
Published: (2026)
Causal Direction from Convergence Time: Faster Training in the True Causal Direction
by: Tamim, Abdulrahman
Published: (2026)
by: Tamim, Abdulrahman
Published: (2026)
Understanding the Nature of Generative AI as Threshold Logic in High-Dimensional Space
by: Levin, Ilya
Published: (2026)
by: Levin, Ilya
Published: (2026)
The Current and Future Perspectives of Zinc Oxide Nanoparticles in the Treatment of Diabetes Mellitus
by: Yousaf, Iqra
Published: (2024)
by: Yousaf, Iqra
Published: (2024)
EARCP: Self-Regulating Coherence-Aware Ensemble Architecture for Sequential Decision Making -- Ensemble Auto-Regule par Coherence et Performance
by: Amega, Mike
Published: (2026)
by: Amega, Mike
Published: (2026)
Frequency Principle: Fourier Analysis Sheds Light on Deep Neural Networks
by: Xu, Zhi-Qin John, et al.
Published: (2019)
by: Xu, Zhi-Qin John, et al.
Published: (2019)
Backpropagation Through Time For Networks With Long-Term Dependencies
by: Bird, George, et al.
Published: (2021)
by: Bird, George, et al.
Published: (2021)
Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning in LLMs
by: Gulati, Aryan, et al.
Published: (2025)
by: Gulati, Aryan, et al.
Published: (2025)
Aligning Inductive Bias for Data-Efficient Generalization in State Space Models
by: Chen, Qiyu, et al.
Published: (2025)
by: Chen, Qiyu, et al.
Published: (2025)
Ambiguous Online Learning
by: Kosoy, Vanessa
Published: (2025)
by: Kosoy, Vanessa
Published: (2025)
Adaptive Discretization in Online Reinforcement Learning
by: Sinclair, Sean R., et al.
Published: (2021)
by: Sinclair, Sean R., et al.
Published: (2021)
Regret Bounds for Robust Online Decision Making
by: Appel, Alexander, et al.
Published: (2025)
by: Appel, Alexander, et al.
Published: (2025)
Superior Scoring Rules for Probabilistic Evaluation of Single-Label Multi-Class Classification Tasks
by: Ahmadian, Rouhollah, et al.
Published: (2024)
by: Ahmadian, Rouhollah, et al.
Published: (2024)
Stringological sequence prediction I: efficient algorithms for predicting highly repetitive sequences
by: Kosoy, Vanessa
Published: (2026)
by: Kosoy, Vanessa
Published: (2026)
CELI: Controller-Embedded Language Model Interactions
by: Wagner, Jan-Samuel, et al.
Published: (2024)
by: Wagner, Jan-Samuel, et al.
Published: (2024)
Murphys Laws of AI Alignment: Why the Gap Always Wins
by: Gaikwad, Madhava
Published: (2025)
by: Gaikwad, Madhava
Published: (2025)
Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization
by: Li, Yin
Published: (2025)
by: Li, Yin
Published: (2025)
Separate Before You Compress: The WWHO Tokenization Architecture
by: Darshana, Kusal
Published: (2026)
by: Darshana, Kusal
Published: (2026)
Tricks and Plug-ins for Gradient Boosting with Transformers
by: Fang, Biyi, et al.
Published: (2025)
by: Fang, Biyi, et al.
Published: (2025)
OpCode-Based Malware Classification Using Machine Learning and Deep Learning Techniques
by: Saini, Varij, et al.
Published: (2025)
by: Saini, Varij, et al.
Published: (2025)
Batched Nonparametric Bandits via k-Nearest Neighbor UCB
by: Arya, Sakshi
Published: (2025)
by: Arya, Sakshi
Published: (2025)
Emotion-Inspired Learning Signals (EILS): A Homeostatic Framework for Adaptive Autonomous Agents
by: Tiwari, Dhruv
Published: (2025)
by: Tiwari, Dhruv
Published: (2025)
Near-Optimal Consistency-Robustness Trade-Offs for Learning-Augmented Online Knapsack Problems
by: Daneshvaramoli, Mohammadreza, et al.
Published: (2024)
by: Daneshvaramoli, Mohammadreza, et al.
Published: (2024)
Agnostic Learning under Targeted Poisoning: Optimal Rates and the Role of Randomness
by: Chornomaz, Bogdan, et al.
Published: (2025)
by: Chornomaz, Bogdan, et al.
Published: (2025)
LAWS: Learning from Actual Workloads Symbolically -- A Self-Certifying Parametrized Cache Architecture for Neural Inference, Robotics, and Edge Deployment
by: Magarshak, Gregory
Published: (2026)
by: Magarshak, Gregory
Published: (2026)
Binarized Neural Networks Converge Toward Algorithmic Simplicity: Empirical Support for the Learning-as-Compression Hypothesis
by: Sakabe, Eduardo Y., et al.
Published: (2025)
by: Sakabe, Eduardo Y., et al.
Published: (2025)
Algorithmic Analysis of Dense Associative Memory: Finite-Size Guarantees and Adversarial Robustness
by: Gaikwad, Madhava
Published: (2026)
by: Gaikwad, Madhava
Published: (2026)
Mitigating Catastrophic Forgetting in Streaming Generative and Predictive Learning via Stateful Replay
by: Du, Wenzhang
Published: (2025)
by: Du, Wenzhang
Published: (2025)
InhibiDistilbert: Knowledge Distillation for a ReLU and Addition-based Transformer
by: Zhang, Tony, et al.
Published: (2025)
by: Zhang, Tony, et al.
Published: (2025)
Think Thrice Before You Speak: Dual knowledge-enhanced Theory-of-Mind Reasoning for Persuasive Agents
by: Ma, Minghui, et al.
Published: (2026)
by: Ma, Minghui, et al.
Published: (2026)
Computational Hardness of Reinforcement Learning with Partial $q^π$-Realizability
by: Karimi, Shayan, et al.
Published: (2025)
by: Karimi, Shayan, et al.
Published: (2025)
Margin in Abstract Spaces
by: Ashlagi, Yair, et al.
Published: (2026)
by: Ashlagi, Yair, et al.
Published: (2026)
Reinforcement Learning in MDPs with Information-Ordered Policies
by: Zhang, Zhongjun, et al.
Published: (2025)
by: Zhang, Zhongjun, et al.
Published: (2025)
Practical Quantum CIM Empowerment via All-Domestic-Core Agentic Large Model
by: Rui, Wang, et al.
Published: (2026)
by: Rui, Wang, et al.
Published: (2026)
Graded Transformers
by: Shaska Sr, Tony
Published: (2025)
by: Shaska Sr, Tony
Published: (2025)
Similar Items
-
Multi-Teacher Ensemble Distillation: A Mathematical Framework for Probability-Domain Knowledge Aggregation
by: Flouro, Aaron R., et al.
Published: (2026) -
Recursive Meta-Distillation: An Axiomatic Framework for Iterative Knowledge Refinement
by: Flouro, Aaron R., et al.
Published: (2026) -
Hallucinations Live in Variance
by: Flouro, Aaron R., et al.
Published: (2026) -
Adaptive Weighting in Knowledge Distillation: An Axiomatic Framework for Multi-Scale Teacher Ensemble Optimization
by: Flouro, Aaron R., et al.
Published: (2026) -
Closed-Form Beta Distribution Estimation from Sparse Statistics with Random Forest Implicit Regularization
by: Landers, Jonathan R.
Published: (2025)