Majority Kernels: An Approach to Leverage Big Model Dynamics for Efficient Small Model Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mazzawi, Hanna, Awasthi, Pranjal, Gonzalvo, Xavi, Ramalingam, Srikumar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Deep Fusion: Efficient Network Training via Pre-trained Initializations
von: Mazzawi, Hanna, et al.
Veröffentlicht: (2023)
von: Mazzawi, Hanna, et al.
Veröffentlicht: (2023)
Grow, Don't Overwrite: Fine-tuning Without Forgetting
von: Adila, Dyah, et al.
Veröffentlicht: (2026)
von: Adila, Dyah, et al.
Veröffentlicht: (2026)
On Distributed Larger-Than-Memory Subset Selection With Pairwise Submodular Functions
von: Böther, Maximilian, et al.
Veröffentlicht: (2024)
von: Böther, Maximilian, et al.
Veröffentlicht: (2024)
Transmuting prompts into weights
von: Mazzawi, Hanna, et al.
Veröffentlicht: (2025)
von: Mazzawi, Hanna, et al.
Veröffentlicht: (2025)
Learning without training: The implicit dynamics of in-context learning
von: Dherin, Benoit, et al.
Veröffentlicht: (2025)
von: Dherin, Benoit, et al.
Veröffentlicht: (2025)
How iteration order influences convergence and stability in deep learning
von: Dherin, Benoit, et al.
Veröffentlicht: (2025)
von: Dherin, Benoit, et al.
Veröffentlicht: (2025)
Learning by solving differential equations
von: Dherin, Benoit, et al.
Veröffentlicht: (2025)
von: Dherin, Benoit, et al.
Veröffentlicht: (2025)
The Limits of Preference Data for Post-Training
von: Zhao, Eric, et al.
Veröffentlicht: (2025)
von: Zhao, Eric, et al.
Veröffentlicht: (2025)
Leveraging GANs For Active Appearance Models Optimized Model Fitting
von: Awasthi, Anurag
Veröffentlicht: (2025)
von: Awasthi, Anurag
Veröffentlicht: (2025)
Sample-Efficient Optimization over Generative Priors via Coarse Learnability
von: Awasthi, Pranjal, et al.
Veröffentlicht: (2025)
von: Awasthi, Pranjal, et al.
Veröffentlicht: (2025)
Analyzing Similarity Metrics for Data Selection for Language Model Pretraining
von: Sam, Dylan, et al.
Veröffentlicht: (2025)
von: Sam, Dylan, et al.
Veröffentlicht: (2025)
Sample, Scrutinize and Scale: Effective Inference-Time Search by Scaling Verification
von: Zhao, Eric, et al.
Veröffentlicht: (2025)
von: Zhao, Eric, et al.
Veröffentlicht: (2025)
From Style to Facts: Mapping the Boundaries of Knowledge Injection with Finetuning
von: Zhao, Eric, et al.
Veröffentlicht: (2025)
von: Zhao, Eric, et al.
Veröffentlicht: (2025)
Agnostic Learning of General ReLU Activation Using Gradient Descent
von: Awasthi, Pranjal, et al.
Veröffentlicht: (2022)
von: Awasthi, Pranjal, et al.
Veröffentlicht: (2022)
Stacking as Accelerated Gradient Descent
von: Agarwal, Naman, et al.
Veröffentlicht: (2024)
von: Agarwal, Naman, et al.
Veröffentlicht: (2024)
Learning Neural Networks with Sparse Activations
von: Awasthi, Pranjal, et al.
Veröffentlicht: (2024)
von: Awasthi, Pranjal, et al.
Veröffentlicht: (2024)
Uncertainty-Aware Tabular Prediction: Evaluating VBLL-Enhanced TabPFN in Safety-Critical Medical Data
von: Ramalingam, Madhushan
Veröffentlicht: (2025)
von: Ramalingam, Madhushan
Veröffentlicht: (2025)
GIST: Greedy Independent Set Thresholding for Max-Min Diversification with Submodular Utility
von: Fahrbach, Matthew, et al.
Veröffentlicht: (2024)
von: Fahrbach, Matthew, et al.
Veröffentlicht: (2024)
Energy Efficient Protein Language Models: Leveraging Small Language Models with LoRA for Controllable Protein Generation
von: Shah, Aayush, et al.
Veröffentlicht: (2024)
von: Shah, Aayush, et al.
Veröffentlicht: (2024)
Language verY Rare for All
von: Merad, Ibrahim, et al.
Veröffentlicht: (2024)
von: Merad, Ibrahim, et al.
Veröffentlicht: (2024)
Big2Small: A Unifying Neural Network Framework for Model Compression
von: Liao, Jing-Xiao, et al.
Veröffentlicht: (2026)
von: Liao, Jing-Xiao, et al.
Veröffentlicht: (2026)
Learnware of Language Models: Specialized Small Language Models Can Do Big
von: Tan, Zhi-Hao, et al.
Veröffentlicht: (2025)
von: Tan, Zhi-Hao, et al.
Veröffentlicht: (2025)
BigMac: A Communication-Efficient Mixture-of-Experts Model Structure for Fast Training and Inference
von: Jin, Zewen, et al.
Veröffentlicht: (2025)
von: Jin, Zewen, et al.
Veröffentlicht: (2025)
Training-Free Generative Modeling via Kernelized Stochastic Interpolants
von: Coeurdoux, Florentin, et al.
Veröffentlicht: (2026)
von: Coeurdoux, Florentin, et al.
Veröffentlicht: (2026)
HQFS: Hybrid Quantum Classical Financial Security with VQC Forecasting, QUBO Annealing, and Audit-Ready Post-Quantum Signing
von: Nayak, Srikumar
Veröffentlicht: (2026)
von: Nayak, Srikumar
Veröffentlicht: (2026)
Named Entity Recognition for Payment Data Using NLP
von: Nayak, Srikumar
Veröffentlicht: (2026)
von: Nayak, Srikumar
Veröffentlicht: (2026)
Calibrated Credit Intelligence: Shift-Robust and Fair Risk Scoring with Bayesian Uncertainty and Gradient Boosting
von: Nayak, Srikumar
Veröffentlicht: (2026)
von: Nayak, Srikumar
Veröffentlicht: (2026)
Fractional Heat Kernel for Semi-Supervised Graph Learning with Small Training Sample Size
von: Bozorgnia, Farid, et al.
Veröffentlicht: (2025)
von: Bozorgnia, Farid, et al.
Veröffentlicht: (2025)
A Little Help Goes a Long Way: Efficient LLM Training by Leveraging Small LMs
von: Rawat, Ankit Singh, et al.
Veröffentlicht: (2024)
von: Rawat, Ankit Singh, et al.
Veröffentlicht: (2024)
Winning Big with Small Models: Knowledge Distillation vs. Self-Training for Reducing Hallucination in Product QA Agents
von: Lewis, Ashley, et al.
Veröffentlicht: (2025)
von: Lewis, Ashley, et al.
Veröffentlicht: (2025)
Efficient Parameter Optimisation for Quantum Kernel Alignment: A Sub-sampling Approach in Variational Training
von: Sahin, M. Emre, et al.
Veröffentlicht: (2024)
von: Sahin, M. Emre, et al.
Veröffentlicht: (2024)
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
Leveraging Kernel Symmetry for Joint Compression and Error Mitigation in Edge Model Transfer
von: Hamadouche, Anis, et al.
Veröffentlicht: (2026)
von: Hamadouche, Anis, et al.
Veröffentlicht: (2026)
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
RLShield: Practical Multi-Agent RL for Financial Cyber Defense with Attack-Surface MDPs and Real-Time Response Orchestration
von: Nayak, Srikumar
Veröffentlicht: (2026)
von: Nayak, Srikumar
Veröffentlicht: (2026)
A Post-Training Enhanced Optimization Approach for Small Language Models
von: Zhai, Keke
Veröffentlicht: (2024)
von: Zhai, Keke
Veröffentlicht: (2024)
Jarzynski Reweighting and Sampling Dynamics for Training Energy-Based Models: Theoretical Analysis of Different Transition Kernels
von: Carbone, Davide
Veröffentlicht: (2025)
von: Carbone, Davide
Veröffentlicht: (2025)
A Machine learning and Empirical Bayesian Approach for Predictive Buying in B2B E-commerce
von: De, Tuhin Subhra, et al.
Veröffentlicht: (2024)
von: De, Tuhin Subhra, et al.
Veröffentlicht: (2024)
MSfusion: A Dynamic Model Splitting Approach for Resource-Constrained Machines to Collaboratively Train Larger Models
von: Xie, Jin, et al.
Veröffentlicht: (2024)
von: Xie, Jin, et al.
Veröffentlicht: (2024)
Small Vocabularies, Big Gains: Pretraining and Tokenization in Time Series Models
von: Roger, Alexis, et al.
Veröffentlicht: (2025)
von: Roger, Alexis, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Deep Fusion: Efficient Network Training via Pre-trained Initializations
von: Mazzawi, Hanna, et al.
Veröffentlicht: (2023) -
Grow, Don't Overwrite: Fine-tuning Without Forgetting
von: Adila, Dyah, et al.
Veröffentlicht: (2026) -
On Distributed Larger-Than-Memory Subset Selection With Pairwise Submodular Functions
von: Böther, Maximilian, et al.
Veröffentlicht: (2024) -
Transmuting prompts into weights
von: Mazzawi, Hanna, et al.
Veröffentlicht: (2025) -
Learning without training: The implicit dynamics of in-context learning
von: Dherin, Benoit, et al.
Veröffentlicht: (2025)