Saved in:
| Main Authors: | Farhat, Sean, Chen, Deming |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2404.03263 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Practical Insights into Knowledge Distillation for Pre-Trained Models
by: Alballa, Norah, et al.
Published: (2024)
by: Alballa, Norah, et al.
Published: (2024)
The Surprising Ineffectiveness of Pre-Trained Visual Representations for Model-Based Reinforcement Learning
by: Schneider, Moritz, et al.
Published: (2024)
by: Schneider, Moritz, et al.
Published: (2024)
Pre-Training Graph Contrastive Masked Autoencoders are Strong Distillers for EEG
by: Wei, Xinxu, et al.
Published: (2024)
by: Wei, Xinxu, et al.
Published: (2024)
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
by: Cazenavette, George, et al.
Published: (2025)
by: Cazenavette, George, et al.
Published: (2025)
World Model Robustness via Surprise Recognition
by: Zollicoffer, Geigh, et al.
Published: (2025)
by: Zollicoffer, Geigh, et al.
Published: (2025)
The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
by: Akyürek, Ekin, et al.
Published: (2024)
by: Akyürek, Ekin, et al.
Published: (2024)
SR-TTT: Surprisal-Aware Residual Test-Time Training
by: P, Swamynathan V
Published: (2026)
by: P, Swamynathan V
Published: (2026)
Toward a Graph Foundation Model: Pre-Training Transformers With Random Walks
by: Tang, Ziyuan, et al.
Published: (2025)
by: Tang, Ziyuan, et al.
Published: (2025)
SpecMemo: Speculative Decoding is in Your Pocket
by: Yildirim, Selin, et al.
Published: (2025)
by: Yildirim, Selin, et al.
Published: (2025)
PTMs-TSCIL Pre-Trained Models Based Class-Incremental Learning
by: Wu, Yuanlong, et al.
Published: (2025)
by: Wu, Yuanlong, et al.
Published: (2025)
A Survey on Time-Series Pre-Trained Models
by: Ma, Qianli, et al.
Published: (2023)
by: Ma, Qianli, et al.
Published: (2023)
IMU-1: Sample-Efficient Pre-training of Small Language Models
by: Grigorev, George
Published: (2026)
by: Grigorev, George
Published: (2026)
Annealing Self-Distillation Rectification Improves Adversarial Training
by: Wu, Yu-Yu, et al.
Published: (2023)
by: Wu, Yu-Yu, et al.
Published: (2023)
The Surprising Effectiveness of Skip-Tuning in Diffusion Sampling
by: Ma, Jiajun, et al.
Published: (2024)
by: Ma, Jiajun, et al.
Published: (2024)
On Surprising Effectiveness of Masking Updates in Adaptive Optimizers
by: Joo, Taejong, et al.
Published: (2026)
by: Joo, Taejong, et al.
Published: (2026)
Winning Big with Small Models: Knowledge Distillation vs. Self-Training for Reducing Hallucination in Product QA Agents
by: Lewis, Ashley, et al.
Published: (2025)
by: Lewis, Ashley, et al.
Published: (2025)
Relational Learning in Pre-Trained Models: A Theory from Hypergraph Recovery Perspective
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities
by: Chandhok, Shivam, et al.
Published: (2025)
by: Chandhok, Shivam, et al.
Published: (2025)
CyclicFL: A Cyclic Model Pre-Training Approach to Efficient Federated Learning
by: Zhang, Pengyu, et al.
Published: (2023)
by: Zhang, Pengyu, et al.
Published: (2023)
Ensemble of Pre-Trained Models for Long-Tailed Trajectory Prediction
by: Thuremella, Divya, et al.
Published: (2025)
by: Thuremella, Divya, et al.
Published: (2025)
Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs
by: Yin, Lu, et al.
Published: (2023)
by: Yin, Lu, et al.
Published: (2023)
Contrastive Language-Image Pre-Training Model based Semantic Communication Performance Optimization
by: Yang, Shaoran, et al.
Published: (2025)
by: Yang, Shaoran, et al.
Published: (2025)
Surprise-Adaptive Intrinsic Motivation for Unsupervised Reinforcement Learning
by: Hugessen, Adriana, et al.
Published: (2024)
by: Hugessen, Adriana, et al.
Published: (2024)
SeqFusion: Sequential Fusion of Pre-Trained Models for Zero-Shot Time-Series Forecasting
by: Huang, Ting-Ji, et al.
Published: (2025)
by: Huang, Ting-Ji, et al.
Published: (2025)
Small Models, Smarter Learning: The Power of Joint Task Training
by: Both, Csaba, et al.
Published: (2025)
by: Both, Csaba, et al.
Published: (2025)
Combining Pre-Trained Models for Enhanced Feature Representation in Reinforcement Learning
by: Piccoli, Elia, et al.
Published: (2025)
by: Piccoli, Elia, et al.
Published: (2025)
Kakugo: Distillation of Low-Resource Languages into Small Language Models
by: Devine, Peter, et al.
Published: (2026)
by: Devine, Peter, et al.
Published: (2026)
Beyond the Proxy: Trajectory-Distilled Guidance for Offline GFlowNet Training
by: Chen, Ruishuo, et al.
Published: (2025)
by: Chen, Ruishuo, et al.
Published: (2025)
Data Efficacy for Language Model Training
by: Dai, Yalun, et al.
Published: (2025)
by: Dai, Yalun, et al.
Published: (2025)
Surprisal Driven $k$-NN for Robust and Interpretable Nonparametric Learning
by: Banerjee, Amartya, et al.
Published: (2023)
by: Banerjee, Amartya, et al.
Published: (2023)
Recommending Pre-Trained Models for IoT Devices
by: Patil, Parth V., et al.
Published: (2024)
by: Patil, Parth V., et al.
Published: (2024)
Pretrained Vision-Language-Action Models are Surprisingly Resistant to Forgetting in Continual Learning
by: Liu, Huihan, et al.
Published: (2026)
by: Liu, Huihan, et al.
Published: (2026)
TB or Not TB: Coverage-Driven Direct Preference Optimization for Verilog Stimulus Generation
by: Nadimi, Bardia, et al.
Published: (2025)
by: Nadimi, Bardia, et al.
Published: (2025)
SmallThinker: A Family of Efficient Large Language Models Natively Trained for Local Deployment
by: Song, Yixin, et al.
Published: (2025)
by: Song, Yixin, et al.
Published: (2025)
Protecting Copyright of Medical Pre-trained Language Models: Training-Free Backdoor Model Watermarking
by: Kong, Cong, et al.
Published: (2024)
by: Kong, Cong, et al.
Published: (2024)
Actor-Critic based Online Data Mixing For Language Model Pre-Training
by: Ma, Jing, et al.
Published: (2025)
by: Ma, Jing, et al.
Published: (2025)
Topology Only Pre-Training: Towards Generalised Multi-Domain Graph Models
by: Davies, Alex O., et al.
Published: (2023)
by: Davies, Alex O., et al.
Published: (2023)
Analyzing Generalization in Pre-Trained Symbolic Regression
by: Voigt, Henrik, et al.
Published: (2025)
by: Voigt, Henrik, et al.
Published: (2025)
Value-Based Pre-Training with Downstream Feedback
by: Ke, Shuqi, et al.
Published: (2026)
by: Ke, Shuqi, et al.
Published: (2026)
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
by: Wu, Yecheng, et al.
Published: (2026)
by: Wu, Yecheng, et al.
Published: (2026)
Similar Items
-
Practical Insights into Knowledge Distillation for Pre-Trained Models
by: Alballa, Norah, et al.
Published: (2024) -
The Surprising Ineffectiveness of Pre-Trained Visual Representations for Model-Based Reinforcement Learning
by: Schneider, Moritz, et al.
Published: (2024) -
Pre-Training Graph Contrastive Masked Autoencoders are Strong Distillers for EEG
by: Wei, Xinxu, et al.
Published: (2024) -
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
by: Cazenavette, George, et al.
Published: (2025) -
World Model Robustness via Surprise Recognition
by: Zollicoffer, Geigh, et al.
Published: (2025)