Transferring Knowledge from Large Foundation Models to Small Downstream Models
Fuente:
arXiv
Saved in:
| Main Authors: | Qiu, Shikai, Han, Boran, Maddix, Danielle C., Zhang, Shuai, Wang, Yuyang, Wilson, Andrew Gordon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding the Implicit Biases of Design Choices for Time Series Foundation Models
by: Yu, Annan, et al.
Published: (2025)
by: Yu, Annan, et al.
Published: (2025)
Enhancing Foundation Models for Time Series Forecasting via Wavelet-based Tokenization
by: Masserano, Luca, et al.
Published: (2024)
by: Masserano, Luca, et al.
Published: (2024)
Large Language Models Are Zero-Shot Time Series Forecasters
by: Gruver, Nate, et al.
Published: (2023)
by: Gruver, Nate, et al.
Published: (2023)
Gradient-Free Generation for Hard-Constrained Systems
by: Cheng, Chaoran, et al.
Published: (2024)
by: Cheng, Chaoran, et al.
Published: (2024)
Understanding Transformers for Time Series: Rank Structure, Flow-of-ranks, and Compressibility
by: Yu, Annan, et al.
Published: (2025)
by: Yu, Annan, et al.
Published: (2025)
Comparing and Contrasting DLWP Backbones on Navier-Stokes and Atmospheric Dynamics
by: Karlbauer, Matthias, et al.
Published: (2024)
by: Karlbauer, Matthias, et al.
Published: (2024)
Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across Scales
by: Qiu, Shikai, et al.
Published: (2025)
by: Qiu, Shikai, et al.
Published: (2025)
Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation Models
by: Zhang, Xiyuan, et al.
Published: (2025)
by: Zhang, Xiyuan, et al.
Published: (2025)
Discovering Bias in Latent Space: An Unsupervised Debiasing Approach
by: Adila, Dyah, et al.
Published: (2024)
by: Adila, Dyah, et al.
Published: (2024)
Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay
by: Marek, Martin, et al.
Published: (2026)
by: Marek, Martin, et al.
Published: (2026)
Bridging Remote Sensors with Multisensor Geospatial Foundation Models
by: Han, Boran, et al.
Published: (2024)
by: Han, Boran, et al.
Published: (2024)
Theoretical Guarantees of Learning Ensembling Strategies with Applications to Time Series Forecasting
by: Hasson, Hilaf, et al.
Published: (2023)
by: Hasson, Hilaf, et al.
Published: (2023)
Compute Better Spent: Replacing Dense Layers with Structured Matrices
by: Qiu, Shikai, et al.
Published: (2024)
by: Qiu, Shikai, et al.
Published: (2024)
Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks
by: Qiu, Shikai, et al.
Published: (2025)
by: Qiu, Shikai, et al.
Published: (2025)
CrEst: Credibility Estimation for Contexts in LLMs via Weak Supervision
by: Adila, Dyah, et al.
Published: (2025)
by: Adila, Dyah, et al.
Published: (2025)
In-Context Clustering with Large Language Models
by: Wang, Ying, et al.
Published: (2025)
by: Wang, Ying, et al.
Published: (2025)
SenTSR-Bench: Thinking with Injected Knowledge for Time-Series Reasoning
by: He, Zelin, et al.
Published: (2026)
by: He, Zelin, et al.
Published: (2026)
When Does Multimodality Lead to Better Time Series Forecasting?
by: Zhang, Xiyuan, et al.
Published: (2025)
by: Zhang, Xiyuan, et al.
Published: (2025)
Using Uncertainty Quantification to Characterize and Improve Out-of-Domain Learning for PDEs
by: Mouli, S. Chandra, et al.
Published: (2024)
by: Mouli, S. Chandra, et al.
Published: (2024)
Adapting to Online Distribution Shifts in Deep Learning: A Black-Box Approach
by: Baby, Dheeraj, et al.
Published: (2025)
by: Baby, Dheeraj, et al.
Published: (2025)
Customizing the Inductive Biases of Softmax Attention using Structured Matrices
by: Kuang, Yilun, et al.
Published: (2025)
by: Kuang, Yilun, et al.
Published: (2025)
End-to-End Probabilistic Framework for Learning with Hard Constraints
by: Utkarsh, Utkarsh, et al.
Published: (2025)
by: Utkarsh, Utkarsh, et al.
Published: (2025)
From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence
by: Finzi, Marc, et al.
Published: (2026)
by: Finzi, Marc, et al.
Published: (2026)
Efficient Table Retrieval and Understanding with Multimodal Large Language Models
by: Xu, Zhuoyan, et al.
Published: (2026)
by: Xu, Zhuoyan, et al.
Published: (2026)
Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific Models
by: Vemulapalli, Raviteja, et al.
Published: (2023)
by: Vemulapalli, Raviteja, et al.
Published: (2023)
Attacking Attention of Foundation Models Disrupts Downstream Tasks
by: Silva, Hondamunige Prasanna, et al.
Published: (2025)
by: Silva, Hondamunige Prasanna, et al.
Published: (2025)
How Well Do Large-Scale Chemical Language Models Transfer to Downstream Tasks?
by: Sagawa, Tatsuya, et al.
Published: (2026)
by: Sagawa, Tatsuya, et al.
Published: (2026)
Harvesting AlphaEarth: Benchmarking the Geospatial Foundation Model for Agricultural Downstream Tasks
by: Ma, Yuchi, et al.
Published: (2025)
by: Ma, Yuchi, et al.
Published: (2025)
Out-of-Distribution Detection Methods Answer the Wrong Questions
by: Li, Yucen Lily, et al.
Published: (2025)
by: Li, Yucen Lily, et al.
Published: (2025)
Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful
by: Marek, Martin, et al.
Published: (2025)
by: Marek, Martin, et al.
Published: (2025)
VocabTailor: Dynamic Vocabulary Selection for Downstream Tasks in Small Language Models
by: Zhang, Hanling, et al.
Published: (2025)
by: Zhang, Hanling, et al.
Published: (2025)
Efficiently Generating Correlated Sample Paths from Multi-step Time Series Foundation Models
by: Baron, Ethan, et al.
Published: (2025)
by: Baron, Ethan, et al.
Published: (2025)
Deep Learning is Not So Mysterious or Different
by: Wilson, Andrew Gordon
Published: (2025)
by: Wilson, Andrew Gordon
Published: (2025)
A More Realistic Evaluation of Cross-Frequency Transfer Learning and Foundation Forecasting Models
by: Olivares, Kin G., et al.
Published: (2025)
by: Olivares, Kin G., et al.
Published: (2025)
Transferable Adversarial Attacks on SAM and Its Downstream Models
by: Xia, Song, et al.
Published: (2024)
by: Xia, Song, et al.
Published: (2024)
EEGFormer: Towards Transferable and Interpretable Large-Scale EEG Foundation Model
by: Chen, Yuqi, et al.
Published: (2024)
by: Chen, Yuqi, et al.
Published: (2024)
Learning Cross-Domain Representations for Transferable Drug Perturbations on Single-Cell Transcriptional Responses
by: Liu, Hui, et al.
Published: (2024)
by: Liu, Hui, et al.
Published: (2024)
Downstream-Pretext Domain Knowledge Traceback for Active Learning
by: Zhang, Beichen, et al.
Published: (2024)
by: Zhang, Beichen, et al.
Published: (2024)
Feature Alignment and Representation Transfer in Knowledge Distillation for Large Language Models
by: Yang, Junjie, et al.
Published: (2025)
by: Yang, Junjie, et al.
Published: (2025)
Searching for Efficient Linear Layers over a Continuous Space of Structured Matrices
by: Potapczynski, Andres, et al.
Published: (2024)
by: Potapczynski, Andres, et al.
Published: (2024)
Similar Items
-
Understanding the Implicit Biases of Design Choices for Time Series Foundation Models
by: Yu, Annan, et al.
Published: (2025) -
Enhancing Foundation Models for Time Series Forecasting via Wavelet-based Tokenization
by: Masserano, Luca, et al.
Published: (2024) -
Large Language Models Are Zero-Shot Time Series Forecasters
by: Gruver, Nate, et al.
Published: (2023) -
Gradient-Free Generation for Hard-Constrained Systems
by: Cheng, Chaoran, et al.
Published: (2024) -
Understanding Transformers for Time Series: Rank Structure, Flow-of-ranks, and Compressibility
by: Yu, Annan, et al.
Published: (2025)