Boomerang Distillation Enables Zero-Shot Model Size Interpolation
Fuente:
arXiv
Saved in:
| Main Authors: | Kangaslahti, Sara, Nayak, Nihal V., Geuter, Jonathan, Fumero, Marco, Locatello, Francesco, Alvarez-Melis, David |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Continuous Language Model Interpolation for Dynamic and Controllable Text Generation
by: Kangaslahti, Sara, et al.
Published: (2024)
by: Kangaslahti, Sara, et al.
Published: (2024)
Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
by: Geuter, Jonathan, et al.
Published: (2025)
by: Geuter, Jonathan, et al.
Published: (2025)
DDEQs: Distributional Deep Equilibrium Models through Wasserstein Gradient Flows
by: Geuter, Jonathan, et al.
Published: (2025)
by: Geuter, Jonathan, et al.
Published: (2025)
Navigating the Latent Space Dynamics of Neural Models
by: Fumero, Marco, et al.
Published: (2025)
by: Fumero, Marco, et al.
Published: (2025)
Out-of-Distribution Detection with Relative Angles
by: Demirel, Berker, et al.
Published: (2024)
by: Demirel, Berker, et al.
Published: (2024)
Stop Training for the Worst: Progressive Unmasking Accelerates Masked Diffusion Training
by: Kim, Jaeyeon, et al.
Published: (2026)
by: Kim, Jaeyeon, et al.
Published: (2026)
Statistical and structural identifiability in representation learning
by: Nelson, Walter, et al.
Published: (2026)
by: Nelson, Walter, et al.
Published: (2026)
A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn't)
by: Nayak, Nihal V., et al.
Published: (2026)
by: Nayak, Nihal V., et al.
Published: (2026)
Latent Functional Maps: a spectral framework for representation alignment
by: Fumero, Marco, et al.
Published: (2024)
by: Fumero, Marco, et al.
Published: (2024)
Latent Space Translation via Inverse Relative Projection
by: Maiorca, Valentino, et al.
Published: (2024)
by: Maiorca, Valentino, et al.
Published: (2024)
Unifying Causal Representation Learning with the Invariance Principle
by: Yao, Dingling, et al.
Published: (2024)
by: Yao, Dingling, et al.
Published: (2024)
MorphGen: Controllable and Morphologically Plausible Generative Cell-Imaging
by: Demirel, Berker, et al.
Published: (2025)
by: Demirel, Berker, et al.
Published: (2025)
Connecting Neural Models Latent Geometries with Relative Geodesic Representations
by: Yu, Hanlin, et al.
Published: (2025)
by: Yu, Hanlin, et al.
Published: (2025)
Learning to Generate Instruction Tuning Datasets for Zero-Shot Task Adaptation
by: Nayak, Nihal V., et al.
Published: (2024)
by: Nayak, Nihal V., et al.
Published: (2024)
Latent Space Translation via Semantic Alignment
by: Maiorca, Valentino, et al.
Published: (2023)
by: Maiorca, Valentino, et al.
Published: (2023)
RoBoN: Routed Online Best-of-n for Test-Time Scaling with Multiple LLMs
by: Geuter, Jonathan, et al.
Published: (2025)
by: Geuter, Jonathan, et al.
Published: (2025)
Hidden Breakthroughs in Language Model Training
by: Kangaslahti, Sara, et al.
Published: (2025)
by: Kangaslahti, Sara, et al.
Published: (2025)
Distributional Dataset Distillation with Subtask Decomposition
by: Qin, Tian, et al.
Published: (2024)
by: Qin, Tian, et al.
Published: (2024)
Causal Learning with the Invariance Principle
by: Montagna, Francesco, et al.
Published: (2026)
by: Montagna, Francesco, et al.
Published: (2026)
Controlling Transient Amplification Improves Long-horizon Rollouts
by: Pervez, Adeel, et al.
Published: (2026)
by: Pervez, Adeel, et al.
Published: (2026)
The Rate-Distortion-Polysemanticity Tradeoff in SAEs
by: Mencattini, Tommaso, et al.
Published: (2026)
by: Mencattini, Tommaso, et al.
Published: (2026)
High-dimensional Analysis of Synthetic Data Selection
by: Rezaei, Parham, et al.
Published: (2025)
by: Rezaei, Parham, et al.
Published: (2025)
A Label is Worth a Thousand Images in Dataset Distillation
by: Qin, Tian, et al.
Published: (2024)
by: Qin, Tian, et al.
Published: (2024)
Fast Forwarding Low-Rank Training
by: Rahamim, Adir, et al.
Published: (2024)
by: Rahamim, Adir, et al.
Published: (2024)
Mechanistic PDE Networks for Discovery of Governing Equations
by: Pervez, Adeel, et al.
Published: (2025)
by: Pervez, Adeel, et al.
Published: (2025)
Addressing Instrument-Outcome Confounding in Mendelian Randomization through Representation Learning
by: Huang, Shimeng, et al.
Published: (2026)
by: Huang, Shimeng, et al.
Published: (2026)
Toward Identifiable Sparse Autoencoders
by: Nelson, Walter, et al.
Published: (2026)
by: Nelson, Walter, et al.
Published: (2026)
Marrying Causal Representation Learning with Dynamical Systems for Science
by: Yao, Dingling, et al.
Published: (2024)
by: Yao, Dingling, et al.
Published: (2024)
Learning Discrete Diffusion of Graphs via Free-Energy Gradient Flows
by: Rancati, Dario, et al.
Published: (2026)
by: Rancati, Dario, et al.
Published: (2026)
Understanding the Role of Functional Diversity in Weight-Ensembling with Ingredient Selection and Multidimensional Scaling
by: Rojas, Alex, et al.
Published: (2024)
by: Rojas, Alex, et al.
Published: (2024)
Strongly Isomorphic Neural Optimal Transport Across Incomparable Spaces
by: Sotiropoulou, Athina, et al.
Published: (2024)
by: Sotiropoulou, Athina, et al.
Published: (2024)
Inverse Depth Scaling From Most Layers Being Similar
by: Liu, Yizhou, et al.
Published: (2026)
by: Liu, Yizhou, et al.
Published: (2026)
Learning Explicit Single-Cell Dynamics Using ODE Representations
by: von Bassewitz, Jan-Philipp, et al.
Published: (2025)
by: von Bassewitz, Jan-Philipp, et al.
Published: (2025)
Exploratory Causal Inference in SAEnce
by: Mencattini, Tommaso, et al.
Published: (2025)
by: Mencattini, Tommaso, et al.
Published: (2025)
Demystifying amortized causal discovery with transformers
by: Montagna, Francesco, et al.
Published: (2024)
by: Montagna, Francesco, et al.
Published: (2024)
Leveraging Zero-Shot Prompting for Efficient Language Model Distillation
by: Vöge, Lukas, et al.
Published: (2024)
by: Vöge, Lukas, et al.
Published: (2024)
Few-Shot Inspired Generative Zero-Shot Learning
by: Shohag, Md Shakil Ahamed, et al.
Published: (2025)
by: Shohag, Md Shakil Ahamed, et al.
Published: (2025)
Efficient Fine-Tuning of Single-Cell Foundation Models Enables Zero-Shot Molecular Perturbation Prediction
by: Maleki, Sepideh, et al.
Published: (2024)
by: Maleki, Sepideh, et al.
Published: (2024)
Zero-Shot Robustification of Zero-Shot Models
by: Adila, Dyah, et al.
Published: (2023)
by: Adila, Dyah, et al.
Published: (2023)
Analyzing Political Text at Scale with Online Tensor LDA
by: Kangaslahti, Sara, et al.
Published: (2025)
by: Kangaslahti, Sara, et al.
Published: (2025)
Similar Items
-
Continuous Language Model Interpolation for Dynamic and Controllable Text Generation
by: Kangaslahti, Sara, et al.
Published: (2024) -
Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
by: Geuter, Jonathan, et al.
Published: (2025) -
DDEQs: Distributional Deep Equilibrium Models through Wasserstein Gradient Flows
by: Geuter, Jonathan, et al.
Published: (2025) -
Navigating the Latent Space Dynamics of Neural Models
by: Fumero, Marco, et al.
Published: (2025) -
Out-of-Distribution Detection with Relative Angles
by: Demirel, Berker, et al.
Published: (2024)