Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Hübotter, Jonas, Wolf, Patrik, Shevchenko, Alexander, Jüni, Dennis, Krause, Andreas, Kur, Gil |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Local Mixtures of Experts: Essentially Free Test-Time Training via Model Merging
by: Bertolissi, Ryo, et al.
Published: (2025)
by: Bertolissi, Ryo, et al.
Published: (2025)
Probabilistic Artificial Intelligence
by: Krause, Andreas, et al.
Published: (2025)
by: Krause, Andreas, et al.
Published: (2025)
Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs
by: Hübotter, Jonas, et al.
Published: (2024)
by: Hübotter, Jonas, et al.
Published: (2024)
Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning
by: Hübotter, Jonas, et al.
Published: (2025)
by: Hübotter, Jonas, et al.
Published: (2025)
Active Few-Shot Fine-Tuning
by: Hübotter, Jonas, et al.
Published: (2024)
by: Hübotter, Jonas, et al.
Published: (2024)
Transductive Active Learning: Theory and Applications
by: Hübotter, Jonas, et al.
Published: (2024)
by: Hübotter, Jonas, et al.
Published: (2024)
DISCOVER: Automated Curricula for Sparse-Reward Reinforcement Learning
by: Diaz-Bone, Leander, et al.
Published: (2025)
by: Diaz-Bone, Leander, et al.
Published: (2025)
Aligning Language Models from User Interactions
by: Buening, Thomas Kleine, et al.
Published: (2026)
by: Buening, Thomas Kleine, et al.
Published: (2026)
Maximizing Prefix-Confidence at Test-Time Efficiently Improves Mathematical Reasoning
by: Otth, Matthias, et al.
Published: (2025)
by: Otth, Matthias, et al.
Published: (2025)
Test-Time Tuned Language Models Enable End-to-end De Novo Molecular Structure Generation from MS/MS Spectra
by: Mismetti, Laura, et al.
Published: (2025)
by: Mismetti, Laura, et al.
Published: (2025)
Toward Understanding BERT-Like Pre-Training for DNA Foundation Models
by: Liang, Chaoqi, et al.
Published: (2023)
by: Liang, Chaoqi, et al.
Published: (2023)
Understanding the Role of Training Data in Test-Time Scaling
by: Javanmard, Adel, et al.
Published: (2025)
by: Javanmard, Adel, et al.
Published: (2025)
Reinforcement Learning via Self-Distillation
by: Hübotter, Jonas, et al.
Published: (2026)
by: Hübotter, Jonas, et al.
Published: (2026)
Revisiting Knowledge Distillation: The Hidden Role of Dataset Size
by: Lanzillotta, Giulia, et al.
Published: (2025)
by: Lanzillotta, Giulia, et al.
Published: (2025)
Towards Graph Foundation Models: Training on Knowledge Graphs Enables Transferability to General Graphs
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
A Theoretical Analysis of Test-Driven Code Generation
by: Menet, Nicolas, et al.
Published: (2026)
by: Menet, Nicolas, et al.
Published: (2026)
KairosHope: A Next-Generation Time-Series Foundation Model for Specialized Classification via Dual-Memory Architecture
by: Balderas, Luis, et al.
Published: (2026)
by: Balderas, Luis, et al.
Published: (2026)
Less is More: Unlocking Specialization of Time Series Foundation Models via Structured Pruning
by: Zhao, Lifan, et al.
Published: (2025)
by: Zhao, Lifan, et al.
Published: (2025)
Test-time Offline Reinforcement Learning on Goal-related Experience
by: Bagatella, Marco, et al.
Published: (2025)
by: Bagatella, Marco, et al.
Published: (2025)
Transformers and Their Roles as Time Series Foundation Models
by: Wu, Dennis, et al.
Published: (2025)
by: Wu, Dennis, et al.
Published: (2025)
Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
by: Chen, Xingwu, et al.
Published: (2025)
by: Chen, Xingwu, et al.
Published: (2025)
Learning Safety Constraints for Large Language Models
by: Chen, Xin, et al.
Published: (2025)
by: Chen, Xin, et al.
Published: (2025)
Towards Neural Scaling Laws for Time Series Foundation Models
by: Yao, Qingren, et al.
Published: (2024)
by: Yao, Qingren, et al.
Published: (2024)
Toward a Graph Foundation Model: Pre-Training Transformers With Random Walks
by: Tang, Ziyuan, et al.
Published: (2025)
by: Tang, Ziyuan, et al.
Published: (2025)
Safe Exploration Using Bayesian World Models and Log-Barrier Optimization
by: As, Yarden, et al.
Published: (2024)
by: As, Yarden, et al.
Published: (2024)
Test-Time Training on Graphs with Large Language Models (LLMs)
by: Zhang, Jiaxin, et al.
Published: (2024)
by: Zhang, Jiaxin, et al.
Published: (2024)
Towards Self-Supervised Foundation Models for Critical Care Time Series
by: Jagd, Katja Naasunnguaq, et al.
Published: (2025)
by: Jagd, Katja Naasunnguaq, et al.
Published: (2025)
Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting
by: Rasul, Kashif, et al.
Published: (2023)
by: Rasul, Kashif, et al.
Published: (2023)
Towards Understanding the Influence of Training Samples on Explanations
by: Artelt, André, et al.
Published: (2024)
by: Artelt, André, et al.
Published: (2024)
Wisdom of Committee: Distilling from Foundation Model to Specialized Application Model
by: Liu, Zichang, et al.
Published: (2024)
by: Liu, Zichang, et al.
Published: (2024)
Majority Voting for Code Generation
by: Launer, Tim, et al.
Published: (2026)
by: Launer, Tim, et al.
Published: (2026)
Towards Theoretical Understandings of Self-Consuming Generative Models
by: Fu, Shi, et al.
Published: (2024)
by: Fu, Shi, et al.
Published: (2024)
TimeDiT: General-purpose Diffusion Transformers for Time Series Foundation Model
by: Cao, Defu, et al.
Published: (2024)
by: Cao, Defu, et al.
Published: (2024)
TiCT: A Synthetically Pre-Trained Foundation Model for Time Series Classification
by: Yeh, Chin-Chia Michael, et al.
Published: (2025)
by: Yeh, Chin-Chia Michael, et al.
Published: (2025)
Active Fine-Tuning of Multi-Task Policies
by: Bagatella, Marco, et al.
Published: (2024)
by: Bagatella, Marco, et al.
Published: (2024)
Leveraging Generic Time Series Foundation Models for EEG Classification
by: Gnassounou, Théo, et al.
Published: (2025)
by: Gnassounou, Théo, et al.
Published: (2025)
Test-Time Training Undermines Safety Guardrails
by: Antonelli, Simone, et al.
Published: (2026)
by: Antonelli, Simone, et al.
Published: (2026)
Test Time Training for Supervised Causal Learning
by: Deng, Zizhen, et al.
Published: (2026)
by: Deng, Zizhen, et al.
Published: (2026)
Distributed Risk-Sensitive Safety Filters for Uncertain Discrete-Time Systems
by: Lederer, Armin, et al.
Published: (2025)
by: Lederer, Armin, et al.
Published: (2025)
Adaptable Cardiovascular Disease Risk Prediction from Heterogeneous Data using Large Language Models
by: Lübeck, Frederike, et al.
Published: (2025)
by: Lübeck, Frederike, et al.
Published: (2025)
Similar Items
-
Local Mixtures of Experts: Essentially Free Test-Time Training via Model Merging
by: Bertolissi, Ryo, et al.
Published: (2025) -
Probabilistic Artificial Intelligence
by: Krause, Andreas, et al.
Published: (2025) -
Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs
by: Hübotter, Jonas, et al.
Published: (2024) -
Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning
by: Hübotter, Jonas, et al.
Published: (2025) -
Active Few-Shot Fine-Tuning
by: Hübotter, Jonas, et al.
Published: (2024)