Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling
Fuente:
arXiv
Salvato in:
| Autori principali: | Goyal, Sachin, Lopez-Paz, David, Ahuja, Kartik |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries
di: Mahajan, Divyat, et al.
Pubblicazione: (2025)
di: Mahajan, Divyat, et al.
Pubblicazione: (2025)
Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification
di: Donhauser, Konstantin, et al.
Pubblicazione: (2025)
di: Donhauser, Konstantin, et al.
Pubblicazione: (2025)
Mode-Conditioning Unlocks Superior Test-Time Scaling
di: Wu, Chen Henry, et al.
Pubblicazione: (2025)
di: Wu, Chen Henry, et al.
Pubblicazione: (2025)
In-Context Learning through the Bayesian Prism
di: Panwar, Madhur, et al.
Pubblicazione: (2023)
di: Panwar, Madhur, et al.
Pubblicazione: (2023)
On Provable Length and Compositional Generalization
di: Ahuja, Kartik, et al.
Pubblicazione: (2024)
di: Ahuja, Kartik, et al.
Pubblicazione: (2024)
DRoP: Distributionally Robust Data Pruning
di: Vysogorets, Artem, et al.
Pubblicazione: (2024)
di: Vysogorets, Artem, et al.
Pubblicazione: (2024)
Interventional Causal Representation Learning
di: Ahuja, Kartik, et al.
Pubblicazione: (2022)
di: Ahuja, Kartik, et al.
Pubblicazione: (2022)
When Should We Introduce Safety Interventions During Pretraining?
di: Sam, Dylan, et al.
Pubblicazione: (2026)
di: Sam, Dylan, et al.
Pubblicazione: (2026)
Retrieval Mechanisms Surpass Long-Context Scaling in Time Series Forecasting
di: Ahuja, Rishi, et al.
Pubblicazione: (2026)
di: Ahuja, Rishi, et al.
Pubblicazione: (2026)
Scaling Laws for Data Filtering -- Data Curation cannot be Compute Agnostic
di: Goyal, Sachin, et al.
Pubblicazione: (2024)
di: Goyal, Sachin, et al.
Pubblicazione: (2024)
Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting
di: Watts, Ishaan, et al.
Pubblicazione: (2026)
di: Watts, Ishaan, et al.
Pubblicazione: (2026)
Pretrain-Test Task Alignment Governs Generalization in In-Context Learning
di: Letey, Mary I., et al.
Pubblicazione: (2025)
di: Letey, Mary I., et al.
Pubblicazione: (2025)
Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance
di: Goyal, Sachin, et al.
Pubblicazione: (2024)
di: Goyal, Sachin, et al.
Pubblicazione: (2024)
ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
di: Gupta, Sharut, et al.
Pubblicazione: (2026)
di: Gupta, Sharut, et al.
Pubblicazione: (2026)
Scaling Atomistic Protein Binder Design with Generative Pretraining and Test-Time Compute
di: Didi, Kieran, et al.
Pubblicazione: (2026)
di: Didi, Kieran, et al.
Pubblicazione: (2026)
Data-Efficient Operator Learning via Unsupervised Pretraining and In-Context Learning
di: Chen, Wuyang, et al.
Pubblicazione: (2024)
di: Chen, Wuyang, et al.
Pubblicazione: (2024)
A Granular Study of Safety Pretraining under Model Abliteration
di: Agnihotri, Shashank, et al.
Pubblicazione: (2025)
di: Agnihotri, Shashank, et al.
Pubblicazione: (2025)
TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining
di: Li, Jeffrey, et al.
Pubblicazione: (2025)
di: Li, Jeffrey, et al.
Pubblicazione: (2025)
Leveraging Pretrained Language Models as Energy Functions for Glauber Dynamics Text Diffusion
di: Kathuria, Tarun, et al.
Pubblicazione: (2026)
di: Kathuria, Tarun, et al.
Pubblicazione: (2026)
Safety Pretraining: Toward the Next Generation of Safe AI
di: Maini, Pratyush, et al.
Pubblicazione: (2025)
di: Maini, Pratyush, et al.
Pubblicazione: (2025)
Narrowing the Focus: Learned Optimizers for Pretrained Models
di: Kristiansen, Gus, et al.
Pubblicazione: (2024)
di: Kristiansen, Gus, et al.
Pubblicazione: (2024)
Towards efficient representation identification in supervised learning
di: Ahuja, Kartik, et al.
Pubblicazione: (2022)
di: Ahuja, Kartik, et al.
Pubblicazione: (2022)
Semantic Differentiation for Tackling Challenges in Watermarking Low-Entropy Constrained Generation Outputs
di: Le, Nghia T., et al.
Pubblicazione: (2026)
di: Le, Nghia T., et al.
Pubblicazione: (2026)
In-Context Data Distillation with TabPFN
di: Ma, Junwei, et al.
Pubblicazione: (2024)
di: Ma, Junwei, et al.
Pubblicazione: (2024)
Test-Time Efficient Pretrained Model Portfolios for Time Series Forecasting
di: Kayaalp, Mert, et al.
Pubblicazione: (2025)
di: Kayaalp, Mert, et al.
Pubblicazione: (2025)
Operationalizing Quantized Disentanglement
di: Barin-Pacela, Vitoria, et al.
Pubblicazione: (2025)
di: Barin-Pacela, Vitoria, et al.
Pubblicazione: (2025)
In-Context Curiosity: Distilling Exploration for Decision-Pretrained Transformers on Bandit Tasks
di: Yang, Huitao, et al.
Pubblicazione: (2025)
di: Yang, Huitao, et al.
Pubblicazione: (2025)
HEART: Emotionally-Driven Test-Time Scaling of Language Models
di: Pinto, Gabriela, et al.
Pubblicazione: (2025)
di: Pinto, Gabriela, et al.
Pubblicazione: (2025)
To Memorize or to Retrieve: Scaling Laws for RAG-Considerate Pretraining
di: Singh, Karan, et al.
Pubblicazione: (2026)
di: Singh, Karan, et al.
Pubblicazione: (2026)
Empowering Time Series Analysis with Large-Scale Multimodal Pretraining
di: Chen, Peng, et al.
Pubblicazione: (2026)
di: Chen, Peng, et al.
Pubblicazione: (2026)
MAP's not dead yet: Uncovering true language model modes by conditioning away degeneracy
di: Yoshida, Davis, et al.
Pubblicazione: (2023)
di: Yoshida, Davis, et al.
Pubblicazione: (2023)
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
di: Bethune, Louis, et al.
Pubblicazione: (2025)
di: Bethune, Louis, et al.
Pubblicazione: (2025)
Learning with Incomplete Context: Linear Contextual Bandits with Pretrained Imputation
di: Yan, Hao, et al.
Pubblicazione: (2025)
di: Yan, Hao, et al.
Pubblicazione: (2025)
United We Pretrain, Divided We Fail! Representation Learning for Time Series by Pretraining on 75 Datasets at Once
di: Kraus, Maurice, et al.
Pubblicazione: (2024)
di: Kraus, Maurice, et al.
Pubblicazione: (2024)
Latent Diffusion Pretraining for Crystal Property Prediction
di: Mukherjee, Shrimon, et al.
Pubblicazione: (2026)
di: Mukherjee, Shrimon, et al.
Pubblicazione: (2026)
NuTime: Numerically Multi-Scaled Embedding for Large-Scale Time-Series Pretraining
di: Lin, Chenguo, et al.
Pubblicazione: (2023)
di: Lin, Chenguo, et al.
Pubblicazione: (2023)
Exact Unlearning of Finetuning Data via Model Merging at Scale
di: Kuo, Kevin, et al.
Pubblicazione: (2025)
di: Kuo, Kevin, et al.
Pubblicazione: (2025)
STEP: Scientific Time-Series Encoder Pretraining via Cross-Domain Distillation
di: Zhang, Chen, et al.
Pubblicazione: (2026)
di: Zhang, Chen, et al.
Pubblicazione: (2026)
Context-Scaling versus Task-Scaling in In-Context Learning
di: Abedsoltan, Amirhesam, et al.
Pubblicazione: (2024)
di: Abedsoltan, Amirhesam, et al.
Pubblicazione: (2024)
Extending Test-Time Scaling: A 3D Perspective with Context, Batch, and Turn
di: Yu, Chao, et al.
Pubblicazione: (2025)
di: Yu, Chao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries
di: Mahajan, Divyat, et al.
Pubblicazione: (2025) -
Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification
di: Donhauser, Konstantin, et al.
Pubblicazione: (2025) -
Mode-Conditioning Unlocks Superior Test-Time Scaling
di: Wu, Chen Henry, et al.
Pubblicazione: (2025) -
In-Context Learning through the Bayesian Prism
di: Panwar, Madhur, et al.
Pubblicazione: (2023) -
On Provable Length and Compositional Generalization
di: Ahuja, Kartik, et al.
Pubblicazione: (2024)