The Finetuner's Fallacy: When to Pretrain with Your Finetuning Data
Fuente:
arXiv
Saved in:
| Main Authors: | Baek, Christina, Monti, Ricardo Pio, Schwab, David, Abbas, Amro, Adiga, Rishabh, Blakeney, Cody, Böther, Maximilian, Burstein, Paul, Carranza, Aldo Gael, Deng, Alvin, Doshi, Parth, Dorna, Vineeth, Fang, Alex, Jiang, Tony, Joshi, Siddharth, Larsen, Brett W., Lee, Jason Chan, Mentzer, Katherine L., Merrick, Luke, Mongstad, Haakon, Pan, Fan, Suri, Anshuman, Teh, Darren, Telanoff, Jason, Urbanek, Jack, Wang, Zhengping, Wills, Josh, Yin, Haoli, Raghunathan, Aditi, Kolter, J. Zico, Gaza, Bogdan, Morcos, Ari, Leavitt, Matthew, Maini, Pratyush |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ÜberWeb: Insights from Multilingual Curation for a 20-Trillion-Token Dataset
by: DatologyAI, et al.
Published: (2026)
by: DatologyAI, et al.
Published: (2026)
DatBench: Discriminative, Faithful, and Efficient VLM Evaluations
by: DatologyAI, et al.
Published: (2026)
by: DatologyAI, et al.
Published: (2026)
Luxical: High-Speed Lexical-Dense Text Embeddings
by: DatologyAI, et al.
Published: (2025)
by: DatologyAI, et al.
Published: (2025)
20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone
by: DatologyAI, et al.
Published: (2026)
by: DatologyAI, et al.
Published: (2026)
BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
by: DatologyAI, et al.
Published: (2025)
by: DatologyAI, et al.
Published: (2025)
Finetuning CLIP to Reason about Pairwise Differences
by: Sam, Dylan, et al.
Published: (2024)
by: Sam, Dylan, et al.
Published: (2024)
Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance
by: Goyal, Sachin, et al.
Published: (2024)
by: Goyal, Sachin, et al.
Published: (2024)
OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics
by: Dorna, Vineeth, et al.
Published: (2025)
by: Dorna, Vineeth, et al.
Published: (2025)
Understanding Hallucinations in Diffusion Models through Mode Interpolation
by: Aithal, Sumukh K, et al.
Published: (2024)
by: Aithal, Sumukh K, et al.
Published: (2024)
When Should We Introduce Safety Interventions During Pretraining?
by: Sam, Dylan, et al.
Published: (2026)
by: Sam, Dylan, et al.
Published: (2026)
T-MARS: Improving Visual Representations by Circumventing Text Feature Learning
by: Maini, Pratyush, et al.
Published: (2023)
by: Maini, Pratyush, et al.
Published: (2023)
Scaling Laws for Data Filtering -- Data Curation cannot be Compute Agnostic
by: Goyal, Sachin, et al.
Published: (2024)
by: Goyal, Sachin, et al.
Published: (2024)
TOFU: A Task of Fictitious Unlearning for LLMs
by: Maini, Pratyush, et al.
Published: (2024)
by: Maini, Pratyush, et al.
Published: (2024)
Rethinking LLM Memorization through the Lens of Adversarial Compression
by: Schwarzschild, Avi, et al.
Published: (2024)
by: Schwarzschild, Avi, et al.
Published: (2024)
Mimetic Initialization of MLPs
by: Trockman, Asher, et al.
Published: (2026)
by: Trockman, Asher, et al.
Published: (2026)
AcceleratedLiNGAM: Learning Causal DAGs at the speed of GPUs
by: Akinwande, Victor, et al.
Published: (2024)
by: Akinwande, Victor, et al.
Published: (2024)
FUSE-ing Language Models: Zero-Shot Adapter Discovery for Prompt Optimization Across Tokenizers
by: Williams, Joshua Nathaniel, et al.
Published: (2024)
by: Williams, Joshua Nathaniel, et al.
Published: (2024)
Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoning
by: Huang, Benhao, et al.
Published: (2026)
by: Huang, Benhao, et al.
Published: (2026)
Why is SAM Robust to Label Noise?
by: Baek, Christina, et al.
Published: (2024)
by: Baek, Christina, et al.
Published: (2024)
Logits-Based Finetuning
by: Li, Jingyao, et al.
Published: (2025)
by: Li, Jingyao, et al.
Published: (2025)
Order Independence With Finetuning
by: Brown, Katrina, et al.
Published: (2025)
by: Brown, Katrina, et al.
Published: (2025)
When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method
by: Zhang, Biao, et al.
Published: (2024)
by: Zhang, Biao, et al.
Published: (2024)
Understanding Optimization in Deep Learning with Central Flows
by: Cohen, Jeremy M., et al.
Published: (2024)
by: Cohen, Jeremy M., et al.
Published: (2024)
Sparse Memory Finetuning as a Low-Forgetting Alternative to LoRA and Full Finetuning
by: Gupta, Prakhar, et al.
Published: (2026)
by: Gupta, Prakhar, et al.
Published: (2026)
Predicting the Performance of Black-box LLMs through Follow-up Queries
by: Sam, Dylan, et al.
Published: (2025)
by: Sam, Dylan, et al.
Published: (2025)
One-Step Diffusion Distillation via Deep Equilibrium Models
by: Geng, Zhengyang, et al.
Published: (2023)
by: Geng, Zhengyang, et al.
Published: (2023)
Diffusing Differentiable Representations
by: Savani, Yash, et al.
Published: (2024)
by: Savani, Yash, et al.
Published: (2024)
UICoder: Finetuning Large Language Models to Generate User Interface Code through Automated Feedback
by: Wu, Jason, et al.
Published: (2024)
by: Wu, Jason, et al.
Published: (2024)
On Finetuning Tabular Foundation Models
by: Rubachev, Ivan, et al.
Published: (2025)
by: Rubachev, Ivan, et al.
Published: (2025)
Learning Dynamics of LLM Finetuning
by: Ren, Yi, et al.
Published: (2024)
by: Ren, Yi, et al.
Published: (2024)
Orthogonal Finetuning Made Scalable
by: Qiu, Zeju, et al.
Published: (2025)
by: Qiu, Zeju, et al.
Published: (2025)
Representation Finetuning for Continual Learning
by: Luo, Haihua, et al.
Published: (2026)
by: Luo, Haihua, et al.
Published: (2026)
Improving Sparse Memory Finetuning
by: Goyal, Satyam, et al.
Published: (2026)
by: Goyal, Satyam, et al.
Published: (2026)
Iterative Finetuning is Mostly Idempotent
by: Roe, Zephaniah, et al.
Published: (2026)
by: Roe, Zephaniah, et al.
Published: (2026)
Exploring Representation Invariance in Finetuning
by: Zu, Wenqiang, et al.
Published: (2025)
by: Zu, Wenqiang, et al.
Published: (2025)
Predicting Emergent Capabilities by Finetuning
by: Snell, Charlie, et al.
Published: (2024)
by: Snell, Charlie, et al.
Published: (2024)
The Hallucination Tax of Reinforcement Finetuning
by: Song, Linxin, et al.
Published: (2025)
by: Song, Linxin, et al.
Published: (2025)
Whisper Finetuning on Nepali Language
by: Rijal, Sanjay, et al.
Published: (2024)
by: Rijal, Sanjay, et al.
Published: (2024)
Learning Dynamics of VLM Finetuning
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
Synthesizing Privacy-Preserving Text Data via Finetuning without Finetuning Billion-Scale LLMs
by: Tan, Bowen, et al.
Published: (2025)
by: Tan, Bowen, et al.
Published: (2025)
Similar Items
-
ÜberWeb: Insights from Multilingual Curation for a 20-Trillion-Token Dataset
by: DatologyAI, et al.
Published: (2026) -
DatBench: Discriminative, Faithful, and Efficient VLM Evaluations
by: DatologyAI, et al.
Published: (2026) -
Luxical: High-Speed Lexical-Dense Text Embeddings
by: DatologyAI, et al.
Published: (2025) -
20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone
by: DatologyAI, et al.
Published: (2026) -
BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
by: DatologyAI, et al.
Published: (2025)