Balancing Synthetic Data and Replay for Enhancing Task-Specific Capabilities

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Spiegelhalter, Urs, Franke, Jörg K. H., Hutter, Frank
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915552912474112
author Spiegelhalter, Urs
Franke, Jörg K. H.
Hutter, Frank
author_facet Spiegelhalter, Urs
Franke, Jörg K. H.
Hutter, Frank
contents Adapting language models to new tasks through continued pretraining faces a fundamental trade-off: models must learn new capabilities while avoiding catastrophic forgetting of existing knowledge. While prior work has studied synthetic data generation techniques, the optimal replay ratios for balancing task performance and knowledge retention under computational constraints remain poorly understood. We present a comprehensive empirical study investigating the interplay between replay ratio configuration and computational budget when adapting language models to new tasks. Using the bAbI reasoning tasks as our target objective, we apply synthetic data generation and systematically evaluate different total token budgets and replay ratio configurations. We analyze their effects on both task mastery and general knowledge retention. Our experiments reveal an optimal configuration that balances task-specific performance with general knowledge retention. Based on our findings, we provide empirically-grounded guidelines for selecting replay ratios based on computational budget, enabling practitioners to achieve strong task adaptation with significantly reduced training costs.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11842
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Balancing Synthetic Data and Replay for Enhancing Task-Specific Capabilities
Spiegelhalter, Urs
Franke, Jörg K. H.
Hutter, Frank
Machine Learning
Computation and Language
Adapting language models to new tasks through continued pretraining faces a fundamental trade-off: models must learn new capabilities while avoiding catastrophic forgetting of existing knowledge. While prior work has studied synthetic data generation techniques, the optimal replay ratios for balancing task performance and knowledge retention under computational constraints remain poorly understood. We present a comprehensive empirical study investigating the interplay between replay ratio configuration and computational budget when adapting language models to new tasks. Using the bAbI reasoning tasks as our target objective, we apply synthetic data generation and systematically evaluate different total token budgets and replay ratio configurations. We analyze their effects on both task mastery and general knowledge retention. Our experiments reveal an optimal configuration that balances task-specific performance with general knowledge retention. Based on our findings, we provide empirically-grounded guidelines for selecting replay ratios based on computational budget, enabling practitioners to achieve strong task adaptation with significantly reduced training costs.
title Balancing Synthetic Data and Replay for Enhancing Task-Specific Capabilities
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2510.11842