Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings
Fuente:
arXiv
Saved in:
| Main Authors: | Jeha, Paul, Sedova, Anastasiia, Béthune, Louis, Seto, Skyler, Frellsen, Jes, Ablin, Pierre, Schluter, Natalie |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Laws for Mixture Pretraining Under Data Constraints
by: Sedova, Anastasiia, et al.
Published: (2026)
by: Sedova, Anastasiia, et al.
Published: (2026)
Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions
by: Sedova, Anastasiia, et al.
Published: (2026)
by: Sedova, Anastasiia, et al.
Published: (2026)
Optimal Splitting of Language Models from Mixtures to Specialized Domains
by: Seto, Skyler, et al.
Published: (2026)
by: Seto, Skyler, et al.
Published: (2026)
Training Bilingual LMs with Data Constraints in the Targeted Language
by: Seto, Skyler, et al.
Published: (2024)
by: Seto, Skyler, et al.
Published: (2024)
Learning Energy-Based Models by Self-normalising the Likelihood
by: Senetaire, Hugo, et al.
Published: (2025)
by: Senetaire, Hugo, et al.
Published: (2025)
Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
by: Grangier, David, et al.
Published: (2024)
by: Grangier, David, et al.
Published: (2024)
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
by: Ablin, Pierre, et al.
Published: (2025)
by: Ablin, Pierre, et al.
Published: (2025)
Completed Hyperparameter Transfer across Modules, Width, Depth, Batch and Duration
by: Mlodozeniec, Bruno, et al.
Published: (2025)
by: Mlodozeniec, Bruno, et al.
Published: (2025)
Variance reduction of diffusion model's gradients with Taylor approximation-based control variate
by: Jeha, Paul, et al.
Published: (2024)
by: Jeha, Paul, et al.
Published: (2024)
Scaling Categorical Flow Maps
by: Davis, Oscar, et al.
Published: (2026)
by: Davis, Oscar, et al.
Published: (2026)
Noise-Response Calibration: A Causal Intervention Protocol for LLM-Judges
by: Khomiakov, Maxim, et al.
Published: (2026)
by: Khomiakov, Maxim, et al.
Published: (2026)
Assessing the Role of Data Quality in Training Bilingual Language Models
by: Seto, Skyler, et al.
Published: (2025)
by: Seto, Skyler, et al.
Published: (2025)
Zoom, Don't Wander: Why Regional Search Outperforms Pareto Reasoning and Global Optimization in Budget-Constrained SBSE
by: Ganguly, Kishan Kumar, et al.
Published: (2026)
by: Ganguly, Kishan Kumar, et al.
Published: (2026)
Supervised Guidance Training for Infinite-Dimensional Diffusion Models
by: Baker, Elizabeth L., et al.
Published: (2026)
by: Baker, Elizabeth L., et al.
Published: (2026)
Debiasing Guidance for Discrete Diffusion with Sequential Monte Carlo
by: Lee, Cheuk Kit, et al.
Published: (2025)
by: Lee, Cheuk Kit, et al.
Published: (2025)
ULF: Unsupervised Labeling Function Correction using Cross-Validation for Weak Supervision
by: Sedova, Anastasiia, et al.
Published: (2022)
by: Sedova, Anastasiia, et al.
Published: (2022)
Large Pre-Training Datasets Don't Always Guarantee Robustness after Fine-Tuning
by: Hwang, Jaedong, et al.
Published: (2024)
by: Hwang, Jaedong, et al.
Published: (2024)
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
by: Bethune, Louis, et al.
Published: (2025)
by: Bethune, Louis, et al.
Published: (2025)
Shielded Diffusion: Generating Novel and Diverse Images using Sparse Repellency
by: Kirchhof, Michael, et al.
Published: (2024)
by: Kirchhof, Michael, et al.
Published: (2024)
DynaMiCS: Fine-tuning LLMs with Performance Constraints using Dynamic Mixtures
by: Gualdoni, Eleonora, et al.
Published: (2026)
by: Gualdoni, Eleonora, et al.
Published: (2026)
An Isotropic Approach to Efficient Uncertainty Quantification with Gradient Norms
by: Grünefeld, Nils, et al.
Published: (2026)
by: Grünefeld, Nils, et al.
Published: (2026)
Internal-Coordinate Density Modelling of Protein Structure: Covariance Matters
by: Arts, Marloes, et al.
Published: (2023)
by: Arts, Marloes, et al.
Published: (2023)
Navigating Uncertainty in Medical Image Segmentation
by: Zepf, Kilian, et al.
Published: (2024)
by: Zepf, Kilian, et al.
Published: (2024)
Fast Sampling for Flows and Diffusions with Lazy and Point Mass Stochastic Interpolants
by: Damsholt, Gabriel, et al.
Published: (2026)
by: Damsholt, Gabriel, et al.
Published: (2026)
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
by: Saada, Thiziri Nait, et al.
Published: (2025)
by: Saada, Thiziri Nait, et al.
Published: (2025)
(Don't) Mind the Gap
by: Frank, Natalie Priebe, et al.
Published: (2025)
by: Frank, Natalie Priebe, et al.
Published: (2025)
MolMiner: Towards Controllable, 3D-Aware, Fragment-Based Molecular Design
by: Ortega-Ochoa, Raul, et al.
Published: (2024)
by: Ortega-Ochoa, Raul, et al.
Published: (2024)
Latent Diffusion for Missing Data
by: Estad, Alberte Heering, et al.
Published: (2026)
by: Estad, Alberte Heering, et al.
Published: (2026)
Order-Agnostic Autoregressive Modelling with Missing Data
by: Peis, Ignacio, et al.
Published: (2026)
by: Peis, Ignacio, et al.
Published: (2026)
GeoFormer: A Multi-Polygon Segmentation Transformer
by: Khomiakov, Maxim, et al.
Published: (2024)
by: Khomiakov, Maxim, et al.
Published: (2024)
The Geometries of Truth Are Orthogonal Across Tasks
by: Azizian, Waiss, et al.
Published: (2025)
by: Azizian, Waiss, et al.
Published: (2025)
When Peers Outperform AI (and When They Don't): Interaction Quality Over Modality
by: Morris, Caitlin, et al.
Published: (2026)
by: Morris, Caitlin, et al.
Published: (2026)
Discriminating Form and Meaning in Multilingual Models with Minimal-Pair ABX Tasks
by: de Seyssel, Maureen, et al.
Published: (2025)
by: de Seyssel, Maureen, et al.
Published: (2025)
Learning with Noisy Labels by Adaptive Gradient-Based Outlier Removal
by: Sedova, Anastasiia, et al.
Published: (2023)
by: Sedova, Anastasiia, et al.
Published: (2023)
The magnetic scalar potential for a rectangular prism
by: James, Berian, et al.
Published: (2025)
by: James, Berian, et al.
Published: (2025)
Hyper-Transforming Latent Diffusion Models
by: Peis, Ignacio, et al.
Published: (2025)
by: Peis, Ignacio, et al.
Published: (2025)
Scaling Laws for Optimal Data Mixtures
by: Shukor, Mustafa, et al.
Published: (2025)
by: Shukor, Mustafa, et al.
Published: (2025)
Explainability as statistical inference
by: Senetaire, Hugo Henri Joseph, et al.
Published: (2022)
by: Senetaire, Hugo Henri Joseph, et al.
Published: (2022)
Dynamic Gradient Alignment for Online Data Mixing
by: Fan, Simin, et al.
Published: (2024)
by: Fan, Simin, et al.
Published: (2024)
Prune, Don't Rebuild: Efficiently Tuning $α$-Reachable Graphs for Nearest Neighbor Search
by: Zhang, Tian, et al.
Published: (2026)
by: Zhang, Tian, et al.
Published: (2026)
Similar Items
-
Scaling Laws for Mixture Pretraining Under Data Constraints
by: Sedova, Anastasiia, et al.
Published: (2026) -
Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions
by: Sedova, Anastasiia, et al.
Published: (2026) -
Optimal Splitting of Language Models from Mixtures to Specialized Domains
by: Seto, Skyler, et al.
Published: (2026) -
Training Bilingual LMs with Data Constraints in the Targeted Language
by: Seto, Skyler, et al.
Published: (2024) -
Learning Energy-Based Models by Self-normalising the Likelihood
by: Senetaire, Hugo, et al.
Published: (2025)