Training Bilingual LMs with Data Constraints in the Targeted Language
Fuente:
arXiv
Saved in:
| Main Authors: | Seto, Skyler, ter Hoeve, Maartje, Bai, Richard He, Schluter, Natalie, Grangier, David |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Assessing the Role of Data Quality in Training Bilingual Language Models
by: Seto, Skyler, et al.
Published: (2025)
by: Seto, Skyler, et al.
Published: (2025)
Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions
by: Sedova, Anastasiia, et al.
Published: (2026)
by: Sedova, Anastasiia, et al.
Published: (2026)
Scaling Laws for Mixture Pretraining Under Data Constraints
by: Sedova, Anastasiia, et al.
Published: (2026)
by: Sedova, Anastasiia, et al.
Published: (2026)
Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
by: Grangier, David, et al.
Published: (2024)
by: Grangier, David, et al.
Published: (2024)
Discriminating Form and Meaning in Multilingual Models with Minimal-Pair ABX Tasks
by: de Seyssel, Maureen, et al.
Published: (2025)
by: de Seyssel, Maureen, et al.
Published: (2025)
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
by: Ablin, Pierre, et al.
Published: (2025)
by: Ablin, Pierre, et al.
Published: (2025)
On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization
by: Lin, Yong, et al.
Published: (2024)
by: Lin, Yong, et al.
Published: (2024)
How Value Induction Reshapes LLM Behaviour
by: Arora, Arnav, et al.
Published: (2026)
by: Arora, Arnav, et al.
Published: (2026)
On the Way to LLM Personalization: Learning to Remember User Conversations
by: Magister, Lucie Charlotte, et al.
Published: (2024)
by: Magister, Lucie Charlotte, et al.
Published: (2024)
Optimal Splitting of Language Models from Mixtures to Specialized Domains
by: Seto, Skyler, et al.
Published: (2026)
by: Seto, Skyler, et al.
Published: (2026)
GrammaMT: Improving Machine Translation with Grammar-Informed In-Context Learning
by: Ramos, Rita, et al.
Published: (2024)
by: Ramos, Rita, et al.
Published: (2024)
Analyzing Dialectical Biases in LLMs for Knowledge and Reasoning Benchmarks
by: Pan, Eileen, et al.
Published: (2025)
by: Pan, Eileen, et al.
Published: (2025)
Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling
by: Maini, Pratyush, et al.
Published: (2024)
by: Maini, Pratyush, et al.
Published: (2024)
Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings
by: Jeha, Paul, et al.
Published: (2026)
by: Jeha, Paul, et al.
Published: (2026)
Analyzing the Effect of Linguistic Similarity on Cross-Lingual Transfer: Tasks and Experimental Setups Matter
by: Blaschke, Verena, et al.
Published: (2025)
by: Blaschke, Verena, et al.
Published: (2025)
Dynamic Gradient Alignment for Online Data Mixing
by: Fan, Simin, et al.
Published: (2024)
by: Fan, Simin, et al.
Published: (2024)
Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIP
by: Huang, Chen, et al.
Published: (2024)
by: Huang, Chen, et al.
Published: (2024)
No Need to Talk: Asynchronous Mixture of Language Models
by: Filippova, Anastasiia, et al.
Published: (2024)
by: Filippova, Anastasiia, et al.
Published: (2024)
Need a Small Specialized Language Model? Plan Early!
by: Grangier, David, et al.
Published: (2024)
by: Grangier, David, et al.
Published: (2024)
mRAKL: Multilingual Retrieval-Augmented Knowledge Graph Construction for Low-Resourced Languages
by: Nigatu, Hellina Hailu, et al.
Published: (2025)
by: Nigatu, Hellina Hailu, et al.
Published: (2025)
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
by: Bethune, Louis, et al.
Published: (2025)
by: Bethune, Louis, et al.
Published: (2025)
Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models
by: Sundar, Anirudh, et al.
Published: (2025)
by: Sundar, Anirudh, et al.
Published: (2025)
Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
by: Damani, Mehul, et al.
Published: (2025)
by: Damani, Mehul, et al.
Published: (2025)
Compositional preference models for aligning LMs
by: Go, Dongyoung, et al.
Published: (2023)
by: Go, Dongyoung, et al.
Published: (2023)
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
by: Saada, Thiziri Nait, et al.
Published: (2025)
by: Saada, Thiziri Nait, et al.
Published: (2025)
A Little Help Goes a Long Way: Efficient LLM Training by Leveraging Small LMs
by: Rawat, Ankit Singh, et al.
Published: (2024)
by: Rawat, Ankit Singh, et al.
Published: (2024)
Enabling Approximate Joint Sampling in Diffusion LMs
by: Bansal, Parikshit, et al.
Published: (2025)
by: Bansal, Parikshit, et al.
Published: (2025)
Entropy-Aligned Decoding of LMs for Better Writing and Reasoning
by: Ahmed, Kareem, et al.
Published: (2026)
by: Ahmed, Kareem, et al.
Published: (2026)
Training a Bilingual Language Model by Mapping Tokens onto a Shared Character Space
by: Rom, Aviad, et al.
Published: (2024)
by: Rom, Aviad, et al.
Published: (2024)
SoftLMs: Efficient Adaptive Low-Rank Approximation of Language Models using Soft-Thresholding Mechanism
by: Bhatnagar, Priyansh, et al.
Published: (2024)
by: Bhatnagar, Priyansh, et al.
Published: (2024)
Exploring Gender Bias in Large Language Models: An In-depth Dive into the German Language
by: Gnadt, Kristin, et al.
Published: (2025)
by: Gnadt, Kristin, et al.
Published: (2025)
Learning How to Ask: Querying LMs with Mixtures of Soft Prompts
by: Qin, Guanghui, et al.
Published: (2021)
by: Qin, Guanghui, et al.
Published: (2021)
Diffusion LMs Can Approximate Optimal Infilling Lengths Implicitly
by: Liu, Hengchang, et al.
Published: (2026)
by: Liu, Hengchang, et al.
Published: (2026)
Kanana: Compute-efficient Bilingual Language Models
by: Kanana LLM Team, et al.
Published: (2025)
by: Kanana LLM Team, et al.
Published: (2025)
Reflect-RL: Two-Player Online RL Fine-Tuning for LMs
by: Zhou, Runlong, et al.
Published: (2024)
by: Zhou, Runlong, et al.
Published: (2024)
Pretraining with hierarchical memories: separating long-tail and common knowledge
by: Pouransari, Hadi, et al.
Published: (2025)
by: Pouransari, Hadi, et al.
Published: (2025)
Enhancing Speech Instruction Understanding and Disambiguation in Robotics via Speech Prosody
by: Sasu, David, et al.
Published: (2025)
by: Sasu, David, et al.
Published: (2025)
Safe Reinforcement Learning with Free-form Natural Language Constraints and Pre-Trained Language Models
by: Lou, Xingzhou, et al.
Published: (2024)
by: Lou, Xingzhou, et al.
Published: (2024)
On the Effect of Purely Synthetic Training Data for Different Automatic Speech Recognition Architectures
by: Hilmes, Benedikt, et al.
Published: (2024)
by: Hilmes, Benedikt, et al.
Published: (2024)
CroissantLLM: A Truly Bilingual French-English Language Model
by: Faysse, Manuel, et al.
Published: (2024)
by: Faysse, Manuel, et al.
Published: (2024)
Similar Items
-
Assessing the Role of Data Quality in Training Bilingual Language Models
by: Seto, Skyler, et al.
Published: (2025) -
Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions
by: Sedova, Anastasiia, et al.
Published: (2026) -
Scaling Laws for Mixture Pretraining Under Data Constraints
by: Sedova, Anastasiia, et al.
Published: (2026) -
Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
by: Grangier, David, et al.
Published: (2024) -
Discriminating Form and Meaning in Multilingual Models with Minimal-Pair ABX Tasks
by: de Seyssel, Maureen, et al.
Published: (2025)