Improving Language Models Trained on Translated Data with Continual Pre-Training and Dictionary Learning Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Boughorbel, Sabri, Parvez, MD Rizwan, Hawasly, Majd |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond the Leaderboard: Understanding Performance Disparities in Large Language Models via Model Diffing
by: Boughorbel, Sabri, et al.
Published: (2025)
by: Boughorbel, Sabri, et al.
Published: (2025)
There Is More to Refusal in Large Language Models than a Single Direction
by: Joad, Faaiz, et al.
Published: (2026)
by: Joad, Faaiz, et al.
Published: (2026)
Self-Consistency from Only Two Samples: CoT-PoT Ensembling for Efficient LLM Reasoning
by: Saparkhan, Raman, et al.
Published: (2026)
by: Saparkhan, Raman, et al.
Published: (2026)
Scaling up Discovery of Latent Concepts in Deep NLP Models
by: Hawasly, Majd, et al.
Published: (2023)
by: Hawasly, Majd, et al.
Published: (2023)
Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data
by: Guo, Xu, et al.
Published: (2026)
by: Guo, Xu, et al.
Published: (2026)
Enhancing Translation Accuracy of Large Language Models through Continual Pre-Training on Parallel Data
by: Kondo, Minato, et al.
Published: (2024)
by: Kondo, Minato, et al.
Published: (2024)
Chain of Evidences and Evidence to Generate: Prompting for Context Grounded and Retrieval Augmented Reasoning
by: Parvez, Md Rizwan
Published: (2024)
by: Parvez, Md Rizwan
Published: (2024)
Learning Dynamics in Continual Pre-Training for Large Language Models
by: Wang, Xingjin, et al.
Published: (2025)
by: Wang, Xingjin, et al.
Published: (2025)
"According to ...": Prompting Language Models Improves Quoting from Pre-Training Data
by: Weller, Orion, et al.
Published: (2023)
by: Weller, Orion, et al.
Published: (2023)
How Do Large Language Models Learn Concepts During Continual Pre-Training?
by: Yao, Barry Menglong, et al.
Published: (2026)
by: Yao, Barry Menglong, et al.
Published: (2026)
Domain-Adaptive Continued Pre-Training of Small Language Models
by: Faroz, Salman
Published: (2025)
by: Faroz, Salman
Published: (2025)
Exploring Alignment in Shared Cross-lingual Spaces
by: Mousi, Basel, et al.
Published: (2024)
by: Mousi, Basel, et al.
Published: (2024)
From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes
by: Goriely, Zébulon, et al.
Published: (2024)
by: Goriely, Zébulon, et al.
Published: (2024)
Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models
by: Zhang, Jingyang, et al.
Published: (2024)
by: Zhang, Jingyang, et al.
Published: (2024)
Hephaestus: Improving Fundamental Agent Capabilities of Large Language Models through Continual Pre-Training
by: Zhuang, Yuchen, et al.
Published: (2025)
by: Zhuang, Yuchen, et al.
Published: (2025)
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
by: Zhang, Charlie, et al.
Published: (2025)
by: Zhang, Charlie, et al.
Published: (2025)
Chain-of-Dictionary Prompting Elicits Translation in Large Language Models
by: Lu, Hongyuan, et al.
Published: (2023)
by: Lu, Hongyuan, et al.
Published: (2023)
Breaking Language Barriers: Cross-Lingual Continual Pre-Training at Scale
by: Zheng, Wenzhen, et al.
Published: (2024)
by: Zheng, Wenzhen, et al.
Published: (2024)
Neural Machine Translation of Clinical Text: An Empirical Investigation into Multilingual Pre-Trained Language Models and Transfer-Learning
by: Han, Lifeng, et al.
Published: (2023)
by: Han, Lifeng, et al.
Published: (2023)
Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data
by: Zhang, Jingyu, et al.
Published: (2024)
by: Zhang, Jingyu, et al.
Published: (2024)
LAraBench: Benchmarking Arabic AI with Large Language Models
by: Abdelali, Ahmed, et al.
Published: (2023)
by: Abdelali, Ahmed, et al.
Published: (2023)
Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre-Training
by: Wu, Yu-Hang, et al.
Published: (2026)
by: Wu, Yu-Hang, et al.
Published: (2026)
LLMeBench: A Flexible Framework for Accelerating LLMs Benchmarking
by: Dalvi, Fahim, et al.
Published: (2023)
by: Dalvi, Fahim, et al.
Published: (2023)
Exploring Forgetting in Large Language Model Pre-Training
by: Liao, Chonghua, et al.
Published: (2024)
by: Liao, Chonghua, et al.
Published: (2024)
Reinforcement Learning on Pre-Training Data
by: Li, Siheng, et al.
Published: (2025)
by: Li, Siheng, et al.
Published: (2025)
Characterizing Learning Curves During Language Model Pre-Training: Learning, Forgetting, and Stability
by: Chang, Tyler A., et al.
Published: (2023)
by: Chang, Tyler A., et al.
Published: (2023)
The Inverse Scaling Effect of Pre-Trained Language Model Surprisal Is Not Due to Data Leakage
by: Oh, Byung-Doh, et al.
Published: (2025)
by: Oh, Byung-Doh, et al.
Published: (2025)
Improving Rare Word Translation With Dictionaries and Attention Masking
by: Sible, Kenneth J., et al.
Published: (2024)
by: Sible, Kenneth J., et al.
Published: (2024)
Explaining Pre-Trained Language Models with Attribution Scores: An Analysis in Low-Resource Settings
by: Zhou, Wei, et al.
Published: (2024)
by: Zhou, Wei, et al.
Published: (2024)
Large Language Model Empowered Recommendation Meets All-domain Continual Pre-Training
by: Ma, Haokai, et al.
Published: (2025)
by: Ma, Haokai, et al.
Published: (2025)
Evolution of Concepts in Language Model Pre-Training
by: Ge, Xuyang, et al.
Published: (2025)
by: Ge, Xuyang, et al.
Published: (2025)
D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models
by: Que, Haoran, et al.
Published: (2024)
by: Que, Haoran, et al.
Published: (2024)
PalmX 2025: The First Shared Task on Benchmarking LLMs on Arabic and Islamic Culture
by: Alwajih, Fakhraddin, et al.
Published: (2025)
by: Alwajih, Fakhraddin, et al.
Published: (2025)
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
by: Abbes, Istabrak, et al.
Published: (2025)
by: Abbes, Istabrak, et al.
Published: (2025)
MiniPLM: Knowledge Distillation for Pre-Training Language Models
by: Gu, Yuxian, et al.
Published: (2024)
by: Gu, Yuxian, et al.
Published: (2024)
Analysing The Impact of Sequence Composition on Language Model Pre-Training
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
Instruction Pre-Training: Language Models are Supervised Multitask Learners
by: Cheng, Daixuan, et al.
Published: (2024)
by: Cheng, Daixuan, et al.
Published: (2024)
ConcEPT: Concept-Enhanced Pre-Training for Language Models
by: Wang, Xintao, et al.
Published: (2024)
by: Wang, Xintao, et al.
Published: (2024)
LoopRPT: Reinforcement Pre-Training for Looped Language Models
by: Tang, Guo, et al.
Published: (2026)
by: Tang, Guo, et al.
Published: (2026)
Quantifying Memorization and Detecting Training Data of Pre-trained Language Models using Japanese Newspaper
by: Ishihara, Shotaro, et al.
Published: (2024)
by: Ishihara, Shotaro, et al.
Published: (2024)
Similar Items
-
Beyond the Leaderboard: Understanding Performance Disparities in Large Language Models via Model Diffing
by: Boughorbel, Sabri, et al.
Published: (2025) -
There Is More to Refusal in Large Language Models than a Single Direction
by: Joad, Faaiz, et al.
Published: (2026) -
Self-Consistency from Only Two Samples: CoT-PoT Ensembling for Efficient LLM Reasoning
by: Saparkhan, Raman, et al.
Published: (2026) -
Scaling up Discovery of Latent Concepts in Deep NLP Models
by: Hawasly, Majd, et al.
Published: (2023) -
Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data
by: Guo, Xu, et al.
Published: (2026)