Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Xu, Peng, Runyu, Tong, Jian, Zhou, Yunhua, Lv, Haijun, Lu, Zhihui, Guo, Qipeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pre-Trained Policy Discriminators are General Reward Models
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
von: Dou, Shihan, et al.
Veröffentlicht: (2025)
Evolution of Concepts in Language Model Pre-Training
von: Ge, Xuyang, et al.
Veröffentlicht: (2025)
von: Ge, Xuyang, et al.
Veröffentlicht: (2025)
Data-free Weight Compress and Denoise for Large Language Models
von: Peng, Runyu, et al.
Veröffentlicht: (2024)
von: Peng, Runyu, et al.
Veröffentlicht: (2024)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
von: Peng, Runyu, et al.
Veröffentlicht: (2026)
von: Peng, Runyu, et al.
Veröffentlicht: (2026)
Explicit Multi-head Attention for Inter-head Interaction in Large Language Models
von: Peng, Runyu, et al.
Veröffentlicht: (2026)
von: Peng, Runyu, et al.
Veröffentlicht: (2026)
Improving Language Models Trained on Translated Data with Continual Pre-Training and Dictionary Learning Analysis
von: Boughorbel, Sabri, et al.
Veröffentlicht: (2024)
von: Boughorbel, Sabri, et al.
Veröffentlicht: (2024)
Synthesize-on-Graph: Knowledgeable Synthetic Data Generation for Continue Pre-training of Large Language Models
von: Ma, Shengjie, et al.
Veröffentlicht: (2025)
von: Ma, Shengjie, et al.
Veröffentlicht: (2025)
LoopRPT: Reinforcement Pre-Training for Looped Language Models
von: Tang, Guo, et al.
Veröffentlicht: (2026)
von: Tang, Guo, et al.
Veröffentlicht: (2026)
"According to ...": Prompting Language Models Improves Quoting from Pre-Training Data
von: Weller, Orion, et al.
Veröffentlicht: (2023)
von: Weller, Orion, et al.
Veröffentlicht: (2023)
MiniPLM: Knowledge Distillation for Pre-Training Language Models
von: Gu, Yuxian, et al.
Veröffentlicht: (2024)
von: Gu, Yuxian, et al.
Veröffentlicht: (2024)
ConcEPT: Concept-Enhanced Pre-Training for Language Models
von: Wang, Xintao, et al.
Veröffentlicht: (2024)
von: Wang, Xintao, et al.
Veröffentlicht: (2024)
Efficient Pre-Training with Token Superposition
von: Peng, Bowen, et al.
Veröffentlicht: (2026)
von: Peng, Bowen, et al.
Veröffentlicht: (2026)
Reinforcement Pre-Training
von: Dong, Qingxiu, et al.
Veröffentlicht: (2025)
von: Dong, Qingxiu, et al.
Veröffentlicht: (2025)
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training
von: Chen, Yangyi, et al.
Veröffentlicht: (2025)
von: Chen, Yangyi, et al.
Veröffentlicht: (2025)
Reinforcement Learning on Pre-Training Data
von: Li, Siheng, et al.
Veröffentlicht: (2025)
von: Li, Siheng, et al.
Veröffentlicht: (2025)
Medical Vision-Language Pre-Training for Brain Abnormalities
von: Monajatipoor, Masoud, et al.
Veröffentlicht: (2024)
von: Monajatipoor, Masoud, et al.
Veröffentlicht: (2024)
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
von: Zhang, Charlie, et al.
Veröffentlicht: (2025)
von: Zhang, Charlie, et al.
Veröffentlicht: (2025)
Exploring Forgetting in Large Language Model Pre-Training
von: Liao, Chonghua, et al.
Veröffentlicht: (2024)
von: Liao, Chonghua, et al.
Veröffentlicht: (2024)
Breaking Language Barriers: Cross-Lingual Continual Pre-Training at Scale
von: Zheng, Wenzhen, et al.
Veröffentlicht: (2024)
von: Zheng, Wenzhen, et al.
Veröffentlicht: (2024)
Long Context Pre-Training with Lighthouse Attention
von: Peng, Bowen, et al.
Veröffentlicht: (2026)
von: Peng, Bowen, et al.
Veröffentlicht: (2026)
Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models
von: Zhang, Jingyang, et al.
Veröffentlicht: (2024)
von: Zhang, Jingyang, et al.
Veröffentlicht: (2024)
Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data
von: Zhang, Jingyu, et al.
Veröffentlicht: (2024)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2024)
Analysing The Impact of Sequence Composition on Language Model Pre-Training
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
Instruction Pre-Training: Language Models are Supervised Multitask Learners
von: Cheng, Daixuan, et al.
Veröffentlicht: (2024)
von: Cheng, Daixuan, et al.
Veröffentlicht: (2024)
Revealing the Inherent Instructability of Pre-Trained Language Models
von: An, Seokhyun, et al.
Veröffentlicht: (2024)
von: An, Seokhyun, et al.
Veröffentlicht: (2024)
Data Darwinism Part I: Unlocking the Value of Scientific Data for Pre-training
von: Qin, Yiwei, et al.
Veröffentlicht: (2026)
von: Qin, Yiwei, et al.
Veröffentlicht: (2026)
FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
von: Patel, Ajay, et al.
Veröffentlicht: (2026)
von: Patel, Ajay, et al.
Veröffentlicht: (2026)
Learning Dynamics in Continual Pre-Training for Large Language Models
von: Wang, Xingjin, et al.
Veröffentlicht: (2025)
von: Wang, Xingjin, et al.
Veröffentlicht: (2025)
Unraveling Emotions with Pre-Trained Models
von: Pajón-Sanmartín, Alejandro, et al.
Veröffentlicht: (2025)
von: Pajón-Sanmartín, Alejandro, et al.
Veröffentlicht: (2025)
On the Limits of Model Merging for Multilinguality in Pre-Training
von: Aycock, Seth, et al.
Veröffentlicht: (2026)
von: Aycock, Seth, et al.
Veröffentlicht: (2026)
Probing Language Models for Pre-training Data Detection
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
The Inverse Scaling Effect of Pre-Trained Language Model Surprisal Is Not Due to Data Leakage
von: Oh, Byung-Doh, et al.
Veröffentlicht: (2025)
von: Oh, Byung-Doh, et al.
Veröffentlicht: (2025)
Investigating and Scaling up Code-Switching for Multilingual Language Model Pre-Training
von: Wang, Zhijun, et al.
Veröffentlicht: (2025)
von: Wang, Zhijun, et al.
Veröffentlicht: (2025)
Rethinking Reflection in Pre-Training
von: AI, Essential, et al.
Veröffentlicht: (2025)
von: AI, Essential, et al.
Veröffentlicht: (2025)
Probing Pre-Trained Language Models for Cross-Cultural Differences in Values
von: Arora, Arnav, et al.
Veröffentlicht: (2022)
von: Arora, Arnav, et al.
Veröffentlicht: (2022)
Explaining Pre-Trained Language Models with Attribution Scores: An Analysis in Low-Resource Settings
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
Release of Pre-Trained Models for the Japanese Language
von: Sawada, Kei, et al.
Veröffentlicht: (2024)
von: Sawada, Kei, et al.
Veröffentlicht: (2024)
Pre-Training Curriculum for Multi-Token Prediction in Language Models
von: Aynetdinov, Ansar, et al.
Veröffentlicht: (2025)
von: Aynetdinov, Ansar, et al.
Veröffentlicht: (2025)
Domain-Adaptive Continued Pre-Training of Small Language Models
von: Faroz, Salman
Veröffentlicht: (2025)
von: Faroz, Salman
Veröffentlicht: (2025)
Development of Pre-Trained Transformer-based Models for the Nepali Language
von: Thapa, Prajwal, et al.
Veröffentlicht: (2024)
von: Thapa, Prajwal, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Pre-Trained Policy Discriminators are General Reward Models
von: Dou, Shihan, et al.
Veröffentlicht: (2025) -
Evolution of Concepts in Language Model Pre-Training
von: Ge, Xuyang, et al.
Veröffentlicht: (2025) -
Data-free Weight Compress and Denoise for Large Language Models
von: Peng, Runyu, et al.
Veröffentlicht: (2024) -
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
von: Peng, Runyu, et al.
Veröffentlicht: (2026) -
Explicit Multi-head Attention for Inter-head Interaction in Large Language Models
von: Peng, Runyu, et al.
Veröffentlicht: (2026)