Rewriting Pre-Training Data Boosts LLM Performance in Math and Code
Fuente:
arXiv
Saved in:
| Main Authors: | Fujii, Kazuki, Tajima, Yukito, Mizuki, Sakae, Kawamura, Masaki, Shimada, Hinari, Shiotani, Taihei, Saito, Koshiro, Oi, Masanari, Nakamura, Taishi, Okamoto, Takumi, Ishida, Shigeki, Hattori, Kakeru, Ma, Youmi, Takamura, Hiroya, Yokota, Rio, Sakuma, Jun, Okazaki, Naoaki |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Building Instruction-Tuning Datasets from Human-Written Instructions with Open-Weight Large Language Models
by: Ma, Youmi, et al.
Published: (2025)
by: Ma, Youmi, et al.
Published: (2025)
Why We Build Local Large Language Models: An Observational Analysis from 35 Japanese and Multilingual LLMs
by: Saito, Koshiro, et al.
Published: (2024)
by: Saito, Koshiro, et al.
Published: (2024)
Continual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capabilities
by: Fujii, Kazuki, et al.
Published: (2024)
by: Fujii, Kazuki, et al.
Published: (2024)
Building a Large Japanese Web Corpus for Large Language Models
by: Okazaki, Naoaki, et al.
Published: (2024)
by: Okazaki, Naoaki, et al.
Published: (2024)
A Japanese Benchmark for Evaluating Social Bias in Reasoning Based on Attribution Theory
by: Shiotani, Taihei, et al.
Published: (2026)
by: Shiotani, Taihei, et al.
Published: (2026)
Synthesizing Instruction-Tuning Datasets with Contrastive Decoding
by: Ichinose, Tatsuya, et al.
Published: (2026)
by: Ichinose, Tatsuya, et al.
Published: (2026)
JUBAKU: An Adversarial Benchmark for Exposing Culturally Grounded Stereotypes in Japanese LLMs
by: Shiotani, Taihei, et al.
Published: (2026)
by: Shiotani, Taihei, et al.
Published: (2026)
From Interpretability to Performance: Optimizing Retrieval Heads for Long-Context Language Models
by: Ma, Youmi, et al.
Published: (2026)
by: Ma, Youmi, et al.
Published: (2026)
Building a Japanese Document-Level Relation Extraction Dataset Assisted by Cross-Lingual Transfer
by: Ma, Youmi, et al.
Published: (2024)
by: Ma, Youmi, et al.
Published: (2024)
Likelihood-based Mitigation of Evaluation Bias in Large Language Models
by: Oi, Masanari, et al.
Published: (2024)
by: Oi, Masanari, et al.
Published: (2024)
Autoregressive Direct Preference Optimization
by: Oi, Masanari, et al.
Published: (2026)
by: Oi, Masanari, et al.
Published: (2026)
QuantumBench: A Benchmark for Quantum Problem Solving
by: Minami, Shunya, et al.
Published: (2025)
by: Minami, Shunya, et al.
Published: (2025)
Sampling-based Pseudo-Likelihood for Membership Inference Attacks
by: Kaneko, Masahiro, et al.
Published: (2024)
by: Kaneko, Masahiro, et al.
Published: (2024)
From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models
by: Oi, Masanari, et al.
Published: (2026)
by: Oi, Masanari, et al.
Published: (2026)
LLM Output Detectability and Task Performance Can be Jointly Optimized
by: Saito, Koshiro, et al.
Published: (2026)
by: Saito, Koshiro, et al.
Published: (2026)
Masked Gated Linear Unit
by: Tajima, Yukito, et al.
Published: (2025)
by: Tajima, Yukito, et al.
Published: (2025)
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
by: Ohi, Masanari, et al.
Published: (2024)
by: Ohi, Masanari, et al.
Published: (2024)
Role of Vanadium-Oxide Layer in Electronic State of Sr$_2$VFeAsO$_{3-δ}$ with Oxygen Deficiency
by: Nakajima, Masamichi, et al.
Published: (2024)
by: Nakajima, Masamichi, et al.
Published: (2024)
JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
by: Maeda, Koki, et al.
Published: (2026)
by: Maeda, Koki, et al.
Published: (2026)
Knowledge of Pretrained Language Models on Surface Information of Tokens
by: Hiraoka, Tatsuya, et al.
Published: (2024)
by: Hiraoka, Tatsuya, et al.
Published: (2024)
Tokenization as Finite-State Transduction
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Agent AI for Finance
by: Chen, Chung-Chi, et al.
Published: (2025)
by: Chen, Chung-Chi, et al.
Published: (2025)
Fermionic theory of nonequilibrium steady states
by: De Giuli, Eric, et al.
Published: (2023)
by: De Giuli, Eric, et al.
Published: (2023)
Statistical mechanics of fluids with hidden degrees of freedom
by: Shimada, Masanari, et al.
Published: (2025)
by: Shimada, Masanari, et al.
Published: (2025)
Prompting for Numerical Sequences: A Case Study on Market Comment Generation
by: Kawarada, Masayuki, et al.
Published: (2024)
by: Kawarada, Masayuki, et al.
Published: (2024)
Multimodal Task Interference: A Benchmark and Analysis of History-Target Mismatch in Multimodal LLMs
by: Kawarada, Masayuki, et al.
Published: (2026)
by: Kawarada, Masayuki, et al.
Published: (2026)
Bias correction of posterior means using MCMC outputs
by: Iba, Yukito
Published: (2024)
by: Iba, Yukito
Published: (2024)
W-Kernel and Its Principal Space for Frequentist Evaluation of Bayesian Estimators
by: Iba, Yukito
Published: (2023)
by: Iba, Yukito
Published: (2023)
On the Optimal Reasoning Length for RL-Trained Language Models
by: Nohara, Daisuke, et al.
Published: (2026)
by: Nohara, Daisuke, et al.
Published: (2026)
Balancing Speed and Stability: The Trade-offs of FP8 vs. BF16 Training in LLMs
by: Fujii, Kazuki, et al.
Published: (2024)
by: Fujii, Kazuki, et al.
Published: (2024)
Adolescent Self‐Inserted Rectal Foreign Body Removed Manually With Topical Anesthetic Only
by: Takuto Shimada, et al.
Published: (2026)
by: Takuto Shimada, et al.
Published: (2026)
Bit-level BPE: Below the byte boundary
by: Moon, Sangwhan, et al.
Published: (2025)
by: Moon, Sangwhan, et al.
Published: (2025)
Drifting Objectives for Refining Discrete Diffusion Language Models
by: Oba, Daisuke, et al.
Published: (2026)
by: Oba, Daisuke, et al.
Published: (2026)
Distributional Properties of Subword Regularization
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Evaluating Gender Bias of Pre-trained Language Models in Natural Language Inference by Considering All Labels
by: Anantaprayoon, Panatchakorn, et al.
Published: (2023)
by: Anantaprayoon, Panatchakorn, et al.
Published: (2023)
Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMs
by: Miyamoto, Sora, et al.
Published: (2026)
by: Miyamoto, Sora, et al.
Published: (2026)
Intent-Aware Self-Correction for Mitigating Social Biases in Large Language Models
by: Anantaprayoon, Panatchakorn, et al.
Published: (2025)
by: Anantaprayoon, Panatchakorn, et al.
Published: (2025)
Diffusion-State Policy Optimization for Masked Diffusion Language Models
by: Oba, Daisuke, et al.
Published: (2026)
by: Oba, Daisuke, et al.
Published: (2026)
OUTFOX: LLM-Generated Essay Detection Through In-Context Learning with Adversarially Generated Examples
by: Koike, Ryuto, et al.
Published: (2023)
by: Koike, Ryuto, et al.
Published: (2023)
How You Prompt Matters! Even Task-Oriented Constraints in Instructions Affect LLM-Generated Text Detection
by: Koike, Ryuto, et al.
Published: (2023)
by: Koike, Ryuto, et al.
Published: (2023)
Similar Items
-
Building Instruction-Tuning Datasets from Human-Written Instructions with Open-Weight Large Language Models
by: Ma, Youmi, et al.
Published: (2025) -
Why We Build Local Large Language Models: An Observational Analysis from 35 Japanese and Multilingual LLMs
by: Saito, Koshiro, et al.
Published: (2024) -
Continual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capabilities
by: Fujii, Kazuki, et al.
Published: (2024) -
Building a Large Japanese Web Corpus for Large Language Models
by: Okazaki, Naoaki, et al.
Published: (2024) -
A Japanese Benchmark for Evaluating Social Bias in Reasoning Based on Attribution Theory
by: Shiotani, Taihei, et al.
Published: (2026)