Register Always Matters: Analysis of LLM Pretraining Data Through the Lens of Language Variation
Fuente:
arXiv
Saved in:
| Main Authors: | Myntti, Amanda, Henriksson, Erik, Laippala, Veronika, Pyysalo, Sampo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Automatic register identification for the open web using multilingual deep learning
by: Henriksson, Erik, et al.
Published: (2024)
by: Henriksson, Erik, et al.
Published: (2024)
Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance
by: Myntti, Amanda, et al.
Published: (2026)
by: Myntti, Amanda, et al.
Published: (2026)
A Survey of Large Language Models for European Languages
by: Ali, Wazir, et al.
Published: (2024)
by: Ali, Wazir, et al.
Published: (2024)
Got Compute, but No Data: Lessons From Post-training a Finnish LLM
by: Zosa, Elaine, et al.
Published: (2025)
by: Zosa, Elaine, et al.
Published: (2025)
Pretraining Finnish ModernBERTs
by: Reunamo, Akseli, et al.
Published: (2025)
by: Reunamo, Akseli, et al.
Published: (2025)
FinerWeb-10BT: Refining Web Data with LLM-Based Line-Level Filtering
by: Henriksson, Erik, et al.
Published: (2025)
by: Henriksson, Erik, et al.
Published: (2025)
Hopes and Fears -- Emotion Distribution in the Topic Landscape of Finnish Parliamentary Speech 2000-2020
by: Ristilä, Anna, et al.
Published: (2026)
by: Ristilä, Anna, et al.
Published: (2026)
An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)
by: Burchell, Laurie, et al.
Published: (2025)
by: Burchell, Laurie, et al.
Published: (2025)
FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models
by: Kytöniemi, Joona, et al.
Published: (2025)
by: Kytöniemi, Joona, et al.
Published: (2025)
Scaling Data-Constrained Language Models
by: Muennighoff, Niklas, et al.
Published: (2023)
by: Muennighoff, Niklas, et al.
Published: (2023)
HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models
by: Oepen, Stephan, et al.
Published: (2025)
by: Oepen, Stephan, et al.
Published: (2025)
Poro 34B and the Blessing of Multilinguality
by: Luukkonen, Risto, et al.
Published: (2024)
by: Luukkonen, Risto, et al.
Published: (2024)
Preference Curriculum: LLMs Should Always Be Pretrained on Their Preferred Data
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
The Impact of Steering Large Language Models with Persona Vectors in Educational Applications
by: Wu, Yongchao, et al.
Published: (2026)
by: Wu, Yongchao, et al.
Published: (2026)
Data-Constrained Synthesis of Training Data for De-Identification
by: Vakili, Thomas, et al.
Published: (2025)
by: Vakili, Thomas, et al.
Published: (2025)
Tracing Persona Vectors Through LLM Pretraining
by: Moskvoretskii, Viktor, et al.
Published: (2026)
by: Moskvoretskii, Viktor, et al.
Published: (2026)
A New Massive Multilingual Dataset for High-Performance Language Technologies
by: de Gibert, Ona, et al.
Published: (2024)
by: de Gibert, Ona, et al.
Published: (2024)
Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
by: Jobanputra, Mayank, et al.
Published: (2025)
by: Jobanputra, Mayank, et al.
Published: (2025)
Is Preference Alignment Always the Best Option to Enhance LLM-Based Translation? An Empirical Analysis
by: Gisserot-Boukhlef, Hippolyte, et al.
Published: (2024)
by: Gisserot-Boukhlef, Hippolyte, et al.
Published: (2024)
RedWhale: An Adapted Korean LLM Through Efficient Continual Pretraining
by: Vo, Anh-Dung, et al.
Published: (2024)
by: Vo, Anh-Dung, et al.
Published: (2024)
Reframing Data Value for Large Language Models Through the Lens of Plausibility
by: Rammal, Mohamad Rida, et al.
Published: (2024)
by: Rammal, Mohamad Rida, et al.
Published: (2024)
Evaluating the Reliability of Self-Explanations in Large Language Models
by: Randl, Korbinian, et al.
Published: (2024)
by: Randl, Korbinian, et al.
Published: (2024)
Steering Large Language Models with Register Analysis for Arbitrary Style Transfer
by: Yang, Xinchen, et al.
Published: (2025)
by: Yang, Xinchen, et al.
Published: (2025)
DistillLens: Symmetric Knowledge Distillation Through Logit Lens
by: Dhakal, Manish, et al.
Published: (2026)
by: Dhakal, Manish, et al.
Published: (2026)
LogitLens4LLMs: Extending Logit Lens Analysis to Modern Large Language Models
by: Wang, Zhenyu
Published: (2025)
by: Wang, Zhenyu
Published: (2025)
Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering
by: Negoita, Vlad, et al.
Published: (2025)
by: Negoita, Vlad, et al.
Published: (2025)
On the importance of Data Scale in Pretraining Arabic Language Models
by: Ghaddar, Abbas, et al.
Published: (2024)
by: Ghaddar, Abbas, et al.
Published: (2024)
Optimizing Pretraining Data Mixtures with LLM-Estimated Utility
by: Held, William, et al.
Published: (2025)
by: Held, William, et al.
Published: (2025)
Always Learning, Always Mixing: Efficient and Simple Data Mixing All The Time
by: Hu, Michael Y., et al.
Published: (2026)
by: Hu, Michael Y., et al.
Published: (2026)
Hermit Kingdom Through the Lens of Multiple Perspectives: A Case Study of LLM Hallucination on North Korea
by: Cho, Eunjung, et al.
Published: (2025)
by: Cho, Eunjung, et al.
Published: (2025)
Analyzing LLMs' Knowledge Boundary Cognition Across Languages Through the Lens of Internal Representations
by: Xiao, Chenghao, et al.
Published: (2025)
by: Xiao, Chenghao, et al.
Published: (2025)
Artificial Impressions: Evaluating Large Language Model Behavior Through the Lens of Trait Impressions
by: Deas, Nicholas, et al.
Published: (2025)
by: Deas, Nicholas, et al.
Published: (2025)
Large Language Model Agents Are Not Always Faithful Self-Evolvers
by: Zhao, Weixiang, et al.
Published: (2026)
by: Zhao, Weixiang, et al.
Published: (2026)
Exploring Concreteness Through a Figurative Lens
by: Ghosh, Saptarshi, et al.
Published: (2026)
by: Ghosh, Saptarshi, et al.
Published: (2026)
QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining
by: Liu, Fengze, et al.
Published: (2025)
by: Liu, Fengze, et al.
Published: (2025)
Data Caricatures: On the Representation of African American Language in Pretraining Corpora
by: Deas, Nicholas, et al.
Published: (2025)
by: Deas, Nicholas, et al.
Published: (2025)
Multilingual Language Model Pretraining using Machine-translated Data
by: Wang, Jiayi, et al.
Published: (2025)
by: Wang, Jiayi, et al.
Published: (2025)
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
by: Messmer, Bettina, et al.
Published: (2025)
by: Messmer, Bettina, et al.
Published: (2025)
Scaling Down Semantic Leakage: Investigating Associative Bias in Smaller Language Models
by: Smilga, Veronika
Published: (2025)
by: Smilga, Veronika
Published: (2025)
Do Large Language Models Adapt to Language Variation across Socioeconomic Status?
by: Bassignana, Elisa, et al.
Published: (2026)
by: Bassignana, Elisa, et al.
Published: (2026)
Similar Items
-
Automatic register identification for the open web using multilingual deep learning
by: Henriksson, Erik, et al.
Published: (2024) -
Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance
by: Myntti, Amanda, et al.
Published: (2026) -
A Survey of Large Language Models for European Languages
by: Ali, Wazir, et al.
Published: (2024) -
Got Compute, but No Data: Lessons From Post-training a Finnish LLM
by: Zosa, Elaine, et al.
Published: (2025) -
Pretraining Finnish ModernBERTs
by: Reunamo, Akseli, et al.
Published: (2025)