Repetition over Diversity: High-Signal Data Filtering for Sample-Efficient German Language Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Aynetdinov, Ansar, Haller, Patrick, Akbik, Alan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pre-Training Curriculum for Multi-Token Prediction in Language Models
by: Aynetdinov, Ansar, et al.
Published: (2025)
by: Aynetdinov, Ansar, et al.
Published: (2025)
NoiseBench: Benchmarking the Impact of Real Label Noise on Named Entity Recognition
by: Merdjanovska, Elena, et al.
Published: (2024)
by: Merdjanovska, Elena, et al.
Published: (2024)
Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements
by: Haller, Patrick, et al.
Published: (2025)
by: Haller, Patrick, et al.
Published: (2025)
SemScore: Automated Evaluation of Instruction-Tuned LLMs based on Semantic Textual Similarity
by: Aynetdinov, Ansar, et al.
Published: (2024)
by: Aynetdinov, Ansar, et al.
Published: (2024)
What Matters in Linearizing Language Models? A Comparative Study of Architecture, Scale, and Task Adaptation
by: Haller, Patrick, et al.
Published: (2025)
by: Haller, Patrick, et al.
Published: (2025)
BabyHGRN: Exploring RNNs for Sample-Efficient Training of Language Models
by: Haller, Patrick, et al.
Published: (2024)
by: Haller, Patrick, et al.
Published: (2024)
Fabricator: An Open Source Toolkit for Generating Labeled Training Data with Teacher LLMs
by: Golde, Jonas, et al.
Published: (2023)
by: Golde, Jonas, et al.
Published: (2023)
From Data to Knowledge: Evaluating How Efficiently Language Models Learn Facts
by: Christoph, Daniel, et al.
Published: (2025)
by: Christoph, Daniel, et al.
Published: (2025)
What Matters When Building Universal Multilingual Named Entity Recognition Models?
by: Golde, Jonas, et al.
Published: (2026)
by: Golde, Jonas, et al.
Published: (2026)
PISA-Bench: The PISA Index as a Multilingual and Multimodal Metric for the Evaluation of Vision-Language Models
by: Haller, Patrick, et al.
Published: (2025)
by: Haller, Patrick, et al.
Published: (2025)
PECC: Problem Extraction and Coding Challenges
by: Haller, Patrick, et al.
Published: (2024)
by: Haller, Patrick, et al.
Published: (2024)
FiNERweb: Datasets and Artifacts for Scalable Multilingual Named Entity Recognition
by: Golde, Jonas, et al.
Published: (2025)
by: Golde, Jonas, et al.
Published: (2025)
MastermindEval: A Simple But Scalable Reasoning Benchmark
by: Golde, Jonas, et al.
Published: (2025)
by: Golde, Jonas, et al.
Published: (2025)
MCSD: An Efficient Language Model with Diverse Fusion
by: Yang, Hua, et al.
Published: (2024)
by: Yang, Hua, et al.
Published: (2024)
A Neural Model for Word Repetition
by: Dager, Daniel, et al.
Published: (2025)
by: Dager, Daniel, et al.
Published: (2025)
Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models
by: Wang, Shuxun, et al.
Published: (2025)
by: Wang, Shuxun, et al.
Published: (2025)
Large Language Models as Quasi-crystals: Coherence Without Repetition in Generative Text
by: Guevara-Vela, Jose Manuel
Published: (2025)
by: Guevara-Vela, Jose Manuel
Published: (2025)
Time-Annealed Perturbation Sampling: Diverse Generation for Diffusion Language Models
by: Wu, Jingxuan, et al.
Published: (2026)
by: Wu, Jingxuan, et al.
Published: (2026)
TexIm FAST: Text-to-Image Representation for Semantic Similarity Evaluation using Transformers
by: Ansar, Wazib, et al.
Published: (2024)
by: Ansar, Wazib, et al.
Published: (2024)
From Transformers to LLMs: A Systematic Survey of Efficiency Considerations in NLP
by: Ansar, Wazib, et al.
Published: (2024)
by: Ansar, Wazib, et al.
Published: (2024)
Fighting Against the Repetitive Training and Sample Dependency Problem in Few-shot Named Entity Recognition
by: Tian, Chang, et al.
Published: (2024)
by: Tian, Chang, et al.
Published: (2024)
Beyond Repetition: Text Simplification and Curriculum Learning for Data-Constrained Pretraining
by: Roque, Matthew Theodore, et al.
Published: (2025)
by: Roque, Matthew Theodore, et al.
Published: (2025)
Free Lunch for Pass@$k$? Low Cost Diverse Sampling for Diffusion Language Models
by: Lamont, Sean, et al.
Published: (2026)
by: Lamont, Sean, et al.
Published: (2026)
Balanced Data Sampling for Language Model Training with Clustering
by: Shao, Yunfan, et al.
Published: (2024)
by: Shao, Yunfan, et al.
Published: (2024)
Domain-Adaptation through Synthetic Data: Fine-Tuning Large Language Models for German Law
by: Bashir, Ali Hamza, et al.
Published: (2026)
by: Bashir, Ali Hamza, et al.
Published: (2026)
Familiarity: Better Evaluation of Zero-Shot Named Entity Recognition by Quantifying Label Shifts in Synthetic Training Data
by: Golde, Jonas, et al.
Published: (2024)
by: Golde, Jonas, et al.
Published: (2024)
Spectral Filters, Dark Signals, and Attention Sinks
by: Cancedda, Nicola
Published: (2024)
by: Cancedda, Nicola
Published: (2024)
Post-training Large Language Models for Diverse High-Quality Responses
by: Chen, Yilei, et al.
Published: (2025)
by: Chen, Yilei, et al.
Published: (2025)
Scaling Data Diversity for Fine-Tuning Language Models in Human Alignment
by: Song, Feifan, et al.
Published: (2024)
by: Song, Feifan, et al.
Published: (2024)
Target-Aware Language Modeling via Granular Data Sampling
by: Chang, Ernie, et al.
Published: (2024)
by: Chang, Ernie, et al.
Published: (2024)
Classifying German Language Proficiency Levels Using Large Language Models
by: Ahlers, Elias-Leander, et al.
Published: (2025)
by: Ahlers, Elias-Leander, et al.
Published: (2025)
Gracefully Filtering Backdoor Samples for Generative Large Language Models without Retraining
by: Wu, Zongru, et al.
Published: (2024)
by: Wu, Zongru, et al.
Published: (2024)
Efficient Evaluation of Large Language Models via Collaborative Filtering
by: Zhong, Xu-Xiang, et al.
Published: (2025)
by: Zhong, Xu-Xiang, et al.
Published: (2025)
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
by: Ali, Mehdi, et al.
Published: (2025)
by: Ali, Mehdi, et al.
Published: (2025)
Aligning Large Language Models to Follow Instructions and Hallucinate Less via Effective Data Filtering
by: Si, Shuzheng, et al.
Published: (2025)
by: Si, Shuzheng, et al.
Published: (2025)
From Language Models over Tokens to Language Models over Characters
by: Vieira, Tim, et al.
Published: (2024)
by: Vieira, Tim, et al.
Published: (2024)
Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
The Overlooked Repetitive Lengthening Form in Sentiment Analysis
by: Wang, Lei, et al.
Published: (2026)
by: Wang, Lei, et al.
Published: (2026)
BEExformer: A Fast Inferencing Binarized Transformer with Early Exits
by: Ansar, Wazib, et al.
Published: (2024)
by: Ansar, Wazib, et al.
Published: (2024)
What Should Baby Models Read? Exploring Sample-Efficient Data Composition on Model Performance
by: Yam, Hong Meng, et al.
Published: (2024)
by: Yam, Hong Meng, et al.
Published: (2024)
Similar Items
-
Pre-Training Curriculum for Multi-Token Prediction in Language Models
by: Aynetdinov, Ansar, et al.
Published: (2025) -
NoiseBench: Benchmarking the Impact of Real Label Noise on Named Entity Recognition
by: Merdjanovska, Elena, et al.
Published: (2024) -
Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements
by: Haller, Patrick, et al.
Published: (2025) -
SemScore: Automated Evaluation of Instruction-Tuned LLMs based on Semantic Textual Similarity
by: Aynetdinov, Ansar, et al.
Published: (2024) -
What Matters in Linearizing Language Models? A Comparative Study of Architecture, Scale, and Task Adaptation
by: Haller, Patrick, et al.
Published: (2025)