Will we run out of data? Limits of LLM scaling based on human-generated data
Fuente:
arXiv
Saved in:
| Main Authors: | Villalobos, Pablo, Ho, Anson, Sevilla, Jaime, Besiroglu, Tamay, Heim, Lennart, Hobbhahn, Marius |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Compute Divide in Machine Learning: A Threat to Academic Contribution and Scrutiny?
by: Besiroglu, Tamay, et al.
Published: (2024)
by: Besiroglu, Tamay, et al.
Published: (2024)
Algorithmic progress in language models
by: Ho, Anson, et al.
Published: (2024)
by: Ho, Anson, et al.
Published: (2024)
Chinchilla Scaling: A replication attempt
by: Besiroglu, Tamay, et al.
Published: (2024)
by: Besiroglu, Tamay, et al.
Published: (2024)
The rising costs of training frontier AI models
by: Cottier, Ben, et al.
Published: (2024)
by: Cottier, Ben, et al.
Published: (2024)
Guided Persona-based AI Surveys: Can we replicate personal mobility preferences at scale using LLMs?
by: Tzachristas, Ioannis, et al.
Published: (2025)
by: Tzachristas, Ioannis, et al.
Published: (2025)
Designing Incident Reporting Systems for Harms from General-Purpose AI
by: Wei, Kevin, et al.
Published: (2025)
by: Wei, Kevin, et al.
Published: (2025)
A semantic embedding space based on large language models for modelling human beliefs
by: Lee, Byunghwee, et al.
Published: (2024)
by: Lee, Byunghwee, et al.
Published: (2024)
What are human values, and how do we align AI to them?
by: Klingefjord, Oliver, et al.
Published: (2024)
by: Klingefjord, Oliver, et al.
Published: (2024)
LLM-based Semantic Augmentation for Harmful Content Detection
by: Meguellati, Elyas, et al.
Published: (2025)
by: Meguellati, Elyas, et al.
Published: (2025)
Counterfactual LLM-based Framework for Measuring Rhetorical Style
by: Qiu, Jingyi, et al.
Published: (2025)
by: Qiu, Jingyi, et al.
Published: (2025)
The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies
by: Zhou, Jiaxu, et al.
Published: (2025)
by: Zhou, Jiaxu, et al.
Published: (2025)
Bias in the Tails: How Name-conditioned Evaluative Framing in Resume Summaries Destabilizes LLM-based Hiring
by: Nghiem, Huy, et al.
Published: (2026)
by: Nghiem, Huy, et al.
Published: (2026)
Unveiling the Truth and Facilitating Change: Towards Agent-based Large-scale Social Movement Simulation
by: Mou, Xinyi, et al.
Published: (2024)
by: Mou, Xinyi, et al.
Published: (2024)
Training LLM-based Tutors to Improve Student Learning Outcomes in Dialogues
by: Scarlatos, Alexander, et al.
Published: (2025)
by: Scarlatos, Alexander, et al.
Published: (2025)
Automatic generation of DRI Statements
by: Flechtner, Maurice
Published: (2025)
by: Flechtner, Maurice
Published: (2025)
Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines
by: Zhang, Peixian, et al.
Published: (2025)
by: Zhang, Peixian, et al.
Published: (2025)
Value Drifts: Tracing Value Alignment During LLM Post-Training
by: Bhatia, Mehar, et al.
Published: (2025)
by: Bhatia, Mehar, et al.
Published: (2025)
Limited Effectiveness of LLM-based Data Augmentation for COVID-19 Misinformation Stance Detection
by: Choi, Eun Cheol, et al.
Published: (2025)
by: Choi, Eun Cheol, et al.
Published: (2025)
Techniques for supercharging academic writing with generative AI
by: Lin, Zhicheng
Published: (2023)
by: Lin, Zhicheng
Published: (2023)
Variance reduction in output from generative AI
by: Xie, Yu, et al.
Published: (2025)
by: Xie, Yu, et al.
Published: (2025)
"In order that" -- a data driven study of symptoms and causes of obsolescence
by: Rudnicka, Karolina
Published: (2025)
by: Rudnicka, Karolina
Published: (2025)
Increased Compute Efficiency and the Diffusion of AI Capabilities
by: Pilz, Konstantin, et al.
Published: (2023)
by: Pilz, Konstantin, et al.
Published: (2023)
AI-generated data contamination erodes pathological variability and diagnostic reliability
by: He, Hongyu, et al.
Published: (2026)
by: He, Hongyu, et al.
Published: (2026)
Training Compute Thresholds: Features and Functions in AI Regulation
by: Heim, Lennart, et al.
Published: (2024)
by: Heim, Lennart, et al.
Published: (2024)
Evaluation Awareness in Language Models Has Limited Effect on Behaviour
by: Knecht, Amelie, et al.
Published: (2026)
by: Knecht, Amelie, et al.
Published: (2026)
TALES: A Taxonomy and Analysis of Cultural Representations in LLM-generated Stories
by: Bhagat, Kirti, et al.
Published: (2025)
by: Bhagat, Kirti, et al.
Published: (2025)
In your own words: computationally identifying interpretable themes in free-text survey data
by: Wang, Jenny S, et al.
Published: (2026)
by: Wang, Jenny S, et al.
Published: (2026)
Using generative AI to support standardization work -- the case of 3GPP
by: Staron, Miroslaw, et al.
Published: (2024)
by: Staron, Miroslaw, et al.
Published: (2024)
Can the capability of Large Language Models be described by human ability? A Meta Study
by: Zan, Mingrui, et al.
Published: (2025)
by: Zan, Mingrui, et al.
Published: (2025)
Beyond Translation: LLM-Based Data Generation for Multilingual Fact-Checking
by: Chung, Yi-Ling, et al.
Published: (2025)
by: Chung, Yi-Ling, et al.
Published: (2025)
ChatGPT for President! Presupposed content in politicians versus GPT-generated texts
by: Garassino, Davide, et al.
Published: (2025)
by: Garassino, Davide, et al.
Published: (2025)
Estimating Idea Production: A Methodological Survey
by: Erdil, Ege, et al.
Published: (2024)
by: Erdil, Ege, et al.
Published: (2024)
Large Language Models can Strategically Deceive their Users when Put Under Pressure
by: Scheurer, Jérémy, et al.
Published: (2023)
by: Scheurer, Jérémy, et al.
Published: (2023)
Stan: An LLM-based thermodynamics course assistant
by: Furst, Eric M., et al.
Published: (2026)
by: Furst, Eric M., et al.
Published: (2026)
Can LLM be a Personalized Judge?
by: Dong, Yijiang River, et al.
Published: (2024)
by: Dong, Yijiang River, et al.
Published: (2024)
Quantifying the Persona Effect in LLM Simulations
by: Hu, Tiancheng, et al.
Published: (2024)
by: Hu, Tiancheng, et al.
Published: (2024)
LLM-MC-Affect: LLM-Based Monte Carlo Modeling of Affective Trajectories and Latent Ambiguity for Interpersonal Dynamic Insight
by: Lin, Yu-Zheng, et al.
Published: (2026)
by: Lin, Yu-Zheng, et al.
Published: (2026)
An LLM Agent for Automatic Geospatial Data Analysis
by: Chen, Yuxing, et al.
Published: (2024)
by: Chen, Yuxing, et al.
Published: (2024)
Multilingual Prompting for Improving LLM Generation Diversity
by: Wang, Qihan, et al.
Published: (2025)
by: Wang, Qihan, et al.
Published: (2025)
The Biased Samaritan: LLM biases in Perceived Kindness
by: Fagan, Jack H, et al.
Published: (2025)
by: Fagan, Jack H, et al.
Published: (2025)
Similar Items
-
The Compute Divide in Machine Learning: A Threat to Academic Contribution and Scrutiny?
by: Besiroglu, Tamay, et al.
Published: (2024) -
Algorithmic progress in language models
by: Ho, Anson, et al.
Published: (2024) -
Chinchilla Scaling: A replication attempt
by: Besiroglu, Tamay, et al.
Published: (2024) -
The rising costs of training frontier AI models
by: Cottier, Ben, et al.
Published: (2024) -
Guided Persona-based AI Surveys: Can we replicate personal mobility preferences at scale using LLMs?
by: Tzachristas, Ioannis, et al.
Published: (2025)