LLM-Generated Natural Language Meets Scaling Laws: New Explorations and Data Augmentation Methods
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zhenhua, Xu, Guang, Ren, Ming |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Will sentiment analysis need subculture? A new data augmentation approach
by: Wang, Zhenhua, et al.
Published: (2023)
by: Wang, Zhenhua, et al.
Published: (2023)
When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method
by: Zhang, Biao, et al.
Published: (2024)
by: Zhang, Biao, et al.
Published: (2024)
Controllable and Diverse Data Augmentation with Large Language Model for Low-Resource Open-Domain Dialogue Generation
by: Liu, Zhenhua, et al.
Published: (2024)
by: Liu, Zhenhua, et al.
Published: (2024)
a1: Steep Test-time Scaling Law via Environment Augmented Generation
by: Mei, Lingrui, et al.
Published: (2025)
by: Mei, Lingrui, et al.
Published: (2025)
InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition
by: Liu, Fengze, et al.
Published: (2026)
by: Liu, Fengze, et al.
Published: (2026)
StatBot.Swiss: Bilingual Open Data Exploration in Natural Language
by: Nooralahzadeh, Farhad, et al.
Published: (2024)
by: Nooralahzadeh, Farhad, et al.
Published: (2024)
Exploration of Marker-Based Approaches in Argument Mining through Augmented Natural Language
by: Das, Nilmadhab, et al.
Published: (2024)
by: Das, Nilmadhab, et al.
Published: (2024)
CAP-LLM: Context-Augmented Personalized Large Language Models for News Headline Generation
by: Wilson, Raymond, et al.
Published: (2025)
by: Wilson, Raymond, et al.
Published: (2025)
Scaling Laws of Synthetic Data for Language Models
by: Qin, Zeyu, et al.
Published: (2025)
by: Qin, Zeyu, et al.
Published: (2025)
Scaling Law for Language Models Training Considering Batch Size
by: Shuai, Xian, et al.
Published: (2024)
by: Shuai, Xian, et al.
Published: (2024)
Scaling Laws for Discriminative Classification in Large Language Models
by: Wyatte, Dean, et al.
Published: (2024)
by: Wyatte, Dean, et al.
Published: (2024)
Correcting Hallucinations in News Summaries: Exploration of Self-Correcting LLM Methods with External Knowledge
by: Vladika, Juraj, et al.
Published: (2025)
by: Vladika, Juraj, et al.
Published: (2025)
Curriculum-style Data Augmentation for LLM-based Metaphor Detection
by: Jia, Kaidi, et al.
Published: (2024)
by: Jia, Kaidi, et al.
Published: (2024)
REX-RAG: Reasoning Exploration with Policy Correction in Retrieval-Augmented Generation
by: Jiang, Wentao, et al.
Published: (2025)
by: Jiang, Wentao, et al.
Published: (2025)
Inference Computation Scaling for Feature Augmentation in Recommendation Systems
by: Liu, Weihao, et al.
Published: (2025)
by: Liu, Weihao, et al.
Published: (2025)
Unveiling Knowledge Utilization Mechanisms in LLM-based Retrieval-Augmented Generation
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
Retrieval-Augmented Generation for Natural Language Processing: A Survey
by: Wu, Shangyu, et al.
Published: (2024)
by: Wu, Shangyu, et al.
Published: (2024)
On-the-fly Denoising for Data Augmentation in Natural Language Understanding
by: Fang, Tianqing, et al.
Published: (2022)
by: Fang, Tianqing, et al.
Published: (2022)
Data Augmentation for Fake Reviews Detection in Multiple Languages and Multiple Domains
by: Liu, Ming, et al.
Published: (2025)
by: Liu, Ming, et al.
Published: (2025)
LLM-DA: Data Augmentation via Large Language Models for Few-Shot Named Entity Recognition
by: Ye, Junjie, et al.
Published: (2024)
by: Ye, Junjie, et al.
Published: (2024)
The Scaling Laws of Skills in LLM Agent Systems
by: Chen, Charles, et al.
Published: (2026)
by: Chen, Charles, et al.
Published: (2026)
When Large Language Models Meet Law: Dual-Lens Taxonomy, Technical Advances, and Ethical Governance
by: Shao, Peizhang, et al.
Published: (2025)
by: Shao, Peizhang, et al.
Published: (2025)
D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models
by: Que, Haoran, et al.
Published: (2024)
by: Que, Haoran, et al.
Published: (2024)
NameBERT: Scaling Name-Based Nationality Classification with LLM-Augmented Open Academic Data
by: Ming, Cong, et al.
Published: (2026)
by: Ming, Cong, et al.
Published: (2026)
Finding the Minimal Parameter Budget for Implicit Reasoning: A Data Complexity Driven Scaling Law for Language Models
by: Wang, Xinyi, et al.
Published: (2025)
by: Wang, Xinyi, et al.
Published: (2025)
Scaling Laws for Post Training Quantized Large Language Models
by: Xu, Zifei, et al.
Published: (2024)
by: Xu, Zifei, et al.
Published: (2024)
Few-shot LLM Synthetic Data with Distribution Matching
by: Ren, Jiyuan, et al.
Published: (2025)
by: Ren, Jiyuan, et al.
Published: (2025)
VTS-LLM: Domain-Adaptive LLM Agent for Enhancing Awareness in Vessel Traffic Services through Natural Language
by: Sun, Sijin, et al.
Published: (2025)
by: Sun, Sijin, et al.
Published: (2025)
A Survey on Data Synthesis and Augmentation for Large Language Models
by: Wang, Ke, et al.
Published: (2024)
by: Wang, Ke, et al.
Published: (2024)
Natural Language Generation in Healthcare: A Review of Methods and Applications
by: Lyu, Mengxian, et al.
Published: (2025)
by: Lyu, Mengxian, et al.
Published: (2025)
Temporal Scaling Law for Large Language Models
by: Xiong, Yizhe, et al.
Published: (2024)
by: Xiong, Yizhe, et al.
Published: (2024)
Scaling Laws for Linear Complexity Language Models
by: Shen, Xuyang, et al.
Published: (2024)
by: Shen, Xuyang, et al.
Published: (2024)
People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection
by: Sen, Indira, et al.
Published: (2023)
by: Sen, Indira, et al.
Published: (2023)
Exploring the Impact of Table-to-Text Methods on Augmenting LLM-based Question Answering with Domain Hybrid Data
by: Min, Dehai, et al.
Published: (2024)
by: Min, Dehai, et al.
Published: (2024)
ChipXplore: Natural Language Exploration of Hardware Designs and Libraries
by: Abdelatty, Manar, et al.
Published: (2024)
by: Abdelatty, Manar, et al.
Published: (2024)
Grounding Natural Language to SQL Translation with Data-Based Self-Explanations
by: Fan, Yuankai, et al.
Published: (2024)
by: Fan, Yuankai, et al.
Published: (2024)
Toward Generalized Cross-Lingual Hateful Language Detection with Web-Scale Data and Ensemble LLM Annotations
by: Dang, Dang H., et al.
Published: (2026)
by: Dang, Dang H., et al.
Published: (2026)
Investigating the Impact of Semi-Supervised Methods with Data Augmentation on Offensive Language Detection in Romanian Language
by: Nicola, Elena-Beatrice, et al.
Published: (2024)
by: Nicola, Elena-Beatrice, et al.
Published: (2024)
A-RAG: Scaling Agentic Retrieval-Augmented Generation via Hierarchical Retrieval Interfaces
by: Du, Mingxuan, et al.
Published: (2026)
by: Du, Mingxuan, et al.
Published: (2026)
Retrieval-Augmented Generation for Natural Language Art Provenance Searches in the Getty Provenance Index
by: Henrickson, Mathew
Published: (2025)
by: Henrickson, Mathew
Published: (2025)
Similar Items
-
Will sentiment analysis need subculture? A new data augmentation approach
by: Wang, Zhenhua, et al.
Published: (2023) -
When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method
by: Zhang, Biao, et al.
Published: (2024) -
Controllable and Diverse Data Augmentation with Large Language Model for Low-Resource Open-Domain Dialogue Generation
by: Liu, Zhenhua, et al.
Published: (2024) -
a1: Steep Test-time Scaling Law via Environment Augmented Generation
by: Mei, Lingrui, et al.
Published: (2025) -
InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition
by: Liu, Fengze, et al.
Published: (2026)