LLM-Generated Natural Language Meets Scaling Laws: New Explorations and Data Augmentation Methods
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zhenhua, Xu, Guang, Ren, Ming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Will sentiment analysis need subculture? A new data augmentation approach
von: Wang, Zhenhua, et al.
Veröffentlicht: (2023)
von: Wang, Zhenhua, et al.
Veröffentlicht: (2023)
When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method
von: Zhang, Biao, et al.
Veröffentlicht: (2024)
von: Zhang, Biao, et al.
Veröffentlicht: (2024)
Controllable and Diverse Data Augmentation with Large Language Model for Low-Resource Open-Domain Dialogue Generation
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024)
a1: Steep Test-time Scaling Law via Environment Augmented Generation
von: Mei, Lingrui, et al.
Veröffentlicht: (2025)
von: Mei, Lingrui, et al.
Veröffentlicht: (2025)
InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition
von: Liu, Fengze, et al.
Veröffentlicht: (2026)
von: Liu, Fengze, et al.
Veröffentlicht: (2026)
StatBot.Swiss: Bilingual Open Data Exploration in Natural Language
von: Nooralahzadeh, Farhad, et al.
Veröffentlicht: (2024)
von: Nooralahzadeh, Farhad, et al.
Veröffentlicht: (2024)
Exploration of Marker-Based Approaches in Argument Mining through Augmented Natural Language
von: Das, Nilmadhab, et al.
Veröffentlicht: (2024)
von: Das, Nilmadhab, et al.
Veröffentlicht: (2024)
CAP-LLM: Context-Augmented Personalized Large Language Models for News Headline Generation
von: Wilson, Raymond, et al.
Veröffentlicht: (2025)
von: Wilson, Raymond, et al.
Veröffentlicht: (2025)
Scaling Laws of Synthetic Data for Language Models
von: Qin, Zeyu, et al.
Veröffentlicht: (2025)
von: Qin, Zeyu, et al.
Veröffentlicht: (2025)
Scaling Law for Language Models Training Considering Batch Size
von: Shuai, Xian, et al.
Veröffentlicht: (2024)
von: Shuai, Xian, et al.
Veröffentlicht: (2024)
Scaling Laws for Discriminative Classification in Large Language Models
von: Wyatte, Dean, et al.
Veröffentlicht: (2024)
von: Wyatte, Dean, et al.
Veröffentlicht: (2024)
Correcting Hallucinations in News Summaries: Exploration of Self-Correcting LLM Methods with External Knowledge
von: Vladika, Juraj, et al.
Veröffentlicht: (2025)
von: Vladika, Juraj, et al.
Veröffentlicht: (2025)
Curriculum-style Data Augmentation for LLM-based Metaphor Detection
von: Jia, Kaidi, et al.
Veröffentlicht: (2024)
von: Jia, Kaidi, et al.
Veröffentlicht: (2024)
REX-RAG: Reasoning Exploration with Policy Correction in Retrieval-Augmented Generation
von: Jiang, Wentao, et al.
Veröffentlicht: (2025)
von: Jiang, Wentao, et al.
Veröffentlicht: (2025)
Inference Computation Scaling for Feature Augmentation in Recommendation Systems
von: Liu, Weihao, et al.
Veröffentlicht: (2025)
von: Liu, Weihao, et al.
Veröffentlicht: (2025)
Unveiling Knowledge Utilization Mechanisms in LLM-based Retrieval-Augmented Generation
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Generation for Natural Language Processing: A Survey
von: Wu, Shangyu, et al.
Veröffentlicht: (2024)
von: Wu, Shangyu, et al.
Veröffentlicht: (2024)
On-the-fly Denoising for Data Augmentation in Natural Language Understanding
von: Fang, Tianqing, et al.
Veröffentlicht: (2022)
von: Fang, Tianqing, et al.
Veröffentlicht: (2022)
Data Augmentation for Fake Reviews Detection in Multiple Languages and Multiple Domains
von: Liu, Ming, et al.
Veröffentlicht: (2025)
von: Liu, Ming, et al.
Veröffentlicht: (2025)
LLM-DA: Data Augmentation via Large Language Models for Few-Shot Named Entity Recognition
von: Ye, Junjie, et al.
Veröffentlicht: (2024)
von: Ye, Junjie, et al.
Veröffentlicht: (2024)
The Scaling Laws of Skills in LLM Agent Systems
von: Chen, Charles, et al.
Veröffentlicht: (2026)
von: Chen, Charles, et al.
Veröffentlicht: (2026)
When Large Language Models Meet Law: Dual-Lens Taxonomy, Technical Advances, and Ethical Governance
von: Shao, Peizhang, et al.
Veröffentlicht: (2025)
von: Shao, Peizhang, et al.
Veröffentlicht: (2025)
D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models
von: Que, Haoran, et al.
Veröffentlicht: (2024)
von: Que, Haoran, et al.
Veröffentlicht: (2024)
NameBERT: Scaling Name-Based Nationality Classification with LLM-Augmented Open Academic Data
von: Ming, Cong, et al.
Veröffentlicht: (2026)
von: Ming, Cong, et al.
Veröffentlicht: (2026)
Finding the Minimal Parameter Budget for Implicit Reasoning: A Data Complexity Driven Scaling Law for Language Models
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
Scaling Laws for Post Training Quantized Large Language Models
von: Xu, Zifei, et al.
Veröffentlicht: (2024)
von: Xu, Zifei, et al.
Veröffentlicht: (2024)
Few-shot LLM Synthetic Data with Distribution Matching
von: Ren, Jiyuan, et al.
Veröffentlicht: (2025)
von: Ren, Jiyuan, et al.
Veröffentlicht: (2025)
VTS-LLM: Domain-Adaptive LLM Agent for Enhancing Awareness in Vessel Traffic Services through Natural Language
von: Sun, Sijin, et al.
Veröffentlicht: (2025)
von: Sun, Sijin, et al.
Veröffentlicht: (2025)
A Survey on Data Synthesis and Augmentation for Large Language Models
von: Wang, Ke, et al.
Veröffentlicht: (2024)
von: Wang, Ke, et al.
Veröffentlicht: (2024)
Natural Language Generation in Healthcare: A Review of Methods and Applications
von: Lyu, Mengxian, et al.
Veröffentlicht: (2025)
von: Lyu, Mengxian, et al.
Veröffentlicht: (2025)
Temporal Scaling Law for Large Language Models
von: Xiong, Yizhe, et al.
Veröffentlicht: (2024)
von: Xiong, Yizhe, et al.
Veröffentlicht: (2024)
Scaling Laws for Linear Complexity Language Models
von: Shen, Xuyang, et al.
Veröffentlicht: (2024)
von: Shen, Xuyang, et al.
Veröffentlicht: (2024)
People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection
von: Sen, Indira, et al.
Veröffentlicht: (2023)
von: Sen, Indira, et al.
Veröffentlicht: (2023)
Exploring the Impact of Table-to-Text Methods on Augmenting LLM-based Question Answering with Domain Hybrid Data
von: Min, Dehai, et al.
Veröffentlicht: (2024)
von: Min, Dehai, et al.
Veröffentlicht: (2024)
ChipXplore: Natural Language Exploration of Hardware Designs and Libraries
von: Abdelatty, Manar, et al.
Veröffentlicht: (2024)
von: Abdelatty, Manar, et al.
Veröffentlicht: (2024)
Grounding Natural Language to SQL Translation with Data-Based Self-Explanations
von: Fan, Yuankai, et al.
Veröffentlicht: (2024)
von: Fan, Yuankai, et al.
Veröffentlicht: (2024)
Toward Generalized Cross-Lingual Hateful Language Detection with Web-Scale Data and Ensemble LLM Annotations
von: Dang, Dang H., et al.
Veröffentlicht: (2026)
von: Dang, Dang H., et al.
Veröffentlicht: (2026)
Investigating the Impact of Semi-Supervised Methods with Data Augmentation on Offensive Language Detection in Romanian Language
von: Nicola, Elena-Beatrice, et al.
Veröffentlicht: (2024)
von: Nicola, Elena-Beatrice, et al.
Veröffentlicht: (2024)
A-RAG: Scaling Agentic Retrieval-Augmented Generation via Hierarchical Retrieval Interfaces
von: Du, Mingxuan, et al.
Veröffentlicht: (2026)
von: Du, Mingxuan, et al.
Veröffentlicht: (2026)
Retrieval-Augmented Generation for Natural Language Art Provenance Searches in the Getty Provenance Index
von: Henrickson, Mathew
Veröffentlicht: (2025)
von: Henrickson, Mathew
Veröffentlicht: (2025)
Ähnliche Einträge
-
Will sentiment analysis need subculture? A new data augmentation approach
von: Wang, Zhenhua, et al.
Veröffentlicht: (2023) -
When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method
von: Zhang, Biao, et al.
Veröffentlicht: (2024) -
Controllable and Diverse Data Augmentation with Large Language Model for Low-Resource Open-Domain Dialogue Generation
von: Liu, Zhenhua, et al.
Veröffentlicht: (2024) -
a1: Steep Test-time Scaling Law via Environment Augmented Generation
von: Mei, Lingrui, et al.
Veröffentlicht: (2025) -
InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition
von: Liu, Fengze, et al.
Veröffentlicht: (2026)