Synthetic Data Generation Using Large Language Models: Advances in Text and Code
Fuente:
arXiv
Guardado en:
| Autores principales: | Nadas, Mihai, Diosan, Laura, Tomescu, Andreea |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TF3-RO-50M: Training Compact Romanian Language Models from Scratch on Synthetic Moral Microfiction
por: Nadas, Mihai Dan, et al.
Publicado: (2026)
por: Nadas, Mihai Dan, et al.
Publicado: (2026)
TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models
por: Nadas, Mihai, et al.
Publicado: (2025)
por: Nadas, Mihai, et al.
Publicado: (2025)
Building Large-Scale English-Romanian Literary Translation Resources with Open Models
por: Nadas, Mihai, et al.
Publicado: (2025)
por: Nadas, Mihai, et al.
Publicado: (2025)
Evaluating Large Language Models for Diacritic Restoration in Romanian Texts: A Comparative Study
por: Nadas, Mihai, et al.
Publicado: (2025)
por: Nadas, Mihai, et al.
Publicado: (2025)
Value-Aware Numerical Representations for Transformer Language Models
por: Dutulescu, Andreea, et al.
Publicado: (2026)
por: Dutulescu, Andreea, et al.
Publicado: (2026)
Training a Large Language Model for Medical Coding Using Privacy-Preserving Synthetic Clinical Data
por: Cook, John, et al.
Publicado: (2026)
por: Cook, John, et al.
Publicado: (2026)
Synthetic Text Generation for Training Large Language Models via Gradient Matching
por: Nguyen, Dang, et al.
Publicado: (2025)
por: Nguyen, Dang, et al.
Publicado: (2025)
German Text Simplification: Finetuning Large Language Models with Semi-Synthetic Data
por: Klöser, Lars, et al.
Publicado: (2024)
por: Klöser, Lars, et al.
Publicado: (2024)
DataGen: Unified Synthetic Dataset Generation via Large Language Models
por: Huang, Yue, et al.
Publicado: (2024)
por: Huang, Yue, et al.
Publicado: (2024)
Evaluating Language Models as Synthetic Data Generators
por: Kim, Seungone, et al.
Publicado: (2024)
por: Kim, Seungone, et al.
Publicado: (2024)
Case2Code: Scalable Synthetic Data for Code Generation
por: Shao, Yunfan, et al.
Publicado: (2024)
por: Shao, Yunfan, et al.
Publicado: (2024)
Data Generation Using Large Language Models for Text Classification: An Empirical Case Study
por: Li, Yinheng, et al.
Publicado: (2024)
por: Li, Yinheng, et al.
Publicado: (2024)
SyntheT2C: Generating Synthetic Data for Fine-Tuning Large Language Models on the Text2Cypher Task
por: Zhong, Ziije, et al.
Publicado: (2024)
por: Zhong, Ziije, et al.
Publicado: (2024)
Synthetic Data Generation for Phrase Break Prediction with Large Language Model
por: Lee, Hoyeon, et al.
Publicado: (2025)
por: Lee, Hoyeon, et al.
Publicado: (2025)
Genetic Instruct: Scaling up Synthetic Generation of Coding Instructions for Large Language Models
por: Majumdar, Somshubra, et al.
Publicado: (2024)
por: Majumdar, Somshubra, et al.
Publicado: (2024)
Novel Preprocessing Technique for Data Embedding in Engineering Code Generation Using Large Language Model
por: Lin, Yu-Chen, et al.
Publicado: (2023)
por: Lin, Yu-Chen, et al.
Publicado: (2023)
Socially Aware Synthetic Data Generation for Suicidal Ideation Detection Using Large Language Models
por: Ghanadian, Hamideh, et al.
Publicado: (2024)
por: Ghanadian, Hamideh, et al.
Publicado: (2024)
Structsum Generation for Faster Text Comprehension
por: Jain, Parag, et al.
Publicado: (2024)
por: Jain, Parag, et al.
Publicado: (2024)
Exploring Mathematical Extrapolation of Large Language Models with Synthetic Data
por: Li, Haolong, et al.
Publicado: (2024)
por: Li, Haolong, et al.
Publicado: (2024)
GCOF: Self-iterative Text Generation for Copywriting Using Large Language Model
por: Zhou, Jianghui, et al.
Publicado: (2024)
por: Zhou, Jianghui, et al.
Publicado: (2024)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
por: Yang, Yue, et al.
Publicado: (2025)
por: Yang, Yue, et al.
Publicado: (2025)
Persona-Based Synthetic Data Generation Using Multi-Stage Conditioning with Large Language Models for Emotion Recognition
por: Inoshita, Keito, et al.
Publicado: (2025)
por: Inoshita, Keito, et al.
Publicado: (2025)
An Extensive Evaluation of Factual Consistency in Large Language Models for Data-to-Text Generation
por: Mahapatra, Joy, et al.
Publicado: (2024)
por: Mahapatra, Joy, et al.
Publicado: (2024)
Generative Text Steganography with Large Language Model
por: Wu, Jiaxuan, et al.
Publicado: (2024)
por: Wu, Jiaxuan, et al.
Publicado: (2024)
Federated Domain-Specific Knowledge Transfer on Large Language Models Using Synthetic Data
por: Li, Haoran, et al.
Publicado: (2024)
por: Li, Haoran, et al.
Publicado: (2024)
Advancing Text Classification with Large Language Models and Neural Attention Mechanisms
por: Lyu, Ning, et al.
Publicado: (2025)
por: Lyu, Ning, et al.
Publicado: (2025)
Private Synthetic Text Generation with Diffusion Models
por: Ochs, Sebastian, et al.
Publicado: (2024)
por: Ochs, Sebastian, et al.
Publicado: (2024)
DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models
por: Huang, Yiming, et al.
Publicado: (2024)
por: Huang, Yiming, et al.
Publicado: (2024)
Forging Time Series with Language: A Large Language Model Approach to Synthetic Data Generation
por: Rousseau, Cécile, et al.
Publicado: (2025)
por: Rousseau, Cécile, et al.
Publicado: (2025)
Can Large Language Models Understand, Reason About, and Generate Code-Switched Text?
por: Winata, Genta Indra, et al.
Publicado: (2026)
por: Winata, Genta Indra, et al.
Publicado: (2026)
Taiyi-Diffusion-XL: Advancing Bilingual Text-to-Image Generation with Large Vision-Language Model Support
por: Wu, Xiaojun, et al.
Publicado: (2024)
por: Wu, Xiaojun, et al.
Publicado: (2024)
Deep Active Learning for Data Mining from Conflict Text Corpora
por: Croicu, Mihai
Publicado: (2024)
por: Croicu, Mihai
Publicado: (2024)
Aligning Large Language Models via Fully Self-Synthetic Data
por: Yin, Shangjian, et al.
Publicado: (2025)
por: Yin, Shangjian, et al.
Publicado: (2025)
Unlocking the Potential of Large Language Models in the Nuclear Industry with Synthetic Data
por: Anwar, Muhammad, et al.
Publicado: (2025)
por: Anwar, Muhammad, et al.
Publicado: (2025)
On the Diversity of Synthetic Data and its Impact on Training Large Language Models
por: Chen, Hao, et al.
Publicado: (2024)
por: Chen, Hao, et al.
Publicado: (2024)
RedStone: Curating General, Code, Math, and QA Data for Large Language Models
por: Chang, Yaoyao, et al.
Publicado: (2024)
por: Chang, Yaoyao, et al.
Publicado: (2024)
The Heap: A Contamination-Free Multilingual Code Dataset for Evaluating Large Language Models
por: Katzy, Jonathan, et al.
Publicado: (2025)
por: Katzy, Jonathan, et al.
Publicado: (2025)
Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models
por: Xu, Ran, et al.
Publicado: (2023)
por: Xu, Ran, et al.
Publicado: (2023)
Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models
por: Golchin, Shahriar, et al.
Publicado: (2023)
por: Golchin, Shahriar, et al.
Publicado: (2023)
SynthTextEval: Synthetic Text Data Generation and Evaluation for High-Stakes Domains
por: Ramesh, Krithika, et al.
Publicado: (2025)
por: Ramesh, Krithika, et al.
Publicado: (2025)
Ejemplares similares
-
TF3-RO-50M: Training Compact Romanian Language Models from Scratch on Synthetic Moral Microfiction
por: Nadas, Mihai Dan, et al.
Publicado: (2026) -
TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models
por: Nadas, Mihai, et al.
Publicado: (2025) -
Building Large-Scale English-Romanian Literary Translation Resources with Open Models
por: Nadas, Mihai, et al.
Publicado: (2025) -
Evaluating Large Language Models for Diacritic Restoration in Romanian Texts: A Comparative Study
por: Nadas, Mihai, et al.
Publicado: (2025) -
Value-Aware Numerical Representations for Transformer Language Models
por: Dutulescu, Andreea, et al.
Publicado: (2026)