DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Patel, Ajay, Raffel, Colin, Callison-Burch, Chris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
von: Patel, Ajay, et al.
Veröffentlicht: (2026)
von: Patel, Ajay, et al.
Veröffentlicht: (2026)
StyleDistance: Stronger Content-Independent Style Embeddings with Synthetic Parallel Examples
von: Patel, Ajay, et al.
Veröffentlicht: (2024)
von: Patel, Ajay, et al.
Veröffentlicht: (2024)
Position: The Most Expensive Part of an LLM should be its Training Data
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2025)
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2025)
Large Language Models Can Self-Improve At Web Agent Tasks
von: Patel, Ajay, et al.
Veröffentlicht: (2024)
von: Patel, Ajay, et al.
Veröffentlicht: (2024)
Low-Resource Authorship Style Transfer: Can Non-Famous Authors Be Imitated?
von: Patel, Ajay, et al.
Veröffentlicht: (2022)
von: Patel, Ajay, et al.
Veröffentlicht: (2022)
Probabilistic Soundness Guarantees in LLM Reasoning Chains
von: You, Weiqiu, et al.
Veröffentlicht: (2025)
von: You, Weiqiu, et al.
Veröffentlicht: (2025)
WithdrarXiv: A Large-Scale Dataset for Retraction Study
von: Rao, Delip, et al.
Veröffentlicht: (2024)
von: Rao, Delip, et al.
Veröffentlicht: (2024)
DataDream: Few-shot Guided Dataset Generation
von: Kim, Jae Myung, et al.
Veröffentlicht: (2024)
von: Kim, Jae Myung, et al.
Veröffentlicht: (2024)
Merging by Matching Models in Task Parameter Subspaces
von: Tam, Derek, et al.
Veröffentlicht: (2023)
von: Tam, Derek, et al.
Veröffentlicht: (2023)
Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap
von: Zhu, Andrew, et al.
Veröffentlicht: (2025)
von: Zhu, Andrew, et al.
Veröffentlicht: (2025)
GenAI Content Detection Task 3: Cross-Domain Machine-Generated Text Detection Challenge
von: Dugan, Liam, et al.
Veröffentlicht: (2025)
von: Dugan, Liam, et al.
Veröffentlicht: (2025)
Autorubric: Unifying Rubric-based LLM Evaluation
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
von: Liang, Hao, et al.
Veröffentlicht: (2025)
von: Liang, Hao, et al.
Veröffentlicht: (2025)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
von: Yang, Yue, et al.
Veröffentlicht: (2025)
von: Yang, Yue, et al.
Veröffentlicht: (2025)
mStyleDistance: Multilingual Style Embeddings and their Evaluation
von: Qiu, Justin, et al.
Veröffentlicht: (2025)
von: Qiu, Justin, et al.
Veröffentlicht: (2025)
The Power of LLM-Generated Synthetic Data for Stance Detection in Online Political Discussions
von: Wagner, Stefan Sylvius, et al.
Veröffentlicht: (2024)
von: Wagner, Stefan Sylvius, et al.
Veröffentlicht: (2024)
Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use
von: Goldie, Anna, et al.
Veröffentlicht: (2025)
von: Goldie, Anna, et al.
Veröffentlicht: (2025)
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
Realistic Evaluation of Model Merging for Compositional Generalization
von: Tam, Derek, et al.
Veröffentlicht: (2024)
von: Tam, Derek, et al.
Veröffentlicht: (2024)
Domain Gating Ensemble Networks for AI-Generated Text Detection
von: Tripathi, Arihant, et al.
Veröffentlicht: (2025)
von: Tripathi, Arihant, et al.
Veröffentlicht: (2025)
What Do Claim Verification Datasets Actually Test? A Reasoning Trace Analysis
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
von: Niklaus, Joel, et al.
Veröffentlicht: (2026)
von: Niklaus, Joel, et al.
Veröffentlicht: (2026)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
von: Yadav, Prateek, et al.
Veröffentlicht: (2023)
von: Yadav, Prateek, et al.
Veröffentlicht: (2023)
A Survey on Data Selection for Language Models
von: Albalak, Alon, et al.
Veröffentlicht: (2024)
von: Albalak, Alon, et al.
Veröffentlicht: (2024)
Scaling Data-Constrained Language Models
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2023)
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2023)
Quality Matters: Evaluating Synthetic Data for Tool-Using LLMs
von: Iskander, Shadi, et al.
Veröffentlicht: (2024)
von: Iskander, Shadi, et al.
Veröffentlicht: (2024)
Simultaneous Masking, Not Prompting Optimization: A Paradigm Shift in Fine-tuning LLMs for Simultaneous Translation
von: Raffel, Matthew, et al.
Veröffentlicht: (2024)
von: Raffel, Matthew, et al.
Veröffentlicht: (2024)
Synthetic vs. Gold: The Role of LLM Generated Labels and Data in Cyberbullying Detection
von: Kazemi, Arefeh, et al.
Veröffentlicht: (2025)
von: Kazemi, Arefeh, et al.
Veröffentlicht: (2025)
First Steps Towards Overhearing LLM Agents: A Case Study With Dungeons & Dragons Gameplay
von: Zhu, Andrew, et al.
Veröffentlicht: (2025)
von: Zhu, Andrew, et al.
Veröffentlicht: (2025)
Balancing Cost and Effectiveness of Synthetic Data Generation Strategies for LLMs
von: Chan, Yung-Chieh, et al.
Veröffentlicht: (2024)
von: Chan, Yung-Chieh, et al.
Veröffentlicht: (2024)
Towards Active Synthetic Data Generation for Finetuning Language Models
von: Kessler, Samuel, et al.
Veröffentlicht: (2025)
von: Kessler, Samuel, et al.
Veröffentlicht: (2025)
ParaGuide: Guided Diffusion Paraphrasers for Plug-and-Play Textual Style Transfer
von: Horvitz, Zachary, et al.
Veröffentlicht: (2023)
von: Horvitz, Zachary, et al.
Veröffentlicht: (2023)
When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method
von: Zhang, Biao, et al.
Veröffentlicht: (2024)
von: Zhang, Biao, et al.
Veröffentlicht: (2024)
BeaverTalk: Oregon State University's IWSLT 2025 Simultaneous Speech Translation System
von: Raffel, Matthew, et al.
Veröffentlicht: (2025)
von: Raffel, Matthew, et al.
Veröffentlicht: (2025)
RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
Synthetic Data Generation for Intersectional Fairness by Leveraging Hierarchical Group Structure
von: Maheshwari, Gaurav, et al.
Veröffentlicht: (2024)
von: Maheshwari, Gaurav, et al.
Veröffentlicht: (2024)
Synthetic Data Generation and Joint Learning for Robust Code-Mixed Translation
von: Kartik, Kartik, et al.
Veröffentlicht: (2024)
von: Kartik, Kartik, et al.
Veröffentlicht: (2024)
Data Cartography for Detecting Memorization Hotspots and Guiding Data Interventions in Generative Models
von: Patel, Laksh, et al.
Veröffentlicht: (2025)
von: Patel, Laksh, et al.
Veröffentlicht: (2025)
LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
von: Zhang, Kangning, et al.
Veröffentlicht: (2025)
von: Zhang, Kangning, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
von: Patel, Ajay, et al.
Veröffentlicht: (2026) -
StyleDistance: Stronger Content-Independent Style Embeddings with Synthetic Parallel Examples
von: Patel, Ajay, et al.
Veröffentlicht: (2024) -
Position: The Most Expensive Part of an LLM should be its Training Data
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2025) -
Large Language Models Can Self-Improve At Web Agent Tasks
von: Patel, Ajay, et al.
Veröffentlicht: (2024) -
Low-Resource Authorship Style Transfer: Can Non-Famous Authors Be Imitated?
von: Patel, Ajay, et al.
Veröffentlicht: (2022)