Generative Data Refinement: Just Ask for Better Data
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Minqi, Araújo, João G. M., Ellsworth, Will, Gooding, Sian, Grefenstette, Edward |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Interaction Dynamics as a Reward Signal for LLMs
by: Gooding, Sian, et al.
Published: (2025)
by: Gooding, Sian, et al.
Published: (2025)
Writing as a testbed for open ended agents
by: Gooding, Sian, et al.
Published: (2025)
by: Gooding, Sian, et al.
Published: (2025)
Better To Ask in English? Evaluating Factual Accuracy of Multilingual LLMs in English and Low-Resource Languages
by: Rohera, Pritika, et al.
Published: (2025)
by: Rohera, Pritika, et al.
Published: (2025)
minimax: Efficient Baselines for Autocurricula in JAX
by: Jiang, Minqi, et al.
Published: (2023)
by: Jiang, Minqi, et al.
Published: (2023)
A Subgoal-driven Framework for Improving Long-Horizon LLM Agents
by: Wang, Taiyi, et al.
Published: (2026)
by: Wang, Taiyi, et al.
Published: (2026)
Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves
by: Deng, Yihe, et al.
Published: (2023)
by: Deng, Yihe, et al.
Published: (2023)
When and What to Ask: AskBench and Rubric-Guided RLVR for LLM Clarification
by: Zhao, Jiale, et al.
Published: (2026)
by: Zhao, Jiale, et al.
Published: (2026)
Social Learning: Towards Collaborative Learning with Large Language Models
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
Sequence-Level Leakage Risk of Training Data in Large Language Models
by: Tiwari, Trishita, et al.
Published: (2024)
by: Tiwari, Trishita, et al.
Published: (2024)
GRIP: Geometric Refinement and Adaptive Information Potential for Data Efficiency
by: Wang, Changhao, et al.
Published: (2026)
by: Wang, Changhao, et al.
Published: (2026)
KVSculpt: KV Cache Compression as Distillation
by: Jiang, Bo, et al.
Published: (2026)
by: Jiang, Bo, et al.
Published: (2026)
Refined Direct Preference Optimization with Synthetic Data for Behavioral Alignment of LLMs
by: Gallego, Víctor
Published: (2024)
by: Gallego, Víctor
Published: (2024)
Better as Generators Than Classifiers: Leveraging LLMs and Synthetic Data for Low-Resource Multilingual Classification
by: Pecher, Branislav, et al.
Published: (2026)
by: Pecher, Branislav, et al.
Published: (2026)
Learning How to Ask: Querying LMs with Mixtures of Soft Prompts
by: Qin, Guanghui, et al.
Published: (2021)
by: Qin, Guanghui, et al.
Published: (2021)
Understanding the Effects of RLHF on LLM Generalisation and Diversity
by: Kirk, Robert, et al.
Published: (2023)
by: Kirk, Robert, et al.
Published: (2023)
Less Finetuning, Better Retrieval: Rethinking LLM Adaptation for Biomedical Retrievers via Synthetic Data and Model Merging
by: Khattab, Sameh, et al.
Published: (2026)
by: Khattab, Sameh, et al.
Published: (2026)
Objective Metrics for Evaluating Large Language Models Using External Data Sources
by: Du, Haoze, et al.
Published: (2025)
by: Du, Haoze, et al.
Published: (2025)
Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use
by: Kumar, Abhijit, et al.
Published: (2026)
by: Kumar, Abhijit, et al.
Published: (2026)
Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
by: Ruis, Laura, et al.
Published: (2024)
by: Ruis, Laura, et al.
Published: (2024)
Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text Classification
by: Kuo, Hsun-Yu, et al.
Published: (2024)
by: Kuo, Hsun-Yu, et al.
Published: (2024)
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows
by: Patel, Ajay, et al.
Published: (2024)
by: Patel, Ajay, et al.
Published: (2024)
Intent-Aware Schema Generation And Refinement For Literature Review Tables
by: Padmakumar, Vishakh, et al.
Published: (2025)
by: Padmakumar, Vishakh, et al.
Published: (2025)
Fine-Tuning Language Models with Just Forward Passes
by: Malladi, Sadhika, et al.
Published: (2023)
by: Malladi, Sadhika, et al.
Published: (2023)
RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?
by: Xu, Haotian, et al.
Published: (2025)
by: Xu, Haotian, et al.
Published: (2025)
Ask Again, Then Fail: Large Language Models' Vacillations in Judgment
by: Xie, Qiming, et al.
Published: (2023)
by: Xie, Qiming, et al.
Published: (2023)
Towards Better Open-Ended Text Generation: A Multicriteria Evaluation Framework
by: Arias, Esteban Garces, et al.
Published: (2024)
by: Arias, Esteban Garces, et al.
Published: (2024)
Better Call SAUL: Fluent and Consistent Language Model Editing with Generation Regularization
by: Wang, Mingyang, et al.
Published: (2024)
by: Wang, Mingyang, et al.
Published: (2024)
Smaller Language Models are Better Black-box Machine-Generated Text Detectors
by: Mireshghallah, Niloofar, et al.
Published: (2023)
by: Mireshghallah, Niloofar, et al.
Published: (2023)
Teaching Models to Understand (but not Generate) High-risk Data
by: Wang, Ryan, et al.
Published: (2025)
by: Wang, Ryan, et al.
Published: (2025)
An Automatic Prompt Generation System for Tabular Data Tasks
by: Akella, Ashlesha, et al.
Published: (2024)
by: Akella, Ashlesha, et al.
Published: (2024)
Data Augmentations for Improved (Large) Language Model Generalization
by: Feder, Amir, et al.
Published: (2023)
by: Feder, Amir, et al.
Published: (2023)
Improving Next Tokens via Second-to-Last Predictions with Generate and Refine
by: Schneider, Johannes
Published: (2024)
by: Schneider, Johannes
Published: (2024)
Cross-Refine: Improving Natural Language Explanation Generation by Learning in Tandem
by: Wang, Qianli, et al.
Published: (2024)
by: Wang, Qianli, et al.
Published: (2024)
Better, Not Just More: Data-Centric Machine Learning for Earth Observation
by: Roscher, Ribana, et al.
Published: (2023)
by: Roscher, Ribana, et al.
Published: (2023)
Reasoning Models Don't Just Think Longer, They Move Differently
by: Gjølbye, Anders, et al.
Published: (2026)
by: Gjølbye, Anders, et al.
Published: (2026)
Is Data Shapley Not Better than Random in Data Selection? Ask NASH
by: Tian, Xiao, et al.
Published: (2026)
by: Tian, Xiao, et al.
Published: (2026)
OptiMer: Optimal Distribution Vector Merging Is Better than Data Mixing for Continual Pre-Training
by: Song, Haiyue, et al.
Published: (2026)
by: Song, Haiyue, et al.
Published: (2026)
Towards Active Synthetic Data Generation for Finetuning Language Models
by: Kessler, Samuel, et al.
Published: (2025)
by: Kessler, Samuel, et al.
Published: (2025)
Seeing to Generalize: How Visual Data Corrects Binding Shortcuts
by: Buzeta, Nicolas, et al.
Published: (2026)
by: Buzeta, Nicolas, et al.
Published: (2026)
Balancing Cost and Effectiveness of Synthetic Data Generation Strategies for LLMs
by: Chan, Yung-Chieh, et al.
Published: (2024)
by: Chan, Yung-Chieh, et al.
Published: (2024)
Similar Items
-
Interaction Dynamics as a Reward Signal for LLMs
by: Gooding, Sian, et al.
Published: (2025) -
Writing as a testbed for open ended agents
by: Gooding, Sian, et al.
Published: (2025) -
Better To Ask in English? Evaluating Factual Accuracy of Multilingual LLMs in English and Low-Resource Languages
by: Rohera, Pritika, et al.
Published: (2025) -
minimax: Efficient Baselines for Autocurricula in JAX
by: Jiang, Minqi, et al.
Published: (2023) -
A Subgoal-driven Framework for Improving Long-Horizon LLM Agents
by: Wang, Taiyi, et al.
Published: (2026)