FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
Fuente:
arXiv
Guardado en:
| Autores principales: | Patel, Ajay, Raffel, Colin, Callison-Burch, Chris |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows
por: Patel, Ajay, et al.
Publicado: (2024)
por: Patel, Ajay, et al.
Publicado: (2024)
StyleDistance: Stronger Content-Independent Style Embeddings with Synthetic Parallel Examples
por: Patel, Ajay, et al.
Publicado: (2024)
por: Patel, Ajay, et al.
Publicado: (2024)
WithdrarXiv: A Large-Scale Dataset for Retraction Study
por: Rao, Delip, et al.
Publicado: (2024)
por: Rao, Delip, et al.
Publicado: (2024)
Large Language Models Can Self-Improve At Web Agent Tasks
por: Patel, Ajay, et al.
Publicado: (2024)
por: Patel, Ajay, et al.
Publicado: (2024)
Low-Resource Authorship Style Transfer: Can Non-Famous Authors Be Imitated?
por: Patel, Ajay, et al.
Publicado: (2022)
por: Patel, Ajay, et al.
Publicado: (2022)
Position: The Most Expensive Part of an LLM should be its Training Data
por: Kandpal, Nikhil, et al.
Publicado: (2025)
por: Kandpal, Nikhil, et al.
Publicado: (2025)
Genetic Instruct: Scaling up Synthetic Generation of Coding Instructions for Large Language Models
por: Majumdar, Somshubra, et al.
Publicado: (2024)
por: Majumdar, Somshubra, et al.
Publicado: (2024)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
por: Tang, Zhengyang, et al.
Publicado: (2024)
por: Tang, Zhengyang, et al.
Publicado: (2024)
Merging by Matching Models in Task Parameter Subspaces
por: Tam, Derek, et al.
Publicado: (2023)
por: Tam, Derek, et al.
Publicado: (2023)
Simultaneous Masking, Not Prompting Optimization: A Paradigm Shift in Fine-tuning LLMs for Simultaneous Translation
por: Raffel, Matthew, et al.
Publicado: (2024)
por: Raffel, Matthew, et al.
Publicado: (2024)
Non-instructional Fine-tuning: Enabling Instruction-Following Capabilities in Pre-trained Language Models without Instruction-Following Data
por: Xie, Juncheng, et al.
Publicado: (2024)
por: Xie, Juncheng, et al.
Publicado: (2024)
Probabilistic Soundness Guarantees in LLM Reasoning Chains
por: You, Weiqiu, et al.
Publicado: (2025)
por: You, Weiqiu, et al.
Publicado: (2025)
Enhancing and Assessing Instruction-Following with Fine-Grained Instruction Variants
por: Yang, Jiuding, et al.
Publicado: (2024)
por: Yang, Jiuding, et al.
Publicado: (2024)
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
por: Wallace, Eric, et al.
Publicado: (2024)
por: Wallace, Eric, et al.
Publicado: (2024)
Scaling Data-Constrained Language Models
por: Muennighoff, Niklas, et al.
Publicado: (2023)
por: Muennighoff, Niklas, et al.
Publicado: (2023)
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
por: Kang, Feiyang, et al.
Publicado: (2024)
por: Kang, Feiyang, et al.
Publicado: (2024)
On the Loss of Context-awareness in General Instruction Fine-tuning
por: Wang, Yihan, et al.
Publicado: (2024)
por: Wang, Yihan, et al.
Publicado: (2024)
mStyleDistance: Multilingual Style Embeddings and their Evaluation
por: Qiu, Justin, et al.
Publicado: (2025)
por: Qiu, Justin, et al.
Publicado: (2025)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
por: Yang, Yue, et al.
Publicado: (2025)
por: Yang, Yue, et al.
Publicado: (2025)
GenAI Content Detection Task 3: Cross-Domain Machine-Generated Text Detection Challenge
por: Dugan, Liam, et al.
Publicado: (2025)
por: Dugan, Liam, et al.
Publicado: (2025)
SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-Tuning
por: He, Yexiao, et al.
Publicado: (2024)
por: He, Yexiao, et al.
Publicado: (2024)
Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale
por: Hammoud, Hasan Abed Al Kader, et al.
Publicado: (2025)
por: Hammoud, Hasan Abed Al Kader, et al.
Publicado: (2025)
Diffusion Language Models Can Perform Many Tasks with Scaling and Instruction-Finetuning
por: Ye, Jiasheng, et al.
Publicado: (2023)
por: Ye, Jiasheng, et al.
Publicado: (2023)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
por: Yadav, Prateek, et al.
Publicado: (2023)
por: Yadav, Prateek, et al.
Publicado: (2023)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
por: Pan, Bowen, et al.
Publicado: (2024)
por: Pan, Bowen, et al.
Publicado: (2024)
What Do Claim Verification Datasets Actually Test? A Reasoning Trace Analysis
por: Rao, Delip, et al.
Publicado: (2026)
por: Rao, Delip, et al.
Publicado: (2026)
Instruction Fine-Tuning: Does Prompt Loss Matter?
por: Huerta-Enochian, Mathew, et al.
Publicado: (2024)
por: Huerta-Enochian, Mathew, et al.
Publicado: (2024)
Joint Selection for Large-Scale Pre-Training Data via Policy Gradient-based Mask Learning
por: Fan, Ziqing, et al.
Publicado: (2025)
por: Fan, Ziqing, et al.
Publicado: (2025)
ParaGuide: Guided Diffusion Paraphrasers for Plug-and-Play Textual Style Transfer
por: Horvitz, Zachary, et al.
Publicado: (2023)
por: Horvitz, Zachary, et al.
Publicado: (2023)
BeaverTalk: Oregon State University's IWSLT 2025 Simultaneous Speech Translation System
por: Raffel, Matthew, et al.
Publicado: (2025)
por: Raffel, Matthew, et al.
Publicado: (2025)
Autorubric: Unifying Rubric-based LLM Evaluation
por: Rao, Delip, et al.
Publicado: (2026)
por: Rao, Delip, et al.
Publicado: (2026)
BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation
por: Rao, Delip, et al.
Publicado: (2026)
por: Rao, Delip, et al.
Publicado: (2026)
Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities
por: Zhang, Junyan, et al.
Publicado: (2025)
por: Zhang, Junyan, et al.
Publicado: (2025)
OpenBezoar: Small, Cost-Effective and Open Models Trained on Mixes of Instruction Data
por: Dissanayake, Chandeepa, et al.
Publicado: (2024)
por: Dissanayake, Chandeepa, et al.
Publicado: (2024)
Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language Models
por: Fang, Yin, et al.
Publicado: (2023)
por: Fang, Yin, et al.
Publicado: (2023)
Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
por: Kang, Feiyang, et al.
Publicado: (2025)
por: Kang, Feiyang, et al.
Publicado: (2025)
Scaling Law for Quantization-Aware Training
por: Chen, Mengzhao, et al.
Publicado: (2025)
por: Chen, Mengzhao, et al.
Publicado: (2025)
Enhancing Event Reasoning in Large Language Models through Instruction Fine-Tuning with Semantic Causal Graphs
por: Bethany, Mazal, et al.
Publicado: (2024)
por: Bethany, Mazal, et al.
Publicado: (2024)
Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap
por: Zhu, Andrew, et al.
Publicado: (2025)
por: Zhu, Andrew, et al.
Publicado: (2025)
Complex Logical Instruction Generation
por: Zhang, Mian, et al.
Publicado: (2025)
por: Zhang, Mian, et al.
Publicado: (2025)
Ejemplares similares
-
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows
por: Patel, Ajay, et al.
Publicado: (2024) -
StyleDistance: Stronger Content-Independent Style Embeddings with Synthetic Parallel Examples
por: Patel, Ajay, et al.
Publicado: (2024) -
WithdrarXiv: A Large-Scale Dataset for Retraction Study
por: Rao, Delip, et al.
Publicado: (2024) -
Large Language Models Can Self-Improve At Web Agent Tasks
por: Patel, Ajay, et al.
Publicado: (2024) -
Low-Resource Authorship Style Transfer: Can Non-Famous Authors Be Imitated?
por: Patel, Ajay, et al.
Publicado: (2022)