Cookbook: A framework for improving LLM generative abilities via programmatic data generating templates
Fuente:
arXiv
Saved in:
| Main Authors: | Narayan, Avanika, Chen, Mayee F., Bhatia, Kush, Ré, Christopher |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry
by: Zhang, Michael, et al.
Published: (2024)
by: Zhang, Michael, et al.
Published: (2024)
Automated Rewards via LLM-Generated Progress Functions
by: Sarukkai, Vishnu, et al.
Published: (2024)
by: Sarukkai, Vishnu, et al.
Published: (2024)
An Information Theoretic Perspective on Agentic System Design
by: He, Shizhe, et al.
Published: (2025)
by: He, Shizhe, et al.
Published: (2025)
Aioli: A Unified Optimization Framework for Language Model Data Mixing
by: Chen, Mayee F., et al.
Published: (2024)
by: Chen, Mayee F., et al.
Published: (2024)
Minions: Cost-efficient Collaboration Between On-device and Cloud Language Models
by: Narayan, Avanika, et al.
Published: (2025)
by: Narayan, Avanika, et al.
Published: (2025)
Olmix: A Framework for Data Mixing Throughout LM Development
by: Chen, Mayee F., et al.
Published: (2026)
by: Chen, Mayee F., et al.
Published: (2026)
Auto-Cypher: Improving LLMs on Cypher generation via LLM-supervised generation-verification framework
by: Tiwari, Aman, et al.
Published: (2024)
by: Tiwari, Aman, et al.
Published: (2024)
Evaluating the fairness of task-adaptive pretraining on unlabeled test data before few-shot text classification
by: Dubey, Kush
Published: (2024)
by: Dubey, Kush
Published: (2024)
Human Texts Are Outliers: Detecting LLM-generated Texts via Out-of-distribution Detection
by: Zeng, Cong, et al.
Published: (2025)
by: Zeng, Cong, et al.
Published: (2025)
OpenJarvis: Personal AI, On Personal Devices
by: Saad-Falcon, Jon, et al.
Published: (2026)
by: Saad-Falcon, Jon, et al.
Published: (2026)
LLM-based feature generation from text for interpretable machine learning
by: Balek, Vojtěch, et al.
Published: (2024)
by: Balek, Vojtěch, et al.
Published: (2024)
What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning
by: Zhu, Yuchang, et al.
Published: (2025)
by: Zhu, Yuchang, et al.
Published: (2025)
Archon: An Architecture Search Framework for Inference-Time Techniques
by: Saad-Falcon, Jon, et al.
Published: (2024)
by: Saad-Falcon, Jon, et al.
Published: (2024)
Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes
by: Arora, Simran, et al.
Published: (2023)
by: Arora, Simran, et al.
Published: (2023)
LLM generation novelty through the lens of semantic similarity
by: Davydov, Philipp, et al.
Published: (2025)
by: Davydov, Philipp, et al.
Published: (2025)
Zero-shot generation of synthetic neurosurgical data with large language models
by: Barr, Austin A., et al.
Published: (2025)
by: Barr, Austin A., et al.
Published: (2025)
Aleph-Alpha-GermanWeb: Improving German-language LLM pre-training with model-based data curation and synthetic data generation
by: Burns, Thomas F, et al.
Published: (2025)
by: Burns, Thomas F, et al.
Published: (2025)
ProdRev: A DNN framework for empowering customers using generative pre-trained transformers
by: Gupta, Aakash, et al.
Published: (2025)
by: Gupta, Aakash, et al.
Published: (2025)
Stylometry recognizes human and LLM-generated texts in short samples
by: Przystalski, Karol, et al.
Published: (2025)
by: Przystalski, Karol, et al.
Published: (2025)
Reveal and Release: Iterative LLM Unlearning with Self-generated Data
by: Xie, Linxi, et al.
Published: (2025)
by: Xie, Linxi, et al.
Published: (2025)
DataComp-LM: In search of the next generation of training sets for language models
by: Li, Jeffrey, et al.
Published: (2024)
by: Li, Jeffrey, et al.
Published: (2024)
Efficient Detection of LLM-generated Texts with a Bayesian Surrogate Model
by: Miao, Yibo, et al.
Published: (2023)
by: Miao, Yibo, et al.
Published: (2023)
SwiftEval: Developing a Language-Specific Benchmark for LLM-generated Code Evaluation
by: Petrukha, Ivan, et al.
Published: (2025)
by: Petrukha, Ivan, et al.
Published: (2025)
Will we run out of data? Limits of LLM scaling based on human-generated data
by: Villalobos, Pablo, et al.
Published: (2022)
by: Villalobos, Pablo, et al.
Published: (2022)
Methods of improving LLM training stability
by: Rybakov, Oleg, et al.
Published: (2024)
by: Rybakov, Oleg, et al.
Published: (2024)
Learning from flowsheets: A generative transformer model for autocompletion of flowsheets
by: Vogel, Gabriel, et al.
Published: (2022)
by: Vogel, Gabriel, et al.
Published: (2022)
Performance of diverse evaluation metrics in NLP-based assessment and text generation of consumer complaints
by: Gao, Peiheng, et al.
Published: (2025)
by: Gao, Peiheng, et al.
Published: (2025)
Generative adversarial networks vs large language models: a comparative study on synthetic tabular data generation
by: Barr, Austin A., et al.
Published: (2025)
by: Barr, Austin A., et al.
Published: (2025)
CRANE: Reasoning with constrained LLM generation
by: Banerjee, Debangshu, et al.
Published: (2025)
by: Banerjee, Debangshu, et al.
Published: (2025)
Machine-generated text detection prevents language model collapse
by: Drayson, George, et al.
Published: (2025)
by: Drayson, George, et al.
Published: (2025)
Smoothie: Label Free Language Model Routing
by: Guha, Neel, et al.
Published: (2024)
by: Guha, Neel, et al.
Published: (2024)
EffiPair: Improving the Efficiency of LLM-generated Code with Relative Contrastive Feedback
by: Hajizadeh, Samira, et al.
Published: (2026)
by: Hajizadeh, Samira, et al.
Published: (2026)
Lightweight reranking for language model generations
by: Jain, Siddhartha, et al.
Published: (2023)
by: Jain, Siddhartha, et al.
Published: (2023)
AIDetx: a compression-based method for identification of machine-learning generated text
by: Almeida, Leonardo, et al.
Published: (2024)
by: Almeida, Leonardo, et al.
Published: (2024)
Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning
by: Sclar, Melanie, et al.
Published: (2024)
by: Sclar, Melanie, et al.
Published: (2024)
Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation
by: Rädsch, Tim, et al.
Published: (2025)
by: Rädsch, Tim, et al.
Published: (2025)
The Transformer Cookbook
by: Yang, Andy, et al.
Published: (2025)
by: Yang, Andy, et al.
Published: (2025)
Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling
by: Fashi, Parsa Ashrafi, et al.
Published: (2026)
by: Fashi, Parsa Ashrafi, et al.
Published: (2026)
Using Machine Learning to Distinguish Human-written from Machine-generated Creative Fiction
by: McGlinchey, Andrea Cristina, et al.
Published: (2024)
by: McGlinchey, Andrea Cristina, et al.
Published: (2024)
SPLAT: A framework for optimised GPU code-generation for SParse reguLar ATtention
by: Gupta, Ahan, et al.
Published: (2024)
by: Gupta, Ahan, et al.
Published: (2024)
Similar Items
-
The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry
by: Zhang, Michael, et al.
Published: (2024) -
Automated Rewards via LLM-Generated Progress Functions
by: Sarukkai, Vishnu, et al.
Published: (2024) -
An Information Theoretic Perspective on Agentic System Design
by: He, Shizhe, et al.
Published: (2025) -
Aioli: A Unified Optimization Framework for Language Model Data Mixing
by: Chen, Mayee F., et al.
Published: (2024) -
Minions: Cost-efficient Collaboration Between On-device and Cloud Language Models
by: Narayan, Avanika, et al.
Published: (2025)