GenQA: Generating Millions of Instructions from a Handful of Prompts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Jiuhai, Qadri, Rifaa, Wen, Yuxin, Jain, Neel, Kirchenbauer, John, Zhou, Tianyi, Goldstein, Tom |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FictionalQA: A Dataset for Studying Memorization and Knowledge Acquisition
von: Kirchenbauer, John, et al.
Veröffentlicht: (2025)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2025)
Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs
von: Synk, Ryan, et al.
Veröffentlicht: (2025)
von: Synk, Ryan, et al.
Veröffentlicht: (2025)
OPTune: Efficient Online Preference Tuning
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs
von: Hans, Abhimanyu, et al.
Veröffentlicht: (2024)
von: Hans, Abhimanyu, et al.
Veröffentlicht: (2024)
A Watermark for Large Language Models
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
Zero-Shot Vision Encoder Grafting via LLM Surrogates
von: Yue, Kaiyu, et al.
Veröffentlicht: (2025)
von: Yue, Kaiyu, et al.
Veröffentlicht: (2025)
Multi-Objective Linguistic Control of Large Language Models
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
von: Geiping, Jonas, et al.
Veröffentlicht: (2025)
von: Geiping, Jonas, et al.
Veröffentlicht: (2025)
Can LLMs Speak For Diverse People? Tuning LLMs via Debate to Generate Controllable Controversial Statements
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
Multi-Token Prediction via Self-Distillation
von: Kirchenbauer, John, et al.
Veröffentlicht: (2026)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2026)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
Instruction Tuning and CoT Prompting for Contextual Medical QA with LLMs
von: Le, Chenqian, et al.
Veröffentlicht: (2025)
von: Le, Chenqian, et al.
Veröffentlicht: (2025)
From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning
von: Li, Ming, et al.
Veröffentlicht: (2023)
von: Li, Ming, et al.
Veröffentlicht: (2023)
On the Reliability of Watermarks for Large Language Models
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
Automated Data Curation for Robust Language Model Fine-Tuning
von: Chen, Jiuhai, et al.
Veröffentlicht: (2024)
von: Chen, Jiuhai, et al.
Veröffentlicht: (2024)
Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
LMD3: Language Model Data Density Dependence
von: Kirchenbauer, John, et al.
Veröffentlicht: (2024)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2024)
ExpertGenQA: Open-ended QA generation in Specialized Domains
von: Shahgir, Haz Sameen, et al.
Veröffentlicht: (2025)
von: Shahgir, Haz Sameen, et al.
Veröffentlicht: (2025)
Few-Shot Prompting for Extractive Quranic QA with Instruction-Tuned LLMs
von: Basem, Mohamed, et al.
Veröffentlicht: (2025)
von: Basem, Mohamed, et al.
Veröffentlicht: (2025)
Antidistillation Fingerprinting
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2026)
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2026)
CorpusQA: A 10 Million Token Benchmark for Corpus-Level Analysis and Reasoning
von: Lu, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Lu, Zhiyuan, et al.
Veröffentlicht: (2026)
P-RAG: Prompt-Enhanced Parametric RAG with LoRA and Selective CoT for Biomedical and Multi-Hop QA
von: Lyu, Xingda, et al.
Veröffentlicht: (2026)
von: Lyu, Xingda, et al.
Veröffentlicht: (2026)
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
von: Jain, Neel, et al.
Veröffentlicht: (2024)
von: Jain, Neel, et al.
Veröffentlicht: (2024)
DynaGuard: A Dynamic Guardian Model With User-Defined Policies
von: Hoover, Monte, et al.
Veröffentlicht: (2025)
von: Hoover, Monte, et al.
Veröffentlicht: (2025)
Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
von: Singh, Siddharth, et al.
Veröffentlicht: (2025)
von: Singh, Siddharth, et al.
Veröffentlicht: (2025)
Structured Chain-of-Thought Prompting for Few-Shot Generation of Content-Grounded QA Conversations
von: Sultan, Md Arafat, et al.
Veröffentlicht: (2024)
von: Sultan, Md Arafat, et al.
Veröffentlicht: (2024)
Scaling Instruction-Tuned LLMs to Million-Token Contexts via Hierarchical Synthetic Data Generation
von: He, Linda, et al.
Veröffentlicht: (2025)
von: He, Linda, et al.
Veröffentlicht: (2025)
Coercing LLMs to do and reveal (almost) anything
von: Geiping, Jonas, et al.
Veröffentlicht: (2024)
von: Geiping, Jonas, et al.
Veröffentlicht: (2024)
DataGen: Unified Synthetic Dataset Generation via Large Language Models
von: Huang, Yue, et al.
Veröffentlicht: (2024)
von: Huang, Yue, et al.
Veröffentlicht: (2024)
From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding
von: Zhu, Chiwei, et al.
Veröffentlicht: (2025)
von: Zhu, Chiwei, et al.
Veröffentlicht: (2025)
Monotonic Paraphrasing Improves Generalization of Language Model Prompting
von: Liu, Qin, et al.
Veröffentlicht: (2024)
von: Liu, Qin, et al.
Veröffentlicht: (2024)
OmniGen2: Towards Instruction-Aligned Multimodal Generation
von: Wu, Chenyuan, et al.
Veröffentlicht: (2025)
von: Wu, Chenyuan, et al.
Veröffentlicht: (2025)
$G^2$-Reader: Dual Evolving Graphs for Multimodal Document QA
von: Du, Yaxin, et al.
Veröffentlicht: (2026)
von: Du, Yaxin, et al.
Veröffentlicht: (2026)
MGen: Millions of Naturally Occurring Generics in Context
von: Cilleruelo, Gustavo, et al.
Veröffentlicht: (2025)
von: Cilleruelo, Gustavo, et al.
Veröffentlicht: (2025)
Replacing Multi-Step Assembly of Data Preparation Pipelines with One-Step LLM Pipeline Generation for Table QA
von: Li, Fengyu, et al.
Veröffentlicht: (2026)
von: Li, Fengyu, et al.
Veröffentlicht: (2026)
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
von: Röttger, Paul, et al.
Veröffentlicht: (2025)
von: Röttger, Paul, et al.
Veröffentlicht: (2025)
Where to show Demos in Your Prompt: A Positional Bias of In-Context Learning
von: Cobbina, Kwesi, et al.
Veröffentlicht: (2025)
von: Cobbina, Kwesi, et al.
Veröffentlicht: (2025)
Continuous QA Learning with Structured Prompts
von: Zheng, Yinhe
Veröffentlicht: (2022)
von: Zheng, Yinhe
Veröffentlicht: (2022)
BioGraphletQA: Knowledge-Anchored Generation of Complex QA Datasets
von: Jonker, Richard A. A., et al.
Veröffentlicht: (2026)
von: Jonker, Richard A. A., et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
FictionalQA: A Dataset for Studying Memorization and Knowledge Acquisition
von: Kirchenbauer, John, et al.
Veröffentlicht: (2025) -
Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs
von: Synk, Ryan, et al.
Veröffentlicht: (2025) -
OPTune: Efficient Online Preference Tuning
von: Chen, Lichang, et al.
Veröffentlicht: (2024) -
Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs
von: Hans, Abhimanyu, et al.
Veröffentlicht: (2024) -
A Watermark for Large Language Models
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)