Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Jijie, Du, Li, Zhao, Hanyu, Zhang, Bo-wen, Wang, Liangdong, Gao, Boyan, Liu, Guang, Lin, Yonghua |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Towards the Information Boundary of Instruction Sets: The Infinity Instruct Subject Technical Report
by: Du, Li, et al.
Published: (2025)
by: Du, Li, et al.
Published: (2025)
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models
by: Liu, Guang, et al.
Published: (2025)
by: Liu, Guang, et al.
Published: (2025)
ReTok: Replacing Tokenizer to Enhance Representation Efficiency in Large Language Model
by: Gu, Shuhao, et al.
Published: (2024)
by: Gu, Shuhao, et al.
Published: (2024)
InCo-DPO: Balancing Distribution Shift and Data Quality for Enhanced Preference Optimization
by: Wang, Yunan, et al.
Published: (2025)
by: Wang, Yunan, et al.
Published: (2025)
AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies
by: Zhang, Bo-Wen, et al.
Published: (2024)
by: Zhang, Bo-Wen, et al.
Published: (2024)
Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
by: Gu, Shuhao, et al.
Published: (2024)
by: Gu, Shuhao, et al.
Published: (2024)
Aquila2 Technical Report
by: Zhang, Bo-Wen, et al.
Published: (2024)
by: Zhang, Bo-Wen, et al.
Published: (2024)
Beyond IID: Optimizing Instruction Learning from the Perspective of Instruction Interaction and Dependency
by: Zhao, Hanyu, et al.
Published: (2024)
by: Zhao, Hanyu, et al.
Published: (2024)
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
by: Jia, Yiming, et al.
Published: (2025)
by: Jia, Yiming, et al.
Published: (2025)
BioInstruct: Instruction Tuning of Large Language Models for Biomedical Natural Language Processing
by: Tran, Hieu, et al.
Published: (2023)
by: Tran, Hieu, et al.
Published: (2023)
CCI3.0-HQ: a large-scale Chinese dataset of high quality designed for pre-training large language models
by: Wang, Liangdong, et al.
Published: (2024)
by: Wang, Liangdong, et al.
Published: (2024)
LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
by: Ren, Huimin, et al.
Published: (2025)
by: Ren, Huimin, et al.
Published: (2025)
Learning to Instruct for Visual Instruction Tuning
by: Zhou, Zhihan, et al.
Published: (2025)
by: Zhou, Zhihan, et al.
Published: (2025)
Are Large Language Models Good Statisticians?
by: Zhu, Yizhang, et al.
Published: (2024)
by: Zhu, Yizhang, et al.
Published: (2024)
InfinityMATH: A Scalable Instruction Tuning Dataset in Programmatic Mathematical Reasoning
by: Zhang, Bo-Wen, et al.
Published: (2024)
by: Zhang, Bo-Wen, et al.
Published: (2024)
RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented Instructions
by: Liu, Wanlong, et al.
Published: (2024)
by: Liu, Wanlong, et al.
Published: (2024)
InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining
by: Wang, Boxin, et al.
Published: (2023)
by: Wang, Boxin, et al.
Published: (2023)
How Large Language Models Need Symbolism
by: Deng, Xiaotie, et al.
Published: (2025)
by: Deng, Xiaotie, et al.
Published: (2025)
InstructAV: Instruction Fine-tuning Large Language Models for Authorship Verification
by: Hu, Yujia, et al.
Published: (2024)
by: Hu, Yujia, et al.
Published: (2024)
MM-Instruct: Generated Visual Instructions for Large Multimodal Model Alignment
by: Liu, Jihao, et al.
Published: (2024)
by: Liu, Jihao, et al.
Published: (2024)
EasyInstruct: An Easy-to-use Instruction Processing Framework for Large Language Models
by: Ou, Yixin, et al.
Published: (2024)
by: Ou, Yixin, et al.
Published: (2024)
Ultra-Fast Language Generation via Discrete Diffusion Divergence Instruct
by: Zheng, Haoyang, et al.
Published: (2025)
by: Zheng, Haoyang, et al.
Published: (2025)
Safer-Instruct: Aligning Language Models with Automated Preference Data
by: Shi, Taiwei, et al.
Published: (2023)
by: Shi, Taiwei, et al.
Published: (2023)
Automatic Instruction Evolving for Large Language Models
by: Zeng, Weihao, et al.
Published: (2024)
by: Zeng, Weihao, et al.
Published: (2024)
Robo-Instruct: Simulator-Augmented Instruction Alignment For Finetuning Code LLMs
by: Hu, Zichao, et al.
Published: (2024)
by: Hu, Zichao, et al.
Published: (2024)
WizardCoder: Empowering Code Large Language Models with Evol-Instruct
by: Luo, Ziyang, et al.
Published: (2023)
by: Luo, Ziyang, et al.
Published: (2023)
Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
InverseCoder: Self-improving Instruction-Tuned Code LLMs with Inverse-Instruct
by: Wu, Yutong, et al.
Published: (2024)
by: Wu, Yutong, et al.
Published: (2024)
ExpertPrompting: Instructing Large Language Models to be Distinguished Experts
by: Xu, Benfeng, et al.
Published: (2023)
by: Xu, Benfeng, et al.
Published: (2023)
Binary Neural Networks for Large Language Model: A Survey
by: Liu, Liangdong, et al.
Published: (2025)
by: Liu, Liangdong, et al.
Published: (2025)
Ada-Instruct: Adapting Instruction Generators for Complex Reasoning
by: Cui, Wanyun, et al.
Published: (2023)
by: Cui, Wanyun, et al.
Published: (2023)
Instruct-Imagen: Image Generation with Multi-modal Instruction
by: Hu, Hexiang, et al.
Published: (2024)
by: Hu, Hexiang, et al.
Published: (2024)
Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis
by: Yuan, Lin, et al.
Published: (2025)
by: Yuan, Lin, et al.
Published: (2025)
Data Selection for Multi-turn Dialogue Instruction Tuning
by: Li, Bo, et al.
Published: (2026)
by: Li, Bo, et al.
Published: (2026)
BASE-Q: Bias and Asymmetric Scaling Enhanced Rotational Quantization for Large Language Models
by: He, Liulu, et al.
Published: (2025)
by: He, Liulu, et al.
Published: (2025)
Self-Instructed Derived Prompt Generation Meets In-Context Learning: Unlocking New Potential of Black-Box LLMs
by: Li, Zhuo, et al.
Published: (2024)
by: Li, Zhuo, et al.
Published: (2024)
InstructCMP: Length Control in Sentence Compression through Instruction-based Large Language Models
by: Juseon-Do, et al.
Published: (2024)
by: Juseon-Do, et al.
Published: (2024)
CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
by: Wang, Peiding, et al.
Published: (2025)
by: Wang, Peiding, et al.
Published: (2025)
InstructCoder: Instruction Tuning Large Language Models for Code Editing
by: Li, Kaixin, et al.
Published: (2023)
by: Li, Kaixin, et al.
Published: (2023)
P-Aligner: Enabling Pre-Alignment of Language Models via Principled Instruction Synthesis
by: Song, Feifan, et al.
Published: (2025)
by: Song, Feifan, et al.
Published: (2025)
Similar Items
-
Scaling Towards the Information Boundary of Instruction Sets: The Infinity Instruct Subject Technical Report
by: Du, Li, et al.
Published: (2025) -
CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models
by: Liu, Guang, et al.
Published: (2025) -
ReTok: Replacing Tokenizer to Enhance Representation Efficiency in Large Language Model
by: Gu, Shuhao, et al.
Published: (2024) -
InCo-DPO: Balancing Distribution Shift and Data Quality for Enhanced Preference Optimization
by: Wang, Yunan, et al.
Published: (2025) -
AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies
by: Zhang, Bo-Wen, et al.
Published: (2024)