Mosaic-IT: Cost-Free Compositional Data Synthesis for Instruction Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Ming, Chen, Pei, Wang, Chenguang, Zhao, Hongyu, Liang, Yijun, Hou, Yupeng, Liu, Fuxiao, Zhou, Tianyi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
InstructGraph: Boosting Large Language Models via Graph-centric Instruction Tuning and Preference Alignment
by: Wang, Jianing, et al.
Published: (2024)
by: Wang, Jianing, et al.
Published: (2024)
From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning
by: Li, Ming, et al.
Published: (2023)
by: Li, Ming, et al.
Published: (2023)
MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning
by: Liu, Fuxiao, et al.
Published: (2023)
by: Liu, Fuxiao, et al.
Published: (2023)
RuleR: Improving LLM Controllability by Rule-based Data Recycling
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness
by: Liang, Yijun, et al.
Published: (2025)
by: Liang, Yijun, et al.
Published: (2025)
Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
by: Liu, Fuxiao, et al.
Published: (2023)
by: Liu, Fuxiao, et al.
Published: (2023)
BenTo: Benchmark Task Reduction with In-Context Transferability
by: Zhao, Hongyu, et al.
Published: (2024)
by: Zhao, Hongyu, et al.
Published: (2024)
CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
CoachLM: Automatic Instruction Revisions Improve the Data Quality in LLM Instruction Tuning
by: Liu, Yilun, et al.
Published: (2023)
by: Liu, Yilun, et al.
Published: (2023)
How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
System-2 Mathematical Reasoning via Enriched Instruction Tuning
by: Cai, Huanqia, et al.
Published: (2024)
by: Cai, Huanqia, et al.
Published: (2024)
Facilitating Multi-turn Function Calling for LLMs via Compositional Instruction Tuning
by: Chen, Mingyang, et al.
Published: (2024)
by: Chen, Mingyang, et al.
Published: (2024)
Contrastive Instruction Tuning
by: Yan, Tianyi Lorena, et al.
Published: (2024)
by: Yan, Tianyi Lorena, et al.
Published: (2024)
Instruction-Tuning Data Synthesis from Scratch via Web Reconstruction
by: Jiang, Yuxin, et al.
Published: (2025)
by: Jiang, Yuxin, et al.
Published: (2025)
ClimateChat: Designing Data and Methods for Instruction Tuning LLMs to Answer Climate Change Queries
by: Chen, Zhou, et al.
Published: (2025)
by: Chen, Zhou, et al.
Published: (2025)
Re-Tuning: Overcoming the Compositionality Limits of Large Language Models with Recursive Tuning
by: Pasewark, Eric, et al.
Published: (2024)
by: Pasewark, Eric, et al.
Published: (2024)
Instruction Following without Instruction Tuning
by: Hewitt, John, et al.
Published: (2024)
by: Hewitt, John, et al.
Published: (2024)
Data Diversity Matters for Robust Instruction Tuning
by: Bukharin, Alexander, et al.
Published: (2023)
by: Bukharin, Alexander, et al.
Published: (2023)
Your Vision-Language Model Itself Is a Strong Filter: Towards High-Quality Instruction Tuning with Data Selection
by: Chen, Ruibo, et al.
Published: (2024)
by: Chen, Ruibo, et al.
Published: (2024)
Can LLMs Speak For Diverse People? Tuning LLMs via Debate to Generate Controllable Controversial Statements
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
Chain-of-Instructions: Compositional Instruction Tuning on Large Language Models
by: Hayati, Shirley Anugrah, et al.
Published: (2024)
by: Hayati, Shirley Anugrah, et al.
Published: (2024)
RECOST: External Knowledge Guided Data-efficient Instruction Tuning
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning
by: Yu, Fangxu, et al.
Published: (2025)
by: Yu, Fangxu, et al.
Published: (2025)
Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning
by: Du, Yanrui, et al.
Published: (2024)
by: Du, Yanrui, et al.
Published: (2024)
Less is More: High-value Data Selection for Visual Instruction Tuning
by: Liu, Zikang, et al.
Published: (2024)
by: Liu, Zikang, et al.
Published: (2024)
DS$^2$-Instruct: Domain-Specific Data Synthesis for Large Language Models Instruction Tuning
by: Xu, Ruiyao, et al.
Published: (2026)
by: Xu, Ruiyao, et al.
Published: (2026)
FedMosaic: Federated Retrieval-Augmented Generation via Parametric Adapters
by: Liang, Zhilin, et al.
Published: (2026)
by: Liang, Zhilin, et al.
Published: (2026)
Importance-Aware Data Selection for Efficient LLM Instruction Tuning
by: Jiang, Tingyu, et al.
Published: (2025)
by: Jiang, Tingyu, et al.
Published: (2025)
Generalizing From Short to Long: Effective Data Synthesis for Long-Context Instruction Tuning
by: Zhu, Wenhao, et al.
Published: (2025)
by: Zhu, Wenhao, et al.
Published: (2025)
MIDB: Multilingual Instruction Data Booster for Enhancing Cultural Equality in Multilingual Instruction Synthesis
by: Liu, Yilun, et al.
Published: (2025)
by: Liu, Yilun, et al.
Published: (2025)
Dual Instruction Tuning with Large Language Models for Mathematical Reasoning
by: Zhou, Yongwei, et al.
Published: (2024)
by: Zhou, Yongwei, et al.
Published: (2024)
MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models
by: Tu, Jianhong, et al.
Published: (2024)
by: Tu, Jianhong, et al.
Published: (2024)
AzeroS: Extending LLM to Speech with Self-Generated Instruction-Free Tuning
by: Shao, Yiwen, et al.
Published: (2025)
by: Shao, Yiwen, et al.
Published: (2025)
GIFT: Guided Fine-Tuning and Transfer for Enhancing Instruction-Tuned Language Models
by: Ruan, Zhiwen, et al.
Published: (2026)
by: Ruan, Zhiwen, et al.
Published: (2026)
Beyond Similarity: A Gradient-based Graph Method for Instruction Tuning Data Selection
by: Zhao, Yang, et al.
Published: (2025)
by: Zhao, Yang, et al.
Published: (2025)
Exploring Format Consistency for Instruction Tuning
by: Liang, Shihao, et al.
Published: (2023)
by: Liang, Shihao, et al.
Published: (2023)
ATLaS: Agent Tuning via Learning Critical Steps
by: Chen, Zhixun, et al.
Published: (2025)
by: Chen, Zhixun, et al.
Published: (2025)
IterSelectTune: An Iterative Training Framework for Efficient Instruction-Tuning Data Selection
by: Song, Jielin, et al.
Published: (2024)
by: Song, Jielin, et al.
Published: (2024)
Similar Items
-
Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning
by: Li, Ming, et al.
Published: (2024) -
Selective Reflection-Tuning: Student-Selected Data Recycling for LLM Instruction-Tuning
by: Li, Ming, et al.
Published: (2024) -
InstructGraph: Boosting Large Language Models via Graph-centric Instruction Tuning and Preference Alignment
by: Wang, Jianing, et al.
Published: (2024) -
From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning
by: Li, Ming, et al.
Published: (2023) -
MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning
by: Liu, Fuxiao, et al.
Published: (2023)