FANNO: Augmenting High-Quality Instruction Data with Open-Sourced LLMs Only
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, He, Su, Junyou, Lun, Tianle, Tao, Yicheng, Zhang, Wenjia, Fan, Zipei, Chen, Guanhua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TAG-INSTRUCT: Controlled Instruction Complexity Enhancement through Structure-based Augmentation
von: Zhu, He, et al.
Veröffentlicht: (2025)
von: Zhu, He, et al.
Veröffentlicht: (2025)
PlanGPT: Enhancing Urban Planning with Tailored Language Model and Efficient Retrieval
von: Zhu, He, et al.
Veröffentlicht: (2024)
von: Zhu, He, et al.
Veröffentlicht: (2024)
PlanGPT-VL: Enhancing Urban Planning with Domain-Specific Vision-Language Models
von: Zhu, He, et al.
Veröffentlicht: (2025)
von: Zhu, He, et al.
Veröffentlicht: (2025)
Anchored Supervised Fine-Tuning
von: Zhu, He, et al.
Veröffentlicht: (2025)
von: Zhu, He, et al.
Veröffentlicht: (2025)
InstructDiff: Domain-Adaptive Data Selection via Differential Entropy for Efficient LLM Fine-Tuning
von: Su, Junyou, et al.
Veröffentlicht: (2026)
von: Su, Junyou, et al.
Veröffentlicht: (2026)
CoachLM: Automatic Instruction Revisions Improve the Data Quality in LLM Instruction Tuning
von: Liu, Yilun, et al.
Veröffentlicht: (2023)
von: Liu, Yilun, et al.
Veröffentlicht: (2023)
Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study
von: Zhu, Yuqi, et al.
Veröffentlicht: (2025)
von: Zhu, Yuqi, et al.
Veröffentlicht: (2025)
Self-adaptive Multimodal Retrieval-Augmented Generation
von: Zhai, Wenjia
Veröffentlicht: (2024)
von: Zhai, Wenjia
Veröffentlicht: (2024)
LLM-Detector: Improving AI-Generated Chinese Text Detection with Open-Source LLM Instruction Tuning
von: Wang, Rongsheng, et al.
Veröffentlicht: (2024)
von: Wang, Rongsheng, et al.
Veröffentlicht: (2024)
Span-level Emotion-Cause-Category Triplet Extraction with Instruction Tuning LLMs and Data Augmentation
von: Li, Xiangju, et al.
Veröffentlicht: (2025)
von: Li, Xiangju, et al.
Veröffentlicht: (2025)
How Do Your Code LLMs Perform? Empowering Code Instruction Tuning with High-Quality Data
von: Wang, Yejie, et al.
Veröffentlicht: (2024)
von: Wang, Yejie, et al.
Veröffentlicht: (2024)
For-Value: Efficient Forward-Only Data Valuation for finetuning LLMs and VLMs
von: Deng, Wenlong, et al.
Veröffentlicht: (2025)
von: Deng, Wenlong, et al.
Veröffentlicht: (2025)
Open (Clinical) LLMs are Sensitive to Instruction Phrasings
von: Arroyo, Alberto Mario Ceballos, et al.
Veröffentlicht: (2024)
von: Arroyo, Alberto Mario Ceballos, et al.
Veröffentlicht: (2024)
MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
von: Nguyen, Huu, et al.
Veröffentlicht: (2025)
von: Nguyen, Huu, et al.
Veröffentlicht: (2025)
Multi-Layer Ranking with Large Language Models for News Source Recommendation
von: Zhang, Wenjia, et al.
Veröffentlicht: (2024)
von: Zhang, Wenjia, et al.
Veröffentlicht: (2024)
Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision
von: Gopal, Shreyas, et al.
Veröffentlicht: (2026)
von: Gopal, Shreyas, et al.
Veröffentlicht: (2026)
ConInstruct: Evaluating Large Language Models on Conflict Detection and Resolution in Instructions
von: He, Xingwei, et al.
Veröffentlicht: (2025)
von: He, Xingwei, et al.
Veröffentlicht: (2025)
CLUES: Collaborative High-Quality Data Selection for LLMs via Training Dynamics
von: Zhao, Wanru, et al.
Veröffentlicht: (2025)
von: Zhao, Wanru, et al.
Veröffentlicht: (2025)
Retrieval Augmented Instruction Tuning for Open NER with Large Language Models
von: Xie, Tingyu, et al.
Veröffentlicht: (2024)
von: Xie, Tingyu, et al.
Veröffentlicht: (2024)
Beyond Human-Only: Evaluating Human-Machine Collaboration for Collecting High-Quality Translation Data
von: Liu, Zhongtao, et al.
Veröffentlicht: (2024)
von: Liu, Zhongtao, et al.
Veröffentlicht: (2024)
DARE-bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science
von: Shu, Fan, et al.
Veröffentlicht: (2026)
von: Shu, Fan, et al.
Veröffentlicht: (2026)
PACIT: Unlocking the Power of Examples for Better In-Context Instruction Tuning
von: Xue, Tianci, et al.
Veröffentlicht: (2023)
von: Xue, Tianci, et al.
Veröffentlicht: (2023)
Your Vision-Language Model Itself Is a Strong Filter: Towards High-Quality Instruction Tuning with Data Selection
von: Chen, Ruibo, et al.
Veröffentlicht: (2024)
von: Chen, Ruibo, et al.
Veröffentlicht: (2024)
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
von: Toshniwal, Shubham, et al.
Veröffentlicht: (2024)
von: Toshniwal, Shubham, et al.
Veröffentlicht: (2024)
Minimum Tuning to Unlock Long Output from LLMs with High Quality Data as the Key
von: Chen, Yingda, et al.
Veröffentlicht: (2024)
von: Chen, Yingda, et al.
Veröffentlicht: (2024)
Scaling Instruction-Tuned LLMs to Million-Token Contexts via Hierarchical Synthetic Data Generation
von: He, Linda, et al.
Veröffentlicht: (2025)
von: He, Linda, et al.
Veröffentlicht: (2025)
ToolBridge: An Open-Source Dataset to Equip LLMs with External Tool Capabilities
von: Jin, Zhenchao, et al.
Veröffentlicht: (2024)
von: Jin, Zhenchao, et al.
Veröffentlicht: (2024)
RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented Instructions
von: Liu, Wanlong, et al.
Veröffentlicht: (2024)
von: Liu, Wanlong, et al.
Veröffentlicht: (2024)
WanJuanSiLu: A High-Quality Open-Source Webtext Dataset for Low-Resource Languages
von: Yu, Jia, et al.
Veröffentlicht: (2025)
von: Yu, Jia, et al.
Veröffentlicht: (2025)
Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
von: Gu, Shuhao, et al.
Veröffentlicht: (2024)
von: Gu, Shuhao, et al.
Veröffentlicht: (2024)
SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution
von: Xie, Chengxing, et al.
Veröffentlicht: (2025)
von: Xie, Chengxing, et al.
Veröffentlicht: (2025)
Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!
von: Zhang, Zhexin, et al.
Veröffentlicht: (2025)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2025)
From Sufficiency to Reflection: Reinforcement-Guided Thinking Quality in Retrieval-Augmented Reasoning for LLMs
von: He, Jie, et al.
Veröffentlicht: (2025)
von: He, Jie, et al.
Veröffentlicht: (2025)
GIFT: Guided Fine-Tuning and Transfer for Enhancing Instruction-Tuned Language Models
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2026)
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2026)
Fabricator: An Open Source Toolkit for Generating Labeled Training Data with Teacher LLMs
von: Golde, Jonas, et al.
Veröffentlicht: (2023)
von: Golde, Jonas, et al.
Veröffentlicht: (2023)
From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
von: Li, Tianle, et al.
Veröffentlicht: (2024)
von: Li, Tianle, et al.
Veröffentlicht: (2024)
OpenSeeker: Democratizing Frontier Search Agents by Fully Open-Sourcing Training Data
von: Du, Yuwen, et al.
Veröffentlicht: (2026)
von: Du, Yuwen, et al.
Veröffentlicht: (2026)
Enhancing Uncertainty Estimation in LLMs with Expectation of Aggregated Internal Belief
von: Xiao, Zeguan, et al.
Veröffentlicht: (2025)
von: Xiao, Zeguan, et al.
Veröffentlicht: (2025)
Graphical Reasoning: LLM-based Semi-Open Relation Extraction
von: Tao, Yicheng, et al.
Veröffentlicht: (2024)
von: Tao, Yicheng, et al.
Veröffentlicht: (2024)
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
von: Chiang, Wei-Lin, et al.
Veröffentlicht: (2024)
von: Chiang, Wei-Lin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TAG-INSTRUCT: Controlled Instruction Complexity Enhancement through Structure-based Augmentation
von: Zhu, He, et al.
Veröffentlicht: (2025) -
PlanGPT: Enhancing Urban Planning with Tailored Language Model and Efficient Retrieval
von: Zhu, He, et al.
Veröffentlicht: (2024) -
PlanGPT-VL: Enhancing Urban Planning with Domain-Specific Vision-Language Models
von: Zhu, He, et al.
Veröffentlicht: (2025) -
Anchored Supervised Fine-Tuning
von: Zhu, He, et al.
Veröffentlicht: (2025) -
InstructDiff: Domain-Adaptive Data Selection via Differential Entropy for Efficient LLM Fine-Tuning
von: Su, Junyou, et al.
Veröffentlicht: (2026)