SELF-GUIDE: Better Task-Specific Instruction Following via Self-Synthetic Finetuning
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhao, Chenyang, Jia, Xueying, Viswanathan, Vijay, Wu, Tongshuang, Neubig, Graham |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Synthetic Multimodal Question Generation
di: Wu, Ian, et al.
Pubblicazione: (2024)
di: Wu, Ian, et al.
Pubblicazione: (2024)
Better Synthetic Data by Retrieving and Transforming Existing Datasets
di: Gandhi, Saumya, et al.
Pubblicazione: (2024)
di: Gandhi, Saumya, et al.
Pubblicazione: (2024)
Better Instruction-Following Through Minimum Bayes Risk
di: Wu, Ian, et al.
Pubblicazione: (2024)
di: Wu, Ian, et al.
Pubblicazione: (2024)
Training Task Experts through Retrieval Based Distillation
di: Ge, Jiaxin, et al.
Pubblicazione: (2024)
di: Ge, Jiaxin, et al.
Pubblicazione: (2024)
Checklists Are Better Than Reward Models For Aligning Language Models
di: Viswanathan, Vijay, et al.
Pubblicazione: (2025)
di: Viswanathan, Vijay, et al.
Pubblicazione: (2025)
Asking What Matters: Reward-Driven Clarification for Software Engineering Tasks
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2026)
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2026)
Light-IF: Endowing LLMs with Generalizable Reasoning via Preview and Self-Checking for Complex Instruction Following
di: Wang, Chenyang, et al.
Pubblicazione: (2025)
di: Wang, Chenyang, et al.
Pubblicazione: (2025)
Training Versatile Coding Agents in Synthetic Environments
di: Zhu, Yiqi, et al.
Pubblicazione: (2025)
di: Zhu, Yiqi, et al.
Pubblicazione: (2025)
Instruction-tuned Language Models are Better Knowledge Learners
di: Jiang, Zhengbao, et al.
Pubblicazione: (2024)
di: Jiang, Zhengbao, et al.
Pubblicazione: (2024)
CoEvol: Constructing Better Responses for Instruction Finetuning through Multi-Agent Cooperation
di: Li, Renhao, et al.
Pubblicazione: (2024)
di: Li, Renhao, et al.
Pubblicazione: (2024)
Effective Strategies for Asynchronous Software Engineering Agents
di: Geng, Jiayi, et al.
Pubblicazione: (2026)
di: Geng, Jiayi, et al.
Pubblicazione: (2026)
Jatmo: Prompt Injection Defense by Task-Specific Finetuning
di: Piet, Julien, et al.
Pubblicazione: (2023)
di: Piet, Julien, et al.
Pubblicazione: (2023)
Countering Catastrophic Forgetting of Large Language Models for Better Instruction Following via Weight-Space Model Merging
di: Lyu, Mengxian, et al.
Pubblicazione: (2026)
di: Lyu, Mengxian, et al.
Pubblicazione: (2026)
Self-Guided Plan Extraction for Instruction-Following Tasks with Goal-Conditional Reinforcement Learning
di: Volovikova, Zoya, et al.
Pubblicazione: (2026)
di: Volovikova, Zoya, et al.
Pubblicazione: (2026)
ImpRIF: Stronger Implicit Reasoning Leads to Better Complex Instruction Following
di: Yang, Yuancheng, et al.
Pubblicazione: (2026)
di: Yang, Yuancheng, et al.
Pubblicazione: (2026)
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
di: Chern, Steffi, et al.
Pubblicazione: (2024)
di: Chern, Steffi, et al.
Pubblicazione: (2024)
MoR: Better Handling Diverse Queries with a Mixture of Sparse, Dense, and Human Retrievers
di: Kalra, Jushaan Singh, et al.
Pubblicazione: (2025)
di: Kalra, Jushaan Singh, et al.
Pubblicazione: (2025)
On Instruction-Finetuning Neural Machine Translation Models
di: Raunak, Vikas, et al.
Pubblicazione: (2024)
di: Raunak, Vikas, et al.
Pubblicazione: (2024)
Improving Multilingual Instruction Finetuning via Linguistically Natural and Diverse Datasets
di: Indurthi, Sathish Reddy, et al.
Pubblicazione: (2024)
di: Indurthi, Sathish Reddy, et al.
Pubblicazione: (2024)
Diffusion Language Models Can Perform Many Tasks with Scaling and Instruction-Finetuning
di: Ye, Jiasheng, et al.
Pubblicazione: (2023)
di: Ye, Jiasheng, et al.
Pubblicazione: (2023)
cAST: Enhancing Code Retrieval-Augmented Generation with Structural Chunking via Abstract Syntax Tree
di: Zhang, Yilin, et al.
Pubblicazione: (2025)
di: Zhang, Yilin, et al.
Pubblicazione: (2025)
Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following
di: Ren, Qingyu, et al.
Pubblicazione: (2025)
di: Ren, Qingyu, et al.
Pubblicazione: (2025)
Self-Review Framework for Enhancing Instruction Following Capability of LLM
di: Park, Sihyun
Pubblicazione: (2025)
di: Park, Sihyun
Pubblicazione: (2025)
MuSC: Improving Complex Instruction Following with Multi-granularity Self-Contrastive Training
di: Huang, Hui, et al.
Pubblicazione: (2025)
di: Huang, Hui, et al.
Pubblicazione: (2025)
BAP v2: An Enhanced Task Framework for Instruction Following in Minecraft Dialogues
di: Jayannavar, Prashant, et al.
Pubblicazione: (2025)
di: Jayannavar, Prashant, et al.
Pubblicazione: (2025)
Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks
di: Xu, Tianze, et al.
Pubblicazione: (2026)
di: Xu, Tianze, et al.
Pubblicazione: (2026)
Task--Specificity Score: Measuring How Much Instructions Really Matter for Supervision
di: Kadasi, Pritam, et al.
Pubblicazione: (2026)
di: Kadasi, Pritam, et al.
Pubblicazione: (2026)
Diverse and Fine-Grained Instruction-Following Ability Exploration with Synthetic Data
di: Gu, Zihui, et al.
Pubblicazione: (2024)
di: Gu, Zihui, et al.
Pubblicazione: (2024)
ClusterFusion: Hybrid Clustering with Embedding Guidance and LLM Adaptation
di: Xu, Yiming, et al.
Pubblicazione: (2025)
di: Xu, Yiming, et al.
Pubblicazione: (2025)
Improving Task Diversity in Label Efficient Supervised Finetuning of LLMs
di: Arabelly, Abhinav, et al.
Pubblicazione: (2025)
di: Arabelly, Abhinav, et al.
Pubblicazione: (2025)
Finetuning Generative Large Language Models with Discrimination Instructions for Knowledge Graph Completion
di: Liu, Yang, et al.
Pubblicazione: (2024)
di: Liu, Yang, et al.
Pubblicazione: (2024)
What Is Wrong with My Model? Identifying Systematic Problems with Semantic Data Slicing
di: Yang, Chenyang, et al.
Pubblicazione: (2024)
di: Yang, Chenyang, et al.
Pubblicazione: (2024)
An Incomplete Loop: Instruction Inference, Instruction Following, and In-context Learning in Language Models
di: Liu, Emmy, et al.
Pubblicazione: (2024)
di: Liu, Emmy, et al.
Pubblicazione: (2024)
Not-Just-Scaling Laws: Towards a Better Understanding of the Downstream Impact of Language Model Design Decisions
di: Liu, Emmy, et al.
Pubblicazione: (2025)
di: Liu, Emmy, et al.
Pubblicazione: (2025)
Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models
di: Sun, Wangtao, et al.
Pubblicazione: (2024)
di: Sun, Wangtao, et al.
Pubblicazione: (2024)
The Instruction Gap: LLMs get lost in Following Instruction
di: Tripathi, Vishesh, et al.
Pubblicazione: (2025)
di: Tripathi, Vishesh, et al.
Pubblicazione: (2025)
SELF: Self-Evolution with Language Feedback
di: Lu, Jianqiao, et al.
Pubblicazione: (2023)
di: Lu, Jianqiao, et al.
Pubblicazione: (2023)
Zero-Shot Classification of Crisis Tweets Using Instruction-Finetuned Large Language Models
di: McDaniel, Emma, et al.
Pubblicazione: (2024)
di: McDaniel, Emma, et al.
Pubblicazione: (2024)
Thinking LLMs: General Instruction Following with Thought Generation
di: Wu, Tianhao, et al.
Pubblicazione: (2024)
di: Wu, Tianhao, et al.
Pubblicazione: (2024)
Robo-Instruct: Simulator-Augmented Instruction Alignment For Finetuning Code LLMs
di: Hu, Zichao, et al.
Pubblicazione: (2024)
di: Hu, Zichao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Synthetic Multimodal Question Generation
di: Wu, Ian, et al.
Pubblicazione: (2024) -
Better Synthetic Data by Retrieving and Transforming Existing Datasets
di: Gandhi, Saumya, et al.
Pubblicazione: (2024) -
Better Instruction-Following Through Minimum Bayes Risk
di: Wu, Ian, et al.
Pubblicazione: (2024) -
Training Task Experts through Retrieval Based Distillation
di: Ge, Jiaxin, et al.
Pubblicazione: (2024) -
Checklists Are Better Than Reward Models For Aligning Language Models
di: Viswanathan, Vijay, et al.
Pubblicazione: (2025)