Can Language Models Follow Multiple Turns of Entangled Instructions?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Chi, Liu, Xin, Wang, Haodong, Li, Shiyang, Yang, Jingfeng, Jiang, Haoming, Wang, Zhengyang, Yin, Qingyu, Qiu, Liang, Yu, Changlong, Gao, Yifan, Li, Zheng, Yin, Bing, Shang, Jingbo, Ji, Heng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915504277422080
author Han, Chi
Liu, Xin
Wang, Haodong
Li, Shiyang
Yang, Jingfeng
Jiang, Haoming
Wang, Zhengyang
Yin, Qingyu
Qiu, Liang
Yu, Changlong
Gao, Yifan
Li, Zheng
Yin, Bing
Shang, Jingbo
Ji, Heng
author_facet Han, Chi
Liu, Xin
Wang, Haodong
Li, Shiyang
Yang, Jingfeng
Jiang, Haoming
Wang, Zhengyang
Yin, Qingyu
Qiu, Liang
Yu, Changlong
Gao, Yifan
Li, Zheng
Yin, Bing
Shang, Jingbo
Ji, Heng
contents Despite significant achievements in improving the instruction-following capabilities of large language models (LLMs), the ability to process multiple potentially entangled or conflicting instructions remains a considerable challenge. Real-world scenarios often require consistency across multiple instructions over time, such as secret privacy, personal preferences, and prioritization, which demand sophisticated abilities to integrate multiple turns and carefully balance competing objectives when instructions intersect or conflict. This work presents a systematic investigation of LLMs' capabilities in handling multiple turns of instructions, covering three levels of difficulty: (1) retrieving information from instructions, (2) tracking and reasoning across turns, and (3) resolving conflicts among instructions. We construct MultiTurnInstruct~with $\sim$1.1K high-quality multi-turn conversations through the human-in-the-loop approach and result in nine capability categories, including statics and dynamics, reasoning, and multitasking. Our finding reveals an intriguing trade-off between different capabilities. While GPT models demonstrate superior memorization, they show reduced effectiveness in privacy-protection tasks requiring selective information withholding. Larger models exhibit stronger reasoning capabilities but still struggle with resolving conflicting instructions. Importantly, these performance gaps cannot be attributed solely to information loss, as models demonstrate strong BLEU scores on memorization tasks. Still, their attention mechanisms fail to integrate multiple related instructions effectively. These findings highlight critical areas for improvement in complex real-world tasks involving multi-turn instructions. Data and codes are released at https://github.com/Glaciohound/Multi-Turn-Instruct.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13222
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Language Models Follow Multiple Turns of Entangled Instructions?
Han, Chi
Liu, Xin
Wang, Haodong
Li, Shiyang
Yang, Jingfeng
Jiang, Haoming
Wang, Zhengyang
Yin, Qingyu
Qiu, Liang
Yu, Changlong
Gao, Yifan
Li, Zheng
Yin, Bing
Shang, Jingbo
Ji, Heng
Computation and Language
Artificial Intelligence
Despite significant achievements in improving the instruction-following capabilities of large language models (LLMs), the ability to process multiple potentially entangled or conflicting instructions remains a considerable challenge. Real-world scenarios often require consistency across multiple instructions over time, such as secret privacy, personal preferences, and prioritization, which demand sophisticated abilities to integrate multiple turns and carefully balance competing objectives when instructions intersect or conflict. This work presents a systematic investigation of LLMs' capabilities in handling multiple turns of instructions, covering three levels of difficulty: (1) retrieving information from instructions, (2) tracking and reasoning across turns, and (3) resolving conflicts among instructions. We construct MultiTurnInstruct~with $\sim$1.1K high-quality multi-turn conversations through the human-in-the-loop approach and result in nine capability categories, including statics and dynamics, reasoning, and multitasking. Our finding reveals an intriguing trade-off between different capabilities. While GPT models demonstrate superior memorization, they show reduced effectiveness in privacy-protection tasks requiring selective information withholding. Larger models exhibit stronger reasoning capabilities but still struggle with resolving conflicting instructions. Importantly, these performance gaps cannot be attributed solely to information loss, as models demonstrate strong BLEU scores on memorization tasks. Still, their attention mechanisms fail to integrate multiple related instructions effectively. These findings highlight critical areas for improvement in complex real-world tasks involving multi-turn instructions. Data and codes are released at https://github.com/Glaciohound/Multi-Turn-Instruct.
title Can Language Models Follow Multiple Turns of Entangled Instructions?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2503.13222