MT-OSC: Path for LLMs that Get Lost in Multi-Turn Conversation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Singh, Jyotika, Tu, Fang, Ballesteros, Miguel, Sun, Weiyi, Ghoshal, Sandip, Yuan, Michelle, Benajiba, Yassine, Ravi, Sujith, Roth, Dan
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910148373512192
author Singh, Jyotika
Tu, Fang
Ballesteros, Miguel
Sun, Weiyi
Ghoshal, Sandip
Yuan, Michelle
Benajiba, Yassine
Ravi, Sujith
Roth, Dan
author_facet Singh, Jyotika
Tu, Fang
Ballesteros, Miguel
Sun, Weiyi
Ghoshal, Sandip
Yuan, Michelle
Benajiba, Yassine
Ravi, Sujith
Roth, Dan
contents Large language models (LLMs) suffer significant performance degradation when user instructions and context are distributed over multiple conversational turns, yet multi-turn (MT) interactions dominate chat interfaces. The routine approach of appending full chat history to prompts rapidly exhausts context windows, leading to increased latency, higher computational costs, and diminishing returns as conversations extend. We introduce MT-OSC, a One-off Sequential Condensation framework that efficiently and automatically condenses chat history in the background without disrupting the user experience. MT-OSC employs a Condenser Agent that uses a few-shot inference-based Condenser and a lightweight Decider to selectively retain essential information, reducing token counts by up to 72% in 10-turn dialogues. Evaluated across 13 state-of-the-art LLMs and diverse multi-turn benchmarks, MT-OSC consistently narrows the multi-turn performance gap - yielding improved or preserved accuracy across datasets while remaining robust to distractors and irrelevant turns. Our results establish MT-OSC as a scalable solution for multi-turn chats, enabling richer context within constrained input spaces, reducing latency and operational cost, while balancing performance.
format Preprint
id arxiv_https___arxiv_org_abs_2604_08782
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MT-OSC: Path for LLMs that Get Lost in Multi-Turn Conversation
Singh, Jyotika
Tu, Fang
Ballesteros, Miguel
Sun, Weiyi
Ghoshal, Sandip
Yuan, Michelle
Benajiba, Yassine
Ravi, Sujith
Roth, Dan
Computation and Language
Large language models (LLMs) suffer significant performance degradation when user instructions and context are distributed over multiple conversational turns, yet multi-turn (MT) interactions dominate chat interfaces. The routine approach of appending full chat history to prompts rapidly exhausts context windows, leading to increased latency, higher computational costs, and diminishing returns as conversations extend. We introduce MT-OSC, a One-off Sequential Condensation framework that efficiently and automatically condenses chat history in the background without disrupting the user experience. MT-OSC employs a Condenser Agent that uses a few-shot inference-based Condenser and a lightweight Decider to selectively retain essential information, reducing token counts by up to 72% in 10-turn dialogues. Evaluated across 13 state-of-the-art LLMs and diverse multi-turn benchmarks, MT-OSC consistently narrows the multi-turn performance gap - yielding improved or preserved accuracy across datasets while remaining robust to distractors and irrelevant turns. Our results establish MT-OSC as a scalable solution for multi-turn chats, enabling richer context within constrained input spaces, reducing latency and operational cost, while balancing performance.
title MT-OSC: Path for LLMs that Get Lost in Multi-Turn Conversation
topic Computation and Language
url https://arxiv.org/abs/2604.08782