CobSeg: Coherence Boundary Modeling for Dialogue Topic Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Sijin, Zhao, Liangbin, Cai, Jiaxiang, Deng, Ming, Luo, Mingyu, Fu, Xiuju
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916063948570624
author Sun, Sijin
Zhao, Liangbin
Cai, Jiaxiang
Deng, Ming
Luo, Mingyu
Fu, Xiuju
author_facet Sun, Sijin
Zhao, Liangbin
Cai, Jiaxiang
Deng, Ming
Luo, Mingyu
Fu, Xiuju
contents Dialogue topic segmentation is critical in many human-AI collaborative applications which requires identifying heterogeneous boundary cues, including lexical transitions near utterance edges and semantic discontinuities across utterances. Existing utterance models often dilute these local lexical signals. We propose CobSeg, a novel multi-branch architecture that separates coherence-level semantic continuity from lexical boundary transitions and recovers both through directional boundary prediction. CobSeg further uses boundary informativeness weighting to emphasize high-utility utterance positions, and incorporates a corpus-derived topic coherence cue with learned combination weights. While CobSeg is evaluated as a compact trainable segmenter under supervised gold-boundary training and a pseudo-label setting with automatically induced boundaries, it performs enhanced boundary prediction without LLM calls during inference. Across five benchmarks, it improves $P_k$ and $W_d$ particularly when local lexical cues are prominent: under gold supervision, it reduces $P_k$ by 0.7 points and $W_d$ by 0.6 points on VHF, and reaches $P_k$ of 1.0 on DialSeg711; with induced boundaries, it reduces $P_k$ by 14.8 points on VHF, by 1.5 points on DialSeg711, and by 1.1 points on TIAGE, outperforming prior non-LLM approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2605_30668
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CobSeg: Coherence Boundary Modeling for Dialogue Topic Segmentation
Sun, Sijin
Zhao, Liangbin
Cai, Jiaxiang
Deng, Ming
Luo, Mingyu
Fu, Xiuju
Computation and Language
Artificial Intelligence
Dialogue topic segmentation is critical in many human-AI collaborative applications which requires identifying heterogeneous boundary cues, including lexical transitions near utterance edges and semantic discontinuities across utterances. Existing utterance models often dilute these local lexical signals. We propose CobSeg, a novel multi-branch architecture that separates coherence-level semantic continuity from lexical boundary transitions and recovers both through directional boundary prediction. CobSeg further uses boundary informativeness weighting to emphasize high-utility utterance positions, and incorporates a corpus-derived topic coherence cue with learned combination weights. While CobSeg is evaluated as a compact trainable segmenter under supervised gold-boundary training and a pseudo-label setting with automatically induced boundaries, it performs enhanced boundary prediction without LLM calls during inference. Across five benchmarks, it improves $P_k$ and $W_d$ particularly when local lexical cues are prominent: under gold supervision, it reduces $P_k$ by 0.7 points and $W_d$ by 0.6 points on VHF, and reaches $P_k$ of 1.0 on DialSeg711; with induced boundaries, it reduces $P_k$ by 14.8 points on VHF, by 1.5 points on DialSeg711, and by 1.1 points on TIAGE, outperforming prior non-LLM approaches.
title CobSeg: Coherence Boundary Modeling for Dialogue Topic Segmentation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.30668