Conversations Gone Awry, But Then? Evaluating Conversational Forecasting Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866908467603701760 |
|---|---|
| author | Tran, Son Quoc Gangavarapu, Tushaar Chernogor, Nicholas Chang, Jonathan P. Danescu-Niculescu-Mizil, Cristian |
| author_facet | Tran, Son Quoc Gangavarapu, Tushaar Chernogor, Nicholas Chang, Jonathan P. Danescu-Niculescu-Mizil, Cristian |
| contents | We often rely on our intuition to anticipate the direction of a conversation. Endowing automated systems with similar foresight can enable them to assist human-human interactions. Recent work on developing models with this predictive capacity has focused on the Conversations Gone Awry (CGA) task: forecasting whether an ongoing conversation will derail. In this work, we revisit this task and introduce the first uniform evaluation framework, creating a benchmark that enables direct and reliable comparisons between different architectures. This allows us to present an up-to-date overview of the current progress in CGA models, in light of recent advancements in language modeling. Our framework also introduces a novel metric that captures a model's ability to revise its forecast as the conversation progresses. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_19470 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Conversations Gone Awry, But Then? Evaluating Conversational Forecasting Models Tran, Son Quoc Gangavarapu, Tushaar Chernogor, Nicholas Chang, Jonathan P. Danescu-Niculescu-Mizil, Cristian Computation and Language Human-Computer Interaction We often rely on our intuition to anticipate the direction of a conversation. Endowing automated systems with similar foresight can enable them to assist human-human interactions. Recent work on developing models with this predictive capacity has focused on the Conversations Gone Awry (CGA) task: forecasting whether an ongoing conversation will derail. In this work, we revisit this task and introduce the first uniform evaluation framework, creating a benchmark that enables direct and reliable comparisons between different architectures. This allows us to present an up-to-date overview of the current progress in CGA models, in light of recent advancements in language modeling. Our framework also introduces a novel metric that captures a model's ability to revise its forecast as the conversation progresses. |
| title | Conversations Gone Awry, But Then? Evaluating Conversational Forecasting Models |
| topic | Computation and Language Human-Computer Interaction |
| url | https://arxiv.org/abs/2507.19470 |