Multi-Lingual Implicit Discourse Relation Recognition with Multi-Label Hierarchical Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Costa, Nelson Filipe, Kosseim, Leila
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911127766564864
author Costa, Nelson Filipe
Kosseim, Leila
author_facet Costa, Nelson Filipe
Kosseim, Leila
contents This paper introduces the first multi-lingual and multi-label classification model for implicit discourse relation recognition (IDRR). Our model, HArch, is evaluated on the recently released DiscoGeM 2.0 corpus and leverages hierarchical dependencies between discourse senses to predict probability distributions across all three sense levels in the PDTB 3.0 framework. We compare several pre-trained encoder backbones and find that RoBERTa-HArch achieves the best performance in English, while XLM-RoBERTa-HArch performs best in the multi-lingual setting. In addition, we compare our fine-tuned models against GPT-4o and Llama-4-Maverick using few-shot prompting across all language configurations. Our results show that our fine-tuned models consistently outperform these LLMs, highlighting the advantages of task-specific fine-tuning over prompting in IDRR. Finally, we report SOTA results on the DiscoGeM 1.0 corpus, further validating the effectiveness of our hierarchical approach.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20712
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-Lingual Implicit Discourse Relation Recognition with Multi-Label Hierarchical Learning
Costa, Nelson Filipe
Kosseim, Leila
Computation and Language
This paper introduces the first multi-lingual and multi-label classification model for implicit discourse relation recognition (IDRR). Our model, HArch, is evaluated on the recently released DiscoGeM 2.0 corpus and leverages hierarchical dependencies between discourse senses to predict probability distributions across all three sense levels in the PDTB 3.0 framework. We compare several pre-trained encoder backbones and find that RoBERTa-HArch achieves the best performance in English, while XLM-RoBERTa-HArch performs best in the multi-lingual setting. In addition, we compare our fine-tuned models against GPT-4o and Llama-4-Maverick using few-shot prompting across all language configurations. Our results show that our fine-tuned models consistently outperform these LLMs, highlighting the advantages of task-specific fine-tuning over prompting in IDRR. Finally, we report SOTA results on the DiscoGeM 1.0 corpus, further validating the effectiveness of our hierarchical approach.
title Multi-Lingual Implicit Discourse Relation Recognition with Multi-Label Hierarchical Learning
topic Computation and Language
url https://arxiv.org/abs/2508.20712