MEIC-DT: Memory-Efficient Incremental Clustering for Long-Text Coreference Resolution with Dual-Threshold Constraints

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Kangyang, Si, Shuzheng, Bai, Yuzhuo, Gao, Cheng, Wang, Zhitong, Huang, Cheng, Shen, Yingli, Han, Yufeng, Li, Wenhao, Kong, Cunliang, Sun, Maosong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917465875808256
author Luo, Kangyang
Si, Shuzheng
Bai, Yuzhuo
Gao, Cheng
Wang, Zhitong
Huang, Cheng
Shen, Yingli
Han, Yufeng
Li, Wenhao
Kong, Cunliang
Sun, Maosong
author_facet Luo, Kangyang
Si, Shuzheng
Bai, Yuzhuo
Gao, Cheng
Wang, Zhitong
Huang, Cheng
Shen, Yingli
Han, Yufeng
Li, Wenhao
Kong, Cunliang
Sun, Maosong
contents In the era of large language models (LLMs), supervised neural methods remain the state-of-the-art (SOTA) for Coreference Resolution. Yet, their full potential is underexplored, particularly in incremental clustering, which faces the critical challenge of balancing efficiency with performance for long texts. To address the limitation, we propose \textbf{MEIC-DT}, a novel dual-threshold, memory-efficient incremental clustering approach based on a lightweight Transformer. MEIC-DT features a dual-threshold constraint mechanism designed to precisely control the Transformer's input scale within a predefined memory budget. This mechanism incorporates a Statistics-Aware Eviction Strategy (\textbf{SAES}), which utilizes distinct statistical profiles from the training and inference phases for intelligent cache management. Furthermore, we introduce an Internal Regularization Policy (\textbf{IRP}) that strategically condenses clusters by selecting the most representative mentions, thereby preserving semantic integrity. Extensive experiments on common benchmarks demonstrate that MEIC-DT achieves highly competitive coreference performance under stringent memory constraints.
format Preprint
id arxiv_https___arxiv_org_abs_2512_24711
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MEIC-DT: Memory-Efficient Incremental Clustering for Long-Text Coreference Resolution with Dual-Threshold Constraints
Luo, Kangyang
Si, Shuzheng
Bai, Yuzhuo
Gao, Cheng
Wang, Zhitong
Huang, Cheng
Shen, Yingli
Han, Yufeng
Li, Wenhao
Kong, Cunliang
Sun, Maosong
Information Retrieval
In the era of large language models (LLMs), supervised neural methods remain the state-of-the-art (SOTA) for Coreference Resolution. Yet, their full potential is underexplored, particularly in incremental clustering, which faces the critical challenge of balancing efficiency with performance for long texts. To address the limitation, we propose \textbf{MEIC-DT}, a novel dual-threshold, memory-efficient incremental clustering approach based on a lightweight Transformer. MEIC-DT features a dual-threshold constraint mechanism designed to precisely control the Transformer's input scale within a predefined memory budget. This mechanism incorporates a Statistics-Aware Eviction Strategy (\textbf{SAES}), which utilizes distinct statistical profiles from the training and inference phases for intelligent cache management. Furthermore, we introduce an Internal Regularization Policy (\textbf{IRP}) that strategically condenses clusters by selecting the most representative mentions, thereby preserving semantic integrity. Extensive experiments on common benchmarks demonstrate that MEIC-DT achieves highly competitive coreference performance under stringent memory constraints.
title MEIC-DT: Memory-Efficient Incremental Clustering for Long-Text Coreference Resolution with Dual-Threshold Constraints
topic Information Retrieval
url https://arxiv.org/abs/2512.24711