Challenges to Evaluating the Generalization of Coreference Resolution Models: A Measurement Modeling Perspective
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Porada, Ian, Olteanu, Alexandra, Suleman, Kaheer, Trischler, Adam, Cheung, Jackie Chi Kit |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Controlled Reevaluation of Coreference Resolution Models
von: Porada, Ian, et al.
Veröffentlicht: (2024)
von: Porada, Ian, et al.
Veröffentlicht: (2024)
Solving the Challenge Set without Solving the Task: On Winograd Schemas as a Test of Pronominal Coreference Resolution
von: Porada, Ian, et al.
Veröffentlicht: (2024)
von: Porada, Ian, et al.
Veröffentlicht: (2024)
The Validity of Coreference-based Evaluations of Natural Language Understanding
von: Porada, Ian
Veröffentlicht: (2026)
von: Porada, Ian
Veröffentlicht: (2026)
SCOPE: Language Models as One-Time Teacher for Hierarchical Planning in Text Environments
von: Lu, Haoye, et al.
Veröffentlicht: (2025)
von: Lu, Haoye, et al.
Veröffentlicht: (2025)
ECBD: Evidence-Centered Benchmark Design for NLP
von: Liu, Yu Lu, et al.
Veröffentlicht: (2024)
von: Liu, Yu Lu, et al.
Veröffentlicht: (2024)
Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations
von: Yu, Lei, et al.
Veröffentlicht: (2024)
von: Yu, Lei, et al.
Veröffentlicht: (2024)
Improving Neural Question Generation using World Knowledge
von: Gupta, Deepak, et al.
Veröffentlicht: (2019)
von: Gupta, Deepak, et al.
Veröffentlicht: (2019)
$\texttt{COSMIC}$: Mutual Information for Task-Agnostic Summarization Evaluation
von: Darrin, Maxime, et al.
Veröffentlicht: (2024)
von: Darrin, Maxime, et al.
Veröffentlicht: (2024)
Can Vision Language Models Be Adaptive in Mathematics Education? A Learner Model-based Rubric Study
von: Gao, Jie, et al.
Veröffentlicht: (2026)
von: Gao, Jie, et al.
Veröffentlicht: (2026)
PreSumm: Predicting Summarization Performance Without Summarizing
von: Koniaev, Steven, et al.
Veröffentlicht: (2025)
von: Koniaev, Steven, et al.
Veröffentlicht: (2025)
Stochastic Chameleons: Irrelevant Context Hallucinations Reveal Class-Based (Mis)Generalization in LLMs
von: Cheng, Ziling, et al.
Veröffentlicht: (2025)
von: Cheng, Ziling, et al.
Veröffentlicht: (2025)
Improving the Calibration of Confidence Scores in Text Generation Using the Output Distribution's Characteristics
von: Flores, Lorenzo Jaime Yu, et al.
Veröffentlicht: (2025)
von: Flores, Lorenzo Jaime Yu, et al.
Veröffentlicht: (2025)
"It was 80% me, 20% AI": Seeking Authenticity in Co-Writing with Large Language Models
von: Hwang, Angel Hsing-Chi, et al.
Veröffentlicht: (2024)
von: Hwang, Angel Hsing-Chi, et al.
Veröffentlicht: (2024)
EasyECR: A Library for Easy Implementation and Evaluation of Event Coreference Resolution Models
von: Li, Yuncong, et al.
Veröffentlicht: (2024)
von: Li, Yuncong, et al.
Veröffentlicht: (2024)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
von: Chehbouni, Khaoula, et al.
Veröffentlicht: (2025)
von: Chehbouni, Khaoula, et al.
Veröffentlicht: (2025)
Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning
von: Flores, Lorenzo Jaime Yu, et al.
Veröffentlicht: (2026)
von: Flores, Lorenzo Jaime Yu, et al.
Veröffentlicht: (2026)
Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation
von: Cheng, Ziling, et al.
Veröffentlicht: (2025)
von: Cheng, Ziling, et al.
Veröffentlicht: (2025)
Interpretable Coreference Resolution Evaluation Using Explicit Semantics
von: Gatti, Bruno, et al.
Veröffentlicht: (2026)
von: Gatti, Bruno, et al.
Veröffentlicht: (2026)
Coreference Resolution for Vietnamese Narrative Texts
von: Tran, Hieu-Dai, et al.
Veröffentlicht: (2025)
von: Tran, Hieu-Dai, et al.
Veröffentlicht: (2025)
Improving LLMs' Learning for Coreference Resolution
von: Gan, Yujian, et al.
Veröffentlicht: (2025)
von: Gan, Yujian, et al.
Veröffentlicht: (2025)
CItruS: Chunked Instruction-aware State Eviction for Long Sequence Modeling
von: Bai, Yu, et al.
Veröffentlicht: (2024)
von: Bai, Yu, et al.
Veröffentlicht: (2024)
Comparing Humans and Models on a Similar Scale: Towards Cognitive Gender Bias Evaluation in Coreference Resolution
von: Lior, Gili, et al.
Veröffentlicht: (2023)
von: Lior, Gili, et al.
Veröffentlicht: (2023)
$(RSA)^2$: A Rhetorical-Strategy-Aware Rational Speech Act Framework for Figurative Language Understanding
von: Piano, Cesare Spinoso-Di, et al.
Veröffentlicht: (2025)
von: Piano, Cesare Spinoso-Di, et al.
Veröffentlicht: (2025)
ÚFAL CorPipe at CRAC 2022: Effectivity of Multilingual Models for Coreference Resolution
von: Straka, Milan, et al.
Veröffentlicht: (2022)
von: Straka, Milan, et al.
Veröffentlicht: (2022)
ThaiCoref: Thai Coreference Resolution Dataset
von: Trakuekul, Pontakorn, et al.
Veröffentlicht: (2024)
von: Trakuekul, Pontakorn, et al.
Veröffentlicht: (2024)
CorPipe at CRAC 2025: Evaluating Multilingual Encoders for Multilingual Coreference Resolution
von: Straka, Milan
Veröffentlicht: (2025)
von: Straka, Milan
Veröffentlicht: (2025)
Identifying and Analyzing Performance-Critical Tokens in Large Language Models
von: Bai, Yu, et al.
Veröffentlicht: (2024)
von: Bai, Yu, et al.
Veröffentlicht: (2024)
BOOKCOREF: Coreference Resolution at Book Scale
von: Martinelli, Giuliano, et al.
Veröffentlicht: (2025)
von: Martinelli, Giuliano, et al.
Veröffentlicht: (2025)
Light Coreference Resolution for Russian with Hierarchical Discourse Features
von: Chistova, Elena, et al.
Veröffentlicht: (2023)
von: Chistova, Elena, et al.
Veröffentlicht: (2023)
Findings of the Third Shared Task on Multilingual Coreference Resolution
von: Novák, Michal, et al.
Veröffentlicht: (2024)
von: Novák, Michal, et al.
Veröffentlicht: (2024)
Reverse Probing: Evaluating Knowledge Transfer via Finetuned Task Embeddings for Coreference Resolution
von: Anikina, Tatiana, et al.
Veröffentlicht: (2025)
von: Anikina, Tatiana, et al.
Veröffentlicht: (2025)
Generating Harder Cross-document Event Coreference Resolution Datasets using Metaphoric Paraphrasing
von: Ahmed, Shafiuddin Rehan, et al.
Veröffentlicht: (2024)
von: Ahmed, Shafiuddin Rehan, et al.
Veröffentlicht: (2024)
Enhancing Coreference Resolution with Pretrained Language Models: Bridging the Gap Between Syntax and Semantics
von: Liu, Xingzu, et al.
Veröffentlicht: (2025)
von: Liu, Xingzu, et al.
Veröffentlicht: (2025)
Investigating Intersectional Bias in Large Language Models using Confidence Disparities in Coreference Resolution
von: Khan, Falaah Arif, et al.
Veröffentlicht: (2025)
von: Khan, Falaah Arif, et al.
Veröffentlicht: (2025)
Cross-Document Contextual Coreference Resolution in Knowledge Graphs
von: Dong, Zhang, et al.
Veröffentlicht: (2025)
von: Dong, Zhang, et al.
Veröffentlicht: (2025)
CorefInst: Leveraging LLMs for Multilingual Coreference Resolution
von: Arslan, Tuğba Pamay, et al.
Veröffentlicht: (2025)
von: Arslan, Tuğba Pamay, et al.
Veröffentlicht: (2025)
BioCoref: Benchmarking Biomedical Coreference Resolution with LLMs
von: Salem, Nourah M, et al.
Veröffentlicht: (2025)
von: Salem, Nourah M, et al.
Veröffentlicht: (2025)
Efficient Seq2seq Coreference Resolution Using Entity Representations
von: Grenander, Matt, et al.
Veröffentlicht: (2025)
von: Grenander, Matt, et al.
Veröffentlicht: (2025)
Okay, Let's Do This! Modeling Event Coreference with Generated Rationales and Knowledge Distillation
von: Nath, Abhijnan, et al.
Veröffentlicht: (2024)
von: Nath, Abhijnan, et al.
Veröffentlicht: (2024)
Testing the Assumptions of Active Learning for Translation Tasks with Few Samples
von: Flores, Lorenzo Jaime Yu, et al.
Veröffentlicht: (2026)
von: Flores, Lorenzo Jaime Yu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
A Controlled Reevaluation of Coreference Resolution Models
von: Porada, Ian, et al.
Veröffentlicht: (2024) -
Solving the Challenge Set without Solving the Task: On Winograd Schemas as a Test of Pronominal Coreference Resolution
von: Porada, Ian, et al.
Veröffentlicht: (2024) -
The Validity of Coreference-based Evaluations of Natural Language Understanding
von: Porada, Ian
Veröffentlicht: (2026) -
SCOPE: Language Models as One-Time Teacher for Hierarchical Planning in Text Environments
von: Lu, Haoye, et al.
Veröffentlicht: (2025) -
ECBD: Evidence-Centered Benchmark Design for NLP
von: Liu, Yu Lu, et al.
Veröffentlicht: (2024)