Solving the Challenge Set without Solving the Task: On Winograd Schemas as a Test of Pronominal Coreference Resolution
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Porada, Ian, Cheung, Jackie Chi Kit |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
A Controlled Reevaluation of Coreference Resolution Models
par: Porada, Ian, et autres
Publié: (2024)
par: Porada, Ian, et autres
Publié: (2024)
Challenges to Evaluating the Generalization of Coreference Resolution Models: A Measurement Modeling Perspective
par: Porada, Ian, et autres
Publié: (2023)
par: Porada, Ian, et autres
Publié: (2023)
The Validity of Coreference-based Evaluations of Natural Language Understanding
par: Porada, Ian
Publié: (2026)
par: Porada, Ian
Publié: (2026)
EvoGrad: A Dynamic Take on the Winograd Schema Challenge with Human Adversaries
par: Sun, Jing Han, et autres
Publié: (2024)
par: Sun, Jing Han, et autres
Publié: (2024)
Picturing Ambiguity: A Visual Twist on the Winograd Schema Challenge
par: Park, Brendan, et autres
Publié: (2024)
par: Park, Brendan, et autres
Publié: (2024)
WSC+: Enhancing The Winograd Schema Challenge Using Tree-of-Experts
par: Zahraei, Pardis Sadat, et autres
Publié: (2024)
par: Zahraei, Pardis Sadat, et autres
Publié: (2024)
Thai Winograd Schemas: A Benchmark for Thai Commonsense Reasoning
par: Artkaew, Phakphum
Publié: (2024)
par: Artkaew, Phakphum
Publié: (2024)
$\texttt{COSMIC}$: Mutual Information for Task-Agnostic Summarization Evaluation
par: Darrin, Maxime, et autres
Publié: (2024)
par: Darrin, Maxime, et autres
Publié: (2024)
Testing the Assumptions of Active Learning for Translation Tasks with Few Samples
par: Flores, Lorenzo Jaime Yu, et autres
Publié: (2026)
par: Flores, Lorenzo Jaime Yu, et autres
Publié: (2026)
PreSumm: Predicting Summarization Performance Without Summarizing
par: Koniaev, Steven, et autres
Publié: (2025)
par: Koniaev, Steven, et autres
Publié: (2025)
Findings of the Third Shared Task on Multilingual Coreference Resolution
par: Novák, Michal, et autres
Publié: (2024)
par: Novák, Michal, et autres
Publié: (2024)
Concept-Reversed Winograd Schema Challenge: Evaluating and Improving Robust Reasoning in Large Language Models via Abstraction
par: Han, Kaiqiao, et autres
Publié: (2024)
par: Han, Kaiqiao, et autres
Publié: (2024)
Stochastic Chameleons: Irrelevant Context Hallucinations Reveal Class-Based (Mis)Generalization in LLMs
par: Cheng, Ziling, et autres
Publié: (2025)
par: Cheng, Ziling, et autres
Publié: (2025)
Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations
par: Yu, Lei, et autres
Publié: (2024)
par: Yu, Lei, et autres
Publié: (2024)
Improving the Calibration of Confidence Scores in Text Generation Using the Output Distribution's Characteristics
par: Flores, Lorenzo Jaime Yu, et autres
Publié: (2025)
par: Flores, Lorenzo Jaime Yu, et autres
Publié: (2025)
Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning
par: Flores, Lorenzo Jaime Yu, et autres
Publié: (2026)
par: Flores, Lorenzo Jaime Yu, et autres
Publié: (2026)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
par: Chehbouni, Khaoula, et autres
Publié: (2025)
par: Chehbouni, Khaoula, et autres
Publié: (2025)
Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation
par: Cheng, Ziling, et autres
Publié: (2025)
par: Cheng, Ziling, et autres
Publié: (2025)
Coreference Resolution for Vietnamese Narrative Texts
par: Tran, Hieu-Dai, et autres
Publié: (2025)
par: Tran, Hieu-Dai, et autres
Publié: (2025)
Improving LLMs' Learning for Coreference Resolution
par: Gan, Yujian, et autres
Publié: (2025)
par: Gan, Yujian, et autres
Publié: (2025)
Reasoning over Object Descriptions Improves Coreference Resolution in Task-Based Dialogue Systems
par: Ijurco, Oier, et autres
Publié: (2026)
par: Ijurco, Oier, et autres
Publié: (2026)
Reverse Probing: Evaluating Knowledge Transfer via Finetuned Task Embeddings for Coreference Resolution
par: Anikina, Tatiana, et autres
Publié: (2025)
par: Anikina, Tatiana, et autres
Publié: (2025)
ThaiCoref: Thai Coreference Resolution Dataset
par: Trakuekul, Pontakorn, et autres
Publié: (2024)
par: Trakuekul, Pontakorn, et autres
Publié: (2024)
$(RSA)^2$: A Rhetorical-Strategy-Aware Rational Speech Act Framework for Figurative Language Understanding
par: Piano, Cesare Spinoso-Di, et autres
Publié: (2025)
par: Piano, Cesare Spinoso-Di, et autres
Publié: (2025)
Findings of the Fourth Shared Task on Multilingual Coreference Resolution: Can LLMs Dethrone Traditional Approaches?
par: Novák, Michal, et autres
Publié: (2025)
par: Novák, Michal, et autres
Publié: (2025)
Findings of the Fifth Shared Task on Multilingual Coreference Resolution: Expanding Datasets for Long-Range Entities
par: Novák, Michal, et autres
Publié: (2026)
par: Novák, Michal, et autres
Publié: (2026)
BOOKCOREF: Coreference Resolution at Book Scale
par: Martinelli, Giuliano, et autres
Publié: (2025)
par: Martinelli, Giuliano, et autres
Publié: (2025)
Light Coreference Resolution for Russian with Hierarchical Discourse Features
par: Chistova, Elena, et autres
Publié: (2023)
par: Chistova, Elena, et autres
Publié: (2023)
Isolating Language-Coding from Problem-Solving: Benchmarking LLMs with PseudoEval
par: Wu, Jiarong, et autres
Publié: (2025)
par: Wu, Jiarong, et autres
Publié: (2025)
StateFlow: Enhancing LLM Task-Solving through State-Driven Workflows
par: Wu, Yiran, et autres
Publié: (2024)
par: Wu, Yiran, et autres
Publié: (2024)
Cross-Document Contextual Coreference Resolution in Knowledge Graphs
par: Dong, Zhang, et autres
Publié: (2025)
par: Dong, Zhang, et autres
Publié: (2025)
CorefInst: Leveraging LLMs for Multilingual Coreference Resolution
par: Arslan, Tuğba Pamay, et autres
Publié: (2025)
par: Arslan, Tuğba Pamay, et autres
Publié: (2025)
BioCoref: Benchmarking Biomedical Coreference Resolution with LLMs
par: Salem, Nourah M, et autres
Publié: (2025)
par: Salem, Nourah M, et autres
Publié: (2025)
Interpretable Coreference Resolution Evaluation Using Explicit Semantics
par: Gatti, Bruno, et autres
Publié: (2026)
par: Gatti, Bruno, et autres
Publié: (2026)
Efficient Seq2seq Coreference Resolution Using Entity Representations
par: Grenander, Matt, et autres
Publié: (2025)
par: Grenander, Matt, et autres
Publié: (2025)
On the Paradoxical Interference between Instruction-Following and Task Solving
par: Qi, Yunjia, et autres
Publié: (2026)
par: Qi, Yunjia, et autres
Publié: (2026)
Decoupling Task-Solving and Output Formatting in LLM Generation
par: Deng, Haikang, et autres
Publié: (2025)
par: Deng, Haikang, et autres
Publié: (2025)
Can Vision Language Models Be Adaptive in Mathematics Education? A Learner Model-based Rubric Study
par: Gao, Jie, et autres
Publié: (2026)
par: Gao, Jie, et autres
Publié: (2026)
ECBD: Evidence-Centered Benchmark Design for NLP
par: Liu, Yu Lu, et autres
Publié: (2024)
par: Liu, Yu Lu, et autres
Publié: (2024)
CItruS: Chunked Instruction-aware State Eviction for Long Sequence Modeling
par: Bai, Yu, et autres
Publié: (2024)
par: Bai, Yu, et autres
Publié: (2024)
Documents similaires
-
A Controlled Reevaluation of Coreference Resolution Models
par: Porada, Ian, et autres
Publié: (2024) -
Challenges to Evaluating the Generalization of Coreference Resolution Models: A Measurement Modeling Perspective
par: Porada, Ian, et autres
Publié: (2023) -
The Validity of Coreference-based Evaluations of Natural Language Understanding
par: Porada, Ian
Publié: (2026) -
EvoGrad: A Dynamic Take on the Winograd Schema Challenge with Human Adversaries
par: Sun, Jing Han, et autres
Publié: (2024) -
Picturing Ambiguity: A Visual Twist on the Winograd Schema Challenge
par: Park, Brendan, et autres
Publié: (2024)