Can LLMs Detect Ambiguous Plural Reference? An Analysis of Split-Antecedent and Mereological Reference
Fuente:
arXiv
Saved in:
| Main Authors: | Anh, Dang, Nouwen, Rick, Poesio, Massimo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Effectiveness of LLMs as Annotators: A Comparative Overview and Empirical Analysis of Direct Representation
by: Pavlovic, Maja, et al.
Published: (2024)
by: Pavlovic, Maja, et al.
Published: (2024)
Data Augmentation for Fake Reviews Detection in Multiple Languages and Multiple Domains
by: Liu, Ming, et al.
Published: (2025)
by: Liu, Ming, et al.
Published: (2025)
VAQUUM: Are Vague Quantifiers Grounded in Visual Data?
by: Wong, Hugh Mee, et al.
Published: (2025)
by: Wong, Hugh Mee, et al.
Published: (2025)
When Models Decide and When They Bind: A Two-Stage Computation for Multiple-Choice Question-Answering
by: Wong, Hugh Mee, et al.
Published: (2026)
by: Wong, Hugh Mee, et al.
Published: (2026)
Making Dialogue Grounding Data Rich: A Three-Tier Data Synthesis Framework for Generalized Referring Expression Comprehension
by: Shao, Juexi, et al.
Published: (2025)
by: Shao, Juexi, et al.
Published: (2025)
Large Language Models as Minecraft Agents
by: Madge, Chris, et al.
Published: (2024)
by: Madge, Chris, et al.
Published: (2024)
Integrating knowledge bases to improve coreference and bridging resolution for the chemical domain
by: Lu, Pengcheng, et al.
Published: (2024)
by: Lu, Pengcheng, et al.
Published: (2024)
A LLM Benchmark based on the Minecraft Builder Dialog Agent Task
by: Madge, Chris, et al.
Published: (2024)
by: Madge, Chris, et al.
Published: (2024)
Understanding The Effect Of Temperature On Alignment With Human Opinions
by: Pavlovic, Maja, et al.
Published: (2024)
by: Pavlovic, Maja, et al.
Published: (2024)
MDC-R: The Minecraft Dialogue Corpus with Reference
by: Madge, Chris, et al.
Published: (2025)
by: Madge, Chris, et al.
Published: (2025)
Improving LLMs' Learning for Coreference Resolution
by: Gan, Yujian, et al.
Published: (2025)
by: Gan, Yujian, et al.
Published: (2025)
Referential ambiguity and clarification requests: comparing human and LLM behaviour
by: Madge, Chris, et al.
Published: (2025)
by: Madge, Chris, et al.
Published: (2025)
Grounded Misunderstandings in Asymmetric Dialogue: A Perspectivist Annotation Scheme for MapTask
by: Li, Nan, et al.
Published: (2025)
by: Li, Nan, et al.
Published: (2025)
Assessing the Reliability of LLMs Annotations in the Context of Demographic Bias and Model Explanation
by: Mohammadi, Hadi, et al.
Published: (2025)
by: Mohammadi, Hadi, et al.
Published: (2025)
An Assessment of Human vs. Model Uncertainty in Soft-Label Learning and Calibration
by: Pavlovic, Maja, et al.
Published: (2026)
by: Pavlovic, Maja, et al.
Published: (2026)
Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
Human Label Variation in Implicit Discourse Relation Recognition
by: Yung, Frances, et al.
Published: (2026)
by: Yung, Frances, et al.
Published: (2026)
Extending Activation Steering to Broad Skills and Multiple Behaviours
by: van der Weij, Teun, et al.
Published: (2024)
by: van der Weij, Teun, et al.
Published: (2024)
Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't
by: Dang, Quy-Anh, et al.
Published: (2025)
by: Dang, Quy-Anh, et al.
Published: (2025)
Factuality or Fiction? Benchmarking Modern LLMs on Ambiguous QA with Citations
by: Patel, Maya, et al.
Published: (2024)
by: Patel, Maya, et al.
Published: (2024)
ClarQ-LLM: A Benchmark for Models Clarifying and Requesting Information in Task-Oriented Dialog
by: Gan, Yujian, et al.
Published: (2024)
by: Gan, Yujian, et al.
Published: (2024)
References Matter: Investigating the Impact of Reference Set Variation on Summarization Evaluation
by: Casola, Silvia, et al.
Published: (2025)
by: Casola, Silvia, et al.
Published: (2025)
When LLMs Struggle: Reference-less Translation Evaluation for Low-resource Languages
by: Sindhujan, Archchana, et al.
Published: (2025)
by: Sindhujan, Archchana, et al.
Published: (2025)
HalluCounter: Reference-free LLM Hallucination Detection in the Wild!
by: Urlana, Ashok, et al.
Published: (2025)
by: Urlana, Ashok, et al.
Published: (2025)
Detecting Reference Errors in Scientific Literature with Large Language Models
by: Zhang, Tianmai M., et al.
Published: (2024)
by: Zhang, Tianmai M., et al.
Published: (2024)
Can VLMs Recall Factual Associations From Visual References?
by: Ashok, Dhananjay, et al.
Published: (2025)
by: Ashok, Dhananjay, et al.
Published: (2025)
Can LLMs Detect Their Own Hallucinations?
by: Kadotani, Sora, et al.
Published: (2025)
by: Kadotani, Sora, et al.
Published: (2025)
Ground Every Sentence: Improving Retrieval-Augmented LLMs with Interleaved Reference-Claim Generation
by: Xia, Sirui, et al.
Published: (2024)
by: Xia, Sirui, et al.
Published: (2024)
PluralLLM: Pluralistic Alignment in LLMs via Federated Learning
by: Srewa, Mahmoud, et al.
Published: (2025)
by: Srewa, Mahmoud, et al.
Published: (2025)
Coordinates from Context: Using LLMs to Ground Complex Location References
by: Masis, Tessa, et al.
Published: (2025)
by: Masis, Tessa, et al.
Published: (2025)
Can Vision Language Models Learn from Visual Demonstrations of Ambiguous Spatial Reasoning?
by: Zhao, Bowen, et al.
Published: (2024)
by: Zhao, Bowen, et al.
Published: (2024)
CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
by: Dong, Qihua, et al.
Published: (2025)
by: Dong, Qihua, et al.
Published: (2025)
Reference-Free Evaluation of Taxonomies
by: Wullschleger, Pascal, et al.
Published: (2025)
by: Wullschleger, Pascal, et al.
Published: (2025)
Evaluating Optimal Reference Translations
by: Zouhar, Vilém, et al.
Published: (2023)
by: Zouhar, Vilém, et al.
Published: (2023)
ReSpark: Leveraging Previous Data Reports as References to Generate New Reports with LLMs
by: Tian, Yuan, et al.
Published: (2025)
by: Tian, Yuan, et al.
Published: (2025)
Where Norms and References Collide: Evaluating LLMs on Normative Reasoning
by: Abrams, Mitchell, et al.
Published: (2026)
by: Abrams, Mitchell, et al.
Published: (2026)
PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs
by: Fu, Xiao, et al.
Published: (2025)
by: Fu, Xiao, et al.
Published: (2025)
Agentic Verification for Ambiguous Query Disambiguation
by: Lee, Youngwon, et al.
Published: (2025)
by: Lee, Youngwon, et al.
Published: (2025)
Polyglot-Lion: Efficient Multilingual ASR for Singapore via Balanced Fine-Tuning of Qwen3-ASR
by: Dang, Quy-Anh, et al.
Published: (2026)
by: Dang, Quy-Anh, et al.
Published: (2026)
Reference-free Hallucination Detection for Large Vision-Language Models
by: Li, Qing, et al.
Published: (2024)
by: Li, Qing, et al.
Published: (2024)
Similar Items
-
The Effectiveness of LLMs as Annotators: A Comparative Overview and Empirical Analysis of Direct Representation
by: Pavlovic, Maja, et al.
Published: (2024) -
Data Augmentation for Fake Reviews Detection in Multiple Languages and Multiple Domains
by: Liu, Ming, et al.
Published: (2025) -
VAQUUM: Are Vague Quantifiers Grounded in Visual Data?
by: Wong, Hugh Mee, et al.
Published: (2025) -
When Models Decide and When They Bind: A Two-Stage Computation for Multiple-Choice Question-Answering
by: Wong, Hugh Mee, et al.
Published: (2026) -
Making Dialogue Grounding Data Rich: A Three-Tier Data Synthesis Framework for Generalized Referring Expression Comprehension
by: Shao, Juexi, et al.
Published: (2025)