RAcQUEt: Unveiling the Dangers of Overlooked Referential Ambiguity in Visual LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Testoni, Alberto, Plank, Barbara, Fernández, Raquel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Asking the Right Question at the Right Time: Human and Model Uncertainty Guidance to Ask Clarification Questions
por: Testoni, Alberto, et al.
Publicado: (2024)
por: Testoni, Alberto, et al.
Publicado: (2024)
Naming, Describing, and Quantifying Visual Objects in Humans and LLMs
por: Testoni, Alberto, et al.
Publicado: (2024)
por: Testoni, Alberto, et al.
Publicado: (2024)
It Depends: Resolving Referential Ambiguity in Minimal Contexts with Commonsense Knowledge
por: Ellinger, Lukas, et al.
Publicado: (2025)
por: Ellinger, Lukas, et al.
Publicado: (2025)
Learning to Ask Informative Questions: Enhancing LLMs with Preference Optimization and Expected Information Gain
por: Mazzaccara, Davide, et al.
Publicado: (2024)
por: Mazzaccara, Davide, et al.
Publicado: (2024)
Interpreting Predictive Probabilities: Model Confidence or Human Label Variation?
por: Baan, Joris, et al.
Publicado: (2024)
por: Baan, Joris, et al.
Publicado: (2024)
Mind the Gap: Benchmarking LLM Uncertainty and Calibration with Specialty-Aware Clinical QA and Reasoning-Based Behavioural Features
por: Testoni, Alberto, et al.
Publicado: (2025)
por: Testoni, Alberto, et al.
Publicado: (2025)
To Know or Not To Know? Analyzing Self-Consistency of Large Language Models under Ambiguity
por: Sedova, Anastasiia, et al.
Publicado: (2024)
por: Sedova, Anastasiia, et al.
Publicado: (2024)
Better Aligned with Survey Respondents or Training Data? Unveiling Political Leanings of LLMs on U.S. Supreme Court Cases
por: Xu, Shanshan, et al.
Publicado: (2025)
por: Xu, Shanshan, et al.
Publicado: (2025)
Do LLMs Give Psychometrically Plausible Responses in Educational Assessments?
por: Säuberli, Andreas, et al.
Publicado: (2025)
por: Säuberli, Andreas, et al.
Publicado: (2025)
Calibrated? Not for Everyone: How Sexual Orientation and Religious Markers Distort LLM Accuracy and Confidence in Medical QA
por: Testoni, Alberto, et al.
Publicado: (2026)
por: Testoni, Alberto, et al.
Publicado: (2026)
Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation
por: Baan, Joris, et al.
Publicado: (2026)
por: Baan, Joris, et al.
Publicado: (2026)
Cognitive Modeling with Scaffolded LLMs: A Case Study of Referential Expression Generation
por: Tsvilodub, Polina, et al.
Publicado: (2024)
por: Tsvilodub, Polina, et al.
Publicado: (2024)
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
por: Bavaresco, Anna, et al.
Publicado: (2024)
por: Bavaresco, Anna, et al.
Publicado: (2024)
Referencing Where to Focus: Improving VisualGrounding with Referential Query
por: Wang, Yabing, et al.
Publicado: (2024)
por: Wang, Yabing, et al.
Publicado: (2024)
Don't Buy it! Reassessing the Ad Understanding Abilities of Contrastive Multimodal Models
por: Bavaresco, A., et al.
Publicado: (2024)
por: Bavaresco, A., et al.
Publicado: (2024)
Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
por: Mondorf, Philipp, et al.
Publicado: (2024)
por: Mondorf, Philipp, et al.
Publicado: (2024)
I Came, I Saw, I Explained: Benchmarking Multimodal LLMs on Figurative Meaning in Memes
por: Zhou, Shijia, et al.
Publicado: (2026)
por: Zhou, Shijia, et al.
Publicado: (2026)
MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs
por: Zhao, Raoyuan, et al.
Publicado: (2025)
por: Zhao, Raoyuan, et al.
Publicado: (2025)
Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set
por: Eichin, Florian, et al.
Publicado: (2025)
por: Eichin, Florian, et al.
Publicado: (2025)
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
por: Mondorf, Philipp, et al.
Publicado: (2024)
por: Mondorf, Philipp, et al.
Publicado: (2024)
Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning
por: Mondorf, Philipp, et al.
Publicado: (2024)
por: Mondorf, Philipp, et al.
Publicado: (2024)
LVLMs are Bad at Overhearing Human Referential Communication
por: Wang, Zhengxiang, et al.
Publicado: (2025)
por: Wang, Zhengxiang, et al.
Publicado: (2025)
MaiBaam Annotation Guidelines
por: Blaschke, Verena, et al.
Publicado: (2024)
por: Blaschke, Verena, et al.
Publicado: (2024)
Towards Self-Referential Analytic Assessment: A Profile-Based Approach to L2 Writing Evaluation with LLMs
por: Bannò, Stefano, et al.
Publicado: (2026)
por: Bannò, Stefano, et al.
Publicado: (2026)
CLIMATELI: Evaluating Entity Linking on Climate Change Data
por: Zhou, Shijia, et al.
Publicado: (2024)
por: Zhou, Shijia, et al.
Publicado: (2024)
Dialetto, ma Quanto Dialetto? Transcribing and Evaluating Dialects on a Continuum
por: Shim, Ryan Soh-Eun, et al.
Publicado: (2024)
por: Shim, Ryan Soh-Eun, et al.
Publicado: (2024)
Exploring the Robustness of Task-oriented Dialogue Systems for Colloquial German Varieties
por: Artemova, Ekaterina, et al.
Publicado: (2024)
por: Artemova, Ekaterina, et al.
Publicado: (2024)
If Probable, Then Acceptable? Understanding Conditional Acceptability Judgments in Large Language Models
por: Orth, Jasmin, et al.
Publicado: (2025)
por: Orth, Jasmin, et al.
Publicado: (2025)
Resource-Lean Lexicon Induction for German Dialects
por: Litschko, Robert, et al.
Publicado: (2026)
por: Litschko, Robert, et al.
Publicado: (2026)
Indirect Question Answering in English, German and Bavarian: A Challenging Task for High- and Low-Resource Languages Alike
por: Winkler, Miriam, et al.
Publicado: (2026)
por: Winkler, Miriam, et al.
Publicado: (2026)
EVADE: LLM-Based Explanation Generation and Validation for Error Detection in NLI
por: Zuo, Longfei, et al.
Publicado: (2025)
por: Zuo, Longfei, et al.
Publicado: (2025)
Add Noise, Tasks, or Layers? MaiNLP at the VarDial 2025 Shared Task on Norwegian Dialectal Slot and Intent Detection
por: Blaschke, Verena, et al.
Publicado: (2025)
por: Blaschke, Verena, et al.
Publicado: (2025)
Standard-to-Dialect Transfer Trends Differ across Text and Speech: A Case Study on Intent and Topic Classification in German Dialects
por: Blaschke, Verena, et al.
Publicado: (2025)
por: Blaschke, Verena, et al.
Publicado: (2025)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
por: Si, Shengyun, et al.
Publicado: (2025)
por: Si, Shengyun, et al.
Publicado: (2025)
From Tools to Teammates: Evaluating LLMs in Multi-Session Coding Interactions
por: Rakotonirina, Nathanaël Carraz, et al.
Publicado: (2025)
por: Rakotonirina, Nathanaël Carraz, et al.
Publicado: (2025)
Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?
por: Bajpai, Ashutosh, et al.
Publicado: (2025)
por: Bajpai, Ashutosh, et al.
Publicado: (2025)
"Seeing the Big through the Small": Can LLMs Approximate Human Judgment Distributions on NLI from a Few Explanations?
por: Chen, Beiduo, et al.
Publicado: (2024)
por: Chen, Beiduo, et al.
Publicado: (2024)
A Comparative Analysis of Ethical and Safety Gaps in LLMs using Relative Danger Coefficient
por: Tereshchenko, Yehor, et al.
Publicado: (2025)
por: Tereshchenko, Yehor, et al.
Publicado: (2025)
Safe in the Future, Dangerous in the Past: Dissecting Temporal and Linguistic Vulnerabilities in LLMs
por: Said, Muhammad Abdullahi, et al.
Publicado: (2025)
por: Said, Muhammad Abdullahi, et al.
Publicado: (2025)
Evaluating Pixel Language Models on Non-Standardized Languages
por: Muñoz-Ortiz, Alberto, et al.
Publicado: (2024)
por: Muñoz-Ortiz, Alberto, et al.
Publicado: (2024)
Ejemplares similares
-
Asking the Right Question at the Right Time: Human and Model Uncertainty Guidance to Ask Clarification Questions
por: Testoni, Alberto, et al.
Publicado: (2024) -
Naming, Describing, and Quantifying Visual Objects in Humans and LLMs
por: Testoni, Alberto, et al.
Publicado: (2024) -
It Depends: Resolving Referential Ambiguity in Minimal Contexts with Commonsense Knowledge
por: Ellinger, Lukas, et al.
Publicado: (2025) -
Learning to Ask Informative Questions: Enhancing LLMs with Preference Optimization and Expected Information Gain
por: Mazzaccara, Davide, et al.
Publicado: (2024) -
Interpreting Predictive Probabilities: Model Confidence or Human Label Variation?
por: Baan, Joris, et al.
Publicado: (2024)