Saved in:
| Main Authors: | Sieker, Judith, Lachenmaier, Clara, Zarrieß, Sina |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2505.22354 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political Questions
by: Lachenmaier, Clara, et al.
Published: (2025)
by: Lachenmaier, Clara, et al.
Published: (2025)
How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models
by: Sieker, Judith, et al.
Published: (2026)
by: Sieker, Judith, et al.
Published: (2026)
Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals unreliable Multi-Turn Behavior in LLMs
by: Lachenmaier, Clara, et al.
Published: (2026)
by: Lachenmaier, Clara, et al.
Published: (2026)
Towards an Analysis of Discourse and Interactional Pragmatic Reasoning Capabilities of Large Language Models
by: Robrecht, Amelie, et al.
Published: (2024)
by: Robrecht, Amelie, et al.
Published: (2024)
Are BabyLMs Deaf to Gricean Maxims? A Pragmatic Evaluation of Sample-efficient Language Models
by: Askari, Raha, et al.
Published: (2025)
by: Askari, Raha, et al.
Published: (2025)
Reference Games as a Testbed for the Alignment of Model Uncertainty and Clarification Requests
by: Ali, Manar, et al.
Published: (2026)
by: Ali, Manar, et al.
Published: (2026)
The Illusion of Competence: Evaluating the Effect of Explanations on Users' Mental Models of Visual Question Answering Systems
by: Sieker, Judith, et al.
Published: (2024)
by: Sieker, Judith, et al.
Published: (2024)
GerPS-Compare: Comparing NER methods for legal norm analysis
by: Bachinger, Sarah T., et al.
Published: (2024)
by: Bachinger, Sarah T., et al.
Published: (2024)
Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions
by: Zhu, Wang Bill, et al.
Published: (2025)
by: Zhu, Wang Bill, et al.
Published: (2025)
Subword models struggle with word learning, but surprisal hides it
by: Bunzeck, Bastian, et al.
Published: (2025)
by: Bunzeck, Bastian, et al.
Published: (2025)
SceneGram: Conceptualizing and Describing Tangrams in Scene Context
by: Junker, Simeon, et al.
Published: (2025)
by: Junker, Simeon, et al.
Published: (2025)
Child-directed speech facilitates production, not comprehension, in BabyLMs
by: Bunzeck, Bastian, et al.
Published: (2026)
by: Bunzeck, Bastian, et al.
Published: (2026)
The Frequency Confound in Language-Model Surprisal and Metaphor Novelty
by: Momen, Omar, et al.
Published: (2026)
by: Momen, Omar, et al.
Published: (2026)
Resilience through Scene Context in Visual Referring Expression Generation
by: Junker, Simeon, et al.
Published: (2024)
by: Junker, Simeon, et al.
Published: (2024)
Rationalizing Transformer Predictions via End-To-End Differentiable Self-Training
by: Brinner, Marc, et al.
Published: (2025)
by: Brinner, Marc, et al.
Published: (2025)
Implicit Causality-biases in humans and LLMs as a tool for benchmarking LLM discourse capabilities
by: Kankowski, Florian, et al.
Published: (2025)
by: Kankowski, Florian, et al.
Published: (2025)
Model Interpretability and Rationale Extraction by Input Mask Optimization
by: Brinner, Marc, et al.
Published: (2025)
by: Brinner, Marc, et al.
Published: (2025)
Efficient Scientific Full Text Classification: The Case of EICAT Impact Assessments
by: Brinner, Marc Felix, et al.
Published: (2025)
by: Brinner, Marc Felix, et al.
Published: (2025)
SemCSE: Semantic Contrastive Sentence Embeddings Using LLM-Generated Summaries For Scientific Abstracts
by: Brinner, Marc, et al.
Published: (2025)
by: Brinner, Marc, et al.
Published: (2025)
SemCSE-Multi: Multifaceted and Decodable Embeddings for Aspect-Specific and Interpretable Scientific Domain Mapping
by: Brinner, Marc, et al.
Published: (2025)
by: Brinner, Marc, et al.
Published: (2025)
Do Construction Distributions Shape Formal Language Learning In German BabyLMs?
by: Bunzeck, Bastian, et al.
Published: (2025)
by: Bunzeck, Bastian, et al.
Published: (2025)
Presupposition and Reasoning in Conditionals: A Theory-Based Study of Humans and LLMs
by: Azin, Tara, et al.
Published: (2026)
by: Azin, Tara, et al.
Published: (2026)
The Presupposition Problem in Representation Genesis
by: Wu, Yiling
Published: (2026)
by: Wu, Yiling
Published: (2026)
Evaluating Reasoning Models for Queries with Presuppositions
by: Sathyanathan, Rose, et al.
Published: (2026)
by: Sathyanathan, Rose, et al.
Published: (2026)
Enhancing Domain-Specific Encoder Models with LLM-Generated Data: How to Leverage Ontologies, and How to Do Without Them
by: Brinner, Marc, et al.
Published: (2025)
by: Brinner, Marc, et al.
Published: (2025)
Evaluating Diversity in Automatic Poetry Generation
by: Chen, Yanran, et al.
Published: (2024)
by: Chen, Yanran, et al.
Published: (2024)
Small Language Models Also Work With Small Vocabularies: Probing the Linguistic Abilities of Grapheme- and Phoneme-Based Baby Llamas
by: Bunzeck, Bastian, et al.
Published: (2024)
by: Bunzeck, Bastian, et al.
Published: (2024)
Safer in Translation? Presupposition Robustness in Indic Languages
by: Palnitkar, Aadi, et al.
Published: (2025)
by: Palnitkar, Aadi, et al.
Published: (2025)
Diverse, but Divisive: LLMs Can Exaggerate Gender Differences in Opinion Related to Harms of Misinformation
by: Neumann, Terrence, et al.
Published: (2024)
by: Neumann, Terrence, et al.
Published: (2024)
The InviTE Corpus: Annotating Invectives in Tudor English Texts for Computational Modeling
by: Spliethoff, Sophie, et al.
Published: (2025)
by: Spliethoff, Sophie, et al.
Published: (2025)
Are Multimodal Large Language Models Pragmatically Competent Listeners in Simple Reference Resolution Tasks?
by: Junker, Simeon, et al.
Published: (2025)
by: Junker, Simeon, et al.
Published: (2025)
Surprisal and Metaphor Novelty Judgments: Moderate Correlations and Divergent Scaling Effects Revealed by Corpus-Based and Synthetic Datasets
by: Momen, Omar, et al.
Published: (2026)
by: Momen, Omar, et al.
Published: (2026)
FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
by: Zhang, Zhehao, et al.
Published: (2025)
by: Zhang, Zhehao, et al.
Published: (2025)
Evaluating Large Language Models for Health-related Queries with Presuppositions
by: Kaur, Navreet, et al.
Published: (2023)
by: Kaur, Navreet, et al.
Published: (2023)
Misinforming LLMs: vulnerabilities, challenges and opportunities
by: Zhou, Bo, et al.
Published: (2024)
by: Zhou, Bo, et al.
Published: (2024)
Dialogue Is Not Enough to Make a Communicative BabyLM (But Neither Is Developmentally Inspired Reinforcement Learning)
by: Padovani, Francesca, et al.
Published: (2025)
by: Padovani, Francesca, et al.
Published: (2025)
If We May De-Presuppose: Robustly Verifying Claims through Presupposition-Free Question Decomposition
by: Dipta, Shubhashis Roy, et al.
Published: (2025)
by: Dipta, Shubhashis Roy, et al.
Published: (2025)
Feedback Friction: LLMs Struggle to Fully Incorporate External Feedback
by: Jiang, Dongwei, et al.
Published: (2025)
by: Jiang, Dongwei, et al.
Published: (2025)
Let's CONFER: A Dataset for Evaluating Natural Language Inference Models on CONditional InFERence and Presupposition
by: Azin, Tara, et al.
Published: (2025)
by: Azin, Tara, et al.
Published: (2025)
Learn and Unlearn: Addressing Misinformation in Multilingual LLMs
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
Similar Items
-
Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political Questions
by: Lachenmaier, Clara, et al.
Published: (2025) -
How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models
by: Sieker, Judith, et al.
Published: (2026) -
Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals unreliable Multi-Turn Behavior in LLMs
by: Lachenmaier, Clara, et al.
Published: (2026) -
Towards an Analysis of Discourse and Interactional Pragmatic Reasoning Capabilities of Large Language Models
by: Robrecht, Amelie, et al.
Published: (2024) -
Are BabyLMs Deaf to Gricean Maxims? A Pragmatic Evaluation of Sample-efficient Language Models
by: Askari, Raha, et al.
Published: (2025)