Implicit Causality-biases in humans and LLMs as a tool for benchmarking LLM discourse capabilities
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kankowski, Florian, Solstad, Torgrim, Zarriess, Sina, Bott, Oliver |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models
von: Sieker, Judith, et al.
Veröffentlicht: (2026)
von: Sieker, Judith, et al.
Veröffentlicht: (2026)
LLMs Struggle to Reject False Presuppositions when Misinformation Stakes are High
von: Sieker, Judith, et al.
Veröffentlicht: (2025)
von: Sieker, Judith, et al.
Veröffentlicht: (2025)
SemCSE: Semantic Contrastive Sentence Embeddings Using LLM-Generated Summaries For Scientific Abstracts
von: Brinner, Marc, et al.
Veröffentlicht: (2025)
von: Brinner, Marc, et al.
Veröffentlicht: (2025)
Subword models struggle with word learning, but surprisal hides it
von: Bunzeck, Bastian, et al.
Veröffentlicht: (2025)
von: Bunzeck, Bastian, et al.
Veröffentlicht: (2025)
SceneGram: Conceptualizing and Describing Tangrams in Scene Context
von: Junker, Simeon, et al.
Veröffentlicht: (2025)
von: Junker, Simeon, et al.
Veröffentlicht: (2025)
Child-directed speech facilitates production, not comprehension, in BabyLMs
von: Bunzeck, Bastian, et al.
Veröffentlicht: (2026)
von: Bunzeck, Bastian, et al.
Veröffentlicht: (2026)
The Frequency Confound in Language-Model Surprisal and Metaphor Novelty
von: Momen, Omar, et al.
Veröffentlicht: (2026)
von: Momen, Omar, et al.
Veröffentlicht: (2026)
Resilience through Scene Context in Visual Referring Expression Generation
von: Junker, Simeon, et al.
Veröffentlicht: (2024)
von: Junker, Simeon, et al.
Veröffentlicht: (2024)
Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals unreliable Multi-Turn Behavior in LLMs
von: Lachenmaier, Clara, et al.
Veröffentlicht: (2026)
von: Lachenmaier, Clara, et al.
Veröffentlicht: (2026)
Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political Questions
von: Lachenmaier, Clara, et al.
Veröffentlicht: (2025)
von: Lachenmaier, Clara, et al.
Veröffentlicht: (2025)
Rationalizing Transformer Predictions via End-To-End Differentiable Self-Training
von: Brinner, Marc, et al.
Veröffentlicht: (2025)
von: Brinner, Marc, et al.
Veröffentlicht: (2025)
Efficient Scientific Full Text Classification: The Case of EICAT Impact Assessments
von: Brinner, Marc Felix, et al.
Veröffentlicht: (2025)
von: Brinner, Marc Felix, et al.
Veröffentlicht: (2025)
Model Interpretability and Rationale Extraction by Input Mask Optimization
von: Brinner, Marc, et al.
Veröffentlicht: (2025)
von: Brinner, Marc, et al.
Veröffentlicht: (2025)
Enhancing Domain-Specific Encoder Models with LLM-Generated Data: How to Leverage Ontologies, and How to Do Without Them
von: Brinner, Marc, et al.
Veröffentlicht: (2025)
von: Brinner, Marc, et al.
Veröffentlicht: (2025)
Do Construction Distributions Shape Formal Language Learning In German BabyLMs?
von: Bunzeck, Bastian, et al.
Veröffentlicht: (2025)
von: Bunzeck, Bastian, et al.
Veröffentlicht: (2025)
SemCSE-Multi: Multifaceted and Decodable Embeddings for Aspect-Specific and Interpretable Scientific Domain Mapping
von: Brinner, Marc, et al.
Veröffentlicht: (2025)
von: Brinner, Marc, et al.
Veröffentlicht: (2025)
Reference Games as a Testbed for the Alignment of Model Uncertainty and Clarification Requests
von: Ali, Manar, et al.
Veröffentlicht: (2026)
von: Ali, Manar, et al.
Veröffentlicht: (2026)
Are BabyLMs Deaf to Gricean Maxims? A Pragmatic Evaluation of Sample-efficient Language Models
von: Askari, Raha, et al.
Veröffentlicht: (2025)
von: Askari, Raha, et al.
Veröffentlicht: (2025)
Evaluating Diversity in Automatic Poetry Generation
von: Chen, Yanran, et al.
Veröffentlicht: (2024)
von: Chen, Yanran, et al.
Veröffentlicht: (2024)
Small Language Models Also Work With Small Vocabularies: Probing the Linguistic Abilities of Grapheme- and Phoneme-Based Baby Llamas
von: Bunzeck, Bastian, et al.
Veröffentlicht: (2024)
von: Bunzeck, Bastian, et al.
Veröffentlicht: (2024)
The InviTE Corpus: Annotating Invectives in Tudor English Texts for Computational Modeling
von: Spliethoff, Sophie, et al.
Veröffentlicht: (2025)
von: Spliethoff, Sophie, et al.
Veröffentlicht: (2025)
Are Multimodal Large Language Models Pragmatically Competent Listeners in Simple Reference Resolution Tasks?
von: Junker, Simeon, et al.
Veröffentlicht: (2025)
von: Junker, Simeon, et al.
Veröffentlicht: (2025)
Surprisal and Metaphor Novelty Judgments: Moderate Correlations and Divergent Scaling Effects Revealed by Corpus-Based and Synthetic Datasets
von: Momen, Omar, et al.
Veröffentlicht: (2026)
von: Momen, Omar, et al.
Veröffentlicht: (2026)
AIDBench: A benchmark for evaluating the authorship identification capability of large language models
von: Wen, Zichen, et al.
Veröffentlicht: (2024)
von: Wen, Zichen, et al.
Veröffentlicht: (2024)
LLMs left, right, and center: Assessing GPT's capabilities to label political bias from web domains
von: Hernandes, Raphael, et al.
Veröffentlicht: (2024)
von: Hernandes, Raphael, et al.
Veröffentlicht: (2024)
CausalGraph2LLM: Evaluating LLMs for Causal Queries
von: Sheth, Ivaxi, et al.
Veröffentlicht: (2024)
von: Sheth, Ivaxi, et al.
Veröffentlicht: (2024)
Dialogue Is Not Enough to Make a Communicative BabyLM (But Neither Is Developmentally Inspired Reinforcement Learning)
von: Padovani, Francesca, et al.
Veröffentlicht: (2025)
von: Padovani, Francesca, et al.
Veröffentlicht: (2025)
Do LLMs exhibit human-like response biases? A case study in survey design
von: Tjuatja, Lindia, et al.
Veröffentlicht: (2023)
von: Tjuatja, Lindia, et al.
Veröffentlicht: (2023)
The Illusion of Competence: Evaluating the Effect of Explanations on Users' Mental Models of Visual Question Answering Systems
von: Sieker, Judith, et al.
Veröffentlicht: (2024)
von: Sieker, Judith, et al.
Veröffentlicht: (2024)
Fin-Bias: Comprehensive Evaluation for LLM Decision-Making under human bias in Finance Domain
von: Hu, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Hu, Xiaoyu, et al.
Veröffentlicht: (2026)
Towards Trustworthy Lexical Simplification: Exploring Safety and Efficiency with Small LLMs
von: Hayakawa, Akio, et al.
Veröffentlicht: (2025)
von: Hayakawa, Akio, et al.
Veröffentlicht: (2025)
A word association network methodology for evaluating implicit biases in LLMs compared to humans
von: Abramski, Katherine, et al.
Veröffentlicht: (2025)
von: Abramski, Katherine, et al.
Veröffentlicht: (2025)
An Expert-grounded benchmark of General Purpose LLMs in LCA
von: Donaldson, Artur, et al.
Veröffentlicht: (2025)
von: Donaldson, Artur, et al.
Veröffentlicht: (2025)
Assessing LLM Reasoning Through Implicit Causal Chain Discovery in Climate Discourse
von: Allein, Liesbeth, et al.
Veröffentlicht: (2025)
von: Allein, Liesbeth, et al.
Veröffentlicht: (2025)
Do LLMs exhibit the same commonsense capabilities across languages?
von: Martínez-Murillo, Ivan, et al.
Veröffentlicht: (2025)
von: Martínez-Murillo, Ivan, et al.
Veröffentlicht: (2025)
Explainable Detection of Implicit Influential Patterns in Conversations via Data Augmentation
von: Abdidizaji, Sina, et al.
Veröffentlicht: (2025)
von: Abdidizaji, Sina, et al.
Veröffentlicht: (2025)
Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks
von: Ailem, Melissa, et al.
Veröffentlicht: (2024)
von: Ailem, Melissa, et al.
Veröffentlicht: (2024)
The role of System 1 and System 2 semantic memory structure in human and LLM biases
von: Abramski, Katherine, et al.
Veröffentlicht: (2026)
von: Abramski, Katherine, et al.
Veröffentlicht: (2026)
Is your LLM trapped in a Mental Set? Investigative study on how mental sets affect the reasoning capabilities of LLMs
von: Haq, Saiful, et al.
Veröffentlicht: (2025)
von: Haq, Saiful, et al.
Veröffentlicht: (2025)
Can the capability of Large Language Models be described by human ability? A Meta Study
von: Zan, Mingrui, et al.
Veröffentlicht: (2025)
von: Zan, Mingrui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models
von: Sieker, Judith, et al.
Veröffentlicht: (2026) -
LLMs Struggle to Reject False Presuppositions when Misinformation Stakes are High
von: Sieker, Judith, et al.
Veröffentlicht: (2025) -
SemCSE: Semantic Contrastive Sentence Embeddings Using LLM-Generated Summaries For Scientific Abstracts
von: Brinner, Marc, et al.
Veröffentlicht: (2025) -
Subword models struggle with word learning, but surprisal hides it
von: Bunzeck, Bastian, et al.
Veröffentlicht: (2025) -
SceneGram: Conceptualizing and Describing Tangrams in Scene Context
von: Junker, Simeon, et al.
Veröffentlicht: (2025)