Evaluating Pragmatic Reasoning in Large Language Models: Evidence from Scalar Diversity
Fuente:
arXiv
Salvato in:
| Autore principale: | Cho, Ye-eun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Continuous Interpretive Steering for Scalar Diversity
di: Cho, Ye-eun
Pubblicazione: (2026)
di: Cho, Ye-eun
Pubblicazione: (2026)
Pragmatic inference of scalar implicature by LLMs
di: Cho, Ye-eun, et al.
Pubblicazione: (2024)
di: Cho, Ye-eun, et al.
Pubblicazione: (2024)
Probing Large Language Models for Scalar Adjective Lexical Semantics and Scalar Diversity Pragmatics
di: Lin, Fangru, et al.
Pubblicazione: (2024)
di: Lin, Fangru, et al.
Pubblicazione: (2024)
Can Vision-Language Models Infer Speaker's Ignorance? The Role of Visual and Linguistic Cues
di: Cho, Ye-eun, et al.
Pubblicazione: (2025)
di: Cho, Ye-eun, et al.
Pubblicazione: (2025)
Pragmatic Competence Evaluation of Large Language Models for the Korean Language
di: Park, Dojun, et al.
Pubblicazione: (2024)
di: Park, Dojun, et al.
Pubblicazione: (2024)
It's Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers
di: Cho, Yong-eun
Pubblicazione: (2026)
di: Cho, Yong-eun
Pubblicazione: (2026)
CEI: A Benchmark for Evaluating Pragmatic Reasoning in Language Models
di: Chun, Jon, et al.
Pubblicazione: (2026)
di: Chun, Jon, et al.
Pubblicazione: (2026)
On the Same Wavelength? Evaluating Pragmatic Reasoning in Language Models across Broad Concepts
di: Qiu, Linlu, et al.
Pubblicazione: (2025)
di: Qiu, Linlu, et al.
Pubblicazione: (2025)
Towards an Analysis of Discourse and Interactional Pragmatic Reasoning Capabilities of Large Language Models
di: Robrecht, Amelie, et al.
Pubblicazione: (2024)
di: Robrecht, Amelie, et al.
Pubblicazione: (2024)
MultiPragEval: Multilingual Pragmatic Evaluation of Large Language Models
di: Park, Dojun, et al.
Pubblicazione: (2024)
di: Park, Dojun, et al.
Pubblicazione: (2024)
On Emergent Social World Models -- Evidence for Functional Integration of Theory of Mind and Pragmatic Reasoning in Language Models
di: Tsvilodub, Polina, et al.
Pubblicazione: (2026)
di: Tsvilodub, Polina, et al.
Pubblicazione: (2026)
The Pragmatic Mind of Machines: Tracing the Emergence of Pragmatic Competence in Large Language Models
di: Yu, Kefan, et al.
Pubblicazione: (2025)
di: Yu, Kefan, et al.
Pubblicazione: (2025)
Diagnosing Moral Reasoning Acquisition in Language Models: Pragmatics and Generalization
di: Liu, Guangliang, et al.
Pubblicazione: (2025)
di: Liu, Guangliang, et al.
Pubblicazione: (2025)
Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges
di: Ma, Bolei, et al.
Pubblicazione: (2025)
di: Ma, Bolei, et al.
Pubblicazione: (2025)
Evaluating Large Language Models as Generative User Simulators for Conversational Recommendation
di: Yoon, Se-eun, et al.
Pubblicazione: (2024)
di: Yoon, Se-eun, et al.
Pubblicazione: (2024)
Aligning Large Language Models with Diverse Political Viewpoints
di: Stammbach, Dominik, et al.
Pubblicazione: (2024)
di: Stammbach, Dominik, et al.
Pubblicazione: (2024)
Measuring Pragmatic Influence in Large Language Model Instructions
di: Geng, Yilin, et al.
Pubblicazione: (2026)
di: Geng, Yilin, et al.
Pubblicazione: (2026)
Cross-Linguistic Transcription and Phonological Representation in the Huìtóngguǎnxì Huáyíyìyǔ
di: Kim, Ji-eun
Pubblicazione: (2026)
di: Kim, Ji-eun
Pubblicazione: (2026)
Are Multimodal Large Language Models Pragmatically Competent Listeners in Simple Reference Resolution Tasks?
di: Junker, Simeon, et al.
Pubblicazione: (2025)
di: Junker, Simeon, et al.
Pubblicazione: (2025)
Social Meaning in Large Language Models: Structure, Magnitude, and Pragmatic Prompting
di: Mühlenbernd, Roland
Pubblicazione: (2026)
di: Mühlenbernd, Roland
Pubblicazione: (2026)
Evaluating Counterfactual Strategic Reasoning in Large Language Models
di: Georgousis, Dimitrios, et al.
Pubblicazione: (2026)
di: Georgousis, Dimitrios, et al.
Pubblicazione: (2026)
Evaluating Accounting Reasoning Capabilities of Large Language Models
di: Zhou, Jie, et al.
Pubblicazione: (2026)
di: Zhou, Jie, et al.
Pubblicazione: (2026)
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
di: Xu, Zixiang, et al.
Pubblicazione: (2025)
di: Xu, Zixiang, et al.
Pubblicazione: (2025)
Are BabyLMs Deaf to Gricean Maxims? A Pragmatic Evaluation of Sample-efficient Language Models
di: Askari, Raha, et al.
Pubblicazione: (2025)
di: Askari, Raha, et al.
Pubblicazione: (2025)
Generative Evaluation of Complex Reasoning in Large Language Models
di: Lin, Haowei, et al.
Pubblicazione: (2025)
di: Lin, Haowei, et al.
Pubblicazione: (2025)
Arrows of Math Reasoning Data Synthesis for Large Language Models: Diversity, Complexity and Correctness
di: Chen, Sirui, et al.
Pubblicazione: (2025)
di: Chen, Sirui, et al.
Pubblicazione: (2025)
EconNLI: Evaluating Large Language Models on Economics Reasoning
di: Guo, Yue, et al.
Pubblicazione: (2024)
di: Guo, Yue, et al.
Pubblicazione: (2024)
Developing a Pragmatic Benchmark for Assessing Korean Legal Language Understanding in Large Language Models
di: Kim, Yeeun, et al.
Pubblicazione: (2024)
di: Kim, Yeeun, et al.
Pubblicazione: (2024)
Unveiling the Limits of Large Language Models in Inferring Pragmatic Meaning from Non-Verbal Responses
di: Eo, Sugyeong, et al.
Pubblicazione: (2026)
di: Eo, Sugyeong, et al.
Pubblicazione: (2026)
Evaluating Large Language Models for Evidence-Based Clinical Question Answering
di: Wang, Can, et al.
Pubblicazione: (2025)
di: Wang, Can, et al.
Pubblicazione: (2025)
Review-driven Personalized Preference Reasoning with Large Language Models for Recommendation
di: Kim, Jieyong, et al.
Pubblicazione: (2024)
di: Kim, Jieyong, et al.
Pubblicazione: (2024)
Evaluating Text Creativity across Diverse Domains: A Dataset and Large Language Model Evaluator
di: Cao, Qian, et al.
Pubblicazione: (2025)
di: Cao, Qian, et al.
Pubblicazione: (2025)
Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks
di: Li, Qintong, et al.
Pubblicazione: (2023)
di: Li, Qintong, et al.
Pubblicazione: (2023)
Beyond Understanding: Evaluating the Pragmatic Gap in LLMs' Cultural Processing of Figurative Language
di: Attia, Mena, et al.
Pubblicazione: (2025)
di: Attia, Mena, et al.
Pubblicazione: (2025)
How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models
di: Sieker, Judith, et al.
Pubblicazione: (2026)
di: Sieker, Judith, et al.
Pubblicazione: (2026)
Understanding Understanding: A Pragmatic Framework Motivated by Large Language Models
di: Leyton-Brown, Kevin, et al.
Pubblicazione: (2024)
di: Leyton-Brown, Kevin, et al.
Pubblicazione: (2024)
Can Large Language Models do Analytical Reasoning?
di: Hu, Yebowen, et al.
Pubblicazione: (2024)
di: Hu, Yebowen, et al.
Pubblicazione: (2024)
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
di: Kamath, Amita, et al.
Pubblicazione: (2026)
di: Kamath, Amita, et al.
Pubblicazione: (2026)
Vision-Language Models Are Not Pragmatically Competent in Referring Expression Generation
di: Ma, Ziqiao, et al.
Pubblicazione: (2025)
di: Ma, Ziqiao, et al.
Pubblicazione: (2025)
PACE: A Pragmatic Agent for Enhancing Communication Efficiency Using Large Language Models
di: Li, Jiaxuan, et al.
Pubblicazione: (2024)
di: Li, Jiaxuan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Continuous Interpretive Steering for Scalar Diversity
di: Cho, Ye-eun
Pubblicazione: (2026) -
Pragmatic inference of scalar implicature by LLMs
di: Cho, Ye-eun, et al.
Pubblicazione: (2024) -
Probing Large Language Models for Scalar Adjective Lexical Semantics and Scalar Diversity Pragmatics
di: Lin, Fangru, et al.
Pubblicazione: (2024) -
Can Vision-Language Models Infer Speaker's Ignorance? The Role of Visual and Linguistic Cues
di: Cho, Ye-eun, et al.
Pubblicazione: (2025) -
Pragmatic Competence Evaluation of Large Language Models for the Korean Language
di: Park, Dojun, et al.
Pubblicazione: (2024)