From Polyester Girlfriends to Blind Mice: Creating the First Pragmatics Understanding Benchmarks for Slovene
Fuente:
arXiv
Guardado en:
| Autores principales: | Brglez, Mojca, Vintar, Špela |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Charting the European LLM Benchmarking Landscape: A New Taxonomy and a Set of Best Practices
por: Vintar, Špela, et al.
Publicado: (2025)
por: Vintar, Špela, et al.
Publicado: (2025)
A Computational Analysis of the Dehumanisation of Migrants from Syria and Ukraine in Slovene News Media
por: Caporusso, Jaya, et al.
Publicado: (2024)
por: Caporusso, Jaya, et al.
Publicado: (2024)
The truth is no diaper: Human and AI-generated associations to emotional words
por: Vintar, Špela, et al.
Publicado: (2025)
por: Vintar, Špela, et al.
Publicado: (2025)
Understand the Implication: Learning to Think for Pragmatic Understanding
por: Sravanthi, Settaluri Lakshmi, et al.
Publicado: (2025)
por: Sravanthi, Settaluri Lakshmi, et al.
Publicado: (2025)
Developing a Pragmatic Benchmark for Assessing Korean Legal Language Understanding in Large Language Models
por: Kim, Yeeun, et al.
Publicado: (2024)
por: Kim, Yeeun, et al.
Publicado: (2024)
Environmental, Social and Governance Sentiment Analysis on Slovene News: A Novel Dataset and Models
por: Dodig, Paula, et al.
Publicado: (2026)
por: Dodig, Paula, et al.
Publicado: (2026)
CEI: A Benchmark for Evaluating Pragmatic Reasoning in Language Models
por: Chun, Jon, et al.
Publicado: (2026)
por: Chun, Jon, et al.
Publicado: (2026)
Understanding Understanding: A Pragmatic Framework Motivated by Large Language Models
por: Leyton-Brown, Kevin, et al.
Publicado: (2024)
por: Leyton-Brown, Kevin, et al.
Publicado: (2024)
Blind Men and the Elephant: Diverse Perspectives on Gender Stereotypes in Benchmark Datasets
por: Zakizadeh, Mahdi, et al.
Publicado: (2025)
por: Zakizadeh, Mahdi, et al.
Publicado: (2025)
From Words to Worlds: Benchmarking Cross-Cultural Cultural Understanding in Machine Translation
por: Han, Bangju, et al.
Publicado: (2026)
por: Han, Bangju, et al.
Publicado: (2026)
From Blind Solvers to Logical Thinkers: Benchmarking LLMs' Logical Integrity on Faulty Mathematical Problems
por: Rahman, A M Muntasir, et al.
Publicado: (2024)
por: Rahman, A M Muntasir, et al.
Publicado: (2024)
Communicating with Speakers and Listeners of Different Pragmatic Levels
por: Naszadi, Kata, et al.
Publicado: (2024)
por: Naszadi, Kata, et al.
Publicado: (2024)
Context-Aware Pragmatic Metacognitive Prompting for Sarcasm Detection
por: Iskandardinata, Michael, et al.
Publicado: (2025)
por: Iskandardinata, Michael, et al.
Publicado: (2025)
Measuring Pragmatic Influence in Large Language Model Instructions
por: Geng, Yilin, et al.
Publicado: (2026)
por: Geng, Yilin, et al.
Publicado: (2026)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
por: Jiang, Botian, et al.
Publicado: (2024)
por: Jiang, Botian, et al.
Publicado: (2024)
Towards AGI A Pragmatic Approach Towards Self Evolving Agent
por: Kar, Indrajit, et al.
Publicado: (2026)
por: Kar, Indrajit, et al.
Publicado: (2026)
Social Meaning in Large Language Models: Structure, Magnitude, and Pragmatic Prompting
por: Mühlenbernd, Roland
Publicado: (2026)
por: Mühlenbernd, Roland
Publicado: (2026)
ALPS: A Diagnostic Challenge Set for Arabic Linguistic & Pragmatic Reasoning
por: Al-Olimat, Hussein S., et al.
Publicado: (2026)
por: Al-Olimat, Hussein S., et al.
Publicado: (2026)
Simple Linguistic Inferences of Large Language Models (LLMs): Blind Spots and Blinds
por: Basmov, Victoria, et al.
Publicado: (2023)
por: Basmov, Victoria, et al.
Publicado: (2023)
RSA-Control: A Pragmatics-Grounded Lightweight Controllable Text Generation Framework
por: Wang, Yifan, et al.
Publicado: (2024)
por: Wang, Yifan, et al.
Publicado: (2024)
Oogiri-Master: Benchmarking Humor Understanding via Oogiri
por: Murakami, Soichiro, et al.
Publicado: (2025)
por: Murakami, Soichiro, et al.
Publicado: (2025)
Pragmatic Inference Chain (PIC) Improving LLMs' Reasoning of Authentic Implicit Toxic Language
por: Chen, Xi, et al.
Publicado: (2025)
por: Chen, Xi, et al.
Publicado: (2025)
PACE: A Pragmatic Agent for Enhancing Communication Efficiency Using Large Language Models
por: Li, Jiaxuan, et al.
Publicado: (2024)
por: Li, Jiaxuan, et al.
Publicado: (2024)
Benchmarking Multimodal LLMs on Recognition and Understanding over Chemical Tables
por: Zhou, Yitong, et al.
Publicado: (2025)
por: Zhou, Yitong, et al.
Publicado: (2025)
Mobile-MMLU: A Mobile Intelligence Language Understanding Benchmark
por: Bsharat, Sondos Mahmoud, et al.
Publicado: (2025)
por: Bsharat, Sondos Mahmoud, et al.
Publicado: (2025)
Chumor 2.0: Towards Benchmarking Chinese Humor Understanding
por: He, Ruiqi, et al.
Publicado: (2024)
por: He, Ruiqi, et al.
Publicado: (2024)
CURE: A Multimodal Benchmark for Clinical Understanding and Retrieval Evaluation
por: Gu, Yannian, et al.
Publicado: (2026)
por: Gu, Yannian, et al.
Publicado: (2026)
Creating a digital poet
por: Tohar, Vered, et al.
Publicado: (2026)
por: Tohar, Vered, et al.
Publicado: (2026)
Pragmatic Reasoning improves LLM Code Generation
por: Cao, Zhuchen, et al.
Publicado: (2025)
por: Cao, Zhuchen, et al.
Publicado: (2025)
Nek Minit: Harnessing Pragmatic Metacognitive Prompting for Explainable Sarcasm Detection of Australian and Indian English
por: Singh, Ishmanbir, et al.
Publicado: (2025)
por: Singh, Ishmanbir, et al.
Publicado: (2025)
Model-Based Data-Centric AI: Bridging the Divide Between Academic Ideals and Industrial Pragmatism
por: Park, Chanjun, et al.
Publicado: (2024)
por: Park, Chanjun, et al.
Publicado: (2024)
Unveiling the Limits of Large Language Models in Inferring Pragmatic Meaning from Non-Verbal Responses
por: Eo, Sugyeong, et al.
Publicado: (2026)
por: Eo, Sugyeong, et al.
Publicado: (2026)
Generative AI, Pragmatics, and Authenticity in Second Language Learning
por: Godwin-Jones`, Robert
Publicado: (2024)
por: Godwin-Jones`, Robert
Publicado: (2024)
Fleurs-SLU: A Massively Multilingual Benchmark for Spoken Language Understanding
por: Schmidt, Fabian David, et al.
Publicado: (2025)
por: Schmidt, Fabian David, et al.
Publicado: (2025)
AraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular Data
por: Alshaikh, Rana, et al.
Publicado: (2025)
por: Alshaikh, Rana, et al.
Publicado: (2025)
NorEval: A Norwegian Language Understanding and Generation Evaluation Benchmark
por: Mikhailov, Vladislav, et al.
Publicado: (2025)
por: Mikhailov, Vladislav, et al.
Publicado: (2025)
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
por: Zhu, Hengchuan, et al.
Publicado: (2025)
por: Zhu, Hengchuan, et al.
Publicado: (2025)
TUMLU: A Unified and Native Language Understanding Benchmark for Turkic Languages
por: Isbarov, Jafar, et al.
Publicado: (2025)
por: Isbarov, Jafar, et al.
Publicado: (2025)
BLUCK: A Benchmark Dataset for Bengali Linguistic Understanding and Cultural Knowledge
por: Kabir, Daeen, et al.
Publicado: (2025)
por: Kabir, Daeen, et al.
Publicado: (2025)
New Semantic Task for the French Spoken Language Understanding MEDIA Benchmark
por: Alavoine, Nadège, et al.
Publicado: (2024)
por: Alavoine, Nadège, et al.
Publicado: (2024)
Ejemplares similares
-
Charting the European LLM Benchmarking Landscape: A New Taxonomy and a Set of Best Practices
por: Vintar, Špela, et al.
Publicado: (2025) -
A Computational Analysis of the Dehumanisation of Migrants from Syria and Ukraine in Slovene News Media
por: Caporusso, Jaya, et al.
Publicado: (2024) -
The truth is no diaper: Human and AI-generated associations to emotional words
por: Vintar, Špela, et al.
Publicado: (2025) -
Understand the Implication: Learning to Think for Pragmatic Understanding
por: Sravanthi, Settaluri Lakshmi, et al.
Publicado: (2025) -
Developing a Pragmatic Benchmark for Assessing Korean Legal Language Understanding in Large Language Models
por: Kim, Yeeun, et al.
Publicado: (2024)