Is the Pope Catholic? Yes, the Pope is Catholic. Generative Evaluation of Non-Literal Intent Resolution in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Yerukola, Akhila, Vaduguru, Saujas, Fried, Daniel, Sap, Maarten |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
by: Yerukola, Akhila, et al.
Published: (2025)
by: Yerukola, Akhila, et al.
Published: (2025)
Generating Pragmatic Examples to Train Neural Program Synthesizers
by: Vaduguru, Saujas, et al.
Published: (2023)
by: Vaduguru, Saujas, et al.
Published: (2023)
Amortizing Pragmatic Program Synthesis with Rankings
by: Pu, Yewen, et al.
Published: (2024)
by: Pu, Yewen, et al.
Published: (2024)
Analyzing Information Sharing and Coordination in Multi-Agent Planning
by: Ou, Tianyue, et al.
Published: (2025)
by: Ou, Tianyue, et al.
Published: (2025)
Identifying & Interactively Refining Ambiguous User Goals for Data Visualization Code Generation
by: İnan, Mert, et al.
Published: (2025)
by: İnan, Mert, et al.
Published: (2025)
Amortizing Pragmatic Program Synthesis with Rankings
by: Pu, Yewen, et al.
Published: (2023)
by: Pu, Yewen, et al.
Published: (2023)
Success and Cost Elicit Convention Formation for Efficient Communication
by: Vaduguru, Saujas, et al.
Published: (2025)
by: Vaduguru, Saujas, et al.
Published: (2025)
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
by: Vijayvargiya, Sanidhya, et al.
Published: (2025)
by: Vijayvargiya, Sanidhya, et al.
Published: (2025)
Social World Models
by: Zhou, Xuhui, et al.
Published: (2025)
by: Zhou, Xuhui, et al.
Published: (2025)
mrCAD: Multimodal Refinement of Computer-aided Designs
by: McCarthy, William P., et al.
Published: (2025)
by: McCarthy, William P., et al.
Published: (2025)
NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models
by: Rao, Abhinav, et al.
Published: (2024)
by: Rao, Abhinav, et al.
Published: (2024)
Out of Style: RAG's Fragility to Linguistic Variation
by: Cao, Tianyu, et al.
Published: (2025)
by: Cao, Tianyu, et al.
Published: (2025)
Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
by: Zheng, Mingqian, et al.
Published: (2026)
by: Zheng, Mingqian, et al.
Published: (2026)
Words Like Knives: Backstory-Personalized Modeling and Detection of Violent Communication
by: Shen, Jocelyn, et al.
Published: (2025)
by: Shen, Jocelyn, et al.
Published: (2025)
Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs
by: Zhou, Xuhui, et al.
Published: (2024)
by: Zhou, Xuhui, et al.
Published: (2024)
PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
by: Kumar, Priyanshu, et al.
Published: (2025)
by: Kumar, Priyanshu, et al.
Published: (2025)
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
by: Zhou, Xuhui, et al.
Published: (2023)
by: Zhou, Xuhui, et al.
Published: (2023)
Difficult Task Yes but Simple Task No: Unveiling the Laziness in Multimodal LLMs
by: Zhao, Sihang, et al.
Published: (2024)
by: Zhao, Sihang, et al.
Published: (2024)
Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification
by: Beniwal, Himanshu, et al.
Published: (2025)
by: Beniwal, Himanshu, et al.
Published: (2025)
GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses
by: Mun, Jimin, et al.
Published: (2026)
by: Mun, Jimin, et al.
Published: (2026)
Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
by: He, Zhonghao, et al.
Published: (2025)
by: He, Zhonghao, et al.
Published: (2025)
The Popes, the Catholic Church and the Transatlantic Enslavement of Black Africans 1418-1839 (Volume 16)
by: Adiele, Pius Onyemechi
Published: (2021)
by: Adiele, Pius Onyemechi
Published: (2021)
Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty
by: Zhou, Kaitlyn, et al.
Published: (2024)
by: Zhou, Kaitlyn, et al.
Published: (2024)
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
by: Mireshghallah, Niloofar, et al.
Published: (2023)
by: Mireshghallah, Niloofar, et al.
Published: (2023)
Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations
by: Park, Eunkyu, et al.
Published: (2025)
by: Park, Eunkyu, et al.
Published: (2025)
SoMi-ToM: Evaluating Multi-Perspective Theory of Mind in Embodied Social Interactions
by: Fan, Xianzhe, et al.
Published: (2025)
by: Fan, Xianzhe, et al.
Published: (2025)
AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents
by: Su, Zhe, et al.
Published: (2024)
by: Su, Zhe, et al.
Published: (2024)
AGENT-CQ: Automatic Generation and Evaluation of Clarifying Questions for Conversational Search with LLMs
by: Siro, Clemencia, et al.
Published: (2024)
by: Siro, Clemencia, et al.
Published: (2024)
Yes, this is what I was looking for! Towards Multi-modal Medical Consultation Concern Summary Generation
by: Tiwari, Abhisek, et al.
Published: (2024)
by: Tiwari, Abhisek, et al.
Published: (2024)
ECCO: Can We Improve Model-Generated Code Efficiency Without Sacrificing Functional Correctness?
by: Waghjale, Siddhant, et al.
Published: (2024)
by: Waghjale, Siddhant, et al.
Published: (2024)
Intent-conditioned and Non-toxic Counterspeech Generation using Multi-Task Instruction Tuning with RLAIF
by: Hengle, Amey, et al.
Published: (2024)
by: Hengle, Amey, et al.
Published: (2024)
Disparities in LLM Reasoning Accuracy and Explanations: A Case Study on African American English
by: Zhou, Runtao, et al.
Published: (2025)
by: Zhou, Runtao, et al.
Published: (2025)
API-Assisted Code Generation for Question Answering on Varied Table Structures
by: Cao, Yihan, et al.
Published: (2023)
by: Cao, Yihan, et al.
Published: (2023)
Training LLMs with Reinforcement Learning for Intent-Aware Personalized Question Answering
by: Amirizaniani, Maryam, et al.
Published: (2026)
by: Amirizaniani, Maryam, et al.
Published: (2026)
Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
by: Liu, Geng, et al.
Published: (2026)
by: Liu, Geng, et al.
Published: (2026)
A Large-Scale Dataset and Citation Intent Classification in Turkish with LLMs
by: Karaca, Kemal Sami, et al.
Published: (2025)
by: Karaca, Kemal Sami, et al.
Published: (2025)
From Intents to Conversations: Generating Intent-Driven Dialogues with Contrastive Learning for Multi-Turn Classification
by: Liu, Junhua, et al.
Published: (2024)
by: Liu, Junhua, et al.
Published: (2024)
CRScore: Grounding Automated Evaluation of Code Review Comments in Code Claims and Smells
by: Naik, Atharva, et al.
Published: (2024)
by: Naik, Atharva, et al.
Published: (2024)
PromptTailor: Multi-turn Intent-Aligned Prompt Synthesis for Lightweight LLMs
by: Xu, Yizhou, et al.
Published: (2025)
by: Xu, Yizhou, et al.
Published: (2025)
Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance
by: Zhou, Kaitlyn, et al.
Published: (2024)
by: Zhou, Kaitlyn, et al.
Published: (2024)
Similar Items
-
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
by: Yerukola, Akhila, et al.
Published: (2025) -
Generating Pragmatic Examples to Train Neural Program Synthesizers
by: Vaduguru, Saujas, et al.
Published: (2023) -
Amortizing Pragmatic Program Synthesis with Rankings
by: Pu, Yewen, et al.
Published: (2024) -
Analyzing Information Sharing and Coordination in Multi-Agent Planning
by: Ou, Tianyue, et al.
Published: (2025) -
Identifying & Interactively Refining Ambiguous User Goals for Data Visualization Code Generation
by: İnan, Mert, et al.
Published: (2025)