AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Su, Zhe, Zhou, Xuhui, Rangreji, Sanketh, Kabra, Anubha, Mendelsohn, Julia, Brahman, Faeze, Sap, Maarten |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Leftover Lunch: Advantage-based Offline Reinforcement Learning for Language Models
by: Baheti, Ashutosh, et al.
Published: (2023)
by: Baheti, Ashutosh, et al.
Published: (2023)
From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models
by: Mendelsohn, Julia, et al.
Published: (2023)
by: Mendelsohn, Julia, et al.
Published: (2023)
Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferences
by: Zheng, Mingqian, et al.
Published: (2025)
by: Zheng, Mingqian, et al.
Published: (2025)
Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs
by: Zhou, Xuhui, et al.
Published: (2024)
by: Zhou, Xuhui, et al.
Published: (2024)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
by: Jung, Jaehun, et al.
Published: (2024)
by: Jung, Jaehun, et al.
Published: (2024)
Examining the Effect of Explanations of AI Privacy Redaction in AI-mediated Interactions
by: Kaushik, Roshni, et al.
Published: (2026)
by: Kaushik, Roshni, et al.
Published: (2026)
1-2-3 Check: Enhancing Contextual Privacy in LLM via Multi-Agent Reasoning
by: Li, Wenkai, et al.
Published: (2025)
by: Li, Wenkai, et al.
Published: (2025)
Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates Hallucinations
by: Chen, Tong, et al.
Published: (2025)
by: Chen, Tong, et al.
Published: (2025)
HAICOSYSTEM: An Ecosystem for Sandboxing Safety Risks in Human-AI Interactions
by: Zhou, Xuhui, et al.
Published: (2024)
by: Zhou, Xuhui, et al.
Published: (2024)
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
by: Vijayvargiya, Sanidhya, et al.
Published: (2025)
by: Vijayvargiya, Sanidhya, et al.
Published: (2025)
Training Proactive and Personalized LLM Agents
by: Sun, Weiwei, et al.
Published: (2025)
by: Sun, Weiwei, et al.
Published: (2025)
Multi-Attribute Constraint Satisfaction via Language Model Rewriting
by: Baheti, Ashutosh, et al.
Published: (2024)
by: Baheti, Ashutosh, et al.
Published: (2024)
OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
by: Vijayvargiya, Sanidhya, et al.
Published: (2025)
by: Vijayvargiya, Sanidhya, et al.
Published: (2025)
Minion: A Technology Probe to Explore How Users Negotiate Harmful Value Conflicts with AI Companions
by: Fan, Xianzhe, et al.
Published: (2024)
by: Fan, Xianzhe, et al.
Published: (2024)
BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data
by: Li, Wenkai, et al.
Published: (2024)
by: Li, Wenkai, et al.
Published: (2024)
On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents
by: Huang, Jen-tse, et al.
Published: (2024)
by: Huang, Jen-tse, et al.
Published: (2024)
Exploring Big Five Personality and AI Capability Effects in LLM-Simulated Negotiation Dialogues
by: Cohen, Myke C., et al.
Published: (2025)
by: Cohen, Myke C., et al.
Published: (2025)
Classifying and Clustering Trading Agents
by: Wilinski, Mateusz, et al.
Published: (2025)
by: Wilinski, Mateusz, et al.
Published: (2025)
TOM-SWE: User Mental Modeling For Software Engineering Agents
by: Zhou, Xuhui, et al.
Published: (2025)
by: Zhou, Xuhui, et al.
Published: (2025)
POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents
by: Zheng, Qiaoyuan, et al.
Published: (2026)
by: Zheng, Qiaoyuan, et al.
Published: (2026)
Creativity Support in the Age of Large Language Models: An Empirical Study Involving Emerging Writers
by: Chakrabarty, Tuhin, et al.
Published: (2023)
by: Chakrabarty, Tuhin, et al.
Published: (2023)
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
by: Jiang, Liwei, et al.
Published: (2024)
by: Jiang, Liwei, et al.
Published: (2024)
Rethinking Theory of Mind Benchmarks for LLMs: Towards A User-Centered Perspective
by: Wang, Qiaosi, et al.
Published: (2025)
by: Wang, Qiaosi, et al.
Published: (2025)
GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses
by: Mun, Jimin, et al.
Published: (2026)
by: Mun, Jimin, et al.
Published: (2026)
Social World Models
by: Zhou, Xuhui, et al.
Published: (2025)
by: Zhou, Xuhui, et al.
Published: (2025)
When Should AI Read the Room? Public Perceptions of Social Intelligence in AI Agents
by: Mathur, Leena, et al.
Published: (2026)
by: Mathur, Leena, et al.
Published: (2026)
SOTOPIA-TOM: Evaluating Information Management in Multi-Agent Interaction with Theory of Mind
by: YS, Yashwanth, et al.
Published: (2026)
by: YS, Yashwanth, et al.
Published: (2026)
ALFA: Aligning LLMs to Ask Good Questions A Case Study in Clinical Reasoning
by: Li, Shuyue Stella, et al.
Published: (2025)
by: Li, Shuyue Stella, et al.
Published: (2025)
Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
by: Zheng, Mingqian, et al.
Published: (2026)
by: Zheng, Mingqian, et al.
Published: (2026)
User-Driven Value Alignment: Understanding Users' Perceptions and Strategies for Addressing Biased and Discriminatory Statements in AI Companions
by: Fan, Xianzhe, et al.
Published: (2024)
by: Fan, Xianzhe, et al.
Published: (2024)
Performance-Efficiency Trade-off for Fashion Image Retrieval
by: Hurtado, Julio, et al.
Published: (2025)
by: Hurtado, Julio, et al.
Published: (2025)
Position: Embodied AI Requires a Privacy-Utility Trade-off
by: Fan, Xiaoliang, et al.
Published: (2026)
by: Fan, Xiaoliang, et al.
Published: (2026)
The Effect of Document Selection on Query-focused Text Analysis
by: Rangreji, Sandesh S, et al.
Published: (2026)
by: Rangreji, Sandesh S, et al.
Published: (2026)
Framing an AI with Values Reduces AI Reliance in AI-supported Writing Tasks
by: Gao, Alice, et al.
Published: (2026)
by: Gao, Alice, et al.
Published: (2026)
Agent Lumos: Unified and Modular Training for Open-Source Language Agents
by: Yin, Da, et al.
Published: (2023)
by: Yin, Da, et al.
Published: (2023)
Tailoring with Targeted Precision: Edit-Based Agents for Open-Domain Procedure Customization
by: Lal, Yash Kumar, et al.
Published: (2023)
by: Lal, Yash Kumar, et al.
Published: (2023)
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
by: Röttger, Paul, et al.
Published: (2025)
by: Röttger, Paul, et al.
Published: (2025)
Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
by: He, Zhonghao, et al.
Published: (2025)
by: He, Zhonghao, et al.
Published: (2025)
SoMi-ToM: Evaluating Multi-Perspective Theory of Mind in Embodied Social Interactions
by: Fan, Xianzhe, et al.
Published: (2025)
by: Fan, Xianzhe, et al.
Published: (2025)
PolygloToxicityPrompts: Multilingual Evaluation of Neural Toxic Degeneration in Large Language Models
by: Jain, Devansh, et al.
Published: (2024)
by: Jain, Devansh, et al.
Published: (2024)
Similar Items
-
Leftover Lunch: Advantage-based Offline Reinforcement Learning for Language Models
by: Baheti, Ashutosh, et al.
Published: (2023) -
From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models
by: Mendelsohn, Julia, et al.
Published: (2023) -
Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferences
by: Zheng, Mingqian, et al.
Published: (2025) -
Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs
by: Zhou, Xuhui, et al.
Published: (2024) -
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
by: Jung, Jaehun, et al.
Published: (2024)