Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong
Fuente:
arXiv
Saved in:
| Main Authors: | Si, Chenglei, Goyal, Navita, Wu, Sherry Tongshuang, Zhao, Chen, Feng, Shi, Daumé III, Hal, Boyd-Graber, Jordan |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Impact of Explanations on Fairness in Human-AI Decision-Making: Protected vs Proxy Features
by: Goyal, Navita, et al.
Published: (2023)
by: Goyal, Navita, et al.
Published: (2023)
Successfully Guiding Humans with Imperfect Instructions by Highlighting Potential Errors and Suggesting Corrections
by: Zhao, Lingjun, et al.
Published: (2024)
by: Zhao, Lingjun, et al.
Published: (2024)
Steering Safely or Off a Cliff? Rethinking Specificity and Robustness in Inference-Time Interventions
by: Goyal, Navita, et al.
Published: (2026)
by: Goyal, Navita, et al.
Published: (2026)
Pragmatics Meets Culture: Culturally-adapted Artwork Description Generation and Evaluation
by: Zhao, Lingjun, et al.
Published: (2026)
by: Zhao, Lingjun, et al.
Published: (2026)
How the Advent of Ubiquitous Large Language Models both Stymie and Turbocharge Dynamic Adversarial Question Generation
by: Sung, Yoo Yeon, et al.
Published: (2024)
by: Sung, Yoo Yeon, et al.
Published: (2024)
Justified or Just Convincing? Error Verifiability as a Dimension of LLM Quality
by: Zhu, Xiaoyuan, et al.
Published: (2026)
by: Zhu, Xiaoyuan, et al.
Published: (2026)
Say It My Way: Exploring Control in Conversational Visual Question Answering with Blind Users
by: Zeraati, Farnaz Zamiri, et al.
Published: (2026)
by: Zeraati, Farnaz Zamiri, et al.
Published: (2026)
Seamful XAI: Operationalizing Seamful Design in Explainable AI
by: Ehsan, Upol, et al.
Published: (2022)
by: Ehsan, Upol, et al.
Published: (2022)
SPHERE: An Evaluation Card for Human-AI Systems
by: Ma, Qianou, et al.
Published: (2025)
by: Ma, Qianou, et al.
Published: (2025)
From Human-Human Collaboration to Human-Agent Collaboration: A Vision, Design Philosophy, and an Empirical Framework for Achieving Successful Partnerships Between Humans and LLM Agents
by: Yao, Bingsheng, et al.
Published: (2026)
by: Yao, Bingsheng, et al.
Published: (2026)
Labeled Interactive Topic Models
by: Seelman, Kyle, et al.
Published: (2023)
by: Seelman, Kyle, et al.
Published: (2023)
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
by: Gor, Maharshi, et al.
Published: (2026)
by: Gor, Maharshi, et al.
Published: (2026)
Which Demographic Features Are Relevant for Individual Fairness Evaluation of U.S. Recidivism Risk Assessment Tools?
by: Nguyen, Tin Trung, et al.
Published: (2025)
by: Nguyen, Tin Trung, et al.
Published: (2025)
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
by: Si, Chenglei, et al.
Published: (2025)
by: Si, Chenglei, et al.
Published: (2025)
Multi-Hop Question Answering: When Can Humans Help, and Where do They Struggle?
by: Su, Jinyan, et al.
Published: (2025)
by: Su, Jinyan, et al.
Published: (2025)
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
by: Si, Chenglei, et al.
Published: (2024)
by: Si, Chenglei, et al.
Published: (2024)
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA
by: Gor, Maharshi, et al.
Published: (2024)
by: Gor, Maharshi, et al.
Published: (2024)
Causal Effect of Group Diversity on Redundancy and Coverage in Peer-Reviewing
by: Goyal, Navita, et al.
Published: (2024)
by: Goyal, Navita, et al.
Published: (2024)
Not Everyone Wins with LLMs: Behavioral Patterns and Pedagogical Implications for AI Literacy in Programmatic Data Science
by: Ma, Qianou, et al.
Published: (2025)
by: Ma, Qianou, et al.
Published: (2025)
Effort-aware Fairness: Incorporating a Philosophy-informed, Human-centered Notion of Effort into Algorithmic Fairness Metrics
by: Nguyen, Tin Trung, et al.
Published: (2025)
by: Nguyen, Tin Trung, et al.
Published: (2025)
When Two Wrongs Don't Make a Right" -- Examining Confirmation Bias and the Role of Time Pressure During Human-AI Collaboration in Computational Pathology
by: Rosbach, Emely, et al.
Published: (2024)
by: Rosbach, Emely, et al.
Published: (2024)
Improving the TENOR of Labeling: Re-evaluating Topic Models for Content Analysis
by: Li, Zongxia, et al.
Published: (2024)
by: Li, Zongxia, et al.
Published: (2024)
Proactive AI Adoption can be Threatening: When Help Backfires
by: Harari, Dana, et al.
Published: (2025)
by: Harari, Dana, et al.
Published: (2025)
When LLMs Help -- and Hurt -- Teaching Assistants in Proof-Based Courses
by: Mahinpei, Romina, et al.
Published: (2026)
by: Mahinpei, Romina, et al.
Published: (2026)
ASL STEM Wiki: Dataset and Benchmark for Interpreting STEM Articles
by: Yin, Kayo, et al.
Published: (2024)
by: Yin, Kayo, et al.
Published: (2024)
Why and When LLM-Based Assistants Can Go Wrong: Investigating the Effectiveness of Prompt-Based Interactions for Software Help-Seeking
by: Khurana, Anjali, et al.
Published: (2024)
by: Khurana, Anjali, et al.
Published: (2024)
Challenges in Trustworthy Human Evaluation of Chatbots
by: Zhao, Wenting, et al.
Published: (2024)
by: Zhao, Wenting, et al.
Published: (2024)
When Humans Don't Feel Like an Option: Contextual Factors That Shape When Older Adults Turn to Conversational AI for Emotional Support
by: Shi, Mengqi, et al.
Published: (2026)
by: Shi, Mengqi, et al.
Published: (2026)
How Voice and Helpfulness Shape Perceptions in Human-Agent Teams
by: Westby, Samuel, et al.
Published: (2023)
by: Westby, Samuel, et al.
Published: (2023)
Evidotes: Integrating Scientific Evidence and Anecdotes to Support Uncertainties Triggered by Peer Health Posts
by: Bali, Shreya, et al.
Published: (2026)
by: Bali, Shreya, et al.
Published: (2026)
Humans Perceive Wrong Narratives from AI Reasoning Texts
by: Levy, Mosh, et al.
Published: (2025)
by: Levy, Mosh, et al.
Published: (2025)
Improving Automated Feedback Systems for Tutor Training in Low-Resource Scenarios through Data Augmentation
by: Xu, Chentianye, et al.
Published: (2025)
by: Xu, Chentianye, et al.
Published: (2025)
What Should We Engineer in Prompts? Training Humans in Requirement-Driven LLM Use
by: Ma, Qianou, et al.
Published: (2024)
by: Ma, Qianou, et al.
Published: (2024)
When Support Escalates Distress: Regulation and Escalation in LLM Responses to Venting and Advice-Seeking
by: Chi, Vivienne Bihe, et al.
Published: (2026)
by: Chi, Vivienne Bihe, et al.
Published: (2026)
Games That Teach, Chats That Convince: Comparing Interactive and Static Formats for Persuasive Learning
by: Alavi, Seyed Hossein, et al.
Published: (2026)
by: Alavi, Seyed Hossein, et al.
Published: (2026)
Analyzing Reluctance to Ask for Help When Cooperating With Robots: Insights to Integrate Artificial Agents in HRC
by: Martin, Ane San, et al.
Published: (2025)
by: Martin, Ane San, et al.
Published: (2025)
Racism, Resistance, and Reddit: How Popular Culture Sparks Online Reckonings
by: Mason, Sherry, et al.
Published: (2025)
by: Mason, Sherry, et al.
Published: (2025)
Behavioral Indicators of Overreliance During Interaction with Conversational Language Models
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
Write on Paper, Wrong in Practice: Why LLMs Still Struggle with Writing Clinical Notes
by: Kupferschmidt, Kristina L., et al.
Published: (2025)
by: Kupferschmidt, Kristina L., et al.
Published: (2025)
The Benefits of Prosociality towards AI Agents: Examining the Effects of Helping AI Agents on Human Well-Being
by: Zhu, Zicheng, et al.
Published: (2025)
by: Zhu, Zicheng, et al.
Published: (2025)
Similar Items
-
The Impact of Explanations on Fairness in Human-AI Decision-Making: Protected vs Proxy Features
by: Goyal, Navita, et al.
Published: (2023) -
Successfully Guiding Humans with Imperfect Instructions by Highlighting Potential Errors and Suggesting Corrections
by: Zhao, Lingjun, et al.
Published: (2024) -
Steering Safely or Off a Cliff? Rethinking Specificity and Robustness in Inference-Time Interventions
by: Goyal, Navita, et al.
Published: (2026) -
Pragmatics Meets Culture: Culturally-adapted Artwork Description Generation and Evaluation
by: Zhao, Lingjun, et al.
Published: (2026) -
How the Advent of Ubiquitous Large Language Models both Stymie and Turbocharge Dynamic Adversarial Question Generation
by: Sung, Yoo Yeon, et al.
Published: (2024)