Navigating Rifts in Human-LLM Grounding: Study and Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Shaikh, Omar, Mozannar, Hussein, Bansal, Gagan, Fourney, Adam, Horvitz, Eric |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When to Show a Suggestion? Integrating Human Feedback in AI-Assisted Programming
by: Mozannar, Hussein, et al.
Published: (2023)
by: Mozannar, Hussein, et al.
Published: (2023)
Reading Between the Lines: Modeling User Behavior and Costs in AI-Assisted Programming
by: Mozannar, Hussein, et al.
Published: (2022)
by: Mozannar, Hussein, et al.
Published: (2022)
Challenges in Human-Agent Communication
by: Bansal, Gagan, et al.
Published: (2024)
by: Bansal, Gagan, et al.
Published: (2024)
Overseeing Agents Without Constant Oversight: Challenges and Opportunities
by: Grunde-McLaughlin, Madeleine, et al.
Published: (2026)
by: Grunde-McLaughlin, Madeleine, et al.
Published: (2026)
Creating General User Models from Computer Use
by: Shaikh, Omar, et al.
Published: (2025)
by: Shaikh, Omar, et al.
Published: (2025)
Grounding Gaps in Language Model Generations
by: Shaikh, Omar, et al.
Published: (2023)
by: Shaikh, Omar, et al.
Published: (2023)
Generation Probabilities Are Not Enough: Uncertainty Highlighting in AI Code Completions
by: Vasconcelos, Helena, et al.
Published: (2023)
by: Vasconcelos, Helena, et al.
Published: (2023)
AutoGen Studio: A No-Code Developer Tool for Building and Debugging Multi-Agent Systems
by: Dibia, Victor, et al.
Published: (2024)
by: Dibia, Victor, et al.
Published: (2024)
Magentic-UI: Towards Human-in-the-loop Agentic Systems
by: Mozannar, Hussein, et al.
Published: (2025)
by: Mozannar, Hussein, et al.
Published: (2025)
CodingGenie: A Proactive LLM-Powered Programming Assistant
by: Zhao, Sebastian, et al.
Published: (2025)
by: Zhao, Sebastian, et al.
Published: (2025)
Interactive Debugging and Steering of Multi-Agent AI Systems
by: Epperson, Will, et al.
Published: (2025)
by: Epperson, Will, et al.
Published: (2025)
Learning Next Action Predictors from Human-Computer Interaction
by: Shaikh, Omar, et al.
Published: (2026)
by: Shaikh, Omar, et al.
Published: (2026)
How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
by: Wang, Zora Zhiruo, et al.
Published: (2025)
by: Wang, Zora Zhiruo, et al.
Published: (2025)
Social Skill Training with Large Language Models
by: Yang, Diyi, et al.
Published: (2024)
by: Yang, Diyi, et al.
Published: (2024)
Navigating the Landscape of Hint Generation Research: From the Past to the Future
by: Jangra, Anubhav, et al.
Published: (2024)
by: Jangra, Anubhav, et al.
Published: (2024)
Aligning Language Models with Demonstrated Feedback
by: Shaikh, Omar, et al.
Published: (2024)
by: Shaikh, Omar, et al.
Published: (2024)
AI-Powered Reminders for Collaborative Tasks: Experiences and Futures
by: Morrison, Katelyn, et al.
Published: (2024)
by: Morrison, Katelyn, et al.
Published: (2024)
Human-LLM Collaborative Feature Engineering for Tabular Data
by: Li, Zhuoyan, et al.
Published: (2026)
by: Li, Zhuoyan, et al.
Published: (2026)
Impact of Large Language Model Assistance on Patients Reading Clinical Notes: A Mixed-Methods Study
by: Mannhardt, Niklas, et al.
Published: (2024)
by: Mannhardt, Niklas, et al.
Published: (2024)
Need Help? Designing Proactive AI Assistants for Programming
by: Chen, Valerie, et al.
Published: (2024)
by: Chen, Valerie, et al.
Published: (2024)
LOGOS: LLM-driven End-to-End Grounded Theory Development and Schema Induction for Qualitative Research
by: Pi, Xinyu, et al.
Published: (2025)
by: Pi, Xinyu, et al.
Published: (2025)
MEGAnno+: A Human-LLM Collaborative Annotation System
by: Kim, Hannah, et al.
Published: (2024)
by: Kim, Hannah, et al.
Published: (2024)
PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A
by: Martin-Boyle, Anna, et al.
Published: (2026)
by: Martin-Boyle, Anna, et al.
Published: (2026)
Human-LLM Collaborative Construction of a Cantonese Emotion Lexicon
by: Zhang, Yusong, et al.
Published: (2024)
by: Zhang, Yusong, et al.
Published: (2024)
PersoDPO: Scalable Preference Optimization for Instruction-Adherent, Persona-Grounded Dialogue via Multi-LLM Evaluation
by: Afzoon, Saleh, et al.
Published: (2026)
by: Afzoon, Saleh, et al.
Published: (2026)
Show or Tell? Modeling the evolution of request-making in Human-LLM conversations
by: Zhu, Shengqi, et al.
Published: (2025)
by: Zhu, Shengqi, et al.
Published: (2025)
Beyond Turn-taking: Introducing Text-based Overlap into Human-LLM Interactions
by: Kim, JiWoo, et al.
Published: (2025)
by: Kim, JiWoo, et al.
Published: (2025)
Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild
by: Mysore, Sheshera, et al.
Published: (2025)
by: Mysore, Sheshera, et al.
Published: (2025)
The LLM Effect: Are Humans Truly Using LLMs, or Are They Being Influenced By Them Instead?
by: Choi, Alexander S., et al.
Published: (2024)
by: Choi, Alexander S., et al.
Published: (2024)
LLM-Augmented Semantic Steering of Text Embedding Projection Spaces
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
Rehearsal: Simulating Conflict to Teach Conflict Resolution
by: Shaikh, Omar, et al.
Published: (2023)
by: Shaikh, Omar, et al.
Published: (2023)
LVLMs and Humans Ground Differently in Referential Communication
by: Zeng, Peter, et al.
Published: (2026)
by: Zeng, Peter, et al.
Published: (2026)
LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces
by: Kirgis, Peter, et al.
Published: (2026)
by: Kirgis, Peter, et al.
Published: (2026)
AI vs. Human Judgment of Content Moderation: LLM-as-a-Judge and Ethics-Based Response Refusals
by: Pasch, Stefan
Published: (2025)
by: Pasch, Stefan
Published: (2025)
Sniff AI: Is My 'Spicy' Your 'Spicy'? Exploring LLM's Perceptual Alignment with Human Smell Experiences
by: Zhong, Shu, et al.
Published: (2024)
by: Zhong, Shu, et al.
Published: (2024)
Benchmarking LLM Tool-Use in the Wild
by: Yu, Peijie, et al.
Published: (2026)
by: Yu, Peijie, et al.
Published: (2026)
Evaluating LLM-Generated Q&A Test: a Student-Centered Study
by: Wróblewska, Anna, et al.
Published: (2025)
by: Wróblewska, Anna, et al.
Published: (2025)
A Comparative Study on Annotation Quality of Crowdsourcing and LLM via Label Aggregation
by: Li, Jiyi
Published: (2024)
by: Li, Jiyi
Published: (2024)
Grounding Emotional Descriptions to Electrovibration Haptic Signals
by: Hu, Guimin, et al.
Published: (2024)
by: Hu, Guimin, et al.
Published: (2024)
"What Are You Really Trying to Do?": Co-Creating Life Goals from Everyday Computer Use
by: Sapkota, Shardul, et al.
Published: (2026)
by: Sapkota, Shardul, et al.
Published: (2026)
Similar Items
-
When to Show a Suggestion? Integrating Human Feedback in AI-Assisted Programming
by: Mozannar, Hussein, et al.
Published: (2023) -
Reading Between the Lines: Modeling User Behavior and Costs in AI-Assisted Programming
by: Mozannar, Hussein, et al.
Published: (2022) -
Challenges in Human-Agent Communication
by: Bansal, Gagan, et al.
Published: (2024) -
Overseeing Agents Without Constant Oversight: Challenges and Opportunities
by: Grunde-McLaughlin, Madeleine, et al.
Published: (2026) -
Creating General User Models from Computer Use
by: Shaikh, Omar, et al.
Published: (2025)