SPHERE: An Evaluation Card for Human-AI Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Qianou, Zhao, Dora, Zhao, Xinran, Si, Chenglei, Yang, Chenyang, Louie, Ryan, Reiter, Ehud, Yang, Diyi, Wu, Tongshuang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What Should We Engineer in Prompts? Training Humans in Requirement-Driven LLM Use
by: Ma, Qianou, et al.
Published: (2024)
by: Ma, Qianou, et al.
Published: (2024)
Not Everyone Wins with LLMs: Behavioral Patterns and Pedagogical Implications for AI Literacy in Programmatic Data Science
by: Ma, Qianou, et al.
Published: (2025)
by: Ma, Qianou, et al.
Published: (2025)
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
by: Si, Chenglei, et al.
Published: (2025)
by: Si, Chenglei, et al.
Published: (2025)
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
by: Si, Chenglei, et al.
Published: (2024)
by: Si, Chenglei, et al.
Published: (2024)
How to Teach Programming in the AI Era? Using LLMs as a Teachable Agent for Debugging
by: Ma, Qianou, et al.
Published: (2023)
by: Ma, Qianou, et al.
Published: (2023)
RECAP: An End-to-End Platform for Capturing, Replaying, and Analyzing AI-Assisted Programming Interactions
by: He, Keyu, et al.
Published: (2026)
by: He, Keyu, et al.
Published: (2026)
Knoll: Creating a Knowledge Ecosystem for Large Language Models
by: Zhao, Dora, et al.
Published: (2025)
by: Zhao, Dora, et al.
Published: (2025)
Evaluation of Human-Understandability of Global Model Explanations using Decision Tree
by: Sivaprasad, Adarsa, et al.
Published: (2023)
by: Sivaprasad, Adarsa, et al.
Published: (2023)
The Rise of AI Companions: Interaction with AI Companions and Psychological Well-being
by: Zhang, Yutong, et al.
Published: (2025)
by: Zhang, Yutong, et al.
Published: (2025)
From Prompts to Reflection: Designing Reflective Play for GenAI Literacy
by: Ma, Qianou, et al.
Published: (2025)
by: Ma, Qianou, et al.
Published: (2025)
Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong
by: Si, Chenglei, et al.
Published: (2023)
by: Si, Chenglei, et al.
Published: (2023)
Behavior Latticing: Inferring User Motivations from Unstructured Interactions
by: Zhao, Dora, et al.
Published: (2026)
by: Zhao, Dora, et al.
Published: (2026)
Can LLM-Simulated Practice and Feedback Upskill Human Counselors? A Randomized Study with 90+ Novice Counselors
by: Louie, Ryan, et al.
Published: (2025)
by: Louie, Ryan, et al.
Published: (2025)
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration
by: Shao, Yijia, et al.
Published: (2024)
by: Shao, Yijia, et al.
Published: (2024)
Orbit: A Framework for Designing and Evaluating Multi-objective Rankers
by: Yang, Chenyang, et al.
Published: (2024)
by: Yang, Chenyang, et al.
Published: (2024)
Wikibench: Community-Driven Data Curation for AI Evaluation on Wikipedia
by: Kuo, Tzu-Sheng, et al.
Published: (2024)
by: Kuo, Tzu-Sheng, et al.
Published: (2024)
Model Cards for AI Teammates: Comparing Human-AI Team Familiarization Methods for High-Stakes Environments
by: Bowers, Ryan, et al.
Published: (2025)
by: Bowers, Ryan, et al.
Published: (2025)
How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
by: Wang, Zora Zhiruo, et al.
Published: (2025)
by: Wang, Zora Zhiruo, et al.
Published: (2025)
Mapping the Spiral of Silence: Surveying Unspoken Opinions in Online Communities
by: Zhao, Dora, et al.
Published: (2025)
by: Zhao, Dora, et al.
Published: (2025)
When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition
by: Li, Karen Jia-Hui, et al.
Published: (2025)
by: Li, Karen Jia-Hui, et al.
Published: (2025)
Position: Towards Bidirectional Human-AI Alignment
by: Shen, Hua, et al.
Published: (2024)
by: Shen, Hua, et al.
Published: (2024)
Roleplay-doh: Enabling Domain-Experts to Create LLM-simulated Patients via Eliciting and Adhering to Principles
by: Louie, Ryan, et al.
Published: (2024)
by: Louie, Ryan, et al.
Published: (2024)
When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration
by: Shi, Quan, et al.
Published: (2025)
by: Shi, Quan, et al.
Published: (2025)
Human-AI Interaction Design Standards
by: Zhao, Chaoyi, et al.
Published: (2025)
by: Zhao, Chaoyi, et al.
Published: (2025)
AI-Induced Human Responsibility (AIHR) in AI-Human teams
by: Nyilasy, Greg, et al.
Published: (2026)
by: Nyilasy, Greg, et al.
Published: (2026)
Towards Human-AI Deliberation: Design and Evaluation of LLM-Empowered Deliberative AI for AI-Assisted Decision-Making
by: Ma, Shuai, et al.
Published: (2024)
by: Ma, Shuai, et al.
Published: (2024)
The AI-DEC: A Card-based Design Method for User-centered AI Explanations
by: Lee, Christine P, et al.
Published: (2024)
by: Lee, Christine P, et al.
Published: (2024)
SparkMe: Adaptive Semi-Structured Interviewing for Qualitative Insight Discovery
by: Anugraha, David, et al.
Published: (2026)
by: Anugraha, David, et al.
Published: (2026)
Generative Experiences for Digital Mental Health Interventions: Evidence from a Randomized Study
by: Bhattacharjee, Ananya, et al.
Published: (2026)
by: Bhattacharjee, Ananya, et al.
Published: (2026)
Learning to Trust: How Humans Mentally Recalibrate AI Confidence Signals
by: Li, ZhaoBin, et al.
Published: (2026)
by: Li, ZhaoBin, et al.
Published: (2026)
Classifying Epistemic Relationships in Human-AI Interaction: An Exploratory Approach
by: Yang, Shengnan, et al.
Published: (2025)
by: Yang, Shengnan, et al.
Published: (2025)
A Two-Phase Visualization System for Continuous Human-AI Collaboration in Sequelae Analysis and Modeling
by: Ouyang, Yang, et al.
Published: (2024)
by: Ouyang, Yang, et al.
Published: (2024)
As Confidence Aligns: Exploring the Effect of AI Confidence on Human Self-confidence in Human-AI Decision Making
by: Li, Jingshu, et al.
Published: (2025)
by: Li, Jingshu, et al.
Published: (2025)
Whose Knowledge Counts? Co-Designing Community-Centered AI Auditing Tools with Educators in Hawai`i
by: Zhao, Dora, et al.
Published: (2026)
by: Zhao, Dora, et al.
Published: (2026)
Multi-Level Feedback Generation with Large Language Models for Empowering Novice Peer Counselors
by: Chaszczewicz, Alicja, et al.
Published: (2024)
by: Chaszczewicz, Alicja, et al.
Published: (2024)
Human-AI Interaction Alignment: Designing, Evaluating, and Evolving Value-Centered AI For Reciprocal Human-AI Futures
by: Shen, Hua, et al.
Published: (2025)
by: Shen, Hua, et al.
Published: (2025)
Human-AI Co-Evolution and Epistemic Collapse: A Dynamical Systems Perspective
by: Wu, Xuening, et al.
Published: (2026)
by: Wu, Xuening, et al.
Published: (2026)
Selenite: Scaffolding Online Sensemaking with Comprehensive Overviews Elicited from Large Language Models
by: Liu, Michael Xieyang, et al.
Published: (2023)
by: Liu, Michael Xieyang, et al.
Published: (2023)
Interoceptive Divergence in Aesthetic Evaluation and Implications for Human-AI Alignment
by: Abe, Yoshia, et al.
Published: (2026)
by: Abe, Yoshia, et al.
Published: (2026)
Evaluating Human-AI Collaboration: A Review and Methodological Framework
by: Fragiadakis, George, et al.
Published: (2024)
by: Fragiadakis, George, et al.
Published: (2024)
Similar Items
-
What Should We Engineer in Prompts? Training Humans in Requirement-Driven LLM Use
by: Ma, Qianou, et al.
Published: (2024) -
Not Everyone Wins with LLMs: Behavioral Patterns and Pedagogical Implications for AI Literacy in Programmatic Data Science
by: Ma, Qianou, et al.
Published: (2025) -
The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
by: Si, Chenglei, et al.
Published: (2025) -
Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
by: Si, Chenglei, et al.
Published: (2024) -
How to Teach Programming in the AI Era? Using LLMs as a Teachable Agent for Debugging
by: Ma, Qianou, et al.
Published: (2023)