Rethinking Model Evaluation as Narrowing the Socio-Technical Gap
Fuente:
arXiv
Guardado en:
| Autores principales: | Liao, Q. Vera, Xiao, Ziang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Generative Echo Chamber? Effects of LLM-Powered Search Systems on Diverse Information Seeking
por: Sharma, Nikhil, et al.
Publicado: (2024)
por: Sharma, Nikhil, et al.
Publicado: (2024)
Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild
por: van der Maden, Willem, et al.
Publicado: (2026)
por: van der Maden, Willem, et al.
Publicado: (2026)
Why is AI not a Panacea for Data Workers? An Interview Study on Human-AI Collaboration in Data Storytelling
por: Li, Haotian, et al.
Publicado: (2023)
por: Li, Haotian, et al.
Publicado: (2023)
From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing Assistants
por: Rismani, Shalaleh, et al.
Publicado: (2026)
por: Rismani, Shalaleh, et al.
Publicado: (2026)
Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies
por: Kim, Sunnie S. Y., et al.
Publicado: (2025)
por: Kim, Sunnie S. Y., et al.
Publicado: (2025)
"I'm Not Sure, But...": Examining the Impact of Large Language Models' Uncertainty Expression on User Reliance and Trust
por: Kim, Sunnie S. Y., et al.
Publicado: (2024)
por: Kim, Sunnie S. Y., et al.
Publicado: (2024)
STAR: SocioTechnical Approach to Red Teaming Language Models
por: Weidinger, Laura, et al.
Publicado: (2024)
por: Weidinger, Laura, et al.
Publicado: (2024)
Red Teaming LLMs as Socio-Technical Practice: From Exploration and Data Creation to Evaluation
por: Garcia, Adriana Alvarado, et al.
Publicado: (2026)
por: Garcia, Adriana Alvarado, et al.
Publicado: (2026)
Generation Probabilities Are Not Enough: Uncertainty Highlighting in AI Code Completions
por: Vasconcelos, Helena, et al.
Publicado: (2023)
por: Vasconcelos, Helena, et al.
Publicado: (2023)
As Confidence Aligns: Exploring the Effect of AI Confidence on Human Self-confidence in Human-AI Decision Making
por: Li, Jingshu, et al.
Publicado: (2025)
por: Li, Jingshu, et al.
Publicado: (2025)
Seamful XAI: Operationalizing Seamful Design in Explainable AI
por: Ehsan, Upol, et al.
Publicado: (2022)
por: Ehsan, Upol, et al.
Publicado: (2022)
"It was 80% me, 20% AI": Seeking Authenticity in Co-Writing with Large Language Models
por: Hwang, Angel Hsing-Chi, et al.
Publicado: (2024)
por: Hwang, Angel Hsing-Chi, et al.
Publicado: (2024)
Understanding the Effects of Miscalibrated AI Confidence on User Trust, Reliance, and Decision Efficacy
por: Li, Jingshu, et al.
Publicado: (2024)
por: Li, Jingshu, et al.
Publicado: (2024)
Evaluating AI for Law: Bridging the Gap with Open-Source Solutions
por: Bhambhoria, Rohan, et al.
Publicado: (2024)
por: Bhambhoria, Rohan, et al.
Publicado: (2024)
Integrating Virtual Reality and Large Language Models for Team-Based Non-Technical Skills Training and Evaluation in the Operating Room
por: Barker, Jacob, et al.
Publicado: (2026)
por: Barker, Jacob, et al.
Publicado: (2026)
When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis
por: Najera, Aisha, et al.
Publicado: (2026)
por: Najera, Aisha, et al.
Publicado: (2026)
In Dialogue with Intelligence: Rethinking Large Language Models as Collective Knowledge
por: Vasilaki, Eleni
Publicado: (2025)
por: Vasilaki, Eleni
Publicado: (2025)
Beyond Permissions: Investigating Mobile Personalization with Simulated Personas
por: Khalilov, Ibrahim, et al.
Publicado: (2025)
por: Khalilov, Ibrahim, et al.
Publicado: (2025)
PsyLite Technical Report
por: Ding, Fangjun, et al.
Publicado: (2025)
por: Ding, Fangjun, et al.
Publicado: (2025)
The Epistemological Consequences of Large Language Models: Rethinking collective intelligence and institutional knowledge
por: Hila, Angjelin
Publicado: (2025)
por: Hila, Angjelin
Publicado: (2025)
Assistance or Disruption? Exploring and Evaluating the Design and Trade-offs of Proactive AI Programming Support
por: Pu, Kevin, et al.
Publicado: (2025)
por: Pu, Kevin, et al.
Publicado: (2025)
INSIGHT: Bridging the Student-Teacher Gap in Times of Large Language Models
por: Thys, Jarne, et al.
Publicado: (2025)
por: Thys, Jarne, et al.
Publicado: (2025)
MITHOS: Interactive Mixed Reality Training to Support Professional Socio-Emotional Interactions at Schools
por: Chehayeb, Lara, et al.
Publicado: (2024)
por: Chehayeb, Lara, et al.
Publicado: (2024)
Guided Reasoning: A Non-Technical Introduction
por: Betz, Gregor
Publicado: (2024)
por: Betz, Gregor
Publicado: (2024)
Meta-Evaluating Local LLMs: Rethinking Performance Metrics for Serious Games
por: Isaza-Giraldo, Andrés, et al.
Publicado: (2025)
por: Isaza-Giraldo, Andrés, et al.
Publicado: (2025)
I Hear, Therefore I Trust: A Socio-Technical Investigation of Humans as Synthetic Speech Detectors
por: Erscoi, Lelia, et al.
Publicado: (2026)
por: Erscoi, Lelia, et al.
Publicado: (2026)
The Who in XAI: How AI Background Shapes Perceptions of AI Explanations
por: Ehsan, Upol, et al.
Publicado: (2021)
por: Ehsan, Upol, et al.
Publicado: (2021)
Technically Love: The Evolution of Human-AI Romance Discourse on Reddit
por: Chang, Tyler, et al.
Publicado: (2026)
por: Chang, Tyler, et al.
Publicado: (2026)
Alignment-Process-Outcome: Rethinking How AIs and Humans Collaborate
por: Li, Haichang, et al.
Publicado: (2026)
por: Li, Haichang, et al.
Publicado: (2026)
From Passive Tool to Socio-cognitive Teammate: A Conceptual Framework for Agentic AI in Human-AI Collaborative Learning
por: Yan, Lixiang
Publicado: (2025)
por: Yan, Lixiang
Publicado: (2025)
Keeping an Eye on AI: A Framework for Effective Human Oversight of AI Systems
por: Gaube, Susanne, et al.
Publicado: (2026)
por: Gaube, Susanne, et al.
Publicado: (2026)
Rethinking Health Agents: From Siloed AI to Collaborative Decision Mediators
por: Chung, Ray-Yuan, et al.
Publicado: (2026)
por: Chung, Ray-Yuan, et al.
Publicado: (2026)
Insights from Railway Professionals: Rethinking Railway assumptions regarding safety and autonomy
por: Hunter, Josh, et al.
Publicado: (2025)
por: Hunter, Josh, et al.
Publicado: (2025)
Rethinking Theory of Mind Benchmarks for LLMs: Towards A User-Centered Perspective
por: Wang, Qiaosi, et al.
Publicado: (2025)
por: Wang, Qiaosi, et al.
Publicado: (2025)
The Widening Gap: The Benefits and Harms of Generative AI for Novice Programmers
por: Prather, James, et al.
Publicado: (2024)
por: Prather, James, et al.
Publicado: (2024)
MASCOT: Towards Multi-Agent Socio-Collaborative Companion Systems
por: Wang, Yiyang, et al.
Publicado: (2026)
por: Wang, Yiyang, et al.
Publicado: (2026)
Disrupting Cognitive Passivity: Rethinking AI-Assisted Data Literacy through Cognitive Alignment
por: Ahn, Yongsu, et al.
Publicado: (2026)
por: Ahn, Yongsu, et al.
Publicado: (2026)
From "AI" to Probabilistic Automation: How Does Anthropomorphization of Technical Systems Descriptions Influence Trust?
por: Inie, Nanna, et al.
Publicado: (2024)
por: Inie, Nanna, et al.
Publicado: (2024)
Designing Conversational AI to Support Think-Aloud Practice in Technical Interview Preparation for CS Students
por: Daryanto, Taufiq, et al.
Publicado: (2025)
por: Daryanto, Taufiq, et al.
Publicado: (2025)
Systematizing LLM Persona Design: A Four-Quadrant Technical Taxonomy for AI Companion Applications
por: Sun, Esther, et al.
Publicado: (2025)
por: Sun, Esther, et al.
Publicado: (2025)
Ejemplares similares
-
Generative Echo Chamber? Effects of LLM-Powered Search Systems on Diverse Information Seeking
por: Sharma, Nikhil, et al.
Publicado: (2024) -
Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild
por: van der Maden, Willem, et al.
Publicado: (2026) -
Why is AI not a Panacea for Data Workers? An Interview Study on Human-AI Collaboration in Data Storytelling
por: Li, Haotian, et al.
Publicado: (2023) -
From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing Assistants
por: Rismani, Shalaleh, et al.
Publicado: (2026) -
Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies
por: Kim, Sunnie S. Y., et al.
Publicado: (2025)