HEDS 3.0: The Human Evaluation Data Sheet Version 3.0
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Belz, Anya, Thomson, Craig |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A User Experience 3.0 (UX 3.0) Paradigm Framework: Designing for Human-Centered AI Experiences
von: Xu, Wei
Veröffentlicht: (2025)
von: Xu, Wei
Veröffentlicht: (2025)
A "User Experience 3.0 (UX 3.0)" Paradigm Framework: User Experience Design for Human-Centered AI Systems
von: Xu, Wei
Veröffentlicht: (2024)
von: Xu, Wei
Veröffentlicht: (2024)
Code Style Sheets: CSS for Code
von: Cohen, Sam, et al.
Veröffentlicht: (2025)
von: Cohen, Sam, et al.
Veröffentlicht: (2025)
Pearmut: Human Evaluation of Translation Made Trivial
von: Zouhar, Vilém, et al.
Veröffentlicht: (2026)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2026)
Context-Aware Monolingual Human Evaluation of Machine Translation
von: Picinini, Silvio, et al.
Veröffentlicht: (2025)
von: Picinini, Silvio, et al.
Veröffentlicht: (2025)
Synchronized Realities: Towards Magic Mobile Experiences through Aligned AR
von: Belz, Jan Henry
Veröffentlicht: (2026)
von: Belz, Jan Henry
Veröffentlicht: (2026)
Kwame 2.0: Human-in-the-Loop Generative AI Teaching Assistant for Large Scale Online Coding Education in Africa
von: Boateng, George, et al.
Veröffentlicht: (2026)
von: Boateng, George, et al.
Veröffentlicht: (2026)
From Seeing it to Experiencing it: Interactive Evaluation of Intersectional Voice Bias in Human-AI Speech Interaction
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale
von: Li, Weiyue, et al.
Veröffentlicht: (2026)
von: Li, Weiyue, et al.
Veröffentlicht: (2026)
IP-Dialog: Evaluating Implicit Personalization in Dialogue Systems with Synthetic Data
von: Peng, Bo, et al.
Veröffentlicht: (2025)
von: Peng, Bo, et al.
Veröffentlicht: (2025)
OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
The Renaissance of Repair: A Timely Opportunity for Fabrication Research
von: Britten, Julian, et al.
Veröffentlicht: (2026)
von: Britten, Julian, et al.
Veröffentlicht: (2026)
AI-native Memory 2.0: Second Me
von: Wei, Jiale, et al.
Veröffentlicht: (2025)
von: Wei, Jiale, et al.
Veröffentlicht: (2025)
Are Humans as Brittle as Large Language Models?
von: Li, Jiahui, et al.
Veröffentlicht: (2025)
von: Li, Jiahui, et al.
Veröffentlicht: (2025)
Evaluating LLMs as Human Surrogates in Controlled Experiments
von: Hoq, Adnan, et al.
Veröffentlicht: (2026)
von: Hoq, Adnan, et al.
Veröffentlicht: (2026)
Modeling Distinct Human Interaction in Web Agents
von: Huq, Faria, et al.
Veröffentlicht: (2026)
von: Huq, Faria, et al.
Veröffentlicht: (2026)
MEGAnno+: A Human-LLM Collaborative Annotation System
von: Kim, Hannah, et al.
Veröffentlicht: (2024)
von: Kim, Hannah, et al.
Veröffentlicht: (2024)
Lexical Indicators of Mind Perception in Human-AI Companionship
von: Banks, Jaime, et al.
Veröffentlicht: (2026)
von: Banks, Jaime, et al.
Veröffentlicht: (2026)
Large Language Models for Virtual Human Gesture Selection
von: Torshizi, Parisa Ghanad, et al.
Veröffentlicht: (2025)
von: Torshizi, Parisa Ghanad, et al.
Veröffentlicht: (2025)
Navigating Rifts in Human-LLM Grounding: Study and Benchmark
von: Shaikh, Omar, et al.
Veröffentlicht: (2025)
von: Shaikh, Omar, et al.
Veröffentlicht: (2025)
Practicing with Language Models Cultivates Human Empathic Communication
von: Kumar, Aakriti, et al.
Veröffentlicht: (2026)
von: Kumar, Aakriti, et al.
Veröffentlicht: (2026)
Human-LLM Collaborative Construction of a Cantonese Emotion Lexicon
von: Zhang, Yusong, et al.
Veröffentlicht: (2024)
von: Zhang, Yusong, et al.
Veröffentlicht: (2024)
Althea: Human-AI Collaboration for Fact-Checking and Critical Reasoning
von: Churina, Svetlana, et al.
Veröffentlicht: (2025)
von: Churina, Svetlana, et al.
Veröffentlicht: (2025)
Learning Next Action Predictors from Human-Computer Interaction
von: Shaikh, Omar, et al.
Veröffentlicht: (2026)
von: Shaikh, Omar, et al.
Veröffentlicht: (2026)
CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
von: Verma, Arnav, et al.
Veröffentlicht: (2025)
von: Verma, Arnav, et al.
Veröffentlicht: (2025)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
von: Vo, Truong, et al.
Veröffentlicht: (2025)
von: Vo, Truong, et al.
Veröffentlicht: (2025)
The QCET Taxonomy of Standard Quality Criterion Names and Definitions for the Evaluation of NLP Systems
von: Belz, Anya, et al.
Veröffentlicht: (2025)
von: Belz, Anya, et al.
Veröffentlicht: (2025)
Characterizing Similarities and Divergences in Conversational Tones in Humans and LLMs by Sampling with People
von: Huang, Dun-Ming, et al.
Veröffentlicht: (2024)
von: Huang, Dun-Ming, et al.
Veröffentlicht: (2024)
Show or Tell? Modeling the evolution of request-making in Human-LLM conversations
von: Zhu, Shengqi, et al.
Veröffentlicht: (2025)
von: Zhu, Shengqi, et al.
Veröffentlicht: (2025)
LLMs as Workers in Human-Computational Algorithms? Replicating Crowdsourcing Pipelines with LLMs
von: Wu, Tongshuang, et al.
Veröffentlicht: (2023)
von: Wu, Tongshuang, et al.
Veröffentlicht: (2023)
PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2026)
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2026)
Exploring Empty Spaces: Human-in-the-Loop Data Augmentation
von: Yeh, Catherine, et al.
Veröffentlicht: (2024)
von: Yeh, Catherine, et al.
Veröffentlicht: (2024)
Human-LLM Collaborative Feature Engineering for Tabular Data
von: Li, Zhuoyan, et al.
Veröffentlicht: (2026)
von: Li, Zhuoyan, et al.
Veröffentlicht: (2026)
On Evaluating Explanation Utility for Human-AI Decision Making in NLP
von: Chaleshtori, Fateme Hashemi, et al.
Veröffentlicht: (2024)
von: Chaleshtori, Fateme Hashemi, et al.
Veröffentlicht: (2024)
Beyond Prompts: Learning from Human Communication for Enhanced AI Intent Alignment
von: Kim, Yoonsu, et al.
Veröffentlicht: (2024)
von: Kim, Yoonsu, et al.
Veröffentlicht: (2024)
Large Language Model-based Human-Agent Collaboration for Complex Task Solving
von: Feng, Xueyang, et al.
Veröffentlicht: (2024)
von: Feng, Xueyang, et al.
Veröffentlicht: (2024)
The LLM Effect: Are Humans Truly Using LLMs, or Are They Being Influenced By Them Instead?
von: Choi, Alexander S., et al.
Veröffentlicht: (2024)
von: Choi, Alexander S., et al.
Veröffentlicht: (2024)
Human-AI Narrative Synthesis to Foster Shared Understanding in Civic Decision-Making
von: Overney, Cassandra, et al.
Veröffentlicht: (2025)
von: Overney, Cassandra, et al.
Veröffentlicht: (2025)
Beyond Turn-taking: Introducing Text-based Overlap into Human-LLM Interactions
von: Kim, JiWoo, et al.
Veröffentlicht: (2025)
von: Kim, JiWoo, et al.
Veröffentlicht: (2025)
Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild
von: Mysore, Sheshera, et al.
Veröffentlicht: (2025)
von: Mysore, Sheshera, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A User Experience 3.0 (UX 3.0) Paradigm Framework: Designing for Human-Centered AI Experiences
von: Xu, Wei
Veröffentlicht: (2025) -
A "User Experience 3.0 (UX 3.0)" Paradigm Framework: User Experience Design for Human-Centered AI Systems
von: Xu, Wei
Veröffentlicht: (2024) -
Code Style Sheets: CSS for Code
von: Cohen, Sam, et al.
Veröffentlicht: (2025) -
Pearmut: Human Evaluation of Translation Made Trivial
von: Zouhar, Vilém, et al.
Veröffentlicht: (2026) -
Context-Aware Monolingual Human Evaluation of Machine Translation
von: Picinini, Silvio, et al.
Veröffentlicht: (2025)