AI-Assisted Systematization for Evaluating GenAI Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Agarwal, Dhruv, Sheng, Emily, Atalla, Chad, Garcia-Gathright, Jean, Mozannar, Hussein, Washington, Hannah, Chouldechova, Alexandra, Barocas, Solon, Wallach, Hanna |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dimensions of Generative AI Evaluation Design
by: Dow, P. Alex, et al.
Published: (2024)
by: Dow, P. Alex, et al.
Published: (2024)
A Framework for Evaluating LLMs Under Task Indeterminacy
by: Guerdan, Luke, et al.
Published: (2024)
by: Guerdan, Luke, et al.
Published: (2024)
A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts
by: Chouldechova, Alexandra, et al.
Published: (2024)
by: Chouldechova, Alexandra, et al.
Published: (2024)
Taxonomizing Representational Harms using Speech Act Theory
by: Corvi, Emily, et al.
Published: (2025)
by: Corvi, Emily, et al.
Published: (2025)
Evaluating Generative AI Systems is a Social Science Measurement Challenge
by: Wallach, Hanna, et al.
Published: (2024)
by: Wallach, Hanna, et al.
Published: (2024)
Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
by: Wallach, Hanna, et al.
Published: (2025)
by: Wallach, Hanna, et al.
Published: (2025)
Validating LLM-as-a-Judge Systems under Rating Indeterminacy
by: Guerdan, Luke, et al.
Published: (2025)
by: Guerdan, Luke, et al.
Published: (2025)
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems
by: Harvey, Emma, et al.
Published: (2025)
by: Harvey, Emma, et al.
Published: (2025)
Effects of Generative AI Errors on User Reliance Across Task Difficulty
by: Anthis, Jacy Reese, et al.
Published: (2026)
by: Anthis, Jacy Reese, et al.
Published: (2026)
Gaps Between Research and Practice When Measuring Representational Harms Caused by LLM-Based Systems
by: Harvey, Emma, et al.
Published: (2024)
by: Harvey, Emma, et al.
Published: (2024)
Comparison requires valid measurement: Rethinking attack success rate comparisons in AI red teaming
by: Chouldechova, Alexandra, et al.
Published: (2026)
by: Chouldechova, Alexandra, et al.
Published: (2026)
GenAI Assisting Medical Training
by: Fritsch, Stefan, et al.
Published: (2024)
by: Fritsch, Stefan, et al.
Published: (2024)
Putting GenAI on Notice: GenAI Exceptionalism and Contract Law
by: Atkinson, David
Published: (2025)
by: Atkinson, David
Published: (2025)
Troubling Taxonomies in GenAI Evaluation
by: Berman, Glen, et al.
Published: (2024)
by: Berman, Glen, et al.
Published: (2024)
GenLens: A Systematic Evaluation of Visual GenAI Model Outputs
by: Lin, Tica, et al.
Published: (2024)
by: Lin, Tica, et al.
Published: (2024)
Reading Between the Lines: Modeling User Behavior and Costs in AI-Assisted Programming
by: Mozannar, Hussein, et al.
Published: (2022)
by: Mozannar, Hussein, et al.
Published: (2022)
When to Show a Suggestion? Integrating Human Feedback in AI-Assisted Programming
by: Mozannar, Hussein, et al.
Published: (2023)
by: Mozannar, Hussein, et al.
Published: (2023)
GenAI Distortion: The Effect of GenAI Fluency and Positive Affect
by: Yang, Xiantong, et al.
Published: (2024)
by: Yang, Xiantong, et al.
Published: (2024)
AI Automatons: AI Systems Intended to Imitate Humans
by: Olteanu, Alexandra, et al.
Published: (2025)
by: Olteanu, Alexandra, et al.
Published: (2025)
Distinguishing Task-Specific and General-Purpose AI in Regulation
by: Wang, Jennifer, et al.
Published: (2025)
by: Wang, Jennifer, et al.
Published: (2025)
Dataset of GenAI-Assisted Information Problem Solving in Education
by: Li, Xinyu, et al.
Published: (2026)
by: Li, Xinyu, et al.
Published: (2026)
Designing Algorithmic Delegates: The Role of Indistinguishability in Human-AI Handoff
by: Greenwood, Sophie, et al.
Published: (2025)
by: Greenwood, Sophie, et al.
Published: (2025)
The Evolving Usage of GenAI by Computing Students
by: Hou, Irene, et al.
Published: (2024)
by: Hou, Irene, et al.
Published: (2024)
Duplicate Detection with GenAI
by: Ormesher, Ian
Published: (2024)
by: Ormesher, Ian
Published: (2024)
Do Responsible AI Artifacts Advance Stakeholder Goals? Four Key Barriers Perceived by Legal and Civil Stakeholders
by: Kawakami, Anna, et al.
Published: (2024)
by: Kawakami, Anna, et al.
Published: (2024)
Lessons for GenAI Literacy From a Field Study of Human-GenAI Augmentation in the Workplace
by: Johri, Aditya, et al.
Published: (2025)
by: Johri, Aditya, et al.
Published: (2025)
GenAI Voice Mode in Programming Education
by: Jacobs, Sven, et al.
Published: (2025)
by: Jacobs, Sven, et al.
Published: (2025)
Can GenAI Move from Individual Use to Collaborative Work? Experiences, Challenges, and Opportunities of Coordinating GenAI into Collaborative Newswork
by: Xiao, Qing, et al.
Published: (2025)
by: Xiao, Qing, et al.
Published: (2025)
Content Platform GenAI Regulation via Compensation
by: Chaimanowong, Wee
Published: (2026)
by: Chaimanowong, Wee
Published: (2026)
Collaborating with GenAI: Incentives and Replacements
by: Taitler, Boaz, et al.
Published: (2025)
by: Taitler, Boaz, et al.
Published: (2025)
GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation
by: Li, Baiqi, et al.
Published: (2024)
by: Li, Baiqi, et al.
Published: (2024)
Haptic Repurposing with GenAI
by: Wang, Haoyu
Published: (2024)
by: Wang, Haoyu
Published: (2024)
What Constitutes a Less Discriminatory Algorithm?
by: Laufer, Benjamin, et al.
Published: (2024)
by: Laufer, Benjamin, et al.
Published: (2024)
GenAI Arena: An Open Evaluation Platform for Generative Models
by: Jiang, Dongfu, et al.
Published: (2024)
by: Jiang, Dongfu, et al.
Published: (2024)
Framework for Adoption of Generative Artificial Intelligence (GenAI) in Education
by: Shailendra, Samar, et al.
Published: (2024)
by: Shailendra, Samar, et al.
Published: (2024)
Terms of (Ab)Use: An Analysis of GenAI Services
by: Pandit, Harshvardhan J., et al.
Published: (2026)
by: Pandit, Harshvardhan J., et al.
Published: (2026)
The Fast and Spurious: Developer Productivity with GenAI
by: Afroz, Sadia, et al.
Published: (2025)
by: Afroz, Sadia, et al.
Published: (2025)
Face Consistency Benchmark for GenAI Video
by: Podstawski, Michal, et al.
Published: (2025)
by: Podstawski, Michal, et al.
Published: (2025)
Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review
by: Pang, Rock Yuren, et al.
Published: (2025)
by: Pang, Rock Yuren, et al.
Published: (2025)
Selective Response Strategies for GenAI
by: Taitler, Boaz, et al.
Published: (2025)
by: Taitler, Boaz, et al.
Published: (2025)
Similar Items
-
Dimensions of Generative AI Evaluation Design
by: Dow, P. Alex, et al.
Published: (2024) -
A Framework for Evaluating LLMs Under Task Indeterminacy
by: Guerdan, Luke, et al.
Published: (2024) -
A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts
by: Chouldechova, Alexandra, et al.
Published: (2024) -
Taxonomizing Representational Harms using Speech Act Theory
by: Corvi, Emily, et al.
Published: (2025) -
Evaluating Generative AI Systems is a Social Science Measurement Challenge
by: Wallach, Hanna, et al.
Published: (2024)