Towards Real-World Validity in Generative AI Benchmarks: Understanding and Designing Domain-Centered Evaluations for Journalism Practitioners
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Charlotte, Hagar, Nick, Nishal, Sachita, Gilbert, Jeremy, Diakopoulos, Nick |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
De-jargonizing Science for Journalists with GPT-4: A Pilot Study
by: Nishal, Sachita, et al.
Published: (2024)
by: Nishal, Sachita, et al.
Published: (2024)
Domain-Specific Evaluation Strategies for AI in Journalism
by: Nishal, Sachita, et al.
Published: (2024)
by: Nishal, Sachita, et al.
Published: (2024)
LLM-Assisted News Discovery in High-Volume Information Streams: A Case Study
by: Hagar, Nick, et al.
Published: (2025)
by: Hagar, Nick, et al.
Published: (2025)
Envisioning the Applications and Implications of Generative AI for News Media
by: Nishal, Sachita, et al.
Published: (2024)
by: Nishal, Sachita, et al.
Published: (2024)
On-Premise AI for the Newsroom: Evaluating Small Language Models for Investigative Document Search
by: Hagar, Nick, et al.
Published: (2025)
by: Hagar, Nick, et al.
Published: (2025)
Design Generative AI for Practitioners: Exploring Interaction Approaches Aligned with Creative Practice
by: Peng, Xiaohan, et al.
Published: (2026)
by: Peng, Xiaohan, et al.
Published: (2026)
OmniActions: Predicting Digital Actions in Response to Real-World Multimodal Sensory Inputs with LLMs
by: Li, Jiahao Nick, et al.
Published: (2024)
by: Li, Jiahao Nick, et al.
Published: (2024)
Beyond Isolation: Towards an Interactionist Perspective on Human Cognitive Bias and AI Bias
by: von Felten, Nick
Published: (2025)
by: von Felten, Nick
Published: (2025)
Towards Apples to Apples for AI Evaluations: From Real-World Use Cases to Evaluation Scenarios
by: Choong, Yee-Yin, et al.
Published: (2026)
by: Choong, Yee-Yin, et al.
Published: (2026)
LLM BiasScope: A Real-Time Bias Analysis Platform for Comparative LLM Evaluation
by: Ghosh, Himel, et al.
Published: (2026)
by: Ghosh, Himel, et al.
Published: (2026)
Situated, Dynamic, and Subjective: Envisioning the Design of Theory-of-Mind-Enabled Everyday AI with Industry Practitioners
by: Wang, Qiaosi, et al.
Published: (2026)
by: Wang, Qiaosi, et al.
Published: (2026)
Envisioning Stakeholder-Action Pairs to Mitigate Negative Impacts of AI: A Participatory Approach to Inform Policy Making
by: Barnett, Julia, et al.
Published: (2025)
by: Barnett, Julia, et al.
Published: (2025)
Scenarios in Computing Research: A Systematic Review of the Use of Scenario Methods for Exploring the Future of Computing Technologies in Society
by: Barnett, Julia, et al.
Published: (2025)
by: Barnett, Julia, et al.
Published: (2025)
Towards Feature Engineering with Human and AI's Knowledge: Understanding Data Science Practitioners' Perceptions in Human&AI-Assisted Feature Engineering Design
by: Zhu, Qian, et al.
Published: (2024)
by: Zhu, Qian, et al.
Published: (2024)
How Human-Centered Explainable AI Interface Are Designed and Evaluated: A Systematic Survey
by: Nguyen, Thu, et al.
Published: (2024)
by: Nguyen, Thu, et al.
Published: (2024)
HCC Is All You Need: Alignment-The Sensible Kind Anyway-Is Just Human-Centered Computing
by: Gilbert, Eric
Published: (2024)
by: Gilbert, Eric
Published: (2024)
Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild
by: van der Maden, Willem, et al.
Published: (2026)
by: van der Maden, Willem, et al.
Published: (2026)
Creativity in the Age of AI: Evaluating the Impact of Generative AI on Design Outputs and Designers' Creative Thinking
by: Fu, Yue, et al.
Published: (2024)
by: Fu, Yue, et al.
Published: (2024)
Biased Minds Meet Biased AI: How Class Imbalance Shapes Appropriate Reliance and Interacts with Human Base Rate Neglect
by: von Felten, Nick, et al.
Published: (2025)
by: von Felten, Nick, et al.
Published: (2025)
Towards Human-AI Deliberation: Design and Evaluation of LLM-Empowered Deliberative AI for AI-Assisted Decision-Making
by: Ma, Shuai, et al.
Published: (2024)
by: Ma, Shuai, et al.
Published: (2024)
Inclusive Practices for Child-Centered AI Design and Testing
by: Dotch, Emani, et al.
Published: (2024)
by: Dotch, Emani, et al.
Published: (2024)
Measuring skill-based uplift from AI in a real biological laboratory
by: Romero-Severson, Ethan Obie, et al.
Published: (2025)
by: Romero-Severson, Ethan Obie, et al.
Published: (2025)
Anticipating Impacts: Using Large-Scale Scenario Writing to Explore Diverse Implications of Generative AI in the News Environment
by: Kieslich, Kimon, et al.
Published: (2023)
by: Kieslich, Kimon, et al.
Published: (2023)
WearBCI Dataset: Understanding and Benchmarking Real-World Wearable Brain-Computer Interfaces Signals
by: Liu, Haoxian, et al.
Published: (2026)
by: Liu, Haoxian, et al.
Published: (2026)
Leveraging Generative AI for Human Understanding: Meta-Requirements and Design Principles for Explanatory AI as a new Paradigm
by: Meske, Christian, et al.
Published: (2025)
by: Meske, Christian, et al.
Published: (2025)
Understanding Critical Thinking in Generative Artificial Intelligence Use: Development, Validation, and Correlates of the Critical Thinking in AI Use Scale
by: Lau, Gabriel R., et al.
Published: (2025)
by: Lau, Gabriel R., et al.
Published: (2025)
Controlling Context: Generative AI at Work in Integrated Circuit Design and Other High-Precision Domains
by: Moss, Emanuel, et al.
Published: (2025)
by: Moss, Emanuel, et al.
Published: (2025)
Rethinking Theory of Mind Benchmarks for LLMs: Towards A User-Centered Perspective
by: Wang, Qiaosi, et al.
Published: (2025)
by: Wang, Qiaosi, et al.
Published: (2025)
ChatCollab: Exploring Collaboration Between Humans and AI Agents in Software Teams
by: Klieger, Benjamin, et al.
Published: (2024)
by: Klieger, Benjamin, et al.
Published: (2024)
Insights Informed Generative AI for Design: Incorporating Real-world Data for Text-to-Image Output
by: Gupta, Richa, et al.
Published: (2025)
by: Gupta, Richa, et al.
Published: (2025)
Toward a Human-Centered AI-assisted Colonoscopy System in Australia
by: Chen, Hsiang-Ting, et al.
Published: (2025)
by: Chen, Hsiang-Ting, et al.
Published: (2025)
Evaluating LLMs for Visualization Generation and Understanding
by: Khan, Saadiq Rauf, et al.
Published: (2025)
by: Khan, Saadiq Rauf, et al.
Published: (2025)
Designing AI for Real Users -- Accessibility Gaps in Retail AI Front-End
by: Puri, Neha, et al.
Published: (2026)
by: Puri, Neha, et al.
Published: (2026)
Understanding the Dataset Practitioners Behind Large Language Model Development
by: Qian, Crystal, et al.
Published: (2024)
by: Qian, Crystal, et al.
Published: (2024)
Towards Experience-Centered AI: A Framework for Integrating Lived Experience in Design and Development
by: Gautam, Sanjana, et al.
Published: (2025)
by: Gautam, Sanjana, et al.
Published: (2025)
Human-AI Interaction Alignment: Designing, Evaluating, and Evolving Value-Centered AI For Reciprocal Human-AI Futures
by: Shen, Hua, et al.
Published: (2025)
by: Shen, Hua, et al.
Published: (2025)
AI for Requirements Engineering: Industry adoption and Practitioner perspectives
by: Rani, Lekshmi Murali, et al.
Published: (2025)
by: Rani, Lekshmi Murali, et al.
Published: (2025)
"I Cannot Write This Because It Violates Our Content Policy": Understanding Content Moderation Policies and User Experiences in Generative AI Products
by: Gao, Lan, et al.
Published: (2025)
by: Gao, Lan, et al.
Published: (2025)
Interaction-Centered Intelligence: Toward Interaction as the Primary Unit of Analysis in Co-Creative AI and Human-AI Systems
by: Davis, Nicholas
Published: (2026)
by: Davis, Nicholas
Published: (2026)
Tracing the Invisible: Understanding Students' Judgment in AI-Supported Design Work
by: Naik, Suchismita, et al.
Published: (2025)
by: Naik, Suchismita, et al.
Published: (2025)
Similar Items
-
De-jargonizing Science for Journalists with GPT-4: A Pilot Study
by: Nishal, Sachita, et al.
Published: (2024) -
Domain-Specific Evaluation Strategies for AI in Journalism
by: Nishal, Sachita, et al.
Published: (2024) -
LLM-Assisted News Discovery in High-Volume Information Streams: A Case Study
by: Hagar, Nick, et al.
Published: (2025) -
Envisioning the Applications and Implications of Generative AI for News Media
by: Nishal, Sachita, et al.
Published: (2024) -
On-Premise AI for the Newsroom: Evaluating Small Language Models for Investigative Document Search
by: Hagar, Nick, et al.
Published: (2025)