The Generative AI Paradox on Evaluation: What It Can Solve, It May Not Evaluate
Fuente:
arXiv
Saved in:
| Main Authors: | Oh, Juhyun, Kim, Eunsu, Cha, Inha, Oh, Alice |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
Uncovering Factor Level Preferences to Improve Human-Model Alignment
by: Oh, Juhyun, et al.
Published: (2024)
by: Oh, Juhyun, et al.
Published: (2024)
Designing and Evaluating Multi-Chatbot Interface for Human-AI Communication: Preliminary Findings from a Persuasion Task
by: Yoon, Sion, et al.
Published: (2024)
by: Yoon, Sion, et al.
Published: (2024)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
Do Ethical AI Principles Matter to Users? A Large-Scale Analysis of User Sentiment and Satisfaction
by: Pasch, Stefan, et al.
Published: (2025)
by: Pasch, Stefan, et al.
Published: (2025)
Can Generative AI Support Patients' & Caregivers' Informational Needs? Towards Task-Centric Evaluation Of AI Systems
by: Rajagopal, Shreya, et al.
Published: (2024)
by: Rajagopal, Shreya, et al.
Published: (2024)
EUDAIMONIA: Evaluating Undesirable Dynamics in AI
by: Huang, Jun Rui, et al.
Published: (2026)
by: Huang, Jun Rui, et al.
Published: (2026)
The Persuasion Paradox: When LLM Explanations Fail to Improve Human-AI Team Performance
by: Cohen, Ruth, et al.
Published: (2026)
by: Cohen, Ruth, et al.
Published: (2026)
Evaluation and Incident Prevention in an Enterprise AI Assistant
by: Maharaj, Akash V., et al.
Published: (2025)
by: Maharaj, Akash V., et al.
Published: (2025)
CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
by: Kim, Tae Soo, et al.
Published: (2025)
by: Kim, Tae Soo, et al.
Published: (2025)
On Evaluating Explanation Utility for Human-AI Decision Making in NLP
by: Chaleshtori, Fateme Hashemi, et al.
Published: (2024)
by: Chaleshtori, Fateme Hashemi, et al.
Published: (2024)
Large Language Models Can Solve Real-World Planning Rigorously with Formal Verification Tools
by: Hao, Yilun, et al.
Published: (2024)
by: Hao, Yilun, et al.
Published: (2024)
Human-AI Interaction Alignment: Designing, Evaluating, and Evolving Value-Centered AI For Reciprocal Human-AI Futures
by: Shen, Hua, et al.
Published: (2025)
by: Shen, Hua, et al.
Published: (2025)
BADGE: BADminton report Generation and Evaluation with LLM
by: Chiang, Shang-Hsuan, et al.
Published: (2024)
by: Chiang, Shang-Hsuan, et al.
Published: (2024)
Evalet: Evaluating Large Language Models through Functional Fragmentation
by: Kim, Tae Soo, et al.
Published: (2025)
by: Kim, Tae Soo, et al.
Published: (2025)
A Risk Ontology for Evaluating AI-Powered Psychotherapy Virtual Agents
by: Steenstra, Ian, et al.
Published: (2025)
by: Steenstra, Ian, et al.
Published: (2025)
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
by: Kim, Tae Soo, et al.
Published: (2023)
by: Kim, Tae Soo, et al.
Published: (2023)
Learning to Generate and Evaluate Fact-checking Explanations with Transformers
by: Feher, Darius, et al.
Published: (2024)
by: Feher, Darius, et al.
Published: (2024)
Just-In-Time Objectives: A General Approach for Specialized AI Interactions
by: Lam, Michelle S., et al.
Published: (2025)
by: Lam, Michelle S., et al.
Published: (2025)
AI-Generated Slides: Are They Good? Can Students Tell?
by: Leinonen, Juho, et al.
Published: (2026)
by: Leinonen, Juho, et al.
Published: (2026)
Can we Debias Social Stereotypes in AI-Generated Images? Examining Text-to-Image Outputs and User Perceptions
by: Barve, Saharsh, et al.
Published: (2025)
by: Barve, Saharsh, et al.
Published: (2025)
Can Good Writing Be Generative? Expert-Level AI Writing Emerges through Fine-Tuning on High-Quality Books
by: Chakrabarty, Tuhin, et al.
Published: (2026)
by: Chakrabarty, Tuhin, et al.
Published: (2026)
Pragmatics Meets Culture: Culturally-adapted Artwork Description Generation and Evaluation
by: Zhao, Lingjun, et al.
Published: (2026)
by: Zhao, Lingjun, et al.
Published: (2026)
One vs. Many: Comprehending Accurate Information from Multiple Erroneous and Inconsistent AI Generations
by: Lee, Yoonjoo, et al.
Published: (2024)
by: Lee, Yoonjoo, et al.
Published: (2024)
Hey GPT, Can You be More Racist? Analysis from Crowdsourced Attempts to Elicit Biased Content from Generative AI
by: Guo, Hangzhi, et al.
Published: (2024)
by: Guo, Hangzhi, et al.
Published: (2024)
Can LLMs Generate Visualizations with Dataless Prompts?
by: Coelho, Darius, et al.
Published: (2024)
by: Coelho, Darius, et al.
Published: (2024)
Using Generative Text Models to Create Qualitative Codebooks for Student Evaluations of Teaching
by: Katz, Andrew, et al.
Published: (2024)
by: Katz, Andrew, et al.
Published: (2024)
VeriLA: A Human-Centered Evaluation Framework for Interpretable Verification of LLM Agent Failures
by: Sung, Yoo Yeon, et al.
Published: (2025)
by: Sung, Yoo Yeon, et al.
Published: (2025)
Human-Centered AI in Multidisciplinary Medical Discussions: Evaluating the Feasibility of a Chat-Based Approach to Case Assessment
by: Sawano, Shinnosuke, et al.
Published: (2025)
by: Sawano, Shinnosuke, et al.
Published: (2025)
SNAP: A Plan-Driven Framework for Controllable Interactive Narrative Generation
by: Bang, Geonwoo, et al.
Published: (2025)
by: Bang, Geonwoo, et al.
Published: (2025)
Flex-TravelPlanner: A Benchmark for Flexible Planning with Language Agents
by: Oh, Juhyun, et al.
Published: (2025)
by: Oh, Juhyun, et al.
Published: (2025)
Interaction Techniques that Encourage Longer Prompts Can Improve Psychological Ownership when Writing with AI
by: Joshi, Nikhita, et al.
Published: (2025)
by: Joshi, Nikhita, et al.
Published: (2025)
ChatBench: From Static Benchmarks to Human-AI Evaluation
by: Chang, Serina, et al.
Published: (2025)
by: Chang, Serina, et al.
Published: (2025)
Mic Drop or Data Flop? Evaluating the Fitness for Purpose of AI Voice Interviewers for Data Collection within Quantitative & Qualitative Research Contexts
by: Tirumala, Shreyas, et al.
Published: (2025)
by: Tirumala, Shreyas, et al.
Published: (2025)
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
by: Gor, Maharshi, et al.
Published: (2026)
by: Gor, Maharshi, et al.
Published: (2026)
When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration
by: Shi, Quan, et al.
Published: (2025)
by: Shi, Quan, et al.
Published: (2025)
Evaluating LLM-Generated Lessons from the Language Learning Students' Perspective: A Short Case Study on Duolingo
by: Catalan, Carlos Rafael, et al.
Published: (2026)
by: Catalan, Carlos Rafael, et al.
Published: (2026)
From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data
by: Kursuncu, Ugur, et al.
Published: (2025)
by: Kursuncu, Ugur, et al.
Published: (2025)
An Evaluation of Estimative Uncertainty in Large Language Models
by: Tang, Zhisheng, et al.
Published: (2024)
by: Tang, Zhisheng, et al.
Published: (2024)
Evaluating the Prompt Steerability of Large Language Models
by: Miehling, Erik, et al.
Published: (2024)
by: Miehling, Erik, et al.
Published: (2024)
Similar Items
-
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
by: Shin, Jisu, et al.
Published: (2025) -
Uncovering Factor Level Preferences to Improve Human-Model Alignment
by: Oh, Juhyun, et al.
Published: (2024) -
Designing and Evaluating Multi-Chatbot Interface for Human-AI Communication: Preliminary Findings from a Persuasion Task
by: Yoon, Sion, et al.
Published: (2024) -
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
by: Shin, Jisu, et al.
Published: (2025) -
Do Ethical AI Principles Matter to Users? A Large-Scale Analysis of User Sentiment and Satisfaction
by: Pasch, Stefan, et al.
Published: (2025)