Facilitating Holistic Evaluations with LLMs: Insights from Scenario-Based Experiments
Fuente:
arXiv
Saved in:
| Main Authors: | Ishida, Toru, Liu, Tongxi, Wang, Hailong, Cheunga, William K. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large Language Models as Partners in Student Essay Evaluation
by: Ishida, Toru, et al.
Published: (2024)
by: Ishida, Toru, et al.
Published: (2024)
Toward AI Systems That Understand Self and Others: A Multi-Phase Inference Framework for Human Cognitive Diversity and World-Model Alignment
by: Takahashi, Toru
Published: (2026)
by: Takahashi, Toru
Published: (2026)
Towards Apples to Apples for AI Evaluations: From Real-World Use Cases to Evaluation Scenarios
by: Choong, Yee-Yin, et al.
Published: (2026)
by: Choong, Yee-Yin, et al.
Published: (2026)
Facilitating Longitudinal Interaction Studies of AI Systems
by: Long, Tao, et al.
Published: (2025)
by: Long, Tao, et al.
Published: (2025)
When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition
by: Li, Karen Jia-Hui, et al.
Published: (2025)
by: Li, Karen Jia-Hui, et al.
Published: (2025)
BLIP: Facilitating the Exploration of Undesirable Consequences of Digital Technologies
by: Pang, Rock Yuren, et al.
Published: (2024)
by: Pang, Rock Yuren, et al.
Published: (2024)
A Comparative Study of Technical Writing Feedback Quality: Evaluating LLMs, SLMs, and Humans in Computer Science Topics
by: Liu, Suqing, et al.
Published: (2025)
by: Liu, Suqing, et al.
Published: (2025)
Public Discourse Sandbox: Facilitating Human and AI Digital Communication Research
by: Radivojevic, Kristina, et al.
Published: (2025)
by: Radivojevic, Kristina, et al.
Published: (2025)
Developer Insights into Designing AI-Based Computer Perception Tools
by: Guhan, Maya, et al.
Published: (2025)
by: Guhan, Maya, et al.
Published: (2025)
From Divergence to Consensus: Evaluating the Role of Large Language Models in Facilitating Agreement through Adaptive Strategies
by: Triantafyllopoulos, Loukas, et al.
Published: (2025)
by: Triantafyllopoulos, Loukas, et al.
Published: (2025)
Harnessing LLMs for Automated Video Content Analysis: An Exploratory Workflow of Short Videos on Depression
by: Liu, Jiaying Lizzy, et al.
Published: (2024)
by: Liu, Jiaying Lizzy, et al.
Published: (2024)
AgentSUMO: An Agentic Framework for Interactive Simulation Scenario Generation in SUMO via Large Language Models
by: Jeong, Minwoo, et al.
Published: (2025)
by: Jeong, Minwoo, et al.
Published: (2025)
Facilitating the Integration of LLMs Into Online Experiments With Simple Chat
by: Schettino, R. Bermudez, et al.
Published: (2025)
by: Schettino, R. Bermudez, et al.
Published: (2025)
Wearable Device-Based Real-Time Monitoring of Physiological Signals: Evaluating Cognitive Load Across Different Tasks
by: He, Ling, et al.
Published: (2024)
by: He, Ling, et al.
Published: (2024)
Epitome: Pioneering an Experimental Platform for AI-Social Science Integration
by: Qu, Jingjing, et al.
Published: (2025)
by: Qu, Jingjing, et al.
Published: (2025)
G-VOILA: Gaze-Facilitated Information Querying in Daily Scenarios
by: Wang, Zeyu, et al.
Published: (2024)
by: Wang, Zeyu, et al.
Published: (2024)
Incorporating AI incident reporting into telecommunications law and policy: Insights from India
by: Agarwal, Avinash, et al.
Published: (2025)
by: Agarwal, Avinash, et al.
Published: (2025)
WaLLM -- Insights from an LLM-Powered Chatbot deployment via WhatsApp
by: Eltigani, Hiba, et al.
Published: (2025)
by: Eltigani, Hiba, et al.
Published: (2025)
MediTools -- Medical Education Powered by LLMs
by: Alshatnawi, Amr, et al.
Published: (2025)
by: Alshatnawi, Amr, et al.
Published: (2025)
Insights from Social Shaping Theory: The Appropriation of Large Language Models in an Undergraduate Programming Course
by: Padiyath, Aadarsh, et al.
Published: (2024)
by: Padiyath, Aadarsh, et al.
Published: (2024)
How Persuasive Could LLMs Be? A First Study Combining Linguistic-Rhetorical Analysis and User Experiments
by: Raffini, Daniel, et al.
Published: (2025)
by: Raffini, Daniel, et al.
Published: (2025)
The Case for "Thick Evaluations" of Cultural Representation in AI
by: Qadri, Rida, et al.
Published: (2025)
by: Qadri, Rida, et al.
Published: (2025)
A Governance and Evaluation Framework for Deterministic, Rule-Based Clinical Decision Support in Empiric Antibiotic Prescribing
by: Gárate, Francisco José, et al.
Published: (2026)
by: Gárate, Francisco José, et al.
Published: (2026)
Hope, Aspirations, and the Impact of LLMs on Female Programming Learners in Afghanistan
by: Behmanush, Hamayoon, et al.
Published: (2025)
by: Behmanush, Hamayoon, et al.
Published: (2025)
Between Myths and Metaphors: Rethinking LLMs for SRH in Conservative Contexts
by: Humayun, Ameemah, et al.
Published: (2025)
by: Humayun, Ameemah, et al.
Published: (2025)
SparkMe: Adaptive Semi-Structured Interviewing for Qualitative Insight Discovery
by: Anugraha, David, et al.
Published: (2026)
by: Anugraha, David, et al.
Published: (2026)
Red Teaming LLMs as Socio-Technical Practice: From Exploration and Data Creation to Evaluation
by: Garcia, Adriana Alvarado, et al.
Published: (2026)
by: Garcia, Adriana Alvarado, et al.
Published: (2026)
"Unlimited Realm of Exploration and Experimentation": Methods and Motivations of AI-Generated Sexual Content Creators
by: Mink, Jaron, et al.
Published: (2026)
by: Mink, Jaron, et al.
Published: (2026)
From Lived Experience to Insight: Unpacking the Psychological Risks of Using AI Conversational Agents
by: Chandra, Mohit, et al.
Published: (2024)
by: Chandra, Mohit, et al.
Published: (2024)
Disability Across Cultures: A Human-Centered Audit of Ableism in Western and Indic LLMs
by: Phutane, Mahika, et al.
Published: (2025)
by: Phutane, Mahika, et al.
Published: (2025)
A Conditional Companion: Lived Experiences of People with Mental Health Disorders Using LLMs
by: Purohit, Aditya Kumar, et al.
Published: (2026)
by: Purohit, Aditya Kumar, et al.
Published: (2026)
ROBOPSY PL[AI]: Using Role-Play to Investigate how LLMs Present Collective Memory
by: Jahrmann, Margarete, et al.
Published: (2025)
by: Jahrmann, Margarete, et al.
Published: (2025)
Artificial Intelligence as a Training Tool in Clinical Psychology: A Comparison of Text-Based and Avatar Simulations
by: Sawah, V. El, et al.
Published: (2025)
by: Sawah, V. El, et al.
Published: (2025)
Evaluating Contextually Personalized Programming Exercises Created with Generative AI
by: Logacheva, Evanfiya, et al.
Published: (2024)
by: Logacheva, Evanfiya, et al.
Published: (2024)
Evaluating Alternative Training Interventions Using Personalized Computational Models of Learning
by: MacLellan, Christopher James, et al.
Published: (2024)
by: MacLellan, Christopher James, et al.
Published: (2024)
Recommendations and Reporting Checklist for Rigorous & Transparent Human Baselines in Model Evaluations
by: Wei, Kevin L., et al.
Published: (2025)
by: Wei, Kevin L., et al.
Published: (2025)
When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis
by: Najera, Aisha, et al.
Published: (2026)
by: Najera, Aisha, et al.
Published: (2026)
Evaluating Generative AI as an Educational Tool for Radiology Resident Report Drafting
by: Verdone, Antonio, et al.
Published: (2025)
by: Verdone, Antonio, et al.
Published: (2025)
Patterns vs. Patients: Evaluating LLMs against Mental Health Professionals on Personality Disorder Diagnosis through First-Person Narratives
by: Drożdż, Karolina, et al.
Published: (2025)
by: Drożdż, Karolina, et al.
Published: (2025)
How Professional Visual Artists are Negotiating Generative AI in the Workplace
by: Jiang, Harry H., et al.
Published: (2026)
by: Jiang, Harry H., et al.
Published: (2026)
Similar Items
-
Large Language Models as Partners in Student Essay Evaluation
by: Ishida, Toru, et al.
Published: (2024) -
Toward AI Systems That Understand Self and Others: A Multi-Phase Inference Framework for Human Cognitive Diversity and World-Model Alignment
by: Takahashi, Toru
Published: (2026) -
Towards Apples to Apples for AI Evaluations: From Real-World Use Cases to Evaluation Scenarios
by: Choong, Yee-Yin, et al.
Published: (2026) -
Facilitating Longitudinal Interaction Studies of AI Systems
by: Long, Tao, et al.
Published: (2025) -
When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition
by: Li, Karen Jia-Hui, et al.
Published: (2025)