EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Tae Soo, Lee, Yoonjoo, Shin, Jamin, Kim, Young-Ho, Kim, Juho |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
by: Kim, Tae Soo, et al.
Published: (2025)
by: Kim, Tae Soo, et al.
Published: (2025)
Evalet: Evaluating Large Language Models through Functional Fragmentation
by: Kim, Tae Soo, et al.
Published: (2025)
by: Kim, Tae Soo, et al.
Published: (2025)
DiscoverLLM: From Executing Intents to Discovering Them
by: Kim, Tae Soo, et al.
Published: (2026)
by: Kim, Tae Soo, et al.
Published: (2026)
One vs. Many: Comprehending Accurate Information from Multiple Erroneous and Inconsistent AI Generations
by: Lee, Yoonjoo, et al.
Published: (2024)
by: Lee, Yoonjoo, et al.
Published: (2024)
Beyond Prompts: Learning from Human Communication for Enhanced AI Intent Alignment
by: Kim, Yoonsu, et al.
Published: (2024)
by: Kim, Yoonsu, et al.
Published: (2024)
VIVID: Human-AI Collaborative Authoring of Vicarious Dialogues from Lecture Videos
by: Choi, Seulgi, et al.
Published: (2024)
by: Choi, Seulgi, et al.
Published: (2024)
GenQuery: Supporting Expressive Visual Search with Generative Models
by: Son, Kihoon, et al.
Published: (2023)
by: Son, Kihoon, et al.
Published: (2023)
"When to Hand Off, When to Work Together": Expanding Human-Agent Co-Creative Collaboration through Concurrent Interaction
by: Son, Kihoon, et al.
Published: (2026)
by: Son, Kihoon, et al.
Published: (2026)
PaperWeaver: Enriching Topical Paper Alerts by Contextualizing Recommended Papers with User-collected Papers
by: Lee, Yoonjoo, et al.
Published: (2024)
by: Lee, Yoonjoo, et al.
Published: (2024)
Unveiling Disparities in Web Task Handling Between Human and Web Agent
by: Son, Kihoon, et al.
Published: (2024)
by: Son, Kihoon, et al.
Published: (2024)
Demystifying Tacit Knowledge in Graphic Design: Characteristics, Instances, Approaches, and Guidelines
by: Son, Kihoon, et al.
Published: (2024)
by: Son, Kihoon, et al.
Published: (2024)
Teach AI How to Code: Using Large Language Models as Teachable Agents for Programming Education
by: Jin, Hyoungwook, et al.
Published: (2023)
by: Jin, Hyoungwook, et al.
Published: (2023)
ClearFairy: Capturing Creative Workflows through Decision Structuring, In-Situ Questioning, and Rationale Inference
by: Son, Kihoon, et al.
Published: (2025)
by: Son, Kihoon, et al.
Published: (2025)
Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards
by: Jung, Minji, et al.
Published: (2026)
by: Jung, Minji, et al.
Published: (2026)
The Adoption and Efficacy of Large Language Models: Evidence From Consumer Complaints in the Financial Industry
by: Shin, Minkyu, et al.
Published: (2023)
by: Shin, Minkyu, et al.
Published: (2023)
Leveraging Large Language Models to Power Chatbots for Collecting User Self-Reported Data
by: Wei, Jing, et al.
Published: (2023)
by: Wei, Jing, et al.
Published: (2023)
Understanding Users' Dissatisfaction with ChatGPT Responses: Types, Resolving Tactics, and the Effect of Knowledge Level
by: Kim, Yoonsu, et al.
Published: (2023)
by: Kim, Yoonsu, et al.
Published: (2023)
MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models
by: Lee, Suhyun, et al.
Published: (2026)
by: Lee, Suhyun, et al.
Published: (2026)
PlanFitting: Personalized Exercise Planning with Large Language Model-driven Conversational Agent
by: Shin, Donghoon, et al.
Published: (2023)
by: Shin, Donghoon, et al.
Published: (2023)
BloomIntent: Automating Search Evaluation with LLM-Generated Fine-Grained User Intents
by: Choi, Yoonseo, et al.
Published: (2025)
by: Choi, Yoonseo, et al.
Published: (2025)
ChaCha: Leveraging Large Language Models to Prompt Children to Share Their Emotions about Personal Events
by: Seo, Woosuk, et al.
Published: (2023)
by: Seo, Woosuk, et al.
Published: (2023)
Designing Prompt Analytics Dashboards to Analyze Student-ChatGPT Interactions in EFL Writing
by: Kim, Minsun, et al.
Published: (2024)
by: Kim, Minsun, et al.
Published: (2024)
VisEval: A Benchmark for Data Visualization in the Era of Large Language Models
by: Chen, Nan, et al.
Published: (2024)
by: Chen, Nan, et al.
Published: (2024)
Applying the Gricean Maxims to a Human-LLM Interaction Cycle: Design Insights from a Participatory Approach
by: Kim, Yoonsu, et al.
Published: (2025)
by: Kim, Yoonsu, et al.
Published: (2025)
Using LLMs to Investigate Correlations of Conversational Follow-up Queries with User Satisfaction
by: Kim, Hyunwoo, et al.
Published: (2024)
by: Kim, Hyunwoo, et al.
Published: (2024)
Designing and Evaluating Multi-Chatbot Interface for Human-AI Communication: Preliminary Findings from a Persuasion Task
by: Yoon, Sion, et al.
Published: (2024)
by: Yoon, Sion, et al.
Published: (2024)
Natural Language Dataset Generation Framework for Visualizations Powered by Large Language Models
by: Ko, Hyung-Kwon, et al.
Published: (2023)
by: Ko, Hyung-Kwon, et al.
Published: (2023)
Evaluating the Prompt Steerability of Large Language Models
by: Miehling, Erik, et al.
Published: (2024)
by: Miehling, Erik, et al.
Published: (2024)
Do LLMs Truly Benefit from Longer Context in Automatic Post-Editing?
by: Kim, Ahrii, et al.
Published: (2026)
by: Kim, Ahrii, et al.
Published: (2026)
LalaEval: A Holistic Human Evaluation Framework for Domain-Specific Large Language Models
by: Sun, Chongyan, et al.
Published: (2024)
by: Sun, Chongyan, et al.
Published: (2024)
Textoshop: Interactions Inspired by Drawing Software to Facilitate Text Editing
by: Masson, Damien, et al.
Published: (2024)
by: Masson, Damien, et al.
Published: (2024)
Beyond Turn-taking: Introducing Text-based Overlap into Human-LLM Interactions
by: Kim, JiWoo, et al.
Published: (2025)
by: Kim, JiWoo, et al.
Published: (2025)
Revealing User Familiarity Bias in Task-Oriented Dialogue via Interactive Evaluation
by: Kim, Takyoung, et al.
Published: (2023)
by: Kim, Takyoung, et al.
Published: (2023)
Chart-to-Experience: Benchmarking Multimodal LLMs for Predicting Experiential Impact of Charts
by: Kim, Seon Gyeom, et al.
Published: (2025)
by: Kim, Seon Gyeom, et al.
Published: (2025)
ELMI: Interactive and Intelligent Sign Language Translation of Lyrics for Song Signing
by: Yoo, Suhyeon, et al.
Published: (2024)
by: Yoo, Suhyeon, et al.
Published: (2024)
MindfulDiary: Harnessing Large Language Model to Support Psychiatric Patients' Journaling
by: Kim, Taewan, et al.
Published: (2023)
by: Kim, Taewan, et al.
Published: (2023)
ModSandbox: Facilitating Online Community Moderation Through Error Prediction and Improvement of Automated Rules
by: Song, Jean Y., et al.
Published: (2022)
by: Song, Jean Y., et al.
Published: (2022)
Evaluation of a Sign Language Avatar on Comprehensibility, User Experience \& Acceptability
by: Wasserroth, Fenya, et al.
Published: (2025)
by: Wasserroth, Fenya, et al.
Published: (2025)
Explainable XR: Understanding User Behaviors of XR Environments using LLM-assisted Analytics Framework
by: Kim, Yoonsang, et al.
Published: (2025)
by: Kim, Yoonsang, et al.
Published: (2025)
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
Similar Items
-
CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
by: Kim, Tae Soo, et al.
Published: (2025) -
Evalet: Evaluating Large Language Models through Functional Fragmentation
by: Kim, Tae Soo, et al.
Published: (2025) -
DiscoverLLM: From Executing Intents to Discovering Them
by: Kim, Tae Soo, et al.
Published: (2026) -
One vs. Many: Comprehending Accurate Information from Multiple Erroneous and Inconsistent AI Generations
by: Lee, Yoonjoo, et al.
Published: (2024) -
Beyond Prompts: Learning from Human Communication for Enhanced AI Intent Alignment
by: Kim, Yoonsu, et al.
Published: (2024)