CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Tae Soo, Lee, Yoonjoo, Park, Yoonah, Kim, Jiho, Kim, Young-Ho, Kim, Juho |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
by: Kim, Tae Soo, et al.
Published: (2023)
by: Kim, Tae Soo, et al.
Published: (2023)
Evalet: Evaluating Large Language Models through Functional Fragmentation
by: Kim, Tae Soo, et al.
Published: (2025)
by: Kim, Tae Soo, et al.
Published: (2025)
One vs. Many: Comprehending Accurate Information from Multiple Erroneous and Inconsistent AI Generations
by: Lee, Yoonjoo, et al.
Published: (2024)
by: Lee, Yoonjoo, et al.
Published: (2024)
DiscoverLLM: From Executing Intents to Discovering Them
by: Kim, Tae Soo, et al.
Published: (2026)
by: Kim, Tae Soo, et al.
Published: (2026)
"When to Hand Off, When to Work Together": Expanding Human-Agent Co-Creative Collaboration through Concurrent Interaction
by: Son, Kihoon, et al.
Published: (2026)
by: Son, Kihoon, et al.
Published: (2026)
PaperWeaver: Enriching Topical Paper Alerts by Contextualizing Recommended Papers with User-collected Papers
by: Lee, Yoonjoo, et al.
Published: (2024)
by: Lee, Yoonjoo, et al.
Published: (2024)
ClearFairy: Capturing Creative Workflows through Decision Structuring, In-Situ Questioning, and Rationale Inference
by: Son, Kihoon, et al.
Published: (2025)
by: Son, Kihoon, et al.
Published: (2025)
Designing and Evaluating Multi-Chatbot Interface for Human-AI Communication: Preliminary Findings from a Persuasion Task
by: Yoon, Sion, et al.
Published: (2024)
by: Yoon, Sion, et al.
Published: (2024)
Beyond Prompts: Learning from Human Communication for Enhanced AI Intent Alignment
by: Kim, Yoonsu, et al.
Published: (2024)
by: Kim, Yoonsu, et al.
Published: (2024)
Aligning LLMs with Individual Preferences via Interaction
by: Wu, Shujin, et al.
Published: (2024)
by: Wu, Shujin, et al.
Published: (2024)
VIVID: Human-AI Collaborative Authoring of Vicarious Dialogues from Lecture Videos
by: Choi, Seulgi, et al.
Published: (2024)
by: Choi, Seulgi, et al.
Published: (2024)
Iffy-Or-Not: Extending the Web to Support the Critical Evaluation of Fallacious Texts
by: Lim, Gionnieve, et al.
Published: (2025)
by: Lim, Gionnieve, et al.
Published: (2025)
GenQuery: Supporting Expressive Visual Search with Generative Models
by: Son, Kihoon, et al.
Published: (2023)
by: Son, Kihoon, et al.
Published: (2023)
AutiHero: Engaging Parents in Creating Personalized, Multi-path Social Narratives for Autistic Children
by: Lee, Jungeun, et al.
Published: (2025)
by: Lee, Jungeun, et al.
Published: (2025)
ValueCompass: A Framework for Measuring Contextual Value Alignment Between Human and LLMs
by: Shen, Hua, et al.
Published: (2024)
by: Shen, Hua, et al.
Published: (2024)
The Generative AI Paradox on Evaluation: What It Can Solve, It May Not Evaluate
by: Oh, Juhyun, et al.
Published: (2024)
by: Oh, Juhyun, et al.
Published: (2024)
Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech
by: Kim, Taesoo, et al.
Published: (2025)
by: Kim, Taesoo, et al.
Published: (2025)
AVIN-Chat: An Audio-Visual Interactive Chatbot System with Emotional State Tuning
by: Park, Chanhyuk, et al.
Published: (2024)
by: Park, Chanhyuk, et al.
Published: (2024)
PlanFitting: Personalized Exercise Planning with Large Language Model-driven Conversational Agent
by: Shin, Donghoon, et al.
Published: (2023)
by: Shin, Donghoon, et al.
Published: (2023)
VeriLA: A Human-Centered Evaluation Framework for Interpretable Verification of LLM Agent Failures
by: Sung, Yoo Yeon, et al.
Published: (2025)
by: Sung, Yoo Yeon, et al.
Published: (2025)
ChaCha: Leveraging Large Language Models to Prompt Children to Share Their Emotions about Personal Events
by: Seo, Woosuk, et al.
Published: (2023)
by: Seo, Woosuk, et al.
Published: (2023)
Evaluating Contextually Personalized Programming Exercises Created with Generative AI
by: Logacheva, Evanfiya, et al.
Published: (2024)
by: Logacheva, Evanfiya, et al.
Published: (2024)
ChoiceMates: Supporting Unfamiliar Online Decision-Making with Multi-Agent Conversational Interactions
by: Park, Jeongeon, et al.
Published: (2023)
by: Park, Jeongeon, et al.
Published: (2023)
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
by: Kim, Jane Paik
Published: (2026)
by: Kim, Jane Paik
Published: (2026)
Benchmarking Large Language Models for Diagnosing Students' Cognitive Skills from Handwritten Math Work
by: Kim, Yoonsu, et al.
Published: (2025)
by: Kim, Yoonsu, et al.
Published: (2025)
Multimodal Transformer Models for Turn-taking Prediction: Effects on Conversational Dynamics of Human-Agent Interaction during Cooperative Gameplay
by: Bae, Young-Ho, et al.
Published: (2025)
by: Bae, Young-Ho, et al.
Published: (2025)
Human-AI Collaborative Taxonomy Construction: A Case Study in Profession-Specific Writing Assistants
by: Lee, Minhwa, et al.
Published: (2024)
by: Lee, Minhwa, et al.
Published: (2024)
ELMI: Interactive and Intelligent Sign Language Translation of Lyrics for Song Signing
by: Yoo, Suhyeon, et al.
Published: (2024)
by: Yoo, Suhyeon, et al.
Published: (2024)
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
Emergent Cognitive Convergence via Implementation: Structured Cognitive Loop Reflecting Four Theories of Mind
by: Kim, Myung Ho
Published: (2025)
by: Kim, Myung Ho
Published: (2025)
Vision-Integrated LLMs for Autonomous Driving Assistance : Human Performance Comparison and Trust Evaluation
by: Kim, Namhee, et al.
Published: (2025)
by: Kim, Namhee, et al.
Published: (2025)
People will agree what I think: Investigating LLM's False Consensus Effect
by: Choi, Junhyuk, et al.
Published: (2024)
by: Choi, Junhyuk, et al.
Published: (2024)
The Adoption and Efficacy of Large Language Models: Evidence From Consumer Complaints in the Financial Industry
by: Shin, Minkyu, et al.
Published: (2023)
by: Shin, Minkyu, et al.
Published: (2023)
CulturalTeaming: AI-Assisted Interactive Red-Teaming for Challenging LLMs' (Lack of) Multicultural Knowledge
by: Chiu, Yu Ying, et al.
Published: (2024)
by: Chiu, Yu Ying, et al.
Published: (2024)
Unveiling Disparities in Web Task Handling Between Human and Web Agent
by: Son, Kihoon, et al.
Published: (2024)
by: Son, Kihoon, et al.
Published: (2024)
Leveraging Large Language Models to Power Chatbots for Collecting User Self-Reported Data
by: Wei, Jing, et al.
Published: (2023)
by: Wei, Jing, et al.
Published: (2023)
Rank-O-ToM: Unlocking Emotional Nuance Ranking to Enhance Affective Theory-of-Mind
by: Kim, JiHyun, et al.
Published: (2025)
by: Kim, JiHyun, et al.
Published: (2025)
Demystifying Tacit Knowledge in Graphic Design: Characteristics, Instances, Approaches, and Guidelines
by: Son, Kihoon, et al.
Published: (2024)
by: Son, Kihoon, et al.
Published: (2024)
Alignment Drift in Multimodal LLMs: A Two-Phase, Longitudinal Evaluation of Harm Across Eight Model Releases
by: Ford, Casey, et al.
Published: (2026)
by: Ford, Casey, et al.
Published: (2026)
Death of the Novel(ty): Beyond n-Gram Novelty as a Metric for Textual Creativity
by: Saakyan, Arkadiy, et al.
Published: (2025)
by: Saakyan, Arkadiy, et al.
Published: (2025)
Similar Items
-
EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined Criteria
by: Kim, Tae Soo, et al.
Published: (2023) -
Evalet: Evaluating Large Language Models through Functional Fragmentation
by: Kim, Tae Soo, et al.
Published: (2025) -
One vs. Many: Comprehending Accurate Information from Multiple Erroneous and Inconsistent AI Generations
by: Lee, Yoonjoo, et al.
Published: (2024) -
DiscoverLLM: From Executing Intents to Discovering Them
by: Kim, Tae Soo, et al.
Published: (2026) -
"When to Hand Off, When to Work Together": Expanding Human-Agent Co-Creative Collaboration through Concurrent Interaction
by: Son, Kihoon, et al.
Published: (2026)