Facial-Expression-Aware Prompting for Empathetic LLM Tutoring

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Shuangquan, Fleig, Laura, Tu, Ruisen, Chi, Philip, Bu, Edmund, Ozel, Melinda, Ma, Junhua, Fei, Teng, de Sa, Virginia R.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917414327812096
author Feng, Shuangquan
Fleig, Laura
Tu, Ruisen
Chi, Philip
Bu, Edmund
Ozel, Melinda
Ma, Junhua
Fei, Teng
de Sa, Virginia R.
author_facet Feng, Shuangquan
Fleig, Laura
Tu, Ruisen
Chi, Philip
Bu, Edmund
Ozel, Melinda
Ma, Junhua
Fei, Teng
de Sa, Virginia R.
contents Large language models (LLMs) enable increasingly capable tutoring-style conversational agents, yet effective tutoring requires sensitivity to learners' affective and cognitive states beyond text alone. Facial expressions provide immediate and practical cues of confusion, frustration, or engagement, but remain underexplored in LLM-driven tutoring. We investigate whether facial-expression-aware signals can improve empathetic tutoring responses through prompt-level integration, without end-to-end retraining. We build a scalable simulated tutoring environment where a student agent exhibits diverse facial behaviors from a large unlabeled facial expression video dataset, and compare four tutor variants: a text-only LLM baseline, a multimodal baseline using a random facial frame, and two Action Unit estimation model (AUM)-based methods that either inject textual AU descriptions or select a peak-expression frame for visual grounding. Across 960 multi-turn conversations spanning three tutor backbones (GPT-5.1, Claude Ops 4.5, and Gemini 2.5 Pro), we evaluate targeted pairwise comparisons with five human raters and an exhaustive AI evaluator. AU-based conditioning consistently improves empathetic responsiveness to facial expressions across all tutor backbones, while AUM-guided peak-frame selection outperforms random-frame visual input. Textual AU abstraction and peak-frame visual injection show model-dependent advantages. Control analyses show that this improvement does not come at the expense of worse pedagogical clarity or responsiveness to textual cues. Finally, AI-human agreement is highest on facial-expression-grounded empathy, supporting scalable AI evaluation for this dimension. Overall, our results show that lightweight, structured facial expression representations can meaningfully enhance empathy in LLM-based tutoring systems with minimal overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2604_15336
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Facial-Expression-Aware Prompting for Empathetic LLM Tutoring
Feng, Shuangquan
Fleig, Laura
Tu, Ruisen
Chi, Philip
Bu, Edmund
Ozel, Melinda
Ma, Junhua
Fei, Teng
de Sa, Virginia R.
Human-Computer Interaction
Artificial Intelligence
Large language models (LLMs) enable increasingly capable tutoring-style conversational agents, yet effective tutoring requires sensitivity to learners' affective and cognitive states beyond text alone. Facial expressions provide immediate and practical cues of confusion, frustration, or engagement, but remain underexplored in LLM-driven tutoring. We investigate whether facial-expression-aware signals can improve empathetic tutoring responses through prompt-level integration, without end-to-end retraining. We build a scalable simulated tutoring environment where a student agent exhibits diverse facial behaviors from a large unlabeled facial expression video dataset, and compare four tutor variants: a text-only LLM baseline, a multimodal baseline using a random facial frame, and two Action Unit estimation model (AUM)-based methods that either inject textual AU descriptions or select a peak-expression frame for visual grounding. Across 960 multi-turn conversations spanning three tutor backbones (GPT-5.1, Claude Ops 4.5, and Gemini 2.5 Pro), we evaluate targeted pairwise comparisons with five human raters and an exhaustive AI evaluator. AU-based conditioning consistently improves empathetic responsiveness to facial expressions across all tutor backbones, while AUM-guided peak-frame selection outperforms random-frame visual input. Textual AU abstraction and peak-frame visual injection show model-dependent advantages. Control analyses show that this improvement does not come at the expense of worse pedagogical clarity or responsiveness to textual cues. Finally, AI-human agreement is highest on facial-expression-grounded empathy, supporting scalable AI evaluation for this dimension. Overall, our results show that lightweight, structured facial expression representations can meaningfully enhance empathy in LLM-based tutoring systems with minimal overhead.
title Facial-Expression-Aware Prompting for Empathetic LLM Tutoring
topic Human-Computer Interaction
Artificial Intelligence
url https://arxiv.org/abs/2604.15336