ACE, Action and Control via Explanations: A Proposal for LLMs to Provide Human-Centered Explainability for Multimodal AI Assistants

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Watkins, Elizabeth Anne, Moss, Emanuel, Manuvinakurike, Ramesh, Shi, Meng, Beckwith, Richard, Raffa, Giuseppe
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917963424071680
author Watkins, Elizabeth Anne
Moss, Emanuel
Manuvinakurike, Ramesh
Shi, Meng
Beckwith, Richard
Raffa, Giuseppe
author_facet Watkins, Elizabeth Anne
Moss, Emanuel
Manuvinakurike, Ramesh
Shi, Meng
Beckwith, Richard
Raffa, Giuseppe
contents In this short paper we address issues related to building multimodal AI systems for human performance support in manufacturing domains. We make two contributions: we first identify challenges of participatory design and training of such systems, and secondly, to address such challenges, we propose the ACE paradigm: "Action and Control via Explanations". Specifically, we suggest that LLMs can be used to produce explanations in the form of human interpretable "semantic frames", which in turn enable end users to provide data the AI system needs to align its multimodal models and representations, including computer vision, automatic speech recognition, and document inputs. ACE, by using LLMs to "explain" using semantic frames, will help the human and the AI system to collaborate, together building a more accurate model of humans activities and behaviors, and ultimately more accurate predictive outputs for better task support, and better outcomes for human users performing manual tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_16466
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ACE, Action and Control via Explanations: A Proposal for LLMs to Provide Human-Centered Explainability for Multimodal AI Assistants
Watkins, Elizabeth Anne
Moss, Emanuel
Manuvinakurike, Ramesh
Shi, Meng
Beckwith, Richard
Raffa, Giuseppe
Human-Computer Interaction
Artificial Intelligence
In this short paper we address issues related to building multimodal AI systems for human performance support in manufacturing domains. We make two contributions: we first identify challenges of participatory design and training of such systems, and secondly, to address such challenges, we propose the ACE paradigm: "Action and Control via Explanations". Specifically, we suggest that LLMs can be used to produce explanations in the form of human interpretable "semantic frames", which in turn enable end users to provide data the AI system needs to align its multimodal models and representations, including computer vision, automatic speech recognition, and document inputs. ACE, by using LLMs to "explain" using semantic frames, will help the human and the AI system to collaborate, together building a more accurate model of humans activities and behaviors, and ultimately more accurate predictive outputs for better task support, and better outcomes for human users performing manual tasks.
title ACE, Action and Control via Explanations: A Proposal for LLMs to Provide Human-Centered Explainability for Multimodal AI Assistants
topic Human-Computer Interaction
Artificial Intelligence
url https://arxiv.org/abs/2503.16466