ReactGenie: A Development Framework for Complex Multimodal Interactions Using Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Jackie Junrui, Shi, Yingtian, Zhang, Yuhan, Li, Karina, Rosli, Daniel Wan, Jain, Anisha, Zhang, Shuning, Li, Tianshi, Landay, James A., Lam, Monica S.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917655198302208
author Yang, Jackie Junrui
Shi, Yingtian
Zhang, Yuhan
Li, Karina
Rosli, Daniel Wan
Jain, Anisha
Zhang, Shuning
Li, Tianshi
Landay, James A.
Lam, Monica S.
author_facet Yang, Jackie Junrui
Shi, Yingtian
Zhang, Yuhan
Li, Karina
Rosli, Daniel Wan
Jain, Anisha
Zhang, Shuning
Li, Tianshi
Landay, James A.
Lam, Monica S.
contents By combining voice and touch interactions, multimodal interfaces can surpass the efficiency of either modality alone. Traditional multimodal frameworks require laborious developer work to support rich multimodal commands where the user's multimodal command involves possibly exponential combinations of actions/function invocations. This paper presents ReactGenie, a programming framework that better separates multimodal input from the computational model to enable developers to create efficient and capable multimodal interfaces with ease. ReactGenie translates multimodal user commands into NLPL (Natural Language Programming Language), a programming language we created, using a neural semantic parser based on large-language models. The ReactGenie runtime interprets the parsed NLPL and composes primitives in the computational model to implement complex user commands. As a result, ReactGenie allows easy implementation and unprecedented richness in commands for end-users of multimodal apps. Our evaluation showed that 12 developers can learn and build a nontrivial ReactGenie application in under 2.5 hours on average. In addition, compared with a traditional GUI, end-users can complete tasks faster and with less task load using ReactGenie apps.
format Preprint
id arxiv_https___arxiv_org_abs_2306_09649
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle ReactGenie: A Development Framework for Complex Multimodal Interactions Using Large Language Models
Yang, Jackie Junrui
Shi, Yingtian
Zhang, Yuhan
Li, Karina
Rosli, Daniel Wan
Jain, Anisha
Zhang, Shuning
Li, Tianshi
Landay, James A.
Lam, Monica S.
Human-Computer Interaction
Computation and Language
By combining voice and touch interactions, multimodal interfaces can surpass the efficiency of either modality alone. Traditional multimodal frameworks require laborious developer work to support rich multimodal commands where the user's multimodal command involves possibly exponential combinations of actions/function invocations. This paper presents ReactGenie, a programming framework that better separates multimodal input from the computational model to enable developers to create efficient and capable multimodal interfaces with ease. ReactGenie translates multimodal user commands into NLPL (Natural Language Programming Language), a programming language we created, using a neural semantic parser based on large-language models. The ReactGenie runtime interprets the parsed NLPL and composes primitives in the computational model to implement complex user commands. As a result, ReactGenie allows easy implementation and unprecedented richness in commands for end-users of multimodal apps. Our evaluation showed that 12 developers can learn and build a nontrivial ReactGenie application in under 2.5 hours on average. In addition, compared with a traditional GUI, end-users can complete tasks faster and with less task load using ReactGenie apps.
title ReactGenie: A Development Framework for Complex Multimodal Interactions Using Large Language Models
topic Human-Computer Interaction
Computation and Language
url https://arxiv.org/abs/2306.09649