Context-Aware Human Behavior Prediction Using Multimodal Large Language Models: Challenges and Insights

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Yuchen, Lerch, Lino, Palmieri, Luigi, Rudenko, Andrey, Koch, Sebastian, Ropinski, Timo, Aiello, Marco
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909657269796864
author Liu, Yuchen
Lerch, Lino
Palmieri, Luigi
Rudenko, Andrey
Koch, Sebastian
Ropinski, Timo
Aiello, Marco
author_facet Liu, Yuchen
Lerch, Lino
Palmieri, Luigi
Rudenko, Andrey
Koch, Sebastian
Ropinski, Timo
Aiello, Marco
contents Predicting human behavior in shared environments is crucial for safe and efficient human-robot interaction. Traditional data-driven methods to that end are pre-trained on domain-specific datasets, activity types, and prediction horizons. In contrast, the recent breakthroughs in Large Language Models (LLMs) promise open-ended cross-domain generalization to describe various human activities and make predictions in any context. In particular, Multimodal LLMs (MLLMs) are able to integrate information from various sources, achieving more contextual awareness and improved scene understanding. The difficulty in applying general-purpose MLLMs directly for prediction stems from their limited capacity for processing large input sequences, sensitivity to prompt design, and expensive fine-tuning. In this paper, we present a systematic analysis of applying pre-trained MLLMs for context-aware human behavior prediction. To this end, we introduce a modular multimodal human activity prediction framework that allows us to benchmark various MLLMs, input variations, In-Context Learning (ICL), and autoregressive techniques. Our evaluation indicates that the best-performing framework configuration is able to reach 92.8% semantic similarity and 66.1% exact label accuracy in predicting human behaviors in the target frame.
format Preprint
id arxiv_https___arxiv_org_abs_2504_00839
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Context-Aware Human Behavior Prediction Using Multimodal Large Language Models: Challenges and Insights
Liu, Yuchen
Lerch, Lino
Palmieri, Luigi
Rudenko, Andrey
Koch, Sebastian
Ropinski, Timo
Aiello, Marco
Robotics
Artificial Intelligence
Predicting human behavior in shared environments is crucial for safe and efficient human-robot interaction. Traditional data-driven methods to that end are pre-trained on domain-specific datasets, activity types, and prediction horizons. In contrast, the recent breakthroughs in Large Language Models (LLMs) promise open-ended cross-domain generalization to describe various human activities and make predictions in any context. In particular, Multimodal LLMs (MLLMs) are able to integrate information from various sources, achieving more contextual awareness and improved scene understanding. The difficulty in applying general-purpose MLLMs directly for prediction stems from their limited capacity for processing large input sequences, sensitivity to prompt design, and expensive fine-tuning. In this paper, we present a systematic analysis of applying pre-trained MLLMs for context-aware human behavior prediction. To this end, we introduce a modular multimodal human activity prediction framework that allows us to benchmark various MLLMs, input variations, In-Context Learning (ICL), and autoregressive techniques. Our evaluation indicates that the best-performing framework configuration is able to reach 92.8% semantic similarity and 66.1% exact label accuracy in predicting human behaviors in the target frame.
title Context-Aware Human Behavior Prediction Using Multimodal Large Language Models: Challenges and Insights
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2504.00839