Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: González-González, Manuela, Belharbi, Soufiane, Zeeshan, Muhammad Osama, Sharafi, Masoumeh, Aslam, Muhammad Haseeb, Sia, Lorenzo, Richet, Nicolas, Pedersoli, Marco, Koerich, Alessandro Lameiras, Bacon, Simon L, Granger, Eric
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915977831120896
author González-González, Manuela
Belharbi, Soufiane
Zeeshan, Muhammad Osama
Sharafi, Masoumeh
Aslam, Muhammad Haseeb
Sia, Lorenzo
Richet, Nicolas
Pedersoli, Marco
Koerich, Alessandro Lameiras
Bacon, Simon L
Granger, Eric
author_facet González-González, Manuela
Belharbi, Soufiane
Zeeshan, Muhammad Osama
Sharafi, Masoumeh
Aslam, Muhammad Haseeb
Sia, Lorenzo
Richet, Nicolas
Pedersoli, Marco
Koerich, Alessandro Lameiras
Bacon, Simon L
Granger, Eric
contents Using behavioural science, health interventions focus on behaviour change by providing a framework to help patients acquire and maintain healthy habits that improve medical outcomes. In-person interventions are costly and difficult to scale, especially in resource-limited regions. Digital health interventions offer a cost-effective approach, potentially supporting independent living and self-management. Automating such interventions, especially through machine learning, has gained considerable attention recently. Ambivalence and hesitancy (A/H) play a primary role for individuals to delay, avoid, or abandon health interventions. A/H are subtle and conflicting emotions that place a person in a state between positive and negative evaluations of a behaviour, or between acceptance and refusal to engage in it. They manifest as affective inconsistency across modalities or within a modality, such as language, facial, vocal expressions, and body language. While experts can be trained to recognize A/H, integrating them into digital health interventions is costly and less effective. Automatic A/H recognition is therefore critical for the personalization and cost-effectiveness of digital health interventions. Here, we explore the application of deep learning models for A/H recognition in videos, a multi-modal task by nature. In particular, this paper covers three learning setups: supervised learning, unsupervised domain adaptation for personalization, and zero-shot inference via large language models (LLMs). Our experiments are conducted on the unique and recently published BAH video dataset for A/H recognition. Our results show limited performance, suggesting that more adapted multi-modal models are required for accurate A/H recognition. Better methods for modeling spatio-temporal and multimodal fusion are necessary to leverage conflicts within/across modalities.
format Preprint
id arxiv_https___arxiv_org_abs_2604_11730
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions
González-González, Manuela
Belharbi, Soufiane
Zeeshan, Muhammad Osama
Sharafi, Masoumeh
Aslam, Muhammad Haseeb
Sia, Lorenzo
Richet, Nicolas
Pedersoli, Marco
Koerich, Alessandro Lameiras
Bacon, Simon L
Granger, Eric
Computer Vision and Pattern Recognition
Human-Computer Interaction
Machine Learning
Using behavioural science, health interventions focus on behaviour change by providing a framework to help patients acquire and maintain healthy habits that improve medical outcomes. In-person interventions are costly and difficult to scale, especially in resource-limited regions. Digital health interventions offer a cost-effective approach, potentially supporting independent living and self-management. Automating such interventions, especially through machine learning, has gained considerable attention recently. Ambivalence and hesitancy (A/H) play a primary role for individuals to delay, avoid, or abandon health interventions. A/H are subtle and conflicting emotions that place a person in a state between positive and negative evaluations of a behaviour, or between acceptance and refusal to engage in it. They manifest as affective inconsistency across modalities or within a modality, such as language, facial, vocal expressions, and body language. While experts can be trained to recognize A/H, integrating them into digital health interventions is costly and less effective. Automatic A/H recognition is therefore critical for the personalization and cost-effectiveness of digital health interventions. Here, we explore the application of deep learning models for A/H recognition in videos, a multi-modal task by nature. In particular, this paper covers three learning setups: supervised learning, unsupervised domain adaptation for personalization, and zero-shot inference via large language models (LLMs). Our experiments are conducted on the unique and recently published BAH video dataset for A/H recognition. Our results show limited performance, suggesting that more adapted multi-modal models are required for accurate A/H recognition. Better methods for modeling spatio-temporal and multimodal fusion are necessary to leverage conflicts within/across modalities.
title Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions
topic Computer Vision and Pattern Recognition
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2604.11730