Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915977831120896 |
|---|---|
| author | González-González, Manuela Belharbi, Soufiane Zeeshan, Muhammad Osama Sharafi, Masoumeh Aslam, Muhammad Haseeb Sia, Lorenzo Richet, Nicolas Pedersoli, Marco Koerich, Alessandro Lameiras Bacon, Simon L Granger, Eric |
| author_facet | González-González, Manuela Belharbi, Soufiane Zeeshan, Muhammad Osama Sharafi, Masoumeh Aslam, Muhammad Haseeb Sia, Lorenzo Richet, Nicolas Pedersoli, Marco Koerich, Alessandro Lameiras Bacon, Simon L Granger, Eric |
| contents | Using behavioural science, health interventions focus on behaviour change by providing a framework to help patients acquire and maintain healthy habits that improve medical outcomes. In-person interventions are costly and difficult to scale, especially in resource-limited regions. Digital health interventions offer a cost-effective approach, potentially supporting independent living and self-management. Automating such interventions, especially through machine learning, has gained considerable attention recently. Ambivalence and hesitancy (A/H) play a primary role for individuals to delay, avoid, or abandon health interventions. A/H are subtle and conflicting emotions that place a person in a state between positive and negative evaluations of a behaviour, or between acceptance and refusal to engage in it. They manifest as affective inconsistency across modalities or within a modality, such as language, facial, vocal expressions, and body language. While experts can be trained to recognize A/H, integrating them into digital health interventions is costly and less effective. Automatic A/H recognition is therefore critical for the personalization and cost-effectiveness of digital health interventions. Here, we explore the application of deep learning models for A/H recognition in videos, a multi-modal task by nature. In particular, this paper covers three learning setups: supervised learning, unsupervised domain adaptation for personalization, and zero-shot inference via large language models (LLMs). Our experiments are conducted on the unique and recently published BAH video dataset for A/H recognition. Our results show limited performance, suggesting that more adapted multi-modal models are required for accurate A/H recognition. Better methods for modeling spatio-temporal and multimodal fusion are necessary to leverage conflicts within/across modalities. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_11730 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions González-González, Manuela Belharbi, Soufiane Zeeshan, Muhammad Osama Sharafi, Masoumeh Aslam, Muhammad Haseeb Sia, Lorenzo Richet, Nicolas Pedersoli, Marco Koerich, Alessandro Lameiras Bacon, Simon L Granger, Eric Computer Vision and Pattern Recognition Human-Computer Interaction Machine Learning Using behavioural science, health interventions focus on behaviour change by providing a framework to help patients acquire and maintain healthy habits that improve medical outcomes. In-person interventions are costly and difficult to scale, especially in resource-limited regions. Digital health interventions offer a cost-effective approach, potentially supporting independent living and self-management. Automating such interventions, especially through machine learning, has gained considerable attention recently. Ambivalence and hesitancy (A/H) play a primary role for individuals to delay, avoid, or abandon health interventions. A/H are subtle and conflicting emotions that place a person in a state between positive and negative evaluations of a behaviour, or between acceptance and refusal to engage in it. They manifest as affective inconsistency across modalities or within a modality, such as language, facial, vocal expressions, and body language. While experts can be trained to recognize A/H, integrating them into digital health interventions is costly and less effective. Automatic A/H recognition is therefore critical for the personalization and cost-effectiveness of digital health interventions. Here, we explore the application of deep learning models for A/H recognition in videos, a multi-modal task by nature. In particular, this paper covers three learning setups: supervised learning, unsupervised domain adaptation for personalization, and zero-shot inference via large language models (LLMs). Our experiments are conducted on the unique and recently published BAH video dataset for A/H recognition. Our results show limited performance, suggesting that more adapted multi-modal models are required for accurate A/H recognition. Better methods for modeling spatio-temporal and multimodal fusion are necessary to leverage conflicts within/across modalities. |
| title | Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions |
| topic | Computer Vision and Pattern Recognition Human-Computer Interaction Machine Learning |
| url | https://arxiv.org/abs/2604.11730 |