Figurative-cum-Commonsense Knowledge Infusion for Multimodal Mental Health Meme Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mazhar, Abdullah, shaik, Zuhair hasan, Srivastava, Aseem, Ruhnke, Polly, Vaddavalli, Lavanya, Katragadda, Sri Keshav, Yadav, Shweta, Akhtar, Md Shad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913666619670528
author Mazhar, Abdullah
shaik, Zuhair hasan
Srivastava, Aseem
Ruhnke, Polly
Vaddavalli, Lavanya
Katragadda, Sri Keshav
Yadav, Shweta
Akhtar, Md Shad
author_facet Mazhar, Abdullah
shaik, Zuhair hasan
Srivastava, Aseem
Ruhnke, Polly
Vaddavalli, Lavanya
Katragadda, Sri Keshav
Yadav, Shweta
Akhtar, Md Shad
contents The expression of mental health symptoms through non-traditional means, such as memes, has gained remarkable attention over the past few years, with users often highlighting their mental health struggles through figurative intricacies within memes. While humans rely on commonsense knowledge to interpret these complex expressions, current Multimodal Language Models (MLMs) struggle to capture these figurative aspects inherent in memes. To address this gap, we introduce a novel dataset, AxiOM, derived from the GAD anxiety questionnaire, which categorizes memes into six fine-grained anxiety symptoms. Next, we propose a commonsense and domain-enriched framework, M3H, to enhance MLMs' ability to interpret figurative language and commonsense knowledge. The overarching goal remains to first understand and then classify the mental health symptoms expressed in memes. We benchmark M3H against 6 competitive baselines (with 20 variations), demonstrating improvements in both quantitative and qualitative metrics, including a detailed human evaluation. We observe a clear improvement of 4.20% and 4.66% on weighted-F1 metric. To assess the generalizability, we perform extensive experiments on a public dataset, RESTORE, for depressive symptom identification, presenting an extensive ablation study that highlights the contribution of each module in both datasets. Our findings reveal limitations in existing models and the advantage of employing commonsense to enhance figurative understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2501_15321
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Figurative-cum-Commonsense Knowledge Infusion for Multimodal Mental Health Meme Classification
Mazhar, Abdullah
shaik, Zuhair hasan
Srivastava, Aseem
Ruhnke, Polly
Vaddavalli, Lavanya
Katragadda, Sri Keshav
Yadav, Shweta
Akhtar, Md Shad
Computation and Language
Social and Information Networks
The expression of mental health symptoms through non-traditional means, such as memes, has gained remarkable attention over the past few years, with users often highlighting their mental health struggles through figurative intricacies within memes. While humans rely on commonsense knowledge to interpret these complex expressions, current Multimodal Language Models (MLMs) struggle to capture these figurative aspects inherent in memes. To address this gap, we introduce a novel dataset, AxiOM, derived from the GAD anxiety questionnaire, which categorizes memes into six fine-grained anxiety symptoms. Next, we propose a commonsense and domain-enriched framework, M3H, to enhance MLMs' ability to interpret figurative language and commonsense knowledge. The overarching goal remains to first understand and then classify the mental health symptoms expressed in memes. We benchmark M3H against 6 competitive baselines (with 20 variations), demonstrating improvements in both quantitative and qualitative metrics, including a detailed human evaluation. We observe a clear improvement of 4.20% and 4.66% on weighted-F1 metric. To assess the generalizability, we perform extensive experiments on a public dataset, RESTORE, for depressive symptom identification, presenting an extensive ablation study that highlights the contribution of each module in both datasets. Our findings reveal limitations in existing models and the advantage of employing commonsense to enhance figurative understanding.
title Figurative-cum-Commonsense Knowledge Infusion for Multimodal Mental Health Meme Classification
topic Computation and Language
Social and Information Networks
url https://arxiv.org/abs/2501.15321