BigTokDetect: A Clinically-Informed Vision-Language Modeling Framework for Detecting Pro-Bigorexia Videos on TikTok

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chu, Minh Duc, Pawar, Kshitij, He, Zihao, Sharifi, Roxanna, Sonnenblick, Ross, Curry, Magdalayna, D'Adamo, Laura, Young, Lindsay, Murray, Stuart B, Lerman, Kristina
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912844840173568
author Chu, Minh Duc
Pawar, Kshitij
He, Zihao
Sharifi, Roxanna
Sonnenblick, Ross
Curry, Magdalayna
D'Adamo, Laura
Young, Lindsay
Murray, Stuart B
Lerman, Kristina
author_facet Chu, Minh Duc
Pawar, Kshitij
He, Zihao
Sharifi, Roxanna
Sonnenblick, Ross
Curry, Magdalayna
D'Adamo, Laura
Young, Lindsay
Murray, Stuart B
Lerman, Kristina
contents Social media platforms face escalating challenges in detecting harmful content that promotes muscle dysmorphic behaviors and cognitions (bigorexia). This content can evade moderation by camouflaging as legitimate fitness advice and disproportionately affects adolescent males. We address this challenge with BigTokDetect, a clinically informed framework for identifying pro-bigorexia content on TikTok. We introduce BigTok, the first expert-annotated multimodal benchmark dataset of over 2,200 TikTok videos labeled by clinical psychiatrists across five categories and eighteen fine-grained subcategories. Comprehensive evaluation of state-of-the-art vision-language models reveals that while commercial zero-shot models achieve the highest accuracy on broad primary categories, supervised fine-tuning enables smaller open-source models to perform better on fine-grained subcategory detection. Ablation studies show that multimodal fusion improves performance by 5 to 15 percent, with video features providing the most discriminative signals. These findings support a grounded moderation approach that automates detection of explicit harms while flagging ambiguous content for human review, and they establish a scalable framework for harm mitigation in emerging mental health domains.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06515
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BigTokDetect: A Clinically-Informed Vision-Language Modeling Framework for Detecting Pro-Bigorexia Videos on TikTok
Chu, Minh Duc
Pawar, Kshitij
He, Zihao
Sharifi, Roxanna
Sonnenblick, Ross
Curry, Magdalayna
D'Adamo, Laura
Young, Lindsay
Murray, Stuart B
Lerman, Kristina
Computer Vision and Pattern Recognition
Social media platforms face escalating challenges in detecting harmful content that promotes muscle dysmorphic behaviors and cognitions (bigorexia). This content can evade moderation by camouflaging as legitimate fitness advice and disproportionately affects adolescent males. We address this challenge with BigTokDetect, a clinically informed framework for identifying pro-bigorexia content on TikTok. We introduce BigTok, the first expert-annotated multimodal benchmark dataset of over 2,200 TikTok videos labeled by clinical psychiatrists across five categories and eighteen fine-grained subcategories. Comprehensive evaluation of state-of-the-art vision-language models reveals that while commercial zero-shot models achieve the highest accuracy on broad primary categories, supervised fine-tuning enables smaller open-source models to perform better on fine-grained subcategory detection. Ablation studies show that multimodal fusion improves performance by 5 to 15 percent, with video features providing the most discriminative signals. These findings support a grounded moderation approach that automates detection of explicit harms while flagging ambiguous content for human review, and they establish a scalable framework for harm mitigation in emerging mental health domains.
title BigTokDetect: A Clinically-Informed Vision-Language Modeling Framework for Detecting Pro-Bigorexia Videos on TikTok
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.06515