SACRED: A Faithful Annotated Multimedia Multimodal Multilingual Dataset for Classifying Connectedness Types in Online Spirituality

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Guan, Qinghao, Pan, Yuchen, Li, Donghao, Zhang, Zishi, Chen, Yiyang, Li, Lu, Canu, Flaminia, Volkart, Emilia, Schneider, Gerold
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912986690486272
author Guan, Qinghao
Pan, Yuchen
Li, Donghao
Zhang, Zishi
Chen, Yiyang
Li, Lu
Canu, Flaminia
Volkart, Emilia
Schneider, Gerold
author_facet Guan, Qinghao
Pan, Yuchen
Li, Donghao
Zhang, Zishi
Chen, Yiyang
Li, Lu
Canu, Flaminia
Volkart, Emilia
Schneider, Gerold
contents In religion and theology studies, spirituality has garnered significant research attention for the reason that it not only transcends culture but offers unique experience to each individual. However, social scientists often rely on limited datasets, which are basically unavailable online. In this study, we collaborated with social scientists to develop a high-quality multimedia multi-modal datasets, \textbf{SACRED}, in which the faithfulness of classification is guaranteed. Using \textbf{SACRED}, we evaluated the performance of 13 popular LLMs as well as traditional rule-based and fine-tuned approaches. The result suggests DeepSeek-V3 model performs well in classifying such abstract concepts (i.e., 79.19\% accuracy in the Quora test set), and the GPT-4o-mini model surpassed the other models in the vision tasks (63.99\% F1 score). Purportedly, this is the first annotated multi-modal dataset from online spirituality communication. Our study also found a new type of connectedness which is valuable for communication science studies.
format Preprint
id arxiv_https___arxiv_org_abs_2603_27331
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SACRED: A Faithful Annotated Multimedia Multimodal Multilingual Dataset for Classifying Connectedness Types in Online Spirituality
Guan, Qinghao
Pan, Yuchen
Li, Donghao
Zhang, Zishi
Chen, Yiyang
Li, Lu
Canu, Flaminia
Volkart, Emilia
Schneider, Gerold
Computation and Language
Multimedia
In religion and theology studies, spirituality has garnered significant research attention for the reason that it not only transcends culture but offers unique experience to each individual. However, social scientists often rely on limited datasets, which are basically unavailable online. In this study, we collaborated with social scientists to develop a high-quality multimedia multi-modal datasets, \textbf{SACRED}, in which the faithfulness of classification is guaranteed. Using \textbf{SACRED}, we evaluated the performance of 13 popular LLMs as well as traditional rule-based and fine-tuned approaches. The result suggests DeepSeek-V3 model performs well in classifying such abstract concepts (i.e., 79.19\% accuracy in the Quora test set), and the GPT-4o-mini model surpassed the other models in the vision tasks (63.99\% F1 score). Purportedly, this is the first annotated multi-modal dataset from online spirituality communication. Our study also found a new type of connectedness which is valuable for communication science studies.
title SACRED: A Faithful Annotated Multimedia Multimodal Multilingual Dataset for Classifying Connectedness Types in Online Spirituality
topic Computation and Language
Multimedia
url https://arxiv.org/abs/2603.27331