SNS-Bench-VL: Benchmarking Multimodal Large Language Models in Social Networking Services

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Hongcheng, Xie, Zheyong, Cao, Shaosheng, Wang, Boyang, Liu, Weiting, Le, Anjie, Li, Lei, Li, Zhoujun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908710274596864
author Guo, Hongcheng
Xie, Zheyong
Cao, Shaosheng
Wang, Boyang
Liu, Weiting
Le, Anjie
Li, Lei
Li, Zhoujun
author_facet Guo, Hongcheng
Xie, Zheyong
Cao, Shaosheng
Wang, Boyang
Liu, Weiting
Le, Anjie
Li, Lei
Li, Zhoujun
contents With the increasing integration of visual and textual content in Social Networking Services (SNS), evaluating the multimodal capabilities of Large Language Models (LLMs) is crucial for enhancing user experience, content understanding, and platform intelligence. Existing benchmarks primarily focus on text-centric tasks, lacking coverage of the multimodal contexts prevalent in modern SNS ecosystems. In this paper, we introduce SNS-Bench-VL, a comprehensive multimodal benchmark designed to assess the performance of Vision-Language LLMs in real-world social media scenarios. SNS-Bench-VL incorporates images and text across 8 multimodal tasks, including note comprehension, user engagement analysis, information retrieval, and personalized recommendation. It comprises 4,001 carefully curated multimodal question-answer pairs, covering single-choice, multiple-choice, and open-ended tasks. We evaluate over 25 state-of-the-art multimodal LLMs, analyzing their performance across tasks. Our findings highlight persistent challenges in multimodal social context comprehension. We hope SNS-Bench-VL will inspire future research towards robust, context-aware, and human-aligned multimodal intelligence for next-generation social networking services.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23065
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SNS-Bench-VL: Benchmarking Multimodal Large Language Models in Social Networking Services
Guo, Hongcheng
Xie, Zheyong
Cao, Shaosheng
Wang, Boyang
Liu, Weiting
Le, Anjie
Li, Lei
Li, Zhoujun
Computation and Language
With the increasing integration of visual and textual content in Social Networking Services (SNS), evaluating the multimodal capabilities of Large Language Models (LLMs) is crucial for enhancing user experience, content understanding, and platform intelligence. Existing benchmarks primarily focus on text-centric tasks, lacking coverage of the multimodal contexts prevalent in modern SNS ecosystems. In this paper, we introduce SNS-Bench-VL, a comprehensive multimodal benchmark designed to assess the performance of Vision-Language LLMs in real-world social media scenarios. SNS-Bench-VL incorporates images and text across 8 multimodal tasks, including note comprehension, user engagement analysis, information retrieval, and personalized recommendation. It comprises 4,001 carefully curated multimodal question-answer pairs, covering single-choice, multiple-choice, and open-ended tasks. We evaluate over 25 state-of-the-art multimodal LLMs, analyzing their performance across tasks. Our findings highlight persistent challenges in multimodal social context comprehension. We hope SNS-Bench-VL will inspire future research towards robust, context-aware, and human-aligned multimodal intelligence for next-generation social networking services.
title SNS-Bench-VL: Benchmarking Multimodal Large Language Models in Social Networking Services
topic Computation and Language
url https://arxiv.org/abs/2505.23065