SoMe: A Realistic Benchmark for LLM-based Social Media Agents

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xue, Dizhan, Cui, Jing, Qian, Shengsheng, Hu, Chuanrui, Xu, Changsheng
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915680408829952
author Xue, Dizhan
Cui, Jing
Qian, Shengsheng
Hu, Chuanrui
Xu, Changsheng
author_facet Xue, Dizhan
Cui, Jing
Qian, Shengsheng
Hu, Chuanrui
Xu, Changsheng
contents Intelligent agents powered by large language models (LLMs) have recently demonstrated impressive capabilities and gained increasing popularity on social media platforms. While LLM agents are reshaping the ecology of social media, there exists a current gap in conducting a comprehensive evaluation of their ability to comprehend media content, understand user behaviors, and make intricate decisions. To address this challenge, we introduce SoMe, a pioneering benchmark designed to evaluate social media agents equipped with various agent tools for accessing and analyzing social media data. SoMe comprises a diverse collection of 8 social media agent tasks, 9,164,284 posts, 6,591 user profiles, and 25,686 reports from various social media platforms and external websites, with 17,869 meticulously annotated task queries. Compared with the existing datasets and benchmarks for social media tasks, SoMe is the first to provide a versatile and realistic platform for LLM-based social media agents to handle diverse social media tasks. By extensive quantitative and qualitative analysis, we provide the first overview insight into the performance of mainstream agentic LLMs in realistic social media environments and identify several limitations. Our evaluation reveals that both the current closed-source and open-source LLMs cannot handle social media agent tasks satisfactorily. SoMe provides a challenging yet meaningful testbed for future social media agents. Our code and data are available at https://github.com/LivXue/SoMe
format Preprint
id arxiv_https___arxiv_org_abs_2512_14720
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SoMe: A Realistic Benchmark for LLM-based Social Media Agents
Xue, Dizhan
Cui, Jing
Qian, Shengsheng
Hu, Chuanrui
Xu, Changsheng
Social and Information Networks
Artificial Intelligence
Computation and Language
Intelligent agents powered by large language models (LLMs) have recently demonstrated impressive capabilities and gained increasing popularity on social media platforms. While LLM agents are reshaping the ecology of social media, there exists a current gap in conducting a comprehensive evaluation of their ability to comprehend media content, understand user behaviors, and make intricate decisions. To address this challenge, we introduce SoMe, a pioneering benchmark designed to evaluate social media agents equipped with various agent tools for accessing and analyzing social media data. SoMe comprises a diverse collection of 8 social media agent tasks, 9,164,284 posts, 6,591 user profiles, and 25,686 reports from various social media platforms and external websites, with 17,869 meticulously annotated task queries. Compared with the existing datasets and benchmarks for social media tasks, SoMe is the first to provide a versatile and realistic platform for LLM-based social media agents to handle diverse social media tasks. By extensive quantitative and qualitative analysis, we provide the first overview insight into the performance of mainstream agentic LLMs in realistic social media environments and identify several limitations. Our evaluation reveals that both the current closed-source and open-source LLMs cannot handle social media agent tasks satisfactorily. SoMe provides a challenging yet meaningful testbed for future social media agents. Our code and data are available at https://github.com/LivXue/SoMe
title SoMe: A Realistic Benchmark for LLM-based Social Media Agents
topic Social and Information Networks
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2512.14720