LLM Agents Predict Social Media Reactions but Do Not Outperform Text Classifiers: Benchmarking Simulation Accuracy Using 120K+ Personas of 1511 Humans
Fuente:
arXiv
Salvato in:
| Autori principali: | Bojic, Ljubisa, Felfernig, Alexander, Dinic, Bojana, Ilic, Velibor, Rettinger, Achim, Mevorah, Vera, Trilling, Damian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Personality testing of Large Language Models: Limited temporal stability, but highlighted prosociality
di: Bodroza, Bojana, et al.
Pubblicazione: (2023)
di: Bodroza, Bojana, et al.
Pubblicazione: (2023)
Towards New Benchmark for AI Alignment & Sentiment Analysis in Socially Important Issues: A Comparative Study of Human and LLMs in the Context of AGI
di: Bojic, Ljubisa, et al.
Pubblicazione: (2025)
di: Bojic, Ljubisa, et al.
Pubblicazione: (2025)
Does GPT-4 surpass human performance in linguistic pragmatics?
di: Bojic, Ljubisa, et al.
Pubblicazione: (2023)
di: Bojic, Ljubisa, et al.
Pubblicazione: (2023)
The Dual Impact of Virtual Reality: Examining the Addictive Potential and Therapeutic Applications of Immersive Media in the Metaverse
di: Bojic, Ljubisa, et al.
Pubblicazione: (2024)
di: Bojic, Ljubisa, et al.
Pubblicazione: (2024)
Maintaining Journalistic Integrity in the Digital Age: A Comprehensive NLP Framework for Evaluating Online News Content
di: Bojic, Ljubisa, et al.
Pubblicazione: (2024)
di: Bojic, Ljubisa, et al.
Pubblicazione: (2024)
CERN for AI: A Theoretical Framework for Autonomous Simulation-Based Artificial Intelligence Testing and Alignment
di: Bojic, Ljubisa, et al.
Pubblicazione: (2023)
di: Bojic, Ljubisa, et al.
Pubblicazione: (2023)
Towards Recommender Systems LLMs Playground (RecSysLLMsP): Exploring Polarization and Engagement in Simulated Social Networks
di: Bojic, Ljubisa, et al.
Pubblicazione: (2025)
di: Bojic, Ljubisa, et al.
Pubblicazione: (2025)
InvBERT: Reconstructing Text from Contextualized Word Embeddings by inverting the BERT pipeline
di: Kugler, Kai, et al.
Pubblicazione: (2021)
di: Kugler, Kai, et al.
Pubblicazione: (2021)
Towards Simulating Social Media Users with LLMs: Evaluating the Operational Validity of Conditioned Comment Prediction
di: Schwager, Nils, et al.
Pubblicazione: (2026)
di: Schwager, Nils, et al.
Pubblicazione: (2026)
Don't Trust Generative Agents to Mimic Communication on Social Networks Unless You Benchmarked their Empirical Realism
di: Münker, Simon, et al.
Pubblicazione: (2025)
di: Münker, Simon, et al.
Pubblicazione: (2025)
Zero-shot prompt-based classification: topic labeling in times of foundation models in German Tweets
di: Münker, Simon, et al.
Pubblicazione: (2024)
di: Münker, Simon, et al.
Pubblicazione: (2024)
Evaluating Large Language Models Against Human Annotators in Latent Content Analysis: Sentiment, Political Leaning, Emotional Intensity, and Sarcasm
di: Bojic, Ljubisa, et al.
Pubblicazione: (2025)
di: Bojic, Ljubisa, et al.
Pubblicazione: (2025)
Incorporation of journalistic approaches into algorithm design
di: Bastian, Mariella, et al.
Pubblicazione: (2025)
di: Bastian, Mariella, et al.
Pubblicazione: (2025)
Identity-Aware Large Language Models require Cultural Reasoning
di: Plum, Alistair, et al.
Pubblicazione: (2025)
di: Plum, Alistair, et al.
Pubblicazione: (2025)
Are generative AI text annotations systematically biased?
di: Stolwijk, Sjoerd B., et al.
Pubblicazione: (2025)
di: Stolwijk, Sjoerd B., et al.
Pubblicazione: (2025)
Next Reply Prediction X Dataset: Linguistic Discrepancies in Naively Generated Content
di: Münker, Simon, et al.
Pubblicazione: (2026)
di: Münker, Simon, et al.
Pubblicazione: (2026)
Recommending Usability Improvements with Multimodal Large Language Models
di: Lubos, Sebastian, et al.
Pubblicazione: (2026)
di: Lubos, Sebastian, et al.
Pubblicazione: (2026)
State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?
di: Pungeršek, Taja Kuzman, et al.
Pubblicazione: (2025)
di: Pungeršek, Taja Kuzman, et al.
Pubblicazione: (2025)
Jacob Grimm und Vuk Karadžić
di: Bojic, Vera
Pubblicazione: (2019)
di: Bojic, Vera
Pubblicazione: (2019)
Reasoner Outperforms: Generative Stance Detection with Rationalization for Social Media
di: Yuan, Jiaqing, et al.
Pubblicazione: (2024)
di: Yuan, Jiaqing, et al.
Pubblicazione: (2024)
Investigating Multimodal Large Language Models to Support Usability Evaluation
di: Lubos, Sebastian, et al.
Pubblicazione: (2025)
di: Lubos, Sebastian, et al.
Pubblicazione: (2025)
Towards LLM-Based Usability Analysis for Recommender User Interfaces
di: Lubos, Sebastian, et al.
Pubblicazione: (2025)
di: Lubos, Sebastian, et al.
Pubblicazione: (2025)
Spatial Priming Outperforms Semantic Prompting: A Grid-Based Approach to Improving LLM Accuracy on Chart Data Extraction
di: Lazarev, Andrei, et al.
Pubblicazione: (2026)
di: Lazarev, Andrei, et al.
Pubblicazione: (2026)
Feature Models
di: Felfernig, Alexander, et al.
Pubblicazione: (2024)
di: Felfernig, Alexander, et al.
Pubblicazione: (2024)
Can AI Outperform Human Experts in Creating Social Media Creatives?
di: Park, Eunkyung, et al.
Pubblicazione: (2024)
di: Park, Eunkyung, et al.
Pubblicazione: (2024)
Conversations with AI Chatbots Increase Short-Term Vaccine Intentions But Do Not Outperform Standard Public Health Messaging
di: Sehgal, Neil K. R., et al.
Pubblicazione: (2025)
di: Sehgal, Neil K. R., et al.
Pubblicazione: (2025)
Adapted Large Language Models Can Outperform Medical Experts in Clinical Text Summarization
di: Van Veen, Dave, et al.
Pubblicazione: (2023)
di: Van Veen, Dave, et al.
Pubblicazione: (2023)
The Dual Personas of Social Media Bots
di: Ng, Lynnette Hui Xian, et al.
Pubblicazione: (2025)
di: Ng, Lynnette Hui Xian, et al.
Pubblicazione: (2025)
Agent-Based Simulations of Online Political Discussions: A Case Study on Elections in Germany
di: Sittar, Abdul, et al.
Pubblicazione: (2025)
di: Sittar, Abdul, et al.
Pubblicazione: (2025)
Machines Do See Color: A Guideline to Classify Different Forms of Racist Discourse in Large Corpora
di: Gordillo, Diana Davila, et al.
Pubblicazione: (2024)
di: Gordillo, Diana Davila, et al.
Pubblicazione: (2024)
Optimizing Robot Programming: Mixed Reality Gripper Control
di: Rettinger, Maximilian, et al.
Pubblicazione: (2025)
di: Rettinger, Maximilian, et al.
Pubblicazione: (2025)
PersonaBooth: Personalized Text-to-Motion Generation
di: Kim, Boeun, et al.
Pubblicazione: (2025)
di: Kim, Boeun, et al.
Pubblicazione: (2025)
Classifying Problem and Solution Framing in Congressional Social Media
di: Melnyk, Misha, et al.
Pubblicazione: (2026)
di: Melnyk, Misha, et al.
Pubblicazione: (2026)
Your Extreme Multi-label Classifier is Secretly a Hierarchical Text Classifier for Free
di: Bertalis, Nerijus, et al.
Pubblicazione: (2024)
di: Bertalis, Nerijus, et al.
Pubblicazione: (2024)
Challenges in Explaining Pretrained Clinical Text Classifiers
di: Miok, Kristian, et al.
Pubblicazione: (2026)
di: Miok, Kristian, et al.
Pubblicazione: (2026)
Prompting Science Report 4: Playing Pretend: Expert Personas Don't Improve Factual Accuracy
di: Basil, Savir, et al.
Pubblicazione: (2025)
di: Basil, Savir, et al.
Pubblicazione: (2025)
Temporal Misalignment in ANN-SNN Conversion and Its Mitigation via Probabilistic Spiking Neurons
di: Bojković, Velibor, et al.
Pubblicazione: (2025)
di: Bojković, Velibor, et al.
Pubblicazione: (2025)
SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages?
di: Li, Senyu, et al.
Pubblicazione: (2025)
di: Li, Senyu, et al.
Pubblicazione: (2025)
Classifying States of the Hopfield Network with Improved Accuracy, Generalization, and Interpretability
di: McAlister, Hayden, et al.
Pubblicazione: (2025)
di: McAlister, Hayden, et al.
Pubblicazione: (2025)
Temperature and Persona Shape LLM Agent Consensus With Minimal Accuracy Gains in Qualitative Coding
di: Borchers, Conrad, et al.
Pubblicazione: (2025)
di: Borchers, Conrad, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Personality testing of Large Language Models: Limited temporal stability, but highlighted prosociality
di: Bodroza, Bojana, et al.
Pubblicazione: (2023) -
Towards New Benchmark for AI Alignment & Sentiment Analysis in Socially Important Issues: A Comparative Study of Human and LLMs in the Context of AGI
di: Bojic, Ljubisa, et al.
Pubblicazione: (2025) -
Does GPT-4 surpass human performance in linguistic pragmatics?
di: Bojic, Ljubisa, et al.
Pubblicazione: (2023) -
The Dual Impact of Virtual Reality: Examining the Addictive Potential and Therapeutic Applications of Immersive Media in the Metaverse
di: Bojic, Ljubisa, et al.
Pubblicazione: (2024) -
Maintaining Journalistic Integrity in the Digital Age: A Comprehensive NLP Framework for Evaluating Online News Content
di: Bojic, Ljubisa, et al.
Pubblicazione: (2024)