Towards New Benchmark for AI Alignment & Sentiment Analysis in Socially Important Issues: A Comparative Study of Human and LLMs in the Context of AGI
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Bojic, Ljubisa, Seychell, Dylan, Cabarkapa, Milan |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Does GPT-4 surpass human performance in linguistic pragmatics?
par: Bojic, Ljubisa, et autres
Publié: (2023)
par: Bojic, Ljubisa, et autres
Publié: (2023)
Evaluating Large Language Models Against Human Annotators in Latent Content Analysis: Sentiment, Political Leaning, Emotional Intensity, and Sarcasm
par: Bojic, Ljubisa, et autres
Publié: (2025)
par: Bojic, Ljubisa, et autres
Publié: (2025)
The Dual Impact of Virtual Reality: Examining the Addictive Potential and Therapeutic Applications of Immersive Media in the Metaverse
par: Bojic, Ljubisa, et autres
Publié: (2024)
par: Bojic, Ljubisa, et autres
Publié: (2024)
LLM Agents Predict Social Media Reactions but Do Not Outperform Text Classifiers: Benchmarking Simulation Accuracy Using 120K+ Personas of 1511 Humans
par: Bojic, Ljubisa, et autres
Publié: (2026)
par: Bojic, Ljubisa, et autres
Publié: (2026)
Towards Recommender Systems LLMs Playground (RecSysLLMsP): Exploring Polarization and Engagement in Simulated Social Networks
par: Bojic, Ljubisa, et autres
Publié: (2025)
par: Bojic, Ljubisa, et autres
Publié: (2025)
CERN for AI: A Theoretical Framework for Autonomous Simulation-Based Artificial Intelligence Testing and Alignment
par: Bojic, Ljubisa, et autres
Publié: (2023)
par: Bojic, Ljubisa, et autres
Publié: (2023)
Are Social Sentiments Inherent in LLMs? An Empirical Study on Extraction of Inter-demographic Sentiments
par: Tanaka, Kunitomo, et autres
Publié: (2024)
par: Tanaka, Kunitomo, et autres
Publié: (2024)
Maintaining Journalistic Integrity in the Digital Age: A Comprehensive NLP Framework for Evaluating Online News Content
par: Bojic, Ljubisa, et autres
Publié: (2024)
par: Bojic, Ljubisa, et autres
Publié: (2024)
Personality testing of Large Language Models: Limited temporal stability, but highlighted prosociality
par: Bodroza, Bojana, et autres
Publié: (2023)
par: Bodroza, Bojana, et autres
Publié: (2023)
AI as a Tool for Fair Journalism: Case Studies from Malta
par: Seychell, Dylan, et autres
Publié: (2024)
par: Seychell, Dylan, et autres
Publié: (2024)
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI
par: Yang, Chao, et autres
Publié: (2024)
par: Yang, Chao, et autres
Publié: (2024)
Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness
par: Alipour, Shayan, et autres
Publié: (2024)
par: Alipour, Shayan, et autres
Publié: (2024)
How Far Are LLMs from Believable AI? A Benchmark for Evaluating the Believability of Human Behavior Simulation
par: Xiao, Yang, et autres
Publié: (2023)
par: Xiao, Yang, et autres
Publié: (2023)
Conversational Alignment with Artificial Intelligence in Context
par: Sterken, Rachel Katharine, et autres
Publié: (2025)
par: Sterken, Rachel Katharine, et autres
Publié: (2025)
Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
par: Li, Ming, et autres
Publié: (2025)
par: Li, Ming, et autres
Publié: (2025)
How Far Are We From AGI: Are LLMs All We Need?
par: Feng, Tao, et autres
Publié: (2024)
par: Feng, Tao, et autres
Publié: (2024)
Sentiment Analysis of Cyberbullying Data in Social Media
par: Susmitha, Arvapalli Sai, et autres
Publié: (2024)
par: Susmitha, Arvapalli Sai, et autres
Publié: (2024)
Comparative Analysis of Image, Video, and Audio Classifiers for Automated News Video Segmentation
par: Attard, Jonathan, et autres
Publié: (2025)
par: Attard, Jonathan, et autres
Publié: (2025)
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
par: Jiang, Han, et autres
Publié: (2025)
par: Jiang, Han, et autres
Publié: (2025)
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
par: Majumdar, Ayan, et autres
Publié: (2025)
par: Majumdar, Ayan, et autres
Publié: (2025)
Expressing Social Emotions: Misalignment Between LLMs and Human Cultural Emotion Norms
par: Bhattacharyya, Sree, et autres
Publié: (2026)
par: Bhattacharyya, Sree, et autres
Publié: (2026)
Imperfectly Cooperative Human-AI Interactions: Comparing the Impacts of Human and AI Attributes in Simulated and User Studies
par: Cohen, Myke C., et autres
Publié: (2026)
par: Cohen, Myke C., et autres
Publié: (2026)
HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning
par: Li, Chance Jiajie, et autres
Publié: (2025)
par: Li, Chance Jiajie, et autres
Publié: (2025)
Human or LLM as Standardized Patients? A Comparative Study for Medical Education
par: Zhang, Bingquan, et autres
Publié: (2025)
par: Zhang, Bingquan, et autres
Publié: (2025)
Evaluating Digital Inclusiveness of Digital Agri-Food Tools Using Large Language Models: A Comparative Analysis Between Human and AI-Based Evaluations
par: Pewinya, Githma, et autres
Publié: (2026)
par: Pewinya, Githma, et autres
Publié: (2026)
Understanding The Effect Of Temperature On Alignment With Human Opinions
par: Pavlovic, Maja, et autres
Publié: (2024)
par: Pavlovic, Maja, et autres
Publié: (2024)
LLMs for Low-Resource Dialect Translation Using Context-Aware Prompting: A Case Study on Sylheti
par: Prama, Tabia Tanzin, et autres
Publié: (2025)
par: Prama, Tabia Tanzin, et autres
Publié: (2025)
Leveraging Explainable AI for LLM Text Attribution: Differentiating Human-Written and Multiple LLMs-Generated Text
par: Najjar, Ayat, et autres
Publié: (2025)
par: Najjar, Ayat, et autres
Publié: (2025)
Are LLMs (Really) Ideological? An IRT-based Analysis and Alignment Tool for Perceived Socio-Economic Bias in LLMs
par: Wachter, Jasmin, et autres
Publié: (2025)
par: Wachter, Jasmin, et autres
Publié: (2025)
The Homogenization Problem in LLMs: Towards Meaningful Diversity in AI Safety
par: Rios-Sialer, Ian
Publié: (2026)
par: Rios-Sialer, Ian
Publié: (2026)
Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through Edits
par: Chakrabarty, Tuhin, et autres
Publié: (2024)
par: Chakrabarty, Tuhin, et autres
Publié: (2024)
Meet Your New Client: Writing Reports for AI -- Benchmarking Information Loss in Market Research Deliverables
par: Simmering, Paul F., et autres
Publié: (2025)
par: Simmering, Paul F., et autres
Publié: (2025)
Which Type of Students can LLMs Act? Investigating Authentic Simulation with Graph-based Human-AI Collaborative System
par: Li, Haoxuan, et autres
Publié: (2025)
par: Li, Haoxuan, et autres
Publié: (2025)
Can AI Debias the News? LLM Interventions Improve Cross-Partisan Receptivity but LLMs Overestimate Their Own Effectiveness
par: Feroz, Faisal, et autres
Publié: (2026)
par: Feroz, Faisal, et autres
Publié: (2026)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
par: Li, Jing-Jing, et autres
Publié: (2026)
par: Li, Jing-Jing, et autres
Publié: (2026)
Polarized Patterns of Language Toxicity and Sentiment of Debunking Posts on Social Media
par: Xu, Wentao, et autres
Publié: (2025)
par: Xu, Wentao, et autres
Publié: (2025)
AI Safety, Alignment, and Ethics (AI SAE)
par: Waldner, Dylan
Publié: (2025)
par: Waldner, Dylan
Publié: (2025)
GermanPartiesQA: Benchmarking Commercial Large Language Models and AI Companions for Political Alignment and Sycophancy
par: Batzner, Jan, et autres
Publié: (2024)
par: Batzner, Jan, et autres
Publié: (2024)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
par: Agarwal, Dhruv, et autres
Publié: (2025)
par: Agarwal, Dhruv, et autres
Publié: (2025)
Large Language Models' Accuracy in Emulating Human Experts' Evaluation of Public Sentiments about Heated Tobacco Products on Social Media
par: Kim, Kwanho, et autres
Publié: (2025)
par: Kim, Kwanho, et autres
Publié: (2025)
Documents similaires
-
Does GPT-4 surpass human performance in linguistic pragmatics?
par: Bojic, Ljubisa, et autres
Publié: (2023) -
Evaluating Large Language Models Against Human Annotators in Latent Content Analysis: Sentiment, Political Leaning, Emotional Intensity, and Sarcasm
par: Bojic, Ljubisa, et autres
Publié: (2025) -
The Dual Impact of Virtual Reality: Examining the Addictive Potential and Therapeutic Applications of Immersive Media in the Metaverse
par: Bojic, Ljubisa, et autres
Publié: (2024) -
LLM Agents Predict Social Media Reactions but Do Not Outperform Text Classifiers: Benchmarking Simulation Accuracy Using 120K+ Personas of 1511 Humans
par: Bojic, Ljubisa, et autres
Publié: (2026) -
Towards Recommender Systems LLMs Playground (RecSysLLMsP): Exploring Polarization and Engagement in Simulated Social Networks
par: Bojic, Ljubisa, et autres
Publié: (2025)