Ideology-Based LLMs for Content Moderation
Fuente:
arXiv
Saved in:
| Main Authors: | Civelli, Stefano, Bernardelle, Pietro, Pratama, Nardiena A., Demartini, Gianluca |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Impact of Persona-based Political Perspectives on Hateful Content Detection
by: Civelli, Stefano, et al.
Published: (2025)
by: Civelli, Stefano, et al.
Published: (2025)
Context Shapes LLMs Retrieval-Augmented Fact-Checking Effectiveness
by: Bernardelle, Pietro, et al.
Published: (2026)
by: Bernardelle, Pietro, et al.
Published: (2026)
SubData: Bridging Heterogeneous Datasets to Enable Theory-Driven Evaluation of Political and Demographic Perspectives in LLMs
by: Bernardelle, Pietro, et al.
Published: (2024)
by: Bernardelle, Pietro, et al.
Published: (2024)
Political Ideology Shifts in Large Language Models
by: Bernardelle, Pietro, et al.
Published: (2025)
by: Bernardelle, Pietro, et al.
Published: (2025)
A Shared Geometry of Difficulty in Multilingual Language Models
by: Civelli, Stefano, et al.
Published: (2026)
by: Civelli, Stefano, et al.
Published: (2026)
Mapping and Influencing the Political Ideology of Large Language Models using Synthetic Personas
by: Bernardelle, Pietro, et al.
Published: (2024)
by: Bernardelle, Pietro, et al.
Published: (2024)
Are Large Language Models Good Data Preprocessors?
by: Meguellati, Elyas, et al.
Published: (2025)
by: Meguellati, Elyas, et al.
Published: (2025)
Political Advertising on Facebook During the 2022 Australian Federal Election: A Social Identity Perspective
by: Civelli, Stefano, et al.
Published: (2025)
by: Civelli, Stefano, et al.
Published: (2025)
Towards Detecting Persuasion on Social Media: From Model Development to Insights on Persuasion Strategies
by: Meguellati, Elyas, et al.
Published: (2025)
by: Meguellati, Elyas, et al.
Published: (2025)
Perception of Visual Content: Differences Between Humans and Foundation Models
by: Pratama, Nardiena A., et al.
Published: (2024)
by: Pratama, Nardiena A., et al.
Published: (2024)
Optimizing LLMs with Direct Preferences: A Data Efficiency Perspective
by: Bernardelle, Pietro, et al.
Published: (2024)
by: Bernardelle, Pietro, et al.
Published: (2024)
LLM-Generated Ads: From Personalization Parity to Persuasion Superiority
by: Meguellati, Elyas, et al.
Published: (2025)
by: Meguellati, Elyas, et al.
Published: (2025)
Personas with Attitudes: Controlling LLMs for Diverse Data Annotation
by: Fröhling, Leon, et al.
Published: (2024)
by: Fröhling, Leon, et al.
Published: (2024)
LLM-based Semantic Augmentation for Harmful Content Detection
by: Meguellati, Elyas, et al.
Published: (2025)
by: Meguellati, Elyas, et al.
Published: (2025)
The Effect of Document Summarization on LLM-Based Relevance Judgments
by: Mohtadi, Samaneh, et al.
Published: (2025)
by: Mohtadi, Samaneh, et al.
Published: (2025)
MiningGPT -- A Domain-Specific Large Language Model for the Mining Industry
by: Demartini, Kurukulasooriya Fernando ana Gianluca
Published: (2024)
by: Demartini, Kurukulasooriya Fernando ana Gianluca
Published: (2024)
Query-Document Dense Vectors for LLM Relevance Judgment Bias Analysis
by: Mohtadi, Samaneh, et al.
Published: (2026)
by: Mohtadi, Samaneh, et al.
Published: (2026)
Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant
by: He, Gaole, et al.
Published: (2025)
by: He, Gaole, et al.
Published: (2025)
SLM-Mod: Small Language Models Surpass LLMs at Content Moderation
by: Zhan, Xianyang, et al.
Published: (2024)
by: Zhan, Xianyang, et al.
Published: (2024)
Leveraging Semantic Type Dependencies for Clinical Named Entity Recognition
by: Le, Linh, et al.
Published: (2025)
by: Le, Linh, et al.
Published: (2025)
Natural Language Processing for the Legal Domain: A Survey of Tasks, Datasets, Models, and Challenges
by: Ariai, Farid, et al.
Published: (2024)
by: Ariai, Farid, et al.
Published: (2024)
Measurement in the Age of LLMs: An Application to Ideological Scaling
by: O'Hagan, Sean, et al.
Published: (2023)
by: O'Hagan, Sean, et al.
Published: (2023)
Graph-Augmented LLMs for Swiss MP Ideology Prediction
by: Yuan, Yifei, et al.
Published: (2026)
by: Yuan, Yifei, et al.
Published: (2026)
Identification of Regulatory Requirements Relevant to Business Processes: A Comparative Study on Generative AI, Embedding-based Ranking, Crowd and Expert-driven Methods
by: Sai, Catherine, et al.
Published: (2024)
by: Sai, Catherine, et al.
Published: (2024)
The Impact of AI Generated Content on Decision Making for Topics Requiring Expertise
by: Li, Shangqian, et al.
Published: (2026)
by: Li, Shangqian, et al.
Published: (2026)
ShieldGemma: Generative AI Content Moderation Based on Gemma
by: Zeng, Wenjun, et al.
Published: (2024)
by: Zeng, Wenjun, et al.
Published: (2024)
Hate Speech Detection with Generalizable Target-aware Fairness
by: Chen, Tong, et al.
Published: (2024)
by: Chen, Tong, et al.
Published: (2024)
Longitudinal Monitoring of LLM Content Moderation of Social Issues
by: Dai, Yunlang, et al.
Published: (2025)
by: Dai, Yunlang, et al.
Published: (2025)
Socio-Culturally Aware Evaluation Framework for LLM-Based Content Moderation
by: Kumar, Shanu, et al.
Published: (2024)
by: Kumar, Shanu, et al.
Published: (2024)
ExpGuard: LLM Content Moderation in Specialized Domains
by: Choi, Minseok, et al.
Published: (2026)
by: Choi, Minseok, et al.
Published: (2026)
Beyond Hate: Differentiating Uncivil and Intolerant Speech in Multimodal Content Moderation
by: Herrmann, Nils A., et al.
Published: (2026)
by: Herrmann, Nils A., et al.
Published: (2026)
HateModerate: Testing Hate Speech Detectors against Content Moderation Policies
by: Zheng, Jiangrui, et al.
Published: (2023)
by: Zheng, Jiangrui, et al.
Published: (2023)
BingoGuard: LLM Content Moderation Tools with Risk Levels
by: Yin, Fan, et al.
Published: (2025)
by: Yin, Fan, et al.
Published: (2025)
The Unappreciated Role of Intent in Algorithmic Moderation of Social Media Content
by: Wang, Xinyu, et al.
Published: (2024)
by: Wang, Xinyu, et al.
Published: (2024)
Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations
by: Hartmann, David, et al.
Published: (2025)
by: Hartmann, David, et al.
Published: (2025)
Identity-related Speech Suppression in Generative AI Content Moderation
by: Proebsting, Grace, et al.
Published: (2024)
by: Proebsting, Grace, et al.
Published: (2024)
Bias Beyond Borders: Political Ideology Evaluation and Steering in Multilingual LLMs
by: Nadeem, Afrozah, et al.
Published: (2026)
by: Nadeem, Afrozah, et al.
Published: (2026)
Legilimens: Practical and Unified Content Moderation for Large Language Model Services
by: Wu, Jialin, et al.
Published: (2024)
by: Wu, Jialin, et al.
Published: (2024)
STAND-Guard: A Small Task-Adaptive Content Moderation Model
by: Wang, Minjia, et al.
Published: (2024)
by: Wang, Minjia, et al.
Published: (2024)
Crowdsourcing Piedmontese to Test LLMs on Non-Standard Orthography
by: Vico, Gianluca, et al.
Published: (2026)
by: Vico, Gianluca, et al.
Published: (2026)
Similar Items
-
The Impact of Persona-based Political Perspectives on Hateful Content Detection
by: Civelli, Stefano, et al.
Published: (2025) -
Context Shapes LLMs Retrieval-Augmented Fact-Checking Effectiveness
by: Bernardelle, Pietro, et al.
Published: (2026) -
SubData: Bridging Heterogeneous Datasets to Enable Theory-Driven Evaluation of Political and Demographic Perspectives in LLMs
by: Bernardelle, Pietro, et al.
Published: (2024) -
Political Ideology Shifts in Large Language Models
by: Bernardelle, Pietro, et al.
Published: (2025) -
A Shared Geometry of Difficulty in Multilingual Language Models
by: Civelli, Stefano, et al.
Published: (2026)