LLM-based Semantic Augmentation for Harmful Content Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Meguellati, Elyas, Zeghina, Assaad, Sadiq, Shazia, Demartini, Gianluca |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM-Generated Ads: From Personalization Parity to Persuasion Superiority
von: Meguellati, Elyas, et al.
Veröffentlicht: (2025)
von: Meguellati, Elyas, et al.
Veröffentlicht: (2025)
Towards Detecting Persuasion on Social Media: From Model Development to Insights on Persuasion Strategies
von: Meguellati, Elyas, et al.
Veröffentlicht: (2025)
von: Meguellati, Elyas, et al.
Veröffentlicht: (2025)
Are Large Language Models Good Data Preprocessors?
von: Meguellati, Elyas, et al.
Veröffentlicht: (2025)
von: Meguellati, Elyas, et al.
Veröffentlicht: (2025)
Identification of Regulatory Requirements Relevant to Business Processes: A Comparative Study on Generative AI, Embedding-based Ranking, Crowd and Expert-driven Methods
von: Sai, Catherine, et al.
Veröffentlicht: (2024)
von: Sai, Catherine, et al.
Veröffentlicht: (2024)
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
The Impact of Persona-based Political Perspectives on Hateful Content Detection
von: Civelli, Stefano, et al.
Veröffentlicht: (2025)
von: Civelli, Stefano, et al.
Veröffentlicht: (2025)
People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection
von: Sen, Indira, et al.
Veröffentlicht: (2023)
von: Sen, Indira, et al.
Veröffentlicht: (2023)
Harmful Suicide Content Detection
von: Park, Kyumin, et al.
Veröffentlicht: (2024)
von: Park, Kyumin, et al.
Veröffentlicht: (2024)
From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content Monitoring
von: Li, Yang, et al.
Veröffentlicht: (2025)
von: Li, Yang, et al.
Veröffentlicht: (2025)
The Hidden Language of Harm: Examining the Role of Emojis in Harmful Online Communication and Content Moderation
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework
von: Gajewska, Ewelina, et al.
Veröffentlicht: (2026)
von: Gajewska, Ewelina, et al.
Veröffentlicht: (2026)
Harmful Speech Detection by Language Models Exhibits Gender-Queer Dialect Bias
von: Dorn, Rebecca, et al.
Veröffentlicht: (2024)
von: Dorn, Rebecca, et al.
Veröffentlicht: (2024)
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems
von: Harvey, Emma, et al.
Veröffentlicht: (2025)
von: Harvey, Emma, et al.
Veröffentlicht: (2025)
Query-Document Dense Vectors for LLM Relevance Judgment Bias Analysis
von: Mohtadi, Samaneh, et al.
Veröffentlicht: (2026)
von: Mohtadi, Samaneh, et al.
Veröffentlicht: (2026)
Ideology-Based LLMs for Content Moderation
von: Civelli, Stefano, et al.
Veröffentlicht: (2025)
von: Civelli, Stefano, et al.
Veröffentlicht: (2025)
Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant
von: He, Gaole, et al.
Veröffentlicht: (2025)
von: He, Gaole, et al.
Veröffentlicht: (2025)
Context Shapes LLMs Retrieval-Augmented Fact-Checking Effectiveness
von: Bernardelle, Pietro, et al.
Veröffentlicht: (2026)
von: Bernardelle, Pietro, et al.
Veröffentlicht: (2026)
Careless Whisper: Speech-to-Text Hallucination Harms
von: Koenecke, Allison, et al.
Veröffentlicht: (2024)
von: Koenecke, Allison, et al.
Veröffentlicht: (2024)
Taxonomizing Representational Harms using Speech Act Theory
von: Corvi, Emily, et al.
Veröffentlicht: (2025)
von: Corvi, Emily, et al.
Veröffentlicht: (2025)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
von: Li, Jing-Jing, et al.
Veröffentlicht: (2026)
von: Li, Jing-Jing, et al.
Veröffentlicht: (2026)
On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025)
von: Ghorbanpour, Faeze, et al.
Veröffentlicht: (2025)
A Capabilities Approach to Studying Bias and Harm in Language Technologies
von: Nigatu, Hellina Hailu, et al.
Veröffentlicht: (2024)
von: Nigatu, Hellina Hailu, et al.
Veröffentlicht: (2024)
Limited Effectiveness of LLM-based Data Augmentation for COVID-19 Misinformation Stance Detection
von: Choi, Eun Cheol, et al.
Veröffentlicht: (2025)
von: Choi, Eun Cheol, et al.
Veröffentlicht: (2025)
Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs
von: Arnaiz-Rodriguez, Adrian, et al.
Veröffentlicht: (2025)
von: Arnaiz-Rodriguez, Adrian, et al.
Veröffentlicht: (2025)
Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?
von: Cyberey, Hannah, et al.
Veröffentlicht: (2024)
von: Cyberey, Hannah, et al.
Veröffentlicht: (2024)
CheckIfExist: Detecting Citation Hallucinations in the Era of AI-Generated Content
von: Abbonato, Diletta
Veröffentlicht: (2026)
von: Abbonato, Diletta
Veröffentlicht: (2026)
MiningGPT -- A Domain-Specific Large Language Model for the Mining Industry
von: Demartini, Kurukulasooriya Fernando ana Gianluca
Veröffentlicht: (2024)
von: Demartini, Kurukulasooriya Fernando ana Gianluca
Veröffentlicht: (2024)
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards
von: Chehbouni, Khaoula, et al.
Veröffentlicht: (2024)
von: Chehbouni, Khaoula, et al.
Veröffentlicht: (2024)
Longitudinal Monitoring of LLM Content Moderation of Social Issues
von: Dai, Yunlang, et al.
Veröffentlicht: (2025)
von: Dai, Yunlang, et al.
Veröffentlicht: (2025)
Exploring Wikipedia Gender Diversity Over Time $\unicode{x2013}$ The Wikipedia Gender Dashboard (WGD)
von: Yunus, Yahya, et al.
Veröffentlicht: (2025)
von: Yunus, Yahya, et al.
Veröffentlicht: (2025)
Leveraging Semantic Type Dependencies for Clinical Named Entity Recognition
von: Le, Linh, et al.
Veröffentlicht: (2025)
von: Le, Linh, et al.
Veröffentlicht: (2025)
Diverse, but Divisive: LLMs Can Exaggerate Gender Differences in Opinion Related to Harms of Misinformation
von: Neumann, Terrence, et al.
Veröffentlicht: (2024)
von: Neumann, Terrence, et al.
Veröffentlicht: (2024)
Political Advertising on Facebook During the 2022 Australian Federal Election: A Social Identity Perspective
von: Civelli, Stefano, et al.
Veröffentlicht: (2025)
von: Civelli, Stefano, et al.
Veröffentlicht: (2025)
HRIPBench: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs
von: Wang, Kaixuan, et al.
Veröffentlicht: (2025)
von: Wang, Kaixuan, et al.
Veröffentlicht: (2025)
Leveraging Lecture Content for Improved Feedback: Explorations with GPT-4 and Retrieval Augmented Generation
von: Jacobs, Sven, et al.
Veröffentlicht: (2024)
von: Jacobs, Sven, et al.
Veröffentlicht: (2024)
Personas with Attitudes: Controlling LLMs for Diverse Data Annotation
von: Fröhling, Leon, et al.
Veröffentlicht: (2024)
von: Fröhling, Leon, et al.
Veröffentlicht: (2024)
What do Large Language Models Say About Animals? Investigating Risks of Animal Harm in Generated Text
von: Kanepajs, Arturs, et al.
Veröffentlicht: (2025)
von: Kanepajs, Arturs, et al.
Veröffentlicht: (2025)
Re-ranking Using Large Language Models for Mitigating Exposure to Harmful Content on Social Media Platforms
von: Oak, Rajvardhan, et al.
Veröffentlicht: (2025)
von: Oak, Rajvardhan, et al.
Veröffentlicht: (2025)
"They are uncultured": Unveiling Covert Harms and Social Threats in LLM Generated Conversations
von: Dammu, Preetam Prabhu Srikar, et al.
Veröffentlicht: (2024)
von: Dammu, Preetam Prabhu Srikar, et al.
Veröffentlicht: (2024)
Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions
von: Borah, Angana, et al.
Veröffentlicht: (2024)
von: Borah, Angana, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LLM-Generated Ads: From Personalization Parity to Persuasion Superiority
von: Meguellati, Elyas, et al.
Veröffentlicht: (2025) -
Towards Detecting Persuasion on Social Media: From Model Development to Insights on Persuasion Strategies
von: Meguellati, Elyas, et al.
Veröffentlicht: (2025) -
Are Large Language Models Good Data Preprocessors?
von: Meguellati, Elyas, et al.
Veröffentlicht: (2025) -
Identification of Regulatory Requirements Relevant to Business Processes: A Comparative Study on Generative AI, Embedding-based Ranking, Crowd and Expert-driven Methods
von: Sai, Catherine, et al.
Veröffentlicht: (2024) -
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
von: Zhang, Chi, et al.
Veröffentlicht: (2025)