Watch Your Language: Investigating Content Moderation with Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kumar, Deepak, AbuHashem, Yousef, Durumeric, Zakir |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Characterizing the MrDeepFakes Sexual Deepfake Marketplace
by: Han, Catherine, et al.
Published: (2024)
by: Han, Catherine, et al.
Published: (2024)
Keys on Doormats: Exposed API Credentials on the Web
by: Demir, Nurullah, et al.
Published: (2026)
by: Demir, Nurullah, et al.
Published: (2026)
Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis
by: Esposito, Matteo, et al.
Published: (2024)
by: Esposito, Matteo, et al.
Published: (2024)
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models
by: Wang, Jiongxiao, et al.
Published: (2023)
by: Wang, Jiongxiao, et al.
Published: (2023)
JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
by: Feng, Yingchaojie, et al.
Published: (2024)
by: Feng, Yingchaojie, et al.
Published: (2024)
A Security Risk Taxonomy for Prompt-Based Interaction With Large Language Models
by: Derner, Erik, et al.
Published: (2023)
by: Derner, Erik, et al.
Published: (2023)
SECURE: Benchmarking Large Language Models for Cybersecurity
by: Bhusal, Dipkamal, et al.
Published: (2024)
by: Bhusal, Dipkamal, et al.
Published: (2024)
"Are You Sure?": An Empirical Study of Human Perception Vulnerability in LLM-Driven Agentic Systems
by: Li, Xinfeng, et al.
Published: (2026)
by: Li, Xinfeng, et al.
Published: (2026)
Human-Centered Privacy Research in the Age of Large Language Models
by: Li, Tianshi, et al.
Published: (2024)
by: Li, Tianshi, et al.
Published: (2024)
InjectLab: A Tactical Framework for Adversarial Threat Modeling Against Large Language Models
by: Howard, Austin
Published: (2025)
by: Howard, Austin
Published: (2025)
The Silicon Psyche: Anthropomorphic Vulnerabilities in Large Language Models
by: Canale, Giuseppe, et al.
Published: (2025)
by: Canale, Giuseppe, et al.
Published: (2025)
Towards Proactive Defense Against Cyber Cognitive Attacks
by: Rushing, Bonnie, et al.
Published: (2025)
by: Rushing, Bonnie, et al.
Published: (2025)
'Debunk-It-Yourself': Health Professionals' Strategies for Responding to Misinformation on TikTok
by: Sharevski, Filipo, et al.
Published: (2024)
by: Sharevski, Filipo, et al.
Published: (2024)
SoK: The Privacy Paradox of Large Language Models: Advancements, Privacy Risks, and Mitigation
by: Shanmugarasa, Yashothara, et al.
Published: (2025)
by: Shanmugarasa, Yashothara, et al.
Published: (2025)
Hacc-Man: An Arcade Game for Jailbreaking LLMs
by: Valentim, Matheus, et al.
Published: (2024)
by: Valentim, Matheus, et al.
Published: (2024)
Emergent misalignment as prompt sensitivity: A research note
by: Wyse, Tim, et al.
Published: (2025)
by: Wyse, Tim, et al.
Published: (2025)
Learned, Lagged, LLM-splained: LLM Responses to End User Security Questions
by: Prakash, Vijay, et al.
Published: (2024)
by: Prakash, Vijay, et al.
Published: (2024)
AI Content Moderation in Therapy Conversations
by: Kim, Jiwon, et al.
Published: (2026)
by: Kim, Jiwon, et al.
Published: (2026)
Decision-Aware Trust Signal Alignment for SOC Alert Triage
by: Chowdhury, Israt Jahan, et al.
Published: (2026)
by: Chowdhury, Israt Jahan, et al.
Published: (2026)
NLP Privacy Risk Identification in Social Media (NLP-PRISM): A Survey
by: Goswami, Dhiman, et al.
Published: (2026)
by: Goswami, Dhiman, et al.
Published: (2026)
Privacy Leakage Overshadowed by Views of AI: A Study on Human Oversight of Privacy in Language Model Agent
by: Zhang, Zhiping, et al.
Published: (2024)
by: Zhang, Zhiping, et al.
Published: (2024)
Can Large Language Models Automate Phishing Warning Explanations? A Controlled Experiment on Effectiveness and User Perception
by: Cau, Federico Maria, et al.
Published: (2025)
by: Cau, Federico Maria, et al.
Published: (2025)
Leveraging Large Language Models for Collective Decision-Making
by: Papachristou, Marios, et al.
Published: (2023)
by: Papachristou, Marios, et al.
Published: (2023)
An Empirical Study of Privacy Leakage Chains via Prompt Injection in Black-Box Chatbot Environments
by: Yang, Hongjang, et al.
Published: (2026)
by: Yang, Hongjang, et al.
Published: (2026)
Toward a Human-Centered Evaluation Framework for Trustworthy LLM-Powered GUI Agents
by: Chen, Chaoran, et al.
Published: (2025)
by: Chen, Chaoran, et al.
Published: (2025)
Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming
by: Inie, Nanna, et al.
Published: (2023)
by: Inie, Nanna, et al.
Published: (2023)
Jaco: An Offline Running Privacy-aware Voice Assistant
by: Bermuth, Daniel, et al.
Published: (2022)
by: Bermuth, Daniel, et al.
Published: (2022)
The Obvious Invisible Threat: LLM-Powered GUI Agents' Vulnerability to Fine-Print Injections
by: Chen, Chaoran, et al.
Published: (2025)
by: Chen, Chaoran, et al.
Published: (2025)
Test Security in Remote Testing Age: Perspectives from Process Data Analytics and AI
by: Hao, Jiangang, et al.
Published: (2024)
by: Hao, Jiangang, et al.
Published: (2024)
Human-in-the-Loop Generation of Adversarial Texts: A Case Study on Tibetan Script
by: Cao, Xi, et al.
Published: (2024)
by: Cao, Xi, et al.
Published: (2024)
Leveraging Prompt-Based Large Language Models: Predicting Pandemic Health Decisions and Outcomes Through Social Media Language
by: Ding, Xiaohan, et al.
Published: (2024)
by: Ding, Xiaohan, et al.
Published: (2024)
Language Model Agents Under Attack: A Cross Model-Benchmark of Profit-Seeking Behaviors in Customer Service
by: Zhang, Jingyu
Published: (2025)
by: Zhang, Jingyu
Published: (2025)
On the Suitability of LLM-Driven Agents for Dark Pattern Audits
by: Sun, Chen, et al.
Published: (2026)
by: Sun, Chen, et al.
Published: (2026)
LLM Novice Uplift on Dual-Use, In Silico Biology Tasks
by: Zhang, Chen Bo Calvin, et al.
Published: (2026)
by: Zhang, Chen Bo Calvin, et al.
Published: (2026)
Assessing LLM Response Quality in the Context of Technology-Facilitated Abuse
by: Prakash, Vijay, et al.
Published: (2026)
by: Prakash, Vijay, et al.
Published: (2026)
LLM Content Moderation and User Satisfaction: Evidence from Response Refusals in Chatbot Arena
by: Pasch, Stefan
Published: (2025)
by: Pasch, Stefan
Published: (2025)
E-Vote Your Conscience: Perceptions of Coercion and Vote Buying, and the Usability of Fake Credentials in Online Voting
by: Merino, Louis-Henri, et al.
Published: (2024)
by: Merino, Louis-Henri, et al.
Published: (2024)
A Deep Dive into Fairness, Bias, Threats, and Privacy in Recommender Systems: Insights and Future Research
by: Roy, Falguni, et al.
Published: (2024)
by: Roy, Falguni, et al.
Published: (2024)
PrivateXR: Defending Privacy Attacks in Extended Reality Through Explainable AI-Guided Differential Privacy
by: Kundu, Ripan Kumar, et al.
Published: (2025)
by: Kundu, Ripan Kumar, et al.
Published: (2025)
Adversarial VR: An Open-Source Testbed for Evaluating Adversarial Robustness of VR Cybersickness Detection and Mitigation
by: Ahmed, Istiak, et al.
Published: (2025)
by: Ahmed, Istiak, et al.
Published: (2025)
Similar Items
-
Characterizing the MrDeepFakes Sexual Deepfake Marketplace
by: Han, Catherine, et al.
Published: (2024) -
Keys on Doormats: Exposed API Credentials on the Web
by: Demir, Nurullah, et al.
Published: (2026) -
Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis
by: Esposito, Matteo, et al.
Published: (2024) -
RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models
by: Wang, Jiongxiao, et al.
Published: (2023) -
JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
by: Feng, Yingchaojie, et al.
Published: (2024)