Why Do Large Language Models Generate Harmful Content?
Fuente:
arXiv
Guardado en:
| Autores principales: | Ganguli, Rajesh, Moraffah, Raha |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Generative Approach to Surrogate-based Black-box Attacks
por: Moraffah, Raha, et al.
Publicado: (2024)
por: Moraffah, Raha, et al.
Publicado: (2024)
Can Large Language Models Infer Causal Relationships from Real-World Text?
por: Saklad, Ryan, et al.
Publicado: (2025)
por: Saklad, Ryan, et al.
Publicado: (2025)
Adversarial Text Purification: A Large Language Model Approach for Defense
por: Moraffah, Raha, et al.
Publicado: (2024)
por: Moraffah, Raha, et al.
Publicado: (2024)
Exploiting Class Probabilities for Black-box Sentence-level Attacks
por: Moraffah, Raha, et al.
Publicado: (2024)
por: Moraffah, Raha, et al.
Publicado: (2024)
Zero-shot LLM-guided Counterfactual Generation: A Case Study on NLP Model Evaluation
por: Bhattacharjee, Amrita, et al.
Publicado: (2024)
por: Bhattacharjee, Amrita, et al.
Publicado: (2024)
EAGLE: A Domain Generalization Framework for AI-generated Text Detection
por: Bhattacharjee, Amrita, et al.
Publicado: (2024)
por: Bhattacharjee, Amrita, et al.
Publicado: (2024)
Causal Feature Selection for Responsible Machine Learning
por: Moraffah, Raha, et al.
Publicado: (2024)
por: Moraffah, Raha, et al.
Publicado: (2024)
Towards LLM-guided Causal Explainability for Black-box Text Classifiers
por: Bhattacharjee, Amrita, et al.
Publicado: (2023)
por: Bhattacharjee, Amrita, et al.
Publicado: (2023)
Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content
por: Bianchi, Federico, et al.
Publicado: (2024)
por: Bianchi, Federico, et al.
Publicado: (2024)
"Glue pizza and eat rocks" -- Exploiting Vulnerabilities in Retrieval-Augmented Generative Models
por: Tan, Zhen, et al.
Publicado: (2024)
por: Tan, Zhen, et al.
Publicado: (2024)
Prefix Probing: Lightweight Harmful Content Detection for Large Language Models
por: Yang, Jirui, et al.
Publicado: (2025)
por: Yang, Jirui, et al.
Publicado: (2025)
Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
por: Orgad, Hadas, et al.
Publicado: (2026)
por: Orgad, Hadas, et al.
Publicado: (2026)
DAGverse: Building Document-Grounded Semantic DAGs from Scientific Papers
por: Wan, Shu, et al.
Publicado: (2026)
por: Wan, Shu, et al.
Publicado: (2026)
A Survey of AI-generated Text Forensic Systems: Detection, Attribution, and Characterization
por: Kumarage, Tharindu, et al.
Publicado: (2024)
por: Kumarage, Tharindu, et al.
Publicado: (2024)
Self-HarmLLM: Can Large Language Model Harm Itself?
por: Kim, Heehwan, et al.
Publicado: (2025)
por: Kim, Heehwan, et al.
Publicado: (2025)
Advancing Harmful Content Detection in Organizational Research: Integrating Large Language Models with Elo Rating System
por: Akben, Mustafa, et al.
Publicado: (2025)
por: Akben, Mustafa, et al.
Publicado: (2025)
Do Large Language Models Need a Content Delivery Network?
por: Cheng, Yihua, et al.
Publicado: (2024)
por: Cheng, Yihua, et al.
Publicado: (2024)
Why Do Language Model Agents Whistleblow?
por: Agrawal, Kushal, et al.
Publicado: (2025)
por: Agrawal, Kushal, et al.
Publicado: (2025)
AI Meets the Classroom: When Do Large Language Models Harm Learning?
por: Lehmann, Matthias, et al.
Publicado: (2024)
por: Lehmann, Matthias, et al.
Publicado: (2024)
Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation
por: Huang, Tiansheng, et al.
Publicado: (2024)
por: Huang, Tiansheng, et al.
Publicado: (2024)
The Wolf Within: Covert Injection of Malice into MLLM Societies via an MLLM Operative
por: Tan, Zhen, et al.
Publicado: (2024)
por: Tan, Zhen, et al.
Publicado: (2024)
Engagement-Driven Content Generation with Large Language Models
por: Coppolillo, Erica, et al.
Publicado: (2024)
por: Coppolillo, Erica, et al.
Publicado: (2024)
Surgery: Mitigating Harmful Fine-Tuning for Large Language Models via Attention Sink
por: Liu, Guozhi, et al.
Publicado: (2026)
por: Liu, Guozhi, et al.
Publicado: (2026)
Bias of AI-Generated Content: An Examination of News Produced by Large Language Models
por: Fang, Xiao, et al.
Publicado: (2023)
por: Fang, Xiao, et al.
Publicado: (2023)
When Harmful Content Gets Camouflaged: Unveiling Perception Failure of LVLMs with CamHarmTI
por: Li, Yanhui, et al.
Publicado: (2025)
por: Li, Yanhui, et al.
Publicado: (2025)
Guiding Large Language Models to Generate Computer-Parsable Content
por: Wang, Jiaye
Publicado: (2024)
por: Wang, Jiaye
Publicado: (2024)
Re-ranking Using Large Language Models for Mitigating Exposure to Harmful Content on Social Media Platforms
por: Oak, Rajvardhan, et al.
Publicado: (2025)
por: Oak, Rajvardhan, et al.
Publicado: (2025)
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
por: Jiang, Yukun, et al.
Publicado: (2026)
por: Jiang, Yukun, et al.
Publicado: (2026)
TFGN: Task-Free, Replay-Free Continual Pre-Training Without Catastrophic Forgetting at LLM Scale
por: Ganguli, Anurup
Publicado: (2026)
por: Ganguli, Anurup
Publicado: (2026)
ChatPCG: Large Language Model-Driven Reward Design for Procedural Content Generation
por: Baek, In-Chang, et al.
Publicado: (2024)
por: Baek, In-Chang, et al.
Publicado: (2024)
Evaluating Language Models for Harmful Manipulation
por: Akbulut, Canfer, et al.
Publicado: (2026)
por: Akbulut, Canfer, et al.
Publicado: (2026)
Safe2Harm: Semantic Isomorphism Attacks for Jailbreaking Large Language Models
por: Yang, Fan
Publicado: (2025)
por: Yang, Fan
Publicado: (2025)
LLM Safety From Within: Detecting Harmful Content with Internal Representations
por: Jiao, Difan, et al.
Publicado: (2026)
por: Jiao, Difan, et al.
Publicado: (2026)
PCGRLLM: Large Language Model-Driven Reward Design for Procedural Content Generation Reinforcement Learning
por: Baek, In-Chang, et al.
Publicado: (2025)
por: Baek, In-Chang, et al.
Publicado: (2025)
Why Larger Language Models Do In-context Learning Differently?
por: Shi, Zhenmei, et al.
Publicado: (2024)
por: Shi, Zhenmei, et al.
Publicado: (2024)
A Deep Dive Into Large Language Model Code Generation Mistakes: What and Why?
por: Chen, QiHong, et al.
Publicado: (2024)
por: Chen, QiHong, et al.
Publicado: (2024)
AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on Harmfulness
por: Chen, Zixin, et al.
Publicado: (2025)
por: Chen, Zixin, et al.
Publicado: (2025)
Moderating Harm: Benchmarking Large Language Models for Cyberbullying Detection in YouTube Comments
por: Muminovic, Amel
Publicado: (2025)
por: Muminovic, Amel
Publicado: (2025)
Code Red! On the Harmfulness of Applying Off-the-shelf Large Language Models to Programming Tasks
por: Al-Kaswan, Ali, et al.
Publicado: (2025)
por: Al-Kaswan, Ali, et al.
Publicado: (2025)
VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples
por: Sun, Qixin, et al.
Publicado: (2025)
por: Sun, Qixin, et al.
Publicado: (2025)
Ejemplares similares
-
A Generative Approach to Surrogate-based Black-box Attacks
por: Moraffah, Raha, et al.
Publicado: (2024) -
Can Large Language Models Infer Causal Relationships from Real-World Text?
por: Saklad, Ryan, et al.
Publicado: (2025) -
Adversarial Text Purification: A Large Language Model Approach for Defense
por: Moraffah, Raha, et al.
Publicado: (2024) -
Exploiting Class Probabilities for Black-box Sentence-level Attacks
por: Moraffah, Raha, et al.
Publicado: (2024) -
Zero-shot LLM-guided Counterfactual Generation: A Case Study on NLP Model Evaluation
por: Bhattacharjee, Amrita, et al.
Publicado: (2024)