Investigating and Alleviating Harm Amplification in LLM Interactions
Fuente:
arXiv
Salvato in:
| Autori principali: | Guo, Ruohao, Xu, Wei, Ritter, Alan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Meta-Tuning LLMs to Leverage Lexical Knowledge for Generalizable Language Style Understanding
di: Guo, Ruohao, et al.
Pubblicazione: (2023)
di: Guo, Ruohao, et al.
Pubblicazione: (2023)
How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit Misinformation
di: Guo, Ruohao, et al.
Pubblicazione: (2025)
di: Guo, Ruohao, et al.
Pubblicazione: (2025)
Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming Attacks
di: Guo, Ruohao, et al.
Pubblicazione: (2025)
di: Guo, Ruohao, et al.
Pubblicazione: (2025)
Tabular Data Understanding with LLMs: A Survey of Recent Advances and Challenges
di: Wu, Xiaofeng, et al.
Pubblicazione: (2025)
di: Wu, Xiaofeng, et al.
Pubblicazione: (2025)
Probabilistic Reasoning with LLMs for k-anonymity Estimation
di: Zheng, Jonathan, et al.
Pubblicazione: (2025)
di: Zheng, Jonathan, et al.
Pubblicazione: (2025)
Constrained Decoding for Cross-lingual Label Projection
di: Le, Duong Minh, et al.
Pubblicazione: (2024)
di: Le, Duong Minh, et al.
Pubblicazione: (2024)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
Distribution-Aware Reward: Reinforcement Learning over Predictive Distributions for LLM Regression
di: Park, Jungsoo, et al.
Pubblicazione: (2026)
di: Park, Jungsoo, et al.
Pubblicazione: (2026)
Language Models can Self-Improve at State-Value Estimation for Better Search
di: Mendes, Ethan, et al.
Pubblicazione: (2025)
di: Mendes, Ethan, et al.
Pubblicazione: (2025)
Having Beer after Prayer? Measuring Cultural Bias in Large Language Models
di: Naous, Tarek, et al.
Pubblicazione: (2023)
di: Naous, Tarek, et al.
Pubblicazione: (2023)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards
di: Chehbouni, Khaoula, et al.
Pubblicazione: (2024)
di: Chehbouni, Khaoula, et al.
Pubblicazione: (2024)
Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning
di: Feng, Weitao, et al.
Pubblicazione: (2025)
di: Feng, Weitao, et al.
Pubblicazione: (2025)
NJUST-KMG at TRAC-2024 Tasks 1 and 2: Offline Harm Potential Identification
di: Wang, Jingyuan, et al.
Pubblicazione: (2024)
di: Wang, Jingyuan, et al.
Pubblicazione: (2024)
Selective Neuron Amplification in Transformer Language Models
di: Akhtar, Ryyan, et al.
Pubblicazione: (2026)
di: Akhtar, Ryyan, et al.
Pubblicazione: (2026)
Anticipatory Evaluation of Language Models
di: Park, Jungsoo, et al.
Pubblicazione: (2025)
di: Park, Jungsoo, et al.
Pubblicazione: (2025)
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
di: Park, Jungsoo, et al.
Pubblicazione: (2025)
di: Park, Jungsoo, et al.
Pubblicazione: (2025)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
di: Chan, Yik Siu, et al.
Pubblicazione: (2025)
di: Chan, Yik Siu, et al.
Pubblicazione: (2025)
Stanceosaurus 2.0: Classifying Stance Towards Russian and Spanish Misinformation
di: Lavrouk, Anton, et al.
Pubblicazione: (2024)
di: Lavrouk, Anton, et al.
Pubblicazione: (2024)
LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey
di: Zou, Henry Peng, et al.
Pubblicazione: (2025)
di: Zou, Henry Peng, et al.
Pubblicazione: (2025)
Investigating Non-Transitivity in LLM-as-a-Judge
di: Xu, Yi, et al.
Pubblicazione: (2025)
di: Xu, Yi, et al.
Pubblicazione: (2025)
Bias Amplification in Language Model Evolution: An Iterated Learning Perspective
di: Ren, Yi, et al.
Pubblicazione: (2024)
di: Ren, Yi, et al.
Pubblicazione: (2024)
STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio
di: Park, Seong-Gyu, et al.
Pubblicazione: (2026)
di: Park, Seong-Gyu, et al.
Pubblicazione: (2026)
Representation Noising: A Defence Mechanism Against Harmful Finetuning
di: Rosati, Domenic, et al.
Pubblicazione: (2024)
di: Rosati, Domenic, et al.
Pubblicazione: (2024)
Harm Amplification in Text-to-Image Models
di: Hao, Susan, et al.
Pubblicazione: (2024)
di: Hao, Susan, et al.
Pubblicazione: (2024)
A Toolbox for Surfacing Health Equity Harms and Biases in Large Language Models
di: Pfohl, Stephen R., et al.
Pubblicazione: (2024)
di: Pfohl, Stephen R., et al.
Pubblicazione: (2024)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
di: Dobre, David, et al.
Pubblicazione: (2025)
di: Dobre, David, et al.
Pubblicazione: (2025)
HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models
di: Lee, Seanie, et al.
Pubblicazione: (2024)
di: Lee, Seanie, et al.
Pubblicazione: (2024)
Initial Investigation of LLM-Assisted Development of Rule-Based Clinical NLP System
di: Shi, Jianlin, et al.
Pubblicazione: (2025)
di: Shi, Jianlin, et al.
Pubblicazione: (2025)
"They are uncultured": Unveiling Covert Harms and Social Threats in LLM Generated Conversations
di: Dammu, Preetam Prabhu Srikar, et al.
Pubblicazione: (2024)
di: Dammu, Preetam Prabhu Srikar, et al.
Pubblicazione: (2024)
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
di: Mendu, Sai Krishna, et al.
Pubblicazione: (2025)
di: Mendu, Sai Krishna, et al.
Pubblicazione: (2025)
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
di: Yu, Yongcan, et al.
Pubblicazione: (2026)
di: Yu, Yongcan, et al.
Pubblicazione: (2026)
GMTRouter: Personalized LLM Router over Multi-turn User Interactions
di: Xie, Encheng, et al.
Pubblicazione: (2025)
di: Xie, Encheng, et al.
Pubblicazione: (2025)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
di: Sheshadri, Abhay, et al.
Pubblicazione: (2024)
di: Sheshadri, Abhay, et al.
Pubblicazione: (2024)
SIFiD: Reassess Summary Factual Inconsistency Detection with LLM
di: Yang, Jiuding, et al.
Pubblicazione: (2024)
di: Yang, Jiuding, et al.
Pubblicazione: (2024)
Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack
di: Huang, Tiansheng, et al.
Pubblicazione: (2024)
di: Huang, Tiansheng, et al.
Pubblicazione: (2024)
Can an LLM Induce a Graph? Investigating Memory Drift and Context Length
di: Yousuf, Raquib Bin, et al.
Pubblicazione: (2025)
di: Yousuf, Raquib Bin, et al.
Pubblicazione: (2025)
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge
di: Chan, Chi-Min, et al.
Pubblicazione: (2025)
di: Chan, Chi-Min, et al.
Pubblicazione: (2025)
Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation
di: Abdelnabi, Sahar, et al.
Pubblicazione: (2023)
di: Abdelnabi, Sahar, et al.
Pubblicazione: (2023)
Harmful Intent as a Geometrically Recoverable Feature of LLM Residual Streams
di: Llorente-Saguer, Isaac
Pubblicazione: (2026)
di: Llorente-Saguer, Isaac
Pubblicazione: (2026)
Documenti analoghi
-
Meta-Tuning LLMs to Leverage Lexical Knowledge for Generalizable Language Style Understanding
di: Guo, Ruohao, et al.
Pubblicazione: (2023) -
How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit Misinformation
di: Guo, Ruohao, et al.
Pubblicazione: (2025) -
Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming Attacks
di: Guo, Ruohao, et al.
Pubblicazione: (2025) -
Tabular Data Understanding with LLMs: A Survey of Recent Advances and Challenges
di: Wu, Xiaofeng, et al.
Pubblicazione: (2025) -
Probabilistic Reasoning with LLMs for k-anonymity Estimation
di: Zheng, Jonathan, et al.
Pubblicazione: (2025)