How Susceptible are Large Language Models to Ideological Manipulation?
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Kai, He, Zihao, Yan, Jun, Shi, Taiwei, Lerman, Kristina |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models
por: Chen, Kai, et al.
Publicado: (2025)
por: Chen, Kai, et al.
Publicado: (2025)
Data Defenses Against Large Language Models
por: Agnew, William, et al.
Publicado: (2024)
por: Agnew, William, et al.
Publicado: (2024)
An Investigation into Misuse of Java Security APIs by Large Language Models
por: Mousavi, Zahra, et al.
Publicado: (2024)
por: Mousavi, Zahra, et al.
Publicado: (2024)
Attacks on Third-Party APIs of Large Language Models
por: Zhao, Wanru, et al.
Publicado: (2024)
por: Zhao, Wanru, et al.
Publicado: (2024)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
por: Murphy, Brendan, et al.
Publicado: (2025)
por: Murphy, Brendan, et al.
Publicado: (2025)
Bridging the Copyright Gap: Do Large Vision-Language Models Recognize and Respect Copyrighted Content?
por: Xu, Naen, et al.
Publicado: (2025)
por: Xu, Naen, et al.
Publicado: (2025)
Phare: A Safety Probe for Large Language Models
por: Jeune, Pierre Le, et al.
Publicado: (2025)
por: Jeune, Pierre Le, et al.
Publicado: (2025)
Beyond Context: Large Language Models' Failure to Grasp Users' Intent
por: Hussain, Ahmed M., et al.
Publicado: (2025)
por: Hussain, Ahmed M., et al.
Publicado: (2025)
TombRaider: Entering the Vault of History to Jailbreak Large Language Models
por: Ding, Junchen, et al.
Publicado: (2025)
por: Ding, Junchen, et al.
Publicado: (2025)
Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
por: An, Bang, et al.
Publicado: (2024)
por: An, Bang, et al.
Publicado: (2024)
SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models
por: Djuhera, Aladin, et al.
Publicado: (2025)
por: Djuhera, Aladin, et al.
Publicado: (2025)
Let's Measure the Elephant in the Room: Facilitating Personalized Automated Analysis of Privacy Policies at Scale
por: Zhao, Rui, et al.
Publicado: (2025)
por: Zhao, Rui, et al.
Publicado: (2025)
SimMark: A Robust Sentence-Level Similarity-Based Watermarking Algorithm for Large Language Models
por: Dabiriaghdam, Amirhossein, et al.
Publicado: (2025)
por: Dabiriaghdam, Amirhossein, et al.
Publicado: (2025)
Simulate and Eliminate: Revoke Backdoors for Generative Large Language Models
por: Li, Haoran, et al.
Publicado: (2024)
por: Li, Haoran, et al.
Publicado: (2024)
Reading Between the Tweets: Deciphering Ideological Stances of Interconnected Mixed-Ideology Communities
por: He, Zihao, et al.
Publicado: (2024)
por: He, Zihao, et al.
Publicado: (2024)
RealHarm: A Collection of Real-World Language Model Application Failures
por: Jeune, Pierre Le, et al.
Publicado: (2025)
por: Jeune, Pierre Le, et al.
Publicado: (2025)
Yet Another Watermark for Large Language Models
por: Bao, Siyuan, et al.
Publicado: (2025)
por: Bao, Siyuan, et al.
Publicado: (2025)
Domain-Independent Deception: A New Taxonomy and Linguistic Analysis
por: Verma, Rakesh M., et al.
Publicado: (2024)
por: Verma, Rakesh M., et al.
Publicado: (2024)
Get my drift? Catching LLM Task Drift with Activation Deltas
por: Abdelnabi, Sahar, et al.
Publicado: (2024)
por: Abdelnabi, Sahar, et al.
Publicado: (2024)
SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization
por: Liu, Houjun, et al.
Publicado: (2026)
por: Liu, Houjun, et al.
Publicado: (2026)
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
por: Gomaa, Amr, et al.
Publicado: (2025)
por: Gomaa, Amr, et al.
Publicado: (2025)
Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
por: Cyberey, Hannah, et al.
Publicado: (2025)
por: Cyberey, Hannah, et al.
Publicado: (2025)
AI Agents May Always Fall for Prompt Injections
por: Abdelnabi, Sahar, et al.
Publicado: (2026)
por: Abdelnabi, Sahar, et al.
Publicado: (2026)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
por: Zhou, Zhenhong, et al.
Publicado: (2024)
por: Zhou, Zhenhong, et al.
Publicado: (2024)
How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?
por: Li, Yuxuan, et al.
Publicado: (2026)
por: Li, Yuxuan, et al.
Publicado: (2026)
InfoFlood: Jailbreaking Large Language Models with Information Overload
por: Yadav, Advait, et al.
Publicado: (2025)
por: Yadav, Advait, et al.
Publicado: (2025)
Learning to Poison Large Language Models for Downstream Manipulation
por: Zhou, Xiangyu, et al.
Publicado: (2024)
por: Zhou, Xiangyu, et al.
Publicado: (2024)
Rethinking Backdoor Detection Evaluation for Language Models
por: Yan, Jun, et al.
Publicado: (2024)
por: Yan, Jun, et al.
Publicado: (2024)
Privacy in Large Language Models: Attacks, Defenses and Future Directions
por: Li, Haoran, et al.
Publicado: (2023)
por: Li, Haoran, et al.
Publicado: (2023)
Leveraging Large Language Models for Preliminary Security Risk Analysis: A Mission-Critical Case Study
por: Esposito, Matteo, et al.
Publicado: (2024)
por: Esposito, Matteo, et al.
Publicado: (2024)
Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models
por: Xu, Zihao, et al.
Publicado: (2024)
por: Xu, Zihao, et al.
Publicado: (2024)
Harnessing Task Overload for Scalable Jailbreak Attacks on Large Language Models
por: Dong, Yiting, et al.
Publicado: (2024)
por: Dong, Yiting, et al.
Publicado: (2024)
Large Language Models as a (Bad) Security Norm in the Context of Regulation and Compliance
por: Ludvigsen, Kaspar Rosager
Publicado: (2025)
por: Ludvigsen, Kaspar Rosager
Publicado: (2025)
$\texttt{ModSCAN}$: Measuring Stereotypical Bias in Large Vision-Language Models from Vision and Language Modalities
por: Jiang, Yukun, et al.
Publicado: (2024)
por: Jiang, Yukun, et al.
Publicado: (2024)
Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
por: Akiri, Charankumar, et al.
Publicado: (2025)
por: Akiri, Charankumar, et al.
Publicado: (2025)
Black-Box Opinion Manipulation Attacks to Retrieval-Augmented Generation of Large Language Models
por: Chen, Zhuo, et al.
Publicado: (2024)
por: Chen, Zhuo, et al.
Publicado: (2024)
Towards Label-Only Membership Inference Attack against Pre-trained Large Language Models
por: He, Yu, et al.
Publicado: (2025)
por: He, Yu, et al.
Publicado: (2025)
Leveraging Large Language Models for Sentiment Analysis: Multi-Modal Analysis of Decentraland's MANA Token
por: Wu, Xintong, et al.
Publicado: (2026)
por: Wu, Xintong, et al.
Publicado: (2026)
The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks
por: Chen, Xiaoyi, et al.
Publicado: (2023)
por: Chen, Xiaoyi, et al.
Publicado: (2023)
MM-AttacKG: A Multimodal Approach to Attack Graph Construction with Large Language Models
por: Zhang, Yongheng, et al.
Publicado: (2025)
por: Zhang, Yongheng, et al.
Publicado: (2025)
Ejemplares similares
-
STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models
por: Chen, Kai, et al.
Publicado: (2025) -
Data Defenses Against Large Language Models
por: Agnew, William, et al.
Publicado: (2024) -
An Investigation into Misuse of Java Security APIs by Large Language Models
por: Mousavi, Zahra, et al.
Publicado: (2024) -
Attacks on Third-Party APIs of Large Language Models
por: Zhao, Wanru, et al.
Publicado: (2024) -
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
por: Murphy, Brendan, et al.
Publicado: (2025)