Can a large language model be a gaslighter?
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Wei, Zhu, Luyao, Song, Yang, Lin, Ruixi, Mao, Rui, You, Yang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Attack and defense techniques in large language models: A survey and new perspectives
di: Liao, Zhiyu, et al.
Pubblicazione: (2025)
di: Liao, Zhiyu, et al.
Pubblicazione: (2025)
An In-Depth Investigation of Data Collection in LLM App Ecosystems
di: Wu, Yuhao, et al.
Pubblicazione: (2024)
di: Wu, Yuhao, et al.
Pubblicazione: (2024)
Generative AI Security: Challenges and Countermeasures
di: Zhu, Banghua, et al.
Pubblicazione: (2024)
di: Zhu, Banghua, et al.
Pubblicazione: (2024)
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
di: Zhang, Andy K., et al.
Pubblicazione: (2024)
di: Zhang, Andy K., et al.
Pubblicazione: (2024)
LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI's ChatGPT Plugins
di: Iqbal, Umar, et al.
Pubblicazione: (2023)
di: Iqbal, Umar, et al.
Pubblicazione: (2023)
Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots
di: Shayegani, Erfan, et al.
Pubblicazione: (2025)
di: Shayegani, Erfan, et al.
Pubblicazione: (2025)
Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
di: Shayegani, Erfan, et al.
Pubblicazione: (2025)
di: Shayegani, Erfan, et al.
Pubblicazione: (2025)
Urania: Differentially Private Insights into AI Use
di: Liu, Daogao, et al.
Pubblicazione: (2025)
di: Liu, Daogao, et al.
Pubblicazione: (2025)
Clio: Privacy-Preserving Insights into Real-World AI Use
di: Tamkin, Alex, et al.
Pubblicazione: (2024)
di: Tamkin, Alex, et al.
Pubblicazione: (2024)
Superficial Safety Alignment Hypothesis
di: Li, Jianwei, et al.
Pubblicazione: (2024)
di: Li, Jianwei, et al.
Pubblicazione: (2024)
What Makes an Evaluation Useful? Common Pitfalls and Best Practices
di: Gekker, Gil, et al.
Pubblicazione: (2025)
di: Gekker, Gil, et al.
Pubblicazione: (2025)
IsolateGPT: An Execution Isolation Architecture for LLM-Based Agentic Systems
di: Wu, Yuhao, et al.
Pubblicazione: (2024)
di: Wu, Yuhao, et al.
Pubblicazione: (2024)
Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods
di: Jang, Yeonwoo, et al.
Pubblicazione: (2025)
di: Jang, Yeonwoo, et al.
Pubblicazione: (2025)
Privacy at a Price: Exploring its Dual Impact on AI Fairness
di: Yang, Mengmeng, et al.
Pubblicazione: (2024)
di: Yang, Mengmeng, et al.
Pubblicazione: (2024)
Optimizing watermarks for large language models
di: Wouters, Bram
Pubblicazione: (2023)
di: Wouters, Bram
Pubblicazione: (2023)
Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses
di: Zheng, Xiaosen, et al.
Pubblicazione: (2024)
di: Zheng, Xiaosen, et al.
Pubblicazione: (2024)
TaeBench: Improving Quality of Toxic Adversarial Examples
di: Zhu, Xuan, et al.
Pubblicazione: (2024)
di: Zhu, Xuan, et al.
Pubblicazione: (2024)
BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
di: Jiang, Fengqing, et al.
Pubblicazione: (2025)
di: Jiang, Fengqing, et al.
Pubblicazione: (2025)
The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation
di: Xu, Rongwu, et al.
Pubblicazione: (2023)
di: Xu, Rongwu, et al.
Pubblicazione: (2023)
Safety Alignment Can Be Not Superficial With Explicit Safety Signals
di: Li, Jianwei, et al.
Pubblicazione: (2025)
di: Li, Jianwei, et al.
Pubblicazione: (2025)
LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
di: Lin, Shi, et al.
Pubblicazione: (2024)
di: Lin, Shi, et al.
Pubblicazione: (2024)
Can LLMs Infer Conversational Agent Users' Personality Traits from Chat History?
di: Cögendez, Derya, et al.
Pubblicazione: (2026)
di: Cögendez, Derya, et al.
Pubblicazione: (2026)
A Public Theory of Distillation Resistance via Constraint-Coupled Reasoning Architectures
di: Wei, Peng, et al.
Pubblicazione: (2026)
di: Wei, Peng, et al.
Pubblicazione: (2026)
How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?
di: Li, Yuxuan, et al.
Pubblicazione: (2026)
di: Li, Yuxuan, et al.
Pubblicazione: (2026)
Improving Your Model Ranking on Chatbot Arena by Vote Rigging
di: Min, Rui, et al.
Pubblicazione: (2025)
di: Min, Rui, et al.
Pubblicazione: (2025)
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection
di: Hu, Xulin, et al.
Pubblicazione: (2026)
di: Hu, Xulin, et al.
Pubblicazione: (2026)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
di: Muhamed, Aashiq, et al.
Pubblicazione: (2025)
di: Muhamed, Aashiq, et al.
Pubblicazione: (2025)
Special Characters Attack: Toward Scalable Training Data Extraction From Large Language Models
di: Bai, Yang, et al.
Pubblicazione: (2024)
di: Bai, Yang, et al.
Pubblicazione: (2024)
SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains
di: Saiem, Bijoy Ahmed, et al.
Pubblicazione: (2024)
di: Saiem, Bijoy Ahmed, et al.
Pubblicazione: (2024)
HeavyWater and SimplexWater: Distortion-Free LLM Watermarks for Low-Entropy Next-Token Predictions
di: Tsur, Dor, et al.
Pubblicazione: (2025)
di: Tsur, Dor, et al.
Pubblicazione: (2025)
Detecting Training Data of Large Language Models via Expectation Maximization
di: Kim, Gyuwan, et al.
Pubblicazione: (2024)
di: Kim, Gyuwan, et al.
Pubblicazione: (2024)
Watermarking Should Be Treated as a Monitoring Primitive
di: Aremu, Toluwani, et al.
Pubblicazione: (2026)
di: Aremu, Toluwani, et al.
Pubblicazione: (2026)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
di: Chu, Junjie, et al.
Pubblicazione: (2024)
di: Chu, Junjie, et al.
Pubblicazione: (2024)
The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks
di: Chen, Xiaoyi, et al.
Pubblicazione: (2023)
di: Chen, Xiaoyi, et al.
Pubblicazione: (2023)
Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data
di: Baumgärtner, Tim, et al.
Pubblicazione: (2024)
di: Baumgärtner, Tim, et al.
Pubblicazione: (2024)
Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT Models
di: Chu, Junjie, et al.
Pubblicazione: (2024)
di: Chu, Junjie, et al.
Pubblicazione: (2024)
Learnable Privacy Neurons Localization in Language Models
di: Chen, Ruizhe, et al.
Pubblicazione: (2024)
di: Chen, Ruizhe, et al.
Pubblicazione: (2024)
Probing the Robustness of Large Language Models Safety to Latent Perturbations
di: Gu, Tianle, et al.
Pubblicazione: (2025)
di: Gu, Tianle, et al.
Pubblicazione: (2025)
AI Propaganda factories with language models
di: Olejnik, Lukasz
Pubblicazione: (2025)
di: Olejnik, Lukasz
Pubblicazione: (2025)
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
di: Yang, Xiaoxue, et al.
Pubblicazione: (2025)
di: Yang, Xiaoxue, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Attack and defense techniques in large language models: A survey and new perspectives
di: Liao, Zhiyu, et al.
Pubblicazione: (2025) -
An In-Depth Investigation of Data Collection in LLM App Ecosystems
di: Wu, Yuhao, et al.
Pubblicazione: (2024) -
Generative AI Security: Challenges and Countermeasures
di: Zhu, Banghua, et al.
Pubblicazione: (2024) -
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
di: Zhang, Andy K., et al.
Pubblicazione: (2024) -
LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI's ChatGPT Plugins
di: Iqbal, Umar, et al.
Pubblicazione: (2023)