Securing Large Language Models: Addressing Bias, Misinformation, and Prompt Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Peng, Benji, Chen, Keyu, Li, Ming, Feng, Pohsun, Bi, Ziqian, Liu, Junyu, Song, Xinyuan, Niu, Qian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mastering AI: Big Data, Deep Learning, and the Evolution of Large Language Models -- Blockchain and Applications
by: Feng, Pohsun, et al.
Published: (2024)
by: Feng, Pohsun, et al.
Published: (2024)
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
by: Peng, Benji, et al.
Published: (2024)
by: Peng, Benji, et al.
Published: (2024)
Deep Learning Model Security: Threats and Defenses
by: Wang, Tianyang, et al.
Published: (2024)
by: Wang, Tianyang, et al.
Published: (2024)
SecPE: Secure Prompt Ensembling for Private and Robust Large Language Models
by: Zhang, Jiawen, et al.
Published: (2025)
by: Zhang, Jiawen, et al.
Published: (2025)
Counterfactual Explainable Incremental Prompt Attack Analysis on Large Language Models
by: Shu, Dong, et al.
Published: (2024)
by: Shu, Dong, et al.
Published: (2024)
System Prompt Extraction Attacks and Defenses in Large Language Models
by: Das, Badhan Chandra, et al.
Published: (2025)
by: Das, Badhan Chandra, et al.
Published: (2025)
ProxyPrompt: Securing System Prompts against Prompt Extraction Attacks
by: Zhuang, Zhixiong, et al.
Published: (2025)
by: Zhuang, Zhixiong, et al.
Published: (2025)
Prompt Inference Attack on Distributed Large Language Model Inference Frameworks
by: Luo, Xinjian, et al.
Published: (2025)
by: Luo, Xinjian, et al.
Published: (2025)
Prompt Inversion Attack against Collaborative Inference of Large Language Models
by: Qu, Wenjie, et al.
Published: (2025)
by: Qu, Wenjie, et al.
Published: (2025)
Securing Large Language Models (LLMs) from Prompt Injection Attacks
by: Suri, Omar Farooq Khan, et al.
Published: (2025)
by: Suri, Omar Farooq Khan, et al.
Published: (2025)
Improving Router Security using BERT
by: Carter, John, et al.
Published: (2026)
by: Carter, John, et al.
Published: (2026)
AttackEval: A Systematic Empirical Study of Prompt Injection Attack Effectiveness Against Large Language Models
by: Wang, Jackson
Published: (2026)
by: Wang, Jackson
Published: (2026)
Prompt Stealing Attacks Against Large Language Models
by: Sha, Zeyang, et al.
Published: (2024)
by: Sha, Zeyang, et al.
Published: (2024)
SoK: Taxonomy and Evaluation of Prompt Security in Large Language Models
by: Hong, Hanbin, et al.
Published: (2025)
by: Hong, Hanbin, et al.
Published: (2025)
Rotated Robustness: A Training-Free Defense against Bit-Flip Attacks on Large Language Models
by: Liu, Deng, et al.
Published: (2026)
by: Liu, Deng, et al.
Published: (2026)
Federated Domain-Specific Knowledge Transfer on Large Language Models Using Synthetic Data
by: Li, Haoran, et al.
Published: (2024)
by: Li, Haoran, et al.
Published: (2024)
Defense Against Prompt Injection Attack by Leveraging Attack Techniques
by: Chen, Yulin, et al.
Published: (2024)
by: Chen, Yulin, et al.
Published: (2024)
Chain-of-Thought Prompting of Large Language Models for Discovering and Fixing Software Vulnerabilities
by: Nong, Yu, et al.
Published: (2024)
by: Nong, Yu, et al.
Published: (2024)
TopicAttack: An Indirect Prompt Injection Attack via Topic Transition
by: Chen, Yulin, et al.
Published: (2025)
by: Chen, Yulin, et al.
Published: (2025)
Inhibitory Attacks on Backdoor-based Fingerprinting for Large Language Models
by: Fu, Hang, et al.
Published: (2026)
by: Fu, Hang, et al.
Published: (2026)
Membership Inference Attacks Against Video Large Language Models
by: Song, Wei, et al.
Published: (2026)
by: Song, Wei, et al.
Published: (2026)
A Novel Evaluation Framework for Assessing Resilience Against Prompt Injection Attacks in Large Language Models
by: Yip, Daniel Wankit, et al.
Published: (2024)
by: Yip, Daniel Wankit, et al.
Published: (2024)
QUIC-Exfil: Exploiting QUIC's Server Preferred Address Feature to Perform Data Exfiltration Attacks
by: Grübl, Thomas, et al.
Published: (2025)
by: Grübl, Thomas, et al.
Published: (2025)
Misleading Large Language Models used (or misused) in Scientific Peer-Reviewing via Hidden Prompt-Injection Attacks
by: Collu, Matteo Gioele, et al.
Published: (2025)
by: Collu, Matteo Gioele, et al.
Published: (2025)
Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security
by: Dai, Muzhi, et al.
Published: (2025)
by: Dai, Muzhi, et al.
Published: (2025)
Large Language Models for Cyber Security
by: Somani, Raunak, et al.
Published: (2025)
by: Somani, Raunak, et al.
Published: (2025)
Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods
by: Chen, Yulin, et al.
Published: (2025)
by: Chen, Yulin, et al.
Published: (2025)
Strengthening Polymorphic Prompt Assembling: Dynamic Separator Generation Against Emerging Prompt Injection Attacks
by: Dorzhiev, Nima, et al.
Published: (2026)
by: Dorzhiev, Nima, et al.
Published: (2026)
Unique Security and Privacy Threats of Large Language Models: A Comprehensive Survey
by: Wang, Shang, et al.
Published: (2024)
by: Wang, Shang, et al.
Published: (2024)
Can Indirect Prompt Injection Attacks Be Detected and Removed?
by: Chen, Yulin, et al.
Published: (2025)
by: Chen, Yulin, et al.
Published: (2025)
Automating Prompt Leakage Attacks on Large Language Models Using Agentic Approach
by: Sternak, Tvrtko, et al.
Published: (2025)
by: Sternak, Tvrtko, et al.
Published: (2025)
Pop Quiz Attack: Black-box Membership Inference Attacks Against Large Language Models
by: Chen, Zeyuan, et al.
Published: (2026)
by: Chen, Zeyuan, et al.
Published: (2026)
SEA: Low-Resource Safety Alignment for Multimodal Large Language Models via Synthetic Embeddings
by: Lu, Weikai, et al.
Published: (2025)
by: Lu, Weikai, et al.
Published: (2025)
Prompt Injection Attacks on Large Language Models in Oncology
by: Clusmann, Jan, et al.
Published: (2024)
by: Clusmann, Jan, et al.
Published: (2024)
Privacy in Large Language Models: Attacks, Defenses and Future Directions
by: Li, Haoran, et al.
Published: (2023)
by: Li, Haoran, et al.
Published: (2023)
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models
by: Hu, Xiaomeng, et al.
Published: (2024)
by: Hu, Xiaomeng, et al.
Published: (2024)
BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
by: Chen, Yulin, et al.
Published: (2024)
by: Chen, Yulin, et al.
Published: (2024)
Securing AI Agents Against Prompt Injection Attacks
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
Towards Robust Multimodal Large Language Models Against Jailbreak Attacks
by: Yin, Ziyi, et al.
Published: (2025)
by: Yin, Ziyi, et al.
Published: (2025)
System Prompt Poisoning: Persistent Attacks on Large Language Models Beyond User Injection
by: Li, Zongze, et al.
Published: (2025)
by: Li, Zongze, et al.
Published: (2025)
Similar Items
-
Mastering AI: Big Data, Deep Learning, and the Evolution of Large Language Models -- Blockchain and Applications
by: Feng, Pohsun, et al.
Published: (2024) -
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
by: Peng, Benji, et al.
Published: (2024) -
Deep Learning Model Security: Threats and Defenses
by: Wang, Tianyang, et al.
Published: (2024) -
SecPE: Secure Prompt Ensembling for Private and Robust Large Language Models
by: Zhang, Jiawen, et al.
Published: (2025) -
Counterfactual Explainable Incremental Prompt Attack Analysis on Large Language Models
by: Shu, Dong, et al.
Published: (2024)