CLASP: Defending Hybrid Large Language Models Against Hidden State Poisoning Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Mercier, Alexandre Le, Demeester, Thomas, Develder, Chris |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hidden State Poisoning Attacks against Mamba-based Language Models
by: Mercier, Alexandre Le, et al.
Published: (2026)
by: Mercier, Alexandre Le, et al.
Published: (2026)
GAMBIT: A Three-Mode Benchmark for Adversarial Robustness in Multi-Agent LLM Collectives
by: Mercier, Alexandre Le, et al.
Published: (2026)
by: Mercier, Alexandre Le, et al.
Published: (2026)
Single- vs. Dual-Prompt Dialogue Generation with LLMs for Job Interviews in Human Resources
by: De Baer, Joachim, et al.
Published: (2025)
by: De Baer, Joachim, et al.
Published: (2025)
SkillMatch: Evaluating Self-supervised Learning of Skill Relatedness
by: Decorte, Jens-Joris, et al.
Published: (2024)
by: Decorte, Jens-Joris, et al.
Published: (2024)
Efficient Text Encoders for Labor Market Analysis
by: Decorte, Jens-Joris, et al.
Published: (2025)
by: Decorte, Jens-Joris, et al.
Published: (2025)
On the Biased Assessment of Expert Finding Systems
by: Decorte, Jens-Joris, et al.
Published: (2024)
by: Decorte, Jens-Joris, et al.
Published: (2024)
Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
by: Zhang, Zhexin, et al.
Published: (2023)
by: Zhang, Zhexin, et al.
Published: (2023)
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
by: Yi, Jingwei, et al.
Published: (2023)
by: Yi, Jingwei, et al.
Published: (2023)
In-Context Learning for Extreme Multi-Label Classification
by: D'Oosterlinck, Karel, et al.
Published: (2024)
by: D'Oosterlinck, Karel, et al.
Published: (2024)
SPML: A DSL for Defending Language Models Against Prompt Attacks
by: Sharma, Reshabh K, et al.
Published: (2024)
by: Sharma, Reshabh K, et al.
Published: (2024)
Defending Large Language Models Against Jailbreak Attacks via In-Decoding Safety-Awareness Probing
by: Zhao, Yinzhi, et al.
Published: (2026)
by: Zhao, Yinzhi, et al.
Published: (2026)
Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks
by: Zhao, Jiawei, et al.
Published: (2024)
by: Zhao, Jiawei, et al.
Published: (2024)
HSF: Defending against Jailbreak Attacks with Hidden State Filtering
by: Qian, Cheng, et al.
Published: (2024)
by: Qian, Cheng, et al.
Published: (2024)
Defending Against Social Engineering Attacks in the Age of LLMs
by: Ai, Lin, et al.
Published: (2024)
by: Ai, Lin, et al.
Published: (2024)
Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks
by: Zhou, Andy, et al.
Published: (2024)
by: Zhou, Andy, et al.
Published: (2024)
Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models
by: Gao, Lang, et al.
Published: (2024)
by: Gao, Lang, et al.
Published: (2024)
Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing
by: Ji, Jiabao, et al.
Published: (2024)
by: Ji, Jiabao, et al.
Published: (2024)
HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States
by: Jiang, Yilei, et al.
Published: (2025)
by: Jiang, Yilei, et al.
Published: (2025)
SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models
by: Afane, Mohamed, et al.
Published: (2025)
by: Afane, Mohamed, et al.
Published: (2025)
Denial-of-Service Poisoning Attacks against Large Language Models
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in Alignment
by: D'Oosterlinck, Karel, et al.
Published: (2024)
by: D'Oosterlinck, Karel, et al.
Published: (2024)
Defending Against Disinformation Attacks in Open-Domain Question Answering
by: Weller, Orion, et al.
Published: (2022)
by: Weller, Orion, et al.
Published: (2022)
Defending Against Indirect Prompt Injection Attacks With Spotlighting
by: Hines, Keegan, et al.
Published: (2024)
by: Hines, Keegan, et al.
Published: (2024)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
by: Liu, Fan, et al.
Published: (2024)
by: Liu, Fan, et al.
Published: (2024)
Robustness of Large Language Models Against Adversarial Attacks
by: Tao, Yiyi, et al.
Published: (2024)
by: Tao, Yiyi, et al.
Published: (2024)
DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts
by: Xin, Yuan, et al.
Published: (2026)
by: Xin, Yuan, et al.
Published: (2026)
Defending LLM Watermarking Against Spoofing Attacks with Contrastive Representation Learning
by: An, Li, et al.
Published: (2025)
by: An, Li, et al.
Published: (2025)
Benchmarking Gaslighting Attacks Against Speech Large Language Models
by: Wu, Jinyang, et al.
Published: (2025)
by: Wu, Jinyang, et al.
Published: (2025)
Defending Against Poisoning Attacks in Federated Learning with Blockchain
by: Dong, Nanqing, et al.
Published: (2023)
by: Dong, Nanqing, et al.
Published: (2023)
Prompt Stealing Attacks Against Large Language Models
by: Sha, Zeyang, et al.
Published: (2024)
by: Sha, Zeyang, et al.
Published: (2024)
A Customer Journey in the Land of Oz: Leveraging the Wizard of Oz Technique to Model Emotions in Customer Service Interactions
by: Labat, Sofie, et al.
Published: (2025)
by: Labat, Sofie, et al.
Published: (2025)
Reasoning-to-Defend: Safety-Aware Reasoning Can Defend Large Language Models from Jailbreaking
by: Zhu, Junda, et al.
Published: (2025)
by: Zhu, Junda, et al.
Published: (2025)
Benchmarking Gaslighting Negation Attacks Against Multimodal Large Language Models
by: Zhu, Bin, et al.
Published: (2025)
by: Zhu, Bin, et al.
Published: (2025)
CtrlRAG: Black-box Document Poisoning Attacks for Retrieval-Augmented Generation of Large Language Models
by: Sui, Runqi
Published: (2025)
by: Sui, Runqi
Published: (2025)
MemPot: Defending Against Memory Extraction Attack with Optimized Honeypots
by: Wang, Yuhao, et al.
Published: (2026)
by: Wang, Yuhao, et al.
Published: (2026)
Composite Backdoor Attacks Against Large Language Models
by: Huang, Hai, et al.
Published: (2023)
by: Huang, Hai, et al.
Published: (2023)
Prompt Injection Attacks in Defended Systems
by: Khomsky, Daniil, et al.
Published: (2024)
by: Khomsky, Daniil, et al.
Published: (2024)
RobustKV: Defending Large Language Models against Jailbreak Attacks via KV Eviction
by: Jiang, Tanqiu, et al.
Published: (2024)
by: Jiang, Tanqiu, et al.
Published: (2024)
Similar Items
-
Hidden State Poisoning Attacks against Mamba-based Language Models
by: Mercier, Alexandre Le, et al.
Published: (2026) -
GAMBIT: A Three-Mode Benchmark for Adversarial Robustness in Multi-Agent LLM Collectives
by: Mercier, Alexandre Le, et al.
Published: (2026) -
Single- vs. Dual-Prompt Dialogue Generation with LLMs for Job Interviews in Human Resources
by: De Baer, Joachim, et al.
Published: (2025) -
SkillMatch: Evaluating Self-supervised Learning of Skill Relatedness
by: Decorte, Jens-Joris, et al.
Published: (2024) -
Efficient Text Encoders for Labor Market Analysis
by: Decorte, Jens-Joris, et al.
Published: (2025)