ControlNET: A Firewall for RAG-based LLM System
Fuente:
arXiv
Guardado en:
| Autores principales: | Yao, Hongwei, Shi, Haoran, Chen, Yidou, Jiang, Yixin, Wang, Cong, Qin, Zhan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
por: Duan, Kaiwen, et al.
Publicado: (2025)
por: Duan, Kaiwen, et al.
Publicado: (2025)
FDINet: Protecting against DNN Model Extraction via Feature Distortion Index
por: Yao, Hongwei, et al.
Publicado: (2023)
por: Yao, Hongwei, et al.
Publicado: (2023)
AttackLLM: LLM-based Attack Pattern Generation for an Industrial Control System
por: Ahmed, Chuadhry Mujeeb
Publicado: (2025)
por: Ahmed, Chuadhry Mujeeb
Publicado: (2025)
Confundo: Learning to Generate Robust Poison for Practical RAG Systems
por: Hu, Haoyang, et al.
Publicado: (2026)
por: Hu, Haoyang, et al.
Publicado: (2026)
MalRAG: A Retrieval-Augmented LLM Framework for Open-set Malicious Traffic Identification
por: Luo, Xiang, et al.
Publicado: (2025)
por: Luo, Xiang, et al.
Publicado: (2025)
FIT-Print: Towards False-claim-resistant Model Ownership Verification via Targeted Fingerprint
por: Shao, Shuo, et al.
Publicado: (2025)
por: Shao, Shuo, et al.
Publicado: (2025)
Explanation as a Watermark: Towards Harmless and Multi-bit Model Ownership Verification via Watermarking Feature Attribution
por: Shao, Shuo, et al.
Publicado: (2024)
por: Shao, Shuo, et al.
Publicado: (2024)
From Firewalls to Frontiers: AI Red-Teaming is a Domain-Specific Evolution of Cyber Red-Teaming
por: Sinha, Anusha, et al.
Publicado: (2025)
por: Sinha, Anusha, et al.
Publicado: (2025)
HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns
por: Shen, Xinyue, et al.
Publicado: (2025)
por: Shen, Xinyue, et al.
Publicado: (2025)
RADAR: Defending RAG Dynamically against Retrieval Corruption
por: Chen, Ziyuan, et al.
Publicado: (2026)
por: Chen, Ziyuan, et al.
Publicado: (2026)
GraphRAG under Fire
por: Liang, Jiacheng, et al.
Publicado: (2025)
por: Liang, Jiacheng, et al.
Publicado: (2025)
Detection and Imputation based Two-Stage Denoising Diffusion Power System Measurement Recovery under Cyber-Physical Uncertainties
por: Pei, Jianhua, et al.
Publicado: (2023)
por: Pei, Jianhua, et al.
Publicado: (2023)
Adaptive Probe-based Steering for Robust LLM Jailbreaking
por: Chen, Junxi, et al.
Publicado: (2026)
por: Chen, Junxi, et al.
Publicado: (2026)
CleanBase: Detecting Malicious Documents in RAG Knowledge Databases
por: Jin, Weifei, et al.
Publicado: (2026)
por: Jin, Weifei, et al.
Publicado: (2026)
Bias Amplification in RAG: Poisoning Knowledge Retrieval to Steer LLMs
por: Wang, Linlin, et al.
Publicado: (2025)
por: Wang, Linlin, et al.
Publicado: (2025)
IstGPT: LLM-based Anomaly Detection for Spatial-Temporal Graph in Industrial Systems
por: Zhang, Yuchen, et al.
Publicado: (2026)
por: Zhang, Yuchen, et al.
Publicado: (2026)
PIDSMaker: Building and Evaluating Provenance-based Intrusion Detection Systems
por: Bilot, Tristan, et al.
Publicado: (2026)
por: Bilot, Tristan, et al.
Publicado: (2026)
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
por: Zhan, Qiusi, et al.
Publicado: (2025)
por: Zhan, Qiusi, et al.
Publicado: (2025)
Auditing Differential Privacy in the Black-Box Setting
por: Shi, Kaining, et al.
Publicado: (2025)
por: Shi, Kaining, et al.
Publicado: (2025)
Ward: Provable RAG Dataset Inference via LLM Watermarks
por: Jovanović, Nikola, et al.
Publicado: (2024)
por: Jovanović, Nikola, et al.
Publicado: (2024)
MERLOT: A Distilled LLM-based Mixture-of-Experts Framework for Scalable Encrypted Traffic Classification
por: Chen, Yuxuan, et al.
Publicado: (2024)
por: Chen, Yuxuan, et al.
Publicado: (2024)
PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models
por: Zou, Wei, et al.
Publicado: (2024)
por: Zou, Wei, et al.
Publicado: (2024)
Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors
por: Chen, Yen-Shan, et al.
Publicado: (2025)
por: Chen, Yen-Shan, et al.
Publicado: (2025)
Enhancing Privacy in ControlNet and Stable Diffusion via Split Learning
por: Yao, Dixi
Publicado: (2024)
por: Yao, Dixi
Publicado: (2024)
ACE: A Security Architecture for LLM-Integrated App Systems
por: Li, Evan, et al.
Publicado: (2025)
por: Li, Evan, et al.
Publicado: (2025)
MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers
por: Wang, Zhiqiang, et al.
Publicado: (2025)
por: Wang, Zhiqiang, et al.
Publicado: (2025)
On Benchmarking Code LLMs for Android Malware Analysis
por: He, Yiling, et al.
Publicado: (2025)
por: He, Yiling, et al.
Publicado: (2025)
CTFusion: A CTF-based Benchmark for LLM Agent Evaluation
por: Lee, Dongjun, et al.
Publicado: (2026)
por: Lee, Dongjun, et al.
Publicado: (2026)
A Systematic Study of Code Obfuscation Against LLM-based Vulnerability Detection
por: Li, Xiao, et al.
Publicado: (2025)
por: Li, Xiao, et al.
Publicado: (2025)
How to make Medical AI Systems safer? Simulating Vulnerabilities, and Threats in Multimodal Medical RAG System
por: Zuo, Kaiwen, et al.
Publicado: (2025)
por: Zuo, Kaiwen, et al.
Publicado: (2025)
Deep Learning-based Anomaly Detection and Log Analysis for Computer Networks
por: Wang, Shuzhan, et al.
Publicado: (2024)
por: Wang, Shuzhan, et al.
Publicado: (2024)
Adaptive Dual-Layer Web Application Firewall (ADL-WAF) Leveraging Machine Learning for Enhanced Anomaly and Threat Detection
por: Sameh, Ahmed, et al.
Publicado: (2025)
por: Sameh, Ahmed, et al.
Publicado: (2025)
Universal Graph Backdoor Defense: A Feature-based Homophily Perspective
por: Pan, Mengting, et al.
Publicado: (2026)
por: Pan, Mengting, et al.
Publicado: (2026)
Voice Jailbreak Attacks Against GPT-4o
por: Shen, Xinyue, et al.
Publicado: (2024)
por: Shen, Xinyue, et al.
Publicado: (2024)
Image-Perfect Imperfections: Safety, Bias, and Authenticity in the Shadow of Text-To-Image Model Evolution
por: Wu, Yixin, et al.
Publicado: (2024)
por: Wu, Yixin, et al.
Publicado: (2024)
G-Safeguard: A Topology-Guided Security Lens and Treatment on LLM-based Multi-agent Systems
por: Wang, Shilong, et al.
Publicado: (2025)
por: Wang, Shilong, et al.
Publicado: (2025)
Privacy-preserving Decision-focused Learning for Multi-energy Systems
por: Zhou, Yangze, et al.
Publicado: (2025)
por: Zhou, Yangze, et al.
Publicado: (2025)
Moss: Proxy Model-based Full-Weight Aggregation in Federated Learning with Heterogeneous Models
por: Cai, Yifeng, et al.
Publicado: (2025)
por: Cai, Yifeng, et al.
Publicado: (2025)
Empirical Perturbation Analysis of Linear System Solvers from a Data Poisoning Perspective
por: Liu, Yixin, et al.
Publicado: (2024)
por: Liu, Yixin, et al.
Publicado: (2024)
Verification of Machine Unlearning is Fragile
por: Zhang, Binchi, et al.
Publicado: (2024)
por: Zhang, Binchi, et al.
Publicado: (2024)
Ejemplares similares
-
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
por: Duan, Kaiwen, et al.
Publicado: (2025) -
FDINet: Protecting against DNN Model Extraction via Feature Distortion Index
por: Yao, Hongwei, et al.
Publicado: (2023) -
AttackLLM: LLM-based Attack Pattern Generation for an Industrial Control System
por: Ahmed, Chuadhry Mujeeb
Publicado: (2025) -
Confundo: Learning to Generate Robust Poison for Practical RAG Systems
por: Hu, Haoyang, et al.
Publicado: (2026) -
MalRAG: A Retrieval-Augmented LLM Framework for Open-set Malicious Traffic Identification
por: Luo, Xiang, et al.
Publicado: (2025)