Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Yihao, Bian, Haocheng, Zhou, Liutong, Wang, Ze, Zhang, Zhaoyi, Kawala, Francois, Dean, Milan, Fischer, Ian, Peng, Yuantao, Tokgozoglu, Noyan, Barrientos, Ivan, Shaik, Riyaaz, Li, Rachel, Venkataraman, Chandru, Far, Reza Shifteh, Pawar, Moses, Sundaranatha, Venkat, Xu, Michael, Chu, Frank |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lightweight Safety Guardrails via Synthetic Data and RL-guided Adversarial Training
by: Ilin, Aleksei, et al.
Published: (2025)
by: Ilin, Aleksei, et al.
Published: (2025)
FedGuard: A Diverse-Byzantine-Robust Mechanism for Federated Learning with Major Malicious Clients
by: Jiang, Haocheng, et al.
Published: (2025)
by: Jiang, Haocheng, et al.
Published: (2025)
Scaling Search Relevance: Augmenting App Store Ranking with LLM-Generated Judgments
by: Christakopoulou, Evangelia, et al.
Published: (2026)
by: Christakopoulou, Evangelia, et al.
Published: (2026)
Topological properties of the [110] SnTe nanowires
by: Kawala, Alicja, et al.
Published: (2024)
by: Kawala, Alicja, et al.
Published: (2024)
FUELVISION: A Multimodal Data Fusion and Multimodel Ensemble Algorithm for Wildfire Fuels Mapping
by: Shaik, Riyaaz Uddien, et al.
Published: (2024)
by: Shaik, Riyaaz Uddien, et al.
Published: (2024)
Incidental interictal epileptiform discharges in infants with nonepileptic events
by: Maria A. Montenegro, et al.
Published: (2025)
by: Maria A. Montenegro, et al.
Published: (2025)
GuardDoor: Safeguarding Against Malicious Diffusion Editing via Protective Backdoors
by: Zeng, Yaopei, et al.
Published: (2025)
by: Zeng, Yaopei, et al.
Published: (2025)
CrossGuard: Safeguarding MLLMs against Joint-Modal Implicit Malicious Attacks
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
GCP: Guarded Collaborative Perception with Spatial-Temporal Aware Malicious Agent Detection
by: Tao, Yihang, et al.
Published: (2025)
by: Tao, Yihang, et al.
Published: (2025)
Scaling Inference-Efficient Language Models
by: Bian, Song, et al.
Published: (2025)
by: Bian, Song, et al.
Published: (2025)
Self‐Rostering for Emergency Career Medical Officers ( CMOs ) and Registrars Within a Small Metropolitan Emergency Department: A Mixed Methods Study on Employee Satisfaction and Implementation Processes
by: Khanh Nguyen, et al.
Published: (2026)
by: Khanh Nguyen, et al.
Published: (2026)
Response: “Incidental interictal epileptiform discharges in infants with nonepileptic events”
by: Maria A. Montenegro, et al.
Published: (2025)
by: Maria A. Montenegro, et al.
Published: (2025)
CP-Guard: Malicious Agent Detection and Defense in Collaborative Bird's Eye View Perception
by: Hu, Senkang, et al.
Published: (2024)
by: Hu, Senkang, et al.
Published: (2024)
CP-Guard+: A New Paradigm for Malicious Agent Detection and Defense in Collaborative Perception
by: Hu, Senkang, et al.
Published: (2025)
by: Hu, Senkang, et al.
Published: (2025)
DiffusionGuard: A Robust Defense Against Malicious Diffusion-based Image Editing
by: Choi, June Suk, et al.
Published: (2024)
by: Choi, June Suk, et al.
Published: (2024)
Detecting Malicious Intents in Smart Contracts with Pre-trained Programming Language Models
by: Huang, Youwei, et al.
Published: (2025)
by: Huang, Youwei, et al.
Published: (2025)
SynthGuard: An Open Platform for Detecting AI-Generated Multimedia with Multimodal LLMs
by: Desai, Shail, et al.
Published: (2025)
by: Desai, Shail, et al.
Published: (2025)
VecIntrinBench: Benchmarking Cross-Architecture Intrinsic Code Migration for RISC-V Vector
by: Han, Liutong, et al.
Published: (2025)
by: Han, Liutong, et al.
Published: (2025)
A Method for Efficient Heterogeneous Parallel Compilation: A Cryptography Case Study
by: Tan, Zhiyuan, et al.
Published: (2024)
by: Tan, Zhiyuan, et al.
Published: (2024)
Mechanism of the Influence of Interlayer Mechanical Properties on the Ballistic Performance of Transparent Armor
by: Jian Zhang, et al.
Published: (2026)
by: Jian Zhang, et al.
Published: (2026)
MalGuard: Towards Real-Time, Accurate, and Actionable Detection of Malicious Packages in PyPI Ecosystem
by: Gao, Xingan, et al.
Published: (2025)
by: Gao, Xingan, et al.
Published: (2025)
GaussReg: Fast 3D Registration with Gaussian Splatting
by: Chang, Jiahao, et al.
Published: (2024)
by: Chang, Jiahao, et al.
Published: (2024)
Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
by: Bian, Song, et al.
Published: (2025)
by: Bian, Song, et al.
Published: (2025)
PersGuard: Preventing Malicious Personalization via Backdoor Attacks on Pre-trained Text-to-Image Diffusion Models
by: Liu, Xinwei, et al.
Published: (2025)
by: Liu, Xinwei, et al.
Published: (2025)
Guarding Against Malicious Biased Threats (GAMBiT): Experimental Design of Cognitive Sensors and Triggers with Behavioral Impact Analysis
by: Beltz, Brandon, et al.
Published: (2025)
by: Beltz, Brandon, et al.
Published: (2025)
INFA-Guard: Mitigating Malicious Propagation via Infection-Aware Safeguarding in LLM-Based Multi-Agent Systems
by: Zhou, Yijin, et al.
Published: (2026)
by: Zhou, Yijin, et al.
Published: (2026)
Can LLMs Deeply Detect Complex Malicious Queries? A Framework for Jailbreaking via Obfuscating Intent
by: Shang, Shang, et al.
Published: (2024)
by: Shang, Shang, et al.
Published: (2024)
Opioid Prescribing Patterns and Characteristics of Hand Injury Presentations to the Emergency Department of a Surgical Referral Centre in Western Sydney
by: Marco Stocca, et al.
Published: (2025)
by: Marco Stocca, et al.
Published: (2025)
A comparison between laboratory‐scale and large‐scale high‐intensity washing of flexible polyethylene packaging waste
by: Ezgi Ceren Boz Noyan, et al.
Published: (2024)
by: Ezgi Ceren Boz Noyan, et al.
Published: (2024)
Modeling User Preferences via Brain-Computer Interfacing
by: Leiva, Luis A., et al.
Published: (2024)
by: Leiva, Luis A., et al.
Published: (2024)
WebGuard++:Interpretable Malicious URL Detection via Bidirectional Fusion of HTML Subgraphs and Multi-Scale Convolutional BERT
by: Tian, Ye, et al.
Published: (2025)
by: Tian, Ye, et al.
Published: (2025)
RegGuard: AI-Powered Retrieval-Enhanced Assistant for Pharmaceutical Regulatory Compliance
by: Yang, Siyuan, et al.
Published: (2026)
by: Yang, Siyuan, et al.
Published: (2026)
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
by: Shen, Xinjie, et al.
Published: (2026)
by: Shen, Xinjie, et al.
Published: (2026)
Majestosas palmeiras: natureza tropical e imaginário no contexto do imperialismo europeu do século XIX
by: Alessandra El Far
Published: (2025)
by: Alessandra El Far
Published: (2025)
Środowiskowa krzywa Kuznetsa jako rama wzajemności świadczeń w polityce prawa
by: Artur Nowak-Far
Published: (2017)
by: Artur Nowak-Far
Published: (2017)
Miejscowości uzdrowiskowe w Austrii, Czechach, Niemczech i na Słowacji: status prawny i regulacyjne determinanty funkcjonowania
by: Artur Nowak-Far
Published: (2018)
by: Artur Nowak-Far
Published: (2018)
Prawnomiędzynarodowe uwarunkowania ochrony krajobrazu w Unii Europejskiej i w jej państwach członkowskich
by: Artur Nowak-Far
Published: (2017)
by: Artur Nowak-Far
Published: (2017)
Krzywa Kuznetsa a wielość jurysdykcji fiskalnych
by: Artur Nowak‑Far
Published: (2014)
by: Artur Nowak‑Far
Published: (2014)
Bilhetes de namoro abertos ao público: mensagens e encontros às escondidas anunciados no Jornal do Commercio (década de 1870)
by: Alessandra El Far
Published: (2017)
by: Alessandra El Far
Published: (2017)
Relación del consumo de alcohol y drogas de los jóvenes españoles con la siniestralidad vial durante la vida recreativa nocturna en tres comunidades autónomas en 2007
by: Amador Calafat Far
Published: (2008)
by: Amador Calafat Far
Published: (2008)
Similar Items
-
Lightweight Safety Guardrails via Synthetic Data and RL-guided Adversarial Training
by: Ilin, Aleksei, et al.
Published: (2025) -
FedGuard: A Diverse-Byzantine-Robust Mechanism for Federated Learning with Major Malicious Clients
by: Jiang, Haocheng, et al.
Published: (2025) -
Scaling Search Relevance: Augmenting App Store Ranking with LLM-Generated Judgments
by: Christakopoulou, Evangelia, et al.
Published: (2026) -
Topological properties of the [110] SnTe nanowires
by: Kawala, Alicja, et al.
Published: (2024) -
FUELVISION: A Multimodal Data Fusion and Multimodel Ensemble Algorithm for Wildfire Fuels Mapping
by: Shaik, Riyaaz Uddien, et al.
Published: (2024)