Benchmarking Safety Risks of Knowledge-Intensive Reasoning under Malicious Knowledge Editing
Fuente:
arXiv
Saved in:
| Main Authors: | Mao, Qinghua, Lin, Xi, Gu, Jinze, Wu, Jun, Li, Siyuan, Chen, Yuliang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GraphSteal: Structural Knowledge Stealing from Graph RAG via Traversal Reconstruction
by: Gu, Jinze, et al.
Published: (2026)
by: Gu, Jinze, et al.
Published: (2026)
CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks
by: Li, Siyuan, et al.
Published: (2026)
by: Li, Siyuan, et al.
Published: (2026)
Analysing Safety Risks in LLMs Fine-Tuned with Pseudo-Malicious Cyber Security Data
by: ElZemity, Adel, et al.
Published: (2025)
by: ElZemity, Adel, et al.
Published: (2025)
Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
by: Fu, Haowei, et al.
Published: (2025)
by: Fu, Haowei, et al.
Published: (2025)
CTIArena: Benchmarking LLM Knowledge and Reasoning Across Heterogeneous Cyber Threat Intelligence
by: Cheng, Yutong, et al.
Published: (2025)
by: Cheng, Yutong, et al.
Published: (2025)
SafeThinker: Reasoning about Risk to Deepen Safety Beyond Shallow Alignment
by: Fang, Xianya, et al.
Published: (2026)
by: Fang, Xianya, et al.
Published: (2026)
HoneyTrap: Deceiving Large Language Model Attackers to Honeypot Traps with Resilient Multi-Agent Defense
by: Li, Siyuan, et al.
Published: (2026)
by: Li, Siyuan, et al.
Published: (2026)
SDD: Self-Degraded Defense against Malicious Fine-tuning
by: Chen, Zixuan, et al.
Published: (2025)
by: Chen, Zixuan, et al.
Published: (2025)
Systematization of Knowledge: Security and Safety in the Model Context Protocol Ecosystem
by: Gaire, Shiva, et al.
Published: (2025)
by: Gaire, Shiva, et al.
Published: (2025)
AttacKG+:Boosting Attack Knowledge Graph Construction with Large Language Models
by: Zhang, Yongheng, et al.
Published: (2024)
by: Zhang, Yongheng, et al.
Published: (2024)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
by: Tu, Shangqing, et al.
Published: (2024)
by: Tu, Shangqing, et al.
Published: (2024)
LinkThief: Combining Generalized Structure Knowledge with Node Similarity for Link Stealing Attack against GNN
by: Zhang, Yuxing, et al.
Published: (2024)
by: Zhang, Yuxing, et al.
Published: (2024)
Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent Reasoning
by: Kereopa-Yorke, Ben, et al.
Published: (2026)
by: Kereopa-Yorke, Ben, et al.
Published: (2026)
Step-by-Step Reasoning Attack: Revealing 'Erased' Knowledge in Large Language Models
by: Sinha, Yash, et al.
Published: (2025)
by: Sinha, Yash, et al.
Published: (2025)
Malla: Demystifying Real-world Large Language Model Integrated Malicious Services
by: Lin, Zilong, et al.
Published: (2024)
by: Lin, Zilong, et al.
Published: (2024)
KnowledgeSG: Privacy-Preserving Synthetic Text Generation with Knowledge Distillation from Server
by: Wang, Wenhao, et al.
Published: (2024)
by: Wang, Wenhao, et al.
Published: (2024)
Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?
by: Yin, Qingyu, et al.
Published: (2025)
by: Yin, Qingyu, et al.
Published: (2025)
Invisible Prompts, Visible Threats: Malicious Font Injection in External Resources for Large Language Models
by: Xiong, Junjie, et al.
Published: (2025)
by: Xiong, Junjie, et al.
Published: (2025)
Reimagining Safety Alignment with An Image
by: Xia, Yifan, et al.
Published: (2025)
by: Xia, Yifan, et al.
Published: (2025)
Untargeted Adversarial Attack on Knowledge Graph Embeddings
by: Zhao, Tianzhe, et al.
Published: (2024)
by: Zhao, Tianzhe, et al.
Published: (2024)
CrossGuard: Safeguarding MLLMs against Joint-Modal Implicit Malicious Attacks
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
by: Tihanyi, Norbert, et al.
Published: (2024)
by: Tihanyi, Norbert, et al.
Published: (2024)
FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments
by: Yang, Zhi, et al.
Published: (2026)
by: Yang, Zhi, et al.
Published: (2026)
Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models
by: Li, Xi, et al.
Published: (2024)
by: Li, Xi, et al.
Published: (2024)
AI Risk Management Should Incorporate Both Safety and Security
by: Qi, Xiangyu, et al.
Published: (2024)
by: Qi, Xiangyu, et al.
Published: (2024)
Quantifying and Defending against Privacy Threats on Federated Knowledge Graph Embedding
by: Hu, Yuke, et al.
Published: (2023)
by: Hu, Yuke, et al.
Published: (2023)
Principle-Guided Verilog Optimization: IP-Safe Knowledge Transfer via Local-Cloud Collaboration
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
Personalized Federated Learning with Adaptive Feature Aggregation and Knowledge Transfer
by: Yin, Keting, et al.
Published: (2024)
by: Yin, Keting, et al.
Published: (2024)
A Novel Approach to Malicious Code Detection Using CNN-BiLSTM and Feature Fusion
by: Zhang, Lixia, et al.
Published: (2024)
by: Zhang, Lixia, et al.
Published: (2024)
A Graph-Attentive LSTM Model for Malicious URL Detection
by: Hossain, Md. Ifthekhar, et al.
Published: (2025)
by: Hossain, Md. Ifthekhar, et al.
Published: (2025)
A Sentence Relation-Based Approach to Sanitizing Malicious Instructions
by: Datta, Soumil, et al.
Published: (2026)
by: Datta, Soumil, et al.
Published: (2026)
Leveraging Large Language Models to Detect npm Malicious Packages
by: Zahan, Nusrat, et al.
Published: (2024)
by: Zahan, Nusrat, et al.
Published: (2024)
How to Backdoor the Knowledge Distillation
by: Wu, Chen, et al.
Published: (2025)
by: Wu, Chen, et al.
Published: (2025)
Safeguarding Multimodal Knowledge Copyright in the RAG-as-a-Service Environment
by: Chen, Tianyu, et al.
Published: (2025)
by: Chen, Tianyu, et al.
Published: (2025)
PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality
by: Li, Nanxi, et al.
Published: (2025)
by: Li, Nanxi, et al.
Published: (2025)
Identification of Malicious Posts on the Dark Web Using Supervised Machine Learning
by: Filho, Sebastião Alves de Jesus, et al.
Published: (2025)
by: Filho, Sebastião Alves de Jesus, et al.
Published: (2025)
Zero-Knowledge Proofs in Sublinear Space
by: Nye, Logan
Published: (2025)
by: Nye, Logan
Published: (2025)
Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting
by: Meng, Yuqiao, et al.
Published: (2025)
by: Meng, Yuqiao, et al.
Published: (2025)
Spatial CAPTCHA: Generatively Benchmarking Spatial Reasoning for Human-Machine Differentiation
by: Kharlamova, Arina, et al.
Published: (2025)
by: Kharlamova, Arina, et al.
Published: (2025)
Safeguarding Text-to-Image Generative Models Against Unauthorized Knowledge Distillation
by: Gao, Yilan, et al.
Published: (2026)
by: Gao, Yilan, et al.
Published: (2026)
Similar Items
-
GraphSteal: Structural Knowledge Stealing from Graph RAG via Traversal Reconstruction
by: Gu, Jinze, et al.
Published: (2026) -
CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks
by: Li, Siyuan, et al.
Published: (2026) -
Analysing Safety Risks in LLMs Fine-Tuned with Pseudo-Malicious Cyber Security Data
by: ElZemity, Adel, et al.
Published: (2025) -
Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
by: Fu, Haowei, et al.
Published: (2025) -
CTIArena: Benchmarking LLM Knowledge and Reasoning Across Heterogeneous Cyber Threat Intelligence
by: Cheng, Yutong, et al.
Published: (2025)