The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Yihao, Wang, Kai, Wu, Jiangrong, Wu, Haolin, Zhou, Yuxuan, Wei, Zeming, Wu, Dongxian, Chen, Xun, Sun, Jun, Sun, Meng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ClawWorm: Self-Propagating Attacks Across LLM Agent Ecosystems
por: Zhang, Yihao, et al.
Publicado: (2026)
por: Zhang, Yihao, et al.
Publicado: (2026)
MILE: A Mutation Testing Framework of In-Context Learning Systems
por: Wei, Zeming, et al.
Publicado: (2024)
por: Wei, Zeming, et al.
Publicado: (2024)
Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models
por: Zhang, Yihao, et al.
Publicado: (2024)
por: Zhang, Yihao, et al.
Publicado: (2024)
Secure LLM Fine-Tuning via Safety-Aware Probing
por: Wu, Chengcan, et al.
Publicado: (2025)
por: Wu, Chengcan, et al.
Publicado: (2025)
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing
por: Wei, Zeming, et al.
Publicado: (2026)
por: Wei, Zeming, et al.
Publicado: (2026)
Automata-Based Steering of Large Language Models for Diverse Structured Generation
por: Luan, Xiaokun, et al.
Publicado: (2025)
por: Luan, Xiaokun, et al.
Publicado: (2025)
Boosting Jailbreak Attack with Momentum
por: Zhang, Yihao, et al.
Publicado: (2024)
por: Zhang, Yihao, et al.
Publicado: (2024)
When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
por: Wang, Kai, et al.
Publicado: (2025)
por: Wang, Kai, et al.
Publicado: (2025)
Control at Stake: Evaluating the Security Landscape of LLM-Driven Email Agents
por: Wu, Jiangrong, et al.
Publicado: (2025)
por: Wu, Jiangrong, et al.
Publicado: (2025)
Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings
por: Zhang, Zhixin, et al.
Publicado: (2025)
por: Zhang, Zhixin, et al.
Publicado: (2025)
Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
por: Yang, Wenkai, et al.
Publicado: (2024)
por: Yang, Wenkai, et al.
Publicado: (2024)
ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction
por: Wei, Zeming, et al.
Publicado: (2025)
por: Wei, Zeming, et al.
Publicado: (2025)
ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
por: Wu, Yiran, et al.
Publicado: (2025)
por: Wu, Yiran, et al.
Publicado: (2025)
Conversations Risk Detection LLMs in Financial Agents via Multi-Stage Generative Rollout
por: Jiang, Xiaotong, et al.
Publicado: (2026)
por: Jiang, Xiaotong, et al.
Publicado: (2026)
The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis
por: Wang, Peiran, et al.
Publicado: (2026)
por: Wang, Peiran, et al.
Publicado: (2026)
Security Attacks on LLM-based Code Completion Tools
por: Cheng, Wen, et al.
Publicado: (2024)
por: Cheng, Wen, et al.
Publicado: (2024)
ChainFuzzer: Greybox Fuzzing for Workflow-Level Multi-Tool Vulnerabilities in LLM Agents
por: Wu, Jiangrong, et al.
Publicado: (2026)
por: Wu, Jiangrong, et al.
Publicado: (2026)
LRCTI: A Large Language Model-Based Framework for Multi-Step Evidence Retrieval and Reasoning in Cyber Threat Intelligence Credibility Verification
por: Tang, Fengxiao, et al.
Publicado: (2025)
por: Tang, Fengxiao, et al.
Publicado: (2025)
From Retrieval to Reasoning: A Framework for Cyber Threat Intelligence NER with Explicit and Adaptive Instructions
por: Peng, Jiaren, et al.
Publicado: (2025)
por: Peng, Jiaren, et al.
Publicado: (2025)
OmniSafeBench-MM: A Unified Benchmark and Toolbox for Multimodal Jailbreak Attack-Defense Evaluation
por: Jia, Xiaojun, et al.
Publicado: (2025)
por: Jia, Xiaojun, et al.
Publicado: (2025)
Is the Digital Forensics and Incident Response Pipeline Ready for Text-Based Threats in LLM Era?
por: Bhandarkar, Avanti, et al.
Publicado: (2024)
por: Bhandarkar, Avanti, et al.
Publicado: (2024)
Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
por: Wu, Yu-Hang, et al.
Publicado: (2025)
por: Wu, Yu-Hang, et al.
Publicado: (2025)
Self-Disguise Attack: Induce the LLM to disguise itself for AIGT detection evasion
por: Zhou, Yinghan, et al.
Publicado: (2025)
por: Zhou, Yinghan, et al.
Publicado: (2025)
HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application Vulnerabilities
por: Ren, Xiaoxue, et al.
Publicado: (2025)
por: Ren, Xiaoxue, et al.
Publicado: (2025)
Exploit the Leak: Understanding Risks in Biometric Matchers
por: Durbet, Axel, et al.
Publicado: (2023)
por: Durbet, Axel, et al.
Publicado: (2023)
Can LLM Infer Risk Information From MCP Server System Logs?
por: Fu, Jiayi, et al.
Publicado: (2025)
por: Fu, Jiayi, et al.
Publicado: (2025)
On the Hidden Costs of Counterfactual Knowledge Training in LLM Unlearning
por: Ye, Xiaotian, et al.
Publicado: (2026)
por: Ye, Xiaotian, et al.
Publicado: (2026)
The Obvious Invisible Threat: LLM-Powered GUI Agents' Vulnerability to Fine-Print Injections
por: Chen, Chaoran, et al.
Publicado: (2025)
por: Chen, Chaoran, et al.
Publicado: (2025)
MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
por: Wen, Rui, et al.
Publicado: (2026)
por: Wen, Rui, et al.
Publicado: (2026)
BraveGuard: From Open-World Threats to Safer Computer-Use Agents
por: Feng, Yunhao, et al.
Publicado: (2026)
por: Feng, Yunhao, et al.
Publicado: (2026)
Exploring the Robustness of In-Context Learning with Noisy Labels
por: Cheng, Chen, et al.
Publicado: (2024)
por: Cheng, Chen, et al.
Publicado: (2024)
Watermarking LLM Agent Trajectories
por: Meng, Wenlong, et al.
Publicado: (2026)
por: Meng, Wenlong, et al.
Publicado: (2026)
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures
por: Zhou, Yukai, et al.
Publicado: (2025)
por: Zhou, Yukai, et al.
Publicado: (2025)
RoMA: Robust Malware Attribution via Byte-level Adversarial Training with Global Perturbations and Adversarial Consistency Regularization
por: Sun, Yuxia, et al.
Publicado: (2025)
por: Sun, Yuxia, et al.
Publicado: (2025)
Generalized Security-Preserving Refinement for Concurrent Systems
por: Sun, Huan, et al.
Publicado: (2025)
por: Sun, Huan, et al.
Publicado: (2025)
Calibrated Adversarial Sampling: Multi-Armed Bandit-Guided Generalization Against Unforeseen Attacks
por: Wang, Rui, et al.
Publicado: (2025)
por: Wang, Rui, et al.
Publicado: (2025)
Triaging Threats to Specialized Guardrails
por: Mo, Wenjie Jacky, et al.
Publicado: (2026)
por: Mo, Wenjie Jacky, et al.
Publicado: (2026)
Resource Consumption Threats in Large Language Models
por: Zhang, Yuanhe, et al.
Publicado: (2026)
por: Zhang, Yuanhe, et al.
Publicado: (2026)
CyberThreat-Eval: Can Large Language Models Automate Real-World Threat Research?
por: Chen, Xiangsen, et al.
Publicado: (2026)
por: Chen, Xiangsen, et al.
Publicado: (2026)
From Perception to Protection: A Developer-Centered Study of Security and Privacy Threats in Extended Reality (XR)
por: Cai, Kunlin, et al.
Publicado: (2025)
por: Cai, Kunlin, et al.
Publicado: (2025)
Ejemplares similares
-
ClawWorm: Self-Propagating Attacks Across LLM Agent Ecosystems
por: Zhang, Yihao, et al.
Publicado: (2026) -
MILE: A Mutation Testing Framework of In-Context Learning Systems
por: Wei, Zeming, et al.
Publicado: (2024) -
Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models
por: Zhang, Yihao, et al.
Publicado: (2024) -
Secure LLM Fine-Tuning via Safety-Aware Probing
por: Wu, Chengcan, et al.
Publicado: (2025) -
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing
por: Wei, Zeming, et al.
Publicado: (2026)