LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Chiyu, Yang, Huiqin, Jiang, Bendong, Zhang, Xiaolei, Zhao, Yiran, Chen, Ruyi, Zhou, Lu, Xu, Xiaogang, Wu, Jiafei, Fang, Liming, Liu, Zhe |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
by: Zhang, Chiyu, et al.
Published: (2025)
by: Zhang, Chiyu, et al.
Published: (2025)
Adversarial Attacks of Vision Tasks in the Past 10 Years: A Survey
by: Zhang, Chiyu, et al.
Published: (2024)
by: Zhang, Chiyu, et al.
Published: (2024)
From Thinker to Society: Security in Hierarchical Autonomy Evolution of AI Agents
by: Zhang, Xiaolei, et al.
Published: (2026)
by: Zhang, Xiaolei, et al.
Published: (2026)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
by: Zeng, Yifan, et al.
Published: (2024)
by: Zeng, Yifan, et al.
Published: (2024)
SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
by: Lee, Hwiwon, et al.
Published: (2025)
by: Lee, Hwiwon, et al.
Published: (2025)
DR-Encoder: Encode Low-rank Gradients with Random Prior for Large Language Models Differentially Privately
by: Wu, Huiwen, et al.
Published: (2024)
by: Wu, Huiwen, et al.
Published: (2024)
RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments
by: Fu, Yuchuan, et al.
Published: (2025)
by: Fu, Yuchuan, et al.
Published: (2025)
Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak Attacks
by: Zhang, Yingjie, et al.
Published: (2025)
by: Zhang, Yingjie, et al.
Published: (2025)
HarmChip: Evaluating Hardware Security Centric LLM Safety via Jailbreak Benchmarking
by: Wang, Zeng, et al.
Published: (2026)
by: Wang, Zeng, et al.
Published: (2026)
AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference
by: Liu, Songyang, et al.
Published: (2026)
by: Liu, Songyang, et al.
Published: (2026)
Code Agent can be an End-to-end System Hacker: Benchmarking Real-world Threats of Computer-use Agent
by: Luo, Weidi, et al.
Published: (2025)
by: Luo, Weidi, et al.
Published: (2025)
Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking
by: Zhang, Junke, et al.
Published: (2026)
by: Zhang, Junke, et al.
Published: (2026)
From Storage to Steering: Memory Control Flow Attacks on LLM Agents
by: Xu, Zhenlin, et al.
Published: (2026)
by: Xu, Zhenlin, et al.
Published: (2026)
SafeSteer: Adaptive Subspace Steering for Efficient Jailbreak Defense in Vision-Language Models
by: Zeng, Xiyu, et al.
Published: (2025)
by: Zeng, Xiyu, et al.
Published: (2025)
Beyond Jailbreak: Unveiling Risks in LLM Applications Arising from Blurred Capability Boundaries
by: Zhang, Yunyi, et al.
Published: (2025)
by: Zhang, Yunyi, et al.
Published: (2025)
Detecting Scams Using Large Language Models
by: Jiang, Liming
Published: (2024)
by: Jiang, Liming
Published: (2024)
Utilizing Large LanguageModels to Detect Privacy Leaks in Mini-App Code
by: Jiang, Liming
Published: (2024)
by: Jiang, Liming
Published: (2024)
SafeMobile: Chain-level Jailbreak Detection and Automated Evaluation for Multimodal Mobile Agents
by: Liang, Siyuan, et al.
Published: (2025)
by: Liang, Siyuan, et al.
Published: (2025)
When LLM Meets DRL: Advancing Jailbreaking Efficiency via DRL-guided Search
by: Chen, Xuan, et al.
Published: (2024)
by: Chen, Xuan, et al.
Published: (2024)
Confusion is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMs
by: Yan, Yu, et al.
Published: (2025)
by: Yan, Yu, et al.
Published: (2025)
JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit
by: He, Zeqing, et al.
Published: (2024)
by: He, Zeqing, et al.
Published: (2024)
From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem
by: Mao, Yanxu, et al.
Published: (2025)
by: Mao, Yanxu, et al.
Published: (2025)
Shell or Nothing: Real-World Benchmarks and Memory-Activated Agents for Automated Penetration Testing
by: Mai, Wuyuao, et al.
Published: (2025)
by: Mai, Wuyuao, et al.
Published: (2025)
Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
by: Wu, Yu-Hang, et al.
Published: (2025)
by: Wu, Yu-Hang, et al.
Published: (2025)
BadLLM-TG: A Backdoor Defender powered by LLM Trigger Generator
by: Zhang, Ruyi, et al.
Published: (2026)
by: Zhang, Ruyi, et al.
Published: (2026)
SpatialJB: How Text Distribution Art Becomes the "Jailbreak Key" for LLM Guardrails
by: Mou, Zhiyi, et al.
Published: (2026)
by: Mou, Zhiyi, et al.
Published: (2026)
Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency
by: Jiang, Yukun, et al.
Published: (2025)
by: Jiang, Yukun, et al.
Published: (2025)
AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency Lens
by: Lu, Lin, et al.
Published: (2024)
by: Lu, Lin, et al.
Published: (2024)
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
by: Chao, Patrick, et al.
Published: (2024)
by: Chao, Patrick, et al.
Published: (2024)
FuzzLLM: A Novel and Universal Fuzzing Framework for Proactively Discovering Jailbreak Vulnerabilities in Large Language Models
by: Yao, Dongyu, et al.
Published: (2023)
by: Yao, Dongyu, et al.
Published: (2023)
Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate
by: Qi, Senmao, et al.
Published: (2025)
by: Qi, Senmao, et al.
Published: (2025)
Siren: A Learning-Based Multi-Turn Attack Framework for Simulating Real-World Human Jailbreak Behaviors
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
RLSpoofer: A Lightweight Evaluator for LLM Watermark Spoofing Resilience
by: Huang, Hanbo, et al.
Published: (2026)
by: Huang, Hanbo, et al.
Published: (2026)
Invisible to Humans, Triggered by Agents: Stealthy Jailbreak Attacks on Mobile Vision-Language Agents
by: Ding, Renhua, et al.
Published: (2025)
by: Ding, Renhua, et al.
Published: (2025)
CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents
by: Fu, Wenjie, et al.
Published: (2026)
by: Fu, Wenjie, et al.
Published: (2026)
Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense
by: Chen, Zejian, et al.
Published: (2026)
by: Chen, Zejian, et al.
Published: (2026)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
by: Zhang, Hanrong, et al.
Published: (2024)
by: Zhang, Hanrong, et al.
Published: (2024)
Beyond Jailbreaking: Auditing Contextual Privacy in LLM Agents
by: Das, Saswat, et al.
Published: (2025)
by: Das, Saswat, et al.
Published: (2025)
JPRO: Automated Multimodal Jailbreaking via Multi-Agent Collaboration Framework
by: Zhou, Yuxuan, et al.
Published: (2025)
by: Zhou, Yuxuan, et al.
Published: (2025)
Similar Items
-
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
by: Zhang, Chiyu, et al.
Published: (2025) -
Adversarial Attacks of Vision Tasks in the Past 10 Years: A Survey
by: Zhang, Chiyu, et al.
Published: (2024) -
From Thinker to Society: Security in Hierarchical Autonomy Evolution of AI Agents
by: Zhang, Xiaolei, et al.
Published: (2026) -
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
by: Zeng, Yifan, et al.
Published: (2024) -
SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
by: Lee, Hwiwon, et al.
Published: (2025)