UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Jiawei, Yang, Shuang, Li, Bo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CHAI: Command Hijacking against embodied AI
por: Burbano, Luis, et al.
Publicado: (2025)
por: Burbano, Luis, et al.
Publicado: (2025)
Seed Hijacking of LLM Sampling and Quantum Random Number Defense
por: You, Ziyang, et al.
Publicado: (2026)
por: You, Ziyang, et al.
Publicado: (2026)
Hijack Vertical Federated Learning Models As One Party
por: Qiu, Pengyu, et al.
Publicado: (2022)
por: Qiu, Pengyu, et al.
Publicado: (2022)
From Firewalls to Frontiers: AI Red-Teaming is a Domain-Specific Evolution of Cyber Red-Teaming
por: Sinha, Anusha, et al.
Publicado: (2025)
por: Sinha, Anusha, et al.
Publicado: (2025)
Red Teaming Large Reasoning Models
por: Chen, Jiawei, et al.
Publicado: (2025)
por: Chen, Jiawei, et al.
Publicado: (2025)
A Unified Framework for LLM Watermarks
por: Gloaguen, Thibaud, et al.
Publicado: (2026)
por: Gloaguen, Thibaud, et al.
Publicado: (2026)
Adaptive Instruction Composition for Automated LLM Red-Teaming
por: Zymet, Jesse, et al.
Publicado: (2026)
por: Zymet, Jesse, et al.
Publicado: (2026)
LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
por: Nie, Yuzhou, et al.
Publicado: (2024)
por: Nie, Yuzhou, et al.
Publicado: (2024)
Capability-Based Scaling Trends for LLM-Based Red-Teaming
por: Panfilov, Alexander, et al.
Publicado: (2025)
por: Panfilov, Alexander, et al.
Publicado: (2025)
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
por: Sharma, Mrinank, et al.
Publicado: (2025)
por: Sharma, Mrinank, et al.
Publicado: (2025)
Leveraging Reinforcement Learning in Red Teaming for Advanced Ransomware Attack Simulations
por: Wang, Cheng, et al.
Publicado: (2024)
por: Wang, Cheng, et al.
Publicado: (2024)
Attack Atlas: A Practitioner's Perspective on Challenges and Pitfalls in Red Teaming GenAI
por: Rawat, Ambrish, et al.
Publicado: (2024)
por: Rawat, Ambrish, et al.
Publicado: (2024)
ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System
por: Liang, Jiacheng, et al.
Publicado: (2026)
por: Liang, Jiacheng, et al.
Publicado: (2026)
GuardReasoner: Towards Reasoning-based LLM Safeguards
por: Liu, Yue, et al.
Publicado: (2025)
por: Liu, Yue, et al.
Publicado: (2025)
Large Empirical Case Study: Go-Explore adapted for AI Red Team Testing
por: Bhatt, Manish, et al.
Publicado: (2025)
por: Bhatt, Manish, et al.
Publicado: (2025)
In-Context Representation Hijacking
por: Yona, Itay, et al.
Publicado: (2025)
por: Yona, Itay, et al.
Publicado: (2025)
Reliable Weak-to-Strong Monitoring of LLM Agents
por: Kale, Neil, et al.
Publicado: (2025)
por: Kale, Neil, et al.
Publicado: (2025)
SafeProtein: Red-Teaming Framework and Benchmark for Protein Foundation Models
por: Fan, Jigang, et al.
Publicado: (2025)
por: Fan, Jigang, et al.
Publicado: (2025)
AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration
por: Zhou, Andy, et al.
Publicado: (2025)
por: Zhou, Andy, et al.
Publicado: (2025)
Your Agent Can Defend Itself against Backdoor Attacks
por: Changjiang, Li, et al.
Publicado: (2025)
por: Changjiang, Li, et al.
Publicado: (2025)
TERD: A Unified Framework for Safeguarding Diffusion Models Against Backdoors
por: Mo, Yichuan, et al.
Publicado: (2024)
por: Mo, Yichuan, et al.
Publicado: (2024)
InfiCoEvalChain: A Blockchain-Based Decentralized Framework for Collaborative LLM Evaluation
por: Yang, Yifan, et al.
Publicado: (2026)
por: Yang, Yifan, et al.
Publicado: (2026)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
por: Kutasov, Jonathan, et al.
Publicado: (2025)
por: Kutasov, Jonathan, et al.
Publicado: (2025)
MIP against Agent: Malicious Image Patches Hijacking Multimodal OS Agents
por: Aichberger, Lukas, et al.
Publicado: (2025)
por: Aichberger, Lukas, et al.
Publicado: (2025)
Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak Prompts
por: Liu, Yi, et al.
Publicado: (2024)
por: Liu, Yi, et al.
Publicado: (2024)
Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-To-Image Generation Models
por: Dong, Yingkai, et al.
Publicado: (2024)
por: Dong, Yingkai, et al.
Publicado: (2024)
Enhancing Cloud Network Resilience via a Robust LLM-Empowered Multi-Agent Reinforcement Learning Framework
por: Peng, Yixiao, et al.
Publicado: (2026)
por: Peng, Yixiao, et al.
Publicado: (2026)
The Autonomy Tax: Defense Training Breaks LLM Agents
por: Li, Shawn, et al.
Publicado: (2026)
por: Li, Shawn, et al.
Publicado: (2026)
RedVisor: Reasoning-Aware Prompt Injection Defense via Zero-Copy KV Cache Reuse
por: Liu, Mingrui, et al.
Publicado: (2026)
por: Liu, Mingrui, et al.
Publicado: (2026)
Benchmark Early and Red Team Often: A Framework for Assessing and Managing Dual-Use Hazards of AI Foundation Models
por: Barrett, Anthony M., et al.
Publicado: (2024)
por: Barrett, Anthony M., et al.
Publicado: (2024)
SUB-PLAY: Adversarial Policies against Partially Observed Multi-Agent Reinforcement Learning Systems
por: Ma, Oubo, et al.
Publicado: (2024)
por: Ma, Oubo, et al.
Publicado: (2024)
FHE-Agent: Automating CKKS Configuration for Practical Encrypted Inference via an LLM-Guided Agentic Framework
por: Xu, Nuo, et al.
Publicado: (2025)
por: Xu, Nuo, et al.
Publicado: (2025)
A Unified Learn-to-Distort-Data Framework for Privacy-Utility Trade-off in Trustworthy Federated Learning
por: Zhang, Xiaojin, et al.
Publicado: (2024)
por: Zhang, Xiaojin, et al.
Publicado: (2024)
ThinkTrap: Denial-of-Service Attacks against Black-box LLM Services via Infinite Thinking
por: Li, Yunzhe, et al.
Publicado: (2025)
por: Li, Yunzhe, et al.
Publicado: (2025)
Throttling Web Agents Using Reasoning Gates
por: Kumar, Abhinav, et al.
Publicado: (2025)
por: Kumar, Abhinav, et al.
Publicado: (2025)
CLASP: Training-Free LLM-Assisted Source Code Watermarking via Semantic-Preserving Transformations
por: Xu, Rui, et al.
Publicado: (2025)
por: Xu, Rui, et al.
Publicado: (2025)
HijackRAG: Hijacking Attacks against Retrieval-Augmented Large Language Models
por: Zhang, Yucheng, et al.
Publicado: (2024)
por: Zhang, Yucheng, et al.
Publicado: (2024)
Predictive Auditing of Hidden Tokens in LLM APIs via Reasoning Length Estimation
por: Wang, Ziyao, et al.
Publicado: (2025)
por: Wang, Ziyao, et al.
Publicado: (2025)
PLeak: Prompt Leaking Attacks against Large Language Model Applications
por: Hui, Bo, et al.
Publicado: (2024)
por: Hui, Bo, et al.
Publicado: (2024)
RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content
por: Yuan, Zhuowen, et al.
Publicado: (2024)
por: Yuan, Zhuowen, et al.
Publicado: (2024)
Ejemplares similares
-
CHAI: Command Hijacking against embodied AI
por: Burbano, Luis, et al.
Publicado: (2025) -
Seed Hijacking of LLM Sampling and Quantum Random Number Defense
por: You, Ziyang, et al.
Publicado: (2026) -
Hijack Vertical Federated Learning Models As One Party
por: Qiu, Pengyu, et al.
Publicado: (2022) -
From Firewalls to Frontiers: AI Red-Teaming is a Domain-Specific Evolution of Cyber Red-Teaming
por: Sinha, Anusha, et al.
Publicado: (2025) -
Red Teaming Large Reasoning Models
por: Chen, Jiawei, et al.
Publicado: (2025)