Throttling Web Agents Using Reasoning Gates
Fuente:
arXiv
Saved in:
| Main Authors: | Kumar, Abhinav, Roh, Jaechul, Naseh, Ali, Houmansadr, Amir, Bagdasarian, Eugene |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Backdooring Bias ($B^2$) into Stable Diffusion Models
by: Naseh, Ali, et al.
Published: (2024)
by: Naseh, Ali, et al.
Published: (2024)
OverThink: Slowdown Attacks on Reasoning LLMs
by: Kumar, Abhinav, et al.
Published: (2025)
by: Kumar, Abhinav, et al.
Published: (2025)
Network-Level Prompt and Trait Leakage in Local Research Agents
by: Jeong, Hyejun, et al.
Published: (2025)
by: Jeong, Hyejun, et al.
Published: (2025)
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
OSLO: One-Shot Label-Only Membership Inference Attacks
by: Peng, Yuefeng, et al.
Published: (2024)
by: Peng, Yuefeng, et al.
Published: (2024)
Multilingual and Multi-Accent Jailbreaking of Audio LLMs
by: Roh, Jaechul, et al.
Published: (2025)
by: Roh, Jaechul, et al.
Published: (2025)
Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs
by: Roh, Jaechul, et al.
Published: (2026)
by: Roh, Jaechul, et al.
Published: (2026)
Riddle Me This! Stealthy Membership Inference for Retrieval-Augmented Generation
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
Diffence: Fencing Membership Privacy With Diffusion Models
by: Peng, Yuefeng, et al.
Published: (2023)
by: Peng, Yuefeng, et al.
Published: (2023)
Can Large Language Models Really Recognize Your Name?
by: Pham, Dzung, et al.
Published: (2025)
by: Pham, Dzung, et al.
Published: (2025)
Self-interpreting Adversarial Images
by: Zhang, Tingwei, et al.
Published: (2024)
by: Zhang, Tingwei, et al.
Published: (2024)
Text-to-Image Models Leave Identifiable Signatures: Implications for Leaderboard Security
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
Exploiting Leaderboards for Large-Scale Distribution of Malicious Models
by: Suri, Anshuman, et al.
Published: (2025)
by: Suri, Anshuman, et al.
Published: (2025)
Trusted Machine Learning Models Unlock Private Inference for Problems Currently Infeasible with Cryptography
by: Shumailov, Ilia, et al.
Published: (2025)
by: Shumailov, Ilia, et al.
Published: (2025)
Contextual Agent Security: A Policy for Every Purpose
by: Tsai, Lillian, et al.
Published: (2025)
by: Tsai, Lillian, et al.
Published: (2025)
PostMark: A Robust Blackbox Watermark for Large Language Models
by: Chang, Yapei, et al.
Published: (2024)
by: Chang, Yapei, et al.
Published: (2024)
WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks
by: Ramesh, Guruprasad Viswanathan, et al.
Published: (2026)
by: Ramesh, Guruprasad Viswanathan, et al.
Published: (2026)
WAREX: Web Agent Reliability Evaluation on Existing Benchmarks
by: Kara, Su, et al.
Published: (2025)
by: Kara, Su, et al.
Published: (2025)
Identifying Models Behind Text-to-Image Leaderboards
by: Naseh, Ali, et al.
Published: (2026)
by: Naseh, Ali, et al.
Published: (2026)
Disappearing Ink: Obfuscation Breaks N-gram Code Watermarks in Theory and Practice
by: Zhang, Gehao, et al.
Published: (2025)
by: Zhang, Gehao, et al.
Published: (2025)
RAIFLE: Reconstruction Attacks on Interaction-based Federated Learning with Adversarial Data Manipulation
by: Pham, Dzung, et al.
Published: (2023)
by: Pham, Dzung, et al.
Published: (2023)
Fake or Compromised? Making Sense of Malicious Clients in Federated Learning
by: Mozaffari, Hamid, et al.
Published: (2024)
by: Mozaffari, Hamid, et al.
Published: (2024)
FameBias: Embedding Manipulation Bias Attack in Text-to-Image Models
by: Roh, Jaechul, et al.
Published: (2024)
by: Roh, Jaechul, et al.
Published: (2024)
UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own Reasoning
by: Zhang, Jiawei, et al.
Published: (2025)
by: Zhang, Jiawei, et al.
Published: (2025)
SPILLage: Agentic Oversharing on the Web
by: Roh, Jaechul, et al.
Published: (2026)
by: Roh, Jaechul, et al.
Published: (2026)
Membership Inference over Diffusion-models-based Synthetic Tabular Data
by: Cheng, Peini, et al.
Published: (2025)
by: Cheng, Peini, et al.
Published: (2025)
Adversarial Illusions in Multi-Modal Embeddings
by: Zhang, Tingwei, et al.
Published: (2023)
by: Zhang, Tingwei, et al.
Published: (2023)
GuardReasoner: Towards Reasoning-based LLM Safeguards
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
by: Liao, Zeyi, et al.
Published: (2024)
by: Liao, Zeyi, et al.
Published: (2024)
Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
by: Huang, Tiansheng, et al.
Published: (2025)
by: Huang, Tiansheng, et al.
Published: (2025)
Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search
by: Luo, Zeren, et al.
Published: (2025)
by: Luo, Zeren, et al.
Published: (2025)
Terrarium: Revisiting the Blackboard for Multi-Agent Safety, Privacy, and Security Studies
by: Nakamura, Mason, et al.
Published: (2025)
by: Nakamura, Mason, et al.
Published: (2025)
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
by: Bazinska, Julia, et al.
Published: (2025)
by: Bazinska, Julia, et al.
Published: (2025)
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
by: Naseh, Ali, et al.
Published: (2025)
by: Naseh, Ali, et al.
Published: (2025)
Machine Learning-Based Security Policy Analysis
by: Jain, Krish, et al.
Published: (2024)
by: Jain, Krish, et al.
Published: (2024)
Beyond the Request: Harnessing HTTP Response Headers for Cross-Browser Web Tracker Classification in an Imbalanced Setting
by: Rieder, Wolf, et al.
Published: (2024)
by: Rieder, Wolf, et al.
Published: (2024)
AgentGuardian: Learning Access Control Policies to Govern AI Agent Behavior
by: Abaev, Nadya, et al.
Published: (2026)
by: Abaev, Nadya, et al.
Published: (2026)
LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
by: Nie, Yuzhou, et al.
Published: (2024)
by: Nie, Yuzhou, et al.
Published: (2024)
Web Phishing Net (WPN): A scalable machine learning approach for real-time phishing campaign detection
by: Zia, Muhammad Fahad, et al.
Published: (2025)
by: Zia, Muhammad Fahad, et al.
Published: (2025)
On The Fragility of Benchmark Contamination Detection in Reasoning Models
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
Similar Items
-
Backdooring Bias ($B^2$) into Stable Diffusion Models
by: Naseh, Ali, et al.
Published: (2024) -
OverThink: Slowdown Attacks on Reasoning LLMs
by: Kumar, Abhinav, et al.
Published: (2025) -
Network-Level Prompt and Trait Leakage in Local Research Agents
by: Jeong, Hyejun, et al.
Published: (2025) -
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
by: Naseh, Ali, et al.
Published: (2025) -
OSLO: One-Shot Label-Only Membership Inference Attacks
by: Peng, Yuefeng, et al.
Published: (2024)