WAREX: Web Agent Reliability Evaluation on Existing Benchmarks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kara, Su, Faisal, Fazle, Nath, Suman |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks
von: Ramesh, Guruprasad Viswanathan, et al.
Veröffentlicht: (2026)
von: Ramesh, Guruprasad Viswanathan, et al.
Veröffentlicht: (2026)
Reliable Weak-to-Strong Monitoring of LLM Agents
von: Kale, Neil, et al.
Veröffentlicht: (2025)
von: Kale, Neil, et al.
Veröffentlicht: (2025)
Throttling Web Agents Using Reasoning Gates
von: Kumar, Abhinav, et al.
Veröffentlicht: (2025)
von: Kumar, Abhinav, et al.
Veröffentlicht: (2025)
Deep Reinforcement Learning for Phishing Detection with Transformer-Based Semantic Features
von: Faisal, Aseer Al
Veröffentlicht: (2025)
von: Faisal, Aseer Al
Veröffentlicht: (2025)
Evaluating the Reliability and Fidelity of Automated Judgment Systems of Large Language Models
von: Biskupski, Tom, et al.
Veröffentlicht: (2026)
von: Biskupski, Tom, et al.
Veröffentlicht: (2026)
Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
von: Bazinska, Julia, et al.
Veröffentlicht: (2025)
von: Bazinska, Julia, et al.
Veröffentlicht: (2025)
Cross-Session Threats in AI Agents: Benchmark, Evaluation, and Algorithms
von: Azarafrooz, Ari
Veröffentlicht: (2026)
von: Azarafrooz, Ari
Veröffentlicht: (2026)
Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
von: Anurin, Andrey, et al.
Veröffentlicht: (2024)
von: Anurin, Andrey, et al.
Veröffentlicht: (2024)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
von: Kutasov, Jonathan, et al.
Veröffentlicht: (2025)
von: Kutasov, Jonathan, et al.
Veröffentlicht: (2025)
EVMbench: Evaluating AI Agents on Smart Contract Security
von: Wang, Justin, et al.
Veröffentlicht: (2026)
von: Wang, Justin, et al.
Veröffentlicht: (2026)
A Backdoor-based Explainable AI Benchmark for High Fidelity Evaluation of Attributions
von: Yang, Peiyu, et al.
Veröffentlicht: (2024)
von: Yang, Peiyu, et al.
Veröffentlicht: (2024)
HybridGuard: Enhancing Minority-Class Intrusion Detection in Dew-Enabled Edge-of-Things Networks
von: Kara, Binayak, et al.
Veröffentlicht: (2025)
von: Kara, Binayak, et al.
Veröffentlicht: (2025)
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
von: Wang, Zhun, et al.
Veröffentlicht: (2025)
von: Wang, Zhun, et al.
Veröffentlicht: (2025)
A Customer Level Fraudulent Activity Detection Benchmark for Enhancing Machine Learning Model Research and Evaluation
von: Jing, Phoebe, et al.
Veröffentlicht: (2024)
von: Jing, Phoebe, et al.
Veröffentlicht: (2024)
Enhancing Reliability in LLM-Based Secure Code Generation
von: Kharma, Mohammed F., et al.
Veröffentlicht: (2026)
von: Kharma, Mohammed F., et al.
Veröffentlicht: (2026)
Content-Aware Attack Detection in LLM Agent Tool-Call Traffic: An Empirical Study of Features, Architectures, and Evaluation Protocols
von: Zavrak, Sultan
Veröffentlicht: (2026)
von: Zavrak, Sultan
Veröffentlicht: (2026)
EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage
von: Liao, Zeyi, et al.
Veröffentlicht: (2024)
von: Liao, Zeyi, et al.
Veröffentlicht: (2024)
Robust and Reliable Early-Stage Website Fingerprinting Attacks via Spatial-Temporal Distribution Analysis
von: Deng, Xinhao, et al.
Veröffentlicht: (2024)
von: Deng, Xinhao, et al.
Veröffentlicht: (2024)
A Fast, Reliable, and Secure Programming Language for LLM Agents with Code Actions
von: Mell, Stephen, et al.
Veröffentlicht: (2025)
von: Mell, Stephen, et al.
Veröffentlicht: (2025)
Enhancing Autonomous Online Intrusion Detection for IoT with Balanced Learning, Reliable Pseudo-Labels, and Lightweight Architectures
von: Afzaal, Hanzala, et al.
Veröffentlicht: (2026)
von: Afzaal, Hanzala, et al.
Veröffentlicht: (2026)
TADP-RME: A Trust-Adaptive Differential Privacy Framework for Enhancing Reliability of Data-Driven Systems
von: Halder, Labani, et al.
Veröffentlicht: (2026)
von: Halder, Labani, et al.
Veröffentlicht: (2026)
Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search
von: Luo, Zeren, et al.
Veröffentlicht: (2025)
von: Luo, Zeren, et al.
Veröffentlicht: (2025)
Beyond the Request: Harnessing HTTP Response Headers for Cross-Browser Web Tracker Classification in an Imbalanced Setting
von: Rieder, Wolf, et al.
Veröffentlicht: (2024)
von: Rieder, Wolf, et al.
Veröffentlicht: (2024)
On The Fragility of Benchmark Contamination Detection in Reasoning Models
von: Wang, Han, et al.
Veröffentlicht: (2025)
von: Wang, Han, et al.
Veröffentlicht: (2025)
LLM Benchmark Datasets Should Be Contamination-Resistant
von: Al-Lawati, Ali, et al.
Veröffentlicht: (2026)
von: Al-Lawati, Ali, et al.
Veröffentlicht: (2026)
AgentGuardian: Learning Access Control Policies to Govern AI Agent Behavior
von: Abaev, Nadya, et al.
Veröffentlicht: (2026)
von: Abaev, Nadya, et al.
Veröffentlicht: (2026)
LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
von: Nie, Yuzhou, et al.
Veröffentlicht: (2024)
von: Nie, Yuzhou, et al.
Veröffentlicht: (2024)
Web Phishing Net (WPN): A scalable machine learning approach for real-time phishing campaign detection
von: Zia, Muhammad Fahad, et al.
Veröffentlicht: (2025)
von: Zia, Muhammad Fahad, et al.
Veröffentlicht: (2025)
DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
von: Wang, Erchi, et al.
Veröffentlicht: (2026)
von: Wang, Erchi, et al.
Veröffentlicht: (2026)
SoK: Benchmarking Poisoning Attacks and Defenses in Federated Learning
von: Zhang, Heyi, et al.
Veröffentlicht: (2025)
von: Zhang, Heyi, et al.
Veröffentlicht: (2025)
Efficient but Vulnerable: Benchmarking and Defending LLM Batch Prompting Attack
von: Yue, Murong, et al.
Veröffentlicht: (2025)
von: Yue, Murong, et al.
Veröffentlicht: (2025)
SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
Security Considerations for Artificial Intelligence Agents
von: Li, Ninghui, et al.
Veröffentlicht: (2026)
von: Li, Ninghui, et al.
Veröffentlicht: (2026)
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
von: Aldahoul, Nouar, et al.
Veröffentlicht: (2025)
von: Aldahoul, Nouar, et al.
Veröffentlicht: (2025)
MalURLBench: A Benchmark Evaluating Agents' Vulnerabilities When Processing Web URLs
von: Kong, Dezhang, et al.
Veröffentlicht: (2026)
von: Kong, Dezhang, et al.
Veröffentlicht: (2026)
No More, No Less: Task Alignment in Terminal Agents
von: Mavali, Sina, et al.
Veröffentlicht: (2026)
von: Mavali, Sina, et al.
Veröffentlicht: (2026)
Compromising Embodied Agents with Contextual Backdoor Attacks
von: Liu, Aishan, et al.
Veröffentlicht: (2024)
von: Liu, Aishan, et al.
Veröffentlicht: (2024)
MAYA: Addressing Inconsistencies in Generative Password Guessing through a Unified Benchmark
von: Corrias, William, et al.
Veröffentlicht: (2025)
von: Corrias, William, et al.
Veröffentlicht: (2025)
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025)
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025)
Detecting Malicious AI Agents Through Simulated Interactions
von: Pi, Yulu, et al.
Veröffentlicht: (2025)
von: Pi, Yulu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks
von: Ramesh, Guruprasad Viswanathan, et al.
Veröffentlicht: (2026) -
Reliable Weak-to-Strong Monitoring of LLM Agents
von: Kale, Neil, et al.
Veröffentlicht: (2025) -
Throttling Web Agents Using Reasoning Gates
von: Kumar, Abhinav, et al.
Veröffentlicht: (2025) -
Deep Reinforcement Learning for Phishing Detection with Transformer-Based Semantic Features
von: Faisal, Aseer Al
Veröffentlicht: (2025) -
Evaluating the Reliability and Fidelity of Automated Judgment Systems of Large Language Models
von: Biskupski, Tom, et al.
Veröffentlicht: (2026)