WebPII: Benchmarking Visual PII Detection for Computer-Use Agents
Fuente:
arXiv
Salvato in:
| Autore principale: | Zhao, Nathan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Membership Inference for Contrastive Pre-training Models with Text-only PII Queries
di: Cheng, Ruoxi, et al.
Pubblicazione: (2026)
di: Cheng, Ruoxi, et al.
Pubblicazione: (2026)
PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding
di: Nakka, Krishna Kanth, et al.
Pubblicazione: (2024)
di: Nakka, Krishna Kanth, et al.
Pubblicazione: (2024)
PIIGuard: Mitigating PII Harvesting under Adversarial Sanitization
di: Liu, Mingshuo, et al.
Pubblicazione: (2026)
di: Liu, Mingshuo, et al.
Pubblicazione: (2026)
Adaptive PII Mitigation Framework for Large Language Models
di: Asthana, Shubhi, et al.
Pubblicazione: (2025)
di: Asthana, Shubhi, et al.
Pubblicazione: (2025)
PII-Bench: Evaluating Query-Aware Privacy Protection Systems
di: Shen, Hao, et al.
Pubblicazione: (2025)
di: Shen, Hao, et al.
Pubblicazione: (2025)
Model Inversion Attacks on Llama 3: Extracting PII from Large Language Models
di: Sivashanmugam, Sathesh P.
Pubblicazione: (2025)
di: Sivashanmugam, Sathesh P.
Pubblicazione: (2025)
VisualLeakBench: Auditing the Fragility of Large Vision-Language Models against PII Leakage and Social Engineering
di: Wang, Youting, et al.
Pubblicazione: (2026)
di: Wang, Youting, et al.
Pubblicazione: (2026)
"I Strongly Suspect This Website Is a Scam": Benchmarking PII Leakage and Detection without Defense in Autonomous Web Agents
di: Roy, Soham, et al.
Pubblicazione: (2026)
di: Roy, Soham, et al.
Pubblicazione: (2026)
Measuring the Accuracy and Effectiveness of PII Removal Services
di: He, Jiahui, et al.
Pubblicazione: (2025)
di: He, Jiahui, et al.
Pubblicazione: (2025)
Mind the Web: The Security of Web Use Agents
di: Shapira, Avishag, et al.
Pubblicazione: (2025)
di: Shapira, Avishag, et al.
Pubblicazione: (2025)
VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents
di: Cao, Tri, et al.
Pubblicazione: (2025)
di: Cao, Tri, et al.
Pubblicazione: (2025)
CAPID: Context-Aware PII Detection for Question-Answering Systems
di: Ponomarenko, Mariia, et al.
Pubblicazione: (2026)
di: Ponomarenko, Mariia, et al.
Pubblicazione: (2026)
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
di: Liu, Yinuo, et al.
Pubblicazione: (2025)
di: Liu, Yinuo, et al.
Pubblicazione: (2025)
PII-VisBench: Evaluating Personally Identifiable Information Safety in Vision Language Models Along a Continuum of Visibility
di: Shahariar, G M, et al.
Pubblicazione: (2026)
di: Shahariar, G M, et al.
Pubblicazione: (2026)
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
di: Evtimov, Ivan, et al.
Pubblicazione: (2025)
di: Evtimov, Ivan, et al.
Pubblicazione: (2025)
PII Jailbreaking in LLMs via Activation Steering Reveals Personal Information Leakage
di: Nakka, Krishna Kanth, et al.
Pubblicazione: (2025)
di: Nakka, Krishna Kanth, et al.
Pubblicazione: (2025)
PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction
di: Garza, Leon, et al.
Pubblicazione: (2025)
di: Garza, Leon, et al.
Pubblicazione: (2025)
The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents
di: Ding, Xuwei, et al.
Pubblicazione: (2026)
di: Ding, Xuwei, et al.
Pubblicazione: (2026)
MalURLBench: A Benchmark Evaluating Agents' Vulnerabilities When Processing Web URLs
di: Kong, Dezhang, et al.
Pubblicazione: (2026)
di: Kong, Dezhang, et al.
Pubblicazione: (2026)
Discovering Universal Activation Directions for PII Leakage in Language Models
di: Marchyok, Leo, et al.
Pubblicazione: (2026)
di: Marchyok, Leo, et al.
Pubblicazione: (2026)
Merger-as-a-Stealer: Stealing Targeted PII from Aligned LLMs with Model Merging
di: Lu, Lin, et al.
Pubblicazione: (2025)
di: Lu, Lin, et al.
Pubblicazione: (2025)
The Art of Building Verifiers for Computer Use Agents
di: Rosset, Corby, et al.
Pubblicazione: (2026)
di: Rosset, Corby, et al.
Pubblicazione: (2026)
SnapGuard: Lightweight Prompt Injection Detection for Screenshot-Based Web Agents
di: Du, Mengyao, et al.
Pubblicazione: (2026)
di: Du, Mengyao, et al.
Pubblicazione: (2026)
WAREX: Web Agent Reliability Evaluation on Existing Benchmarks
di: Kara, Su, et al.
Pubblicazione: (2025)
di: Kara, Su, et al.
Pubblicazione: (2025)
WIPI: A New Web Threat for LLM-Driven Web Agents
di: Wu, Fangzhou, et al.
Pubblicazione: (2024)
di: Wu, Fangzhou, et al.
Pubblicazione: (2024)
Temporal UI State Inconsistency in Desktop GUI Agents: Formalizing and Defending Against TOCTOU Attacks on Computer-Use Agents
di: Xu, Wenpeng
Pubblicazione: (2026)
di: Xu, Wenpeng
Pubblicazione: (2026)
WebTrap Park: An Automated Platform for Systematic Security Evaluation of Web Agents
di: Wu, Xinyi, et al.
Pubblicazione: (2026)
di: Wu, Xinyi, et al.
Pubblicazione: (2026)
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
di: Wang, Xilong, et al.
Pubblicazione: (2026)
di: Wang, Xilong, et al.
Pubblicazione: (2026)
Multi-Agent Penetration Testing AI for the Web
di: David, Isaac, et al.
Pubblicazione: (2025)
di: David, Isaac, et al.
Pubblicazione: (2025)
SEAL-Tag: Self-Tag Evidence Aggregation with Probabilistic Circuits for PII-Safe Retrieval-Augmented Generation
di: Xie, Jin, et al.
Pubblicazione: (2026)
di: Xie, Jin, et al.
Pubblicazione: (2026)
AWE: Adaptive Agents for Dynamic Web Penetration Testing
di: Jaswal, Akshat Singh, et al.
Pubblicazione: (2026)
di: Jaswal, Akshat Singh, et al.
Pubblicazione: (2026)
DECEPTICON: How Dark Patterns Manipulate Web Agents
di: Cuvin, Phil, et al.
Pubblicazione: (2025)
di: Cuvin, Phil, et al.
Pubblicazione: (2025)
LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injection
di: Zhao, Lei, et al.
Pubblicazione: (2026)
di: Zhao, Lei, et al.
Pubblicazione: (2026)
DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain
di: Huang, Enhao, et al.
Pubblicazione: (2025)
di: Huang, Enhao, et al.
Pubblicazione: (2025)
Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces
di: Zhang, Yilin, et al.
Pubblicazione: (2026)
di: Zhang, Yilin, et al.
Pubblicazione: (2026)
Cross-Modal Content Optimization for Steering Web Agent Preferences
di: Jiang, Tanqiu, et al.
Pubblicazione: (2025)
di: Jiang, Tanqiu, et al.
Pubblicazione: (2025)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
di: Cao, Tri, et al.
Pubblicazione: (2026)
di: Cao, Tri, et al.
Pubblicazione: (2026)
AI-Governed Agent Architecture for Web-Trustworthy Tokenization of Alternative Assets
di: Borjigin, Ailiya, et al.
Pubblicazione: (2025)
di: Borjigin, Ailiya, et al.
Pubblicazione: (2025)
AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use
di: Yang, Chenglin
Pubblicazione: (2026)
di: Yang, Chenglin
Pubblicazione: (2026)
Atomicity for Agents: Exposing, Exploiting, and Mitigating TOCTOU Vulnerabilities in Browser-Use Agents
di: Jiang, Linxi, et al.
Pubblicazione: (2026)
di: Jiang, Linxi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Membership Inference for Contrastive Pre-training Models with Text-only PII Queries
di: Cheng, Ruoxi, et al.
Pubblicazione: (2026) -
PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding
di: Nakka, Krishna Kanth, et al.
Pubblicazione: (2024) -
PIIGuard: Mitigating PII Harvesting under Adversarial Sanitization
di: Liu, Mingshuo, et al.
Pubblicazione: (2026) -
Adaptive PII Mitigation Framework for Large Language Models
di: Asthana, Shubhi, et al.
Pubblicazione: (2025) -
PII-Bench: Evaluating Query-Aware Privacy Protection Systems
di: Shen, Hao, et al.
Pubblicazione: (2025)