Next-Gen CAPTCHAs: Leveraging the Cognitive Gap for Scalable and Diverse GUI-Agent Defense
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Jiacheng, Luo, Yaxin, Cui, Jiacheng, Shang, Xinyi, Zhao, Xiaohan, Shen, Zhiqiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMSurgeon: Diagnosing Data Mixture of Large Language Models
by: Luo, Yaxin, et al.
Published: (2026)
by: Luo, Yaxin, et al.
Published: (2026)
Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
by: Liu, Jiacheng, et al.
Published: (2026)
by: Liu, Jiacheng, et al.
Published: (2026)
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
by: Luo, Yaxin, et al.
Published: (2025)
by: Luo, Yaxin, et al.
Published: (2025)
Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting
by: Zhao, Xiaohan, et al.
Published: (2026)
by: Zhao, Xiaohan, et al.
Published: (2026)
FADRM: Fast and Accurate Data Residual Matching for Dataset Distillation
by: Cui, Jiacheng, et al.
Published: (2025)
by: Cui, Jiacheng, et al.
Published: (2025)
Dataset Distillation via Committee Voting
by: Cui, Jiacheng, et al.
Published: (2025)
by: Cui, Jiacheng, et al.
Published: (2025)
Are Machines Better at Complex Reasoning? Unveiling Human-Machine Inference Gaps in Entailment Verification
by: Sanyal, Soumya, et al.
Published: (2024)
by: Sanyal, Soumya, et al.
Published: (2024)
GenRES: Rethinking Evaluation for Generative Relation Extraction in the Era of Large Language Models
by: Jiang, Pengcheng, et al.
Published: (2024)
by: Jiang, Pengcheng, et al.
Published: (2024)
A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1
by: Li, Zhaoyi, et al.
Published: (2025)
by: Li, Zhaoyi, et al.
Published: (2025)
Chain Association-based Attacking and Shielding Natural Language Processing Systems
by: Huang, Jiacheng, et al.
Published: (2024)
by: Huang, Jiacheng, et al.
Published: (2024)
IAE: Irony-based Adversarial Examples for Sentiment Analysis Systems
by: Yi, Xiaoyin, et al.
Published: (2024)
by: Yi, Xiaoyin, et al.
Published: (2024)
Co-EPG: A Framework for Co-Evolution of Planning and Grounding in Autonomous GUI Agents
by: Zhao, Yuan, et al.
Published: (2025)
by: Zhao, Yuan, et al.
Published: (2025)
META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI
by: Sun, Liangtai, et al.
Published: (2022)
by: Sun, Liangtai, et al.
Published: (2022)
RNG: Reducing Multi-level Noise and Multi-grained Semantic Gap for Joint Multimodal Aspect-Sentiment Analysis
by: Liu, Yaxin, et al.
Published: (2024)
by: Liu, Yaxin, et al.
Published: (2024)
BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanism
by: Wu, Qinzhuo, et al.
Published: (2025)
by: Wu, Qinzhuo, et al.
Published: (2025)
AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management
by: Tian, Shizuo, et al.
Published: (2025)
by: Tian, Shizuo, et al.
Published: (2025)
Fast and Scalable Analytical Diffusion
by: Shang, Xinyi, et al.
Published: (2026)
by: Shang, Xinyi, et al.
Published: (2026)
InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners
by: Liu, Yuhang, et al.
Published: (2025)
by: Liu, Yuhang, et al.
Published: (2025)
Exploring 3D Dataset Pruning
by: Zhao, Xiaohan, et al.
Published: (2026)
by: Zhao, Xiaohan, et al.
Published: (2026)
Adaptive Milestone Reward for GUI Agents
by: Zheng, Congmin, et al.
Published: (2026)
by: Zheng, Congmin, et al.
Published: (2026)
Mobile-MMLU: A Mobile Intelligence Language Understanding Benchmark
by: Bsharat, Sondos Mahmoud, et al.
Published: (2025)
by: Bsharat, Sondos Mahmoud, et al.
Published: (2025)
InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training
by: Zhang, Ziyun, et al.
Published: (2026)
by: Zhang, Ziyun, et al.
Published: (2026)
Bridging Context Gaps: Leveraging Coreference Resolution for Long Contextual Understanding
by: Liu, Yanming, et al.
Published: (2024)
by: Liu, Yanming, et al.
Published: (2024)
AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks
by: Siska, Charlotte, et al.
Published: (2025)
by: Siska, Charlotte, et al.
Published: (2025)
A Lightweight Framework for Trigger-Guided LoRA-Based Self-Adaptation in LLMs
by: Wei, Jiacheng, et al.
Published: (2025)
by: Wei, Jiacheng, et al.
Published: (2025)
Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents
by: Xu, Haiyang, et al.
Published: (2026)
by: Xu, Haiyang, et al.
Published: (2026)
Leveraging Group Relative Policy Optimization to Advance Large Language Models in Traditional Chinese Medicine
by: Xie, Jiacheng, et al.
Published: (2025)
by: Xie, Jiacheng, et al.
Published: (2025)
Enhancing In-Hospital Mortality Prediction Using Multi-Representational Learning with LLM-Generated Expert Summaries
by: Battula, Harshavardhan, et al.
Published: (2024)
by: Battula, Harshavardhan, et al.
Published: (2024)
From Masks to Pixels and Meaning: A New Taxonomy, Benchmark, and Metrics for VLM Image Tampering
by: Shang, Xinyi, et al.
Published: (2026)
by: Shang, Xinyi, et al.
Published: (2026)
Paper2Agent: Reimagining Research Papers As Interactive and Reliable AI Agents
by: Miao, Jiacheng, et al.
Published: (2025)
by: Miao, Jiacheng, et al.
Published: (2025)
Mobile GUI Agents under Real-world Threats: Are We There Yet?
by: Liu, Guohong, et al.
Published: (2025)
by: Liu, Guohong, et al.
Published: (2025)
Scalable Environments Drive Generalizable Agents
by: Zhang, Jiayi, et al.
Published: (2026)
by: Zhang, Jiayi, et al.
Published: (2026)
DRAG: Distilling RAG for SLMs from LLMs to Transfer Knowledge and Mitigate Hallucination via Evidence and Graph-based Distillation
by: Chen, Jennifer, et al.
Published: (2025)
by: Chen, Jennifer, et al.
Published: (2025)
D-GARA: A Dynamic Benchmarking Framework for GUI Agent Robustness in Real-World Anomalies
by: Chen, Sen, et al.
Published: (2025)
by: Chen, Sen, et al.
Published: (2025)
ToolLibGen: Scalable Automatic Tool Creation and Aggregation for LLM Reasoning
by: Yue, Murong, et al.
Published: (2025)
by: Yue, Murong, et al.
Published: (2025)
Improving Academic Skills Assessment with NLP and Ensemble Learning
by: Huang, Xinyi, et al.
Published: (2024)
by: Huang, Xinyi, et al.
Published: (2024)
ProgRM: Build Better GUI Agents with Progress Rewards
by: Zhang, Danyang, et al.
Published: (2025)
by: Zhang, Danyang, et al.
Published: (2025)
MAGE: Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory
by: Wang, Yuhui, et al.
Published: (2026)
by: Wang, Yuhui, et al.
Published: (2026)
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
by: Wu, Qianhui, et al.
Published: (2025)
by: Wu, Qianhui, et al.
Published: (2025)
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
by: Tang, Fei, et al.
Published: (2026)
by: Tang, Fei, et al.
Published: (2026)
Similar Items
-
LLMSurgeon: Diagnosing Data Mixture of Large Language Models
by: Luo, Yaxin, et al.
Published: (2026) -
Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
by: Liu, Jiacheng, et al.
Published: (2026) -
Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
by: Luo, Yaxin, et al.
Published: (2025) -
Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting
by: Zhao, Xiaohan, et al.
Published: (2026) -
FADRM: Fast and Accurate Data Residual Matching for Dataset Distillation
by: Cui, Jiacheng, et al.
Published: (2025)