MCA-Bench: A Multimodal Benchmark for Evaluating CAPTCHA Robustness Against VLM-based Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Zonglin, Xue, Yule, Feng, Yaoyao, Wang, Xiaolong, Song, Yiren |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AnatoMaskGAN: GNN-Driven Slice Feature Fusion and Noise Augmentation for Medical Semantic Image Synthesis
by: Wu, Zonglin, et al.
Published: (2025)
by: Wu, Zonglin, et al.
Published: (2025)
WPDA: Frequency-based Backdoor Attack with Wavelet Packet Decomposition
by: Song, Zhengyao, et al.
Published: (2024)
by: Song, Zhengyao, et al.
Published: (2024)
Fingerprint Membership and Identity Inference Against Generative Adversarial Networks
by: Cavasin, Saverio, et al.
Published: (2024)
by: Cavasin, Saverio, et al.
Published: (2024)
DVGBench: Implicit-to-Explicit Visual Grounding Benchmark in UAV Imagery with Large Vision-Language Models
by: Zhou, Yue, et al.
Published: (2026)
by: Zhou, Yue, et al.
Published: (2026)
AVadCLIP: Audio-Visual Collaboration for Robust Video Anomaly Detection
by: Wu, Peng, et al.
Published: (2025)
by: Wu, Peng, et al.
Published: (2025)
Model Agnostic Defense against Adversarial Patch Attacks on Object Detection in Unmanned Aerial Vehicles
by: Pathak, Saurabh, et al.
Published: (2024)
by: Pathak, Saurabh, et al.
Published: (2024)
Robust Multi-Source Covid-19 Detection in CT Images
by: Pritha, Asmita Yuki, et al.
Published: (2026)
by: Pritha, Asmita Yuki, et al.
Published: (2026)
MdaIF: Robust One-Stop Multi-Degradation-Aware Image Fusion with Language-Driven Semantics
by: Li, Jing, et al.
Published: (2025)
by: Li, Jing, et al.
Published: (2025)
Dynamic Brightness Adaptation for Robust Multi-modal Image Fusion
by: Sun, Yiming, et al.
Published: (2024)
by: Sun, Yiming, et al.
Published: (2024)
Ego-Motion Aware Target Prediction Module for Robust Multi-Object Tracking
by: Mahdian, Navid, et al.
Published: (2024)
by: Mahdian, Navid, et al.
Published: (2024)
TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models
by: Wang, Zeqing, et al.
Published: (2025)
by: Wang, Zeqing, et al.
Published: (2025)
Evaluating the Significance of Outdoor Advertising from Driver's Perspective Using Computer Vision
by: Černeková, Zuzana, et al.
Published: (2023)
by: Černeková, Zuzana, et al.
Published: (2023)
Doodle Your Keypoints: Sketch-Based Few-Shot Keypoint Detection
by: Maity, Subhajit, et al.
Published: (2025)
by: Maity, Subhajit, et al.
Published: (2025)
VersaGen: Unleashing Versatile Visual Control for Text-to-Image Synthesis
by: Chen, Zhipeng, et al.
Published: (2024)
by: Chen, Zhipeng, et al.
Published: (2024)
Data Augmentation in Earth Observation: A Diffusion Model Approach
by: Sousa, Tiago, et al.
Published: (2024)
by: Sousa, Tiago, et al.
Published: (2024)
DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response
by: Wang, Junjue, et al.
Published: (2025)
by: Wang, Junjue, et al.
Published: (2025)
Traffic Scene Small Target Detection Method Based on YOLOv8n-SPTS Model for Autonomous Driving
by: Wu, Songhan
Published: (2025)
by: Wu, Songhan
Published: (2025)
Tiny-YOLOSAM: Fast Hybrid Image Segmentation
by: Xu, Kenneth, et al.
Published: (2025)
by: Xu, Kenneth, et al.
Published: (2025)
Orientation-conditioned Facial Texture Mapping for Video-based Facial Remote Photoplethysmography Estimation
by: Cantrill, Sam, et al.
Published: (2024)
by: Cantrill, Sam, et al.
Published: (2024)
Cross-View-Prediction: Exploring Contrastive Feature for Hyperspectral Image Classification
by: Zhang, Anyu, et al.
Published: (2022)
by: Zhang, Anyu, et al.
Published: (2022)
SkeletonX: Data-Efficient Skeleton-based Action Recognition via Cross-sample Feature Aggregation
by: Zhang, Zongye, et al.
Published: (2025)
by: Zhang, Zongye, et al.
Published: (2025)
Gaze-Guided Learning: Avoiding Shortcut Bias in Visual Classification
by: Li, Jiahang, et al.
Published: (2025)
by: Li, Jiahang, et al.
Published: (2025)
StyleDrive: Towards Driving-Style Aware Benchmarking of End-To-End Autonomous Driving
by: Hao, Ruiyang, et al.
Published: (2025)
by: Hao, Ruiyang, et al.
Published: (2025)
DVLA-RL: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning
by: Li, Wenhao, et al.
Published: (2026)
by: Li, Wenhao, et al.
Published: (2026)
Rip Current Segmentation: A Novel Benchmark and YOLOv8 Baseline Results
by: Dumitriu, Andrei, et al.
Published: (2025)
by: Dumitriu, Andrei, et al.
Published: (2025)
Reference Dataset and Benchmark for Reconstructing Laser Parameters from On-axis Video in Powder Bed Fusion of Bulk Stainless Steel
by: Blanc, Cyril, et al.
Published: (2024)
by: Blanc, Cyril, et al.
Published: (2024)
RipVIS: Rip Currents Video Instance Segmentation Benchmark for Beach Monitoring and Safety
by: Dumitriu, Andrei, et al.
Published: (2025)
by: Dumitriu, Andrei, et al.
Published: (2025)
SPEAK: Speech-Driven Pose and Emotion-Adjustable Talking Head Generation
by: Cai, Changpeng, et al.
Published: (2024)
by: Cai, Changpeng, et al.
Published: (2024)
Human-Centric Perception for Child Sexual Abuse Imagery
by: Laranjeira, Camila, et al.
Published: (2026)
by: Laranjeira, Camila, et al.
Published: (2026)
Synthetic Image Detection with CLIP: Understanding and Assessing Predictive Cues
by: Willi, Marco, et al.
Published: (2026)
by: Willi, Marco, et al.
Published: (2026)
Investigation of cardinality classification for bacterial colony counting using explainable artificial intelligence
by: Zheng, Minghua, et al.
Published: (2026)
by: Zheng, Minghua, et al.
Published: (2026)
A Multi-purpose Tracking Framework for Salmon Welfare Monitoring in Challenging Environments
by: Høgstedt, Espen Uri, et al.
Published: (2025)
by: Høgstedt, Espen Uri, et al.
Published: (2025)
Learning to count small and clustered objects with application to bacterial colonies
by: Zheng, Minghua, et al.
Published: (2026)
by: Zheng, Minghua, et al.
Published: (2026)
Decoupled Sensitivity-Consistency Learning for Weakly Supervised Video Anomaly Detection
by: Zheng, Hantao, et al.
Published: (2026)
by: Zheng, Hantao, et al.
Published: (2026)
AOI-SSL: Self-Supervised Framework for Efficient Segmentation of Wire-bonded Semiconductors In Optical Inspection
by: Figueira, Joaquín, et al.
Published: (2026)
by: Figueira, Joaquín, et al.
Published: (2026)
DIsoN: Decentralized Isolation Networks for Out-of-Distribution Detection in Medical Imaging
by: Wagner, Felix, et al.
Published: (2025)
by: Wagner, Felix, et al.
Published: (2025)
Detailed Evaluation of Modern Machine Learning Approaches for Optic Plastics Sorting
by: Maheshkar, Vaishali, et al.
Published: (2025)
by: Maheshkar, Vaishali, et al.
Published: (2025)
NumeriKontrol: Adding Numeric Control to Diffusion Transformers for Instruction-based Image Editing
by: Xu, Zhenyu, et al.
Published: (2025)
by: Xu, Zhenyu, et al.
Published: (2025)
Scalable and Realistic Virtual Try-on Application for Foundation Makeup with Kubelka-Munk Theory
by: Pang, Hui, et al.
Published: (2025)
by: Pang, Hui, et al.
Published: (2025)
Cost Savings from Automatic Quality Assessment of Generated Images
by: Giro-i-Nieto, Xavier, et al.
Published: (2025)
by: Giro-i-Nieto, Xavier, et al.
Published: (2025)
Similar Items
-
AnatoMaskGAN: GNN-Driven Slice Feature Fusion and Noise Augmentation for Medical Semantic Image Synthesis
by: Wu, Zonglin, et al.
Published: (2025) -
WPDA: Frequency-based Backdoor Attack with Wavelet Packet Decomposition
by: Song, Zhengyao, et al.
Published: (2024) -
Fingerprint Membership and Identity Inference Against Generative Adversarial Networks
by: Cavasin, Saverio, et al.
Published: (2024) -
DVGBench: Implicit-to-Explicit Visual Grounding Benchmark in UAV Imagery with Large Vision-Language Models
by: Zhou, Yue, et al.
Published: (2026) -
AVadCLIP: Audio-Visual Collaboration for Robust Video Anomaly Detection
by: Wu, Peng, et al.
Published: (2025)