Robustness Evaluation of OCR-based Visual Document Understanding under Multi-Modal Adversarial Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Tien, Dong Nguyen, Le, Dung D. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adversarial Training with OCR Modality Perturbation for Scene-Text Visual Question Answering
by: Shen, Zhixuan, et al.
Published: (2024)
by: Shen, Zhixuan, et al.
Published: (2024)
ThyroidEffi 1.0: A Cost-Effective System for High-Performance Multi-Class Thyroid Carcinoma Classification
by: Pham-Ngoc, Hai, et al.
Published: (2025)
by: Pham-Ngoc, Hai, et al.
Published: (2025)
TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
by: Liu, Yuliang, et al.
Published: (2024)
by: Liu, Yuliang, et al.
Published: (2024)
Adversarial Attack for RGB-Event based Visual Object Tracking
by: Chen, Qiang, et al.
Published: (2025)
by: Chen, Qiang, et al.
Published: (2025)
Lost in OCR Translation? Vision-Based Approaches to Robust Document Retrieval
by: Most, Alexander, et al.
Published: (2025)
by: Most, Alexander, et al.
Published: (2025)
Steal Now and Attack Later: Evaluating Robustness of Object Detection against Black-box Adversarial Attacks
by: Chen, Erh-Chung, et al.
Published: (2024)
by: Chen, Erh-Chung, et al.
Published: (2024)
Emotion Loss Attacking: Adversarial Attack Perception for Skeleton based on Multi-dimensional Features
by: Liu, Feng, et al.
Published: (2024)
by: Liu, Feng, et al.
Published: (2024)
One Noise to Rule Them All: Multi-View Adversarial Attacks with Universal Perturbation
by: Ergezer, Mehmet, et al.
Published: (2024)
by: Ergezer, Mehmet, et al.
Published: (2024)
MonkeyOCR v1.5 Technical Report: Unlocking Robust Document Parsing for Complex Patterns
by: Zhang, Jiarui, et al.
Published: (2025)
by: Zhang, Jiarui, et al.
Published: (2025)
OCR is All you need: Importing Multi-Modality into Image-based Defect Detection System
by: Hsu, Chih-Chung, et al.
Published: (2024)
by: Hsu, Chih-Chung, et al.
Published: (2024)
Adversarial Attack Against Images Classification based on Generative Adversarial Networks
by: Yang, Yahe
Published: (2024)
by: Yang, Yahe
Published: (2024)
SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative Refinement
by: Jain, Chelsi, et al.
Published: (2025)
by: Jain, Chelsi, et al.
Published: (2025)
Comparing Deep Neural Network for Multi-Label ECG Diagnosis From Scanned ECG
by: Nguyen, Cuong V., et al.
Published: (2025)
by: Nguyen, Cuong V., et al.
Published: (2025)
MolX: Enhancing Large Language Models for Molecular Understanding With A Multi-Modal Extension
by: Le, Khiem, et al.
Published: (2024)
by: Le, Khiem, et al.
Published: (2024)
Shedding Light on VLN Robustness: A Black-box Framework for Indoor Lighting-based Adversarial Attack
by: Li, Chenyang, et al.
Published: (2025)
by: Li, Chenyang, et al.
Published: (2025)
CDUPatch: Color-Driven Universal Adversarial Patch Attack for Dual-Modal Visible-Infrared Detectors
by: Long, Jiahuan, et al.
Published: (2025)
by: Long, Jiahuan, et al.
Published: (2025)
Beyond Attack Success Rate: A Multi-Metric Evaluation of Adversarial Transferability in Medical Imaging Models
by: Curl, Emily, et al.
Published: (2026)
by: Curl, Emily, et al.
Published: (2026)
VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding
by: Wu, Yinghao, et al.
Published: (2026)
by: Wu, Yinghao, et al.
Published: (2026)
HDC: Hierarchical Distillation for Multi-level Noisy Consistency in Semi-Supervised Fetal Ultrasound Segmentation
by: Le, Tran Quoc Khanh, et al.
Published: (2025)
by: Le, Tran Quoc Khanh, et al.
Published: (2025)
Enhancing Adversarial Transferability in Visual-Language Pre-training Models via Local Shuffle and Sample-based Attack
by: Liu, Xin, et al.
Published: (2025)
by: Liu, Xin, et al.
Published: (2025)
Concept-based Adversarial Attack: a Probabilistic Perspective
by: Zhang, Andi, et al.
Published: (2025)
by: Zhang, Andi, et al.
Published: (2025)
Confidence-Aware Document OCR Error Detection
by: Hemmer, Arthur, et al.
Published: (2024)
by: Hemmer, Arthur, et al.
Published: (2024)
Filtered-ViT: A Robust Defense Against Multiple Adversarial Patch Attacks
by: Khanal, Aja, et al.
Published: (2025)
by: Khanal, Aja, et al.
Published: (2025)
B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
Rethinking Gradient-based Adversarial Attacks on Point Cloud Classification
by: Chen, Jun, et al.
Published: (2025)
by: Chen, Jun, et al.
Published: (2025)
MADTempo: An Interactive System for Multi-Event Temporal Video Retrieval with Query Augmentation
by: Vu, Huu-An, et al.
Published: (2025)
by: Vu, Huu-An, et al.
Published: (2025)
Probing the Robustness of Vision-Language Pretrained Models: A Multimodal Adversarial Attack Approach
by: Guan, Jiwei, et al.
Published: (2024)
by: Guan, Jiwei, et al.
Published: (2024)
Securing Vision-Language Models with a Robust Encoder Against Jailbreak and Adversarial Attacks
by: Hossain, Md Zarif, et al.
Published: (2024)
by: Hossain, Md Zarif, et al.
Published: (2024)
Enhancing Object Detection Robustness: Detecting and Restoring Confidence in the Presence of Adversarial Patch Attacks
by: Kazoom, Roie, et al.
Published: (2024)
by: Kazoom, Roie, et al.
Published: (2024)
Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding
by: Tang, Jiaqi, et al.
Published: (2025)
by: Tang, Jiaqi, et al.
Published: (2025)
Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments
by: Nguyen, Van Quang
Published: (2026)
by: Nguyen, Van Quang
Published: (2026)
Evaluating OCR performance on food packaging labels in South Africa
by: Nagayi, Mayimunah, et al.
Published: (2025)
by: Nagayi, Mayimunah, et al.
Published: (2025)
Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT
by: Dong, Zhuobai, et al.
Published: (2025)
by: Dong, Zhuobai, et al.
Published: (2025)
On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression
by: Zhang, Xinwei, et al.
Published: (2026)
by: Zhang, Xinwei, et al.
Published: (2026)
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
by: Ding, Yihao, et al.
Published: (2025)
by: Ding, Yihao, et al.
Published: (2025)
OCR-Quality: A Human-Annotated Dataset for OCR Quality Assessment
by: Zhang, Yulong
Published: (2025)
by: Zhang, Yulong
Published: (2025)
PerspectiveNet: Multi-View Perception for Dynamic Scene Understanding
by: Nguyen, Vinh
Published: (2024)
by: Nguyen, Vinh
Published: (2024)
IGL-DT: Iterative Global-Local Feature Learning with Dual-Teacher Semantic Segmentation Framework under Limited Annotation Scheme
by: Tran, Dinh Dai Quan, et al.
Published: (2025)
by: Tran, Dinh Dai Quan, et al.
Published: (2025)
Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR
by: Zhong, Yufeng, et al.
Published: (2025)
by: Zhong, Yufeng, et al.
Published: (2025)
Architecture-Agnostic Modality-Isolated Gated Fusion for Robust Multi-Modal Prostate MRI Segmentation
by: Shu, Yongbo, et al.
Published: (2026)
by: Shu, Yongbo, et al.
Published: (2026)
Similar Items
-
Adversarial Training with OCR Modality Perturbation for Scene-Text Visual Question Answering
by: Shen, Zhixuan, et al.
Published: (2024) -
ThyroidEffi 1.0: A Cost-Effective System for High-Performance Multi-Class Thyroid Carcinoma Classification
by: Pham-Ngoc, Hai, et al.
Published: (2025) -
TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
by: Liu, Yuliang, et al.
Published: (2024) -
Adversarial Attack for RGB-Event based Visual Object Tracking
by: Chen, Qiang, et al.
Published: (2025) -
Lost in OCR Translation? Vision-Based Approaches to Robust Document Retrieval
by: Most, Alexander, et al.
Published: (2025)