Skeletonization-Based Adversarial Perturbations on Large Vision Language Model's Mathematical Text Recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Yoshida, Masatomo, Namura, Haruto, Adami, Nicola, Okuda, Masahiro |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SkeletonX: Data-Efficient Skeleton-based Action Recognition via Cross-sample Feature Aggregation
por: Zhang, Zongye, et al.
Publicado: (2025)
por: Zhang, Zongye, et al.
Publicado: (2025)
DVGBench: Implicit-to-Explicit Visual Grounding Benchmark in UAV Imagery with Large Vision-Language Models
por: Zhou, Yue, et al.
Publicado: (2026)
por: Zhou, Yue, et al.
Publicado: (2026)
TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models
por: Wang, Zeqing, et al.
Publicado: (2025)
por: Wang, Zeqing, et al.
Publicado: (2025)
SAM Encoder Breach by Adversarial Simplicial Complex Triggers Downstream Model Failures
por: Qin, Yi, et al.
Publicado: (2025)
por: Qin, Yi, et al.
Publicado: (2025)
EarthVL: A Progressive Earth Vision-Language Understanding and Generation Framework
por: Wang, Junjue, et al.
Publicado: (2026)
por: Wang, Junjue, et al.
Publicado: (2026)
MetaCloak-JPEG: JPEG-Robust Adversarial Perturbation for Preventing Unauthorized DreamBooth-Based Deepfake Generation
por: Fardin, Tanjim Rahaman, et al.
Publicado: (2026)
por: Fardin, Tanjim Rahaman, et al.
Publicado: (2026)
Model Agnostic Defense against Adversarial Patch Attacks on Object Detection in Unmanned Aerial Vehicles
por: Pathak, Saurabh, et al.
Publicado: (2024)
por: Pathak, Saurabh, et al.
Publicado: (2024)
DVLA-RL: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning
por: Li, Wenhao, et al.
Publicado: (2026)
por: Li, Wenhao, et al.
Publicado: (2026)
DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response
por: Wang, Junjue, et al.
Publicado: (2025)
por: Wang, Junjue, et al.
Publicado: (2025)
Fingerprint Membership and Identity Inference Against Generative Adversarial Networks
por: Cavasin, Saverio, et al.
Publicado: (2024)
por: Cavasin, Saverio, et al.
Publicado: (2024)
Group Activity Recognition using Unreliable Tracked Pose
por: Thilakarathne, Haritha, et al.
Publicado: (2024)
por: Thilakarathne, Haritha, et al.
Publicado: (2024)
How do Foundation Models Compare to Skeleton-Based Approaches for Gesture Recognition in Human-Robot Interaction?
por: Käs, Stephanie, et al.
Publicado: (2025)
por: Käs, Stephanie, et al.
Publicado: (2025)
VersaGen: Unleashing Versatile Visual Control for Text-to-Image Synthesis
por: Chen, Zhipeng, et al.
Publicado: (2024)
por: Chen, Zhipeng, et al.
Publicado: (2024)
Evaluating the Significance of Outdoor Advertising from Driver's Perspective Using Computer Vision
por: Černeková, Zuzana, et al.
Publicado: (2023)
por: Černeková, Zuzana, et al.
Publicado: (2023)
Facial Attribute Based Text Guided Face Anonymization
por: Muştu, Mustafa İzzet, et al.
Publicado: (2025)
por: Muştu, Mustafa İzzet, et al.
Publicado: (2025)
Scaling Large Vision-Language Models for Enhanced Multimodal Comprehension In Biomedical Image Analysis
por: Umeike, Robinson, et al.
Publicado: (2025)
por: Umeike, Robinson, et al.
Publicado: (2025)
VT-FSL: Bridging Vision and Text with LLMs for Few-Shot Learning
por: Li, Wenhao, et al.
Publicado: (2025)
por: Li, Wenhao, et al.
Publicado: (2025)
Data Augmentation in Earth Observation: A Diffusion Model Approach
por: Sousa, Tiago, et al.
Publicado: (2024)
por: Sousa, Tiago, et al.
Publicado: (2024)
Doodle Your Keypoints: Sketch-Based Few-Shot Keypoint Detection
por: Maity, Subhajit, et al.
Publicado: (2025)
por: Maity, Subhajit, et al.
Publicado: (2025)
Video-Based Human Pose Regression via Decoupled Space-Time Aggregation
por: He, Jijie, et al.
Publicado: (2024)
por: He, Jijie, et al.
Publicado: (2024)
HyperFM: An Efficient Hyperspectral Foundation Model with Spectral Grouping
por: Tushar, Zahid Hassan, et al.
Publicado: (2026)
por: Tushar, Zahid Hassan, et al.
Publicado: (2026)
Efficient Diffusion Models: A Comprehensive Survey from Principles to Practices
por: Ma, Zhiyuan, et al.
Publicado: (2024)
por: Ma, Zhiyuan, et al.
Publicado: (2024)
Traffic Scene Small Target Detection Method Based on YOLOv8n-SPTS Model for Autonomous Driving
por: Wu, Songhan
Publicado: (2025)
por: Wu, Songhan
Publicado: (2025)
MdaIF: Robust One-Stop Multi-Degradation-Aware Image Fusion with Language-Driven Semantics
por: Li, Jing, et al.
Publicado: (2025)
por: Li, Jing, et al.
Publicado: (2025)
Diffusion-Based Ukrainian Handwritten Text Generation with Cross-Domain Style Transfer
por: Ahitoliev, Andrii, et al.
Publicado: (2026)
por: Ahitoliev, Andrii, et al.
Publicado: (2026)
Grounding Synthetic Data Generation With Vision and Language Models
por: Çağlar, Ümit Mert, et al.
Publicado: (2026)
por: Çağlar, Ümit Mert, et al.
Publicado: (2026)
Rapid Adaptation of Earth Observation Foundation Models for Segmentation
por: Selvam, Karthick Panner, et al.
Publicado: (2024)
por: Selvam, Karthick Panner, et al.
Publicado: (2024)
DCT-HistoTransformer: Efficient Lightweight Vision Transformer with DCT Integration for histopathological image analysis
por: Ranjbar, Mahtab, et al.
Publicado: (2024)
por: Ranjbar, Mahtab, et al.
Publicado: (2024)
Optimal Blackjack Strategy Recommender: A Comprehensive Study on Computer Vision Integration for Enhanced Gameplay
por: Gupta, Krishnanshu, et al.
Publicado: (2024)
por: Gupta, Krishnanshu, et al.
Publicado: (2024)
SPEAK: Speech-Driven Pose and Emotion-Adjustable Talking Head Generation
por: Cai, Changpeng, et al.
Publicado: (2024)
por: Cai, Changpeng, et al.
Publicado: (2024)
Human-Centric Perception for Child Sexual Abuse Imagery
por: Laranjeira, Camila, et al.
Publicado: (2026)
por: Laranjeira, Camila, et al.
Publicado: (2026)
Synthetic Image Detection with CLIP: Understanding and Assessing Predictive Cues
por: Willi, Marco, et al.
Publicado: (2026)
por: Willi, Marco, et al.
Publicado: (2026)
Robust Multi-Source Covid-19 Detection in CT Images
por: Pritha, Asmita Yuki, et al.
Publicado: (2026)
por: Pritha, Asmita Yuki, et al.
Publicado: (2026)
Investigation of cardinality classification for bacterial colony counting using explainable artificial intelligence
por: Zheng, Minghua, et al.
Publicado: (2026)
por: Zheng, Minghua, et al.
Publicado: (2026)
A Multi-purpose Tracking Framework for Salmon Welfare Monitoring in Challenging Environments
por: Høgstedt, Espen Uri, et al.
Publicado: (2025)
por: Høgstedt, Espen Uri, et al.
Publicado: (2025)
Learning to count small and clustered objects with application to bacterial colonies
por: Zheng, Minghua, et al.
Publicado: (2026)
por: Zheng, Minghua, et al.
Publicado: (2026)
AVadCLIP: Audio-Visual Collaboration for Robust Video Anomaly Detection
por: Wu, Peng, et al.
Publicado: (2025)
por: Wu, Peng, et al.
Publicado: (2025)
Decoupled Sensitivity-Consistency Learning for Weakly Supervised Video Anomaly Detection
por: Zheng, Hantao, et al.
Publicado: (2026)
por: Zheng, Hantao, et al.
Publicado: (2026)
AOI-SSL: Self-Supervised Framework for Efficient Segmentation of Wire-bonded Semiconductors In Optical Inspection
por: Figueira, Joaquín, et al.
Publicado: (2026)
por: Figueira, Joaquín, et al.
Publicado: (2026)
Practical End-to-End Optical Music Recognition for Pianoform Music
por: Mayer, Jiří, et al.
Publicado: (2024)
por: Mayer, Jiří, et al.
Publicado: (2024)
Ejemplares similares
-
SkeletonX: Data-Efficient Skeleton-based Action Recognition via Cross-sample Feature Aggregation
por: Zhang, Zongye, et al.
Publicado: (2025) -
DVGBench: Implicit-to-Explicit Visual Grounding Benchmark in UAV Imagery with Large Vision-Language Models
por: Zhou, Yue, et al.
Publicado: (2026) -
TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models
por: Wang, Zeqing, et al.
Publicado: (2025) -
SAM Encoder Breach by Adversarial Simplicial Complex Triggers Downstream Model Failures
por: Qin, Yi, et al.
Publicado: (2025) -
EarthVL: A Progressive Earth Vision-Language Understanding and Generation Framework
por: Wang, Junjue, et al.
Publicado: (2026)