SCAM: A Real-World Typographic Robustness Evaluation for Multimodal Foundation Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Westerhoff, Justus, Purelku, Erblina, Hackstein, Jakob, Loos, Jonas, Pinetzki, Leo, Rodner, Erik, Hufe, Lorenz |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
di: Hufe, Lorenz, et al.
Pubblicazione: (2025)
di: Hufe, Lorenz, et al.
Pubblicazione: (2025)
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
di: Dreyer, Maximilian, et al.
Pubblicazione: (2024)
di: Dreyer, Maximilian, et al.
Pubblicazione: (2024)
Latent Diffusion U-Net Representations Contain Positional Embeddings and Anomalies
di: Loos, Jonas, et al.
Pubblicazione: (2025)
di: Loos, Jonas, et al.
Pubblicazione: (2025)
Imbalanced Classification through the Lens of Spurious Correlations
di: Hackstein, Jakob, et al.
Pubblicazione: (2025)
di: Hackstein, Jakob, et al.
Pubblicazione: (2025)
On the Domain Robustness of Contrastive Vision-Language Models
di: Koddenbrock, Mario, et al.
Pubblicazione: (2025)
di: Koddenbrock, Mario, et al.
Pubblicazione: (2025)
Unveiling Typographic Deceptions: Insights of the Typographic Vulnerability in Large Vision-Language Model
di: Cheng, Hao, et al.
Pubblicazione: (2024)
di: Cheng, Hao, et al.
Pubblicazione: (2024)
Feedback-driven object detection and iterative model improvement
di: Tenckhoff, Sönke, et al.
Pubblicazione: (2024)
di: Tenckhoff, Sönke, et al.
Pubblicazione: (2024)
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments
di: Cao, Yue, et al.
Pubblicazione: (2024)
di: Cao, Yue, et al.
Pubblicazione: (2024)
Typographic Text Generation with Off-the-Shelf Diffusion Model
di: Peong, KhayTze, et al.
Pubblicazione: (2024)
di: Peong, KhayTze, et al.
Pubblicazione: (2024)
Exploring Masked Autoencoders for Sensor-Agnostic Image Retrieval in Remote Sensing
di: Hackstein, Jakob, et al.
Pubblicazione: (2024)
di: Hackstein, Jakob, et al.
Pubblicazione: (2024)
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models
di: Waseda, Futa, et al.
Pubblicazione: (2025)
di: Waseda, Futa, et al.
Pubblicazione: (2025)
Automatic Text Box Placement for Supporting Typographic Design
di: Muraoka, Jun, et al.
Pubblicazione: (2025)
di: Muraoka, Jun, et al.
Pubblicazione: (2025)
Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models
di: Cheng, Hao, et al.
Pubblicazione: (2024)
di: Cheng, Hao, et al.
Pubblicazione: (2024)
Adversarial Robustness of AI-Generated Image Detectors in the Real World
di: Mavali, Sina, et al.
Pubblicazione: (2024)
di: Mavali, Sina, et al.
Pubblicazione: (2024)
Beyond Pixels: Semantic-aware Typographic Attack for Geo-Privacy Protection
di: Zhu, Jiayi, et al.
Pubblicazione: (2025)
di: Zhu, Jiayi, et al.
Pubblicazione: (2025)
Exploring Typographic Visual Prompts Injection Threats in Cross-Modality Generation Models
di: Cheng, Hao, et al.
Pubblicazione: (2025)
di: Cheng, Hao, et al.
Pubblicazione: (2025)
Depth Supervised Neural Surface Reconstruction from Airborne Imagery
di: Hackstein, Vincent, et al.
Pubblicazione: (2024)
di: Hackstein, Vincent, et al.
Pubblicazione: (2024)
MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks
di: Chen, Jiacheng, et al.
Pubblicazione: (2024)
di: Chen, Jiacheng, et al.
Pubblicazione: (2024)
Multi-View Foundation Models
di: Segre, Leo, et al.
Pubblicazione: (2025)
di: Segre, Leo, et al.
Pubblicazione: (2025)
Boosting General Trimap-free Matting in the Real-World Image
di: Zhao, Leo Shan Wenzhang Zhou Grace
Pubblicazione: (2024)
di: Zhao, Leo Shan Wenzhang Zhou Grace
Pubblicazione: (2024)
Hybrid Attention for Robust RGB-T Pedestrian Detection in Real-World Conditions
di: Rathinam, Arunkumar, et al.
Pubblicazione: (2024)
di: Rathinam, Arunkumar, et al.
Pubblicazione: (2024)
Robust Weight Imprinting: Insights from Neural Collapse and Proxy-Based Aggregation
di: Westerhoff, Justus, et al.
Pubblicazione: (2025)
di: Westerhoff, Justus, et al.
Pubblicazione: (2025)
HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents
di: X, Tencent Robotics, et al.
Pubblicazione: (2026)
di: X, Tencent Robotics, et al.
Pubblicazione: (2026)
Toward Robust Multimodal Learning using Multimodal Foundational Models
di: Zhao, Xianbing, et al.
Pubblicazione: (2024)
di: Zhao, Xianbing, et al.
Pubblicazione: (2024)
Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models
di: Balakrishnan, Ravikumar, et al.
Pubblicazione: (2026)
di: Balakrishnan, Ravikumar, et al.
Pubblicazione: (2026)
Is Visual in-Context Learning for Compositional Medical Tasks within Reach?
di: Reiß, Simon, et al.
Pubblicazione: (2025)
di: Reiß, Simon, et al.
Pubblicazione: (2025)
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
di: Hong, Jack, et al.
Pubblicazione: (2025)
di: Hong, Jack, et al.
Pubblicazione: (2025)
VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing
di: Deng, Andong, et al.
Pubblicazione: (2026)
di: Deng, Andong, et al.
Pubblicazione: (2026)
Learning to Fuse: Modality-Aware Adaptive Scheduling for Robust Multimodal Foundation Models
di: Bennett, Liam, et al.
Pubblicazione: (2025)
di: Bennett, Liam, et al.
Pubblicazione: (2025)
A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning
di: Chen, Tianle, et al.
Pubblicazione: (2026)
di: Chen, Tianle, et al.
Pubblicazione: (2026)
Foundation Model-oriented Robustness: Robust Image Model Evaluation with Pretrained Models
di: Zhang, Peiyan, et al.
Pubblicazione: (2023)
di: Zhang, Peiyan, et al.
Pubblicazione: (2023)
Boosting Multi-View Stereo with Depth Foundation Model in the Absence of Real-World Labels
di: Zhu, Jie, et al.
Pubblicazione: (2025)
di: Zhu, Jie, et al.
Pubblicazione: (2025)
Percept, Chat, and then Adapt: Multimodal Knowledge Transfer of Foundation Models for Open-World Video Recognition
di: Chen, Boyu, et al.
Pubblicazione: (2024)
di: Chen, Boyu, et al.
Pubblicazione: (2024)
Are Foundation Models Ready for Industrial Defect Recognition? A Reality Check on Real-World Data
di: Baeuerle, Simon, et al.
Pubblicazione: (2025)
di: Baeuerle, Simon, et al.
Pubblicazione: (2025)
WorldEval: World Model as Real-World Robot Policies Evaluator
di: Li, Yaxuan, et al.
Pubblicazione: (2025)
di: Li, Yaxuan, et al.
Pubblicazione: (2025)
One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations
di: Balakrishnan, Ravikumar, et al.
Pubblicazione: (2026)
di: Balakrishnan, Ravikumar, et al.
Pubblicazione: (2026)
Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks
di: Qraitem, Maan, et al.
Pubblicazione: (2024)
di: Qraitem, Maan, et al.
Pubblicazione: (2024)
Bridging the Gap Between Multimodal Foundation Models and World Models
di: He, Xuehai
Pubblicazione: (2025)
di: He, Xuehai
Pubblicazione: (2025)
On the Real-World Adversarial Robustness of Real-Time Semantic Segmentation Models for Autonomous Driving
di: Rossolini, Giulio, et al.
Pubblicazione: (2022)
di: Rossolini, Giulio, et al.
Pubblicazione: (2022)
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
di: Zhang, Yi-Fan, et al.
Pubblicazione: (2024)
di: Zhang, Yi-Fan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
di: Hufe, Lorenz, et al.
Pubblicazione: (2025) -
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
di: Dreyer, Maximilian, et al.
Pubblicazione: (2024) -
Latent Diffusion U-Net Representations Contain Positional Embeddings and Anomalies
di: Loos, Jonas, et al.
Pubblicazione: (2025) -
Imbalanced Classification through the Lens of Spurious Correlations
di: Hackstein, Jakob, et al.
Pubblicazione: (2025) -
On the Domain Robustness of Contrastive Vision-Language Models
di: Koddenbrock, Mario, et al.
Pubblicazione: (2025)