SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Yue, Xing, Yun, Zhang, Jie, Lin, Di, Zhang, Tianwei, Tsang, Ivor, Liu, Yang, Guo, Qing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IRAD: Implicit Representation-driven Image Resampling against Adversarial Attacks
by: Cao, Yue, et al.
Published: (2023)
by: Cao, Yue, et al.
Published: (2023)
MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents
by: Xing, Yun, et al.
Published: (2024)
by: Xing, Yun, et al.
Published: (2024)
Visible Yet Unreadable: A Systematic Blind Spot of Vision Language Models Across Writing Systems
by: Zhang, Jie, et al.
Published: (2025)
by: Zhang, Jie, et al.
Published: (2025)
DepthVanish: Optimizing Adversarial Interval Structures for Stereo-Depth-Invisible Patches
by: Xing, Yun, et al.
Published: (2025)
by: Xing, Yun, et al.
Published: (2025)
Boosting Transferability in Vision-Language Attacks via Diversification along the Intersection Region of Adversarial Trajectory
by: Gao, Sensen, et al.
Published: (2024)
by: Gao, Sensen, et al.
Published: (2024)
Unveiling Typographic Deceptions: Insights of the Typographic Vulnerability in Large Vision-Language Model
by: Cheng, Hao, et al.
Published: (2024)
by: Cheng, Hao, et al.
Published: (2024)
Vision-Language Integration for Zero-Shot Scene Understanding in Real-World Environments
by: Rajiv, Manjunath Prasad Holenarasipura, et al.
Published: (2025)
by: Rajiv, Manjunath Prasad Holenarasipura, et al.
Published: (2025)
Semantic-Aligned Adversarial Evolution Triangle for High-Transferability Vision-Language Attack
by: Jia, Xiaojun, et al.
Published: (2024)
by: Jia, Xiaojun, et al.
Published: (2024)
HC$^2$L: Hybrid and Cooperative Contrastive Learning for Cross-lingual Spoken Language Understanding
by: Xing, Bowen, et al.
Published: (2024)
by: Xing, Bowen, et al.
Published: (2024)
Unifying Watermarking via Dimension-Aware Mapping
by: Meng, Jiale, et al.
Published: (2026)
by: Meng, Jiale, et al.
Published: (2026)
Time-variant Image Inpainting via Interactive Distribution Transition Estimation
by: Xing, Yun, et al.
Published: (2025)
by: Xing, Yun, et al.
Published: (2025)
SceneComplete: Open-World 3D Scene Completion in Cluttered Real World Environments for Robot Manipulation
by: Agarwal, Aditya, et al.
Published: (2024)
by: Agarwal, Aditya, et al.
Published: (2024)
FocusDD: Real-World Scene Infusion for Robust Dataset Distillation
by: Hu, Youbing, et al.
Published: (2025)
by: Hu, Youbing, et al.
Published: (2025)
Advancing Analytic Class-Incremental Learning through Vision-Language Calibration
by: Zhao, Binyu, et al.
Published: (2026)
by: Zhao, Binyu, et al.
Published: (2026)
FOCUS: Frequency-Optimized Conditioning of DiffUSion Models for mitigating catastrophic forgetting during Test-Time Adaptation
by: Tjio, Gabriel, et al.
Published: (2025)
by: Tjio, Gabriel, et al.
Published: (2025)
Real-World Scene Recovery for Scattering-Degraded Images Using Spatial and Frequency Priors
by: Liu, Yun, et al.
Published: (2025)
by: Liu, Yun, et al.
Published: (2025)
Beyond Pixels: Semantic-aware Typographic Attack for Geo-Privacy Protection
by: Zhu, Jiayi, et al.
Published: (2025)
by: Zhu, Jiayi, et al.
Published: (2025)
Generating Multimodal Driving Scenes via Next-Scene Prediction
by: Wu, Yanhao, et al.
Published: (2025)
by: Wu, Yanhao, et al.
Published: (2025)
Reasoning Robustness of LLMs to Adversarial Typographical Errors
by: Gan, Esther, et al.
Published: (2024)
by: Gan, Esther, et al.
Published: (2024)
SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
by: Jia, Baoxiong, et al.
Published: (2024)
by: Jia, Baoxiong, et al.
Published: (2024)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
by: Wang, Zhaowei, et al.
Published: (2024)
by: Wang, Zhaowei, et al.
Published: (2024)
Coherence-guided Preference Disentanglement for Cross-domain Recommendations
by: Xiang, Zongyi, et al.
Published: (2024)
by: Xiang, Zongyi, et al.
Published: (2024)
Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models
by: Cheng, Hao, et al.
Published: (2024)
by: Cheng, Hao, et al.
Published: (2024)
Causal Scene Narration with Runtime Safety Supervision for Vision-Language-Action Driving
by: Li, Yun, et al.
Published: (2026)
by: Li, Yun, et al.
Published: (2026)
Energy-based Models are Zero-Shot Planners for Compositional Scene Rearrangement
by: Gkanatsios, Nikolaos, et al.
Published: (2023)
by: Gkanatsios, Nikolaos, et al.
Published: (2023)
UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes
by: Ni, Shuo, et al.
Published: (2025)
by: Ni, Shuo, et al.
Published: (2025)
SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining
by: Li, Yue, et al.
Published: (2025)
by: Li, Yue, et al.
Published: (2025)
Transductive Reward Inference on Graph
by: Qu, Bohao, et al.
Published: (2024)
by: Qu, Bohao, et al.
Published: (2024)
AnthroTAP: Learning Point Tracking with Real-World Motion
by: Kim, Inès Hyeonsu, et al.
Published: (2025)
by: Kim, Inès Hyeonsu, et al.
Published: (2025)
SCAM: A Real-World Typographic Robustness Evaluation for Multimodal Foundation Models
by: Westerhoff, Justus, et al.
Published: (2025)
by: Westerhoff, Justus, et al.
Published: (2025)
Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving
by: Gao, Haoxiang, et al.
Published: (2025)
by: Gao, Haoxiang, et al.
Published: (2025)
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
by: Qi, Zhangyang, et al.
Published: (2025)
by: Qi, Zhangyang, et al.
Published: (2025)
Disrupting Vision-Language Model-Driven Navigation Services via Adversarial Object Fusion
by: Xie, Chunlong, et al.
Published: (2025)
by: Xie, Chunlong, et al.
Published: (2025)
Open World Scene Graph Generation using Vision Language Models
by: Dutta, Amartya, et al.
Published: (2025)
by: Dutta, Amartya, et al.
Published: (2025)
MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models
by: Feng, Jun, et al.
Published: (2025)
by: Feng, Jun, et al.
Published: (2025)
Off-dynamics Conditional Diffusion Planners
by: Ng, Wen Zheng Terence, et al.
Published: (2024)
by: Ng, Wen Zheng Terence, et al.
Published: (2024)
TravelPlanner: A Benchmark for Real-World Planning with Language Agents
by: Xie, Jian, et al.
Published: (2024)
by: Xie, Jian, et al.
Published: (2024)
Dynamic Context-Aware Scene Reasoning Using Vision-Language Alignment in Zero-Shot Real-World Scenarios
by: Rajiv, Manjunath Prasad Holenarasipura, et al.
Published: (2025)
by: Rajiv, Manjunath Prasad Holenarasipura, et al.
Published: (2025)
Leveraging Scene Context with Dual Networks for Sequential User Behavior Modeling
by: Chen, Xu, et al.
Published: (2025)
by: Chen, Xu, et al.
Published: (2025)
DynamicPAE: Generating Scene-Aware Physical Adversarial Examples in Real-Time
by: Hu, Jin, et al.
Published: (2024)
by: Hu, Jin, et al.
Published: (2024)
Similar Items
-
IRAD: Implicit Representation-driven Image Resampling against Adversarial Attacks
by: Cao, Yue, et al.
Published: (2023) -
MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents
by: Xing, Yun, et al.
Published: (2024) -
Visible Yet Unreadable: A Systematic Blind Spot of Vision Language Models Across Writing Systems
by: Zhang, Jie, et al.
Published: (2025) -
DepthVanish: Optimizing Adversarial Interval Structures for Stereo-Depth-Invisible Patches
by: Xing, Yun, et al.
Published: (2025) -
Boosting Transferability in Vision-Language Attacks via Diversification along the Intersection Region of Adversarial Trajectory
by: Gao, Sensen, et al.
Published: (2024)