On the Adversarial Robustness of Discrete Image Tokenizers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bhagwatkar, Rishika, Rish, Irina, Flammarion, Nicolas, Croce, Francesco |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Adversarially Robust Vision-Language Models: Insights from Design Choices and Prompt Formatting Techniques
von: Bhagwatkar, Rishika, et al.
Veröffentlicht: (2024)
von: Bhagwatkar, Rishika, et al.
Veröffentlicht: (2024)
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
von: Schlarmann, Christian, et al.
Veröffentlicht: (2025)
von: Schlarmann, Christian, et al.
Veröffentlicht: (2025)
CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
von: Bhagwatkar, Rishika, et al.
Veröffentlicht: (2025)
von: Bhagwatkar, Rishika, et al.
Veröffentlicht: (2025)
Image Tiling for High-Resolution Reasoning: Balancing Local Detail with Global Context
von: de Margerie, Anatole Jacquin, et al.
Veröffentlicht: (2025)
von: de Margerie, Anatole Jacquin, et al.
Veröffentlicht: (2025)
Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models
von: Schlarmann, Christian, et al.
Veröffentlicht: (2024)
von: Schlarmann, Christian, et al.
Veröffentlicht: (2024)
Robustness Tokens: Towards Adversarial Robustness of Transformers
von: Pulfer, Brian, et al.
Veröffentlicht: (2025)
von: Pulfer, Brian, et al.
Veröffentlicht: (2025)
On the Out-of-Distribution Generalization of Reasoning in Multimodal LLMs for Simple Visual Planning Tasks
von: Neuhaus, Yannic, et al.
Veröffentlicht: (2026)
von: Neuhaus, Yannic, et al.
Veröffentlicht: (2026)
Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing
von: Wang, Weixing, et al.
Veröffentlicht: (2025)
von: Wang, Weixing, et al.
Veröffentlicht: (2025)
MacTok: Robust Continuous Tokenization for Image Generation
von: Zeng, Hengyu, et al.
Veröffentlicht: (2026)
von: Zeng, Hengyu, et al.
Veröffentlicht: (2026)
Parameter Interpolation Adversarial Training for Robust Image Classification
von: Liu, Xin, et al.
Veröffentlicht: (2025)
von: Liu, Xin, et al.
Veröffentlicht: (2025)
CODA: Repurposing Continuous VAEs for Discrete Tokenization
von: Liu, Zeyu, et al.
Veröffentlicht: (2025)
von: Liu, Zeyu, et al.
Veröffentlicht: (2025)
Fine-Tuning Adversarially-Robust Transformers for Single-Image Dehazing
von: Vasilescu, Vlad, et al.
Veröffentlicht: (2025)
von: Vasilescu, Vlad, et al.
Veröffentlicht: (2025)
MIMIR: Masked Image Modeling for Mutual Information-based Adversarial Robustness
von: Xu, Xiaoyun, et al.
Veröffentlicht: (2023)
von: Xu, Xiaoyun, et al.
Veröffentlicht: (2023)
Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2025)
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2025)
TITAN: Query-Token based Domain Adaptive Adversarial Learning
von: Ashraf, Tajamul, et al.
Veröffentlicht: (2025)
von: Ashraf, Tajamul, et al.
Veröffentlicht: (2025)
Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
von: Kim, Dongwon, et al.
Veröffentlicht: (2026)
von: Kim, Dongwon, et al.
Veröffentlicht: (2026)
On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression
von: Zhang, Xinwei, et al.
Veröffentlicht: (2026)
von: Zhang, Xinwei, et al.
Veröffentlicht: (2026)
Robust SAM: On the Adversarial Robustness of Vision Foundation Models
von: Long, Jiahuan, et al.
Veröffentlicht: (2025)
von: Long, Jiahuan, et al.
Veröffentlicht: (2025)
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization
von: Tan, Zhentao, et al.
Veröffentlicht: (2024)
von: Tan, Zhentao, et al.
Veröffentlicht: (2024)
InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression
von: Ye, Haotian, et al.
Veröffentlicht: (2025)
von: Ye, Haotian, et al.
Veröffentlicht: (2025)
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
von: Qu, Liao, et al.
Veröffentlicht: (2024)
von: Qu, Liao, et al.
Veröffentlicht: (2024)
TIMA: Text-Image Mutual Awareness for Balancing Zero-Shot Adversarial Robustness and Generalization Ability
von: Ma, Fengji, et al.
Veröffentlicht: (2024)
von: Ma, Fengji, et al.
Veröffentlicht: (2024)
GaussianToken: An Effective Image Tokenizer with 2D Gaussian Splatting
von: Dong, Jiajun, et al.
Veröffentlicht: (2025)
von: Dong, Jiajun, et al.
Veröffentlicht: (2025)
Adversarially-Refined VQ-GAN with Dense Motion Tokenization for Spatio-Temporal Heatmaps
von: Maldonado, Gabriel, et al.
Veröffentlicht: (2025)
von: Maldonado, Gabriel, et al.
Veröffentlicht: (2025)
Homogeneous Tokenizer Matters: Homogeneous Visual Tokenizer for Remote Sensing Image Understanding
von: Shao, Run, et al.
Veröffentlicht: (2024)
von: Shao, Run, et al.
Veröffentlicht: (2024)
Negative Token Merging: Image-based Adversarial Feature Guidance
von: Singh, Jaskirat, et al.
Veröffentlicht: (2024)
von: Singh, Jaskirat, et al.
Veröffentlicht: (2024)
CIARD: Cyclic Iterative Adversarial Robustness Distillation
von: Lu, Liming, et al.
Veröffentlicht: (2025)
von: Lu, Liming, et al.
Veröffentlicht: (2025)
On Inherent Adversarial Robustness of Active Vision Systems
von: Mukherjee, Amitangshu, et al.
Veröffentlicht: (2024)
von: Mukherjee, Amitangshu, et al.
Veröffentlicht: (2024)
Revisiting the Robust Generalization of Adversarial Prompt Tuning
von: Yang, Fan, et al.
Veröffentlicht: (2024)
von: Yang, Fan, et al.
Veröffentlicht: (2024)
Adversarial Attack Against Images Classification based on Generative Adversarial Networks
von: Yang, Yahe
Veröffentlicht: (2024)
von: Yang, Yahe
Veröffentlicht: (2024)
Frequency Autoregressive Image Generation with Continuous Tokens
von: Yu, Hu, et al.
Veröffentlicht: (2025)
von: Yu, Hu, et al.
Veröffentlicht: (2025)
Hita: Holistic Tokenizer for Autoregressive Image Generation
von: Zheng, Anlin, et al.
Veröffentlicht: (2025)
von: Zheng, Anlin, et al.
Veröffentlicht: (2025)
Scaling Image Tokenizers with Grouped Spherical Quantization
von: Wang, Jiangtao, et al.
Veröffentlicht: (2024)
von: Wang, Jiangtao, et al.
Veröffentlicht: (2024)
Scalable Image Tokenization with Index Backpropagation Quantization
von: Shi, Fengyuan, et al.
Veröffentlicht: (2024)
von: Shi, Fengyuan, et al.
Veröffentlicht: (2024)
CHIRP: A Fine-Grained Benchmark for Open-Ended Response Evaluation in Vision-Language Models
von: Roger, Alexis, et al.
Veröffentlicht: (2025)
von: Roger, Alexis, et al.
Veröffentlicht: (2025)
Assessing Robustness via Score-Based Adversarial Image Generation
von: Kollovieh, Marcel, et al.
Veröffentlicht: (2023)
von: Kollovieh, Marcel, et al.
Veröffentlicht: (2023)
VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector Quantization
von: Zhang, Yiwei, et al.
Veröffentlicht: (2024)
von: Zhang, Yiwei, et al.
Veröffentlicht: (2024)
Discriminative Class Tokens for Text-to-Image Diffusion Models
von: Schwartz, Idan, et al.
Veröffentlicht: (2023)
von: Schwartz, Idan, et al.
Veröffentlicht: (2023)
xT: Nested Tokenization for Larger Context in Large Images
von: Gupta, Ritwik, et al.
Veröffentlicht: (2024)
von: Gupta, Ritwik, et al.
Veröffentlicht: (2024)
TMCIR: Token Merge Benefits Composed Image Retrieval
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards Adversarially Robust Vision-Language Models: Insights from Design Choices and Prompt Formatting Techniques
von: Bhagwatkar, Rishika, et al.
Veröffentlicht: (2024) -
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
von: Schlarmann, Christian, et al.
Veröffentlicht: (2025) -
CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
von: Bhagwatkar, Rishika, et al.
Veröffentlicht: (2025) -
Image Tiling for High-Resolution Reasoning: Balancing Local Detail with Global Context
von: de Margerie, Anatole Jacquin, et al.
Veröffentlicht: (2025) -
Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models
von: Schlarmann, Christian, et al.
Veröffentlicht: (2024)