An Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Luo, Zhi, Yuan, Zenghui, Wei, Wenqi, Liu, Daizong, Zhou, Pan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Video Is Not Worth a Thousand Words
di: Pollard, Sam, et al.
Pubblicazione: (2025)
di: Pollard, Sam, et al.
Pubblicazione: (2025)
Vript: A Video Is Worth Thousands of Words
di: Yang, Dongjie, et al.
Pubblicazione: (2024)
di: Yang, Dongjie, et al.
Pubblicazione: (2024)
Words Worth a Thousand Pictures: Measuring and Understanding Perceptual Variability in Text-to-Image Generation
di: Tang, Raphael, et al.
Pubblicazione: (2024)
di: Tang, Raphael, et al.
Pubblicazione: (2024)
WordVIS: A Color Worth A Thousand Words
di: Khan, Umar, et al.
Pubblicazione: (2024)
di: Khan, Umar, et al.
Pubblicazione: (2024)
One Image is Worth a Thousand Words: A Usability Preservable Text-Image Collaborative Erasing Framework
di: Li, Feiran, et al.
Pubblicazione: (2025)
di: Li, Feiran, et al.
Pubblicazione: (2025)
Not Every Image is Worth a Thousand Words: Quantifying Originality in Stable Diffusion
di: Haviv, Adi, et al.
Pubblicazione: (2024)
di: Haviv, Adi, et al.
Pubblicazione: (2024)
A LoRA is Worth a Thousand Pictures
di: Liu, Chenxi, et al.
Pubblicazione: (2024)
di: Liu, Chenxi, et al.
Pubblicazione: (2024)
A Label is Worth a Thousand Images in Dataset Distillation
di: Qin, Tian, et al.
Pubblicazione: (2024)
di: Qin, Tian, et al.
Pubblicazione: (2024)
Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models
di: Wang, Jiayu, et al.
Pubblicazione: (2024)
di: Wang, Jiayu, et al.
Pubblicazione: (2024)
A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends
di: Liu, Daizong, et al.
Pubblicazione: (2024)
di: Liu, Daizong, et al.
Pubblicazione: (2024)
Is a Picture Worth a Thousand Words? Adaptive Multimodal Fact-Checking with Visual Evidence Necessity
di: Jung, Jaeyoon, et al.
Pubblicazione: (2026)
di: Jung, Jaeyoon, et al.
Pubblicazione: (2026)
A Task is Worth One Word: Learning with Task Prompts for High-Quality Versatile Image Inpainting
di: Zhuang, Junhao, et al.
Pubblicazione: (2023)
di: Zhuang, Junhao, et al.
Pubblicazione: (2023)
Hard-Label Black-Box Attacks on 3D Point Clouds
di: Liu, Daizong, et al.
Pubblicazione: (2024)
di: Liu, Daizong, et al.
Pubblicazione: (2024)
Video Is Worth a Thousand Images: Exploring the Latest Trends in Long Video Generation
di: Waseem, Faraz, et al.
Pubblicazione: (2024)
di: Waseem, Faraz, et al.
Pubblicazione: (2024)
A Survey on Text-guided 3D Visual Grounding: Elements, Recent Advances, and Future Directions
di: Liu, Daizong, et al.
Pubblicazione: (2024)
di: Liu, Daizong, et al.
Pubblicazione: (2024)
Malware Detection in Docker Containers: An Image is Worth a Thousand Logs
di: Nousias, Akis, et al.
Pubblicazione: (2025)
di: Nousias, Akis, et al.
Pubblicazione: (2025)
Improving the Transferability of 3D Point Cloud Attack via Spectral-aware Admix and Optimization Designs
di: Hu, Shiyu, et al.
Pubblicazione: (2024)
di: Hu, Shiyu, et al.
Pubblicazione: (2024)
A Scene is Worth a Thousand Features: Feed-Forward Camera Localization from a Collection of Image Features
di: Barroso-Laguna, Axel, et al.
Pubblicazione: (2025)
di: Barroso-Laguna, Axel, et al.
Pubblicazione: (2025)
Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images
di: Gao, Kuofeng, et al.
Pubblicazione: (2024)
di: Gao, Kuofeng, et al.
Pubblicazione: (2024)
A Picture is Worth a Thousand Words? An Empirical Study of Aggregation Strategies for Visual Financial Document Retrieval
di: Lim, Ho Hung, et al.
Pubblicazione: (2026)
di: Lim, Ho Hung, et al.
Pubblicazione: (2026)
An Embedding is Worth a Thousand Noisy Labels
di: Di Salvo, Francesco, et al.
Pubblicazione: (2024)
di: Di Salvo, Francesco, et al.
Pubblicazione: (2024)
Joint Top-Down and Bottom-Up Frameworks for 3D Visual Grounding
di: Liu, Yang, et al.
Pubblicazione: (2024)
di: Liu, Yang, et al.
Pubblicazione: (2024)
Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding
di: Huang, Wencan, et al.
Pubblicazione: (2025)
di: Huang, Wencan, et al.
Pubblicazione: (2025)
Modeling Thousands of Human Annotators for Generalizable Text-to-Image Person Re-identification
di: Jiang, Jiayu, et al.
Pubblicazione: (2025)
di: Jiang, Jiayu, et al.
Pubblicazione: (2025)
Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
di: Wang, Feng, et al.
Pubblicazione: (2025)
di: Wang, Feng, et al.
Pubblicazione: (2025)
A Thousand Words or An Image: Studying the Influence of Persona Modality in Multimodal LLMs
di: Broomfield, Julius, et al.
Pubblicazione: (2025)
di: Broomfield, Julius, et al.
Pubblicazione: (2025)
An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models
di: Luo, Haochen, et al.
Pubblicazione: (2024)
di: Luo, Haochen, et al.
Pubblicazione: (2024)
From Words to Wavelengths: VLMs for Few-Shot Multispectral Object Detection
di: Nkegoum, Manuel, et al.
Pubblicazione: (2025)
di: Nkegoum, Manuel, et al.
Pubblicazione: (2025)
A Picture is Worth a Thousand Prompts? Efficacy of Iterative Human-Driven Prompt Refinement in Image Regeneration Tasks
di: Trinh, Khoi, et al.
Pubblicazione: (2025)
di: Trinh, Khoi, et al.
Pubblicazione: (2025)
Images are Worth Variable Length of Representations
di: Mao, Lingjun, et al.
Pubblicazione: (2025)
di: Mao, Lingjun, et al.
Pubblicazione: (2025)
From Part to Whole: 3D Generative World Model with an Adaptive Structural Hierarchy
di: Du, Bi'an, et al.
Pubblicazione: (2026)
di: Du, Bi'an, et al.
Pubblicazione: (2026)
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types
di: Chen, Jiankang, et al.
Pubblicazione: (2025)
di: Chen, Jiankang, et al.
Pubblicazione: (2025)
Energy-Latency Manipulation of Multi-modal Large Language Models via Verbose Samples
di: Gao, Kuofeng, et al.
Pubblicazione: (2024)
di: Gao, Kuofeng, et al.
Pubblicazione: (2024)
Track Any Peppers: Weakly Supervised Sweet Pepper Tracking Using VLMs
di: Lim, Jia Syuen, et al.
Pubblicazione: (2024)
di: Lim, Jia Syuen, et al.
Pubblicazione: (2024)
From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs
di: Cao, Ang, et al.
Pubblicazione: (2025)
di: Cao, Ang, et al.
Pubblicazione: (2025)
When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs
di: Dai, Aobotao, et al.
Pubblicazione: (2025)
di: Dai, Aobotao, et al.
Pubblicazione: (2025)
An Image is Worth 32 Tokens for Reconstruction and Generation
di: Yu, Qihang, et al.
Pubblicazione: (2024)
di: Yu, Qihang, et al.
Pubblicazione: (2024)
Rethinking Transferable Adversarial Attacks on Point Clouds from a Compact Subspace Perspective
di: Tang, Keke, et al.
Pubblicazione: (2026)
di: Tang, Keke, et al.
Pubblicazione: (2026)
RealCustom: Narrowing Real Text Word for Real-Time Open-Domain Text-to-Image Customization
di: Huang, Mengqi, et al.
Pubblicazione: (2024)
di: Huang, Mengqi, et al.
Pubblicazione: (2024)
An Item is Worth a Prompt: Versatile Image Editing with Disentangled Control
di: Feng, Aosong, et al.
Pubblicazione: (2024)
di: Feng, Aosong, et al.
Pubblicazione: (2024)
Documenti analoghi
-
A Video Is Not Worth a Thousand Words
di: Pollard, Sam, et al.
Pubblicazione: (2025) -
Vript: A Video Is Worth Thousands of Words
di: Yang, Dongjie, et al.
Pubblicazione: (2024) -
Words Worth a Thousand Pictures: Measuring and Understanding Perceptual Variability in Text-to-Image Generation
di: Tang, Raphael, et al.
Pubblicazione: (2024) -
WordVIS: A Color Worth A Thousand Words
di: Khan, Umar, et al.
Pubblicazione: (2024) -
One Image is Worth a Thousand Words: A Usability Preservable Text-Image Collaborative Erasing Framework
di: Li, Feiran, et al.
Pubblicazione: (2025)