Hidden in Plain Sight: Evaluating Abstract Shape Recognition in Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Hemmat, Arshia, Davies, Adam, Lamb, Tom A., Yuan, Jianhao, Torr, Philip, Khakzar, Ashkan, Pinto, Francesco |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Not Just Pretty Pictures: Toward Interventional Data Augmentation Using Text-to-Image Generators
di: Yuan, Jianhao, et al.
Pubblicazione: (2022)
di: Yuan, Jianhao, et al.
Pubblicazione: (2022)
How Visual Representations Map to Language Feature Space in Multimodal LLMs
di: Venhoff, Constantin, et al.
Pubblicazione: (2025)
di: Venhoff, Constantin, et al.
Pubblicazione: (2025)
Learning Visual Prompts for Guiding the Attention of Vision Transformers
di: Rezaei, Razieh, et al.
Pubblicazione: (2024)
di: Rezaei, Razieh, et al.
Pubblicazione: (2024)
Articulate3D: Zero-Shot Text-Driven 3D Object Posing
di: Deb, Oishi, et al.
Pubblicazione: (2025)
di: Deb, Oishi, et al.
Pubblicazione: (2025)
Learnable Sparsity for Vision Generative Models
di: Zhang, Yang, et al.
Pubblicazione: (2024)
di: Zhang, Yang, et al.
Pubblicazione: (2024)
Latent Guard: a Safety Framework for Text-to-image Generation
di: Liu, Runtao, et al.
Pubblicazione: (2024)
di: Liu, Runtao, et al.
Pubblicazione: (2024)
Hidden Meanings in Plain Sight: RebusBench for Evaluating Cognitive Visual Reasoning
di: Kasaei, Seyed Amir, et al.
Pubblicazione: (2026)
di: Kasaei, Seyed Amir, et al.
Pubblicazione: (2026)
Towards Understanding Multimodal Fine-Tuning: Spatial Features
di: Naghashyar, Lachin, et al.
Pubblicazione: (2026)
di: Naghashyar, Lachin, et al.
Pubblicazione: (2026)
Minimalist Concept Erasure in Generative Models
di: Zhang, Yang, et al.
Pubblicazione: (2025)
di: Zhang, Yang, et al.
Pubblicazione: (2025)
Hidden in Plain Sight -- Class Competition Focuses Attribution Maps
di: Walter, Nils Philipp, et al.
Pubblicazione: (2025)
di: Walter, Nils Philipp, et al.
Pubblicazione: (2025)
LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference
di: Yuan, Jianhao, et al.
Pubblicazione: (2025)
di: Yuan, Jianhao, et al.
Pubblicazione: (2025)
The Cognitive Revolution in Interpretability: From Explaining Behavior to Interpreting Representations and Algorithms
di: Davies, Adam, et al.
Pubblicazione: (2024)
di: Davies, Adam, et al.
Pubblicazione: (2024)
TDGNet: Hallucination Detection in Diffusion Language Models via Temporal Dynamic Graphs
di: Hemmat, Arshia, et al.
Pubblicazione: (2026)
di: Hemmat, Arshia, et al.
Pubblicazione: (2026)
Hidden in Plain Sight: Undetectable Adversarial Bias Attacks on Vulnerable Patient Populations
di: Kulkarni, Pranav, et al.
Pubblicazione: (2024)
di: Kulkarni, Pranav, et al.
Pubblicazione: (2024)
A Survey on Transferability of Adversarial Examples across Deep Neural Networks
di: Gu, Jindong, et al.
Pubblicazione: (2023)
di: Gu, Jindong, et al.
Pubblicazione: (2023)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
di: Liu, Runtao, et al.
Pubblicazione: (2024)
di: Liu, Runtao, et al.
Pubblicazione: (2024)
As Firm As Their Foundations: Can open-sourced foundation models be used to create adversarial examples for downstream tasks?
di: Hu, Anjun, et al.
Pubblicazione: (2024)
di: Hu, Anjun, et al.
Pubblicazione: (2024)
An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models
di: Luo, Haochen, et al.
Pubblicazione: (2024)
di: Luo, Haochen, et al.
Pubblicazione: (2024)
Shape and Texture Recognition in Large Vision-Language Models
di: Eppel, Sagi, et al.
Pubblicazione: (2025)
di: Eppel, Sagi, et al.
Pubblicazione: (2025)
GlyphPattern: An Abstract Pattern Recognition Benchmark for Vision-Language Models
di: Wu, Zixuan, et al.
Pubblicazione: (2024)
di: Wu, Zixuan, et al.
Pubblicazione: (2024)
AR as an Evaluation Playground: Bridging Metrics and Visual Perception of Computer Vision Models
di: Ganj, Ashkan, et al.
Pubblicazione: (2025)
di: Ganj, Ashkan, et al.
Pubblicazione: (2025)
Focus On This, Not That! Steering LLMs with Adaptive Feature Specification
di: Lamb, Tom A., et al.
Pubblicazione: (2024)
di: Lamb, Tom A., et al.
Pubblicazione: (2024)
Seeing the Abstract: Translating the Abstract Language for Vision Language Models
di: Talon, Davide, et al.
Pubblicazione: (2025)
di: Talon, Davide, et al.
Pubblicazione: (2025)
kNN-CLIP: Retrieval Enables Training-Free Segmentation on Continually Expanding Large Vocabularies
di: Gui, Zhongrui, et al.
Pubblicazione: (2024)
di: Gui, Zhongrui, et al.
Pubblicazione: (2024)
Towards Interpreting Visual Information Processing in Vision-Language Models
di: Neo, Clement, et al.
Pubblicazione: (2024)
di: Neo, Clement, et al.
Pubblicazione: (2024)
Vision-Language Models Do Not Understand Negation
di: Alhamoud, Kumail, et al.
Pubblicazione: (2025)
di: Alhamoud, Kumail, et al.
Pubblicazione: (2025)
SpatialBot: Precise Spatial Understanding with Vision Language Models
di: Cai, Wenxiao, et al.
Pubblicazione: (2024)
di: Cai, Wenxiao, et al.
Pubblicazione: (2024)
The Shape of Sight: A Homological Framework for Unifying Visual Perception
di: Li, Xin
Pubblicazione: (2018)
di: Li, Xin
Pubblicazione: (2018)
Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions
di: Qu, Yiting, et al.
Pubblicazione: (2025)
di: Qu, Yiting, et al.
Pubblicazione: (2025)
Bridging Hidden States in Vision-Language Models
di: Fein-Ashley, Benjamin, et al.
Pubblicazione: (2025)
di: Fein-Ashley, Benjamin, et al.
Pubblicazione: (2025)
BlindSight: Harnessing Sparsity for Efficient Vision-Language Models
di: Srikrishnan, Tharun Adithya, et al.
Pubblicazione: (2025)
di: Srikrishnan, Tharun Adithya, et al.
Pubblicazione: (2025)
LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
di: Man, Yunze, et al.
Pubblicazione: (2025)
di: Man, Yunze, et al.
Pubblicazione: (2025)
Evaluating Vision-Language Models for Emotion Recognition
di: Bhattacharyya, Sree, et al.
Pubblicazione: (2025)
di: Bhattacharyya, Sree, et al.
Pubblicazione: (2025)
SynArtifact: Classifying and Alleviating Artifacts in Synthetic Images via Vision-Language Model
di: Cao, Bin, et al.
Pubblicazione: (2024)
di: Cao, Bin, et al.
Pubblicazione: (2024)
Hiding Faces in Plain Sight: Defending DeepFakes by Disrupting Face Detection
di: Zhu, Delong, et al.
Pubblicazione: (2024)
di: Zhu, Delong, et al.
Pubblicazione: (2024)
Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs
di: Liu, Xiaoyuan, et al.
Pubblicazione: (2024)
di: Liu, Xiaoyuan, et al.
Pubblicazione: (2024)
Manipulation Facing Threats: Evaluating Physical Vulnerabilities in End-to-End Vision Language Action Models
di: Cheng, Hao, et al.
Pubblicazione: (2024)
di: Cheng, Hao, et al.
Pubblicazione: (2024)
Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images
di: Gao, Kuofeng, et al.
Pubblicazione: (2024)
di: Gao, Kuofeng, et al.
Pubblicazione: (2024)
Learning to Hear by Seeing: It's Time for Vision Language Models to Understand Artistic Emotion from Sight and Sound
di: Zhang, Dengming, et al.
Pubblicazione: (2025)
di: Zhang, Dengming, et al.
Pubblicazione: (2025)
Finetune Like You Pretrain: Boosting Zero-shot Adversarial Robustness in Vision-language Models
di: Xing, Songlong, et al.
Pubblicazione: (2026)
di: Xing, Songlong, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Not Just Pretty Pictures: Toward Interventional Data Augmentation Using Text-to-Image Generators
di: Yuan, Jianhao, et al.
Pubblicazione: (2022) -
How Visual Representations Map to Language Feature Space in Multimodal LLMs
di: Venhoff, Constantin, et al.
Pubblicazione: (2025) -
Learning Visual Prompts for Guiding the Attention of Vision Transformers
di: Rezaei, Razieh, et al.
Pubblicazione: (2024) -
Articulate3D: Zero-Shot Text-Driven 3D Object Posing
di: Deb, Oishi, et al.
Pubblicazione: (2025) -
Learnable Sparsity for Vision Generative Models
di: Zhang, Yang, et al.
Pubblicazione: (2024)