Hidden in Plain Sight: Evaluating Abstract Shape Recognition in Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hemmat, Arshia, Davies, Adam, Lamb, Tom A., Yuan, Jianhao, Torr, Philip, Khakzar, Ashkan, Pinto, Francesco |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Not Just Pretty Pictures: Toward Interventional Data Augmentation Using Text-to-Image Generators
von: Yuan, Jianhao, et al.
Veröffentlicht: (2022)
von: Yuan, Jianhao, et al.
Veröffentlicht: (2022)
How Visual Representations Map to Language Feature Space in Multimodal LLMs
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
Learning Visual Prompts for Guiding the Attention of Vision Transformers
von: Rezaei, Razieh, et al.
Veröffentlicht: (2024)
von: Rezaei, Razieh, et al.
Veröffentlicht: (2024)
Articulate3D: Zero-Shot Text-Driven 3D Object Posing
von: Deb, Oishi, et al.
Veröffentlicht: (2025)
von: Deb, Oishi, et al.
Veröffentlicht: (2025)
Learnable Sparsity for Vision Generative Models
von: Zhang, Yang, et al.
Veröffentlicht: (2024)
von: Zhang, Yang, et al.
Veröffentlicht: (2024)
Latent Guard: a Safety Framework for Text-to-image Generation
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
Hidden Meanings in Plain Sight: RebusBench for Evaluating Cognitive Visual Reasoning
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2026)
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2026)
Towards Understanding Multimodal Fine-Tuning: Spatial Features
von: Naghashyar, Lachin, et al.
Veröffentlicht: (2026)
von: Naghashyar, Lachin, et al.
Veröffentlicht: (2026)
Minimalist Concept Erasure in Generative Models
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
Hidden in Plain Sight -- Class Competition Focuses Attribution Maps
von: Walter, Nils Philipp, et al.
Veröffentlicht: (2025)
von: Walter, Nils Philipp, et al.
Veröffentlicht: (2025)
LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference
von: Yuan, Jianhao, et al.
Veröffentlicht: (2025)
von: Yuan, Jianhao, et al.
Veröffentlicht: (2025)
The Cognitive Revolution in Interpretability: From Explaining Behavior to Interpreting Representations and Algorithms
von: Davies, Adam, et al.
Veröffentlicht: (2024)
von: Davies, Adam, et al.
Veröffentlicht: (2024)
TDGNet: Hallucination Detection in Diffusion Language Models via Temporal Dynamic Graphs
von: Hemmat, Arshia, et al.
Veröffentlicht: (2026)
von: Hemmat, Arshia, et al.
Veröffentlicht: (2026)
Hidden in Plain Sight: Undetectable Adversarial Bias Attacks on Vulnerable Patient Populations
von: Kulkarni, Pranav, et al.
Veröffentlicht: (2024)
von: Kulkarni, Pranav, et al.
Veröffentlicht: (2024)
A Survey on Transferability of Adversarial Examples across Deep Neural Networks
von: Gu, Jindong, et al.
Veröffentlicht: (2023)
von: Gu, Jindong, et al.
Veröffentlicht: (2023)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
As Firm As Their Foundations: Can open-sourced foundation models be used to create adversarial examples for downstream tasks?
von: Hu, Anjun, et al.
Veröffentlicht: (2024)
von: Hu, Anjun, et al.
Veröffentlicht: (2024)
An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models
von: Luo, Haochen, et al.
Veröffentlicht: (2024)
von: Luo, Haochen, et al.
Veröffentlicht: (2024)
Shape and Texture Recognition in Large Vision-Language Models
von: Eppel, Sagi, et al.
Veröffentlicht: (2025)
von: Eppel, Sagi, et al.
Veröffentlicht: (2025)
GlyphPattern: An Abstract Pattern Recognition Benchmark for Vision-Language Models
von: Wu, Zixuan, et al.
Veröffentlicht: (2024)
von: Wu, Zixuan, et al.
Veröffentlicht: (2024)
AR as an Evaluation Playground: Bridging Metrics and Visual Perception of Computer Vision Models
von: Ganj, Ashkan, et al.
Veröffentlicht: (2025)
von: Ganj, Ashkan, et al.
Veröffentlicht: (2025)
Focus On This, Not That! Steering LLMs with Adaptive Feature Specification
von: Lamb, Tom A., et al.
Veröffentlicht: (2024)
von: Lamb, Tom A., et al.
Veröffentlicht: (2024)
Seeing the Abstract: Translating the Abstract Language for Vision Language Models
von: Talon, Davide, et al.
Veröffentlicht: (2025)
von: Talon, Davide, et al.
Veröffentlicht: (2025)
kNN-CLIP: Retrieval Enables Training-Free Segmentation on Continually Expanding Large Vocabularies
von: Gui, Zhongrui, et al.
Veröffentlicht: (2024)
von: Gui, Zhongrui, et al.
Veröffentlicht: (2024)
Towards Interpreting Visual Information Processing in Vision-Language Models
von: Neo, Clement, et al.
Veröffentlicht: (2024)
von: Neo, Clement, et al.
Veröffentlicht: (2024)
Vision-Language Models Do Not Understand Negation
von: Alhamoud, Kumail, et al.
Veröffentlicht: (2025)
von: Alhamoud, Kumail, et al.
Veröffentlicht: (2025)
SpatialBot: Precise Spatial Understanding with Vision Language Models
von: Cai, Wenxiao, et al.
Veröffentlicht: (2024)
von: Cai, Wenxiao, et al.
Veröffentlicht: (2024)
The Shape of Sight: A Homological Framework for Unifying Visual Perception
von: Li, Xin
Veröffentlicht: (2018)
von: Li, Xin
Veröffentlicht: (2018)
Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions
von: Qu, Yiting, et al.
Veröffentlicht: (2025)
von: Qu, Yiting, et al.
Veröffentlicht: (2025)
Bridging Hidden States in Vision-Language Models
von: Fein-Ashley, Benjamin, et al.
Veröffentlicht: (2025)
von: Fein-Ashley, Benjamin, et al.
Veröffentlicht: (2025)
BlindSight: Harnessing Sparsity for Efficient Vision-Language Models
von: Srikrishnan, Tharun Adithya, et al.
Veröffentlicht: (2025)
von: Srikrishnan, Tharun Adithya, et al.
Veröffentlicht: (2025)
LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
von: Man, Yunze, et al.
Veröffentlicht: (2025)
von: Man, Yunze, et al.
Veröffentlicht: (2025)
Evaluating Vision-Language Models for Emotion Recognition
von: Bhattacharyya, Sree, et al.
Veröffentlicht: (2025)
von: Bhattacharyya, Sree, et al.
Veröffentlicht: (2025)
SynArtifact: Classifying and Alleviating Artifacts in Synthetic Images via Vision-Language Model
von: Cao, Bin, et al.
Veröffentlicht: (2024)
von: Cao, Bin, et al.
Veröffentlicht: (2024)
Hiding Faces in Plain Sight: Defending DeepFakes by Disrupting Face Detection
von: Zhu, Delong, et al.
Veröffentlicht: (2024)
von: Zhu, Delong, et al.
Veröffentlicht: (2024)
Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs
von: Liu, Xiaoyuan, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoyuan, et al.
Veröffentlicht: (2024)
Manipulation Facing Threats: Evaluating Physical Vulnerabilities in End-to-End Vision Language Action Models
von: Cheng, Hao, et al.
Veröffentlicht: (2024)
von: Cheng, Hao, et al.
Veröffentlicht: (2024)
Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
Learning to Hear by Seeing: It's Time for Vision Language Models to Understand Artistic Emotion from Sight and Sound
von: Zhang, Dengming, et al.
Veröffentlicht: (2025)
von: Zhang, Dengming, et al.
Veröffentlicht: (2025)
Finetune Like You Pretrain: Boosting Zero-shot Adversarial Robustness in Vision-language Models
von: Xing, Songlong, et al.
Veröffentlicht: (2026)
von: Xing, Songlong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Not Just Pretty Pictures: Toward Interventional Data Augmentation Using Text-to-Image Generators
von: Yuan, Jianhao, et al.
Veröffentlicht: (2022) -
How Visual Representations Map to Language Feature Space in Multimodal LLMs
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025) -
Learning Visual Prompts for Guiding the Attention of Vision Transformers
von: Rezaei, Razieh, et al.
Veröffentlicht: (2024) -
Articulate3D: Zero-Shot Text-Driven 3D Object Posing
von: Deb, Oishi, et al.
Veröffentlicht: (2025) -
Learnable Sparsity for Vision Generative Models
von: Zhang, Yang, et al.
Veröffentlicht: (2024)