Visual Anagrams Reveal Hidden Differences in Holistic Shape Processing Across Vision Models
Fuente:
arXiv
Saved in:
| Main Authors: | Doshi, Fenil R., Fel, Thomas, Konkle, Talia, Alvarez, George |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bi-Orthogonal Factor Decomposition for Vision Transformers
by: Doshi, Fenil R., et al.
Published: (2026)
by: Doshi, Fenil R., et al.
Published: (2026)
Feature Accentuation: Revealing 'What' Features Respond to in Natural Images
by: Hamblin, Chris, et al.
Published: (2024)
by: Hamblin, Chris, et al.
Published: (2024)
Sparks of Explainability: Recent Advancements in Explaining Large Vision Models
by: Fel, Thomas
Published: (2025)
by: Fel, Thomas
Published: (2025)
Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry
by: Fel, Thomas, et al.
Published: (2025)
by: Fel, Thomas, et al.
Published: (2025)
Understanding Visual Feature Reliance through the Lens of Complexity
by: Fel, Thomas, et al.
Published: (2024)
by: Fel, Thomas, et al.
Published: (2024)
Understanding Inhibition Through Maximally Tense Images
by: Hamblin, Chris, et al.
Published: (2024)
by: Hamblin, Chris, et al.
Published: (2024)
VHELM: A Holistic Evaluation of Vision Language Models
by: Lee, Tony, et al.
Published: (2024)
by: Lee, Tony, et al.
Published: (2024)
Visual IRL for Human-Like Robotic Manipulation
by: Asali, Ehsan, et al.
Published: (2024)
by: Asali, Ehsan, et al.
Published: (2024)
HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model
by: Lee, Youngwan, et al.
Published: (2025)
by: Lee, Youngwan, et al.
Published: (2025)
Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
by: Yasunaga, Michihiro, et al.
Published: (2025)
by: Yasunaga, Michihiro, et al.
Published: (2025)
GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models
by: Shen, Guanxi
Published: (2025)
by: Shen, Guanxi
Published: (2025)
Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models
by: Fel, Thomas, et al.
Published: (2025)
by: Fel, Thomas, et al.
Published: (2025)
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
by: Bohacek, Matyas, et al.
Published: (2025)
by: Bohacek, Matyas, et al.
Published: (2025)
Block-Recurrent Dynamics in Vision Transformers
by: Jacobs, Mozes, et al.
Published: (2025)
by: Jacobs, Mozes, et al.
Published: (2025)
Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation
by: Zhang, Ziheng, et al.
Published: (2025)
by: Zhang, Ziheng, et al.
Published: (2025)
FOVI: A biologically-inspired foveated interface for deep vision models
by: Blauch, Nicholas M., et al.
Published: (2026)
by: Blauch, Nicholas M., et al.
Published: (2026)
A Geometric Unification of Concept Learning with Concept Cones
by: Rocchi--Henry, Alexandre, et al.
Published: (2025)
by: Rocchi--Henry, Alexandre, et al.
Published: (2025)
Unlocking Feature Visualization for Deeper Networks with MAgnitude Constrained Optimization
by: Fel, Thomas, et al.
Published: (2023)
by: Fel, Thomas, et al.
Published: (2023)
Diffusion-based Visual Anagram as Multi-task Learning
by: Xu, Zhiyuan, et al.
Published: (2024)
by: Xu, Zhiyuan, et al.
Published: (2024)
Back to the Baseline: Examining Baseline Effects on Explainability Metrics
by: Picard, Agustin Martin, et al.
Published: (2025)
by: Picard, Agustin Martin, et al.
Published: (2025)
Same Answer, Different Representations: Hidden instability in VLMs
by: Wani, Farooq Ahmad, et al.
Published: (2026)
by: Wani, Farooq Ahmad, et al.
Published: (2026)
Interpreting Physics in Video World Models
by: Joseph, Sonia, et al.
Published: (2026)
by: Joseph, Sonia, et al.
Published: (2026)
Visual Anagrams: Generating Multi-View Optical Illusions with Diffusion Models
by: Geng, Daniel, et al.
Published: (2023)
by: Geng, Daniel, et al.
Published: (2023)
Hidden Clones: Exposing and Fixing Family Bias in Vision-Language Model Ensembles
by: Bugaud, Zacharie
Published: (2026)
by: Bugaud, Zacharie
Published: (2026)
SHAPE : Self-Improved Visual Preference Alignment by Iteratively Generating Holistic Winner
by: Chen, Kejia, et al.
Published: (2025)
by: Chen, Kejia, et al.
Published: (2025)
Revealing Multi-View Hallucination in Large Vision-Language Models
by: Park, Wooje, et al.
Published: (2026)
by: Park, Wooje, et al.
Published: (2026)
Freeze and Reveal: Exposing Modality Bias in Vision-Language Models
by: Kavuri, Vivek Hruday, et al.
Published: (2025)
by: Kavuri, Vivek Hruday, et al.
Published: (2025)
From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning
by: Jiao, Pengkun, et al.
Published: (2024)
by: Jiao, Pengkun, et al.
Published: (2024)
The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering
by: Li, Zhuowei, et al.
Published: (2025)
by: Li, Zhuowei, et al.
Published: (2025)
Grid2Matrix: Revealing Digital Agnosia in Vision-Language Models
by: Zhang, Yunkai, et al.
Published: (2026)
by: Zhang, Yunkai, et al.
Published: (2026)
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
by: Zhang, Jihai, et al.
Published: (2025)
by: Zhang, Jihai, et al.
Published: (2025)
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
by: Tang, Zhenwei, et al.
Published: (2025)
by: Tang, Zhenwei, et al.
Published: (2025)
Attention Hijacking: Response Manipulation Across Queries in Vision-Language Models
by: Wang, Zhiqiang, et al.
Published: (2026)
by: Wang, Zhiqiang, et al.
Published: (2026)
Brain2Text Decoding Model Reveals the Neural Mechanisms of Visual Semantic Processing
by: Feng, Feihan, et al.
Published: (2025)
by: Feng, Feihan, et al.
Published: (2025)
Text Speaks Louder than Vision: ASCII Art Reveals Textual Biases in Vision-Language Models
by: Wang, Zhaochen, et al.
Published: (2025)
by: Wang, Zhaochen, et al.
Published: (2025)
Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities
by: Wang, Liuyi, et al.
Published: (2025)
by: Wang, Liuyi, et al.
Published: (2025)
Holistic Surgical Phase Recognition with Hierarchical Input Dependent State Space Models
by: Wu, Haoyang, et al.
Published: (2025)
by: Wu, Haoyang, et al.
Published: (2025)
Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models
by: Chae, Hyunsik, et al.
Published: (2025)
by: Chae, Hyunsik, et al.
Published: (2025)
Discovering Hidden Visual Concepts Beyond Linguistic Input in Infant Learning
by: Ke, Xueyi, et al.
Published: (2025)
by: Ke, Xueyi, et al.
Published: (2025)
Understanding Visual Concepts Across Models
by: Trabucco, Brandon, et al.
Published: (2024)
by: Trabucco, Brandon, et al.
Published: (2024)
Similar Items
-
Bi-Orthogonal Factor Decomposition for Vision Transformers
by: Doshi, Fenil R., et al.
Published: (2026) -
Feature Accentuation: Revealing 'What' Features Respond to in Natural Images
by: Hamblin, Chris, et al.
Published: (2024) -
Sparks of Explainability: Recent Advancements in Explaining Large Vision Models
by: Fel, Thomas
Published: (2025) -
Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski Geometry
by: Fel, Thomas, et al.
Published: (2025) -
Understanding Visual Feature Reliance through the Lens of Complexity
by: Fel, Thomas, et al.
Published: (2024)