Correlates of Image Memorability in Vision Encoders: Activations, Attention Entropy, Patch Uniformity and Autoencoder Losses
Fuente:
arXiv
Saved in:
| Main Authors: | Takmaz, Ece, Gatt, Albert, Dotlacil, Jakub |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Model Merging to Maintain Language-Only Performance in Developmentally Plausible Multimodal Models
by: Takmaz, Ece, et al.
Published: (2025)
by: Takmaz, Ece, et al.
Published: (2025)
Common Objects Out of Context (COOCo): Investigating Multimodal Context and Semantic Scene Violations in Referential Communication
by: Merlo, Filippo, et al.
Published: (2025)
by: Merlo, Filippo, et al.
Published: (2025)
Decoding Emotions in Abstract Art: Cognitive Plausibility of CLIP in Recognizing Color-Emotion Associations
by: Widhoelzl, Hanna-Sophia, et al.
Published: (2024)
by: Widhoelzl, Hanna-Sophia, et al.
Published: (2024)
Describing Images $\textit{Fast and Slow}$: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes
by: Takmaz, Ece, et al.
Published: (2024)
by: Takmaz, Ece, et al.
Published: (2024)
Localizing Memorization in SSL Vision Encoders
by: Wang, Wenhao, et al.
Published: (2024)
by: Wang, Wenhao, et al.
Published: (2024)
Modeling Visual Memorability Assessment with Autoencoders Reveals Characteristics of Memorable Images
by: Bagheri, Elham, et al.
Published: (2024)
by: Bagheri, Elham, et al.
Published: (2024)
From Image Captioning to Visual Storytelling
by: Passadakis, Admitos, et al.
Published: (2025)
by: Passadakis, Admitos, et al.
Published: (2025)
Rényi Attention Entropy for Patch Pruning
by: Aizawa, Hiroaki, et al.
Published: (2026)
by: Aizawa, Hiroaki, et al.
Published: (2026)
Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?
by: Song, Yingjin, et al.
Published: (2025)
by: Song, Yingjin, et al.
Published: (2025)
Context-aware Visual Storytelling with Visual Prefix Tuning and Contrastive Learning
by: Song, Yingjin, et al.
Published: (2024)
by: Song, Yingjin, et al.
Published: (2024)
REN: Fast and Efficient Region Encodings from Patch-Based Image Encoders
by: Khosla, Savya, et al.
Published: (2025)
by: Khosla, Savya, et al.
Published: (2025)
Patch-enhanced Mask Encoder Prompt Image Generation
by: Xu, Shusong, et al.
Published: (2024)
by: Xu, Shusong, et al.
Published: (2024)
Patch-Based Stochastic Attention for Image Editing
by: Cherel, Nicolas, et al.
Published: (2022)
by: Cherel, Nicolas, et al.
Published: (2022)
Rethinking Patch Dependence for Masked Autoencoders
by: Fu, Letian, et al.
Published: (2024)
by: Fu, Letian, et al.
Published: (2024)
VideoPatchCore: An Effective Method to Memorize Normality for Video Anomaly Detection
by: Ahn, Sunghyun, et al.
Published: (2024)
by: Ahn, Sunghyun, et al.
Published: (2024)
Activation Quantization of Vision Encoders Needs Prefixing Registers
by: Kim, Seunghyeon, et al.
Published: (2025)
by: Kim, Seunghyeon, et al.
Published: (2025)
Harnessing The Power of Attention For Patch-Based Biomedical Image Classification
by: Habib, Gousia, et al.
Published: (2024)
by: Habib, Gousia, et al.
Published: (2024)
Unforgettable Lessons from Forgettable Images: Intra-Class Memorability Matters in Computer Vision
by: Jing, Jie, et al.
Published: (2024)
by: Jing, Jie, et al.
Published: (2024)
General Vision Encoder Features as Guidance in Medical Image Registration
by: Kögl, Fryderyk, et al.
Published: (2024)
by: Kögl, Fryderyk, et al.
Published: (2024)
UNIT: Unifying Image and Text Recognition in One Vision Encoder
by: Zhu, Yi, et al.
Published: (2024)
by: Zhu, Yi, et al.
Published: (2024)
Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders
by: Rasekh, Ali, et al.
Published: (2025)
by: Rasekh, Ali, et al.
Published: (2025)
PA-Attack: Guiding Gray-Box Attacks on LVLM Vision Encoders with Prototypes and Attention
by: Mei, Hefei, et al.
Published: (2026)
by: Mei, Hefei, et al.
Published: (2026)
Vision Augmentation Prediction Autoencoder with Attention Design (VAPAAD)
by: Yin, Yiqiao
Published: (2024)
by: Yin, Yiqiao
Published: (2024)
Rediscovering BCE Loss for Uniform Classification
by: Li, Qiufu, et al.
Published: (2024)
by: Li, Qiufu, et al.
Published: (2024)
Causal Attribution via Activation Patching
by: Izadi, Amirmohammad, et al.
Published: (2026)
by: Izadi, Amirmohammad, et al.
Published: (2026)
Uniformity First: Uniformity-aware Test-time Adaptation of Vision-language Models against Image Corruption
by: Adachi, Kazuki, et al.
Published: (2025)
by: Adachi, Kazuki, et al.
Published: (2025)
SelfMedHPM: Self Pre-training With Hard Patches Mining Masked Autoencoders For Medical Image Segmentation
by: Lv, Yunhao, et al.
Published: (2025)
by: Lv, Yunhao, et al.
Published: (2025)
Patch Pruning Strategy Based on Robust Statistical Measures of Attention Weight Diversity in Vision Transformers
by: Igaue, Yuki, et al.
Published: (2025)
by: Igaue, Yuki, et al.
Published: (2025)
Exploring Local Memorization in Diffusion Models via Bright Ending Attention
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
Patch-Level Glioblastoma Subregion Classification with a Contrastive Learning-Based Encoder
by: Zhang, Juexin, et al.
Published: (2025)
by: Zhang, Juexin, et al.
Published: (2025)
CDPDNet: Integrating Text Guidance with Hybrid Vision Encoders for Medical Image Segmentation
by: Wu, Jiong, et al.
Published: (2025)
by: Wu, Jiong, et al.
Published: (2025)
PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders
by: Cavagnero, Niccolò, et al.
Published: (2026)
by: Cavagnero, Niccolò, et al.
Published: (2026)
VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
by: Wang, Zihu, et al.
Published: (2025)
by: Wang, Zihu, et al.
Published: (2025)
Beyond Independent Frames: Latent Attention Masked Autoencoders for Multi-View Echocardiography
by: Böhi, Simon, et al.
Published: (2026)
by: Böhi, Simon, et al.
Published: (2026)
Attention-Guided Masked Autoencoders For Learning Image Representations
by: Sick, Leon, et al.
Published: (2024)
by: Sick, Leon, et al.
Published: (2024)
Seeing Eye to AI: Comparing Human Gaze and Model Attention in Video Memorability
by: Kumar, Prajneya, et al.
Published: (2023)
by: Kumar, Prajneya, et al.
Published: (2023)
Patch-wise Auto-Encoder for Visual Anomaly Detection
by: Cui, Yajie, et al.
Published: (2023)
by: Cui, Yajie, et al.
Published: (2023)
Locating Demographic Bias at the Attention-Head Level in CLIP's Vision Encoder
by: Yasser, Alaa, et al.
Published: (2026)
by: Yasser, Alaa, et al.
Published: (2026)
You Don't Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models
by: Zhao, Kairan, et al.
Published: (2026)
by: Zhao, Kairan, et al.
Published: (2026)
Improving Image Coding for Machines through Optimizing Encoder via Auxiliary Loss
by: Iino, Kei, et al.
Published: (2024)
by: Iino, Kei, et al.
Published: (2024)
Similar Items
-
Model Merging to Maintain Language-Only Performance in Developmentally Plausible Multimodal Models
by: Takmaz, Ece, et al.
Published: (2025) -
Common Objects Out of Context (COOCo): Investigating Multimodal Context and Semantic Scene Violations in Referential Communication
by: Merlo, Filippo, et al.
Published: (2025) -
Decoding Emotions in Abstract Art: Cognitive Plausibility of CLIP in Recognizing Color-Emotion Associations
by: Widhoelzl, Hanna-Sophia, et al.
Published: (2024) -
Describing Images $\textit{Fast and Slow}$: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes
by: Takmaz, Ece, et al.
Published: (2024) -
Localizing Memorization in SSL Vision Encoders
by: Wang, Wenhao, et al.
Published: (2024)