Disentanglement and Compositionality of Letter Identity and Letter Position in Variational Auto-Encoder Vision Models
Fuente:
arXiv
Saved in:
| Main Authors: | Bianchi, Bruno, Agrawal, Aakash, Dehaene, Stanislas, Chemla, Emmanuel, Lakretz, Yair |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cracking the neural code for word recognition in convolutional neural networks
by: Agrawal, Aakash, et al.
Published: (2024)
by: Agrawal, Aakash, et al.
Published: (2024)
DiA-gnostic VLVAE: Disentangled Alignment-Constrained Vision Language Variational AutoEncoder for Robust Radiology Reporting with Missing Modalities
by: Shaik, Nagur Shareef, et al.
Published: (2025)
by: Shaik, Nagur Shareef, et al.
Published: (2025)
A Neural Model for Word Repetition
by: Dager, Daniel, et al.
Published: (2025)
by: Dager, Daniel, et al.
Published: (2025)
Auto-Comp: An Automated Pipeline for Scalable Compositional Probing of Contrastive Vision-Language Models
by: Sbrolli, Cristian, et al.
Published: (2026)
by: Sbrolli, Cristian, et al.
Published: (2026)
Patch-wise Auto-Encoder for Visual Anomaly Detection
by: Cui, Yajie, et al.
Published: (2023)
by: Cui, Yajie, et al.
Published: (2023)
Physics Informed Capsule Enhanced Variational AutoEncoder for Underwater Image Enhancement
by: Martinel, Niki, et al.
Published: (2025)
by: Martinel, Niki, et al.
Published: (2025)
SAEN-BGS: Energy-Efficient Spiking AutoEncoder Network for Background Subtraction
by: Zhang, Zhixuan, et al.
Published: (2025)
by: Zhang, Zhixuan, et al.
Published: (2025)
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
by: Diao, Haiwen, et al.
Published: (2025)
by: Diao, Haiwen, et al.
Published: (2025)
Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
by: Wang, Yizhou, et al.
Published: (2025)
by: Wang, Yizhou, et al.
Published: (2025)
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion
by: Chen, Jiuhai, et al.
Published: (2024)
by: Chen, Jiuhai, et al.
Published: (2024)
StyleAutoEncoder for manipulating image attributes using pre-trained StyleGAN
by: Bedychaj, Andrzej, et al.
Published: (2024)
by: Bedychaj, Andrzej, et al.
Published: (2024)
HaloAE: An HaloNet based Local Transformer Auto-Encoder for Anomaly Detection and Localization
by: Mathian, E., et al.
Published: (2022)
by: Mathian, E., et al.
Published: (2022)
SelfSwapper: Self-Supervised Face Swapping via Shape Agnostic Masked AutoEncoder
by: Lee, Jaeseong, et al.
Published: (2024)
by: Lee, Jaeseong, et al.
Published: (2024)
Generation of Indian Sign Language Letters, Numbers, and Words
by: Yadav, Ajeet Kumar, et al.
Published: (2025)
by: Yadav, Ajeet Kumar, et al.
Published: (2025)
Securing Vision-Language Models with a Robust Encoder Against Jailbreak and Adversarial Attacks
by: Hossain, Md Zarif, et al.
Published: (2024)
by: Hossain, Md Zarif, et al.
Published: (2024)
When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
by: Ortu, Francesco, et al.
Published: (2025)
by: Ortu, Francesco, et al.
Published: (2025)
Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
by: Yang, Yi, et al.
Published: (2025)
by: Yang, Yi, et al.
Published: (2025)
Visual Words Meet BM25: Sparse Auto-Encoder Visual Word Scoring for Image Retrieval
by: Han, Donghoon, et al.
Published: (2026)
by: Han, Donghoon, et al.
Published: (2026)
GrabDAE: An Innovative Framework for Unsupervised Domain Adaptation Utilizing Grab-Mask and Denoise Auto-Encoder
by: Chen, Junzhou, et al.
Published: (2024)
by: Chen, Junzhou, et al.
Published: (2024)
Blind to Position, Biased in Language: Probing Mid-Layer Representational Bias in Vision-Language Encoders for Zero-Shot Language-Grounded Spatial Understanding
by: An, Na Min, et al.
Published: (2025)
by: An, Na Min, et al.
Published: (2025)
Agglomerating Large Vision Encoders via Distillation for VFSS Segmentation
by: Zeng, Chengxi, et al.
Published: (2025)
by: Zeng, Chengxi, et al.
Published: (2025)
A Mixed Diet Makes DINO An Omnivorous Vision Encoder
by: Kabra, Rishabh, et al.
Published: (2026)
by: Kabra, Rishabh, et al.
Published: (2026)
AutoBench-V: Can Large Vision-Language Models Benchmark Themselves?
by: Bao, Han, et al.
Published: (2024)
by: Bao, Han, et al.
Published: (2024)
SAEdit: Token-level control for continuous image editing via Sparse AutoEncoder
by: Kamenetsky, Ronen, et al.
Published: (2025)
by: Kamenetsky, Ronen, et al.
Published: (2025)
AutoCLIP: Auto-tuning Zero-Shot Classifiers for Vision-Language Models
by: Metzen, Jan Hendrik, et al.
Published: (2023)
by: Metzen, Jan Hendrik, et al.
Published: (2023)
Vanishing Depth: A Depth Adapter with Positional Depth Encoding for Generalized Image Encoders
by: Koch, Paul, et al.
Published: (2025)
by: Koch, Paul, et al.
Published: (2025)
Self-Attention Based Multi-Scale Graph Auto-Encoder Network of 3D Meshes
by: Nazir, Saqib, et al.
Published: (2025)
by: Nazir, Saqib, et al.
Published: (2025)
Discovering Failure Modes in Vision-Language Models using RL
by: Jain, Kanishk, et al.
Published: (2026)
by: Jain, Kanishk, et al.
Published: (2026)
Variational Encoder--Multi-Decoder (VE-MD) for Privacy-by-functional-design (Group) Emotion Recognition
by: Augusma, Anderson, et al.
Published: (2026)
by: Augusma, Anderson, et al.
Published: (2026)
Quantization with Unified Adaptive Distillation to enable multi-LoRA based one-for-all Generative Vision Models on edge
by: Vajrala, Sowmya, et al.
Published: (2026)
by: Vajrala, Sowmya, et al.
Published: (2026)
VITAL: Vision-Encoder-centered Pre-training for LMMs in Visual Quality Assessment
by: Jia, Ziheng, et al.
Published: (2025)
by: Jia, Ziheng, et al.
Published: (2025)
Prompting Large Vision-Language Models for Compositional Reasoning
by: Ossowski, Timothy, et al.
Published: (2024)
by: Ossowski, Timothy, et al.
Published: (2024)
TAP-VL: Text Layout-Aware Pre-training for Enriched Vision-Language Models
by: Fhima, Jonathan, et al.
Published: (2024)
by: Fhima, Jonathan, et al.
Published: (2024)
VLM-AutoDrive: Post-Training Vision-Language Models for Safety-Critical Autonomous Driving Events
by: Bhat, Mohammad Qazim, et al.
Published: (2026)
by: Bhat, Mohammad Qazim, et al.
Published: (2026)
Weierstrass Positional Encoding for Vision Transformers
by: Xin, Zhihang, et al.
Published: (2026)
by: Xin, Zhihang, et al.
Published: (2026)
EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture
by: Feng, Wenfeng, et al.
Published: (2025)
by: Feng, Wenfeng, et al.
Published: (2025)
An Examination of Offline-Trained Encoders in Vision-Based Deep Reinforcement Learning for Autonomous Driving
by: Mohammed, Shawan, et al.
Published: (2024)
by: Mohammed, Shawan, et al.
Published: (2024)
FROST-Drive: Scalable and Efficient End-to-End Driving with a Frozen Vision Encoder
by: Dong, Zeyu, et al.
Published: (2026)
by: Dong, Zeyu, et al.
Published: (2026)
L-VAE: Variational Auto-Encoder with Learnable Beta for Disentangled Representation
by: Ozcan, Hazal Mogultay, et al.
Published: (2025)
by: Ozcan, Hazal Mogultay, et al.
Published: (2025)
UniFusion: Vision-Language Model as Unified Encoder in Image Generation
by: Li, Kevin, et al.
Published: (2025)
by: Li, Kevin, et al.
Published: (2025)
Similar Items
-
Cracking the neural code for word recognition in convolutional neural networks
by: Agrawal, Aakash, et al.
Published: (2024) -
DiA-gnostic VLVAE: Disentangled Alignment-Constrained Vision Language Variational AutoEncoder for Robust Radiology Reporting with Missing Modalities
by: Shaik, Nagur Shareef, et al.
Published: (2025) -
A Neural Model for Word Repetition
by: Dager, Daniel, et al.
Published: (2025) -
Auto-Comp: An Automated Pipeline for Scalable Compositional Probing of Contrastive Vision-Language Models
by: Sbrolli, Cristian, et al.
Published: (2026) -
Patch-wise Auto-Encoder for Visual Anomaly Detection
by: Cui, Yajie, et al.
Published: (2023)