Beyond Semantics: Disentangling Information Scope in Sparse Autoencoders for CLIP
Fuente:
arXiv
Saved in:
| Main Authors: | Ro, Yusung, Choi, Jaehyun, Kim, Junmo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cross-Domain Semantic Segmentation on Inconsistent Taxonomy using VLMs
by: Lim, Jeongkee, et al.
Published: (2024)
by: Lim, Jeongkee, et al.
Published: (2024)
Learning Neural Deformation Representation for 4D Dynamic Shape Generation
by: Han, Gyojin, et al.
Published: (2026)
by: Han, Gyojin, et al.
Published: (2026)
Modeling Stereo-Confidence Out of the End-to-End Stereo-Matching Network via Disparity Plane Sweep
by: Lee, Jae Young, et al.
Published: (2024)
by: Lee, Jae Young, et al.
Published: (2024)
Stereo-Matching Knowledge Distilled Monocular Depth Estimation Filtered by Multiple Disparity Consistency
by: Ka, Woonghyun, et al.
Published: (2024)
by: Ka, Woonghyun, et al.
Published: (2024)
PRISM: Video Dataset Condensation with Progressive Refinement and Insertion for Sparse Motion
by: Choi, Jaehyun, et al.
Published: (2025)
by: Choi, Jaehyun, et al.
Published: (2025)
AH-OCDA: Amplitude-based Curriculum Learning and Hopfield Segmentation Model for Open Compound Domain Adaptation
by: Choi, Jaehyun, et al.
Published: (2024)
by: Choi, Jaehyun, et al.
Published: (2024)
Self-supervised One-Stage Learning for RF-based Multi-Person Pose Estimation
by: Shin, Seunghwan, et al.
Published: (2025)
by: Shin, Seunghwan, et al.
Published: (2025)
Unlocking the Capabilities of Masked Generative Models for Image Synthesis via Self-Guidance
by: Hur, Jiwan, et al.
Published: (2024)
by: Hur, Jiwan, et al.
Published: (2024)
Frequency-Aware Token Reduction for Efficient Vision Transformer
by: Lee, Dong-Jae, et al.
Published: (2025)
by: Lee, Dong-Jae, et al.
Published: (2025)
Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models
by: Morelli, Fabian, et al.
Published: (2026)
by: Morelli, Fabian, et al.
Published: (2026)
Interpreting CLIP with Hierarchical Sparse Autoencoders
by: Zaigrajew, Vladimir, et al.
Published: (2025)
by: Zaigrajew, Vladimir, et al.
Published: (2025)
ConceptPrism: Concept Disentanglement in Personalized Diffusion Models via Residual Token Optimization
by: Kim, Minseo, et al.
Published: (2026)
by: Kim, Minseo, et al.
Published: (2026)
Do Vision Models Encode Object-Level Semantic Relatedness? A Cognitive Psychology-Inspired Benchmark
by: Lee, Hansang, et al.
Published: (2017)
by: Lee, Hansang, et al.
Published: (2017)
Dream to Generalize: Zero-Shot Model-Based Reinforcement Learning for Unseen Visual Distractions
by: Ha, Jeongsoo, et al.
Published: (2025)
by: Ha, Jeongsoo, et al.
Published: (2025)
DAM: Domain-Aware Module for Multi-Domain Dataset Condensation
by: Choi, Jaehyun, et al.
Published: (2025)
by: Choi, Jaehyun, et al.
Published: (2025)
Self-supervised Transformation Learning for Equivariant Representations
by: Yu, Jaemyung, et al.
Published: (2025)
by: Yu, Jaemyung, et al.
Published: (2025)
Self-Predictive Dynamics for Generalization of Vision-based Reinforcement Learning
by: Kim, Kyungsoo, et al.
Published: (2025)
by: Kim, Kyungsoo, et al.
Published: (2025)
SPACE-CLIP: Spatial Perception via Adaptive CLIP Embeddings for Monocular Depth Estimation
by: Cho, Taewan, et al.
Published: (2026)
by: Cho, Taewan, et al.
Published: (2026)
Efficiently Disentangling CLIP for Multi-Object Perception
by: Rawlekar, Samyak, et al.
Published: (2025)
by: Rawlekar, Samyak, et al.
Published: (2025)
Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling
by: Park, Sungjune, et al.
Published: (2025)
by: Park, Sungjune, et al.
Published: (2025)
Concept Steerers: Leveraging K-Sparse Autoencoders for Test-Time Controllable Generations
by: Kim, Dahye, et al.
Published: (2025)
by: Kim, Dahye, et al.
Published: (2025)
SSG: Scaled Spatial Guidance for Multi-Scale Visual Autoregressive Generation
by: Shin, Youngwoo, et al.
Published: (2026)
by: Shin, Youngwoo, et al.
Published: (2026)
Instruct-4DGS: Efficient Dynamic Scene Editing via 4D Gaussian-based Static-Dynamic Separation
by: Kwon, Joohyun, et al.
Published: (2025)
by: Kwon, Joohyun, et al.
Published: (2025)
ConceptScope: Characterizing Dataset Bias via Disentangled Visual Concepts
by: Choi, Jinho, et al.
Published: (2025)
by: Choi, Jinho, et al.
Published: (2025)
ReCLIP++: Learn to Rectify the Bias of CLIP for Unsupervised Semantic Segmentation
by: Wang, Jingyun, et al.
Published: (2024)
by: Wang, Jingyun, et al.
Published: (2024)
B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding
by: Choi, Changho, et al.
Published: (2025)
by: Choi, Changho, et al.
Published: (2025)
Inspecting Explainability of Transformer Models with Additional Statistical Information
by: Nguyen, Hoang C., et al.
Published: (2023)
by: Nguyen, Hoang C., et al.
Published: (2023)
Cross-Axis Feature Fusion with Joint-Wise Motion Difference Prediction for Text-Based 3D Human Motion Editing
by: Han, Gyojin, et al.
Published: (2026)
by: Han, Gyojin, et al.
Published: (2026)
Boundary-Recovering Network for Temporal Action Detection
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
Deciphering the Role of Representation Disentanglement: Investigating Compositional Generalization in CLIP Models
by: Abbasi, Reza, et al.
Published: (2024)
by: Abbasi, Reza, et al.
Published: (2024)
Query-Calibrated Segmental Admission for Descriptor-Agnostic LiDAR Loop Closure in Repetitive Environments
by: Kim, Jaehyun, et al.
Published: (2025)
by: Kim, Jaehyun, et al.
Published: (2025)
The Rescue Effect: Spatio-Semantic Early Exit Bypasses Quantization Collapse in CLIP
by: Nam, Kahyeon, et al.
Published: (2026)
by: Nam, Kahyeon, et al.
Published: (2026)
Causal Interpretation of Sparse Autoencoder Features in Vision
by: Han, Sangyu, et al.
Published: (2025)
by: Han, Sangyu, et al.
Published: (2025)
CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic Segmentation
by: Zhang, Dengke, et al.
Published: (2024)
by: Zhang, Dengke, et al.
Published: (2024)
CLIP-SENet: CLIP-based Semantic Enhancement Network for Vehicle Re-identification
by: Lu, Liping, et al.
Published: (2025)
by: Lu, Liping, et al.
Published: (2025)
Disentangled Diffusion Autoencoder for Harmonization of Multi-site Neuroimaging Data
by: Ijishakin, Ayodeji, et al.
Published: (2024)
by: Ijishakin, Ayodeji, et al.
Published: (2024)
CLIP-MUSED: CLIP-Guided Multi-Subject Visual Neural Information Semantic Decoding
by: Zhou, Qiongyi, et al.
Published: (2024)
by: Zhou, Qiongyi, et al.
Published: (2024)
Leveraging Image Augmentation for Object Manipulation: Towards Interpretable Controllability in Object-Centric Learning
by: Kim, Jinwoo, et al.
Published: (2023)
by: Kim, Jinwoo, et al.
Published: (2023)
Disentanglement with Factor Quantized Variational Autoencoders
by: Baykal, Gulcin, et al.
Published: (2024)
by: Baykal, Gulcin, et al.
Published: (2024)
SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels
by: Singh, Darshan, et al.
Published: (2024)
by: Singh, Darshan, et al.
Published: (2024)
Similar Items
-
Cross-Domain Semantic Segmentation on Inconsistent Taxonomy using VLMs
by: Lim, Jeongkee, et al.
Published: (2024) -
Learning Neural Deformation Representation for 4D Dynamic Shape Generation
by: Han, Gyojin, et al.
Published: (2026) -
Modeling Stereo-Confidence Out of the End-to-End Stereo-Matching Network via Disparity Plane Sweep
by: Lee, Jae Young, et al.
Published: (2024) -
Stereo-Matching Knowledge Distilled Monocular Depth Estimation Filtered by Multiple Disparity Consistency
by: Ka, Woonghyun, et al.
Published: (2024) -
PRISM: Video Dataset Condensation with Progressive Refinement and Insertion for Sparse Motion
by: Choi, Jaehyun, et al.
Published: (2025)