Causal Interpretation of Sparse Autoencoder Features in Vision
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Han, Sangyu, Kim, Yearim, Kwak, Nojun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Respect the model: Fine-grained and Robust Explanation with Sharing Ratio Decomposition
von: Han, Sangyu, et al.
Veröffentlicht: (2024)
von: Han, Sangyu, et al.
Veröffentlicht: (2024)
From Local to Global to Mechanistic: An iERF-Centered Unified Framework for Interpreting Vision Models
von: Kim, Yearim, et al.
Veröffentlicht: (2026)
von: Kim, Yearim, et al.
Veröffentlicht: (2026)
Decompose the model: Mechanistic interpretability in image models with Generalized Integrated Gradients (GIG)
von: Kim, Yearim, et al.
Veröffentlicht: (2024)
von: Kim, Yearim, et al.
Veröffentlicht: (2024)
PedaCo-Gen: Scaffolding Pedagogical Agency in Human-AI Collaborative Video Authoring
von: Baek, Injun, et al.
Veröffentlicht: (2026)
von: Baek, Injun, et al.
Veröffentlicht: (2026)
Bi-ICE: An Inner Interpretable Framework for Image Classification via Bi-directional Interactions between Concept and Input Embeddings
von: Hong, Jinyung, et al.
Veröffentlicht: (2024)
von: Hong, Jinyung, et al.
Veröffentlicht: (2024)
MSG Score: Automated Video Verification for Reliable Multi-Scene Generation
von: Yoon, Daewon, et al.
Veröffentlicht: (2024)
von: Yoon, Daewon, et al.
Veröffentlicht: (2024)
VDPP: Video Depth Post-Processing for Speed and Scalability
von: Yoon, Daewon, et al.
Veröffentlicht: (2026)
von: Yoon, Daewon, et al.
Veröffentlicht: (2026)
Unlocking the Potential of Unlabeled Data in Semi-Supervised Domain Generalization
von: Lee, Dongkwan, et al.
Veröffentlicht: (2025)
von: Lee, Dongkwan, et al.
Veröffentlicht: (2025)
An X-Ray Is Worth 15 Features: Sparse Autoencoders for Interpretable Radiology Report Generation
von: Abdulaal, Ahmed, et al.
Veröffentlicht: (2024)
von: Abdulaal, Ahmed, et al.
Veröffentlicht: (2024)
ReSpinQuant: Efficient Layer-Wise LLM Quantization via Subspace Residual Rotation Approximation
von: Kim, Suyoung, et al.
Veröffentlicht: (2026)
von: Kim, Suyoung, et al.
Veröffentlicht: (2026)
The Role of Teacher Calibration in Knowledge Distillation
von: Kim, Suyoung, et al.
Veröffentlicht: (2025)
von: Kim, Suyoung, et al.
Veröffentlicht: (2025)
Unleash the Potential of CLIP for Video Highlight Detection
von: Han, Donghoon, et al.
Veröffentlicht: (2024)
von: Han, Donghoon, et al.
Veröffentlicht: (2024)
Real-Time Intuitive AI Drawing System for Collaboration: Enhancing Human Creativity through Formal and Contextual Intent Integration
von: Song, Jookyung, et al.
Veröffentlicht: (2025)
von: Song, Jookyung, et al.
Veröffentlicht: (2025)
CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image Models
von: Lee, Junhoo, et al.
Veröffentlicht: (2026)
von: Lee, Junhoo, et al.
Veröffentlicht: (2026)
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
von: Pach, Mateusz, et al.
Veröffentlicht: (2025)
von: Pach, Mateusz, et al.
Veröffentlicht: (2025)
MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline
von: Han, Donghoon, et al.
Veröffentlicht: (2024)
von: Han, Donghoon, et al.
Veröffentlicht: (2024)
Interpreting CLIP with Hierarchical Sparse Autoencoders
von: Zaigrajew, Vladimir, et al.
Veröffentlicht: (2025)
von: Zaigrajew, Vladimir, et al.
Veröffentlicht: (2025)
SPARC: Concept-Aligned Sparse Autoencoders for Cross-Model and Cross-Modal Interpretability
von: Nasiri-Sarvi, Ali, et al.
Veröffentlicht: (2025)
von: Nasiri-Sarvi, Ali, et al.
Veröffentlicht: (2025)
LouvreSAE: Sparse Autoencoders for Interpretable and Controllable Style Transfer
von: Panda, Raina, et al.
Veröffentlicht: (2025)
von: Panda, Raina, et al.
Veröffentlicht: (2025)
SAUCE: Selective Concept Unlearning in Vision-Language Models with Sparse Autoencoders
von: Li, Qing, et al.
Veröffentlicht: (2025)
von: Li, Qing, et al.
Veröffentlicht: (2025)
S3D: Sketch-Driven 3D Model Generation
von: Song, Hail, et al.
Veröffentlicht: (2025)
von: Song, Hail, et al.
Veröffentlicht: (2025)
Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models
von: Yeung, Calvin, et al.
Veröffentlicht: (2026)
von: Yeung, Calvin, et al.
Veröffentlicht: (2026)
ConcatPlexer: Additional Dim1 Batching for Faster ViTs
von: Han, Donghoon, et al.
Veröffentlicht: (2023)
von: Han, Donghoon, et al.
Veröffentlicht: (2023)
TIDE : Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation
von: Huang, Victor Shea-Jay, et al.
Veröffentlicht: (2025)
von: Huang, Victor Shea-Jay, et al.
Veröffentlicht: (2025)
What's Making That Sound Right Now? Video-centric Audio-Visual Localization
von: Choi, Hahyeon, et al.
Veröffentlicht: (2025)
von: Choi, Hahyeon, et al.
Veröffentlicht: (2025)
Interpretable and Testable Vision Features via Sparse Autoencoders
von: Stevens, Samuel, et al.
Veröffentlicht: (2025)
von: Stevens, Samuel, et al.
Veröffentlicht: (2025)
LUCID-SAE: Learning Unified Vision-Language Sparse Codes for Interpretable Concept Discovery
von: Gu, Difei, et al.
Veröffentlicht: (2026)
von: Gu, Difei, et al.
Veröffentlicht: (2026)
SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse Autoencoders
von: Cassano, Enrico, et al.
Veröffentlicht: (2025)
von: Cassano, Enrico, et al.
Veröffentlicht: (2025)
Vision-Language Introspection: Mitigating Overconfident Hallucinations in MLLMs via Interpretable Bi-Causal Steering
von: Liu, Shuliang, et al.
Veröffentlicht: (2026)
von: Liu, Shuliang, et al.
Veröffentlicht: (2026)
Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering
von: Chatzoudis, Gerasimos, et al.
Veröffentlicht: (2025)
von: Chatzoudis, Gerasimos, et al.
Veröffentlicht: (2025)
Synergistic Integration of Coordinate Network and Tensorial Feature for Improving Neural Radiance Fields from Sparse Inputs
von: Kim, Mingyu, et al.
Veröffentlicht: (2024)
von: Kim, Mingyu, et al.
Veröffentlicht: (2024)
EA: An Event Autoencoder for High-Speed Vision Sensing
von: Islam, Riadul, et al.
Veröffentlicht: (2025)
von: Islam, Riadul, et al.
Veröffentlicht: (2025)
Vision Augmentation Prediction Autoencoder with Attention Design (VAPAAD)
von: Yin, Yiqiao
Veröffentlicht: (2024)
von: Yin, Yiqiao
Veröffentlicht: (2024)
SAVE: Sparse Autoencoder-Driven Visual Information Enhancement for Mitigating Object Hallucination
von: Park, Sangha, et al.
Veröffentlicht: (2025)
von: Park, Sangha, et al.
Veröffentlicht: (2025)
Conservative Generator, Progressive Discriminator: Coordination of Adversaries in Few-shot Incremental Image Synthesis
von: Kong, Chaerin, et al.
Veröffentlicht: (2022)
von: Kong, Chaerin, et al.
Veröffentlicht: (2022)
Steering Sparse Autoencoder Latents to Control Dynamic Head Pruning in Vision Transformers (Student Abstract)
von: Lee, Yousung, et al.
Veröffentlicht: (2026)
von: Lee, Yousung, et al.
Veröffentlicht: (2026)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
von: Hua, Zhenglin, et al.
Veröffentlicht: (2025)
von: Hua, Zhenglin, et al.
Veröffentlicht: (2025)
Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders
von: Dokme, Atahan, et al.
Veröffentlicht: (2026)
von: Dokme, Atahan, et al.
Veröffentlicht: (2026)
SparseFormer: Detecting Objects in HRW Shots via Sparse Vision Transformer
von: Li, Wenxi, et al.
Veröffentlicht: (2025)
von: Li, Wenxi, et al.
Veröffentlicht: (2025)
MiniMaxAD: A Lightweight Autoencoder for Feature-Rich Anomaly Detection
von: Wang, Fengjie, et al.
Veröffentlicht: (2024)
von: Wang, Fengjie, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Respect the model: Fine-grained and Robust Explanation with Sharing Ratio Decomposition
von: Han, Sangyu, et al.
Veröffentlicht: (2024) -
From Local to Global to Mechanistic: An iERF-Centered Unified Framework for Interpreting Vision Models
von: Kim, Yearim, et al.
Veröffentlicht: (2026) -
Decompose the model: Mechanistic interpretability in image models with Generalized Integrated Gradients (GIG)
von: Kim, Yearim, et al.
Veröffentlicht: (2024) -
PedaCo-Gen: Scaffolding Pedagogical Agency in Human-AI Collaborative Video Authoring
von: Baek, Injun, et al.
Veröffentlicht: (2026) -
Bi-ICE: An Inner Interpretable Framework for Image Classification via Bi-directional Interactions between Concept and Input Embeddings
von: Hong, Jinyung, et al.
Veröffentlicht: (2024)