Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)
Fuente:
arXiv
Guardado en:
| Autores principales: | Bhalla, Usha, Oesterling, Alex, Srinivas, Suraj, Calmon, Flavio P., Lakkaraju, Himabindu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent Interpretability
por: Bhalla, Usha, et al.
Publicado: (2023)
por: Bhalla, Usha, et al.
Publicado: (2023)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
por: Bhalla, Usha, et al.
Publicado: (2025)
por: Bhalla, Usha, et al.
Publicado: (2025)
All Roads Lead to Rome? Exploring Representational Similarities Between Latent Spaces of Generative Image Models
por: Badrinath, Charumathi, et al.
Publicado: (2024)
por: Badrinath, Charumathi, et al.
Publicado: (2024)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
por: Li, Aaron J., et al.
Publicado: (2025)
por: Li, Aaron J., et al.
Publicado: (2025)
Towards Unifying Interpretability and Control: Evaluation via Intervention
por: Bhalla, Usha, et al.
Publicado: (2024)
por: Bhalla, Usha, et al.
Publicado: (2024)
Operationalizing the Blueprint for an AI Bill of Rights: Recommendations for Practitioners, Researchers, and Policy Makers
por: Oesterling, Alex, et al.
Publicado: (2024)
por: Oesterling, Alex, et al.
Publicado: (2024)
Which Models have Perceptually-Aligned Gradients? An Explanation via Off-Manifold Robustness
por: Srinivas, Suraj, et al.
Publicado: (2023)
por: Srinivas, Suraj, et al.
Publicado: (2023)
Interpretability Needs a New Paradigm
por: Madsen, Andreas, et al.
Publicado: (2024)
por: Madsen, Andreas, et al.
Publicado: (2024)
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
por: Zhang, Shichang, et al.
Publicado: (2025)
por: Zhang, Shichang, et al.
Publicado: (2025)
Inference-Time Reward Hacking in Large Language Models
por: Khalaf, Hadi, et al.
Publicado: (2025)
por: Khalaf, Hadi, et al.
Publicado: (2025)
Soft Best-of-n Sampling for Model Alignment
por: Verdun, Claudio Mayrink, et al.
Publicado: (2025)
por: Verdun, Claudio Mayrink, et al.
Publicado: (2025)
Multi-Group Proportional Representation for Text-to-Image Models
por: Jung, Sangwon, et al.
Publicado: (2025)
por: Jung, Sangwon, et al.
Publicado: (2025)
Fair Machine Unlearning: Data Removal while Mitigating Disparities
por: Oesterling, Alex, et al.
Publicado: (2023)
por: Oesterling, Alex, et al.
Publicado: (2023)
Interpreting CLIP with Hierarchical Sparse Autoencoders
por: Zaigrajew, Vladimir, et al.
Publicado: (2025)
por: Zaigrajew, Vladimir, et al.
Publicado: (2025)
Characterizing Data Point Vulnerability via Average-Case Robustness
por: Han, Tessa, et al.
Publicado: (2023)
por: Han, Tessa, et al.
Publicado: (2023)
Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
por: Kulkarni, Akshay, et al.
Publicado: (2025)
por: Kulkarni, Akshay, et al.
Publicado: (2025)
Hierarchical Concept Embedding & Pursuit for Interpretable Image Classification
por: Nguyen, Nghia, et al.
Publicado: (2026)
por: Nguyen, Nghia, et al.
Publicado: (2026)
CASL: Concept-Aligned Sparse Latents for Interpreting Diffusion Models
por: He, Zhenghao, et al.
Publicado: (2026)
por: He, Zhenghao, et al.
Publicado: (2026)
Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment
por: Thasarathan, Harrish, et al.
Publicado: (2025)
por: Thasarathan, Harrish, et al.
Publicado: (2025)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
por: Kim, Eunji, et al.
Publicado: (2024)
por: Kim, Eunji, et al.
Publicado: (2024)
Disentangling Latent Embeddings with Sparse Linear Concept Subspaces (SLiCS)
por: Li, Zhi, et al.
Publicado: (2025)
por: Li, Zhi, et al.
Publicado: (2025)
CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling
por: Wang, Xinze, et al.
Publicado: (2025)
por: Wang, Xinze, et al.
Publicado: (2025)
Bi-ICE: An Inner Interpretable Framework for Image Classification via Bi-directional Interactions between Concept and Input Embeddings
por: Hong, Jinyung, et al.
Publicado: (2024)
por: Hong, Jinyung, et al.
Publicado: (2024)
Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive Learning
por: Qin, Chuan, et al.
Publicado: (2026)
por: Qin, Chuan, et al.
Publicado: (2026)
Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models
por: Morelli, Fabian, et al.
Publicado: (2026)
por: Morelli, Fabian, et al.
Publicado: (2026)
Optimizing CLIP Models for Image Retrieval with Maintained Joint-Embedding Alignment
por: Schall, Konstantin, et al.
Publicado: (2024)
por: Schall, Konstantin, et al.
Publicado: (2024)
Towards Interpretable Soft Prompts
por: Patel, Oam, et al.
Publicado: (2025)
por: Patel, Oam, et al.
Publicado: (2025)
From Segments to Concepts: Interpretable Image Classification via Concept-Guided Segmentation
por: Eisenberg, Ran, et al.
Publicado: (2025)
por: Eisenberg, Ran, et al.
Publicado: (2025)
CLIP-PAE: Projection-Augmentation Embedding to Extract Relevant Features for a Disentangled, Interpretable, and Controllable Text-Guided Face Manipulation
por: Zhou, Chenliang, et al.
Publicado: (2022)
por: Zhou, Chenliang, et al.
Publicado: (2022)
Memorization In Stable Diffusion Is Unexpectedly Driven by CLIP Embeddings
por: Kim, Bumjun, et al.
Publicado: (2026)
por: Kim, Bumjun, et al.
Publicado: (2026)
From Weights to Concepts: Data-Free Interpretability of CLIP via Singular Vector Decomposition
por: Gentile, Francesco, et al.
Publicado: (2026)
por: Gentile, Francesco, et al.
Publicado: (2026)
Isotropic3D: Image-to-3D Generation Based on a Single CLIP Embedding
por: Liu, Pengkun, et al.
Publicado: (2024)
por: Liu, Pengkun, et al.
Publicado: (2024)
Interpreting CLIP: Insights on the Robustness to ImageNet Distribution Shifts
por: Crabbé, Jonathan, et al.
Publicado: (2023)
por: Crabbé, Jonathan, et al.
Publicado: (2023)
Enhancing Interpretability of Sparse Latent Representations with Class Information
por: Abiz, Farshad Sangari, et al.
Publicado: (2025)
por: Abiz, Farshad Sangari, et al.
Publicado: (2025)
Sparse Autoencoders for Interpretable Medical Image Representation Learning
por: Wesp, Philipp, et al.
Publicado: (2026)
por: Wesp, Philipp, et al.
Publicado: (2026)
Conceptualizing Embeddings: Sparse Disentanglement for Vision-Language Models
por: Kubaty, Piotr, et al.
Publicado: (2026)
por: Kubaty, Piotr, et al.
Publicado: (2026)
DeCLIP: Decoding CLIP representations for deepfake localization
por: Smeu, Stefan, et al.
Publicado: (2024)
por: Smeu, Stefan, et al.
Publicado: (2024)
Concept-Centric Token Interpretation for Vector-Quantized Generative Models
por: Yang, Tianze, et al.
Publicado: (2025)
por: Yang, Tianze, et al.
Publicado: (2025)
Interpretable Generative Models through Post-hoc Concept Bottlenecks
por: Kulkarni, Akshay, et al.
Publicado: (2025)
por: Kulkarni, Akshay, et al.
Publicado: (2025)
ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features
por: Helbling, Alec, et al.
Publicado: (2025)
por: Helbling, Alec, et al.
Publicado: (2025)
Ejemplares similares
-
Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent Interpretability
por: Bhalla, Usha, et al.
Publicado: (2023) -
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
por: Bhalla, Usha, et al.
Publicado: (2025) -
All Roads Lead to Rome? Exploring Representational Similarities Between Latent Spaces of Generative Image Models
por: Badrinath, Charumathi, et al.
Publicado: (2024) -
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
por: Li, Aaron J., et al.
Publicado: (2025) -
Towards Unifying Interpretability and Control: Evaluation via Intervention
por: Bhalla, Usha, et al.
Publicado: (2024)