Concept Gradient: Concept-based Interpretation Without Linear Assumption
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bai, Andrew, Yeh, Chih-Kuan, Ravikumar, Pradeep, Lin, Neil Y. C., Hsieh, Cho-Jui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An Efficient Rehearsal Scheme for Catastrophic Forgetting Mitigation during Multi-stage Fine-tuning
von: Bai, Andrew, et al.
Veröffentlicht: (2024)
von: Bai, Andrew, et al.
Veröffentlicht: (2024)
Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models
von: Bai, Andrew, et al.
Veröffentlicht: (2025)
von: Bai, Andrew, et al.
Veröffentlicht: (2025)
CLUE: Concept-Level Uncertainty Estimation for Large Language Models
von: Wang, Yu-Hsiang, et al.
Veröffentlicht: (2024)
von: Wang, Yu-Hsiang, et al.
Veröffentlicht: (2024)
Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation Models
von: Rajendran, Goutham, et al.
Veröffentlicht: (2024)
von: Rajendran, Goutham, et al.
Veröffentlicht: (2024)
A Unifying Framework for Unsupervised Concept Extraction
von: Squires, Chandler, et al.
Veröffentlicht: (2026)
von: Squires, Chandler, et al.
Veröffentlicht: (2026)
Unlabeled Data Improves Fine-Grained Image Zero-shot Classification with Multimodal LLMs
von: Hong, Yunqi, et al.
Veröffentlicht: (2025)
von: Hong, Yunqi, et al.
Veröffentlicht: (2025)
Data Attribution for Diffusion Models: Timestep-induced Bias in Influence Estimation
von: Xie, Tong, et al.
Veröffentlicht: (2024)
von: Xie, Tong, et al.
Veröffentlicht: (2024)
Online Continuous Hyperparameter Optimization for Generalized Linear Contextual Bandits
von: Kang, Yue, et al.
Veröffentlicht: (2023)
von: Kang, Yue, et al.
Veröffentlicht: (2023)
On the Loss of Context-awareness in General Instruction Fine-tuning
von: Wang, Yihan, et al.
Veröffentlicht: (2024)
von: Wang, Yihan, et al.
Veröffentlicht: (2024)
Embedding Space Selection for Detecting Memorization and Fingerprinting in Generative Models
von: He, Jack, et al.
Veröffentlicht: (2024)
von: He, Jack, et al.
Veröffentlicht: (2024)
Concept Bottleneck Models Without Predefined Concepts
von: Schrodi, Simon, et al.
Veröffentlicht: (2024)
von: Schrodi, Simon, et al.
Veröffentlicht: (2024)
Hierarchical Concept-based Interpretable Models
von: Hill, Oscar, et al.
Veröffentlicht: (2026)
von: Hill, Oscar, et al.
Veröffentlicht: (2026)
Tree of Concepts: Interpretable Continual Learners in Non-Stationary Clinical Domains
von: Cho, Dongkyu, et al.
Veröffentlicht: (2026)
von: Cho, Dongkyu, et al.
Veröffentlicht: (2026)
Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
von: Hindupur, Sai Sumedh R., et al.
Veröffentlicht: (2025)
von: Hindupur, Sai Sumedh R., et al.
Veröffentlicht: (2025)
CAT: Interpretable Concept-based Taylor Additive Models
von: Duong, Viet, et al.
Veröffentlicht: (2024)
von: Duong, Viet, et al.
Veröffentlicht: (2024)
ICED: Concept-level Machine Unlearning via Interpretable Concept Decomposition
von: Lin, Shen, et al.
Veröffentlicht: (2026)
von: Lin, Shen, et al.
Veröffentlicht: (2026)
Efficient Frameworks for Generalized Low-Rank Matrix Bandit Problems
von: Kang, Yue, et al.
Veröffentlicht: (2024)
von: Kang, Yue, et al.
Veröffentlicht: (2024)
Low-rank Matrix Bandits with Heavy-tailed Rewards
von: Kang, Yue, et al.
Veröffentlicht: (2024)
von: Kang, Yue, et al.
Veröffentlicht: (2024)
On the Origins of Linear Representations in Large Language Models
von: Jiang, Yibo, et al.
Veröffentlicht: (2024)
von: Jiang, Yibo, et al.
Veröffentlicht: (2024)
Unlearning Concepts in Diffusion Model via Concept Domain Correction and Concept Preserving Gradient
von: Wu, Yongliang, et al.
Veröffentlicht: (2024)
von: Wu, Yongliang, et al.
Veröffentlicht: (2024)
Identifying General Mechanism Shifts in Linear Causal Representations
von: Chen, Tianyu, et al.
Veröffentlicht: (2024)
von: Chen, Tianyu, et al.
Veröffentlicht: (2024)
Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)
von: Bhalla, Usha, et al.
Veröffentlicht: (2024)
von: Bhalla, Usha, et al.
Veröffentlicht: (2024)
Linear Adversarial Concept Erasure
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
Explaining Concept Shift with Interpretable Feature Attribution
von: Lyu, Ruiqi, et al.
Veröffentlicht: (2025)
von: Lyu, Ruiqi, et al.
Veröffentlicht: (2025)
Interpretable Reward Modeling with Active Concept Bottlenecks
von: Laguna, Sonia, et al.
Veröffentlicht: (2025)
von: Laguna, Sonia, et al.
Veröffentlicht: (2025)
ConceptCaps: a Distilled Concept Dataset for Interpretability in Music Models
von: Sienkiewicz, Bruno, et al.
Veröffentlicht: (2026)
von: Sienkiewicz, Bruno, et al.
Veröffentlicht: (2026)
Interpretable Concept-Based Memory Reasoning
von: Debot, David, et al.
Veröffentlicht: (2024)
von: Debot, David, et al.
Veröffentlicht: (2024)
Leakage and Interpretability in Concept-Based Models
von: Parisini, Enrico, et al.
Veröffentlicht: (2025)
von: Parisini, Enrico, et al.
Veröffentlicht: (2025)
Interpretable Prognostics with Concept Bottleneck Models
von: Forest, Florent, et al.
Veröffentlicht: (2024)
von: Forest, Florent, et al.
Veröffentlicht: (2024)
What Does Preference Learning Recover from Pairwise Comparison Data?
von: Pukdee, Rattana, et al.
Veröffentlicht: (2026)
von: Pukdee, Rattana, et al.
Veröffentlicht: (2026)
Provably Robust Training of Quantum Circuit Classifiers Against Parameter Noise
von: Tecot, Lucas, et al.
Veröffentlicht: (2025)
von: Tecot, Lucas, et al.
Veröffentlicht: (2025)
From Segments to Concepts: Interpretable Image Classification via Concept-Guided Segmentation
von: Eisenberg, Ran, et al.
Veröffentlicht: (2025)
von: Eisenberg, Ran, et al.
Veröffentlicht: (2025)
Understanding Augmentation-based Self-Supervised Representation Learning via RKHS Approximation and Regression
von: Zhai, Runtian, et al.
Veröffentlicht: (2023)
von: Zhai, Runtian, et al.
Veröffentlicht: (2023)
Interpretability for Multimodal Emotion Recognition using Concept Activation Vectors
von: Asokan, Ashish Ramayee, et al.
Veröffentlicht: (2022)
von: Asokan, Ashish Ramayee, et al.
Veröffentlicht: (2022)
Exploiting Interpretable Capabilities with Concept-Enhanced Diffusion and Prototype Networks
von: Carballo-Castro, Alba, et al.
Veröffentlicht: (2024)
von: Carballo-Castro, Alba, et al.
Veröffentlicht: (2024)
Federated Concept-Based Models: Interpretable models with distributed supervision
von: Fenoglio, Dario, et al.
Veröffentlicht: (2026)
von: Fenoglio, Dario, et al.
Veröffentlicht: (2026)
Interpretable Neural-Symbolic Concept Reasoning
von: Barbiero, Pietro, et al.
Veröffentlicht: (2023)
von: Barbiero, Pietro, et al.
Veröffentlicht: (2023)
Preserving Task-Relevant Information Under Linear Concept Removal
von: Holstege, Floris, et al.
Veröffentlicht: (2025)
von: Holstege, Floris, et al.
Veröffentlicht: (2025)
Interpretable Concept Bottlenecks to Align Reinforcement Learning Agents
von: Delfosse, Quentin, et al.
Veröffentlicht: (2024)
von: Delfosse, Quentin, et al.
Veröffentlicht: (2024)
Self-explaining Neural Network with Concept-based Explanations for ICU Mortality Prediction
von: Kumar, Sayantan, et al.
Veröffentlicht: (2021)
von: Kumar, Sayantan, et al.
Veröffentlicht: (2021)
Ähnliche Einträge
-
An Efficient Rehearsal Scheme for Catastrophic Forgetting Mitigation during Multi-stage Fine-tuning
von: Bai, Andrew, et al.
Veröffentlicht: (2024) -
Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models
von: Bai, Andrew, et al.
Veröffentlicht: (2025) -
CLUE: Concept-Level Uncertainty Estimation for Large Language Models
von: Wang, Yu-Hsiang, et al.
Veröffentlicht: (2024) -
Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation Models
von: Rajendran, Goutham, et al.
Veröffentlicht: (2024) -
A Unifying Framework for Unsupervised Concept Extraction
von: Squires, Chandler, et al.
Veröffentlicht: (2026)