Certified Circuits: Stability Guarantees for Mechanistic Circuits
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Anani, Alaa, Lorenz, Tobias, Schiele, Bernt, Fritz, Mario, Fischer, Jonas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pixel-level Certified Explanations via Randomized Smoothing
von: Anani, Alaa, et al.
Veröffentlicht: (2025)
von: Anani, Alaa, et al.
Veröffentlicht: (2025)
Adaptive Hierarchical Certification for Segmentation using Randomized Smoothing
von: Anani, Alaa, et al.
Veröffentlicht: (2024)
von: Anani, Alaa, et al.
Veröffentlicht: (2024)
FaCT: Faithful Concept Traces for Explaining Neural Network Decisions
von: Parchami-Araghi, Amin, et al.
Veröffentlicht: (2025)
von: Parchami-Araghi, Amin, et al.
Veröffentlicht: (2025)
CFM: Language-aligned Concept Foundation Model for Vision
von: Wittenmayer, Kai, et al.
Veröffentlicht: (2026)
von: Wittenmayer, Kai, et al.
Veröffentlicht: (2026)
VITAL: More Understandable Feature Visualization through Distribution Alignment and Relevant Information Flow
von: Gorgun, Ada, et al.
Veröffentlicht: (2025)
von: Gorgun, Ada, et al.
Veröffentlicht: (2025)
Improving Alignment and Robustness with Circuit Breakers
von: Zou, Andy, et al.
Veröffentlicht: (2024)
von: Zou, Andy, et al.
Veröffentlicht: (2024)
Better Understanding Differences in Attribution Methods via Systematic Evaluations
von: Rao, Sukrut, et al.
Veröffentlicht: (2023)
von: Rao, Sukrut, et al.
Veröffentlicht: (2023)
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
von: Che, Liwei, et al.
Veröffentlicht: (2026)
von: Che, Liwei, et al.
Veröffentlicht: (2026)
MM-TS: Multi-Modal Temperature and Margin Schedules for Contrastive Learning with Long-Tail Data
von: Sheludzko, Siarhei, et al.
Veröffentlicht: (2026)
von: Sheludzko, Siarhei, et al.
Veröffentlicht: (2026)
Locating Demographic Bias at the Attention-Head Level in CLIP's Vision Encoder
von: Yasser, Alaa, et al.
Veröffentlicht: (2026)
von: Yasser, Alaa, et al.
Veröffentlicht: (2026)
Sum of Group Error Differences: A Critical Examination of Bias Evaluation in Biometric Verification and a Dual-Metric Measure
von: Elobaid, Alaa, et al.
Veröffentlicht: (2024)
von: Elobaid, Alaa, et al.
Veröffentlicht: (2024)
B-cosification: Transforming Deep Neural Networks to be Inherently Interpretable
von: Arya, Shreyash, et al.
Veröffentlicht: (2024)
von: Arya, Shreyash, et al.
Veröffentlicht: (2024)
Discover-then-Name: Task-Agnostic Concept Bottlenecks via Automated Concept Discovery
von: Rao, Sukrut, et al.
Veröffentlicht: (2024)
von: Rao, Sukrut, et al.
Veröffentlicht: (2024)
What Matters for Scalable and Robust Learning in End-to-End Driving Planners?
von: Holtz, David, et al.
Veröffentlicht: (2026)
von: Holtz, David, et al.
Veröffentlicht: (2026)
Studying How to Efficiently and Effectively Guide Models with Explanations
von: Rao, Sukrut, et al.
Veröffentlicht: (2023)
von: Rao, Sukrut, et al.
Veröffentlicht: (2023)
Good Teachers Explain: Explanation-Enhanced Knowledge Distillation
von: Parchami-Araghi, Amin, et al.
Veröffentlicht: (2024)
von: Parchami-Araghi, Amin, et al.
Veröffentlicht: (2024)
Granular Concept Circuits: Toward a Fine-Grained Circuit Discovery for Concept Representations
von: Kwon, Dahee, et al.
Veröffentlicht: (2025)
von: Kwon, Dahee, et al.
Veröffentlicht: (2025)
Samba: Synchronized Set-of-Sequences Modeling for Multiple Object Tracking
von: Segu, Mattia, et al.
Veröffentlicht: (2024)
von: Segu, Mattia, et al.
Veröffentlicht: (2024)
Automatic Discovery of Visual Circuits
von: Rajaram, Achyuta, et al.
Veröffentlicht: (2024)
von: Rajaram, Achyuta, et al.
Veröffentlicht: (2024)
Toward Understanding Unlearning Difficulty: A Mechanistic Perspective and Circuit-Guided Difficulty Metric
von: Cheng, Jiali, et al.
Veröffentlicht: (2026)
von: Cheng, Jiali, et al.
Veröffentlicht: (2026)
Optimising for Interpretability: Convolutional Dynamic Alignment Networks
von: Böhle, Moritz, et al.
Veröffentlicht: (2021)
von: Böhle, Moritz, et al.
Veröffentlicht: (2021)
B-cos Alignment for Inherently Interpretable CNNs and Vision Transformers
von: Böhle, Moritz, et al.
Veröffentlicht: (2023)
von: Böhle, Moritz, et al.
Veröffentlicht: (2023)
Temporal Concept Dynamics in Diffusion Models via Prompt-Conditioned Interventions
von: Gorgun, Ada, et al.
Veröffentlicht: (2025)
von: Gorgun, Ada, et al.
Veröffentlicht: (2025)
Interpretable 3D Neural Object Volumes for Robust Conceptual Reasoning
von: Pham, Nhi, et al.
Veröffentlicht: (2025)
von: Pham, Nhi, et al.
Veröffentlicht: (2025)
Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers
von: Żukowska, Nina, et al.
Veröffentlicht: (2026)
von: Żukowska, Nina, et al.
Veröffentlicht: (2026)
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
von: Hufe, Lorenz, et al.
Veröffentlicht: (2025)
von: Hufe, Lorenz, et al.
Veröffentlicht: (2025)
Where Reliability Lives in Vision-Language Models: A Mechanistic Study of Attention, Hidden States, and Causal Circuits
von: Mann, Logan, et al.
Veröffentlicht: (2026)
von: Mann, Logan, et al.
Veröffentlicht: (2026)
PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding
von: Kuzucu, Selim, et al.
Veröffentlicht: (2026)
von: Kuzucu, Selim, et al.
Veröffentlicht: (2026)
AnalogRetriever: Learning Cross-Modal Representations for Analog Circuit Retrieval
von: Wang, Yihan, et al.
Veröffentlicht: (2026)
von: Wang, Yihan, et al.
Veröffentlicht: (2026)
PCB-Vision: A Multiscene RGB-Hyperspectral Benchmark Dataset of Printed Circuit Boards
von: Arbash, Elias, et al.
Veröffentlicht: (2024)
von: Arbash, Elias, et al.
Veröffentlicht: (2024)
DAVE: Distribution-aware Attribution via ViT Gradient Decomposition
von: Wróbel, Adam, et al.
Veröffentlicht: (2026)
von: Wróbel, Adam, et al.
Veröffentlicht: (2026)
SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs
von: Avogaro, Niccolo, et al.
Veröffentlicht: (2026)
von: Avogaro, Niccolo, et al.
Veröffentlicht: (2026)
TSOM: Small Object Motion Detection Neural Network Inspired by Avian Visual Circuit
von: Hu, Pignge, et al.
Veröffentlicht: (2024)
von: Hu, Pignge, et al.
Veröffentlicht: (2024)
CoE: Chain-of-Explanation via Automatic Visual Concept Circuit Description and Polysemanticity Quantification
von: Yu, Wenlong, et al.
Veröffentlicht: (2025)
von: Yu, Wenlong, et al.
Veröffentlicht: (2025)
An Embarrassingly Simple Baseline for Imbalanced Semi-Supervised Learning
von: Chen, Hao, et al.
Veröffentlicht: (2022)
von: Chen, Hao, et al.
Veröffentlicht: (2022)
Are Multimodal LLMs Ready for Clinical Dermatology? A Real-World Evaluation in Dermatology
von: Jiang, Roy, et al.
Veröffentlicht: (2026)
von: Jiang, Roy, et al.
Veröffentlicht: (2026)
Mass Concept Erasure in Diffusion Models with Concept Hierarchy
von: Tu, Jiahang, et al.
Veröffentlicht: (2026)
von: Tu, Jiahang, et al.
Veröffentlicht: (2026)
EDU-CIRCUIT-HW: Evaluating Multimodal Large Language Models on Real-World University-Level STEM Student Handwritten Solutions
von: Sun, Weiyu, et al.
Veröffentlicht: (2026)
von: Sun, Weiyu, et al.
Veröffentlicht: (2026)
Aesthetics as Structural Harm: Algorithmic Lookism Across Text-to-Image Generation and Classification
von: Doh, Miriam, et al.
Veröffentlicht: (2026)
von: Doh, Miriam, et al.
Veröffentlicht: (2026)
OrthoEraser: Coupled-Neuron Orthogonal Projection for Concept Erasure
von: Shi, Chuancheng, et al.
Veröffentlicht: (2026)
von: Shi, Chuancheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Pixel-level Certified Explanations via Randomized Smoothing
von: Anani, Alaa, et al.
Veröffentlicht: (2025) -
Adaptive Hierarchical Certification for Segmentation using Randomized Smoothing
von: Anani, Alaa, et al.
Veröffentlicht: (2024) -
FaCT: Faithful Concept Traces for Explaining Neural Network Decisions
von: Parchami-Araghi, Amin, et al.
Veröffentlicht: (2025) -
CFM: Language-aligned Concept Foundation Model for Vision
von: Wittenmayer, Kai, et al.
Veröffentlicht: (2026) -
VITAL: More Understandable Feature Visualization through Distribution Alignment and Relevant Information Flow
von: Gorgun, Ada, et al.
Veröffentlicht: (2025)