Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Żukowska, Nina, Stammer, Wolfgang, Schiele, Bernt, Fischer, Jonas |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Semantic Delta: An Interpretable Signal Differentiating Human and LLMs Dialogue
by: Scantamburlo, Riccardo, et al.
Published: (2026)
by: Scantamburlo, Riccardo, et al.
Published: (2026)
Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations
by: Mahale, Ajay Pravin
Published: (2026)
by: Mahale, Ajay Pravin
Published: (2026)
What Makes a Good AI Review? Concern-Level Diagnostics for AI Peer Review
by: Jin, Ming
Published: (2026)
by: Jin, Ming
Published: (2026)
Attribution-based Explanations for Markov Decision Processes
by: Kobialka, Paul, et al.
Published: (2026)
by: Kobialka, Paul, et al.
Published: (2026)
Counterfactual Strategies for Markov Decision Processes
by: Kobialka, Paul, et al.
Published: (2025)
by: Kobialka, Paul, et al.
Published: (2025)
Short-Term Gain, Long-Term Fragility: AI Labor Substitution and the Erosion of Sustainable Capability
by: Rohde, Wolfgang
Published: (2026)
by: Rohde, Wolfgang
Published: (2026)
Certified Circuits: Stability Guarantees for Mechanistic Circuits
by: Anani, Alaa, et al.
Published: (2026)
by: Anani, Alaa, et al.
Published: (2026)
An Exploration of Default Images in Text-to-Image Generation
by: Simonen, Hannu, et al.
Published: (2025)
by: Simonen, Hannu, et al.
Published: (2025)
Prompt Engineer: Analyzing Hard and Soft Skill Requirements in the AI Job Market
by: Vu, An, et al.
Published: (2025)
by: Vu, An, et al.
Published: (2025)
Mapping the Challenges of HCI: An Application and Evaluation of ChatGPT for Mining Insights at Scale
by: Oppenlaender, Jonas, et al.
Published: (2023)
by: Oppenlaender, Jonas, et al.
Published: (2023)
Defining and Quantifying Creative Behavior in Popular Image Generators
by: Ramaswamy, Aditi, et al.
Published: (2025)
by: Ramaswamy, Aditi, et al.
Published: (2025)
Evaluating Vision Transformer Models for Visual Quality Control in Industrial Manufacturing
by: Alber, Miriam, et al.
Published: (2024)
by: Alber, Miriam, et al.
Published: (2024)
A Survey on Semantic Modeling for Building Energy Management
by: Aniakor, Miracle, et al.
Published: (2024)
by: Aniakor, Miracle, et al.
Published: (2024)
DEXTER: Diffusion-Guided EXplanations with TExtual Reasoning for Vision Models
by: Carnemolla, Simone, et al.
Published: (2025)
by: Carnemolla, Simone, et al.
Published: (2025)
Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
by: Nainani, Jatin, et al.
Published: (2024)
by: Nainani, Jatin, et al.
Published: (2024)
Real-Time Band-Grouped Vocal Denoising Using Sigmoid-Driven Ideal Ratio Masking
by: Williams, Daniel
Published: (2026)
by: Williams, Daniel
Published: (2026)
Exploring Model Invariance with Discrete Search for Ultra-Low-Bit Quantization
by: Wen, Yuqiao, et al.
Published: (2025)
by: Wen, Yuqiao, et al.
Published: (2025)
EBBS: An Ensemble with Bi-Level Beam Search for Zero-Shot Machine Translation
by: Wen, Yuqiao, et al.
Published: (2024)
by: Wen, Yuqiao, et al.
Published: (2024)
Average Attention Transformers and Arithmetic Circuits
by: Ehrmuth, Lena, et al.
Published: (2026)
by: Ehrmuth, Lena, et al.
Published: (2026)
Balanced Anomaly-guided Ego-graph Diffusion Model for Inductive Graph Anomaly Detection
by: Wei, Chunyu, et al.
Published: (2026)
by: Wei, Chunyu, et al.
Published: (2026)
Knowledge Distillation for Large Language Models
by: La Torre, Alejandro Paredes, et al.
Published: (2026)
by: La Torre, Alejandro Paredes, et al.
Published: (2026)
Revisiting Clustering of Neural Bandits: Selective Reinitialization for Mitigating Loss of Plasticity
by: Su, Zhiyuan, et al.
Published: (2025)
by: Su, Zhiyuan, et al.
Published: (2025)
Deepfake Technology Unveiled: The Commoditization of AI and Its Impact on Digital Trust
by: Popa, Claudiu, et al.
Published: (2025)
by: Popa, Claudiu, et al.
Published: (2025)
SCOR: A Framework for Responsible AI Innovation in Digital Ecosystems
by: Torkestani, Mohammad Saleh, et al.
Published: (2025)
by: Torkestani, Mohammad Saleh, et al.
Published: (2025)
Establishing Best Practices for Building Rigorous Agentic Benchmarks
by: Zhu, Yuxuan, et al.
Published: (2025)
by: Zhu, Yuxuan, et al.
Published: (2025)
Positional encoding is not the same as context: A study on positional encoding for sequential recommendation
by: Lopez-Avila, Alejo, et al.
Published: (2024)
by: Lopez-Avila, Alejo, et al.
Published: (2024)
The QCET Taxonomy of Standard Quality Criterion Names and Definitions for the Evaluation of NLP Systems
by: Belz, Anya, et al.
Published: (2025)
by: Belz, Anya, et al.
Published: (2025)
Machine Learning for Quantum Computing Specialists
by: Goldsmith, Daniel, et al.
Published: (2024)
by: Goldsmith, Daniel, et al.
Published: (2024)
Integrated Water Resource Management in the Segura Hydrographic Basin: An Artificial Intelligence Approach
by: Otamendi, Urtzi, et al.
Published: (2024)
by: Otamendi, Urtzi, et al.
Published: (2024)
Balancing the Scales: A Comprehensive Study on Tackling Class Imbalance in Binary Classification
by: Abdelhamid, Mohamed, et al.
Published: (2024)
by: Abdelhamid, Mohamed, et al.
Published: (2024)
Challenges for Responsible AI Design and Workflow Integration in Healthcare: A Case Study of Automatic Feeding Tube Qualification in Radiology
by: Thieme, Anja, et al.
Published: (2024)
by: Thieme, Anja, et al.
Published: (2024)
Skeleton-based sign language recognition using a dual-stream spatio-temporal dynamic graph convolutional network
by: Liu, Liangjin, et al.
Published: (2025)
by: Liu, Liangjin, et al.
Published: (2025)
A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models
by: Rai, Daking, et al.
Published: (2024)
by: Rai, Daking, et al.
Published: (2024)
Emergence of Fragility in LLM-based Social Networks: the Case of Moltbook
by: Sodano, Luca, et al.
Published: (2026)
by: Sodano, Luca, et al.
Published: (2026)
L3AC: Towards a Lightweight and Lossless Audio Codec
by: Zhai, Linwei, et al.
Published: (2025)
by: Zhai, Linwei, et al.
Published: (2025)
PrivSpike: Employing Homomorphic Encryption for Private Inference of Deep Spiking Neural Networks
by: Njungle, Nges Brian, et al.
Published: (2025)
by: Njungle, Nges Brian, et al.
Published: (2025)
Does GPT Really Get It? A Hierarchical Scale to Quantify Human vs AI's Understanding of Algorithms
by: Reid, Mirabel, et al.
Published: (2024)
by: Reid, Mirabel, et al.
Published: (2024)
A Framework for Collaborating a Large Language Model Tool in Brainstorming for Triggering Creative Thoughts
by: Chang, Hung-Fu, et al.
Published: (2024)
by: Chang, Hung-Fu, et al.
Published: (2024)
Heterogeneous network and graph attention auto-encoder for LncRNA-disease association prediction
by: Liu, Jin-Xing, et al.
Published: (2024)
by: Liu, Jin-Xing, et al.
Published: (2024)
Guiding Multimodal Large Language Models with Blind and Low Vision People Visual Questions for Proactive Visual Interpretations
by: Penuela, Ricardo Gonzalez, et al.
Published: (2025)
by: Penuela, Ricardo Gonzalez, et al.
Published: (2025)
Similar Items
-
Semantic Delta: An Interpretable Signal Differentiating Human and LLMs Dialogue
by: Scantamburlo, Riccardo, et al.
Published: (2026) -
Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations
by: Mahale, Ajay Pravin
Published: (2026) -
What Makes a Good AI Review? Concern-Level Diagnostics for AI Peer Review
by: Jin, Ming
Published: (2026) -
Attribution-based Explanations for Markov Decision Processes
by: Kobialka, Paul, et al.
Published: (2026) -
Counterfactual Strategies for Markov Decision Processes
by: Kobialka, Paul, et al.
Published: (2025)