Mechanistic Interpretability of Fine-Tuned Vision Transformers on Distorted Images: Decoding Attention Head Behavior for Transparent and Trustworthy AI
Fuente:
arXiv
Saved in:
| Main Author: | Bahador, Nooshin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Localized Definitions and Distributed Reasoning: A Proof-of-Concept Mechanistic Interpretability Study via Activation Patching
by: Bahador, Nooshin
Published: (2025)
by: Bahador, Nooshin
Published: (2025)
Chirp Localization via Fine-Tuned Transformer Model: A Proof-of-Concept Study
by: Bahador, Nooshin, et al.
Published: (2025)
by: Bahador, Nooshin, et al.
Published: (2025)
Semi-Supervised Anomaly Detection Pipeline for SOZ Localization Using Ictal-Related Chirp
by: Bahador, Nooshin, et al.
Published: (2025)
by: Bahador, Nooshin, et al.
Published: (2025)
Transparent, Evaluable, and Accessible Data Agents: A Proof-of-Concept Framework
by: Bahador, Nooshin
Published: (2025)
by: Bahador, Nooshin
Published: (2025)
Vision Transformers Exhibit Human-Like Biases: Evidence of Orientation and Color Selectivity, Categorical Perception, and Phase Transitions
by: Bahador, Nooshin
Published: (2025)
by: Bahador, Nooshin
Published: (2025)
Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability
by: Yan, Ge, et al.
Published: (2025)
by: Yan, Ge, et al.
Published: (2025)
Towards Mechanistic Interpretability of Graph Transformers via Attention Graphs
by: El, Batu, et al.
Published: (2025)
by: El, Batu, et al.
Published: (2025)
Even Heads Fix Odd Errors: Mechanistic Discovery and Surgical Repair in Transformer Attention
by: Sandoval, Gustavo
Published: (2025)
by: Sandoval, Gustavo
Published: (2025)
Panoptic Pairwise Distortion Graph
by: Janjua, Muhammad Kamran, et al.
Published: (2026)
by: Janjua, Muhammad Kamran, et al.
Published: (2026)
nnterp: A Standardized Interface for Mechanistic Interpretability of Transformers
by: Dumas, Clément
Published: (2025)
by: Dumas, Clément
Published: (2025)
Mechanistic Interpretability for Transformer-based Time Series Classification
by: Kalnāre, Matīss, et al.
Published: (2025)
by: Kalnāre, Matīss, et al.
Published: (2025)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
Towards Unified Attribution in Explainable AI, Data-Centric AI, and Mechanistic Interpretability
by: Zhang, Shichang, et al.
Published: (2025)
by: Zhang, Shichang, et al.
Published: (2025)
Transparent AI: The Case for Interpretability and Explainability
by: Ramachandram, Dhanesh, et al.
Published: (2025)
by: Ramachandram, Dhanesh, et al.
Published: (2025)
Ravan: Multi-Head Low-Rank Adaptation for Federated Fine-Tuning
by: Raje, Arian, et al.
Published: (2025)
by: Raje, Arian, et al.
Published: (2025)
Decoding Style: Efficient Fine-Tuning of LLMs for Image-Guided Outfit Recommendation with Preference
by: Forouzandehmehr, Najmeh, et al.
Published: (2024)
by: Forouzandehmehr, Najmeh, et al.
Published: (2024)
Massive Activations in Graph Neural Networks: Decoding Attention for Domain-Dependent Interpretability
by: Bini, Lorenzo, et al.
Published: (2024)
by: Bini, Lorenzo, et al.
Published: (2024)
Challenges in Mechanistically Interpreting Model Representations
by: Golechha, Satvik, et al.
Published: (2024)
by: Golechha, Satvik, et al.
Published: (2024)
Graph Attention Network for Lane-Wise and Topology-Invariant Intersection Traffic Simulation
by: Yousefzadeh, Nooshin, et al.
Published: (2024)
by: Yousefzadeh, Nooshin, et al.
Published: (2024)
Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
by: Nainani, Jatin, et al.
Published: (2024)
by: Nainani, Jatin, et al.
Published: (2024)
Meta-Models: An Architecture for Decoding LLM Behaviors Through Interpreted Embeddings and Natural Language
by: Costarelli, Anthony, et al.
Published: (2024)
by: Costarelli, Anthony, et al.
Published: (2024)
MINAR: Mechanistic Interpretability for Neural Algorithmic Reasoning
by: He, Jesse, et al.
Published: (2026)
by: He, Jesse, et al.
Published: (2026)
Classifying Clinical Outcome of Epilepsy Patients with Ictal Chirp Embeddings
by: Bahador, Nooshin, et al.
Published: (2025)
by: Bahador, Nooshin, et al.
Published: (2025)
Distilling Linearized Behavior into Non-Linear Fine-Tuning for Effective Task Arithmetic
by: Sommariva, Thomas, et al.
Published: (2026)
by: Sommariva, Thomas, et al.
Published: (2026)
The Anxiety of Influence: Bloom Filters in Transformer Attention Heads
by: Balogh, Peter
Published: (2026)
by: Balogh, Peter
Published: (2026)
Trustworthy AI Must Account for Interactions
by: Cresswell, Jesse C.
Published: (2025)
by: Cresswell, Jesse C.
Published: (2025)
Eliciting Fine-Tuned Transformer Capabilities via Inference-Time Techniques
by: Sharma, Asankhaya
Published: (2025)
by: Sharma, Asankhaya
Published: (2025)
Adaptive Ensembles of Fine-Tuned Transformers for LLM-Generated Text Detection
by: Lai, Zhixin, et al.
Published: (2024)
by: Lai, Zhixin, et al.
Published: (2024)
Dispatch-Aware Ragged Attention for Pruned Vision Transformers
by: Abdellatif, Seifeldin, et al.
Published: (2026)
by: Abdellatif, Seifeldin, et al.
Published: (2026)
Adaptive Prompt Tuning: Vision Guided Prompt Tuning with Cross-Attention for Fine-Grained Few-Shot Learning
by: Brouwer, Eric, et al.
Published: (2024)
by: Brouwer, Eric, et al.
Published: (2024)
Tuning for Trustworthiness -- Balancing Performance and Explanation Consistency in Neural Network Optimization
by: Hinterleitner, Alexander, et al.
Published: (2025)
by: Hinterleitner, Alexander, et al.
Published: (2025)
HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization
by: Williams, Jorge L. Ruiz
Published: (2026)
by: Williams, Jorge L. Ruiz
Published: (2026)
AttenGluco: Multimodal Transformer-Based Blood Glucose Forecasting on AI-READI Dataset
by: Farahmand, Ebrahim, et al.
Published: (2025)
by: Farahmand, Ebrahim, et al.
Published: (2025)
RecurFormer: Not All Transformer Heads Need Self-Attention
by: Yan, Ruiqing, et al.
Published: (2024)
by: Yan, Ruiqing, et al.
Published: (2024)
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
by: Chen, Yilong, et al.
Published: (2024)
by: Chen, Yilong, et al.
Published: (2024)
Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video
by: Joseph, Sonia, et al.
Published: (2025)
by: Joseph, Sonia, et al.
Published: (2025)
QFlash: Bridging Quantization and Memory Efficiency in Vision Transformer Attention
by: Oh, Sehyeon, et al.
Published: (2026)
by: Oh, Sehyeon, et al.
Published: (2026)
Self-Tuning Sparse Attention: Multi-Fidelity Hyperparameter Optimization for Transformer Acceleration
by: Dev, Arundhathi, et al.
Published: (2026)
by: Dev, Arundhathi, et al.
Published: (2026)
RPO: Fine-Tuning Visual Generative Models via Rich Vision-Language Preferences
by: Zhao, Hanyang, et al.
Published: (2025)
by: Zhao, Hanyang, et al.
Published: (2025)
A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models
by: Lin, Zihao, et al.
Published: (2025)
by: Lin, Zihao, et al.
Published: (2025)
Similar Items
-
Localized Definitions and Distributed Reasoning: A Proof-of-Concept Mechanistic Interpretability Study via Activation Patching
by: Bahador, Nooshin
Published: (2025) -
Chirp Localization via Fine-Tuned Transformer Model: A Proof-of-Concept Study
by: Bahador, Nooshin, et al.
Published: (2025) -
Semi-Supervised Anomaly Detection Pipeline for SOZ Localization Using Ictal-Related Chirp
by: Bahador, Nooshin, et al.
Published: (2025) -
Transparent, Evaluable, and Accessible Data Agents: A Proof-of-Concept Framework
by: Bahador, Nooshin
Published: (2025) -
Vision Transformers Exhibit Human-Like Biases: Evidence of Orientation and Color Selectivity, Categorical Perception, and Phase Transitions
by: Bahador, Nooshin
Published: (2025)