Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits
Fuente:
arXiv
Saved in:
| Main Authors: | Ahmad, Areeb, Joshi, Abhinav, Modi, Ashutosh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Calibration Across Layers: Understanding Calibration Evolution in LLMs
by: Joshi, Abhinav, et al.
Published: (2025)
by: Joshi, Abhinav, et al.
Published: (2025)
COLD: Causal reasOning in cLosed Daily activities
by: Joshi, Abhinav, et al.
Published: (2024)
by: Joshi, Abhinav, et al.
Published: (2024)
Towards Quantifying Commonsense Reasoning with Mechanistic Insights
by: Joshi, Abhinav, et al.
Published: (2025)
by: Joshi, Abhinav, et al.
Published: (2025)
Geometry of Decision Making in Language Models
by: Joshi, Abhinav, et al.
Published: (2025)
by: Joshi, Abhinav, et al.
Published: (2025)
POSESTITCH-SLT: Linguistically Inspired Pose-Stitching for End-to-End Sign Language Translation
by: Joshi, Abhinav, et al.
Published: (2025)
by: Joshi, Abhinav, et al.
Published: (2025)
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
by: Joshi, Abhinav, et al.
Published: (2024)
by: Joshi, Abhinav, et al.
Published: (2024)
IL-TUR: Benchmark for Indian Legal Text Understanding and Reasoning
by: Joshi, Abhinav, et al.
Published: (2024)
by: Joshi, Abhinav, et al.
Published: (2024)
Calibration vs Decision Making: Revisiting the Reliability Paradox in Unlearned Language Models
by: Shukla, Divyaksh, et al.
Published: (2026)
by: Shukla, Divyaksh, et al.
Published: (2026)
IITK at SemEval-2024 Task 2: Exploring the Capabilities of LLMs for Safe Biomedical Natural Language Inference for Clinical Trials
by: Mandal, Shreyasi, et al.
Published: (2024)
by: Mandal, Shreyasi, et al.
Published: (2024)
LoRMA: Low-Rank Multiplicative Adaptation for LLMs
by: Bihany, Harsh, et al.
Published: (2025)
by: Bihany, Harsh, et al.
Published: (2025)
IITK at SemEval-2024 Task 4: Hierarchical Embeddings for Detection of Persuasion Techniques in Memes
by: Chikoti, Shreenaga, et al.
Published: (2024)
by: Chikoti, Shreenaga, et al.
Published: (2024)
IITK at SemEval-2024 Task 10: Who is the speaker? Improving Emotion Recognition and Flip Reasoning in Conversations via Speaker Embeddings
by: Patel, Shubham, et al.
Published: (2024)
by: Patel, Shubham, et al.
Published: (2024)
iSign: A Benchmark for Indian Sign Language Processing
by: Joshi, Abhinav, et al.
Published: (2024)
by: Joshi, Abhinav, et al.
Published: (2024)
Circuit Insights: Towards Interpretability Beyond Activations
by: Golimblevskaia, Elena, et al.
Published: (2025)
by: Golimblevskaia, Elena, et al.
Published: (2025)
IITK at SemEval-2024 Task 1: Contrastive Learning and Autoencoders for Semantic Textual Relatedness in Multilingual Texts
by: Basak, Udvas, et al.
Published: (2024)
by: Basak, Udvas, et al.
Published: (2024)
BookSQL: A Large Scale Text-to-SQL Dataset for Accounting Domain
by: Kumar, Rahul, et al.
Published: (2024)
by: Kumar, Rahul, et al.
Published: (2024)
SVFT: Parameter-Efficient Fine-Tuning with Singular Vectors
by: Lingam, Vijay, et al.
Published: (2024)
by: Lingam, Vijay, et al.
Published: (2024)
CoMuMDR: Code-mixed Multi-modal Multi-domain corpus for Discourse paRsing in conversations
by: Shukla, Divyaksh, et al.
Published: (2025)
by: Shukla, Divyaksh, et al.
Published: (2025)
Generation and De-Identification of Indian Clinical Discharge Summaries using LLMs
by: Singh, Sanjeet, et al.
Published: (2024)
by: Singh, Sanjeet, et al.
Published: (2024)
IL-PCSR: Legal Corpus for Prior Case and Statute Retrieval
by: Paul, Shounak, et al.
Published: (2025)
by: Paul, Shounak, et al.
Published: (2025)
Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation
by: Zhang, Wenbo, et al.
Published: (2025)
by: Zhang, Wenbo, et al.
Published: (2025)
Transformer Circuit Faithfulness Metrics are not Robust
by: Miller, Joseph, et al.
Published: (2024)
by: Miller, Joseph, et al.
Published: (2024)
Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs
by: Pepper, Keenan, et al.
Published: (2026)
by: Pepper, Keenan, et al.
Published: (2026)
Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression
by: Xiao, Hanqi, et al.
Published: (2025)
by: Xiao, Hanqi, et al.
Published: (2025)
AMUSED: A Multi-Stream Vector Representation Method for Use in Natural Dialogue
by: Kumar, Gaurav, et al.
Published: (2019)
by: Kumar, Gaurav, et al.
Published: (2019)
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models
by: Lv, Ang, et al.
Published: (2024)
by: Lv, Ang, et al.
Published: (2024)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
by: Kim, Geonhee, et al.
Published: (2024)
by: Kim, Geonhee, et al.
Published: (2024)
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
by: Marks, Samuel, et al.
Published: (2024)
by: Marks, Samuel, et al.
Published: (2024)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
by: Lan, Michael, et al.
Published: (2023)
by: Lan, Michael, et al.
Published: (2023)
Beyond Transfer Accuracy: Faithful Circuits for Controlled Low-Resource Adaptation
by: Nur'aini, Khumaisa, et al.
Published: (2026)
by: Nur'aini, Khumaisa, et al.
Published: (2026)
Does Transformer Interpretability Transfer to RNNs?
by: Paulo, Gonçalo, et al.
Published: (2024)
by: Paulo, Gonçalo, et al.
Published: (2024)
Efficient Automated Circuit Discovery in Transformers using Contextual Decomposition
by: Hsu, Aliyah R., et al.
Published: (2024)
by: Hsu, Aliyah R., et al.
Published: (2024)
Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian Distribution
by: Zhao, Haiyan, et al.
Published: (2024)
by: Zhao, Haiyan, et al.
Published: (2024)
HLDC: Hindi Legal Documents Corpus
by: Kapoor, Arnav, et al.
Published: (2022)
by: Kapoor, Arnav, et al.
Published: (2022)
Discovering Interpretable Algorithms by Decompiling Transformers to RASP
by: Huang, Xinting, et al.
Published: (2026)
by: Huang, Xinting, et al.
Published: (2026)
Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers
by: Adhikari, Rabin
Published: (2025)
by: Adhikari, Rabin
Published: (2025)
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
by: Su, Jingtong, et al.
Published: (2025)
by: Su, Jingtong, et al.
Published: (2025)
Dimensionality Reduction in Sentence Transformer Vector Databases with Fast Fourier Transform
by: Bulgakov, Vitaly, et al.
Published: (2024)
by: Bulgakov, Vitaly, et al.
Published: (2024)
Transformers as Support Vector Machines
by: Tarzanagh, Davoud Ataee, et al.
Published: (2023)
by: Tarzanagh, Davoud Ataee, et al.
Published: (2023)
The Dual-Stream Transformer: Channelized Architecture for Interpretable Language Modeling
by: Kerce, J. Clayton, et al.
Published: (2026)
by: Kerce, J. Clayton, et al.
Published: (2026)
Similar Items
-
Calibration Across Layers: Understanding Calibration Evolution in LLMs
by: Joshi, Abhinav, et al.
Published: (2025) -
COLD: Causal reasOning in cLosed Daily activities
by: Joshi, Abhinav, et al.
Published: (2024) -
Towards Quantifying Commonsense Reasoning with Mechanistic Insights
by: Joshi, Abhinav, et al.
Published: (2025) -
Geometry of Decision Making in Language Models
by: Joshi, Abhinav, et al.
Published: (2025) -
POSESTITCH-SLT: Linguistically Inspired Pose-Stitching for End-to-End Sign Language Translation
by: Joshi, Abhinav, et al.
Published: (2025)