Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs
Fuente:
arXiv
Saved in:
| Main Authors: | Pepper, Keenan, McKenzie, Alex, Pop, Florin, Servaes, Stijn, Leitgab, Martin, Vaiana, Mike, Rosenblatt, Judd, Graziano, Michael S. A., de Lucena, Diogo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Endogenous Resistance to Activation Steering in Language Models
by: McKenzie, Alex, et al.
Published: (2026)
by: McKenzie, Alex, et al.
Published: (2026)
Rethinking harmless refusals when fine-tuning foundation models
by: Pop, Florin, et al.
Published: (2024)
by: Pop, Florin, et al.
Published: (2024)
Unexpected Benefits of Self-Modeling in Neural Systems
by: Premakumar, Vickram N., et al.
Published: (2024)
by: Premakumar, Vickram N., et al.
Published: (2024)
Towards Safe and Honest AI Agents with Neural Self-Other Overlap
by: Carauleanu, Marc, et al.
Published: (2024)
by: Carauleanu, Marc, et al.
Published: (2024)
Large Language Models Report Subjective Experience Under Self-Referential Processing
by: Berg, Cameron, et al.
Published: (2025)
by: Berg, Cameron, et al.
Published: (2025)
Momentum Point-Perplexity Mechanics in Large Language Models
by: Tomaz, Lorenzo, et al.
Published: (2025)
by: Tomaz, Lorenzo, et al.
Published: (2025)
Self-Ablating Transformers: More Interpretability, Less Sparsity
by: Ferrao, Jeremias, et al.
Published: (2025)
by: Ferrao, Jeremias, et al.
Published: (2025)
Prototype-Guided and Lightweight Adapters for Inherent Interpretation and Generalisation in Federated Learning
by: Mensah, Samuel Ofosu, et al.
Published: (2025)
by: Mensah, Samuel Ofosu, et al.
Published: (2025)
Improved Lattice QCD $B_c\to J/ψ$ Vector, Axial-Vector, and Tensor Form Factors
by: Harrison, Judd
Published: (2025)
by: Harrison, Judd
Published: (2025)
Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
by: Minder, Julian, et al.
Published: (2025)
by: Minder, Julian, et al.
Published: (2025)
TDHook: A Lightweight Framework for Interpretability
by: Poupart, Yoann
Published: (2025)
by: Poupart, Yoann
Published: (2025)
A Lightweight and Interpretable Deepfakes Detection Framework
by: Farooq, Muhammad Umar, et al.
Published: (2025)
by: Farooq, Muhammad Umar, et al.
Published: (2025)
Analyzing Inter‐Hemispheric Climate Change Asymmetries With a Cointegrated Vector Autoregression
by: Graziano Moramarco
Published: (2025)
by: Graziano Moramarco
Published: (2025)
Air Traffic Controller Task Demand via Graph Neural Networks: An Interpretable Approach to Airspace Complexity
by: Henderson, Edward, et al.
Published: (2025)
by: Henderson, Edward, et al.
Published: (2025)
Explainability-Driven Leaf Disease Classification Using Adversarial Training and Knowledge Distillation
by: Echim, Sebastian-Vasile, et al.
Published: (2023)
by: Echim, Sebastian-Vasile, et al.
Published: (2023)
INSIGHT: An Interpretable Neural Vision-Language Framework for Reasoning of Generative Artifacts
by: Bagaria, Anshul
Published: (2025)
by: Bagaria, Anshul
Published: (2025)
Fast and Interpretable Autoregressive Estimation with Neural Network Backpropagation
by: Lucena, Anaísa, et al.
Published: (2026)
by: Lucena, Anaísa, et al.
Published: (2026)
Gram2Vec: An Interpretable Document Vectorizer
by: Zeng, Peter, et al.
Published: (2024)
by: Zeng, Peter, et al.
Published: (2024)
Interpretable Prediction and Feature Selection for Survival Analysis
by: Van Ness, Mike, et al.
Published: (2024)
by: Van Ness, Mike, et al.
Published: (2024)
Improve Two‐year Interpreter Training Programs for Educational Interpreters
by: Sara Simpson, et al.
Published: (2024)
by: Sara Simpson, et al.
Published: (2024)
Automated Motion Artifact Check for MRI (AutoMAC-MRI): An Interpretable Framework for Motion Artifact Detection and Severity Assessment
by: Jerald, Antony, et al.
Published: (2025)
by: Jerald, Antony, et al.
Published: (2025)
Partially-elementary end extensions of countable models of set theory
by: McKenzie, Zachiri
Published: (2024)
by: McKenzie, Zachiri
Published: (2024)
IT Students Career Confidence and Career Identity During COVID-19
by: McKenzie, Sophie
Published: (2025)
by: McKenzie, Sophie
Published: (2025)
The set-theoretic Kaufmann-Clote question
by: McKenzie, Zachiri
Published: (2025)
by: McKenzie, Zachiri
Published: (2025)
Migration, remittances, poverty, and human capital : conceptual and empirical challenges / David McKenzie, Marcin J. Sasin
by: McKenzie, David
Published: (2007)
by: McKenzie, David
Published: (2007)
Self-selection patterns in Mexico-U.S. migration : the role of migration networks / David McKenzie, Hillel Rapoport
by: McKenzie, David
Published: (2007)
by: McKenzie, David
Published: (2007)
A land of milk and honey with streets paved with gold : do emigrants have over-optimistic expectations about incomes abroad? / David McKenzie, John Gibson, Steven Stillman
by: McKenzie, David
Published: (2007)
by: McKenzie, David
Published: (2007)
How important is selection? : experimental versus non-experimental measures of the income gains from migration / David McKenzie, John Gibson, Steven Stillman
by: McKenzie, David
Published: (2006)
by: McKenzie, David
Published: (2006)
Can migration reduce educational attainment? : evidence from Mexico / David McKenzie, Hillel Rapoport
by: McKenzie, David
Published: (2006)
by: McKenzie, David
Published: (2006)
Libraries of the Future.
by: McKenzie, Jamie
Published: (1996)
by: McKenzie, Jamie
Published: (1996)
A Window on to the World: Newspaper Collecting in the National Library of Australia.
by: McKenzie, Amelia
Published: (1999)
by: McKenzie, Amelia
Published: (1999)
Artifact for paper: Abstract Interpretation of Temporal Safety Effects of Higher Order Programs
by: Nicola, Mihai, et al.
Published: (2025)
by: Nicola, Mihai, et al.
Published: (2025)
SPIRE: Structure-Preserving Interpretable Retrieval of Evidence
by: Rainey, Mike, et al.
Published: (2026)
by: Rainey, Mike, et al.
Published: (2026)
Label Forensics: Interpreting Hard Labels in Black-Box Text Classifier
by: Du, Mengyao, et al.
Published: (2025)
by: Du, Mengyao, et al.
Published: (2025)
Inv-Adapter: ID Customization Generation via Image Inversion and Lightweight Adapter
by: Xing, Peng, et al.
Published: (2024)
by: Xing, Peng, et al.
Published: (2024)
T-Gated Adapter: A Lightweight Temporal Adapter for Vision-Language Medical Segmentation
by: Khadka, Pranjal
Published: (2026)
by: Khadka, Pranjal
Published: (2026)
ArtifactLens: Hundreds of Labels Are Enough for Artifact Detection with VLMs
by: Burgess, James, et al.
Published: (2026)
by: Burgess, James, et al.
Published: (2026)
Interpretable Syntactic Representations Enable Hierarchical Word Vectors
by: Silwal, Biraj
Published: (2024)
by: Silwal, Biraj
Published: (2024)
VRAIL: Vectorized Reward-based Attribution for Interpretable Learning
by: Kim, Jina, et al.
Published: (2025)
by: Kim, Jina, et al.
Published: (2025)
A Lightweight Generative Model for Interpretable Subject-level Prediction
by: Mauri, Chiara, et al.
Published: (2023)
by: Mauri, Chiara, et al.
Published: (2023)
Similar Items
-
Endogenous Resistance to Activation Steering in Language Models
by: McKenzie, Alex, et al.
Published: (2026) -
Rethinking harmless refusals when fine-tuning foundation models
by: Pop, Florin, et al.
Published: (2024) -
Unexpected Benefits of Self-Modeling in Neural Systems
by: Premakumar, Vickram N., et al.
Published: (2024) -
Towards Safe and Honest AI Agents with Neural Self-Other Overlap
by: Carauleanu, Marc, et al.
Published: (2024) -
Large Language Models Report Subjective Experience Under Self-Referential Processing
by: Berg, Cameron, et al.
Published: (2025)