Interpretability as Alignment: Making Internal Understanding a Design Principle
Fuente:
arXiv
Saved in:
| Main Authors: | Sengupta, Aadit, Seth, Pratinav, Sankarapu, Vinay Kumar |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging the Gap in XAI-Why Reliable Metrics Matter for Explainability and Compliance
by: Seth, Pratinav, et al.
Published: (2025)
by: Seth, Pratinav, et al.
Published: (2025)
xai_evals : A Framework for Evaluating Post-Hoc Local Explanation Methods
by: Seth, Pratinav, et al.
Published: (2025)
by: Seth, Pratinav, et al.
Published: (2025)
Interpretability-Aware Pruning for Efficient Medical Image Analysis
by: Malik, Nikita, et al.
Published: (2025)
by: Malik, Nikita, et al.
Published: (2025)
Forgetting That Sticks: Quantization-Permanent Unlearning via Circuit Attribution
by: Sadhu, Saisab, et al.
Published: (2026)
by: Sadhu, Saisab, et al.
Published: (2026)
$C$-$ΔΘ$: Circuit-Restricted Weight Arithmetic for Selective Refusal
by: Kasliwal, Aditya, et al.
Published: (2026)
by: Kasliwal, Aditya, et al.
Published: (2026)
Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands
by: Seth, Pratinav, et al.
Published: (2026)
by: Seth, Pratinav, et al.
Published: (2026)
Enhancing short-term traffic prediction by integrating trends and fluctuations with attention mechanism
by: Das, Adway, et al.
Published: (2025)
by: Das, Adway, et al.
Published: (2025)
Orion-Bix: Bi-Axial Attention for Tabular In-Context Learning
by: Bouadi, Mohamed, et al.
Published: (2025)
by: Bouadi, Mohamed, et al.
Published: (2025)
Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning
by: Bouadi, Mohamed, et al.
Published: (2025)
by: Bouadi, Mohamed, et al.
Published: (2025)
Trust-Based Incentive Mechanisms in Semi-Decentralized Federated Learning Systems
by: Shrestha, Ajay Kumar
Published: (2026)
by: Shrestha, Ajay Kumar
Published: (2026)
NEURODNAAI: Neural pipeline approaches for the advancing dna-based information storage as a sustainable digital medium using deep learning framework
by: Thakur, Rakesh, et al.
Published: (2025)
by: Thakur, Rakesh, et al.
Published: (2025)
XBTorch: A Unified Framework for Modeling and Co-Design of Crossbar-Based Deep Learning Accelerators
by: Yousuf, Osama, et al.
Published: (2026)
by: Yousuf, Osama, et al.
Published: (2026)
MenTeR: A fully-automated Multi-agenT workflow for end-to-end RF/Analog Circuits Netlist Design
by: Chen, Pin-Han, et al.
Published: (2025)
by: Chen, Pin-Han, et al.
Published: (2025)
Exploring the Decentraland Economy: Multifaceted Parcel Attributes, Key Insights, and Benchmarking
by: Jha, Dipika, et al.
Published: (2024)
by: Jha, Dipika, et al.
Published: (2024)
Self Distillation via Iterative Constructive Perturbations
by: Dave, Maheak, et al.
Published: (2025)
by: Dave, Maheak, et al.
Published: (2025)
TabTune: A Unified Library for Inference and Fine-Tuning Tabular Foundation Models
by: Tanna, Aditya, et al.
Published: (2025)
by: Tanna, Aditya, et al.
Published: (2025)
Distilling Tabular Foundation Models for Structured Health Data
by: Tanna, Aditya, et al.
Published: (2026)
by: Tanna, Aditya, et al.
Published: (2026)
CMOS + stochastic nanomagnets: heterogeneous computers for probabilistic inference and learning
by: Singh, Nihal Sanjay, et al.
Published: (2023)
by: Singh, Nihal Sanjay, et al.
Published: (2023)
RMAAT: Astrocyte-Inspired Memory Compression and Replay for Efficient Long-Context Transformers
by: Mia, Md Zesun Ahmed, et al.
Published: (2026)
by: Mia, Md Zesun Ahmed, et al.
Published: (2026)
Delving Deeper Into Astromorphic Transformers
by: Mia, Md Zesun Ahmed, et al.
Published: (2023)
by: Mia, Md Zesun Ahmed, et al.
Published: (2023)
Hardware-Adaptive and Superlinear-Capacity Memristor-based Associative Memory
by: He, Chengping, et al.
Published: (2025)
by: He, Chengping, et al.
Published: (2025)
Artificial Intelligence for Personalized Prediction of Alzheimer's Disease Progression: A Survey of Methods, Data Challenges, and Future Directions
by: Koksalmis, Gulsah Hancerliogullari, et al.
Published: (2025)
by: Koksalmis, Gulsah Hancerliogullari, et al.
Published: (2025)
LLMs meet Federated Learning for Scalable and Secure IoT Management
by: Otoum, Yazan, et al.
Published: (2025)
by: Otoum, Yazan, et al.
Published: (2025)
LeForecast: Enterprise Hybrid Forecast by Time Series Intelligence
by: Tan, Zheng, et al.
Published: (2025)
by: Tan, Zheng, et al.
Published: (2025)
M2RU: Memristive Minion Recurrent Unit for On-Chip Continual Learning at the Edge
by: Zyarah, Abdullah M., et al.
Published: (2025)
by: Zyarah, Abdullah M., et al.
Published: (2025)
LPCVAE: A Conditional VAE with Long-Term Dependency and Probabilistic Time-Frequency Fusion for Time Series Anomaly Detection
by: Cheng, Hanchang, et al.
Published: (2025)
by: Cheng, Hanchang, et al.
Published: (2025)
Keeping Medical AI Healthy and Trustworthy: A Review of Detection and Correction Methods for System Degradation
by: Guan, Hao, et al.
Published: (2025)
by: Guan, Hao, et al.
Published: (2025)
SurvUnc: A Meta-Model Based Uncertainty Quantification Framework for Survival Analysis
by: Liu, Yu, et al.
Published: (2025)
by: Liu, Yu, et al.
Published: (2025)
Coordinating Ride-Pooling with Public Transit using Reward-Guided Conservative Q-Learning: An Offline Training and Online Fine-Tuning Reinforcement Learning Framework
by: Hu, Yulong, et al.
Published: (2025)
by: Hu, Yulong, et al.
Published: (2025)
Accurate AI-Driven Emergency Vehicle Location Tracking in Healthcare ITS Digital Twin
by: Al-Shareeda, Sarah, et al.
Published: (2025)
by: Al-Shareeda, Sarah, et al.
Published: (2025)
Evaluating the Energy Efficiency of NPU-Accelerated Machine Learning Inference on Embedded Microcontrollers
by: Fanariotis, Anastasios, et al.
Published: (2025)
by: Fanariotis, Anastasios, et al.
Published: (2025)
Cyber Physical Awareness via Intent-Driven Threat Assessment: Enhanced Space Networks with Intershell Links
by: Cetin, Selen Gecgel, et al.
Published: (2025)
by: Cetin, Selen Gecgel, et al.
Published: (2025)
Fully analogue in-memory neural computing via quantum tunneling effect
by: Li, Songyuan, et al.
Published: (2025)
by: Li, Songyuan, et al.
Published: (2025)
Learning with Calibration: Exploring Test-Time Computing of Spatio-Temporal Forecasting
by: Chen, Wei, et al.
Published: (2025)
by: Chen, Wei, et al.
Published: (2025)
Open-Source LLM-Driven Federated Transformer for Predictive IoV Management
by: Otoum, Yazan, et al.
Published: (2025)
by: Otoum, Yazan, et al.
Published: (2025)
Polynomial Neural Sheaf Diffusion: A Spectral Filtering Approach on Cellular Sheaves
by: Borgi, Alessio, et al.
Published: (2025)
by: Borgi, Alessio, et al.
Published: (2025)
Bringing Multi-Modal Multi-Task Federated Foundation Models to Education Domain: Prospects and Challenges
by: Borazjani, Kasra, et al.
Published: (2025)
by: Borazjani, Kasra, et al.
Published: (2025)
CORPGEN: Simulating Corporate Environments with Autonomous Digital Employees in Multi-Horizon Task Environments
by: Jaye, Abubakarr, et al.
Published: (2026)
by: Jaye, Abubakarr, et al.
Published: (2026)
Concurrent Self-testing of Neural Networks Using Uncertainty Fingerprint
by: Ahmed, Soyed Tuhin, et al.
Published: (2024)
by: Ahmed, Soyed Tuhin, et al.
Published: (2024)
Intrinsic Voltage Offsets in Memcapacitive Bio-Membranes Enable High-Performance Physical Reservoir Computing
by: Mohamed, Ahmed S., et al.
Published: (2024)
by: Mohamed, Ahmed S., et al.
Published: (2024)
Similar Items
-
Bridging the Gap in XAI-Why Reliable Metrics Matter for Explainability and Compliance
by: Seth, Pratinav, et al.
Published: (2025) -
xai_evals : A Framework for Evaluating Post-Hoc Local Explanation Methods
by: Seth, Pratinav, et al.
Published: (2025) -
Interpretability-Aware Pruning for Efficient Medical Image Analysis
by: Malik, Nikita, et al.
Published: (2025) -
Forgetting That Sticks: Quantization-Permanent Unlearning via Circuit Attribution
by: Sadhu, Saisab, et al.
Published: (2026) -
$C$-$ΔΘ$: Circuit-Restricted Weight Arithmetic for Selective Refusal
by: Kasliwal, Aditya, et al.
Published: (2026)