A Multi-Level Causal Intervention Framework for Mechanistic Interpretability in Variational Autoencoders
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Roy, Dip, Misra, Rajiv, Singh, Sanjay Kumar, Roy, Anisha |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Posterior-Calibrated Causal Circuits in Variational Autoencoders: Why Image-Domain Interpretability Fails on Tabular Data
von: Roy, Dip, et al.
Veröffentlicht: (2026)
von: Roy, Dip, et al.
Veröffentlicht: (2026)
Detection Without Correction: A Robust Asymmetry in Activation-Based Hallucination Probing
von: Roy, Dip, et al.
Veröffentlicht: (2026)
von: Roy, Dip, et al.
Veröffentlicht: (2026)
Fundamental Limits of Neural Network Sparsification: Evidence from Catastrophic Interpretability Collapse
von: Roy, Dip, et al.
Veröffentlicht: (2026)
von: Roy, Dip, et al.
Veröffentlicht: (2026)
MemGuard-Alpha: Detecting and Filtering Memorization-Contaminated Signals in LLM-Based Financial Forecasting via Membership Inference and Cross-Model Disagreement
von: Roy, Anisha, et al.
Veröffentlicht: (2026)
von: Roy, Anisha, et al.
Veröffentlicht: (2026)
Mechanistic Interpretability of Diffusion Models: Circuit-Level Analysis and Causal Validation
von: Roy, Dip
Veröffentlicht: (2025)
von: Roy, Dip
Veröffentlicht: (2025)
Bayesian Autoencoder for Medical Anomaly Detection: Uncertainty-Aware Approach for Brain 2 MRI Analysis
von: Roy, Dip
Veröffentlicht: (2025)
von: Roy, Dip
Veröffentlicht: (2025)
ARDDQN: Attention Recurrent Double Deep Q-Network for UAV Coverage Path Planning and Data Harvesting
von: Kumar, Praveen, et al.
Veröffentlicht: (2024)
von: Kumar, Praveen, et al.
Veröffentlicht: (2024)
Group Equivariance Meets Mechanistic Interpretability: Equivariant Sparse Autoencoders
von: Erdogan, Ege, et al.
Veröffentlicht: (2025)
von: Erdogan, Ege, et al.
Veröffentlicht: (2025)
A Delta-Aware Orchestration Framework for Scalable Multi-Agent Edge Computing
von: Singh, Samaresh Kumar, et al.
Veröffentlicht: (2026)
von: Singh, Samaresh Kumar, et al.
Veröffentlicht: (2026)
Mechanistic Interpretability with Sparse Autoencoder Neural Operators
von: Tolooshams, Bahareh, et al.
Veröffentlicht: (2025)
von: Tolooshams, Bahareh, et al.
Veröffentlicht: (2025)
A Latent Space Correlation-Aware Autoencoder for Anomaly Detection in Skewed Data
von: Roy, Padmaksha
Veröffentlicht: (2023)
von: Roy, Padmaksha
Veröffentlicht: (2023)
Innovative Framework for Early Estimation of Mental Disorder Scores to Enable Timely Interventions
von: Singh, Himanshi, et al.
Veröffentlicht: (2025)
von: Singh, Himanshi, et al.
Veröffentlicht: (2025)
Resurrecting the Salmon: Rethinking Mechanistic Interpretability with Domain-Specific Sparse Autoencoders
von: O'Neill, Charles, et al.
Veröffentlicht: (2025)
von: O'Neill, Charles, et al.
Veröffentlicht: (2025)
Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models
von: Zhou, Hanhan, et al.
Veröffentlicht: (2026)
von: Zhou, Hanhan, et al.
Veröffentlicht: (2026)
Quantum Generative Adversarial Autoencoders: Learning latent representations for quantum data generation
von: Raj, Naipunnya, et al.
Veröffentlicht: (2025)
von: Raj, Naipunnya, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
von: Tahimic, Kriz, et al.
Veröffentlicht: (2025)
von: Tahimic, Kriz, et al.
Veröffentlicht: (2025)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
von: Cho, Hakaze, et al.
Veröffentlicht: (2025)
von: Cho, Hakaze, et al.
Veröffentlicht: (2025)
Causal Dynamic Variational Autoencoder for Counterfactual Regression in Longitudinal Data
von: Bouchattaoui, Mouad El, et al.
Veröffentlicht: (2023)
von: Bouchattaoui, Mouad El, et al.
Veröffentlicht: (2023)
Step-Level Sparse Autoencoder for Reasoning Process Interpretation
von: Yang, Xuan, et al.
Veröffentlicht: (2026)
von: Yang, Xuan, et al.
Veröffentlicht: (2026)
The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
von: Sutter, Denis, et al.
Veröffentlicht: (2025)
von: Sutter, Denis, et al.
Veröffentlicht: (2025)
Variational Autoencoder for Calibration: A New Approach
von: Barrett, Travis, et al.
Veröffentlicht: (2025)
von: Barrett, Travis, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders
von: Lehn-Schiøler, William, et al.
Veröffentlicht: (2026)
von: Lehn-Schiøler, William, et al.
Veröffentlicht: (2026)
HeartBeatAI: An Interpretable and Robust Deep Learning Framework for Multi-Label ECG Arrhythmia Detection
von: Gupta, Shubham, et al.
Veröffentlicht: (2026)
von: Gupta, Shubham, et al.
Veröffentlicht: (2026)
Linking Model Intervention to Causal Interpretation in Model Explanation
von: Cheng, Debo, et al.
Veröffentlicht: (2024)
von: Cheng, Debo, et al.
Veröffentlicht: (2024)
Competition is the key: A Game Theoretic Causal Discovery Approach
von: Roy, Amartya, et al.
Veröffentlicht: (2025)
von: Roy, Amartya, et al.
Veröffentlicht: (2025)
Downstream Task Guided Masking Learning in Masked Autoencoders Using Multi-Level Optimization
von: Guo, Han, et al.
Veröffentlicht: (2024)
von: Guo, Han, et al.
Veröffentlicht: (2024)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
von: Wang, Xu, et al.
Veröffentlicht: (2026)
von: Wang, Xu, et al.
Veröffentlicht: (2026)
Disentanglement of Sources in a Multi-Stream Variational Autoencoder
von: Boukun, Veranika, et al.
Veröffentlicht: (2025)
von: Boukun, Veranika, et al.
Veröffentlicht: (2025)
Decentralized Multi-Agent Swarms for Autonomous Grid Security in Industrial IoT: A Consensus-based Approach
von: Singh, Samaresh Kumar, et al.
Veröffentlicht: (2026)
von: Singh, Samaresh Kumar, et al.
Veröffentlicht: (2026)
Correlating Variational Autoencoders Natively For Multi-View Imputation
von: Orme, Ella S. C., et al.
Veröffentlicht: (2024)
von: Orme, Ella S. C., et al.
Veröffentlicht: (2024)
Adaptive Interface-PINNs (AdaI-PINNs): An Efficient Physics-informed Neural Networks Framework for Interface Problems
von: Roy, Sumanta, et al.
Veröffentlicht: (2024)
von: Roy, Sumanta, et al.
Veröffentlicht: (2024)
Learning Mixtures of Unknown Causal Interventions
von: Kumar, Abhinav, et al.
Veröffentlicht: (2024)
von: Kumar, Abhinav, et al.
Veröffentlicht: (2024)
Variational Autoencoders for Efficient Simulation-Based Inference
von: Nautiyal, Mayank, et al.
Veröffentlicht: (2024)
von: Nautiyal, Mayank, et al.
Veröffentlicht: (2024)
DEM: A Distilled Explanation Model for Interpretable Anomaly Detection in Physiological Sensor Networks
von: Singh, Jyotirmoy, et al.
Veröffentlicht: (2026)
von: Singh, Jyotirmoy, et al.
Veröffentlicht: (2026)
A Multi-directional Meta-Learning Framework for Class-Generalizable Anomaly Detection
von: Roy, Padmaksha, et al.
Veröffentlicht: (2026)
von: Roy, Padmaksha, et al.
Veröffentlicht: (2026)
ACTIVA: Amortized Causal Effect Estimation via Transformer-based Variational Autoencoder
von: Sauter, Andreas, et al.
Veröffentlicht: (2025)
von: Sauter, Andreas, et al.
Veröffentlicht: (2025)
Robust Topology Optimization Using Multi-Fidelity Variational Autoencoders
von: Gladstone, Rini Jasmine, et al.
Veröffentlicht: (2021)
von: Gladstone, Rini Jasmine, et al.
Veröffentlicht: (2021)
Graph Variate Neural Networks
von: Roy, Om, et al.
Veröffentlicht: (2025)
von: Roy, Om, et al.
Veröffentlicht: (2025)
Open Problems in Mechanistic Interpretability
von: Sharkey, Lee, et al.
Veröffentlicht: (2025)
von: Sharkey, Lee, et al.
Veröffentlicht: (2025)
Exemplar Partitioning for Mechanistic Interpretability
von: Rumbelow, Jessica
Veröffentlicht: (2026)
von: Rumbelow, Jessica
Veröffentlicht: (2026)
Ähnliche Einträge
-
Posterior-Calibrated Causal Circuits in Variational Autoencoders: Why Image-Domain Interpretability Fails on Tabular Data
von: Roy, Dip, et al.
Veröffentlicht: (2026) -
Detection Without Correction: A Robust Asymmetry in Activation-Based Hallucination Probing
von: Roy, Dip, et al.
Veröffentlicht: (2026) -
Fundamental Limits of Neural Network Sparsification: Evidence from Catastrophic Interpretability Collapse
von: Roy, Dip, et al.
Veröffentlicht: (2026) -
MemGuard-Alpha: Detecting and Filtering Memorization-Contaminated Signals in LLM-Based Financial Forecasting via Membership Inference and Cross-Model Disagreement
von: Roy, Anisha, et al.
Veröffentlicht: (2026) -
Mechanistic Interpretability of Diffusion Models: Circuit-Level Analysis and Causal Validation
von: Roy, Dip
Veröffentlicht: (2025)