Bilinear Convolution Decomposition for Causal RL Interpretability
Fuente:
arXiv
Saved in:
| Main Authors: | Oozeer, Narmeen, Erisken, Sinem, Rigg, Alice |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Weight-based Decomposition: A Case for Bilinear MLPs
by: Pearce, Michael T., et al.
Published: (2024)
by: Pearce, Michael T., et al.
Published: (2024)
Distribution-Aware Feature Selection for SAEs
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
DreamReader: An Interpretability Toolkit for Text-to-Image Models
by: Prakash, Nirmalendu, et al.
Published: (2026)
by: Prakash, Nirmalendu, et al.
Published: (2026)
Spectral Superposition: A Theory of Feature Geometry
by: Ivanov, Georgi, et al.
Published: (2026)
by: Ivanov, Georgi, et al.
Published: (2026)
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
by: Roytburg, Dani, et al.
Published: (2025)
by: Roytburg, Dani, et al.
Published: (2025)
MAEBE: Multi-Agent Emergent Behavior Framework
by: Erisken, Sinem, et al.
Published: (2025)
by: Erisken, Sinem, et al.
Published: (2025)
Understanding and Mitigating Dataset Corruption in LLM Steering
by: Anderson, Cullen, et al.
Published: (2026)
by: Anderson, Cullen, et al.
Published: (2026)
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
by: Roytburg, Dani, et al.
Published: (2026)
by: Roytburg, Dani, et al.
Published: (2026)
Approximating Human Preferences Using a Multi-Judge Learned System
by: Sprejer, Eitán, et al.
Published: (2025)
by: Sprejer, Eitán, et al.
Published: (2025)
BECAUSE: Bilinear Causal Representation for Generalizable Offline Model-based Reinforcement Learning
by: Lin, Haohong, et al.
Published: (2024)
by: Lin, Haohong, et al.
Published: (2024)
Position: Require Frontier AI Labs To Release Small "Analog" Models
by: Upadhyay, Shriyash, et al.
Published: (2025)
by: Upadhyay, Shriyash, et al.
Published: (2025)
Coefficient Decomposition for Spectral Graph Convolution
by: Huang, Feng, et al.
Published: (2024)
by: Huang, Feng, et al.
Published: (2024)
Correlating Time Series with Interpretable Convolutional Kernels
by: Chen, Xinyu, et al.
Published: (2024)
by: Chen, Xinyu, et al.
Published: (2024)
Bilinear MLPs enable weight-based mechanistic interpretability
by: Pearce, Michael T., et al.
Published: (2024)
by: Pearce, Michael T., et al.
Published: (2024)
CRITS: Convolutional Rectifier for Interpretable Time Series Classification
by: Kuratomi, Alejandro, et al.
Published: (2025)
by: Kuratomi, Alejandro, et al.
Published: (2025)
Toward Temporal Causal Representation Learning with Tensor Decomposition
by: Chen, Jianhong, et al.
Published: (2025)
by: Chen, Jianhong, et al.
Published: (2025)
Towards Empowerment Gain through Causal Structure Learning in Model-Based RL
by: Cao, Hongye, et al.
Published: (2025)
by: Cao, Hongye, et al.
Published: (2025)
DecoKAN: Interpretable Decomposition for Forecasting Cryptocurrency Market Dynamics
by: Gao, Yuan, et al.
Published: (2025)
by: Gao, Yuan, et al.
Published: (2025)
Beyond Johnson-Lindenstrauss: Uniform Bounds for Sketched Bilinear Forms
by: Deb, Rohan, et al.
Published: (2025)
by: Deb, Rohan, et al.
Published: (2025)
AICRN: Attention-Integrated Convolutional Residual Network for Interpretable Electrocardiogram Analysis
by: Jayakody, J. M. I. H., et al.
Published: (2025)
by: Jayakody, J. M. I. H., et al.
Published: (2025)
Causality-Aware Local Interpretable Model-Agnostic Explanations
by: Cinquini, Martina, et al.
Published: (2022)
by: Cinquini, Martina, et al.
Published: (2022)
Out-of-Distribution Adaptation in Offline RL: Counterfactual Reasoning via Causal Normalizing Flows
by: Cho, Minjae, et al.
Published: (2024)
by: Cho, Minjae, et al.
Published: (2024)
EDformer: Embedded Decomposition Transformer for Interpretable Multivariate Time Series Predictions
by: Chakraborty, Sanjay, et al.
Published: (2024)
by: Chakraborty, Sanjay, et al.
Published: (2024)
Network-Aware Bilinear Tokenization for Brain Functional Connectivity Representation Learning
by: Milecki, Leo, et al.
Published: (2026)
by: Milecki, Leo, et al.
Published: (2026)
Bilinear representation mitigates reversal curse and enables consistent model editing
by: Kim, Dong-Kyum, et al.
Published: (2025)
by: Kim, Dong-Kyum, et al.
Published: (2025)
Interpretable Diffusion via Information Decomposition
by: Kong, Xianghao, et al.
Published: (2023)
by: Kong, Xianghao, et al.
Published: (2023)
LFA applied to CNNs: Efficient Singular Value Decomposition of Convolutional Mappings by Local Fourier Analysis
by: van Betteray, Antonia, et al.
Published: (2025)
by: van Betteray, Antonia, et al.
Published: (2025)
DRExplainer: Quantifiable Interpretability in Drug Response Prediction with Directed Graph Convolutional Network
by: Shi, Haoyuan, et al.
Published: (2024)
by: Shi, Haoyuan, et al.
Published: (2024)
A Recursive Decomposition Framework for Causal Structure Learning in the Presence of Latent Variables
by: Li, Zheng, et al.
Published: (2026)
by: Li, Zheng, et al.
Published: (2026)
SphUnc: Hyperspherical Uncertainty Decomposition and Causal Identification via Information Geometry
by: Fu, Rong, et al.
Published: (2026)
by: Fu, Rong, et al.
Published: (2026)
Disentangling Causal Substructures for Interpretable and Generalizable Drug Synergy Prediction
by: Luo, Yi, et al.
Published: (2025)
by: Luo, Yi, et al.
Published: (2025)
Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders
by: Oldfield, James, et al.
Published: (2025)
by: Oldfield, James, et al.
Published: (2025)
Preserving Bilinear Weight Spectra with a Signed and Shrunk Quadratic Activation Function
by: Abohwo, Jason, et al.
Published: (2025)
by: Abohwo, Jason, et al.
Published: (2025)
Causal Rule Forest: Toward Interpretable and Precise Treatment Effect Estimation
by: Hsu, Chan, et al.
Published: (2024)
by: Hsu, Chan, et al.
Published: (2024)
Structured Temporal Causality for Interpretable Multivariate Time Series Anomaly Detection
by: Cho, Dongchan, et al.
Published: (2025)
by: Cho, Dongchan, et al.
Published: (2025)
CARL: Causality-guided Architecture Representation Learning for an Interpretable Performance Predictor
by: Ji, Han, et al.
Published: (2025)
by: Ji, Han, et al.
Published: (2025)
A Unified Frequency Domain Decomposition Framework for Interpretable and Robust Time Series Forecasting
by: He, Cheng, et al.
Published: (2025)
by: He, Cheng, et al.
Published: (2025)
Causally-Aware Spatio-Temporal Multi-Graph Convolution Network for Accurate and Reliable Traffic Prediction
by: Dong, Pingping, et al.
Published: (2024)
by: Dong, Pingping, et al.
Published: (2024)
MIBP-Cert: Certified Training against Data Perturbations with Mixed-Integer Bilinear Programs
by: Lorenz, Tobias, et al.
Published: (2024)
by: Lorenz, Tobias, et al.
Published: (2024)
Similar Items
-
Weight-based Decomposition: A Case for Bilinear MLPs
by: Pearce, Michael T., et al.
Published: (2024) -
Distribution-Aware Feature Selection for SAEs
by: Oozeer, Narmeen, et al.
Published: (2025) -
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
by: Oozeer, Narmeen, et al.
Published: (2025) -
DreamReader: An Interpretability Toolkit for Text-to-Image Models
by: Prakash, Nirmalendu, et al.
Published: (2026) -
Spectral Superposition: A Theory of Feature Geometry
by: Ivanov, Georgi, et al.
Published: (2026)