Position: Mechanistic Interpretability Must Disclose Identification Assumptions for Causal Claims
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Zezheng, Liu, Fengming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
von: Sarkar, Nilesh, et al.
Veröffentlicht: (2026)
von: Sarkar, Nilesh, et al.
Veröffentlicht: (2026)
Matryoshka Policy Gradient for Entropy-Regularized RL: Convergence and Global Optimality
von: Ged, François, et al.
Veröffentlicht: (2023)
von: Ged, François, et al.
Veröffentlicht: (2023)
On Privacy Leakage in Tabular Diffusion Models: Influential Factors, Attacker Knowledge, and Metrics
von: Shafieinejad, Masoumeh, et al.
Veröffentlicht: (2026)
von: Shafieinejad, Masoumeh, et al.
Veröffentlicht: (2026)
JAM: Controllable and Responsible Text Generation via Causal Reasoning and Latent Vector Manipulation
von: Huang, Yingbing, et al.
Veröffentlicht: (2025)
von: Huang, Yingbing, et al.
Veröffentlicht: (2025)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
Adaptive Latent-Space Constraints in Personalized Federated Learning
von: Ayromlou, Sana, et al.
Veröffentlicht: (2025)
von: Ayromlou, Sana, et al.
Veröffentlicht: (2025)
The Matching Principle: A Geometric Theory of Loss Functions for Nuisance-Robust Representation Learning
von: Rajput, Vishal
Veröffentlicht: (2026)
von: Rajput, Vishal
Veröffentlicht: (2026)
Safe Continual Reinforcement Learning Methods for Nonstationary Environments. Towards a Survey of the State of the Art
von: Tomashevskiy, Timofey
Veröffentlicht: (2026)
von: Tomashevskiy, Timofey
Veröffentlicht: (2026)
AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models
von: Keeman, Michael
Veröffentlicht: (2026)
von: Keeman, Michael
Veröffentlicht: (2026)
Correcting Stochastic Update Bias in Preconditioned Language Model Optimizers
von: Nayak, Nikhil, et al.
Veröffentlicht: (2026)
von: Nayak, Nikhil, et al.
Veröffentlicht: (2026)
Unpacking Hateful Memes: Presupposed Context and False Claims
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
von: Wang, Zhen, et al.
Veröffentlicht: (2025)
von: Wang, Zhen, et al.
Veröffentlicht: (2025)
Thinking Machines: Mathematical Reasoning in the Age of LLMs
von: Asperti, Andrea, et al.
Veröffentlicht: (2025)
von: Asperti, Andrea, et al.
Veröffentlicht: (2025)
The Efficiency Attenuation Phenomenon: A Computational Challenge to the Language of Thought Hypothesis
von: Zhang, Di
Veröffentlicht: (2026)
von: Zhang, Di
Veröffentlicht: (2026)
An Automatic Text Classification Method Based on Hierarchical Taxonomies, Neural Networks and Document Embedding: The NETHIC Tool
von: Lomasto, Luigi, et al.
Veröffentlicht: (2026)
von: Lomasto, Luigi, et al.
Veröffentlicht: (2026)
Visual Categorization Across Minds and Models: Cognitive Analysis of Human Labeling and Neuro-Symbolic Integration
von: Kabgere, Chethana Prasad
Veröffentlicht: (2025)
von: Kabgere, Chethana Prasad
Veröffentlicht: (2025)
Convolutional Neural Networks Can (Meta-)Learn the Same-Different Relation
von: Gupta, Max, et al.
Veröffentlicht: (2025)
von: Gupta, Max, et al.
Veröffentlicht: (2025)
Interpretability-Guided Bi-objective Optimization: Aligning Accuracy and Explainability
von: Fouladi, Kasra, et al.
Veröffentlicht: (2026)
von: Fouladi, Kasra, et al.
Veröffentlicht: (2026)
When Your Own Output Becomes Your Training Data: Noise-to-Meaning Loops and a Formal RSI Trigger
von: Ando, Rintaro
Veröffentlicht: (2025)
von: Ando, Rintaro
Veröffentlicht: (2025)
Lost or Hidden? A Concept-Level Forgetting in Supervised Continual Learning
von: Filus, Katarzyna, et al.
Veröffentlicht: (2026)
von: Filus, Katarzyna, et al.
Veröffentlicht: (2026)
Robust DDoS-Attack Classification with 3D CNNs Against Adversarial Methods
von: Bragg, Landon, et al.
Veröffentlicht: (2025)
von: Bragg, Landon, et al.
Veröffentlicht: (2025)
Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations
von: Mahale, Ajay Pravin
Veröffentlicht: (2026)
von: Mahale, Ajay Pravin
Veröffentlicht: (2026)
Position Paper: Bounded Alignment: What (Not) To Expect From AGI Agents
von: Minai, Ali A.
Veröffentlicht: (2025)
von: Minai, Ali A.
Veröffentlicht: (2025)
Choosing DAG Models Using Markov and Minimal Edge Count in the Absence of Ground Truth
von: Ramsey, Joseph D., et al.
Veröffentlicht: (2024)
von: Ramsey, Joseph D., et al.
Veröffentlicht: (2024)
Interpretability Can Be Actionable
von: Orgad, Hadas, et al.
Veröffentlicht: (2026)
von: Orgad, Hadas, et al.
Veröffentlicht: (2026)
Structural Plasticity as Active Inference: A Biologically-Inspired Architecture for Homeostatic Control
von: Hill, Brennen A.
Veröffentlicht: (2025)
von: Hill, Brennen A.
Veröffentlicht: (2025)
Why Geometric Continuity Emerges in Deep Neural Networks: Residual Connections and Rotational Symmetry Breaking
von: Jeong, Kyungwon, et al.
Veröffentlicht: (2026)
von: Jeong, Kyungwon, et al.
Veröffentlicht: (2026)
Active Inference with a Self-Prior in the Mirror-Mark Task
von: Kim, Dongmin, et al.
Veröffentlicht: (2026)
von: Kim, Dongmin, et al.
Veröffentlicht: (2026)
Emergence of Goal-Directed Behaviors via Active Inference with Self-Prior
von: Kim, Dongmin, et al.
Veröffentlicht: (2025)
von: Kim, Dongmin, et al.
Veröffentlicht: (2025)
ProbeScale: Probing Analysis to Optimize Neural Scaling Laws for Efficient Small Language Model Inference
von: Das, Sourav
Veröffentlicht: (2026)
von: Das, Sourav
Veröffentlicht: (2026)
Unsolvability Ceiling in Multi-LLM Routing: An Empirical Study of Evaluation Artifacts
von: Garg, Saloni, et al.
Veröffentlicht: (2026)
von: Garg, Saloni, et al.
Veröffentlicht: (2026)
Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference
von: Mathew, Aby Mammen
Veröffentlicht: (2026)
von: Mathew, Aby Mammen
Veröffentlicht: (2026)
SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulator
von: Zhou, Xueyang, et al.
Veröffentlicht: (2025)
von: Zhou, Xueyang, et al.
Veröffentlicht: (2025)
Emergence of Self-Awareness in Artificial Systems: A Minimalist Three-Layer Approach to Artificial Consciousness
von: Iida, Kurando
Veröffentlicht: (2025)
von: Iida, Kurando
Veröffentlicht: (2025)
From Imitation to Interaction: Mastering Game of Schnapsen with Shallow Reinforcement Learning
von: Klačan, Ján, et al.
Veröffentlicht: (2026)
von: Klačan, Ján, et al.
Veröffentlicht: (2026)
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
von: Badshah, Sher, et al.
Veröffentlicht: (2024)
von: Badshah, Sher, et al.
Veröffentlicht: (2024)
Training Language Models to Win Debates with Self-Play Improves Judge Accuracy
von: Arnesen, Samuel, et al.
Veröffentlicht: (2024)
von: Arnesen, Samuel, et al.
Veröffentlicht: (2024)
Revisiting Parameter-Based Knowledge Editing in Large Language Models: Theoretical Limits and Empirical Evidence
von: Ren, Wanying, et al.
Veröffentlicht: (2026)
von: Ren, Wanying, et al.
Veröffentlicht: (2026)
Improving LLM Agent Planning with In-Context Learning via Atomic Fact Augmentation and Lookahead Search
von: Holt, Samuel, et al.
Veröffentlicht: (2025)
von: Holt, Samuel, et al.
Veröffentlicht: (2025)
Multi-Level Fusion Graph Neural Network for Molecule Property Prediction
von: Liu, XiaYu, et al.
Veröffentlicht: (2025)
von: Liu, XiaYu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
von: Sarkar, Nilesh, et al.
Veröffentlicht: (2026) -
Matryoshka Policy Gradient for Entropy-Regularized RL: Convergence and Global Optimality
von: Ged, François, et al.
Veröffentlicht: (2023) -
On Privacy Leakage in Tabular Diffusion Models: Influential Factors, Attacker Knowledge, and Metrics
von: Shafieinejad, Masoumeh, et al.
Veröffentlicht: (2026) -
JAM: Controllable and Responsible Text Generation via Causal Reasoning and Latent Vector Manipulation
von: Huang, Yingbing, et al.
Veröffentlicht: (2025) -
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)