Sparse Feature Coactivation Reveals Causal Semantic Modules in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Deng, Ruixuan, Hu, Xiaoyang, Gilberti, Miles, Storks, Shane, Taxali, Aman, Angstadt, Mike, Sripada, Chandra, Chai, Joyce |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Discovering Properties of Inflectional Morphology in Neural Emergent Communication
by: Gilberti, Miles, et al.
Published: (2025)
by: Gilberti, Miles, et al.
Published: (2025)
Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional Properties
by: Yu, Keunwoo Peter, et al.
Published: (2023)
by: Yu, Keunwoo Peter, et al.
Published: (2023)
SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models
by: Torres-Fonseca, Josue, et al.
Published: (2026)
by: Torres-Fonseca, Josue, et al.
Published: (2026)
Transparent and Coherent Procedural Mistake Detection
by: Storks, Shane, et al.
Published: (2024)
by: Storks, Shane, et al.
Published: (2024)
Distributional Semantics Tracing: A Framework for Explaining Hallucinations in Large Language Models
by: Bhatia, Gagan, et al.
Published: (2025)
by: Bhatia, Gagan, et al.
Published: (2025)
Application of Time-Aware PC algorithm to compute Causal Functional Connectivity in Alzheimer's Disease from fMRI data
by: Biswas, Rahul, et al.
Published: (2023)
by: Biswas, Rahul, et al.
Published: (2023)
Coactive-Staggered Feature in Weyl Materials for Enhancing the Anomalous Nernst Conductivity
by: Ivanov, Vsevolod, et al.
Published: (2022)
by: Ivanov, Vsevolod, et al.
Published: (2022)
Neurophysiological and Behavioral Effects of Micro- and Nanoplastics in Aquatic Organisms.
by: Belanger, Rachelle M, et al.
Published: (2026)
by: Belanger, Rachelle M, et al.
Published: (2026)
Resting‐State Coactivation Patterns of Language Reorganization in Brain Tumors
by: Antonio Napolitano, et al.
Published: (2026)
by: Antonio Napolitano, et al.
Published: (2026)
Unveiling Language-Specific Features in Large Language Models via Sparse Autoencoders
by: Deng, Boyi, et al.
Published: (2025)
by: Deng, Boyi, et al.
Published: (2025)
Applications of Large Language Model Reasoning in Feature Generation
by: Chandra, Dharani
Published: (2025)
by: Chandra, Dharani
Published: (2025)
Revealing Multimodal Causality with Large Language Models
by: Li, Jin, et al.
Published: (2025)
by: Li, Jin, et al.
Published: (2025)
Dissecting Chronos: Sparse Autoencoders Reveal Causal Feature Hierarchies in Time Series Foundation Models
by: Mishra, Anurag
Published: (2026)
by: Mishra, Anurag
Published: (2026)
Walking, Rolling, and Beyond: First-Principles and RL Locomotion on a TARS-Inspired Robot
by: Sripada, Aditya, et al.
Published: (2025)
by: Sripada, Aditya, et al.
Published: (2025)
V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Models
by: Wang, Qidong, et al.
Published: (2025)
by: Wang, Qidong, et al.
Published: (2025)
CITS: Nonparametric Statistical Causal Modeling for High-Resolution Neural Time Series
by: Biswas, Rahul, et al.
Published: (2025)
by: Biswas, Rahul, et al.
Published: (2025)
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models
by: Yu, Keunwoo Peter, et al.
Published: (2025)
by: Yu, Keunwoo Peter, et al.
Published: (2025)
Causal Language Control in Multilingual Transformers via Sparse Feature Steering
by: Chou, Cheng-Ting, et al.
Published: (2025)
by: Chou, Cheng-Ting, et al.
Published: (2025)
Towards A Holistic Landscape of Situated Theory of Mind in Large Language Models
by: Ma, Ziqiao, et al.
Published: (2023)
by: Ma, Ziqiao, et al.
Published: (2023)
Scene Exploration by Vision-Language Models
by: Sripada, Venkatesh, et al.
Published: (2024)
by: Sripada, Venkatesh, et al.
Published: (2024)
Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models
by: Deng, Boyi, et al.
Published: (2026)
by: Deng, Boyi, et al.
Published: (2026)
Identifying and Mitigating Gender Cues in Academic Recommendation Letters: An Interpretability Case Study
by: Alexander, Charlotte S., et al.
Published: (2026)
by: Alexander, Charlotte S., et al.
Published: (2026)
Conflict Adaptation in Vision-Language Models
by: Hu, Xiaoyang
Published: (2025)
by: Hu, Xiaoyang
Published: (2025)
Causal Interpretation of Sparse Autoencoder Features in Vision
by: Han, Sangyu, et al.
Published: (2025)
by: Han, Sangyu, et al.
Published: (2025)
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
by: Marks, Samuel, et al.
Published: (2024)
by: Marks, Samuel, et al.
Published: (2024)
Dimensional Coactivation for Representational Consistency in Frozen Vision Foundation Models
by: Saddik, Izaldein Al-Zyoud Abdulmotaleb El
Published: (2026)
by: Saddik, Izaldein Al-Zyoud Abdulmotaleb El
Published: (2026)
Can Large Language Models Infer Causal Relationships from Real-World Text?
by: Saklad, Ryan, et al.
Published: (2025)
by: Saklad, Ryan, et al.
Published: (2025)
Babysit A Language Model From Scratch: Interactive Language Learning by Trials and Demonstrations
by: Ma, Ziqiao, et al.
Published: (2024)
by: Ma, Ziqiao, et al.
Published: (2024)
Leveraging Large Language Models for Rare Disease Named Entity Recognition
by: Xi, Nan Miles, et al.
Published: (2025)
by: Xi, Nan Miles, et al.
Published: (2025)
Sparse Autoencoders Reveal Temporal Difference Learning in Large Language Models
by: Demircan, Can, et al.
Published: (2024)
by: Demircan, Can, et al.
Published: (2024)
A Large-Scale Empirical Comparison of Meta-Learners and Causal Forests for Heterogeneous Treatment Effect Estimation in Marketing Uplift Modeling
by: Singh, Aman
Published: (2026)
by: Singh, Aman
Published: (2026)
LinkGPT: Teaching Large Language Models To Predict Missing Links
by: He, Zhongmou, et al.
Published: (2024)
by: He, Zhongmou, et al.
Published: (2024)
Do Language Models Encode Semantic Relations? Probing and Sparse Feature Analysis
by: Diera, Andor, et al.
Published: (2026)
by: Diera, Andor, et al.
Published: (2026)
Ephaptic Neuro-Modulation
by: Chawla, Aman
Published: (2025)
by: Chawla, Aman
Published: (2025)
SFC: Shared Feature Calibration in Weakly Supervised Semantic Segmentation
by: Zhao, Xinqiao, et al.
Published: (2024)
by: Zhao, Xinqiao, et al.
Published: (2024)
GROUNDHOG: Grounding Large Language Models to Holistic Segmentation
by: Zhang, Yichi, et al.
Published: (2024)
by: Zhang, Yichi, et al.
Published: (2024)
Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Models
by: Swann, Aiden, et al.
Published: (2026)
by: Swann, Aiden, et al.
Published: (2026)
A Survey of Pipeline Tools for Data Engineering
by: Mbata, Anthony, et al.
Published: (2024)
by: Mbata, Anthony, et al.
Published: (2024)
Improving Factual Accuracy of Neural Table-to-Text Output by Addressing Input Problems in ToTTo
by: Sundararajan, Barkavi, et al.
Published: (2024)
by: Sundararajan, Barkavi, et al.
Published: (2024)
Input Matters: Evaluating Input Structure's Impact on LLM Summaries of Sports Play-by-Play
by: Sundararajan, Barkavi, et al.
Published: (2025)
by: Sundararajan, Barkavi, et al.
Published: (2025)
Similar Items
-
Discovering Properties of Inflectional Morphology in Neural Emergent Communication
by: Gilberti, Miles, et al.
Published: (2025) -
Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional Properties
by: Yu, Keunwoo Peter, et al.
Published: (2023) -
SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models
by: Torres-Fonseca, Josue, et al.
Published: (2026) -
Transparent and Coherent Procedural Mistake Detection
by: Storks, Shane, et al.
Published: (2024) -
Distributional Semantics Tracing: A Framework for Explaining Hallucinations in Large Language Models
by: Bhatia, Gagan, et al.
Published: (2025)