Automatically Finding Rule-Based Neurons in OthelloGPT
Fuente:
arXiv
Saved in:
| Main Authors: | Singh, Aditya, Wen, Zihang, Medicherla, Srujananjali, Karvonen, Adam, Rager, Can |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
by: Karvonen, Adam, et al.
Published: (2024)
by: Karvonen, Adam, et al.
Published: (2024)
An Adversarial Example for Direct Logit Attribution: Memory Management in GELU-4L
by: Janiak, Jett, et al.
Published: (2023)
by: Janiak, Jett, et al.
Published: (2023)
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
by: Karvonen, Adam, et al.
Published: (2024)
by: Karvonen, Adam, et al.
Published: (2024)
Robustly Improving LLM Fairness in Realistic Settings via Interpretability
by: Karvonen, Adam, et al.
Published: (2025)
by: Karvonen, Adam, et al.
Published: (2025)
Discovering Forbidden Topics in Language Models
by: Rager, Can, et al.
Published: (2025)
by: Rager, Can, et al.
Published: (2025)
Finding NeMo: Localizing Neurons Responsible For Memorization in Diffusion Models
by: Hintersdorf, Dominik, et al.
Published: (2024)
by: Hintersdorf, Dominik, et al.
Published: (2024)
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
by: Bussmann, Bart, et al.
Published: (2025)
by: Bussmann, Bart, et al.
Published: (2025)
Automatic Extraction of Linguistic Description from Fuzzy Rule Base
by: Siminski, Krzysztof, et al.
Published: (2024)
by: Siminski, Krzysztof, et al.
Published: (2024)
MiniZero: Comparative Analysis of AlphaZero and MuZero on Go, Othello, and Atari Games
by: Wu, Ti-Rong, et al.
Published: (2023)
by: Wu, Ti-Rong, et al.
Published: (2023)
Automatically Finding Reward Model Biases
by: Wang, Atticus, et al.
Published: (2026)
by: Wang, Atticus, et al.
Published: (2026)
Universal Neurons in GPT2 Language Models
by: Gurnee, Wes, et al.
Published: (2024)
by: Gurnee, Wes, et al.
Published: (2024)
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
by: Marks, Samuel, et al.
Published: (2024)
by: Marks, Samuel, et al.
Published: (2024)
Enabling Regional Explainability by Automatic and Model-agnostic Rule Extraction
by: Chen, Yu, et al.
Published: (2024)
by: Chen, Yu, et al.
Published: (2024)
Decomposing Attention To Find Context-Sensitive Neurons
by: Gibson, Alex
Published: (2025)
by: Gibson, Alex
Published: (2025)
Neuron-Anchored Rule Extraction for Large Language Models via Contrastive Hierarchical Ablation
by: Sovrano, Francesco, et al.
Published: (2026)
by: Sovrano, Francesco, et al.
Published: (2026)
DiFR: Inference Verification Despite Nondeterminism
by: Karvonen, Adam, et al.
Published: (2025)
by: Karvonen, Adam, et al.
Published: (2025)
Finding Culture-Sensitive Neurons in Vision-Language Models
by: Zhao, Xiutian, et al.
Published: (2025)
by: Zhao, Xiutian, et al.
Published: (2025)
Do Sparse Autoencoders Capture Concept Manifolds?
by: Bhalla, Usha, et al.
Published: (2026)
by: Bhalla, Usha, et al.
Published: (2026)
TractRLFusion: A GPT-Based Multi-Critic Policy Fusion Framework for Fiber Tractography
by: Joshi, Ankita, et al.
Published: (2026)
by: Joshi, Ankita, et al.
Published: (2026)
Universal Neurons in GPT-2: Emergence, Persistence, and Functional Impact
by: Nandan, Advey, et al.
Published: (2025)
by: Nandan, Advey, et al.
Published: (2025)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
by: Casademunt, Helena, et al.
Published: (2025)
by: Casademunt, Helena, et al.
Published: (2025)
Dynamic Interpretability for Model Comparison via Decision Rules
by: Rida, Adam, et al.
Published: (2023)
by: Rida, Adam, et al.
Published: (2023)
Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers
by: Ranganath, Aditya
Published: (2026)
by: Ranganath, Aditya
Published: (2026)
NeuronSeek: On Stability and Expressivity of Task-driven Neurons
by: Pei, Hanyu, et al.
Published: (2025)
by: Pei, Hanyu, et al.
Published: (2025)
Can GPT Redefine Medical Understanding? Evaluating GPT on Biomedical Machine Reading Comprehension
by: Vatsal, Shubham, et al.
Published: (2024)
by: Vatsal, Shubham, et al.
Published: (2024)
CFT-RAG: An Entity Tree Based Retrieval Augmented Generation Algorithm With Cuckoo Filter
by: Li, Zihang, et al.
Published: (2025)
by: Li, Zihang, et al.
Published: (2025)
Negation Neglect: When models fail to learn negations in training
by: Mayne, Harry, et al.
Published: (2026)
by: Mayne, Harry, et al.
Published: (2026)
Fantastic Pretraining Optimizers and Where to Find Them
by: Wen, Kaiyue, et al.
Published: (2025)
by: Wen, Kaiyue, et al.
Published: (2025)
Shapley Neuron Values for Continual Learning: Which Neurons Matter Most?
by: Vahedifar, Mohammad Ali, et al.
Published: (2026)
by: Vahedifar, Mohammad Ali, et al.
Published: (2026)
Scale When Needed: Adaptive Neuron-level Mixed Precision Quantization Aware Training
by: Varshney, Ayush K., et al.
Published: (2026)
by: Varshney, Ayush K., et al.
Published: (2026)
The Quest for the Right Mediator: Surveying Mechanistic Interpretability Through the Lens of Causal Mediation Analysis
by: Mueller, Aaron, et al.
Published: (2024)
by: Mueller, Aaron, et al.
Published: (2024)
Can GPT Improve the State of Prior Authorization via Guideline Based Automated Question Answering?
by: Vatsal, Shubham, et al.
Published: (2024)
by: Vatsal, Shubham, et al.
Published: (2024)
Whispers of Many Shores: Cultural Alignment through Collaborative Cultural Expertise
by: Feng, Shuai, et al.
Published: (2025)
by: Feng, Shuai, et al.
Published: (2025)
SAFR: Neuron Redistribution for Interpretability
by: Chang, Ruidi, et al.
Published: (2025)
by: Chang, Ruidi, et al.
Published: (2025)
Artificial Kuramoto Oscillatory Neurons
by: Miyato, Takeru, et al.
Published: (2024)
by: Miyato, Takeru, et al.
Published: (2024)
Rethinking the Function of Neurons in KANs
by: Altarabichi, Mohammed Ghaith
Published: (2024)
by: Altarabichi, Mohammed Ghaith
Published: (2024)
TE2Rules: Explaining Tree Ensembles using Rules
by: Lal, G Roshan, et al.
Published: (2022)
by: Lal, G Roshan, et al.
Published: (2022)
MAnchors: Memorization-Based Acceleration of Anchors via Rule Reuse and Transformation
by: Yu, Haonan, et al.
Published: (2025)
by: Yu, Haonan, et al.
Published: (2025)
Incorporating Expert Rules into Neural Networks in the Framework of Concept-Based Learning
by: Konstantinov, Andrei V., et al.
Published: (2024)
by: Konstantinov, Andrei V., et al.
Published: (2024)
Model Fusion via Neuron Transplantation
by: Öz, Muhammed, et al.
Published: (2025)
by: Öz, Muhammed, et al.
Published: (2025)
Similar Items
-
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models
by: Karvonen, Adam, et al.
Published: (2024) -
An Adversarial Example for Direct Logit Attribution: Memory Management in GELU-4L
by: Janiak, Jett, et al.
Published: (2023) -
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
by: Karvonen, Adam, et al.
Published: (2024) -
Robustly Improving LLM Fairness in Realistic Settings via Interpretability
by: Karvonen, Adam, et al.
Published: (2025) -
Discovering Forbidden Topics in Language Models
by: Rager, Can, et al.
Published: (2025)