Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Zhengfu, Ge, Xuyang, Tang, Qiong, Sun, Tianxiang, Cheng, Qinyuan, Qiu, Xipeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
von: Wang, Junxuan, et al.
Veröffentlicht: (2025)
von: Wang, Junxuan, et al.
Veröffentlicht: (2025)
Towards Universality: Studying Mechanistic Similarity Across Language Model Architectures
von: Wang, Junxuan, et al.
Veröffentlicht: (2024)
von: Wang, Junxuan, et al.
Veröffentlicht: (2024)
A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle
von: Zhou, Guancheng, et al.
Veröffentlicht: (2026)
von: Zhou, Guancheng, et al.
Veröffentlicht: (2026)
Automatically Identifying Local and Global Circuits with Linear Computation Graphs
von: Ge, Xuyang, et al.
Veröffentlicht: (2024)
von: Ge, Xuyang, et al.
Veröffentlicht: (2024)
Agent Alignment in Evolving Social Norms
von: Li, Shimin, et al.
Veröffentlicht: (2024)
von: Li, Shimin, et al.
Veröffentlicht: (2024)
Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition
von: He, Zhengfu, et al.
Veröffentlicht: (2025)
von: He, Zhengfu, et al.
Veröffentlicht: (2025)
Evolution of Concepts in Language Model Pre-Training
von: Ge, Xuyang, et al.
Veröffentlicht: (2025)
von: Ge, Xuyang, et al.
Veröffentlicht: (2025)
Can AI Assistants Know What They Don't Know?
von: Cheng, Qinyuan, et al.
Veröffentlicht: (2024)
von: Cheng, Qinyuan, et al.
Veröffentlicht: (2024)
Tracing the Thought of a Grandmaster-level Chess-Playing Transformer
von: Lin, Rui, et al.
Veröffentlicht: (2026)
von: Lin, Rui, et al.
Veröffentlicht: (2026)
MechRL: Reinforcement Learning Agents Perform Circuit Discovery for Mechanistic Interpretability
von: Khadka, Barsat
Veröffentlicht: (2026)
von: Khadka, Barsat
Veröffentlicht: (2026)
Formal Mechanistic Interpretability: Automated Circuit Discovery with Provable Guarantees
von: Hadad, Itamar, et al.
Veröffentlicht: (2026)
von: Hadad, Itamar, et al.
Veröffentlicht: (2026)
In-Memory Learning: A Declarative Learning Framework for Large Language Models
von: Wang, Bo, et al.
Veröffentlicht: (2024)
von: Wang, Bo, et al.
Veröffentlicht: (2024)
Automatically Finding Rule-Based Neurons in OthelloGPT
von: Singh, Aditya, et al.
Veröffentlicht: (2025)
von: Singh, Aditya, et al.
Veröffentlicht: (2025)
LLM can Achieve Self-Regulation via Hyperparameter Aware Generation
von: Wang, Siyin, et al.
Veröffentlicht: (2024)
von: Wang, Siyin, et al.
Veröffentlicht: (2024)
DenoSent: A Denoising Objective for Self-Supervised Sentence Representation Learning
von: Wang, Xinghao, et al.
Veröffentlicht: (2024)
von: Wang, Xinghao, et al.
Veröffentlicht: (2024)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025)
A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima
von: Tang, Yiming, et al.
Veröffentlicht: (2025)
von: Tang, Yiming, et al.
Veröffentlicht: (2025)
Dynamic and Generalizable Process Reward Modeling
von: Yin, Zhangyue, et al.
Veröffentlicht: (2025)
von: Yin, Zhangyue, et al.
Veröffentlicht: (2025)
Evaluating Brain-Inspired Modular Training in Automated Circuit Discovery for Mechanistic Interpretability
von: Nainani, Jatin
Veröffentlicht: (2024)
von: Nainani, Jatin
Veröffentlicht: (2024)
Unified Active Retrieval for Retrieval Augmented Generation
von: Cheng, Qinyuan, et al.
Veröffentlicht: (2024)
von: Cheng, Qinyuan, et al.
Veröffentlicht: (2024)
Othello is Solved
von: Takizawa, Hiroki
Veröffentlicht: (2023)
von: Takizawa, Hiroki
Veröffentlicht: (2023)
Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
von: He, Zhengfu, et al.
Veröffentlicht: (2024)
von: He, Zhengfu, et al.
Veröffentlicht: (2024)
DILA: Dictionary Label Attention for Mechanistic Interpretability in High-dimensional Multi-label Medical Coding Prediction
von: Wu, John, et al.
Veröffentlicht: (2024)
von: Wu, John, et al.
Veröffentlicht: (2024)
How does GPT-2 Predict Acronyms? Extracting and Understanding a Circuit via Mechanistic Interpretability
von: García-Carrasco, Jorge, et al.
Veröffentlicht: (2024)
von: García-Carrasco, Jorge, et al.
Veröffentlicht: (2024)
How to Mitigate Overfitting in Weak-to-strong Generalization?
von: Shi, Junhao, et al.
Veröffentlicht: (2025)
von: Shi, Junhao, et al.
Veröffentlicht: (2025)
Scaling Laws for Fact Memorization of Large Language Models
von: Lu, Xingyu, et al.
Veröffentlicht: (2024)
von: Lu, Xingyu, et al.
Veröffentlicht: (2024)
Othello entre gêneros
von: ROBERTO MOREIRA
Veröffentlicht: (2008)
von: ROBERTO MOREIRA
Veröffentlicht: (2008)
Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
von: Nainani, Jatin, et al.
Veröffentlicht: (2024)
von: Nainani, Jatin, et al.
Veröffentlicht: (2024)
Aggregation of Reasoning: A Hierarchical Framework for Enhancing Answer Selection in Large Language Models
von: Yin, Zhangyue, et al.
Veröffentlicht: (2024)
von: Yin, Zhangyue, et al.
Veröffentlicht: (2024)
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
von: Mishra, Anurag
Veröffentlicht: (2025)
von: Mishra, Anurag
Veröffentlicht: (2025)
RLoop: An Self-Improving Framework for Reinforcement Learning with Iterative Policy Initialization
von: Zhiyuan, Zeng, et al.
Veröffentlicht: (2025)
von: Zhiyuan, Zeng, et al.
Veröffentlicht: (2025)
La indianización de Othello
von: Genoveva Castro
Veröffentlicht: (2012)
von: Genoveva Castro
Veröffentlicht: (2012)
World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
von: Che, Liwei, et al.
Veröffentlicht: (2026)
von: Che, Liwei, et al.
Veröffentlicht: (2026)
Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2024)
LLatrieval: LLM-Verified Retrieval for Verifiable Generation
von: Li, Xiaonan, et al.
Veröffentlicht: (2023)
von: Li, Xiaonan, et al.
Veröffentlicht: (2023)
Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance
von: Ye, Jiasheng, et al.
Veröffentlicht: (2024)
von: Ye, Jiasheng, et al.
Veröffentlicht: (2024)
Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers
von: Żukowska, Nina, et al.
Veröffentlicht: (2026)
von: Żukowska, Nina, et al.
Veröffentlicht: (2026)
Localized Definitions and Distributed Reasoning: A Proof-of-Concept Mechanistic Interpretability Study via Activation Patching
von: Bahador, Nooshin
Veröffentlicht: (2025)
von: Bahador, Nooshin
Veröffentlicht: (2025)
Function Induction and Task Generalization: An Interpretability Study with Off-by-One Addition
von: Ye, Qinyuan, et al.
Veröffentlicht: (2025)
von: Ye, Qinyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
von: Wang, Junxuan, et al.
Veröffentlicht: (2025) -
Towards Universality: Studying Mechanistic Similarity Across Language Model Architectures
von: Wang, Junxuan, et al.
Veröffentlicht: (2024) -
A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle
von: Zhou, Guancheng, et al.
Veröffentlicht: (2026) -
Automatically Identifying Local and Global Circuits with Linear Computation Graphs
von: Ge, Xuyang, et al.
Veröffentlicht: (2024) -
Agent Alignment in Evolving Social Norms
von: Li, Shimin, et al.
Veröffentlicht: (2024)