Saved in:
| Main Authors: | He, Zhengfu, Ge, Xuyang, Tang, Qiong, Sun, Tianxiang, Cheng, Qinyuan, Qiu, Xipeng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2402.12201 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Universality: Studying Mechanistic Similarity Across Language Model Architectures
by: Wang, Junxuan, et al.
Published: (2024)
by: Wang, Junxuan, et al.
Published: (2024)
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
by: Wang, Junxuan, et al.
Published: (2025)
by: Wang, Junxuan, et al.
Published: (2025)
Automatically Identifying Local and Global Circuits with Linear Computation Graphs
by: Ge, Xuyang, et al.
Published: (2024)
by: Ge, Xuyang, et al.
Published: (2024)
A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle
by: Zhou, Guancheng, et al.
Published: (2026)
by: Zhou, Guancheng, et al.
Published: (2026)
Agent Alignment in Evolving Social Norms
by: Li, Shimin, et al.
Published: (2024)
by: Li, Shimin, et al.
Published: (2024)
Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition
by: He, Zhengfu, et al.
Published: (2025)
by: He, Zhengfu, et al.
Published: (2025)
Evolution of Concepts in Language Model Pre-Training
by: Ge, Xuyang, et al.
Published: (2025)
by: Ge, Xuyang, et al.
Published: (2025)
Can AI Assistants Know What They Don't Know?
by: Cheng, Qinyuan, et al.
Published: (2024)
by: Cheng, Qinyuan, et al.
Published: (2024)
Tracing the Thought of a Grandmaster-level Chess-Playing Transformer
by: Lin, Rui, et al.
Published: (2026)
by: Lin, Rui, et al.
Published: (2026)
LLM can Achieve Self-Regulation via Hyperparameter Aware Generation
by: Wang, Siyin, et al.
Published: (2024)
by: Wang, Siyin, et al.
Published: (2024)
In-Memory Learning: A Declarative Learning Framework for Large Language Models
by: Wang, Bo, et al.
Published: (2024)
by: Wang, Bo, et al.
Published: (2024)
Automatically Finding Rule-Based Neurons in OthelloGPT
by: Singh, Aditya, et al.
Published: (2025)
by: Singh, Aditya, et al.
Published: (2025)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
by: Zeng, Zhiyuan, et al.
Published: (2025)
by: Zeng, Zhiyuan, et al.
Published: (2025)
DenoSent: A Denoising Objective for Self-Supervised Sentence Representation Learning
by: Wang, Xinghao, et al.
Published: (2024)
by: Wang, Xinghao, et al.
Published: (2024)
Dynamic and Generalizable Process Reward Modeling
by: Yin, Zhangyue, et al.
Published: (2025)
by: Yin, Zhangyue, et al.
Published: (2025)
MechRL: Reinforcement Learning Agents Perform Circuit Discovery for Mechanistic Interpretability
by: Khadka, Barsat
Published: (2026)
by: Khadka, Barsat
Published: (2026)
Formal Mechanistic Interpretability: Automated Circuit Discovery with Provable Guarantees
by: Hadad, Itamar, et al.
Published: (2026)
by: Hadad, Itamar, et al.
Published: (2026)
Unified Active Retrieval for Retrieval Augmented Generation
by: Cheng, Qinyuan, et al.
Published: (2024)
by: Cheng, Qinyuan, et al.
Published: (2024)
Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
by: He, Zhengfu, et al.
Published: (2024)
by: He, Zhengfu, et al.
Published: (2024)
Othello is Solved
by: Takizawa, Hiroki
Published: (2023)
by: Takizawa, Hiroki
Published: (2023)
A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima
by: Tang, Yiming, et al.
Published: (2025)
by: Tang, Yiming, et al.
Published: (2025)
How to Mitigate Overfitting in Weak-to-strong Generalization?
by: Shi, Junhao, et al.
Published: (2025)
by: Shi, Junhao, et al.
Published: (2025)
Scaling Laws for Fact Memorization of Large Language Models
by: Lu, Xingyu, et al.
Published: (2024)
by: Lu, Xingyu, et al.
Published: (2024)
Evaluating Brain-Inspired Modular Training in Automated Circuit Discovery for Mechanistic Interpretability
by: Nainani, Jatin
Published: (2024)
by: Nainani, Jatin
Published: (2024)
Aggregation of Reasoning: A Hierarchical Framework for Enhancing Answer Selection in Large Language Models
by: Yin, Zhangyue, et al.
Published: (2024)
by: Yin, Zhangyue, et al.
Published: (2024)
Othello entre gêneros
by: ROBERTO MOREIRA
Published: (2008)
by: ROBERTO MOREIRA
Published: (2008)
La indianización de Othello
by: Genoveva Castro
Published: (2012)
by: Genoveva Castro
Published: (2012)
DILA: Dictionary Label Attention for Mechanistic Interpretability in High-dimensional Multi-label Medical Coding Prediction
by: Wu, John, et al.
Published: (2024)
by: Wu, John, et al.
Published: (2024)
World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning
by: Wang, Siyin, et al.
Published: (2025)
by: Wang, Siyin, et al.
Published: (2025)
How does GPT-2 Predict Acronyms? Extracting and Understanding a Circuit via Mechanistic Interpretability
by: García-Carrasco, Jorge, et al.
Published: (2024)
by: García-Carrasco, Jorge, et al.
Published: (2024)
RLoop: An Self-Improving Framework for Reinforcement Learning with Iterative Policy Initialization
by: Zhiyuan, Zeng, et al.
Published: (2025)
by: Zhiyuan, Zeng, et al.
Published: (2025)
LLatrieval: LLM-Verified Retrieval for Verifiable Generation
by: Li, Xiaonan, et al.
Published: (2023)
by: Li, Xiaonan, et al.
Published: (2023)
Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance
by: Ye, Jiasheng, et al.
Published: (2024)
by: Ye, Jiasheng, et al.
Published: (2024)
Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective
by: Zeng, Zhiyuan, et al.
Published: (2024)
by: Zeng, Zhiyuan, et al.
Published: (2024)
Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
by: Nainani, Jatin, et al.
Published: (2024)
by: Nainani, Jatin, et al.
Published: (2024)
Revisiting the Othello World Model Hypothesis
by: Yuan, Yifei, et al.
Published: (2025)
by: Yuan, Yifei, et al.
Published: (2025)
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
by: Mishra, Anurag
Published: (2025)
by: Mishra, Anurag
Published: (2025)
R3-RAG: Learning Step-by-Step Reasoning and Retrieval for LLMs via Reinforcement Learning
by: Li, Yuan, et al.
Published: (2025)
by: Li, Yuan, et al.
Published: (2025)
VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search
by: Wang, Yikun, et al.
Published: (2025)
by: Wang, Yikun, et al.
Published: (2025)
CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation
by: Deng, Ruifan, et al.
Published: (2025)
by: Deng, Ruifan, et al.
Published: (2025)
Similar Items
-
Towards Universality: Studying Mechanistic Similarity Across Language Model Architectures
by: Wang, Junxuan, et al.
Published: (2024) -
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
by: Wang, Junxuan, et al.
Published: (2025) -
Automatically Identifying Local and Global Circuits with Linear Computation Graphs
by: Ge, Xuyang, et al.
Published: (2024) -
A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle
by: Zhou, Guancheng, et al.
Published: (2026) -
Agent Alignment in Evolving Social Norms
by: Li, Shimin, et al.
Published: (2024)