Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
Fuente:
arXiv
Salvato in:
| Autori principali: | Furuta, Hiroki, Minegishi, Gouki, Iwasawa, Yusuke, Matsuo, Yutaka |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
di: Minegishi, Gouki, et al.
Pubblicazione: (2025)
di: Minegishi, Gouki, et al.
Pubblicazione: (2025)
Understanding Emergent Misalignment via Feature Superposition Geometry
di: Minegishi, Gouki, et al.
Pubblicazione: (2026)
di: Minegishi, Gouki, et al.
Pubblicazione: (2026)
Bridging Lottery Ticket and Grokking: Understanding Grokking from Inner Structure of Networks
di: Minegishi, Gouki, et al.
Pubblicazione: (2023)
di: Minegishi, Gouki, et al.
Pubblicazione: (2023)
Safe Transformer: An Explicit Safety Bit For Interpretable And Controllable Alignment
di: Feng, Jingyuan, et al.
Pubblicazione: (2026)
di: Feng, Jingyuan, et al.
Pubblicazione: (2026)
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
di: Minegishi, Gouki, et al.
Pubblicazione: (2025)
di: Minegishi, Gouki, et al.
Pubblicazione: (2025)
Topology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph Properties
di: Minegishi, Gouki, et al.
Pubblicazione: (2025)
di: Minegishi, Gouki, et al.
Pubblicazione: (2025)
Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
di: Matsutani, Kohsei, et al.
Pubblicazione: (2026)
di: Matsutani, Kohsei, et al.
Pubblicazione: (2026)
RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
di: Matsutani, Kohsei, et al.
Pubblicazione: (2025)
di: Matsutani, Kohsei, et al.
Pubblicazione: (2025)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
di: Gambardella, Andrew, et al.
Pubblicazione: (2024)
di: Gambardella, Andrew, et al.
Pubblicazione: (2024)
Residual Koopman Spectral Profiling for Predicting and Preventing Transformer Training Instability
di: Kim, Bum Jun, et al.
Pubblicazione: (2026)
di: Kim, Bum Jun, et al.
Pubblicazione: (2026)
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
di: Gambardella, Andrew, et al.
Pubblicazione: (2025)
di: Gambardella, Andrew, et al.
Pubblicazione: (2025)
Exposing Limitations of Language Model Agents in Sequential-Task Compositions on the Web
di: Furuta, Hiroki, et al.
Pubblicazione: (2023)
di: Furuta, Hiroki, et al.
Pubblicazione: (2023)
Mechanism of Task-oriented Information Removal in In-context Learning
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
di: Cao, Qi, et al.
Pubblicazione: (2026)
di: Cao, Qi, et al.
Pubblicazione: (2026)
Thinking While Listening: Fast-Slow Recurrence for Long-Horizon Sequential Modeling
di: Takashiro, Shota, et al.
Pubblicazione: (2026)
di: Takashiro, Shota, et al.
Pubblicazione: (2026)
C-voting: Confidence-Based Test-Time Voting without Explicit Energy Functions
di: Kubo, Kenji, et al.
Pubblicazione: (2026)
di: Kubo, Kenji, et al.
Pubblicazione: (2026)
WorldPack: Compressed Memory Improves Spatial Consistency in Video World Modeling
di: Oshima, Yuta, et al.
Pubblicazione: (2025)
di: Oshima, Yuta, et al.
Pubblicazione: (2025)
Self-Harmony: Learning to Harmonize Self-Supervision and Self-Play in Test-Time Reinforcement Learning
di: Wang, Ru, et al.
Pubblicazione: (2025)
di: Wang, Ru, et al.
Pubblicazione: (2025)
Multimodal Web Navigation with Instruction-Finetuned Foundation Models
di: Furuta, Hiroki, et al.
Pubblicazione: (2023)
di: Furuta, Hiroki, et al.
Pubblicazione: (2023)
Collective Intelligence for 2D Push Manipulations with Mobile Robots
di: Kuroki, So, et al.
Pubblicazione: (2022)
di: Kuroki, So, et al.
Pubblicazione: (2022)
A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
di: Gur, Izzeddin, et al.
Pubblicazione: (2023)
di: Gur, Izzeddin, et al.
Pubblicazione: (2023)
ADOPT: Modified Adam Can Converge with Any $β_2$ with the Optimal Rate
di: Taniguchi, Shohei, et al.
Pubblicazione: (2024)
di: Taniguchi, Shohei, et al.
Pubblicazione: (2024)
Geometric-Averaged Preference Optimization for Soft Preference Labels
di: Furuta, Hiroki, et al.
Pubblicazione: (2024)
di: Furuta, Hiroki, et al.
Pubblicazione: (2024)
Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback
di: Furuta, Hiroki, et al.
Pubblicazione: (2024)
di: Furuta, Hiroki, et al.
Pubblicazione: (2024)
Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models
di: Gu, Xiaojie, et al.
Pubblicazione: (2026)
di: Gu, Xiaojie, et al.
Pubblicazione: (2026)
NeuralGrok: Accelerate Grokking by Neural Gradient Transformation
di: Zhou, Xinyu, et al.
Pubblicazione: (2025)
di: Zhou, Xinyu, et al.
Pubblicazione: (2025)
Feature Repulsion and Spectral Lock-in: An Empirical Study of Two-Layer Network Grokking
di: Xu, Yongzhong
Pubblicazione: (2026)
di: Xu, Yongzhong
Pubblicazione: (2026)
GenORM: Generalizable One-shot Rope Manipulation with Parameter-Aware Policy
di: Kuroki, So, et al.
Pubblicazione: (2023)
di: Kuroki, So, et al.
Pubblicazione: (2023)
Grokking as Structural Inference: Transformers Need Bayesian Lottery Tickets
di: Hidajat, Kai, et al.
Pubblicazione: (2026)
di: Hidajat, Kai, et al.
Pubblicazione: (2026)
Topological Signatures of Grokking
di: Tang, Yifan, et al.
Pubblicazione: (2026)
di: Tang, Yifan, et al.
Pubblicazione: (2026)
Enhancing Unimodal Latent Representations in Multimodal VAEs through Iterative Amortized Inference
di: Oshima, Yuta, et al.
Pubblicazione: (2024)
di: Oshima, Yuta, et al.
Pubblicazione: (2024)
A Comprehensive Survey on Physical Risk Control in the Era of Foundation Model-enabled Robotics
di: Kojima, Takeshi, et al.
Pubblicazione: (2025)
di: Kojima, Takeshi, et al.
Pubblicazione: (2025)
MedRECT: A Medical Reasoning Benchmark for Error Correction in Clinical Texts
di: Iwase, Naoto, et al.
Pubblicazione: (2025)
di: Iwase, Naoto, et al.
Pubblicazione: (2025)
Interpreting Multi-Attribute Confounding through Numerical Attributes in Large Language Models
di: Takagi, Hirohane, et al.
Pubblicazione: (2025)
di: Takagi, Hirohane, et al.
Pubblicazione: (2025)
Controlling Grokking with Nonlinearity and Data Symmetry
di: Salah, Ahmed, et al.
Pubblicazione: (2024)
di: Salah, Ahmed, et al.
Pubblicazione: (2024)
Grokking Explained: A Statistical Phenomenon
di: Carvalho, Breno W., et al.
Pubblicazione: (2025)
di: Carvalho, Breno W., et al.
Pubblicazione: (2025)
Grokking in Linear Models for Logistic Regression
di: Das, Nataraj, et al.
Pubblicazione: (2026)
di: Das, Nataraj, et al.
Pubblicazione: (2026)
Large Language Models as Theory of Mind Aware Generative Agents with Counterfactual Reflection
di: Yang, Bo, et al.
Pubblicazione: (2025)
di: Yang, Bo, et al.
Pubblicazione: (2025)
Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond
di: Jeffares, Alan, et al.
Pubblicazione: (2024)
di: Jeffares, Alan, et al.
Pubblicazione: (2024)
Grokfast: Accelerated Grokking by Amplifying Slow Gradients
di: Lee, Jaerin, et al.
Pubblicazione: (2024)
di: Lee, Jaerin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
di: Minegishi, Gouki, et al.
Pubblicazione: (2025) -
Understanding Emergent Misalignment via Feature Superposition Geometry
di: Minegishi, Gouki, et al.
Pubblicazione: (2026) -
Bridging Lottery Ticket and Grokking: Understanding Grokking from Inner Structure of Networks
di: Minegishi, Gouki, et al.
Pubblicazione: (2023) -
Safe Transformer: An Explicit Safety Bit For Interpretable And Controllable Alignment
di: Feng, Jingyuan, et al.
Pubblicazione: (2026) -
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
di: Minegishi, Gouki, et al.
Pubblicazione: (2025)