Understanding Emergent Misalignment via Feature Superposition Geometry
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Minegishi, Gouki, Furuta, Hiroki, Kojima, Takeshi, Iwasawa, Yusuke, Matsuo, Yutaka |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
von: Furuta, Hiroki, et al.
Veröffentlicht: (2024)
von: Furuta, Hiroki, et al.
Veröffentlicht: (2024)
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
Topology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph Properties
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
von: Matsutani, Kohsei, et al.
Veröffentlicht: (2026)
von: Matsutani, Kohsei, et al.
Veröffentlicht: (2026)
Safe Transformer: An Explicit Safety Bit For Interpretable And Controllable Alignment
von: Feng, Jingyuan, et al.
Veröffentlicht: (2026)
von: Feng, Jingyuan, et al.
Veröffentlicht: (2026)
Bridging Lottery Ticket and Grokking: Understanding Grokking from Inner Structure of Networks
von: Minegishi, Gouki, et al.
Veröffentlicht: (2023)
von: Minegishi, Gouki, et al.
Veröffentlicht: (2023)
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
von: Matsutani, Kohsei, et al.
Veröffentlicht: (2025)
von: Matsutani, Kohsei, et al.
Veröffentlicht: (2025)
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
von: Gambardella, Andrew, et al.
Veröffentlicht: (2025)
von: Gambardella, Andrew, et al.
Veröffentlicht: (2025)
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
von: Cao, Qi, et al.
Veröffentlicht: (2026)
von: Cao, Qi, et al.
Veröffentlicht: (2026)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
von: Gambardella, Andrew, et al.
Veröffentlicht: (2024)
von: Gambardella, Andrew, et al.
Veröffentlicht: (2024)
Exposing Limitations of Language Model Agents in Sequential-Task Compositions on the Web
von: Furuta, Hiroki, et al.
Veröffentlicht: (2023)
von: Furuta, Hiroki, et al.
Veröffentlicht: (2023)
Residual Koopman Spectral Profiling for Predicting and Preventing Transformer Training Instability
von: Kim, Bum Jun, et al.
Veröffentlicht: (2026)
von: Kim, Bum Jun, et al.
Veröffentlicht: (2026)
Mechanism of Task-oriented Information Removal in In-context Learning
von: Cho, Hakaze, et al.
Veröffentlicht: (2025)
von: Cho, Hakaze, et al.
Veröffentlicht: (2025)
Thinking While Listening: Fast-Slow Recurrence for Long-Horizon Sequential Modeling
von: Takashiro, Shota, et al.
Veröffentlicht: (2026)
von: Takashiro, Shota, et al.
Veröffentlicht: (2026)
C-voting: Confidence-Based Test-Time Voting without Explicit Energy Functions
von: Kubo, Kenji, et al.
Veröffentlicht: (2026)
von: Kubo, Kenji, et al.
Veröffentlicht: (2026)
Which Programming Language and What Features at Pre-training Stage Affect Downstream Logical Inference Performance?
von: Uchiyama, Fumiya, et al.
Veröffentlicht: (2024)
von: Uchiyama, Fumiya, et al.
Veröffentlicht: (2024)
A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
von: Gur, Izzeddin, et al.
Veröffentlicht: (2023)
von: Gur, Izzeddin, et al.
Veröffentlicht: (2023)
A Comprehensive Survey on Physical Risk Control in the Era of Foundation Model-enabled Robotics
von: Kojima, Takeshi, et al.
Veröffentlicht: (2025)
von: Kojima, Takeshi, et al.
Veröffentlicht: (2025)
WorldPack: Compressed Memory Improves Spatial Consistency in Video World Modeling
von: Oshima, Yuta, et al.
Veröffentlicht: (2025)
von: Oshima, Yuta, et al.
Veröffentlicht: (2025)
ClinDet-Bench: Beyond Abstention, Evaluating Judgment Determinability of LLMs in Clinical Decision-Making
von: Watanabe, Yusuke, et al.
Veröffentlicht: (2026)
von: Watanabe, Yusuke, et al.
Veröffentlicht: (2026)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
von: Takashiro, Shota, et al.
Veröffentlicht: (2026)
von: Takashiro, Shota, et al.
Veröffentlicht: (2026)
Self-Harmony: Learning to Harmonize Self-Supervision and Self-Play in Test-Time Reinforcement Learning
von: Wang, Ru, et al.
Veröffentlicht: (2025)
von: Wang, Ru, et al.
Veröffentlicht: (2025)
Spectral Superposition: A Theory of Feature Geometry
von: Ivanov, Georgi, et al.
Veröffentlicht: (2026)
von: Ivanov, Georgi, et al.
Veröffentlicht: (2026)
Multimodal Web Navigation with Instruction-Finetuned Foundation Models
von: Furuta, Hiroki, et al.
Veröffentlicht: (2023)
von: Furuta, Hiroki, et al.
Veröffentlicht: (2023)
Collective Intelligence for 2D Push Manipulations with Mobile Robots
von: Kuroki, So, et al.
Veröffentlicht: (2022)
von: Kuroki, So, et al.
Veröffentlicht: (2022)
Model Organisms for Emergent Misalignment
von: Turner, Edward, et al.
Veröffentlicht: (2025)
von: Turner, Edward, et al.
Veröffentlicht: (2025)
Convergent Linear Representations of Emergent Misalignment
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
Persona Features Control Emergent Misalignment
von: Wang, Miles, et al.
Veröffentlicht: (2025)
von: Wang, Miles, et al.
Veröffentlicht: (2025)
ADOPT: Modified Adam Can Converge with Any $β_2$ with the Optimal Rate
von: Taniguchi, Shohei, et al.
Veröffentlicht: (2024)
von: Taniguchi, Shohei, et al.
Veröffentlicht: (2024)
Geometric-Averaged Preference Optimization for Soft Preference Labels
von: Furuta, Hiroki, et al.
Veröffentlicht: (2024)
von: Furuta, Hiroki, et al.
Veröffentlicht: (2024)
Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback
von: Furuta, Hiroki, et al.
Veröffentlicht: (2024)
von: Furuta, Hiroki, et al.
Veröffentlicht: (2024)
In-Training Defenses against Emergent Misalignment in Language Models
von: Kaczér, David, et al.
Veröffentlicht: (2025)
von: Kaczér, David, et al.
Veröffentlicht: (2025)
On Implications of Scaling Laws on Feature Superposition
von: Katta, Pavan
Veröffentlicht: (2024)
von: Katta, Pavan
Veröffentlicht: (2024)
GenORM: Generalizable One-shot Rope Manipulation with Parameter-Aware Policy
von: Kuroki, So, et al.
Veröffentlicht: (2023)
von: Kuroki, So, et al.
Veröffentlicht: (2023)
Shared Parameter Subspaces and Cross-Task Linearity in Emergently Misaligned Behavior
von: Arturi, Daniel Aarao Reis, et al.
Veröffentlicht: (2025)
von: Arturi, Daniel Aarao Reis, et al.
Veröffentlicht: (2025)
Decomposing Behavioral Phase Transitions in LLMs: Order Parameters for Emergent Misalignment
von: Arnold, Julian, et al.
Veröffentlicht: (2025)
von: Arnold, Julian, et al.
Veröffentlicht: (2025)
From Narrow Unlearning to Emergent Misalignment: Causes, Consequences, and Containment in LLMs
von: Mushtaq, Erum, et al.
Veröffentlicht: (2025)
von: Mushtaq, Erum, et al.
Veröffentlicht: (2025)
From Data Statistics to Feature Geometry: How Correlations Shape Superposition
von: Prieto, Lucas, et al.
Veröffentlicht: (2026)
von: Prieto, Lucas, et al.
Veröffentlicht: (2026)
Enhancing Unimodal Latent Representations in Multimodal VAEs through Iterative Amortized Inference
von: Oshima, Yuta, et al.
Veröffentlicht: (2024)
von: Oshima, Yuta, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
von: Furuta, Hiroki, et al.
Veröffentlicht: (2024) -
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025) -
Topology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph Properties
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025) -
Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
von: Matsutani, Kohsei, et al.
Veröffentlicht: (2026) -
Safe Transformer: An Explicit Safety Bit For Interpretable And Controllable Alignment
von: Feng, Jingyuan, et al.
Veröffentlicht: (2026)