Residual Koopman Spectral Profiling for Predicting and Preventing Transformer Training Instability
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Bum Jun, Taniguchi, Shohei, Kawano, Makoto, Iwasawa, Yusuke, Matsuo, Yutaka |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unlocking Noise-Resistant Vision: Key Architectural Secrets for Robust Models
by: Kim, Bum Jun, et al.
Published: (2025)
by: Kim, Bum Jun, et al.
Published: (2025)
QuadNorm: Resolution-Robust Normalization for Neural Operators
by: Kim, Bum Jun, et al.
Published: (2026)
by: Kim, Bum Jun, et al.
Published: (2026)
Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
by: Furuta, Hiroki, et al.
Published: (2024)
by: Furuta, Hiroki, et al.
Published: (2024)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
by: Gambardella, Andrew, et al.
Published: (2024)
by: Gambardella, Andrew, et al.
Published: (2024)
Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
by: Matsutani, Kohsei, et al.
Published: (2026)
by: Matsutani, Kohsei, et al.
Published: (2026)
Safe Transformer: An Explicit Safety Bit For Interpretable And Controllable Alignment
by: Feng, Jingyuan, et al.
Published: (2026)
by: Feng, Jingyuan, et al.
Published: (2026)
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
by: Gambardella, Andrew, et al.
Published: (2025)
by: Gambardella, Andrew, et al.
Published: (2025)
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
by: Minegishi, Gouki, et al.
Published: (2025)
by: Minegishi, Gouki, et al.
Published: (2025)
Understanding Emergent Misalignment via Feature Superposition Geometry
by: Minegishi, Gouki, et al.
Published: (2026)
by: Minegishi, Gouki, et al.
Published: (2026)
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
by: Cao, Qi, et al.
Published: (2026)
by: Cao, Qi, et al.
Published: (2026)
Thinking While Listening: Fast-Slow Recurrence for Long-Horizon Sequential Modeling
by: Takashiro, Shota, et al.
Published: (2026)
by: Takashiro, Shota, et al.
Published: (2026)
C-voting: Confidence-Based Test-Time Voting without Explicit Energy Functions
by: Kubo, Kenji, et al.
Published: (2026)
by: Kubo, Kenji, et al.
Published: (2026)
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
by: Minegishi, Gouki, et al.
Published: (2025)
by: Minegishi, Gouki, et al.
Published: (2025)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
by: Takashiro, Shota, et al.
Published: (2026)
by: Takashiro, Shota, et al.
Published: (2026)
Self-Harmony: Learning to Harmonize Self-Supervision and Self-Play in Test-Time Reinforcement Learning
by: Wang, Ru, et al.
Published: (2025)
by: Wang, Ru, et al.
Published: (2025)
Position Encoding with Random Float Sampling Enhances Length Generalization of Transformers
by: Shimizu, Atsushi, et al.
Published: (2026)
by: Shimizu, Atsushi, et al.
Published: (2026)
Bridging Lottery Ticket and Grokking: Understanding Grokking from Inner Structure of Networks
by: Minegishi, Gouki, et al.
Published: (2023)
by: Minegishi, Gouki, et al.
Published: (2023)
GenORM: Generalizable One-shot Rope Manipulation with Parameter-Aware Policy
by: Kuroki, So, et al.
Published: (2023)
by: Kuroki, So, et al.
Published: (2023)
Configuring Data Augmentations to Reduce Variance Shift in Positional Embedding of Vision Transformers
by: Kim, Bum Jun, et al.
Published: (2024)
by: Kim, Bum Jun, et al.
Published: (2024)
Enhancing Unimodal Latent Representations in Multimodal VAEs through Iterative Amortized Inference
by: Oshima, Yuta, et al.
Published: (2024)
by: Oshima, Yuta, et al.
Published: (2024)
A Comprehensive Survey on Physical Risk Control in the Era of Foundation Model-enabled Robotics
by: Kojima, Takeshi, et al.
Published: (2025)
by: Kojima, Takeshi, et al.
Published: (2025)
CSRA: Controlled Spectral Residual Augmentation for Robust Sepsis Prediction
by: Guo, Honglin, et al.
Published: (2026)
by: Guo, Honglin, et al.
Published: (2026)
Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology
by: Fernando, Jesseba, et al.
Published: (2026)
by: Fernando, Jesseba, et al.
Published: (2026)
Retrieve-Augmented Generation for Speeding up Diffusion Policy without Additional Training
by: Odonchimed, Sodtavilan, et al.
Published: (2025)
by: Odonchimed, Sodtavilan, et al.
Published: (2025)
Learnable Koopman-Enhanced Transformer-Based Time Series Forecasting with Spectral Control
by: Forootani, Ali, et al.
Published: (2026)
by: Forootani, Ali, et al.
Published: (2026)
Large Language Models as Theory of Mind Aware Generative Agents with Counterfactual Reflection
by: Yang, Bo, et al.
Published: (2025)
by: Yang, Bo, et al.
Published: (2025)
Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models
by: Gu, Xiaojie, et al.
Published: (2026)
by: Gu, Xiaojie, et al.
Published: (2026)
Does "Do Differentiable Simulators Give Better Policy Gradients?'' Give Better Policy Gradients?
by: Onoda, Ku, et al.
Published: (2026)
by: Onoda, Ku, et al.
Published: (2026)
Topology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph Properties
by: Minegishi, Gouki, et al.
Published: (2025)
by: Minegishi, Gouki, et al.
Published: (2025)
Stochastic Subsampling With Average Pooling
by: Kim, Bum Jun, et al.
Published: (2024)
by: Kim, Bum Jun, et al.
Published: (2024)
The Spectral Lifecycle of Transformer Training: Transient Compression Waves, Persistent Spectral Gradients, and the Q/K--V Asymmetry
by: Liu, Yi
Published: (2026)
by: Liu, Yi
Published: (2026)
Exposing Limitations of Language Model Agents in Sequential-Task Compositions on the Web
by: Furuta, Hiroki, et al.
Published: (2023)
by: Furuta, Hiroki, et al.
Published: (2023)
SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces
by: Oshima, Yuta, et al.
Published: (2024)
by: Oshima, Yuta, et al.
Published: (2024)
ClinDet-Bench: Beyond Abstention, Evaluating Judgment Determinability of LLMs in Clinical Decision-Making
by: Watanabe, Yusuke, et al.
Published: (2026)
by: Watanabe, Yusuke, et al.
Published: (2026)
The Fourier Spectral Transformer Networks For Efficient and Generalizable Nonlinear PDEs Prediction
by: Li, Beibei
Published: (2025)
by: Li, Beibei
Published: (2025)
Electrocardiogram Classification with Transformers Using Koopman and Wavelet Features
by: Ghosh, Sucheta, et al.
Published: (2026)
by: Ghosh, Sucheta, et al.
Published: (2026)
Residual Feature Integration is Sufficient to Prevent Negative Transfer
by: Xu, Yichen, et al.
Published: (2025)
by: Xu, Yichen, et al.
Published: (2025)
Predictive Modeling and Explainable AI for Veterinary Safety Profiles, Residue Assessment, and Health Outcomes Using Real-World Data and Physicochemical Properties
by: Sholehrasa, Hossein, et al.
Published: (2025)
by: Sholehrasa, Hossein, et al.
Published: (2025)
The Disappearance of Timestep Embedding in Modern Time-Dependent Neural Networks
by: Kim, Bum Jun, et al.
Published: (2024)
by: Kim, Bum Jun, et al.
Published: (2024)
Koopman Invariants as Drivers of Emergent Time-Series Clustering in Joint-Embedding Predictive Architectures
by: Ruiz-Morales, Pablo, et al.
Published: (2025)
by: Ruiz-Morales, Pablo, et al.
Published: (2025)
Similar Items
-
Unlocking Noise-Resistant Vision: Key Architectural Secrets for Robust Models
by: Kim, Bum Jun, et al.
Published: (2025) -
QuadNorm: Resolution-Robust Normalization for Neural Operators
by: Kim, Bum Jun, et al.
Published: (2026) -
Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
by: Furuta, Hiroki, et al.
Published: (2024) -
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
by: Gambardella, Andrew, et al.
Published: (2024) -
Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
by: Matsutani, Kohsei, et al.
Published: (2026)