Residual Koopman Spectral Profiling for Predicting and Preventing Transformer Training Instability
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Bum Jun, Taniguchi, Shohei, Kawano, Makoto, Iwasawa, Yusuke, Matsuo, Yutaka |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Unlocking Noise-Resistant Vision: Key Architectural Secrets for Robust Models
di: Kim, Bum Jun, et al.
Pubblicazione: (2025)
di: Kim, Bum Jun, et al.
Pubblicazione: (2025)
QuadNorm: Resolution-Robust Normalization for Neural Operators
di: Kim, Bum Jun, et al.
Pubblicazione: (2026)
di: Kim, Bum Jun, et al.
Pubblicazione: (2026)
Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
di: Furuta, Hiroki, et al.
Pubblicazione: (2024)
di: Furuta, Hiroki, et al.
Pubblicazione: (2024)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
di: Gambardella, Andrew, et al.
Pubblicazione: (2024)
di: Gambardella, Andrew, et al.
Pubblicazione: (2024)
Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
di: Matsutani, Kohsei, et al.
Pubblicazione: (2026)
di: Matsutani, Kohsei, et al.
Pubblicazione: (2026)
Safe Transformer: An Explicit Safety Bit For Interpretable And Controllable Alignment
di: Feng, Jingyuan, et al.
Pubblicazione: (2026)
di: Feng, Jingyuan, et al.
Pubblicazione: (2026)
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
di: Gambardella, Andrew, et al.
Pubblicazione: (2025)
di: Gambardella, Andrew, et al.
Pubblicazione: (2025)
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
di: Minegishi, Gouki, et al.
Pubblicazione: (2025)
di: Minegishi, Gouki, et al.
Pubblicazione: (2025)
Understanding Emergent Misalignment via Feature Superposition Geometry
di: Minegishi, Gouki, et al.
Pubblicazione: (2026)
di: Minegishi, Gouki, et al.
Pubblicazione: (2026)
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
di: Cao, Qi, et al.
Pubblicazione: (2026)
di: Cao, Qi, et al.
Pubblicazione: (2026)
Thinking While Listening: Fast-Slow Recurrence for Long-Horizon Sequential Modeling
di: Takashiro, Shota, et al.
Pubblicazione: (2026)
di: Takashiro, Shota, et al.
Pubblicazione: (2026)
C-voting: Confidence-Based Test-Time Voting without Explicit Energy Functions
di: Kubo, Kenji, et al.
Pubblicazione: (2026)
di: Kubo, Kenji, et al.
Pubblicazione: (2026)
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
di: Minegishi, Gouki, et al.
Pubblicazione: (2025)
di: Minegishi, Gouki, et al.
Pubblicazione: (2025)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
di: Takashiro, Shota, et al.
Pubblicazione: (2026)
di: Takashiro, Shota, et al.
Pubblicazione: (2026)
Self-Harmony: Learning to Harmonize Self-Supervision and Self-Play in Test-Time Reinforcement Learning
di: Wang, Ru, et al.
Pubblicazione: (2025)
di: Wang, Ru, et al.
Pubblicazione: (2025)
Position Encoding with Random Float Sampling Enhances Length Generalization of Transformers
di: Shimizu, Atsushi, et al.
Pubblicazione: (2026)
di: Shimizu, Atsushi, et al.
Pubblicazione: (2026)
Bridging Lottery Ticket and Grokking: Understanding Grokking from Inner Structure of Networks
di: Minegishi, Gouki, et al.
Pubblicazione: (2023)
di: Minegishi, Gouki, et al.
Pubblicazione: (2023)
GenORM: Generalizable One-shot Rope Manipulation with Parameter-Aware Policy
di: Kuroki, So, et al.
Pubblicazione: (2023)
di: Kuroki, So, et al.
Pubblicazione: (2023)
Configuring Data Augmentations to Reduce Variance Shift in Positional Embedding of Vision Transformers
di: Kim, Bum Jun, et al.
Pubblicazione: (2024)
di: Kim, Bum Jun, et al.
Pubblicazione: (2024)
Enhancing Unimodal Latent Representations in Multimodal VAEs through Iterative Amortized Inference
di: Oshima, Yuta, et al.
Pubblicazione: (2024)
di: Oshima, Yuta, et al.
Pubblicazione: (2024)
A Comprehensive Survey on Physical Risk Control in the Era of Foundation Model-enabled Robotics
di: Kojima, Takeshi, et al.
Pubblicazione: (2025)
di: Kojima, Takeshi, et al.
Pubblicazione: (2025)
CSRA: Controlled Spectral Residual Augmentation for Robust Sepsis Prediction
di: Guo, Honglin, et al.
Pubblicazione: (2026)
di: Guo, Honglin, et al.
Pubblicazione: (2026)
Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology
di: Fernando, Jesseba, et al.
Pubblicazione: (2026)
di: Fernando, Jesseba, et al.
Pubblicazione: (2026)
Retrieve-Augmented Generation for Speeding up Diffusion Policy without Additional Training
di: Odonchimed, Sodtavilan, et al.
Pubblicazione: (2025)
di: Odonchimed, Sodtavilan, et al.
Pubblicazione: (2025)
Learnable Koopman-Enhanced Transformer-Based Time Series Forecasting with Spectral Control
di: Forootani, Ali, et al.
Pubblicazione: (2026)
di: Forootani, Ali, et al.
Pubblicazione: (2026)
Large Language Models as Theory of Mind Aware Generative Agents with Counterfactual Reflection
di: Yang, Bo, et al.
Pubblicazione: (2025)
di: Yang, Bo, et al.
Pubblicazione: (2025)
Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models
di: Gu, Xiaojie, et al.
Pubblicazione: (2026)
di: Gu, Xiaojie, et al.
Pubblicazione: (2026)
Does "Do Differentiable Simulators Give Better Policy Gradients?'' Give Better Policy Gradients?
di: Onoda, Ku, et al.
Pubblicazione: (2026)
di: Onoda, Ku, et al.
Pubblicazione: (2026)
Topology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph Properties
di: Minegishi, Gouki, et al.
Pubblicazione: (2025)
di: Minegishi, Gouki, et al.
Pubblicazione: (2025)
Stochastic Subsampling With Average Pooling
di: Kim, Bum Jun, et al.
Pubblicazione: (2024)
di: Kim, Bum Jun, et al.
Pubblicazione: (2024)
The Spectral Lifecycle of Transformer Training: Transient Compression Waves, Persistent Spectral Gradients, and the Q/K--V Asymmetry
di: Liu, Yi
Pubblicazione: (2026)
di: Liu, Yi
Pubblicazione: (2026)
Exposing Limitations of Language Model Agents in Sequential-Task Compositions on the Web
di: Furuta, Hiroki, et al.
Pubblicazione: (2023)
di: Furuta, Hiroki, et al.
Pubblicazione: (2023)
SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces
di: Oshima, Yuta, et al.
Pubblicazione: (2024)
di: Oshima, Yuta, et al.
Pubblicazione: (2024)
ClinDet-Bench: Beyond Abstention, Evaluating Judgment Determinability of LLMs in Clinical Decision-Making
di: Watanabe, Yusuke, et al.
Pubblicazione: (2026)
di: Watanabe, Yusuke, et al.
Pubblicazione: (2026)
The Fourier Spectral Transformer Networks For Efficient and Generalizable Nonlinear PDEs Prediction
di: Li, Beibei
Pubblicazione: (2025)
di: Li, Beibei
Pubblicazione: (2025)
Electrocardiogram Classification with Transformers Using Koopman and Wavelet Features
di: Ghosh, Sucheta, et al.
Pubblicazione: (2026)
di: Ghosh, Sucheta, et al.
Pubblicazione: (2026)
Residual Feature Integration is Sufficient to Prevent Negative Transfer
di: Xu, Yichen, et al.
Pubblicazione: (2025)
di: Xu, Yichen, et al.
Pubblicazione: (2025)
Predictive Modeling and Explainable AI for Veterinary Safety Profiles, Residue Assessment, and Health Outcomes Using Real-World Data and Physicochemical Properties
di: Sholehrasa, Hossein, et al.
Pubblicazione: (2025)
di: Sholehrasa, Hossein, et al.
Pubblicazione: (2025)
The Disappearance of Timestep Embedding in Modern Time-Dependent Neural Networks
di: Kim, Bum Jun, et al.
Pubblicazione: (2024)
di: Kim, Bum Jun, et al.
Pubblicazione: (2024)
Koopman Invariants as Drivers of Emergent Time-Series Clustering in Joint-Embedding Predictive Architectures
di: Ruiz-Morales, Pablo, et al.
Pubblicazione: (2025)
di: Ruiz-Morales, Pablo, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Unlocking Noise-Resistant Vision: Key Architectural Secrets for Robust Models
di: Kim, Bum Jun, et al.
Pubblicazione: (2025) -
QuadNorm: Resolution-Robust Normalization for Neural Operators
di: Kim, Bum Jun, et al.
Pubblicazione: (2026) -
Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
di: Furuta, Hiroki, et al.
Pubblicazione: (2024) -
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
di: Gambardella, Andrew, et al.
Pubblicazione: (2024) -
Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
di: Matsutani, Kohsei, et al.
Pubblicazione: (2026)