Feature Starvation as Geometric Instability in Sparse Autoencoders
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chaudhry, Faris, Yano, Keisuke, Monod, Anthea |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Trajectory-Restricted Optimization Conditions and Geometry-Aware Linear Convergence
von: Chaudhry, Faris, et al.
Veröffentlicht: (2026)
von: Chaudhry, Faris, et al.
Veröffentlicht: (2026)
The Geometry of Projection Heads: Conditioning, Invariance, and Collapse
von: Chaudhry, Faris
Veröffentlicht: (2026)
von: Chaudhry, Faris
Veröffentlicht: (2026)
Asymptotic and Finite-Time Guarantees for Langevin-Based Temperature Annealing in InfoNCE
von: Chaudhry, Faris
Veröffentlicht: (2026)
von: Chaudhry, Faris
Veröffentlicht: (2026)
Riemannian Neural Optimal Transport
von: Micheli, Alessandro, et al.
Veröffentlicht: (2026)
von: Micheli, Alessandro, et al.
Veröffentlicht: (2026)
Multi-Objective Optimization for Sparse Deep Multi-Task Learning
von: Hotegni, S. S., et al.
Veröffentlicht: (2023)
von: Hotegni, S. S., et al.
Veröffentlicht: (2023)
Geometric Neural Operators (GNPs) for Data-Driven Deep Learning of Non-Euclidean Operators
von: Quackenbush, Blaine, et al.
Veröffentlicht: (2024)
von: Quackenbush, Blaine, et al.
Veröffentlicht: (2024)
Hereditary Geometric Meta-RL: Nonlocal Generalization via Task Symmetries
von: Nitschke, Paul, et al.
Veröffentlicht: (2026)
von: Nitschke, Paul, et al.
Veröffentlicht: (2026)
A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models
von: Han, X. Y., et al.
Veröffentlicht: (2025)
von: Han, X. Y., et al.
Veröffentlicht: (2025)
Global Convergence and Rich Feature Learning in $L$-Layer Infinite-Width Neural Networks under $μ$P Parametrization
von: Chen, Zixiang, et al.
Veröffentlicht: (2025)
von: Chen, Zixiang, et al.
Veröffentlicht: (2025)
Probabilistic Geometric Alignment via Bayesian Latent Transport for Domain-Adaptive Foundation Models
von: Aueawatthanaphisut, Aueaphum, et al.
Veröffentlicht: (2026)
von: Aueawatthanaphisut, Aueaphum, et al.
Veröffentlicht: (2026)
Feature-aligned N-BEATS with Sinkhorn divergence
von: Lee, Joonhun, et al.
Veröffentlicht: (2023)
von: Lee, Joonhun, et al.
Veröffentlicht: (2023)
Geometric Data Valuation via Leverage Scores
von: Mendoza-Smith, Rodrigo
Veröffentlicht: (2025)
von: Mendoza-Smith, Rodrigo
Veröffentlicht: (2025)
Straight-Through meets Sparse Recovery: the Support Exploration Algorithm
von: Mohamed, Mimoun, et al.
Veröffentlicht: (2023)
von: Mohamed, Mimoun, et al.
Veröffentlicht: (2023)
Mean-Field Control on Sparse Graphs: From Local Limits to GNNs via Neighborhood Distributions
von: Schmidt, Tobias, et al.
Veröffentlicht: (2026)
von: Schmidt, Tobias, et al.
Veröffentlicht: (2026)
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
von: Chen, Siyu, et al.
Veröffentlicht: (2024)
An Approximate Ascent Approach To Prove Convergence of PPO
von: Doering, Leif, et al.
Veröffentlicht: (2026)
von: Doering, Leif, et al.
Veröffentlicht: (2026)
On a Gradient Approach to Chebyshev Center Problems with Applications to Function Learning
von: Raghuvanshi, Abhinav, et al.
Veröffentlicht: (2026)
von: Raghuvanshi, Abhinav, et al.
Veröffentlicht: (2026)
Delightful Distributed Policy Gradient
von: Osband, Ian
Veröffentlicht: (2026)
von: Osband, Ian
Veröffentlicht: (2026)
Optimal Pattern Detection Tree for Symbolic Rule-Based Classification
von: Hong, Young-Chae, et al.
Veröffentlicht: (2026)
von: Hong, Young-Chae, et al.
Veröffentlicht: (2026)
Predictive and Prescriptive AI toward Optimizing Wildfire Suppression
von: Boussioux, Leonard, et al.
Veröffentlicht: (2026)
von: Boussioux, Leonard, et al.
Veröffentlicht: (2026)
$γ$-weakly $θ$-up-concavity: A Unified Framework for Non-Convex Optimization Beyond DR-Submodular and OSS Functions
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2026)
von: Pedramfar, Mohammad, et al.
Veröffentlicht: (2026)
From Soliloquy to Agora: Memory-Enhanced LLM Agents with Decentralized Debate for Optimization Modeling
von: Lin, Jianghao, et al.
Veröffentlicht: (2026)
von: Lin, Jianghao, et al.
Veröffentlicht: (2026)
Hierarchical Mixture-of-Experts with Two-Stage Optimization
von: Molodtsov, Gleb, et al.
Veröffentlicht: (2026)
von: Molodtsov, Gleb, et al.
Veröffentlicht: (2026)
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
von: Huang, Yu, et al.
Veröffentlicht: (2026)
von: Huang, Yu, et al.
Veröffentlicht: (2026)
Budget-aware Auto Optimizer Configurator
von: Liu, Kang, et al.
Veröffentlicht: (2026)
von: Liu, Kang, et al.
Veröffentlicht: (2026)
Optimizer-Induced Mode Connectivity: From AdamW to Muon
von: Zhang, Fangzhao, et al.
Veröffentlicht: (2026)
von: Zhang, Fangzhao, et al.
Veröffentlicht: (2026)
Efficient and provably convergent end-to-end training of deep neural networks with linear constraints
von: Yang, Zonglin, et al.
Veröffentlicht: (2026)
von: Yang, Zonglin, et al.
Veröffentlicht: (2026)
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less
von: Liu, Yuxing, et al.
Veröffentlicht: (2026)
von: Liu, Yuxing, et al.
Veröffentlicht: (2026)
Cactus: Accelerating Auto-Regressive Decoding with Constrained Acceptance Speculative Sampling
von: Hao, Yongchang, et al.
Veröffentlicht: (2026)
von: Hao, Yongchang, et al.
Veröffentlicht: (2026)
The Newton-Muon Optimizer
von: Du, Zhehang, et al.
Veröffentlicht: (2026)
von: Du, Zhehang, et al.
Veröffentlicht: (2026)
Delightful Policy Gradient
von: Osband, Ian
Veröffentlicht: (2026)
von: Osband, Ian
Veröffentlicht: (2026)
Constraint-Anchored Attribution: Feasibility-Certified Counterfactuals and Bonferroni-PAC Sufficient Subsets for Neural CO Policies
von: Lafifi, Sohaib
Veröffentlicht: (2026)
von: Lafifi, Sohaib
Veröffentlicht: (2026)
Centrality-Based Pruning for Efficient Echo State Networks
von: Laudari, Sudip
Veröffentlicht: (2026)
von: Laudari, Sudip
Veröffentlicht: (2026)
Demystifying Manifold Constraints in LLM Pre-training
von: An, Kang, et al.
Veröffentlicht: (2026)
von: An, Kang, et al.
Veröffentlicht: (2026)
A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints
von: Nie, Chengyi, et al.
Veröffentlicht: (2026)
von: Nie, Chengyi, et al.
Veröffentlicht: (2026)
Muon Dynamics as a Spectral Wasserstein Flow
von: Peyré, Gabriel
Veröffentlicht: (2026)
von: Peyré, Gabriel
Veröffentlicht: (2026)
Bilevel Optimization over Saddle Points of Zero-Sum Markov Games
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
MINTS: Minimalist Thompson Sampling
von: Wang, Kaizheng
Veröffentlicht: (2026)
von: Wang, Kaizheng
Veröffentlicht: (2026)
MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models
von: Huang, Feihu, et al.
Veröffentlicht: (2026)
von: Huang, Feihu, et al.
Veröffentlicht: (2026)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
von: Kim, Jihwan, et al.
Veröffentlicht: (2026)
von: Kim, Jihwan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Trajectory-Restricted Optimization Conditions and Geometry-Aware Linear Convergence
von: Chaudhry, Faris, et al.
Veröffentlicht: (2026) -
The Geometry of Projection Heads: Conditioning, Invariance, and Collapse
von: Chaudhry, Faris
Veröffentlicht: (2026) -
Asymptotic and Finite-Time Guarantees for Langevin-Based Temperature Annealing in InfoNCE
von: Chaudhry, Faris
Veröffentlicht: (2026) -
Riemannian Neural Optimal Transport
von: Micheli, Alessandro, et al.
Veröffentlicht: (2026) -
Multi-Objective Optimization for Sparse Deep Multi-Task Learning
von: Hotegni, S. S., et al.
Veröffentlicht: (2023)