Extending $μ$P: Spectral Conditions for Feature Learning Across Optimizers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gupta, Akshita, Ngom, Marieme, Foreman, Sam, Vishwanath, Venkatram |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scalify: scale propagation for efficient low-precision LLM training
von: Balança, Paul, et al.
Veröffentlicht: (2024)
von: Balança, Paul, et al.
Veröffentlicht: (2024)
MAcPNN: Mutual Assisted Learning on Data Streams with Temporal Dependence
von: Giannini, Federico, et al.
Veröffentlicht: (2026)
von: Giannini, Federico, et al.
Veröffentlicht: (2026)
The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth
von: Henry, James
Veröffentlicht: (2026)
von: Henry, James
Veröffentlicht: (2026)
The Curious Case of In-Training Compression of State Space Models
von: Chahine, Makram, et al.
Veröffentlicht: (2025)
von: Chahine, Makram, et al.
Veröffentlicht: (2025)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
von: Basu, Abhinaba
Veröffentlicht: (2026)
von: Basu, Abhinaba
Veröffentlicht: (2026)
How Pruning Reshapes Features: Sparse Autoencoder Analysis of Weight-Pruned Language Models
von: Borobia, Hector, et al.
Veröffentlicht: (2026)
von: Borobia, Hector, et al.
Veröffentlicht: (2026)
ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language Models
von: Danieli, Federico, et al.
Veröffentlicht: (2025)
von: Danieli, Federico, et al.
Veröffentlicht: (2025)
A Practical Guide to Streaming Continual Learning
von: Cossu, Andrea, et al.
Veröffentlicht: (2026)
von: Cossu, Andrea, et al.
Veröffentlicht: (2026)
Probing for Representation Manifolds in Superposition
von: Modell, Alexander
Veröffentlicht: (2026)
von: Modell, Alexander
Veröffentlicht: (2026)
The Origins of Representation Manifolds in Large Language Models
von: Modell, Alexander, et al.
Veröffentlicht: (2025)
von: Modell, Alexander, et al.
Veröffentlicht: (2025)
Don't Look Back in Anger: MAGIC Net for Streaming Continual Learning with Temporal Dependence
von: Giannini, Federico, et al.
Veröffentlicht: (2026)
von: Giannini, Federico, et al.
Veröffentlicht: (2026)
ProactBench: Beyond What The User Asked For
von: Harfi, Sepehr, et al.
Veröffentlicht: (2026)
von: Harfi, Sepehr, et al.
Veröffentlicht: (2026)
SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization
von: Lee, Wooin, et al.
Veröffentlicht: (2026)
von: Lee, Wooin, et al.
Veröffentlicht: (2026)
Cache-to-Cache: Direct Semantic Communication Between Large Language Models
von: Fu, Tianyu, et al.
Veröffentlicht: (2025)
von: Fu, Tianyu, et al.
Veröffentlicht: (2025)
Strategic Doctrine Language Models (sdLM): A Learning-System Framework for Doctrinal Consistency and Geopolitical Forecasting
von: Imanov, Olaf Yunus Laitinen, et al.
Veröffentlicht: (2026)
von: Imanov, Olaf Yunus Laitinen, et al.
Veröffentlicht: (2026)
cPNN: Continuous Progressive Neural Networks for Evolving Streaming Time Series
von: Giannini, Federico, et al.
Veröffentlicht: (2026)
von: Giannini, Federico, et al.
Veröffentlicht: (2026)
Correcting Gradient-Based Circuit Localization via Interaction-Aware Backpropagation
von: Edin, Joakim, et al.
Veröffentlicht: (2025)
von: Edin, Joakim, et al.
Veröffentlicht: (2025)
Geometric Evolution Maps: Extracting Stable Concept Probes from Transformer Residual Streams
von: Henry, James
Veröffentlicht: (2026)
von: Henry, James
Veröffentlicht: (2026)
ProbeScale: Probing Analysis to Optimize Neural Scaling Laws for Efficient Small Language Model Inference
von: Das, Sourav
Veröffentlicht: (2026)
von: Das, Sourav
Veröffentlicht: (2026)
Optimized Gradient Clipping for Noisy Label Learning
von: Ye, Xichen, et al.
Veröffentlicht: (2024)
von: Ye, Xichen, et al.
Veröffentlicht: (2024)
Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
von: Imanov, Olaf Yunus Laitinen
Veröffentlicht: (2026)
von: Imanov, Olaf Yunus Laitinen
Veröffentlicht: (2026)
Latent Object Permanence: Topological Phase Transitions, Free-Energy Principles, and Renormalization Group Flows in Deep Transformer Manifolds
von: Alpay, Faruk, et al.
Veröffentlicht: (2026)
von: Alpay, Faruk, et al.
Veröffentlicht: (2026)
Inference acceleration for large language models using "stairs" assisted greedy generation
von: Grigaliūnas, Domas, et al.
Veröffentlicht: (2024)
von: Grigaliūnas, Domas, et al.
Veröffentlicht: (2024)
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
von: Sarkar, Nilesh, et al.
Veröffentlicht: (2026)
von: Sarkar, Nilesh, et al.
Veröffentlicht: (2026)
DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
von: Wang, Zhen, et al.
Veröffentlicht: (2025)
von: Wang, Zhen, et al.
Veröffentlicht: (2025)
Streaming Continual Learning for Unified Adaptive Intelligence in Dynamic Environments
von: Giannini, Federico, et al.
Veröffentlicht: (2026)
von: Giannini, Federico, et al.
Veröffentlicht: (2026)
Approaching I/O-optimality for Approximate Attention
von: Papp, Pál András, et al.
Veröffentlicht: (2026)
von: Papp, Pál András, et al.
Veröffentlicht: (2026)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
von: Breneur, Oleksandr Marchenko, et al.
Veröffentlicht: (2026)
von: Breneur, Oleksandr Marchenko, et al.
Veröffentlicht: (2026)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
DIVERSED: Relaxed Speculative Decoding via Dynamic Ensemble Verification
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
CWM: An Open-Weights LLM for Research on Code Generation with World Models
von: FAIR CodeGen team, et al.
Veröffentlicht: (2025)
von: FAIR CodeGen team, et al.
Veröffentlicht: (2025)
Stick to your Role! Stability of Personal Values Expressed in Large Language Models
von: Kovač, Grgur, et al.
Veröffentlicht: (2024)
von: Kovač, Grgur, et al.
Veröffentlicht: (2024)
FedRot-LoRA: Mitigating Rotational Misalignment in Federated LoRA
von: Zhang, Haoran, et al.
Veröffentlicht: (2026)
von: Zhang, Haoran, et al.
Veröffentlicht: (2026)
Unsolvability Ceiling in Multi-LLM Routing: An Empirical Study of Evaluation Artifacts
von: Garg, Saloni, et al.
Veröffentlicht: (2026)
von: Garg, Saloni, et al.
Veröffentlicht: (2026)
Research on a hybrid LSTM-CNN-Attention model for text-based web content classification
von: Kuz, Mykola, et al.
Veröffentlicht: (2025)
von: Kuz, Mykola, et al.
Veröffentlicht: (2025)
Linguistic Collapse: Neural Collapse in (Large) Language Models
von: Wu, Robert, et al.
Veröffentlicht: (2024)
von: Wu, Robert, et al.
Veröffentlicht: (2024)
TaylorShift: Shifting the Complexity of Self-Attention from Squared to Linear (and Back) using Taylor-Softmax
von: Nauen, Tobias Christian, et al.
Veröffentlicht: (2024)
von: Nauen, Tobias Christian, et al.
Veröffentlicht: (2024)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
von: Viveiros, André G., et al.
Veröffentlicht: (2025)
von: Viveiros, André G., et al.
Veröffentlicht: (2025)
Why Geometric Continuity Emerges in Deep Neural Networks: Residual Connections and Rotational Symmetry Breaking
von: Jeong, Kyungwon, et al.
Veröffentlicht: (2026)
von: Jeong, Kyungwon, et al.
Veröffentlicht: (2026)
DRO-InstructZero: Distributionally Robust Prompt Optimization for Large Language Models
von: Li, Yangyang
Veröffentlicht: (2025)
von: Li, Yangyang
Veröffentlicht: (2025)
Ähnliche Einträge
-
Scalify: scale propagation for efficient low-precision LLM training
von: Balança, Paul, et al.
Veröffentlicht: (2024) -
MAcPNN: Mutual Assisted Learning on Data Streams with Temporal Dependence
von: Giannini, Federico, et al.
Veröffentlicht: (2026) -
The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth
von: Henry, James
Veröffentlicht: (2026) -
The Curious Case of In-Training Compression of State Space Models
von: Chahine, Makram, et al.
Veröffentlicht: (2025) -
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
von: Basu, Abhinaba
Veröffentlicht: (2026)