Emergent Low-Rank Training Dynamics in MLPs with Smooth Activations
Fuente:
arXiv
Guardado en:
| Autores principales: | Xu, Alec S., Yaras, Can, Asato, Matthew, Qu, Qing, Balzano, Laura |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Compressible Dynamics in Deep Overparameterized Low-Rank Learning & Adaptation
por: Yaras, Can, et al.
Publicado: (2024)
por: Yaras, Can, et al.
Publicado: (2024)
Out-of-Distribution Generalization of In-Context Learning: A Low-Dimensional Subspace Perspective
por: Kwon, Soo Min, et al.
Publicado: (2025)
por: Kwon, Soo Min, et al.
Publicado: (2025)
MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured Attention
por: Yaras, Can, et al.
Publicado: (2025)
por: Yaras, Can, et al.
Publicado: (2025)
Explaining and Mitigating the Modality Gap in Contrastive Multimodal Learning
por: Yaras, Can, et al.
Publicado: (2024)
por: Yaras, Can, et al.
Publicado: (2024)
An Overview of Low-Rank Structures in the Training and Adaptation of Large Models
por: Balzano, Laura, et al.
Publicado: (2025)
por: Balzano, Laura, et al.
Publicado: (2025)
Linearly Separable Features in Shallow Nonlinear Networks: Width Scales Polynomially with Intrinsic Data Dimension
por: Xu, Alec S., et al.
Publicado: (2025)
por: Xu, Alec S., et al.
Publicado: (2025)
Training MLPs on Graphs without Supervision
por: Wang, Zehong, et al.
Publicado: (2024)
por: Wang, Zehong, et al.
Publicado: (2024)
From Latent Space to Training Data: Explainable Specialization in Minimal MLPs
por: Alba, Enrique, et al.
Publicado: (2026)
por: Alba, Enrique, et al.
Publicado: (2026)
CoLA: Compute-Efficient Pre-Training of LLMs via Low-Rank Activation
por: Liu, Ziyue, et al.
Publicado: (2025)
por: Liu, Ziyue, et al.
Publicado: (2025)
ELAS: Efficient Pre-Training of Low-Rank Large Language Models via 2:4 Activation Sparsity
por: Li, Jiaxi, et al.
Publicado: (2026)
por: Li, Jiaxi, et al.
Publicado: (2026)
Probe-Free Low-Rank Activation Intervention
por: Jiang, Chonghe, et al.
Publicado: (2025)
por: Jiang, Chonghe, et al.
Publicado: (2025)
SimMLP: Training MLPs on Graphs without Supervision
por: Wang, Zehong, et al.
Publicado: (2024)
por: Wang, Zehong, et al.
Publicado: (2024)
Constructing Efficient Fact-Storing MLPs for Transformers
por: Dugan, Owen, et al.
Publicado: (2025)
por: Dugan, Owen, et al.
Publicado: (2025)
Understanding Deep Representation Learning via Layerwise Feature Compression and Discrimination
por: Wang, Peng, et al.
Publicado: (2023)
por: Wang, Peng, et al.
Publicado: (2023)
Rotated Runtime Smooth: Training-Free Activation Smoother for accurate INT4 inference
por: Yi, Ke, et al.
Publicado: (2024)
por: Yi, Ke, et al.
Publicado: (2024)
Harnessing Orthogonality to Train Low-Rank Neural Networks
por: Coquelin, Daniel, et al.
Publicado: (2024)
por: Coquelin, Daniel, et al.
Publicado: (2024)
Memory-Efficient Fine-Tuning via Low-Rank Activation Compression
por: Shi, Jiang-Xin, et al.
Publicado: (2025)
por: Shi, Jiang-Xin, et al.
Publicado: (2025)
MDMLP-EIA: Multi-domain Dynamic MLPs with Energy Invariant Attention for Time Series Forecasting
por: Zhang, Hu, et al.
Publicado: (2025)
por: Zhang, Hu, et al.
Publicado: (2025)
Exploring Dynamic Properties of Backdoor Training Through Information Bottleneck
por: Liu, Xinyu, et al.
Publicado: (2025)
por: Liu, Xinyu, et al.
Publicado: (2025)
Weight-based Decomposition: A Case for Bilinear MLPs
por: Pearce, Michael T., et al.
Publicado: (2024)
por: Pearce, Michael T., et al.
Publicado: (2024)
Federated Dynamical Low-Rank Training with Global Loss Convergence Guarantees
por: Schotthöfer, Steffen, et al.
Publicado: (2024)
por: Schotthöfer, Steffen, et al.
Publicado: (2024)
Global Low-Rank, Local Full-Rank: The Holographic Encoding of Learned Algorithms
por: Xu, Yongzhong
Publicado: (2026)
por: Xu, Yongzhong
Publicado: (2026)
Channel-Wise MLPs Improve the Generalization of Recurrent Convolutional Networks
por: Breslow, Nathan
Publicado: (2025)
por: Breslow, Nathan
Publicado: (2025)
GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection
por: Su, DiJia, et al.
Publicado: (2025)
por: Su, DiJia, et al.
Publicado: (2025)
Optimizer-Induced Low-Dimensional Drift and Transverse Dynamics in Transformer Training
por: Xu, Yongzhong
Publicado: (2026)
por: Xu, Yongzhong
Publicado: (2026)
VQGraph: Rethinking Graph Representation Space for Bridging GNNs and MLPs
por: Yang, Ling, et al.
Publicado: (2023)
por: Yang, Ling, et al.
Publicado: (2023)
Joint Tensor-Train Parameterization for Efficient and Expressive Low-Rank Adaptation
por: Qi, Jun, et al.
Publicado: (2025)
por: Qi, Jun, et al.
Publicado: (2025)
Mimetic Initialization of MLPs
por: Trockman, Asher, et al.
Publicado: (2026)
por: Trockman, Asher, et al.
Publicado: (2026)
Heuristic Methods are Good Teachers to Distill MLPs for Graph Link Prediction
por: Qin, Zongyue, et al.
Publicado: (2025)
por: Qin, Zongyue, et al.
Publicado: (2025)
Diffusion-Assisted Distillation for Self-Supervised Graph Representation Learning with MLPs
por: Ahn, Seong Jin, et al.
Publicado: (2025)
por: Ahn, Seong Jin, et al.
Publicado: (2025)
Low-Rank Quantization-Aware Training for LLMs
por: Bondarenko, Yelysei, et al.
Publicado: (2024)
por: Bondarenko, Yelysei, et al.
Publicado: (2024)
Rank-Aware Spectral Bounds on Attention Logits for Stable Low-Precision Training
por: Emadi, Seyed Morteza
Publicado: (2026)
por: Emadi, Seyed Morteza
Publicado: (2026)
InfoQuant: Shaping Activation Distributions for Low-Bit LLM Quantization
por: Li, Ke, et al.
Publicado: (2026)
por: Li, Ke, et al.
Publicado: (2026)
Modular addition without black-boxes: Compressing explanations of MLPs that compute numerical integration
por: Yip, Chun Hei, et al.
Publicado: (2024)
por: Yip, Chun Hei, et al.
Publicado: (2024)
Learning to Model Graph Structural Information on MLPs via Graph Structure Self-Contrasting
por: Wu, Lirong, et al.
Publicado: (2024)
por: Wu, Lirong, et al.
Publicado: (2024)
Edge-free but Structure-aware: Prototype-Guided Knowledge Distillation from GNNs to MLPs
por: Wu, Taiqiang, et al.
Publicado: (2023)
por: Wu, Taiqiang, et al.
Publicado: (2023)
Q3R: Quadratic Reweighted Rank Regularizer for Effective Low-Rank Training
por: Ghosh, Ipsita, et al.
Publicado: (2025)
por: Ghosh, Ipsita, et al.
Publicado: (2025)
Training Acceleration of Low-Rank Decomposed Networks using Sequential Freezing and Rank Quantization
por: Hajimolahoseini, Habib, et al.
Publicado: (2023)
por: Hajimolahoseini, Habib, et al.
Publicado: (2023)
Activation Sensitivity as a Unifying Principle for Post-Training Quantization
por: Xu, Bruce Changlong
Publicado: (2026)
por: Xu, Bruce Changlong
Publicado: (2026)
Lotus: Efficient LLM Training by Randomized Low-Rank Gradient Projection with Adaptive Subspace Switching
por: Miao, Tianhao, et al.
Publicado: (2026)
por: Miao, Tianhao, et al.
Publicado: (2026)
Ejemplares similares
-
Compressible Dynamics in Deep Overparameterized Low-Rank Learning & Adaptation
por: Yaras, Can, et al.
Publicado: (2024) -
Out-of-Distribution Generalization of In-Context Learning: A Low-Dimensional Subspace Perspective
por: Kwon, Soo Min, et al.
Publicado: (2025) -
MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured Attention
por: Yaras, Can, et al.
Publicado: (2025) -
Explaining and Mitigating the Modality Gap in Contrastive Multimodal Learning
por: Yaras, Can, et al.
Publicado: (2024) -
An Overview of Low-Rank Structures in the Training and Adaptation of Large Models
por: Balzano, Laura, et al.
Publicado: (2025)