Communication Dynamics Neural Networks: FFT-Diagonalized Layers for Improved Hessian Conditioning at Reduced Parameter Count
Fuente:
arXiv
Guardado en:
| Autor principal: | Pan, Lurong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Revisiting Scalable Hessian Diagonal Approximations for Applications in Reinforcement Learning
por: Elsayed, Mohamed, et al.
Publicado: (2024)
por: Elsayed, Mohamed, et al.
Publicado: (2024)
Dynamic Layer Tying for Parameter-Efficient Transformers
por: Hay, Tamir David, et al.
Publicado: (2024)
por: Hay, Tamir David, et al.
Publicado: (2024)
Dynamic Continual Learning: Harnessing Parameter Uncertainty for Improved Network Adaptation
por: Angelini, Christopher, et al.
Publicado: (2025)
por: Angelini, Christopher, et al.
Publicado: (2025)
HELENE: Hessian Layer-wise Clipping and Gradient Annealing for Accelerating Fine-tuning LLM with Zeroth-order Optimization
por: Zhao, Huaqin, et al.
Publicado: (2024)
por: Zhao, Huaqin, et al.
Publicado: (2024)
Diagonal Adaptive Non-local Observables on Quantum Neural Networks
por: Tseng, Huan-Hsin, et al.
Publicado: (2026)
por: Tseng, Huan-Hsin, et al.
Publicado: (2026)
LAPLEX: The FFT of Learnable Laplace Kernels
por: Struski, Łukasz, et al.
Publicado: (2026)
por: Struski, Łukasz, et al.
Publicado: (2026)
Parameter Efficient Mamba Tuning via Projector-targeted Diagonal-centric Linear Transformation
por: Ham, Seokil, et al.
Publicado: (2024)
por: Ham, Seokil, et al.
Publicado: (2024)
Symmetry in Neural Network Parameter Spaces
por: Zhao, Bo, et al.
Publicado: (2025)
por: Zhao, Bo, et al.
Publicado: (2025)
FFT-based Dynamic Subspace Selection for Low-Rank Adaptive Optimization of Large Language Models
por: Modoranu, Ionut-Vlad, et al.
Publicado: (2025)
por: Modoranu, Ionut-Vlad, et al.
Publicado: (2025)
HeRo-Q: A General Framework for Stable Low Bit Quantization via Hessian Conditioning
por: Zhang, Jinhao Zhang Yunquan, et al.
Publicado: (2026)
por: Zhang, Jinhao Zhang Yunquan, et al.
Publicado: (2026)
Flow to Learn: Flow Matching on Neural Network Parameters
por: Saragih, Daniel, et al.
Publicado: (2025)
por: Saragih, Daniel, et al.
Publicado: (2025)
Ray-Tracing for Conditionally Activated Neural Networks
por: Gallicchio, Claudio, et al.
Publicado: (2025)
por: Gallicchio, Claudio, et al.
Publicado: (2025)
Beyond Parameter Count: Implicit Bias in Soft Mixture of Experts
por: Chung, Youngseog, et al.
Publicado: (2024)
por: Chung, Youngseog, et al.
Publicado: (2024)
Layer Embedding Deep Fusion Graph Neural Network
por: Xu, Taihua, et al.
Publicado: (2026)
por: Xu, Taihua, et al.
Publicado: (2026)
Compressing Neural Networks Using Tensor Networks with Exponentially Fewer Variational Parameters
por: Qing, Yong, et al.
Publicado: (2023)
por: Qing, Yong, et al.
Publicado: (2023)
Bayesian Neural Network For Personalized Federated Learning Parameter Selection
por: Luo, Mengen, et al.
Publicado: (2024)
por: Luo, Mengen, et al.
Publicado: (2024)
Latent Communication in Artificial Neural Networks
por: Moschella, Luca
Publicado: (2024)
por: Moschella, Luca
Publicado: (2024)
Symbol Correctness in Deep Neural Networks Containing Symbolic Layers
por: Bembenek, Aaron, et al.
Publicado: (2024)
por: Bembenek, Aaron, et al.
Publicado: (2024)
Hessian-Free Online Certified Unlearning
por: Qiao, Xinbao, et al.
Publicado: (2024)
por: Qiao, Xinbao, et al.
Publicado: (2024)
Conditional LoRA Parameter Generation
por: Jin, Xiaolong, et al.
Publicado: (2024)
por: Jin, Xiaolong, et al.
Publicado: (2024)
INSPIRE-GNN: Intelligent Sensor Placement to Improve Sparse Bicycling Network Prediction via Reinforcement Learning Boosted Graph Neural Networks
por: Gupta, Mohit, et al.
Publicado: (2025)
por: Gupta, Mohit, et al.
Publicado: (2025)
SBS: Enhancing Parameter-Efficiency of Neural Representations for Neural Networks via Spectral Bias Suppression
por: Xie, Qihu, et al.
Publicado: (2025)
por: Xie, Qihu, et al.
Publicado: (2025)
Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset
por: Galashov, Alexandre, et al.
Publicado: (2024)
por: Galashov, Alexandre, et al.
Publicado: (2024)
Neural Inhibition Improves Dynamic Routing and Mixture of Experts
por: Zou, Will Y., et al.
Publicado: (2025)
por: Zou, Will Y., et al.
Publicado: (2025)
Flow-Induced Diagonal Gaussian Processes
por: Lin, Moule, et al.
Publicado: (2025)
por: Lin, Moule, et al.
Publicado: (2025)
Neural Network Parameter-optimization of Gaussian pmDAGs
por: Saremi, Mehrzad
Publicado: (2023)
por: Saremi, Mehrzad
Publicado: (2023)
SnareNet: Flexible Repair Layers for Neural Networks with Hard Constraints
por: Chu, Ya-Chi, et al.
Publicado: (2026)
por: Chu, Ya-Chi, et al.
Publicado: (2026)
The Hidden Power of Normalization Layers in Neural Networks: Exponential Capacity Control
por: Than, Khoat
Publicado: (2025)
por: Than, Khoat
Publicado: (2025)
SplitQuant: Layer Splitting for Low-Bit Neural Network Quantization
por: Song, Jaewoo, et al.
Publicado: (2025)
por: Song, Jaewoo, et al.
Publicado: (2025)
Reduce Computational Complexity for Convolutional Layers by Skipping Zeros
por: Zhang, Zhiyi, et al.
Publicado: (2023)
por: Zhang, Zhiyi, et al.
Publicado: (2023)
DynaMoE: Dynamic Token-Level Expert Activation with Layer-Wise Adaptive Capacity for Mixture-of-Experts Neural Networks
por: Gülmez, Gökdeniz
Publicado: (2026)
por: Gülmez, Gökdeniz
Publicado: (2026)
On the Spatiotemporal Dynamics of Generalization in Neural Networks
por: Wei, Zichao
Publicado: (2026)
por: Wei, Zichao
Publicado: (2026)
Wormhole Dynamics in Deep Neural Networks
por: Lai, Yen-Lung, et al.
Publicado: (2025)
por: Lai, Yen-Lung, et al.
Publicado: (2025)
How to Provably Improve Return Conditioned Supervised Learning?
por: Liu, Zhishuai, et al.
Publicado: (2025)
por: Liu, Zhishuai, et al.
Publicado: (2025)
Epi$^2$-Net: Advancing Epidemic Dynamics Forecasting with Physics-Inspired Neural Networks
por: Sun, Rui, et al.
Publicado: (2025)
por: Sun, Rui, et al.
Publicado: (2025)
NdLinear: Preserving Multi-Dimensional Structure for Parameter-Efficient Neural Networks
por: Reneau, Alex, et al.
Publicado: (2025)
por: Reneau, Alex, et al.
Publicado: (2025)
Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization
por: Chaubard, Francois, et al.
Publicado: (2025)
por: Chaubard, Francois, et al.
Publicado: (2025)
Puppet-CNN: Continuous Parameter Dynamics for Input-Adaptive Convolutional Networks
por: Xing, Yucheng, et al.
Publicado: (2024)
por: Xing, Yucheng, et al.
Publicado: (2024)
Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency
por: Dwyer, Joe
Publicado: (2026)
por: Dwyer, Joe
Publicado: (2026)
Graph-Conditioned Mixture of Graph Neural Network Experts for Traffic Forecasting
por: Ghaffari, Amirhossein, et al.
Publicado: (2026)
por: Ghaffari, Amirhossein, et al.
Publicado: (2026)
Ejemplares similares
-
Revisiting Scalable Hessian Diagonal Approximations for Applications in Reinforcement Learning
por: Elsayed, Mohamed, et al.
Publicado: (2024) -
Dynamic Layer Tying for Parameter-Efficient Transformers
por: Hay, Tamir David, et al.
Publicado: (2024) -
Dynamic Continual Learning: Harnessing Parameter Uncertainty for Improved Network Adaptation
por: Angelini, Christopher, et al.
Publicado: (2025) -
HELENE: Hessian Layer-wise Clipping and Gradient Annealing for Accelerating Fine-tuning LLM with Zeroth-order Optimization
por: Zhao, Huaqin, et al.
Publicado: (2024) -
Diagonal Adaptive Non-local Observables on Quantum Neural Networks
por: Tseng, Huan-Hsin, et al.
Publicado: (2026)