Scaling Recurrent Neural Networks to a Billion Parameters with Zero-Order Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Chaubard, Francois, Kochenderfer, Mykel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Gradient Averaging in Parallel Optimization: Improved Robustness through Gradient Agreement Filtering
by: Chaubard, Francois, et al.
Published: (2024)
by: Chaubard, Francois, et al.
Published: (2024)
Graph Q-Learning for Combinatorial Optimization
by: Dax, Victoria M., et al.
Published: (2024)
by: Dax, Victoria M., et al.
Published: (2024)
Hierarchical Zero-Order Optimization for Deep Neural Networks
by: Cao, Sansheng, et al.
Published: (2026)
by: Cao, Sansheng, et al.
Published: (2026)
Failure Probability Estimation for Black-Box Autonomous Systems using State-Dependent Importance Sampling Proposals
by: Delecki, Harrison, et al.
Published: (2024)
by: Delecki, Harrison, et al.
Published: (2024)
Zono-Conformal Prediction: Zonotope-Based Uncertainty Quantification for Regression and Classification Tasks
by: Lützow, Laura, et al.
Published: (2025)
by: Lützow, Laura, et al.
Published: (2025)
Imperfect World Models are Exploitable
by: Bhamidipaty, Logan Mondal, et al.
Published: (2026)
by: Bhamidipaty, Logan Mondal, et al.
Published: (2026)
Billion-Scale Graph Foundation Models
by: Bechler-Speicher, Maya, et al.
Published: (2026)
by: Bechler-Speicher, Maya, et al.
Published: (2026)
Recurrent Diffusion for Large-Scale Parameter Generation
by: Wang, Kai, et al.
Published: (2025)
by: Wang, Kai, et al.
Published: (2025)
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
LPS-GNN : Deploying Graph Neural Networks on Graphs with 100-Billion Edges
by: Cheng, Xu, et al.
Published: (2025)
by: Cheng, Xu, et al.
Published: (2025)
Enhanced Importance Sampling through Latent Space Exploration in Normalizing Flows
by: Kruse, Liam A., et al.
Published: (2025)
by: Kruse, Liam A., et al.
Published: (2025)
Zenith: Scaling up Ranking Models for Billion-scale Livestreaming Recommendation
by: Zhang, Ruifeng, et al.
Published: (2026)
by: Zhang, Ruifeng, et al.
Published: (2026)
Signal-Adaptive Trust Regions for Gradient-Free Optimization of Recurrent Spiking Neural Networks
by: Li, Jinhao, et al.
Published: (2026)
by: Li, Jinhao, et al.
Published: (2026)
Generative System Dynamics in Recurrent Neural Networks
by: Casoni, Michele, et al.
Published: (2025)
by: Casoni, Michele, et al.
Published: (2025)
Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts
by: Shi, Xiaoming, et al.
Published: (2024)
by: Shi, Xiaoming, et al.
Published: (2024)
Communication-Efficient Byzantine-Resilient Federated Zero-Order Optimization
by: Neto, Afonso de Sá Delgado, et al.
Published: (2024)
by: Neto, Afonso de Sá Delgado, et al.
Published: (2024)
ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parameters
by: Hansen-Estruch, Philippe, et al.
Published: (2026)
by: Hansen-Estruch, Philippe, et al.
Published: (2026)
Each Prompt Matters: Scaling Reinforcement Learning Without Wasting Rollouts on Hundred-Billion-Scale MoE
by: Zeng, Anxiang, et al.
Published: (2025)
by: Zeng, Anxiang, et al.
Published: (2025)
BARNN: A Bayesian Autoregressive and Recurrent Neural Network
by: Coscia, Dario, et al.
Published: (2025)
by: Coscia, Dario, et al.
Published: (2025)
Symmetry in Neural Network Parameter Spaces
by: Zhao, Bo, et al.
Published: (2025)
by: Zhao, Bo, et al.
Published: (2025)
Spectral Higher-Order Neural Networks
by: Peri, Gianluca, et al.
Published: (2026)
by: Peri, Gianluca, et al.
Published: (2026)
Fast Training of Recurrent Neural Networks with Stationary State Feedbacks
by: Caillon, Paul, et al.
Published: (2025)
by: Caillon, Paul, et al.
Published: (2025)
GroverGPT: A Large Language Model with 8 Billion Parameters for Quantum Searching
by: Wang, Haoran, et al.
Published: (2024)
by: Wang, Haoran, et al.
Published: (2024)
MathPile: A Billion-Token-Scale Pretraining Corpus for Math
by: Wang, Zengzhi, et al.
Published: (2023)
by: Wang, Zengzhi, et al.
Published: (2023)
Modular Boundaries in Recurrent Neural Networks
by: Tanner, Jacob, et al.
Published: (2023)
by: Tanner, Jacob, et al.
Published: (2023)
Quantized Approximately Orthogonal Recurrent Neural Networks
by: Foucault, Armand, et al.
Published: (2024)
by: Foucault, Armand, et al.
Published: (2024)
Spectral Theory for Edge Pruning in Asynchronous Recurrent Graph Neural Networks
by: Bessone, Nicolas
Published: (2025)
by: Bessone, Nicolas
Published: (2025)
Dimer-Enhanced Optimization: A First-Order Approach to Escaping Saddle Points in Neural Network Training
by: Hu, Yue, et al.
Published: (2025)
by: Hu, Yue, et al.
Published: (2025)
Identifying Information-Transfer Nodes in a Recurrent Neural Network Reveals Dynamic Representations
by: Hintze, Arend, et al.
Published: (2025)
by: Hintze, Arend, et al.
Published: (2025)
On the Optimizer Dependence of Neural Scaling Laws
by: Ramani, Vansh, et al.
Published: (2026)
by: Ramani, Vansh, et al.
Published: (2026)
SCOUT: A Lightweight Framework for Scenario Coverage Assessment in Autonomous Driving
by: Yildiz, Anil, et al.
Published: (2025)
by: Yildiz, Anil, et al.
Published: (2025)
Flow to Learn: Flow Matching on Neural Network Parameters
by: Saragih, Daniel, et al.
Published: (2025)
by: Saragih, Daniel, et al.
Published: (2025)
Topological Neural Networks: Mitigating the Bottlenecks of Graph Neural Networks via Higher-Order Interactions
by: Giusti, Lorenzo
Published: (2024)
by: Giusti, Lorenzo
Published: (2024)
Language Models as Zero-shot Lossless Gradient Compressors: Towards General Neural Parameter Prior Models
by: Wang, Hui-Po, et al.
Published: (2024)
by: Wang, Hui-Po, et al.
Published: (2024)
DouRN: Improving DouZero by Residual Neural Networks
by: Chen, Yiquan, et al.
Published: (2024)
by: Chen, Yiquan, et al.
Published: (2024)
Compressing Neural Networks Using Tensor Networks with Exponentially Fewer Variational Parameters
by: Qing, Yong, et al.
Published: (2023)
by: Qing, Yong, et al.
Published: (2023)
Recurrent Aggregators in Neural Algorithmic Reasoning
by: Xu, Kaijia, et al.
Published: (2024)
by: Xu, Kaijia, et al.
Published: (2024)
Bayesian Neural Network For Personalized Federated Learning Parameter Selection
by: Luo, Mengen, et al.
Published: (2024)
by: Luo, Mengen, et al.
Published: (2024)
IRNN: Innovation-driven Recurrent Neural Network for Time-Series Data Modeling and Prediction
by: Zhou, Yifan, et al.
Published: (2025)
by: Zhou, Yifan, et al.
Published: (2025)
Uncertainty-Aware Deep Attention Recurrent Neural Network for Heterogeneous Time Series Imputation
by: Qian, Linglong, et al.
Published: (2024)
by: Qian, Linglong, et al.
Published: (2024)
Similar Items
-
Beyond Gradient Averaging in Parallel Optimization: Improved Robustness through Gradient Agreement Filtering
by: Chaubard, Francois, et al.
Published: (2024) -
Graph Q-Learning for Combinatorial Optimization
by: Dax, Victoria M., et al.
Published: (2024) -
Hierarchical Zero-Order Optimization for Deep Neural Networks
by: Cao, Sansheng, et al.
Published: (2026) -
Failure Probability Estimation for Black-Box Autonomous Systems using State-Dependent Importance Sampling Proposals
by: Delecki, Harrison, et al.
Published: (2024) -
Zono-Conformal Prediction: Zonotope-Based Uncertainty Quantification for Regression and Classification Tasks
by: Lützow, Laura, et al.
Published: (2025)