Accelerating Training with Neuron Interaction and Nowcasting Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Knyazev, Boris, Moudgil, Abhinav, Lajoie, Guillaume, Belilovsky, Eugene, Lacoste-Julien, Simon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Celo2: Towards Learned Optimization Free Lunch
by: Moudgil, Abhinav, et al.
Published: (2026)
by: Moudgil, Abhinav, et al.
Published: (2026)
Celo: Training Versatile Learned Optimizers on a Compute Diet
by: Moudgil, Abhinav, et al.
Published: (2025)
by: Moudgil, Abhinav, et al.
Published: (2025)
Meta-learning Optimizers for Communication-Efficient Learning
by: Joseph, Charles-Étienne, et al.
Published: (2023)
by: Joseph, Charles-Étienne, et al.
Published: (2023)
Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients
by: Legate, Gwen, et al.
Published: (2025)
by: Legate, Gwen, et al.
Published: (2025)
Efficient Refusal Ablation in LLM through Optimal Transport
by: Nanfack, Geraldin, et al.
Published: (2026)
by: Nanfack, Geraldin, et al.
Published: (2026)
Not Only the Last-Layer Features for Spurious Correlations: All Layer Deep Feature Reweighting
by: Hameed, Humza Wajid, et al.
Published: (2024)
by: Hameed, Humza Wajid, et al.
Published: (2024)
ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training
by: Nabli, Adel, et al.
Published: (2024)
by: Nabli, Adel, et al.
Published: (2024)
Layerwise LQR for Geometry-Aware Optimization of Deep Networks
by: Dufort-Labbé, Simon, et al.
Published: (2026)
by: Dufort-Labbé, Simon, et al.
Published: (2026)
Beyond Distribution Sharpening: The Importance of Task Rewards
by: Mittal, Sarthak, et al.
Published: (2026)
by: Mittal, Sarthak, et al.
Published: (2026)
PyLO: Towards Accessible Learned Optimizers in PyTorch
by: Janson, Paul, et al.
Published: (2025)
by: Janson, Paul, et al.
Published: (2025)
Less is More: Undertraining Experts Improves Model Upcycling
by: Horoi, Stefan, et al.
Published: (2025)
by: Horoi, Stefan, et al.
Published: (2025)
Non-Uniform Parameter-Wise Model Merging
by: Camacho, Albert Manuel Orozco, et al.
Published: (2024)
by: Camacho, Albert Manuel Orozco, et al.
Published: (2024)
Model Parallelism With Subnetwork Data Parallelism
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
On the Identifiability of Quantized Factors
by: Barin-Pacela, Vitória, et al.
Published: (2023)
by: Barin-Pacela, Vitória, et al.
Published: (2023)
Iterative Amortized Inference: Unifying In-Context Learning and Learned Optimizers
by: Mittal, Sarthak, et al.
Published: (2025)
by: Mittal, Sarthak, et al.
Published: (2025)
In-Context Parametric Inference: Point or Distribution Estimators?
by: Mittal, Sarthak, et al.
Published: (2025)
by: Mittal, Sarthak, et al.
Published: (2025)
DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
Tight Lower Bounds and Improved Convergence in Performative Prediction
by: Khorsandi, Pedram, et al.
Published: (2024)
by: Khorsandi, Pedram, et al.
Published: (2024)
Harmony in Diversity: Merging Neural Networks with Canonical Correlation Analysis
by: Horoi, Stefan, et al.
Published: (2024)
by: Horoi, Stefan, et al.
Published: (2024)
Z-Error Loss for Training Neural Networks
by: Godin, Guillaume
Published: (2025)
by: Godin, Guillaume
Published: (2025)
Graph Neural Networks for Learning Equivariant Representations of Neural Networks
by: Kofinas, Miltiadis, et al.
Published: (2024)
by: Kofinas, Miltiadis, et al.
Published: (2024)
A Complexity-Based Theory of Compositionality
by: Elmoznino, Eric, et al.
Published: (2024)
by: Elmoznino, Eric, et al.
Published: (2024)
Amortized In-Context Bayesian Posterior Estimation
by: Mittal, Sarthak, et al.
Published: (2025)
by: Mittal, Sarthak, et al.
Published: (2025)
Training Verification-Friendly Neural Networks via Neuron Behavior Consistency
by: Liu, Zongxin, et al.
Published: (2024)
by: Liu, Zongxin, et al.
Published: (2024)
MAD-SmaAt-GNet: A Multimodal Advection-Guided Neural Network for Precipitation Nowcasting
by: van Wonderen, Samuel, et al.
Published: (2026)
by: van Wonderen, Samuel, et al.
Published: (2026)
Surrogate Neural Networks Local Stability for Aircraft Predictive Maintenance
by: Ducoffe, Mélanie, et al.
Published: (2024)
by: Ducoffe, Mélanie, et al.
Published: (2024)
Stable Attention Response for Reliable Precipitation Nowcasting
by: Wen, Penghui, et al.
Published: (2026)
by: Wen, Penghui, et al.
Published: (2026)
(Almost) Free Modality Stitching of Foundation Models
by: Singh, Jaisidh, et al.
Published: (2025)
by: Singh, Jaisidh, et al.
Published: (2025)
A Diffusion-Contrastive Graph Neural Network with Virtual Nodes for Wind Nowcasting in Unobserved Regions
by: Shi, Jie, et al.
Published: (2026)
by: Shi, Jie, et al.
Published: (2026)
$μ$LO: Compute-Efficient Meta-Generalization of Learned Optimizers
by: Thérien, Benjamin, et al.
Published: (2024)
by: Thérien, Benjamin, et al.
Published: (2024)
Feasible Learning
by: Ramirez, Juan, et al.
Published: (2025)
by: Ramirez, Juan, et al.
Published: (2025)
Discrete, compositional, and symbolic representations through attractor dynamics
by: Nam, Andrew, et al.
Published: (2023)
by: Nam, Andrew, et al.
Published: (2023)
TN-SHAP-G: Graph-Structured Tensor Network Surrogates for Shapley Values and Interactions
by: Heidari, Farzaneh, et al.
Published: (2026)
by: Heidari, Farzaneh, et al.
Published: (2026)
Optimizing Sensory Neurons: Nonlinear Attention Mechanisms for Accelerated Convergence in Permutation-Invariant Neural Networks for Reinforcement Learning
by: Muzaffar, Junaid, et al.
Published: (2025)
by: Muzaffar, Junaid, et al.
Published: (2025)
Does learning the right latent variables necessarily improve in-context learning?
by: Mittal, Sarthak, et al.
Published: (2024)
by: Mittal, Sarthak, et al.
Published: (2024)
Next-Token Prediction Should be Ambiguity-Sensitive: A Meta-Learning Perspective
by: Gagnon, Leo, et al.
Published: (2025)
by: Gagnon, Leo, et al.
Published: (2025)
Precipitation Nowcasting Using Physics Informed Discriminator Generative Models
by: Yin, Junzhe, et al.
Published: (2024)
by: Yin, Junzhe, et al.
Published: (2024)
Extreme Precipitation Nowcasting using Transformer-based Generative Models
by: Meo, Cristian, et al.
Published: (2024)
by: Meo, Cristian, et al.
Published: (2024)
Beyond MSE: Improving Precipitation Nowcasting with Multi-Quantile Regression
by: van Nieuwkoop, Gijs, et al.
Published: (2026)
by: van Nieuwkoop, Gijs, et al.
Published: (2026)
PIANO: Physics-informed Dual Neural Operator for Precipitation Nowcasting
by: Chin, Seokhyun, et al.
Published: (2025)
by: Chin, Seokhyun, et al.
Published: (2025)
Similar Items
-
Celo2: Towards Learned Optimization Free Lunch
by: Moudgil, Abhinav, et al.
Published: (2026) -
Celo: Training Versatile Learned Optimizers on a Compute Diet
by: Moudgil, Abhinav, et al.
Published: (2025) -
Meta-learning Optimizers for Communication-Efficient Learning
by: Joseph, Charles-Étienne, et al.
Published: (2023) -
Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients
by: Legate, Gwen, et al.
Published: (2025) -
Efficient Refusal Ablation in LLM through Optimal Transport
by: Nanfack, Geraldin, et al.
Published: (2026)