MotherNet: Fast Training and Inference via Hyper-Network Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Müller, Andreas, Curino, Carlo, Ramakrishnan, Raghu |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FSC-Net: Fast-Slow Consolidation Networks for Continual Learning
by: Gorrim, Mohamed El
Published: (2025)
by: Gorrim, Mohamed El
Published: (2025)
Inference-Time Machine Unlearning via Gated Activation Redirection
by: Turani, Vinícius Conte, et al.
Published: (2026)
by: Turani, Vinícius Conte, et al.
Published: (2026)
Fusing Rewards and Preferences in Reinforcement Learning
by: Khorasani, Sadegh, et al.
Published: (2025)
by: Khorasani, Sadegh, et al.
Published: (2025)
HGCN(O): A Self-Tuning GCN HyperModel Toolkit for Outcome Prediction in Event-Sequence Data
by: Wang, Fang, et al.
Published: (2025)
by: Wang, Fang, et al.
Published: (2025)
ReBoot: Encrypted Training of Deep Neural Networks with CKKS Bootstrapping
by: Pirillo, Alberto, et al.
Published: (2025)
by: Pirillo, Alberto, et al.
Published: (2025)
New Paradigm of Adversarial Training: Releasing Accuracy-Robustness Trade-Off via Dummy Class
by: Wang, Yanyun, et al.
Published: (2024)
by: Wang, Yanyun, et al.
Published: (2024)
Deep Variational Inference Symbolic Regression
by: Butterworth, James, et al.
Published: (2026)
by: Butterworth, James, et al.
Published: (2026)
Faster Predictive Coding Networks via Better Initialization
by: Pinchetti, Luca, et al.
Published: (2026)
by: Pinchetti, Luca, et al.
Published: (2026)
CASSANDRA: Programmatic and Probabilistic Learning and Inference for Stochastic World Modeling
by: Lymperopoulos, Panagiotis, et al.
Published: (2026)
by: Lymperopoulos, Panagiotis, et al.
Published: (2026)
TensorLens: End-to-End Transformer Analysis via High-Order Attention Tensors
by: Atad, Ido Andrew, et al.
Published: (2026)
by: Atad, Ido Andrew, et al.
Published: (2026)
A Simple Approximate Bayesian Inference Neural Surrogate for Stochastic Petri Net Models
by: Manu, Bright Kwaku, et al.
Published: (2025)
by: Manu, Bright Kwaku, et al.
Published: (2025)
BOND: License to Train with Black-Box Functions
by: Clark, Andrew, et al.
Published: (2025)
by: Clark, Andrew, et al.
Published: (2025)
CPT: Competence-progressive Training Strategy for Few-shot Node Classification
by: Yan, Qilong, et al.
Published: (2024)
by: Yan, Qilong, et al.
Published: (2024)
Fine-grained Attention in Hierarchical Transformers for Tabular Time-series
by: Azorin, Raphael, et al.
Published: (2024)
by: Azorin, Raphael, et al.
Published: (2024)
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
by: Hedar, Abdel-Rahman, et al.
Published: (2024)
The Final-Stage Bottleneck: A Systematic Dissection of the R-Learner for Network Causal Inference
by: Sairam, S, et al.
Published: (2025)
by: Sairam, S, et al.
Published: (2025)
Adaptive Epsilon Adversarial Training for Robust Gravitational Wave Parameter Estimation Using Normalizing Flows
by: Yang, Yiqian, et al.
Published: (2024)
by: Yang, Yiqian, et al.
Published: (2024)
Chunked TabPFN: Exact Training-Free In-Context Learning for Long-Context Tabular Data
by: Sergazinov, Renat, et al.
Published: (2025)
by: Sergazinov, Renat, et al.
Published: (2025)
Normalization Layer Per-Example Gradients are Sufficient to Predict Gradient Noise Scale in Transformers
by: Gray, Gavia, et al.
Published: (2024)
by: Gray, Gavia, et al.
Published: (2024)
Self-Expanding Neural Networks
by: Mitchell, Rupert, et al.
Published: (2023)
by: Mitchell, Rupert, et al.
Published: (2023)
Learning To Help: Training Models to Assist Legacy Devices
by: Wu, Yu, et al.
Published: (2024)
by: Wu, Yu, et al.
Published: (2024)
Versatile Ordering Network: An Attention-based Neural Network for Ordering Across Scales and Quality Metrics
by: Yu, Zehua, et al.
Published: (2024)
by: Yu, Zehua, et al.
Published: (2024)
Discrete Latent Structure in Neural Networks
by: Niculae, Vlad, et al.
Published: (2023)
by: Niculae, Vlad, et al.
Published: (2023)
HEHRGNN: A Unified Embedding Model for Knowledge Graphs with Hyperedges and Hyper-Relational Edges
by: Rajagopalamenon, Rajesh, et al.
Published: (2026)
by: Rajagopalamenon, Rajesh, et al.
Published: (2026)
Implicit Regularization and Generalization in Overparameterized Neural Networks
by: Johannsen, Zeran
Published: (2026)
by: Johannsen, Zeran
Published: (2026)
Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability
by: Bakish, Yarden, et al.
Published: (2025)
by: Bakish, Yarden, et al.
Published: (2025)
Strengthening the Internal Adversarial Robustness in Lifted Neural Networks
by: Zach, Christopher
Published: (2025)
by: Zach, Christopher
Published: (2025)
The Bayesian Confidence (BACON) Estimator for Deep Neural Networks
by: Kee, Patrick D., et al.
Published: (2024)
by: Kee, Patrick D., et al.
Published: (2024)
Zero-Shot Context Generalization in Reinforcement Learning from Few Training Contexts
by: Chapman, James, et al.
Published: (2025)
by: Chapman, James, et al.
Published: (2025)
Newtonian and Lagrangian Neural Networks: A Comparison Towards Efficient Inverse Dynamics Identification
by: Trinh, Minh, et al.
Published: (2025)
by: Trinh, Minh, et al.
Published: (2025)
Expressivity of Graph Neural Networks Through the Lens of Adversarial Robustness
by: Campi, Francesco, et al.
Published: (2023)
by: Campi, Francesco, et al.
Published: (2023)
Learning Useful Representations of Recurrent Neural Network Weight Matrices
by: Herrmann, Vincent, et al.
Published: (2024)
by: Herrmann, Vincent, et al.
Published: (2024)
Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales
by: Salfati, Samuel
Published: (2026)
by: Salfati, Samuel
Published: (2026)
Real-Time Pulsatile Flow Prediction for Realistic, Diverse Intracranial Aneurysm Morphologies using a Graph Transformer and Steady-Flow Data Augmentation
by: Sheng, Yiying, et al.
Published: (2026)
by: Sheng, Yiying, et al.
Published: (2026)
Robustness of Spatio-temporal Graph Neural Networks for Fault Location in Partially Observable Distribution Grids
by: Karabulut, Burak, et al.
Published: (2026)
by: Karabulut, Burak, et al.
Published: (2026)
Training Artificial Neural Networks by Coordinate Search Algorithm
by: Rokhsatyazdi, Ehsan, et al.
Published: (2024)
by: Rokhsatyazdi, Ehsan, et al.
Published: (2024)
Playing Hex and Counter Wargames using Reinforcement Learning and Recurrent Neural Networks
by: Palma, Guilherme, et al.
Published: (2025)
by: Palma, Guilherme, et al.
Published: (2025)
Almost Equivariance via Lie Algebra Convolutions
by: McNeela, Daniel
Published: (2023)
by: McNeela, Daniel
Published: (2023)
Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
by: Moghadasi, Mahdi Naser, et al.
Published: (2026)
by: Moghadasi, Mahdi Naser, et al.
Published: (2026)
Explainable Graph Representation Learning via Graph Pattern Analysis
by: Wang, Xudong, et al.
Published: (2025)
by: Wang, Xudong, et al.
Published: (2025)
Similar Items
-
FSC-Net: Fast-Slow Consolidation Networks for Continual Learning
by: Gorrim, Mohamed El
Published: (2025) -
Inference-Time Machine Unlearning via Gated Activation Redirection
by: Turani, Vinícius Conte, et al.
Published: (2026) -
Fusing Rewards and Preferences in Reinforcement Learning
by: Khorasani, Sadegh, et al.
Published: (2025) -
HGCN(O): A Self-Tuning GCN HyperModel Toolkit for Outcome Prediction in Event-Sequence Data
by: Wang, Fang, et al.
Published: (2025) -
ReBoot: Encrypted Training of Deep Neural Networks with CKKS Bootstrapping
by: Pirillo, Alberto, et al.
Published: (2025)