MotherNet: Fast Training and Inference via Hyper-Network Transformers
Fuente:
arXiv
Guardado en:
| Autores principales: | Müller, Andreas, Curino, Carlo, Ramakrishnan, Raghu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FSC-Net: Fast-Slow Consolidation Networks for Continual Learning
por: Gorrim, Mohamed El
Publicado: (2025)
por: Gorrim, Mohamed El
Publicado: (2025)
Inference-Time Machine Unlearning via Gated Activation Redirection
por: Turani, Vinícius Conte, et al.
Publicado: (2026)
por: Turani, Vinícius Conte, et al.
Publicado: (2026)
Fusing Rewards and Preferences in Reinforcement Learning
por: Khorasani, Sadegh, et al.
Publicado: (2025)
por: Khorasani, Sadegh, et al.
Publicado: (2025)
HGCN(O): A Self-Tuning GCN HyperModel Toolkit for Outcome Prediction in Event-Sequence Data
por: Wang, Fang, et al.
Publicado: (2025)
por: Wang, Fang, et al.
Publicado: (2025)
ReBoot: Encrypted Training of Deep Neural Networks with CKKS Bootstrapping
por: Pirillo, Alberto, et al.
Publicado: (2025)
por: Pirillo, Alberto, et al.
Publicado: (2025)
New Paradigm of Adversarial Training: Releasing Accuracy-Robustness Trade-Off via Dummy Class
por: Wang, Yanyun, et al.
Publicado: (2024)
por: Wang, Yanyun, et al.
Publicado: (2024)
Deep Variational Inference Symbolic Regression
por: Butterworth, James, et al.
Publicado: (2026)
por: Butterworth, James, et al.
Publicado: (2026)
Faster Predictive Coding Networks via Better Initialization
por: Pinchetti, Luca, et al.
Publicado: (2026)
por: Pinchetti, Luca, et al.
Publicado: (2026)
CASSANDRA: Programmatic and Probabilistic Learning and Inference for Stochastic World Modeling
por: Lymperopoulos, Panagiotis, et al.
Publicado: (2026)
por: Lymperopoulos, Panagiotis, et al.
Publicado: (2026)
TensorLens: End-to-End Transformer Analysis via High-Order Attention Tensors
por: Atad, Ido Andrew, et al.
Publicado: (2026)
por: Atad, Ido Andrew, et al.
Publicado: (2026)
A Simple Approximate Bayesian Inference Neural Surrogate for Stochastic Petri Net Models
por: Manu, Bright Kwaku, et al.
Publicado: (2025)
por: Manu, Bright Kwaku, et al.
Publicado: (2025)
BOND: License to Train with Black-Box Functions
por: Clark, Andrew, et al.
Publicado: (2025)
por: Clark, Andrew, et al.
Publicado: (2025)
CPT: Competence-progressive Training Strategy for Few-shot Node Classification
por: Yan, Qilong, et al.
Publicado: (2024)
por: Yan, Qilong, et al.
Publicado: (2024)
Fine-grained Attention in Hierarchical Transformers for Tabular Time-series
por: Azorin, Raphael, et al.
Publicado: (2024)
por: Azorin, Raphael, et al.
Publicado: (2024)
Deep Memory Search: A Metaheuristic Approach for Optimizing Heuristic Search
por: Hedar, Abdel-Rahman, et al.
Publicado: (2024)
por: Hedar, Abdel-Rahman, et al.
Publicado: (2024)
The Final-Stage Bottleneck: A Systematic Dissection of the R-Learner for Network Causal Inference
por: Sairam, S, et al.
Publicado: (2025)
por: Sairam, S, et al.
Publicado: (2025)
Adaptive Epsilon Adversarial Training for Robust Gravitational Wave Parameter Estimation Using Normalizing Flows
por: Yang, Yiqian, et al.
Publicado: (2024)
por: Yang, Yiqian, et al.
Publicado: (2024)
Chunked TabPFN: Exact Training-Free In-Context Learning for Long-Context Tabular Data
por: Sergazinov, Renat, et al.
Publicado: (2025)
por: Sergazinov, Renat, et al.
Publicado: (2025)
Normalization Layer Per-Example Gradients are Sufficient to Predict Gradient Noise Scale in Transformers
por: Gray, Gavia, et al.
Publicado: (2024)
por: Gray, Gavia, et al.
Publicado: (2024)
Self-Expanding Neural Networks
por: Mitchell, Rupert, et al.
Publicado: (2023)
por: Mitchell, Rupert, et al.
Publicado: (2023)
Learning To Help: Training Models to Assist Legacy Devices
por: Wu, Yu, et al.
Publicado: (2024)
por: Wu, Yu, et al.
Publicado: (2024)
Versatile Ordering Network: An Attention-based Neural Network for Ordering Across Scales and Quality Metrics
por: Yu, Zehua, et al.
Publicado: (2024)
por: Yu, Zehua, et al.
Publicado: (2024)
Discrete Latent Structure in Neural Networks
por: Niculae, Vlad, et al.
Publicado: (2023)
por: Niculae, Vlad, et al.
Publicado: (2023)
HEHRGNN: A Unified Embedding Model for Knowledge Graphs with Hyperedges and Hyper-Relational Edges
por: Rajagopalamenon, Rajesh, et al.
Publicado: (2026)
por: Rajagopalamenon, Rajesh, et al.
Publicado: (2026)
Implicit Regularization and Generalization in Overparameterized Neural Networks
por: Johannsen, Zeran
Publicado: (2026)
por: Johannsen, Zeran
Publicado: (2026)
Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability
por: Bakish, Yarden, et al.
Publicado: (2025)
por: Bakish, Yarden, et al.
Publicado: (2025)
Strengthening the Internal Adversarial Robustness in Lifted Neural Networks
por: Zach, Christopher
Publicado: (2025)
por: Zach, Christopher
Publicado: (2025)
The Bayesian Confidence (BACON) Estimator for Deep Neural Networks
por: Kee, Patrick D., et al.
Publicado: (2024)
por: Kee, Patrick D., et al.
Publicado: (2024)
Zero-Shot Context Generalization in Reinforcement Learning from Few Training Contexts
por: Chapman, James, et al.
Publicado: (2025)
por: Chapman, James, et al.
Publicado: (2025)
Newtonian and Lagrangian Neural Networks: A Comparison Towards Efficient Inverse Dynamics Identification
por: Trinh, Minh, et al.
Publicado: (2025)
por: Trinh, Minh, et al.
Publicado: (2025)
Expressivity of Graph Neural Networks Through the Lens of Adversarial Robustness
por: Campi, Francesco, et al.
Publicado: (2023)
por: Campi, Francesco, et al.
Publicado: (2023)
Learning Useful Representations of Recurrent Neural Network Weight Matrices
por: Herrmann, Vincent, et al.
Publicado: (2024)
por: Herrmann, Vincent, et al.
Publicado: (2024)
Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales
por: Salfati, Samuel
Publicado: (2026)
por: Salfati, Samuel
Publicado: (2026)
Real-Time Pulsatile Flow Prediction for Realistic, Diverse Intracranial Aneurysm Morphologies using a Graph Transformer and Steady-Flow Data Augmentation
por: Sheng, Yiying, et al.
Publicado: (2026)
por: Sheng, Yiying, et al.
Publicado: (2026)
Robustness of Spatio-temporal Graph Neural Networks for Fault Location in Partially Observable Distribution Grids
por: Karabulut, Burak, et al.
Publicado: (2026)
por: Karabulut, Burak, et al.
Publicado: (2026)
Training Artificial Neural Networks by Coordinate Search Algorithm
por: Rokhsatyazdi, Ehsan, et al.
Publicado: (2024)
por: Rokhsatyazdi, Ehsan, et al.
Publicado: (2024)
Playing Hex and Counter Wargames using Reinforcement Learning and Recurrent Neural Networks
por: Palma, Guilherme, et al.
Publicado: (2025)
por: Palma, Guilherme, et al.
Publicado: (2025)
Almost Equivariance via Lie Algebra Convolutions
por: McNeela, Daniel
Publicado: (2023)
por: McNeela, Daniel
Publicado: (2023)
Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
por: Moghadasi, Mahdi Naser, et al.
Publicado: (2026)
por: Moghadasi, Mahdi Naser, et al.
Publicado: (2026)
Explainable Graph Representation Learning via Graph Pattern Analysis
por: Wang, Xudong, et al.
Publicado: (2025)
por: Wang, Xudong, et al.
Publicado: (2025)
Ejemplares similares
-
FSC-Net: Fast-Slow Consolidation Networks for Continual Learning
por: Gorrim, Mohamed El
Publicado: (2025) -
Inference-Time Machine Unlearning via Gated Activation Redirection
por: Turani, Vinícius Conte, et al.
Publicado: (2026) -
Fusing Rewards and Preferences in Reinforcement Learning
por: Khorasani, Sadegh, et al.
Publicado: (2025) -
HGCN(O): A Self-Tuning GCN HyperModel Toolkit for Outcome Prediction in Event-Sequence Data
por: Wang, Fang, et al.
Publicado: (2025) -
ReBoot: Encrypted Training of Deep Neural Networks with CKKS Bootstrapping
por: Pirillo, Alberto, et al.
Publicado: (2025)