Improved Generalization of Weight Space Networks via Augmentations
Fuente:
arXiv
Saved in:
| Main Authors: | Shamsian, Aviv, Navon, Aviv, Zhang, David W., Zhang, Yan, Fetaya, Ethan, Chechik, Gal, Maron, Haggai |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Equivariant Deep Weight Space Alignment
by: Navon, Aviv, et al.
Published: (2023)
by: Navon, Aviv, et al.
Published: (2023)
Go Beyond Your Means: Unlearning with Per-Sample Gradient Orthogonalization
by: Shamsian, Aviv, et al.
Published: (2025)
by: Shamsian, Aviv, et al.
Published: (2025)
PromptEvolver: Prompt Inversion through Evolutionary Optimization in Natural-Language Space
by: Buchnick, Asaf, et al.
Published: (2026)
by: Buchnick, Asaf, et al.
Published: (2026)
GradMetaNet: An Equivariant Architecture for Learning on Gradients
by: Gelberg, Yoav, et al.
Published: (2025)
by: Gelberg, Yoav, et al.
Published: (2025)
Multi Task Inverse Reinforcement Learning for Common Sense Reward
by: Glazer, Neta, et al.
Published: (2024)
by: Glazer, Neta, et al.
Published: (2024)
Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR
by: Segal-Feldman, Yael, et al.
Published: (2024)
by: Segal-Feldman, Yael, et al.
Published: (2024)
Spanning the Visual Analogy Space with a Weight Basis of LoRAs
by: Manor, Hila, et al.
Published: (2026)
by: Manor, Hila, et al.
Published: (2026)
Efficient GNN Training Through Structure-Aware Randomized Mini-Batching
by: Balaji, Vignesh, et al.
Published: (2025)
by: Balaji, Vignesh, et al.
Published: (2025)
Drax: Speech Recognition with Discrete Flow Matching
by: Navon, Aviv, et al.
Published: (2025)
by: Navon, Aviv, et al.
Published: (2025)
Keyword-Guided Adaptation of Automatic Speech Recognition
by: Shamsian, Aviv, et al.
Published: (2024)
by: Shamsian, Aviv, et al.
Published: (2024)
Beyond Transcription: Mechanistic Interpretability in ASR
by: Glazer, Neta, et al.
Published: (2025)
by: Glazer, Neta, et al.
Published: (2025)
WhisperNER: Unified Open Named Entity and Speech Recognition
by: Ayache, Gil, et al.
Published: (2024)
by: Ayache, Gil, et al.
Published: (2024)
FedSelect: Personalized Federated Learning with Customized Selection of Parameters for Fine-Tuning
by: Tamirisa, Rishub, et al.
Published: (2024)
by: Tamirisa, Rishub, et al.
Published: (2024)
A Comparison of Methods for Neural Network Aggregation
by: Pomerat, John, et al.
Published: (2023)
by: Pomerat, John, et al.
Published: (2023)
FlowTSE: Target Speaker Extraction with Flow Matching
by: Navon, Aviv, et al.
Published: (2025)
by: Navon, Aviv, et al.
Published: (2025)
Learning from Historical Activations in Graph Neural Networks
by: Galron, Yaniv, et al.
Published: (2026)
by: Galron, Yaniv, et al.
Published: (2026)
GRANOLA: Adaptive Normalization for Graph Neural Networks
by: Eliasof, Moshe, et al.
Published: (2024)
by: Eliasof, Moshe, et al.
Published: (2024)
Understanding and Improving Laplacian Positional Encodings For Temporal GNNs
by: Galron, Yaniv, et al.
Published: (2025)
by: Galron, Yaniv, et al.
Published: (2025)
On the Expressive Power of Permutation-Equivariant Weight-Space Networks
by: Dayan, Adir, et al.
Published: (2026)
by: Dayan, Adir, et al.
Published: (2026)
Conformal Prediction of Classifiers with Many Classes based on Noisy Labels
by: Penso, Coby, et al.
Published: (2025)
by: Penso, Coby, et al.
Published: (2025)
More Test-Time Compute Can Hurt: Overestimation Bias in LLM Beam Search
by: Dalal, Gal, et al.
Published: (2026)
by: Dalal, Gal, et al.
Published: (2026)
Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism
by: Bick, Aviv, et al.
Published: (2025)
by: Bick, Aviv, et al.
Published: (2025)
Meta Reinforcement Learning with Finite Training Tasks -- a Density Estimation Approach
by: Rimon, Zohar, et al.
Published: (2022)
by: Rimon, Zohar, et al.
Published: (2022)
MoGU: Mixture-of-Gaussians with Uncertainty-based Gating for Time Series Forecasting
by: Aviv, Gilad, et al.
Published: (2025)
by: Aviv, Gilad, et al.
Published: (2025)
A Graph Meta-Network for Learning on Kolmogorov-Arnold Networks
by: Bar-Shalom, Guy, et al.
Published: (2026)
by: Bar-Shalom, Guy, et al.
Published: (2026)
Graph Metanetworks for Processing Diverse Neural Architectures
by: Lim, Derek, et al.
Published: (2023)
by: Lim, Derek, et al.
Published: (2023)
Retrieval-Aware Distillation for Transformer-SSM Hybrids
by: Bick, Aviv, et al.
Published: (2026)
by: Bick, Aviv, et al.
Published: (2026)
Policy Gradient with Tree Expansion
by: Dalal, Gal, et al.
Published: (2023)
by: Dalal, Gal, et al.
Published: (2023)
The Empirical Impact of Neural Parameter Symmetries, or Lack Thereof
by: Lim, Derek, et al.
Published: (2024)
by: Lim, Derek, et al.
Published: (2024)
Bayesian Uncertainty for Gradient Aggregation in Multi-Task Learning
by: Achituve, Idan, et al.
Published: (2024)
by: Achituve, Idan, et al.
Published: (2024)
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
by: Glazer, Neta, et al.
Published: (2025)
by: Glazer, Neta, et al.
Published: (2025)
Assessing Image Quality Using a Simple Generative Representation
by: Raviv, Simon, et al.
Published: (2024)
by: Raviv, Simon, et al.
Published: (2024)
LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading
by: Yemini, Yochai, et al.
Published: (2023)
by: Yemini, Yochai, et al.
Published: (2023)
Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing
by: Bick, Aviv, et al.
Published: (2025)
by: Bick, Aviv, et al.
Published: (2025)
Joint Embedding vs Reconstruction: Provable Benefits of Latent Space Prediction for Self Supervised Learning
by: Van Assel, Hugues, et al.
Published: (2025)
by: Van Assel, Hugues, et al.
Published: (2025)
Adversarial Attacks in Weight-Space Classifiers
by: Shor, Tamir, et al.
Published: (2025)
by: Shor, Tamir, et al.
Published: (2025)
Few-Shot Task Learning through Inverse Generative Modeling
by: Netanyahu, Aviv, et al.
Published: (2024)
by: Netanyahu, Aviv, et al.
Published: (2024)
Diverse Sampling in Diffusion Models with Marginal Preserving Particle Guidance
by: Vinograd, Gal, et al.
Published: (2026)
by: Vinograd, Gal, et al.
Published: (2026)
Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models
by: Bick, Aviv, et al.
Published: (2024)
by: Bick, Aviv, et al.
Published: (2024)
Training-Free Consistent Text-to-Image Generation
by: Tewel, Yoad, et al.
Published: (2024)
by: Tewel, Yoad, et al.
Published: (2024)
Similar Items
-
Equivariant Deep Weight Space Alignment
by: Navon, Aviv, et al.
Published: (2023) -
Go Beyond Your Means: Unlearning with Per-Sample Gradient Orthogonalization
by: Shamsian, Aviv, et al.
Published: (2025) -
PromptEvolver: Prompt Inversion through Evolutionary Optimization in Natural-Language Space
by: Buchnick, Asaf, et al.
Published: (2026) -
GradMetaNet: An Equivariant Architecture for Learning on Gradients
by: Gelberg, Yoav, et al.
Published: (2025) -
Multi Task Inverse Reinforcement Learning for Common Sense Reward
by: Glazer, Neta, et al.
Published: (2024)