When Bias Meets Trainability: Connecting Theories of Initialization
Fuente:
arXiv
Saved in:
| Main Authors: | Bassi, Alberto, Baity-Jesi, Marco, Lucchi, Aurelien, Albert, Carlo, Francazi, Emanuele |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Initial Guessing Bias: How Untrained Networks Favor Some Classes
by: Francazi, Emanuele, et al.
Published: (2023)
by: Francazi, Emanuele, et al.
Published: (2023)
A Theoretical Analysis of the Learning Dynamics under Class Imbalance
by: Francazi, Emanuele, et al.
Published: (2022)
by: Francazi, Emanuele, et al.
Published: (2022)
Where You Place the Norm Matters: From Prejudiced to Neutral Initializations
by: Francazi, Emanuele, et al.
Published: (2025)
by: Francazi, Emanuele, et al.
Published: (2025)
Producing Plankton Classifiers that are Robust to Dataset Shift
by: Chen, Cheng, et al.
Published: (2024)
by: Chen, Cheng, et al.
Published: (2024)
Optimization Guarantees for Square-Root Natural-Gradient Variational Inference
by: Kumar, Navish, et al.
Published: (2025)
by: Kumar, Navish, et al.
Published: (2025)
Conjugate Learning Theory: Uncovering the Mechanisms of Trainability and Generalization in Deep Neural Networks
by: Qi, Binchuan
Published: (2026)
by: Qi, Binchuan
Published: (2026)
SageBwd: A Trainable Low-bit Attention
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
A Unified Noise-Curvature View of Loss of Trainability
by: Baveja, Gunbir Singh, et al.
Published: (2025)
by: Baveja, Gunbir Singh, et al.
Published: (2025)
The Majority Vote Paradigm Shift: When Popular Meets Optimal
by: Purificato, Antonio, et al.
Published: (2025)
by: Purificato, Antonio, et al.
Published: (2025)
Agentic Unlearning: When LLM Agent Meets Machine Unlearning
by: Wang, Bin, et al.
Published: (2026)
by: Wang, Bin, et al.
Published: (2026)
When Drafts Evolve: Speculative Decoding Meets Online Learning
by: Qian, Yu-Yang, et al.
Published: (2026)
by: Qian, Yu-Yang, et al.
Published: (2026)
Teasing Apart Architecture and Initial Weights as Sources of Inductive Bias in Neural Networks
by: Bencomo, Gianluca, et al.
Published: (2025)
by: Bencomo, Gianluca, et al.
Published: (2025)
Natively Trainable Sparse Attention for Hierarchical Point Cloud Datasets
by: Lapautre, Nicolas, et al.
Published: (2025)
by: Lapautre, Nicolas, et al.
Published: (2025)
Playing the Lottery With Concave Regularizers for Sparse Trainable Neural Networks
by: Fracastoro, Giulia, et al.
Published: (2025)
by: Fracastoro, Giulia, et al.
Published: (2025)
Scalable Decision Focused Learning via Online Trainable Surrogates
by: Signorelli, Gaetano, et al.
Published: (2025)
by: Signorelli, Gaetano, et al.
Published: (2025)
On the Trainability of Masked Diffusion Language Models via Blockwise Locality
by: Wang, Yuxiang, et al.
Published: (2026)
by: Wang, Yuxiang, et al.
Published: (2026)
Self-Attribution Bias: When AI Monitors Go Easy on Themselves
by: Khullar, Dipika, et al.
Published: (2026)
by: Khullar, Dipika, et al.
Published: (2026)
When Continue Learning Meets Multimodal Large Language Model: A Survey
by: Huo, Yukang, et al.
Published: (2025)
by: Huo, Yukang, et al.
Published: (2025)
Trainable Dynamic Mask Sparse Attention
by: Shi, Jingze, et al.
Published: (2025)
by: Shi, Jingze, et al.
Published: (2025)
RewriteNets: End-to-End Trainable String-Rewriting for Generative Sequence Modeling
by: Vejendla, Harshil
Published: (2026)
by: Vejendla, Harshil
Published: (2026)
Gradients: When Markets Meet Fine-tuning -- A Distributed Approach to Model Optimisation
by: Subia-Waud, Christopher
Published: (2025)
by: Subia-Waud, Christopher
Published: (2025)
Decoupling Spatio-Temporal Prediction: When Lightweight Large Models Meet Adaptive Hypergraphs
by: Chen, Jiawen, et al.
Published: (2025)
by: Chen, Jiawen, et al.
Published: (2025)
When Reasoning Meets Compression: Understanding the Effects of LLMs Compression on Large Reasoning Models
by: Zhang, Nan, et al.
Published: (2025)
by: Zhang, Nan, et al.
Published: (2025)
When Does Closeness in Distribution Imply Representational Similarity? An Identifiability Perspective
by: Nielsen, Beatrix M. G., et al.
Published: (2025)
by: Nielsen, Beatrix M. G., et al.
Published: (2025)
Trainable and Explainable Simplicial Map Neural Networks
by: Paluzo-Hidalgo, Eduardo, et al.
Published: (2023)
by: Paluzo-Hidalgo, Eduardo, et al.
Published: (2023)
When Graph Neural Network Meets Causality: Opportunities, Methodologies and An Outlook
by: Jiang, Wenzhao, et al.
Published: (2023)
by: Jiang, Wenzhao, et al.
Published: (2023)
Mapping the Edge of Chaos: Fractal-Like Boundaries in The Trainability of Decoder-Only Transformer Models
by: Torkamandi, Bahman
Published: (2025)
by: Torkamandi, Bahman
Published: (2025)
When Model Meets New Normals: Test-time Adaptation for Unsupervised Time-series Anomaly Detection
by: Kim, Dongmin, et al.
Published: (2023)
by: Kim, Dongmin, et al.
Published: (2023)
Projected Compression: Trainable Projection for Efficient Transformer Compression
by: Stefaniak, Maciej, et al.
Published: (2025)
by: Stefaniak, Maciej, et al.
Published: (2025)
Sharp Generalization Bounds for Foundation Models with Asymmetric Randomized Low-Rank Adapters
by: Kratsios, Anastasis, et al.
Published: (2025)
by: Kratsios, Anastasis, et al.
Published: (2025)
Reinforcement Learning Jazz Improvisation: When Music Meets Game Theory
by: Tapiavala, Vedant, et al.
Published: (2024)
by: Tapiavala, Vedant, et al.
Published: (2024)
It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
by: Behrouz, Ali, et al.
Published: (2025)
by: Behrouz, Ali, et al.
Published: (2025)
HATA: Trainable and Hardware-Efficient Hash-Aware Top-k Attention for Scalable Large Model Inference
by: Gong, Ping, et al.
Published: (2025)
by: Gong, Ping, et al.
Published: (2025)
When Noisy Labels Meet Class Imbalance on Graphs: A Graph Augmentation Method with LLM and Pseudo Label
by: Xia, Riting, et al.
Published: (2025)
by: Xia, Riting, et al.
Published: (2025)
When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference
by: Ye, Wen, et al.
Published: (2025)
by: Ye, Wen, et al.
Published: (2025)
When Demonstrations Meet Generative World Models: A Maximum Likelihood Framework for Offline Inverse Reinforcement Learning
by: Zeng, Siliang, et al.
Published: (2023)
by: Zeng, Siliang, et al.
Published: (2023)
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
by: Ramesh, Shyam Sundhar, et al.
Published: (2026)
by: Ramesh, Shyam Sundhar, et al.
Published: (2026)
Weightless Neural Networks for Continuously Trainable Personalized Recommendation Systems
by: Latif, Rafayel, et al.
Published: (2025)
by: Latif, Rafayel, et al.
Published: (2025)
Class Imbalance in Anomaly Detection: Learning from an Exactly Solvable Model
by: Pezzicoli, F. S., et al.
Published: (2025)
by: Pezzicoli, F. S., et al.
Published: (2025)
When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs
by: Jeong, Soyeong, et al.
Published: (2025)
by: Jeong, Soyeong, et al.
Published: (2025)
Similar Items
-
Initial Guessing Bias: How Untrained Networks Favor Some Classes
by: Francazi, Emanuele, et al.
Published: (2023) -
A Theoretical Analysis of the Learning Dynamics under Class Imbalance
by: Francazi, Emanuele, et al.
Published: (2022) -
Where You Place the Norm Matters: From Prejudiced to Neutral Initializations
by: Francazi, Emanuele, et al.
Published: (2025) -
Producing Plankton Classifiers that are Robust to Dataset Shift
by: Chen, Cheng, et al.
Published: (2024) -
Optimization Guarantees for Square-Root Natural-Gradient Variational Inference
by: Kumar, Navish, et al.
Published: (2025)