A Geometric Analysis of Sign-Magnitude Asymmetry in a ReLU + RMSNorm Block under Ternary Quantization
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Dong, Lei |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Upper Bound of Bayesian Generalization Error in Partial Concept Bottleneck Model (CBM): Partial CBM outperforms naive CBM
par: Hayashi, Naoki, et autres
Publié: (2024)
par: Hayashi, Naoki, et autres
Publié: (2024)
Tropical Geometric Tools for Machine Learning: the TML package
par: Barnhill, David, et autres
Publié: (2023)
par: Barnhill, David, et autres
Publié: (2023)
A Complete Symmetry Classification of Shallow ReLU Networks
par: Ramakrishnan, Pranavkrishnan
Publié: (2026)
par: Ramakrishnan, Pranavkrishnan
Publié: (2026)
Tverberg's theorem and multi-class support vector machines
par: Soberón, Pablo
Publié: (2024)
par: Soberón, Pablo
Publié: (2024)
Identifiability of Deep Polynomial Neural Networks
par: Usevich, Konstantin, et autres
Publié: (2025)
par: Usevich, Konstantin, et autres
Publié: (2025)
The Selective G-Bispectrum and its Inversion: Applications to G-Invariant Networks
par: Mataigne, Simon, et autres
Publié: (2024)
par: Mataigne, Simon, et autres
Publié: (2024)
Combinatorial Regularity for Relatively Perfect Discrete Morse Gradient Vector Fields of ReLU Neural Networks
par: Brooks, Robyn, et autres
Publié: (2024)
par: Brooks, Robyn, et autres
Publié: (2024)
Diffeomorphic Measure Matching with Kernels for Generative Modeling
par: Pandey, Biraj, et autres
Publié: (2024)
par: Pandey, Biraj, et autres
Publié: (2024)
CoxKAN: Kolmogorov-Arnold Networks for Interpretable, High-Performance Survival Analysis
par: Knottenbelt, William, et autres
Publié: (2024)
par: Knottenbelt, William, et autres
Publié: (2024)
Global law of conjugate kernel random matrices with heavy-tailed weights
par: Guionnet, Alice, et autres
Publié: (2025)
par: Guionnet, Alice, et autres
Publié: (2025)
On the algorithmic construction of deep ReLU networks
par: Huybrechs, Daan
Publié: (2025)
par: Huybrechs, Daan
Publié: (2025)
Geometry and Optimization of Shallow Polynomial Networks
par: Arjevani, Yossi, et autres
Publié: (2025)
par: Arjevani, Yossi, et autres
Publié: (2025)
Estimating the Local Learning Coefficient at Scale
par: Furman, Zach, et autres
Publié: (2024)
par: Furman, Zach, et autres
Publié: (2024)
Supervised Guidance Training for Infinite-Dimensional Diffusion Models
par: Baker, Elizabeth L., et autres
Publié: (2026)
par: Baker, Elizabeth L., et autres
Publié: (2026)
The Inhibitor: ReLU and Addition-Based Attention for Efficient Transformers under Fully Homomorphic Encryption on the Torus
par: Brännvall, Rickard, et autres
Publié: (2023)
par: Brännvall, Rickard, et autres
Publié: (2023)
Sigma Flows for Image and Data Labeling and Learning Structured Prediction
par: Cassel, Jonas, et autres
Publié: (2024)
par: Cassel, Jonas, et autres
Publié: (2024)
G-Mapper: Learning a Cover in the Mapper Construction
par: Alvarado, Enrique, et autres
Publié: (2023)
par: Alvarado, Enrique, et autres
Publié: (2023)
Algebraic Machine Learning with an Application to Chemistry
par: Sai, Ezzeddine El, et autres
Publié: (2022)
par: Sai, Ezzeddine El, et autres
Publié: (2022)
Manifold Learning with Sparse Regularised Optimal Transport
par: Zhang, Stephen, et autres
Publié: (2023)
par: Zhang, Stephen, et autres
Publié: (2023)
Jacobian-Velocity Bounds for Deployment Risk Under Covariate Drift
par: Landers, Jonathan R.
Publié: (2026)
par: Landers, Jonathan R.
Publié: (2026)
An Asymptotic Equation Linking WAIC and WBIC in Singular Models
par: Hayashi, Naoki, et autres
Publié: (2025)
par: Hayashi, Naoki, et autres
Publié: (2025)
Near-optimal estimates for the $\ell^p$-Lipschitz constants of deep random ReLU neural networks
par: Dirksen, Sjoerd, et autres
Publié: (2025)
par: Dirksen, Sjoerd, et autres
Publié: (2025)
Covering Numbers for Deep ReLU Networks with Applications to Function Approximation and Nonparametric Regression
par: Ou, Weigutian, et autres
Publié: (2024)
par: Ou, Weigutian, et autres
Publié: (2024)
Equidistribution-based training of Free Knot Splines and ReLU Neural Networks
par: Appella, Simone, et autres
Publié: (2024)
par: Appella, Simone, et autres
Publié: (2024)
Mathematical Theory of Collinearity Effects on Machine Learning Variable Importance Measures
par: Bladen, Kelvyn K., et autres
Publié: (2025)
par: Bladen, Kelvyn K., et autres
Publié: (2025)
Evaluating the Quality of the Quantified Uncertainty for (Re)Calibration of Data-Driven Regression Models
par: Wibbeke, Jelke, et autres
Publié: (2025)
par: Wibbeke, Jelke, et autres
Publié: (2025)
Algebraic Study of Discrete Imsetal Models
par: Alkeswani, Amira
Publié: (2026)
par: Alkeswani, Amira
Publié: (2026)
Distributed Sparse Linear Regression under Communication Constraints
par: Fonseca, Rodney, et autres
Publié: (2023)
par: Fonseca, Rodney, et autres
Publié: (2023)
The Local Learning Coefficient: A Singularity-Aware Complexity Measure
par: Lau, Edmund, et autres
Publié: (2023)
par: Lau, Edmund, et autres
Publié: (2023)
Bayesian ICA with super-Gaussian Source Priors
par: Datta, Jyotishka, et autres
Publié: (2024)
par: Datta, Jyotishka, et autres
Publié: (2024)
Is ReLU Adversarially Robust?
par: Sooksatra, Korn, et autres
Publié: (2024)
par: Sooksatra, Korn, et autres
Publié: (2024)
Evaluation of the impact of expert knowledge: How decision support scores impact the effectiveness of automatic knowledge-driven feature engineering (aKDFE)
par: Björneld, Olof, et autres
Publié: (2025)
par: Björneld, Olof, et autres
Publié: (2025)
Unraveling Media Perspectives: A Comprehensive Methodology Combining Large Language Models, Topic Modeling, Sentiment Analysis, and Ontology Learning to Analyse Media Bias
par: Jähde, Orlando, et autres
Publié: (2025)
par: Jähde, Orlando, et autres
Publié: (2025)
Scalable non-separable spatio-temporal Gaussian process models for large-scale short-term weather prediction
par: Gyger, Tim, et autres
Publié: (2026)
par: Gyger, Tim, et autres
Publié: (2026)
The Positivity of the Neural Tangent Kernel
par: Carvalho, Luís, et autres
Publié: (2024)
par: Carvalho, Luís, et autres
Publié: (2024)
Deep Adaptive Dimension Reduction for Bayesian Inference in Inverse Problems
par: Wang, Yueyang, et autres
Publié: (2026)
par: Wang, Yueyang, et autres
Publié: (2026)
Comparison of Machine Learning Classification Algorithms and Application to the Framingham Heart Study
par: Kahouadji, Nabil
Publié: (2024)
par: Kahouadji, Nabil
Publié: (2024)
Tropical toric maximum likelihood estimation
par: Boniface, Emma, et autres
Publié: (2024)
par: Boniface, Emma, et autres
Publié: (2024)
Boosted generalized normal distributions: Integrating machine learning with operations knowledge
par: Gurlek, Ragip, et autres
Publié: (2024)
par: Gurlek, Ragip, et autres
Publié: (2024)
Multilevel Sampling in Algebraic Statistics
par: Kirk, Nathan, et autres
Publié: (2025)
par: Kirk, Nathan, et autres
Publié: (2025)
Documents similaires
-
Upper Bound of Bayesian Generalization Error in Partial Concept Bottleneck Model (CBM): Partial CBM outperforms naive CBM
par: Hayashi, Naoki, et autres
Publié: (2024) -
Tropical Geometric Tools for Machine Learning: the TML package
par: Barnhill, David, et autres
Publié: (2023) -
A Complete Symmetry Classification of Shallow ReLU Networks
par: Ramakrishnan, Pranavkrishnan
Publié: (2026) -
Tverberg's theorem and multi-class support vector machines
par: Soberón, Pablo
Publié: (2024) -
Identifiability of Deep Polynomial Neural Networks
par: Usevich, Konstantin, et autres
Publié: (2025)