Beyond Softmax: A Natural Parameterization for Categorical Random Variables
Fuente:
arXiv
Saved in:
| Main Authors: | Manenti, Alessandro, Alippi, Cesare |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Latent Graph Structures and their Uncertainty
by: Manenti, Alessandro, et al.
Published: (2024)
by: Manenti, Alessandro, et al.
Published: (2024)
SWING: Unlocking Implicit Graph Representations for Graph Random Features
by: Manenti, Alessandro, et al.
Published: (2026)
by: Manenti, Alessandro, et al.
Published: (2026)
Assessment of Spatio-Temporal Predictors in the Presence of Missing and Heterogeneous Data
by: Zambon, Daniele, et al.
Published: (2023)
by: Zambon, Daniele, et al.
Published: (2023)
The Unreasonable Effectiveness of Randomized Representations in Online Continual Graph Learning
by: Donghi, Giovanni, et al.
Published: (2025)
by: Donghi, Giovanni, et al.
Published: (2025)
Hierarchical Message-Passing Policies for Multi-Agent Reinforcement Learning
by: Marzi, Tommaso, et al.
Published: (2025)
by: Marzi, Tommaso, et al.
Published: (2025)
Graph State-Space Models and Latent Relational Inference
by: Zambon, Daniele, et al.
Published: (2023)
by: Zambon, Daniele, et al.
Published: (2023)
Graph-based Time Series Clustering for End-to-End Hierarchical Forecasting
by: Cini, Andrea, et al.
Published: (2023)
by: Cini, Andrea, et al.
Published: (2023)
Beyond Softmax and Entropy: Convergence Rates of Policy Gradients with f-SoftArgmax Parameterization & Coupled Regularization
by: Labbi, Safwan, et al.
Published: (2026)
by: Labbi, Safwan, et al.
Published: (2026)
Position: Current Benchmarking Hinders Real Progress in Deep Learning for Time Series Forecasting
by: Moretti, Valentina, et al.
Published: (2025)
by: Moretti, Valentina, et al.
Published: (2025)
Temporal Graph ODEs for Irregularly-Sampled Time Series
by: Gravina, Alessio, et al.
Published: (2024)
by: Gravina, Alessio, et al.
Published: (2024)
Change Detection in Graph Streams by Learning Graph Embeddings on Constant-Curvature Manifolds
by: Grattarola, Daniele, et al.
Published: (2018)
by: Grattarola, Daniele, et al.
Published: (2018)
Feudal Graph Reinforcement Learning
by: Marzi, Tommaso, et al.
Published: (2023)
by: Marzi, Tommaso, et al.
Published: (2023)
Graph-based Forecasting with Missing Data through Spatiotemporal Downsampling
by: Marisca, Ivan, et al.
Published: (2024)
by: Marisca, Ivan, et al.
Published: (2024)
DRAN: A Distribution and Relation Adaptive Network for Spatio-temporal Forecasting
by: Zou, Xiaobei, et al.
Published: (2025)
by: Zou, Xiaobei, et al.
Published: (2025)
FX-DARTS: Designing Topology-unconstrained Architectures with Differentiable Architecture Search and Entropy-based Super-network Shrinking
by: Rao, Xuan, et al.
Published: (2025)
by: Rao, Xuan, et al.
Published: (2025)
Graph Deep Learning for Time Series Forecasting
by: Cini, Andrea, et al.
Published: (2023)
by: Cini, Andrea, et al.
Published: (2023)
Online Continual Graph Learning
by: Donghi, Giovanni, et al.
Published: (2025)
by: Donghi, Giovanni, et al.
Published: (2025)
Over-squashing in Spatiotemporal Graph Neural Networks
by: Marisca, Ivan, et al.
Published: (2025)
by: Marisca, Ivan, et al.
Published: (2025)
Understanding Pooling in Graph Neural Networks
by: Grattarola, Daniele, et al.
Published: (2021)
by: Grattarola, Daniele, et al.
Published: (2021)
Hierarchical Representation Learning in Graph Neural Networks with Node Decimation Pooling
by: Bianchi, Filippo Maria, et al.
Published: (2019)
by: Bianchi, Filippo Maria, et al.
Published: (2019)
On the Regularization of Learnable Embeddings for Time Series Forecasting
by: Butera, Luca, et al.
Published: (2024)
by: Butera, Luca, et al.
Published: (2024)
Why Do Time Series Models Need Long Context Windows?
by: Butera, Luca, et al.
Published: (2026)
by: Butera, Luca, et al.
Published: (2026)
Object-Centric Relational Representations for Image Generation
by: Butera, Luca, et al.
Published: (2023)
by: Butera, Luca, et al.
Published: (2023)
Convergence Rates for Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2025)
by: Nguyen, Huy, et al.
Published: (2025)
A Hybrid Active-Passive Approach to Imbalanced Nonstationary Data Stream Classification
by: Malialis, Kleanthis, et al.
Published: (2022)
by: Malialis, Kleanthis, et al.
Published: (2022)
Beyond Softmax: A New Perspective on Gradient Bandits
by: Melo, Emerson, et al.
Published: (2025)
by: Melo, Emerson, et al.
Published: (2025)
On Least Square Estimation in Softmax Gating Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
PIF: Anomaly detection via preference embedding
by: Leveni, Filippo, et al.
Published: (2025)
by: Leveni, Filippo, et al.
Published: (2025)
Hashing for Structure-based Anomaly Detection
by: Leveni, Filippo, et al.
Published: (2025)
by: Leveni, Filippo, et al.
Published: (2025)
Preference Isolation Forest for Structure-based Anomaly Detection
by: Leveni, Filippo, et al.
Published: (2025)
by: Leveni, Filippo, et al.
Published: (2025)
PeakWeather: MeteoSwiss Weather Station Measurements for Spatiotemporal Deep Learning
by: Zambon, Daniele, et al.
Published: (2025)
by: Zambon, Daniele, et al.
Published: (2025)
Optimal Categorical Instrumental Variables
by: Wiemann, Thomas
Published: (2023)
by: Wiemann, Thomas
Published: (2023)
Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant
by: Lv, Qi, et al.
Published: (2025)
by: Lv, Qi, et al.
Published: (2025)
Relational Conformal Prediction for Correlated Time Series
by: Cini, Andrea, et al.
Published: (2025)
by: Cini, Andrea, et al.
Published: (2025)
Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning
by: Xie, Zixuan, et al.
Published: (2026)
by: Xie, Zixuan, et al.
Published: (2026)
On Expert Estimation in Hierarchical Mixture of Experts: Beyond Softmax Gating Functions
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts
by: Yan, Fanqi, et al.
Published: (2025)
by: Yan, Fanqi, et al.
Published: (2025)
Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of Experts
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
Graph-based Virtual Sensing from Sparse and Partial Multivariate Observations
by: De Felice, Giovanni, et al.
Published: (2024)
by: De Felice, Giovanni, et al.
Published: (2024)
Is Softmax Loss All You Need? A Principled Analysis of Softmax-family Loss
by: Pu, Yuanhao, et al.
Published: (2026)
by: Pu, Yuanhao, et al.
Published: (2026)
Similar Items
-
Learning Latent Graph Structures and their Uncertainty
by: Manenti, Alessandro, et al.
Published: (2024) -
SWING: Unlocking Implicit Graph Representations for Graph Random Features
by: Manenti, Alessandro, et al.
Published: (2026) -
Assessment of Spatio-Temporal Predictors in the Presence of Missing and Heterogeneous Data
by: Zambon, Daniele, et al.
Published: (2023) -
The Unreasonable Effectiveness of Randomized Representations in Online Continual Graph Learning
by: Donghi, Giovanni, et al.
Published: (2025) -
Hierarchical Message-Passing Policies for Multi-Agent Reinforcement Learning
by: Marzi, Tommaso, et al.
Published: (2025)