FAdam: Adam is a natural gradient optimizer using diagonal empirical Fisher information
Fuente:
arXiv
Saved in:
| Main Author: | Hwang, Dongseong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Simple Convergence Proof of Adam From a Sign-like Descent Perspective
by: Peng, Hanyang, et al.
Published: (2025)
by: Peng, Hanyang, et al.
Published: (2025)
WarpAdam: A new Adam optimizer based on Meta-Learning approach
by: Pan, Chengxi, et al.
Published: (2024)
by: Pan, Chengxi, et al.
Published: (2024)
Entropy-informed Decoding: Adaptive Information-Driven Branching
by: Evans, Benjamin Patrick, et al.
Published: (2026)
by: Evans, Benjamin Patrick, et al.
Published: (2026)
LZ Penalty: An information-theoretic repetition penalty for autoregressive language models
by: Ginart, Antonio A., et al.
Published: (2025)
by: Ginart, Antonio A., et al.
Published: (2025)
A Novel Double Pruning method for Imbalanced Data using Information Entropy and Roulette Wheel Selection for Breast Cancer Diagnosis
by: Bacha, Soufiane, et al.
Published: (2025)
by: Bacha, Soufiane, et al.
Published: (2025)
On the optimization dynamics of RLVR: Gradient gap and step size thresholds
by: Suk, Joe, et al.
Published: (2025)
by: Suk, Joe, et al.
Published: (2025)
Persistent Entropy as a Detector of Phase Transitions
by: Rucco, Matteo
Published: (2026)
by: Rucco, Matteo
Published: (2026)
Tackling Federated Unlearning as a Parameter Estimation Problem
by: Balordi, Antonio, et al.
Published: (2025)
by: Balordi, Antonio, et al.
Published: (2025)
Redundancy as a Structural Information Principle for Learning and Generalization
by: Bi, Yuda, et al.
Published: (2025)
by: Bi, Yuda, et al.
Published: (2025)
Contextual Control without Memory Growth in a Context-Switching Task
by: Kim, Song-Ju
Published: (2026)
by: Kim, Song-Ju
Published: (2026)
Laplace Sample Information: Data Informativeness Through a Bayesian Lens
by: Kaiser, Johannes, et al.
Published: (2025)
by: Kaiser, Johannes, et al.
Published: (2025)
Flexible Variational Information Bottleneck: Achieving Diverse Compression with a Single Training
by: Kudo, Sota, et al.
Published: (2024)
by: Kudo, Sota, et al.
Published: (2024)
Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective
by: Yang, Jiaming, et al.
Published: (2026)
by: Yang, Jiaming, et al.
Published: (2026)
Enhancing User Throughput in Multi-panel mmWave Radio Access Networks for Beam-based MU-MIMO Using a DRL Method
by: Hashemi, Ramin, et al.
Published: (2026)
by: Hashemi, Ramin, et al.
Published: (2026)
Intent-Aware DRL-Based NOMA Uplink Dynamic Scheduler for IIoT
by: Mostafa, Salwa, et al.
Published: (2024)
by: Mostafa, Salwa, et al.
Published: (2024)
Trustworthy Actionable Perturbations
by: Friedbaum, Jesse, et al.
Published: (2024)
by: Friedbaum, Jesse, et al.
Published: (2024)
Random Aggregate Beamforming for Over-the-Air Federated Learning in Large-Scale Networks
by: Xu, Chunmei, et al.
Published: (2024)
by: Xu, Chunmei, et al.
Published: (2024)
Partial Information Decomposition for Data Interpretability and Feature Selection
by: Westphal, Charles, et al.
Published: (2024)
by: Westphal, Charles, et al.
Published: (2024)
MambaJSCC: Adaptive Deep Joint Source-Channel Coding with Generalized State Space Model
by: Wu, Tong, et al.
Published: (2024)
by: Wu, Tong, et al.
Published: (2024)
Adaptive $k$-nearest neighbor classifier based on the local estimation of the shape operator
by: Levada, Alexandre Luís Magalhães, et al.
Published: (2024)
by: Levada, Alexandre Luís Magalhães, et al.
Published: (2024)
Combinatorial Multi-armed Bandits: Arm Selection via Group Testing
by: Mukherjee, Arpan, et al.
Published: (2024)
by: Mukherjee, Arpan, et al.
Published: (2024)
Hierarchical Over-the-Air Federated Learning with Awareness of Interference and Data Heterogeneity
by: Azimi-Abarghouyi, Seyed Mohammad, et al.
Published: (2024)
by: Azimi-Abarghouyi, Seyed Mohammad, et al.
Published: (2024)
Can Kernel Methods Explain How the Data Affects Neural Collapse?
by: Kothapalli, Vignesh, et al.
Published: (2024)
by: Kothapalli, Vignesh, et al.
Published: (2024)
Dynamic Multi-Network Mining of Tensor Time Series
by: Obata, Kohei, et al.
Published: (2024)
by: Obata, Kohei, et al.
Published: (2024)
ISR: Invertible Symbolic Regression
by: Tohme, Tony, et al.
Published: (2024)
by: Tohme, Tony, et al.
Published: (2024)
Understanding Transformer Architecture through Continuous Dynamics: A Partial Differential Equation Perspective
by: Zhang, Yukun, et al.
Published: (2024)
by: Zhang, Yukun, et al.
Published: (2024)
A Simple Model of Inference Scaling Laws
by: Levi, Noam
Published: (2024)
by: Levi, Noam
Published: (2024)
An Effective Information Theoretic Framework for Channel Pruning
by: Chen, Yihao, et al.
Published: (2024)
by: Chen, Yihao, et al.
Published: (2024)
WDMoE: Wireless Distributed Large Language Models with Mixture of Experts
by: Xue, Nan, et al.
Published: (2024)
by: Xue, Nan, et al.
Published: (2024)
IBB Traffic Graph Data: Benchmarking and Road Traffic Prediction Model
by: Olug, Eren, et al.
Published: (2024)
by: Olug, Eren, et al.
Published: (2024)
Semantic Text Transmission via Prediction with Small Language Models: Cost-Similarity Trade-off
by: Madhabhavi, Bhavani A, et al.
Published: (2024)
by: Madhabhavi, Bhavani A, et al.
Published: (2024)
Generalization and Informativeness of Conformal Prediction
by: Zecchin, Matteo, et al.
Published: (2024)
by: Zecchin, Matteo, et al.
Published: (2024)
Accelerating Error Correction Code Transformers
by: Levy, Matan, et al.
Published: (2024)
by: Levy, Matan, et al.
Published: (2024)
An Information Criterion for Controlled Disentanglement of Multimodal Data
by: Wang, Chenyu, et al.
Published: (2024)
by: Wang, Chenyu, et al.
Published: (2024)
Learning in Convolutional Neural Networks Accelerated by Transfer Entropy
by: Moldovan, Adrian, et al.
Published: (2024)
by: Moldovan, Adrian, et al.
Published: (2024)
Three Quantization Regimes for ReLU Networks
by: Ou, Weigutian, et al.
Published: (2024)
by: Ou, Weigutian, et al.
Published: (2024)
Provable Privacy Advantages of Decentralized Federated Learning via Distributed Optimization
by: Yu, Wenrui, et al.
Published: (2024)
by: Yu, Wenrui, et al.
Published: (2024)
Localized Adaptive Risk Control
by: Zecchin, Matteo, et al.
Published: (2024)
by: Zecchin, Matteo, et al.
Published: (2024)
Softmax is not Enough (for Sharp Size Generalisation)
by: Veličković, Petar, et al.
Published: (2024)
by: Veličković, Petar, et al.
Published: (2024)
The Causal Information Bottleneck and Optimal Causal Variable Abstractions
by: Simoes, Francisco N. F. Q., et al.
Published: (2024)
by: Simoes, Francisco N. F. Q., et al.
Published: (2024)
Similar Items
-
Simple Convergence Proof of Adam From a Sign-like Descent Perspective
by: Peng, Hanyang, et al.
Published: (2025) -
WarpAdam: A new Adam optimizer based on Meta-Learning approach
by: Pan, Chengxi, et al.
Published: (2024) -
Entropy-informed Decoding: Adaptive Information-Driven Branching
by: Evans, Benjamin Patrick, et al.
Published: (2026) -
LZ Penalty: An information-theoretic repetition penalty for autoregressive language models
by: Ginart, Antonio A., et al.
Published: (2025) -
A Novel Double Pruning method for Imbalanced Data using Information Entropy and Roulette Wheel Selection for Breast Cancer Diagnosis
by: Bacha, Soufiane, et al.
Published: (2025)