A Trainable Optimizer
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Ruiqi, Klabjan, Diego |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Divergence Results and Convergence of a Variance Reduced Version of ADAM
von: Wang, Ruiqi, et al.
Veröffentlicht: (2022)
von: Wang, Ruiqi, et al.
Veröffentlicht: (2022)
A Mirror Descent Perspective of Smoothed Sign Descent
von: Wang, Shuyang, et al.
Veröffentlicht: (2024)
von: Wang, Shuyang, et al.
Veröffentlicht: (2024)
Non-Convex Optimization with Spectral Radius Regularization
von: Sandler, Adam, et al.
Veröffentlicht: (2021)
von: Sandler, Adam, et al.
Veröffentlicht: (2021)
Rank-Accuracy Trade-off for LoRA: A Gradient-Flow Analysis
von: Rushka, Michael, et al.
Veröffentlicht: (2026)
von: Rushka, Michael, et al.
Veröffentlicht: (2026)
On the Convergence Rate of LoRA Gradient Descent
von: Mu, Siqiao, et al.
Veröffentlicht: (2025)
von: Mu, Siqiao, et al.
Veröffentlicht: (2025)
Descend or Rewind? Stochastic Gradient Descent Unlearning
von: Mu, Siqiao, et al.
Veröffentlicht: (2025)
von: Mu, Siqiao, et al.
Veröffentlicht: (2025)
On the Second-Order Convergence of Biased Policy Gradient Algorithms
von: Mu, Siqiao, et al.
Veröffentlicht: (2023)
von: Mu, Siqiao, et al.
Veröffentlicht: (2023)
IW-GAE: Importance Weighted Group Accuracy Estimation for Improved Calibration and Model Selection in Unsupervised Domain Adaptation
von: Joo, Taejong, et al.
Veröffentlicht: (2023)
von: Joo, Taejong, et al.
Veröffentlicht: (2023)
Rewind-to-Delete: Certified Machine Unlearning for Nonconvex Functions
von: Mu, Siqiao, et al.
Veröffentlicht: (2024)
von: Mu, Siqiao, et al.
Veröffentlicht: (2024)
Improving self-training under distribution shifts via anchored confidence with theoretical guarantees
von: Joo, Taejong, et al.
Veröffentlicht: (2024)
von: Joo, Taejong, et al.
Veröffentlicht: (2024)
Technical Debt in In-Context Learning: Diminishing Efficiency in Long Context
von: Joo, Taejong, et al.
Veröffentlicht: (2025)
von: Joo, Taejong, et al.
Veröffentlicht: (2025)
Communication-Efficient Federated Low-Rank Update Algorithm and its Connection to Implicit Regularization
von: Park, Haemin, et al.
Veröffentlicht: (2024)
von: Park, Haemin, et al.
Veröffentlicht: (2024)
LanFL: Differentially Private Federated Learning with Large Language Models using Synthetic Samples
von: Wu, Huiyu, et al.
Veröffentlicht: (2024)
von: Wu, Huiyu, et al.
Veröffentlicht: (2024)
Regret Bounds and Reinforcement Learning Exploration of EXP-based Algorithms
von: Xu, Mengfan, et al.
Veröffentlicht: (2020)
von: Xu, Mengfan, et al.
Veröffentlicht: (2020)
Decentralized Blockchain-based Robust Multi-agent Multi-armed Bandit
von: Xu, Mengfan, et al.
Veröffentlicht: (2024)
von: Xu, Mengfan, et al.
Veröffentlicht: (2024)
A Primal-Dual Algorithm for Hybrid Federated Learning
von: Overman, Tom, et al.
Veröffentlicht: (2022)
von: Overman, Tom, et al.
Veröffentlicht: (2022)
Federated Automated Feature Engineering
von: Overman, Tom, et al.
Veröffentlicht: (2024)
von: Overman, Tom, et al.
Veröffentlicht: (2024)
IIFE: Interaction Information Based Automated Feature Engineering
von: Overman, Tom, et al.
Veröffentlicht: (2024)
von: Overman, Tom, et al.
Veröffentlicht: (2024)
FedGA-Tree: Federated Decision Tree using Genetic Algorithm
von: Nguyen, Anh V, et al.
Veröffentlicht: (2025)
von: Nguyen, Anh V, et al.
Veröffentlicht: (2025)
Conditional Hierarchical Bayesian Tucker Decomposition for Genetic Data Analysis
von: Sandler, Adam, et al.
Veröffentlicht: (2019)
von: Sandler, Adam, et al.
Veröffentlicht: (2019)
Continuous-Time Analysis of Federated Averaging
von: Overman, Tom, et al.
Veröffentlicht: (2025)
von: Overman, Tom, et al.
Veröffentlicht: (2025)
Topic Analysis with Side Information: A Neural-Augmented LDA Approach
von: Fang, Biyi, et al.
Veröffentlicht: (2025)
von: Fang, Biyi, et al.
Veröffentlicht: (2025)
Multi-Layer Attention-Based Explainability via Transformers for Tabular Data
von: Gavito, Andrea Treviño, et al.
Veröffentlicht: (2023)
von: Gavito, Andrea Treviño, et al.
Veröffentlicht: (2023)
Video to Video Generative Adversarial Network for Few-shot Learning Based on Policy Gradient
von: Ma, Yintai, et al.
Veröffentlicht: (2024)
von: Ma, Yintai, et al.
Veröffentlicht: (2024)
Unsupervised Video Summarization via Iterative Training and Simplified GAN
von: Li, Hanqing, et al.
Veröffentlicht: (2023)
von: Li, Hanqing, et al.
Veröffentlicht: (2023)
Geometric Preconditioning and Curriculum Optimization for Trainable Variational Quantum Regression
von: Meng, Qingyu, et al.
Veröffentlicht: (2026)
von: Meng, Qingyu, et al.
Veröffentlicht: (2026)
Differentiable Calibration of Inexact Stochastic Simulation Models via Kernel Score Minimization
von: Su, Ziwei, et al.
Veröffentlicht: (2024)
von: Su, Ziwei, et al.
Veröffentlicht: (2024)
Hybrid FedGraph: An efficient hybrid federated learning algorithm using graph convolutional neural network
von: Jang, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Jang, Jaeyeon, et al.
Veröffentlicht: (2024)
Trainable Transformer in Transformer
von: Panigrahi, Abhishek, et al.
Veröffentlicht: (2023)
von: Panigrahi, Abhishek, et al.
Veröffentlicht: (2023)
A Trainable Centrality Framework for Modern Data
von: Vu, Minh Duc, et al.
Veröffentlicht: (2025)
von: Vu, Minh Duc, et al.
Veröffentlicht: (2025)
Maestro: Uncovering Low-Rank Structures via Trainable Decomposition
von: Horvath, Samuel, et al.
Veröffentlicht: (2023)
von: Horvath, Samuel, et al.
Veröffentlicht: (2023)
Survival Analysis as Imprecise Classification with Trainable Kernels
von: Konstantinov, Andrei V., et al.
Veröffentlicht: (2025)
von: Konstantinov, Andrei V., et al.
Veröffentlicht: (2025)
SageBwd: A Trainable Low-bit Attention
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
Trainable Weight Averaging: Accelerating Training and Improving Generalization
von: Li, Tao, et al.
Veröffentlicht: (2022)
von: Li, Tao, et al.
Veröffentlicht: (2022)
Trainable Bitwise Soft Quantization for Input Feature Compression
von: Schrödter, Karsten, et al.
Veröffentlicht: (2026)
von: Schrödter, Karsten, et al.
Veröffentlicht: (2026)
Trainability issues in quantum policy gradients
von: Sequeira, André, et al.
Veröffentlicht: (2024)
von: Sequeira, André, et al.
Veröffentlicht: (2024)
Neural Network Training via Stochastic Alternating Minimization with Trainable Step Sizes
von: Yan, Chengcheng, et al.
Veröffentlicht: (2025)
von: Yan, Chengcheng, et al.
Veröffentlicht: (2025)
Trainable Dynamic Mask Sparse Attention
von: Shi, Jingze, et al.
Veröffentlicht: (2025)
von: Shi, Jingze, et al.
Veröffentlicht: (2025)
A Unified Noise-Curvature View of Loss of Trainability
von: Baveja, Gunbir Singh, et al.
Veröffentlicht: (2025)
von: Baveja, Gunbir Singh, et al.
Veröffentlicht: (2025)
Characterizing Trainability of Instantaneous Quantum Polynomial Circuit Born Machines
von: Shen, Kevin, et al.
Veröffentlicht: (2026)
von: Shen, Kevin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Divergence Results and Convergence of a Variance Reduced Version of ADAM
von: Wang, Ruiqi, et al.
Veröffentlicht: (2022) -
A Mirror Descent Perspective of Smoothed Sign Descent
von: Wang, Shuyang, et al.
Veröffentlicht: (2024) -
Non-Convex Optimization with Spectral Radius Regularization
von: Sandler, Adam, et al.
Veröffentlicht: (2021) -
Rank-Accuracy Trade-off for LoRA: A Gradient-Flow Analysis
von: Rushka, Michael, et al.
Veröffentlicht: (2026) -
On the Convergence Rate of LoRA Gradient Descent
von: Mu, Siqiao, et al.
Veröffentlicht: (2025)