Saved in:
| Main Authors: | Sadrtdinov, Ildus, Kodryan, Maxim, Pokonechny, Eduard, Lobacheva, Ekaterina, Vetrov, Dmitry |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2410.22113 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
To Stay or Not to Stay in the Pre-train Basin: Insights on Ensembling in Transfer Learning
by: Sadrtdinov, Ildus, et al.
Published: (2023)
by: Sadrtdinov, Ildus, et al.
Published: (2023)
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training
by: Sadrtdinov, Ildus, et al.
Published: (2025)
by: Sadrtdinov, Ildus, et al.
Published: (2025)
Can Stationary Distributions of Scale-Invariant Neural Networks Be Described by the Thermodynamics of an Ideal Gas?
by: Sadrtdinov, Ildus, et al.
Published: (2025)
by: Sadrtdinov, Ildus, et al.
Published: (2025)
Why Gaussian Diffusion Models Fail on Discrete Data and How to Prevent It?
by: Shabalin, Alexander, et al.
Published: (2026)
by: Shabalin, Alexander, et al.
Published: (2026)
Unsupervised Process Reward Models
by: Gadetsky, Artyom, et al.
Published: (2026)
by: Gadetsky, Artyom, et al.
Published: (2026)
SDE Matching: Scalable and Simulation-Free Training of Latent Stochastic Differential Equations
by: Bartosh, Grigory, et al.
Published: (2025)
by: Bartosh, Grigory, et al.
Published: (2025)
Neural Diffusion Models
by: Bartosh, Grigory, et al.
Published: (2023)
by: Bartosh, Grigory, et al.
Published: (2023)
Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates
by: Kovalev, Dmitry, et al.
Published: (2025)
by: Kovalev, Dmitry, et al.
Published: (2025)
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning
by: Mircea, Andrei, et al.
Published: (2025)
by: Mircea, Andrei, et al.
Published: (2025)
Regularized Distribution Matching Distillation for One-step Unpaired Image-to-Image Translation
by: Rakitin, Denis, et al.
Published: (2024)
by: Rakitin, Denis, et al.
Published: (2024)
Generative Flow Networks as Entropy-Regularized RL
by: Tiapkin, Daniil, et al.
Published: (2023)
by: Tiapkin, Daniil, et al.
Published: (2023)
Streaming Generation of Co-Speech Gestures via Accelerated Rolling Diffusion
by: Vu, Evgeniia, et al.
Published: (2025)
by: Vu, Evgeniia, et al.
Published: (2025)
Neural Flow Diffusion Models: Learnable Forward Process for Improved Diffusion Modelling
by: Bartosh, Grigory, et al.
Published: (2024)
by: Bartosh, Grigory, et al.
Published: (2024)
Improving GFlowNets with Monte Carlo Tree Search
by: Morozov, Nikita, et al.
Published: (2024)
by: Morozov, Nikita, et al.
Published: (2024)
Do Large Language Models Reason Causally Like Us? Even Better?
by: Dettki, Hanna M., et al.
Published: (2025)
by: Dettki, Hanna M., et al.
Published: (2025)
Tight and Efficient Upper Bound on Spectral Norm of Convolutional Layers
by: Grishina, Ekaterina, et al.
Published: (2024)
by: Grishina, Ekaterina, et al.
Published: (2024)
On Linear Convergence in Smooth Convex-Concave Bilinearly-Coupled Saddle-Point Optimization: Lower Bounds and Optimal Algorithms
by: Kovalev, Dmitry, et al.
Published: (2024)
by: Kovalev, Dmitry, et al.
Published: (2024)
Nesterov Finds GRAAL: Optimal and Adaptive Gradient Method for Convex Optimization
by: Borodich, Ekaterina, et al.
Published: (2025)
by: Borodich, Ekaterina, et al.
Published: (2025)
ProcrustesGPT: Compressing LLMs with Structured Matrices and Orthogonal Transformations
by: Grishina, Ekaterina, et al.
Published: (2025)
by: Grishina, Ekaterina, et al.
Published: (2025)
Adaptive Destruction Processes for Diffusion Samplers
by: Gritsaev, Timofei, et al.
Published: (2025)
by: Gritsaev, Timofei, et al.
Published: (2025)
Guided Star-Shaped Masked Diffusion
by: Meshchaninov, Viacheslav, et al.
Published: (2025)
by: Meshchaninov, Viacheslav, et al.
Published: (2025)
GLGENN: A Novel Parameter-Light Equivariant Neural Networks Architecture Based on Clifford Geometric Algebras
by: Filimoshina, Ekaterina, et al.
Published: (2025)
by: Filimoshina, Ekaterina, et al.
Published: (2025)
Generalization error of spectral algorithms
by: Velikanov, Maksim, et al.
Published: (2024)
by: Velikanov, Maksim, et al.
Published: (2024)
Diffusion on language model encodings for protein sequence generation
by: Meshchaninov, Viacheslav, et al.
Published: (2024)
by: Meshchaninov, Viacheslav, et al.
Published: (2024)
Where Do Reasoning Models Refuse?
by: Yamaguchi, Kureha, et al.
Published: (2025)
by: Yamaguchi, Kureha, et al.
Published: (2025)
Gradual Optimization Learning for Conformational Energy Minimization
by: Tsypin, Artem, et al.
Published: (2023)
by: Tsypin, Artem, et al.
Published: (2023)
What Can Grokking Teach Us About Learning Under Nonstationarity?
by: Lyle, Clare, et al.
Published: (2025)
by: Lyle, Clare, et al.
Published: (2025)
Lower Bounds and Optimal Algorithms for Non-Smooth Convex Decentralized Optimization over Time-Varying Networks
by: Kovalev, Dmitry, et al.
Published: (2024)
by: Kovalev, Dmitry, et al.
Published: (2024)
Do LLMs Understand Romanian Driving Laws? A Study on Multimodal and Fine-Tuned Question Answering
by: Barbu, Eduard, et al.
Published: (2025)
by: Barbu, Eduard, et al.
Published: (2025)
Where Do the Joules Go? Diagnosing Inference Energy Consumption
by: Chung, Jae-Won, et al.
Published: (2026)
by: Chung, Jae-Won, et al.
Published: (2026)
Motivating Next-Gen Accelerators with Flexible (N:M) Activation Sparsity via Benchmarking Lightweight Post-Training Sparsification Approaches
by: Alanova, Shirin, et al.
Published: (2025)
by: Alanova, Shirin, et al.
Published: (2025)
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks
by: Sharifloo, Amir Molzam, et al.
Published: (2025)
by: Sharifloo, Amir Molzam, et al.
Published: (2025)
How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability
by: Im, Shawn, et al.
Published: (2026)
by: Im, Shawn, et al.
Published: (2026)
Communication Compression for Byzantine Robust Learning: New Efficient Algorithms and Improved Rates
by: Rammal, Ahmad, et al.
Published: (2023)
by: Rammal, Ahmad, et al.
Published: (2023)
HOTA: Hamiltonian framework for Optimal Transport Advection
by: Buzun, Nazar, et al.
Published: (2025)
by: Buzun, Nazar, et al.
Published: (2025)
See Beyond a Single View: Multi-Attribution Learning Leads to Better Conversion Rate Prediction
by: Chen, Sishuo, et al.
Published: (2025)
by: Chen, Sishuo, et al.
Published: (2025)
A Neural Operator based on Dynamic Mode Decomposition
by: Sakovich, Nikita, et al.
Published: (2025)
by: Sakovich, Nikita, et al.
Published: (2025)
Assessing Library Instruction: Where It Has Been and Where Is It Taking Us?
by: Kenney, Donald J.
Published: (1987)
by: Kenney, Donald J.
Published: (1987)
Uncertainty Quantification for Data-Driven Machine Learning Models in Nuclear Engineering Applications: Where We Are and What Do We Need?
by: Wu, Xu, et al.
Published: (2025)
by: Wu, Xu, et al.
Published: (2025)
SANIA: Polyak-type Optimization Framework Leads to Scale Invariant Stochastic Algorithms
by: Abdukhakimov, Farshed, et al.
Published: (2023)
by: Abdukhakimov, Farshed, et al.
Published: (2023)
Similar Items
-
To Stay or Not to Stay in the Pre-train Basin: Insights on Ensembling in Transfer Learning
by: Sadrtdinov, Ildus, et al.
Published: (2023) -
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training
by: Sadrtdinov, Ildus, et al.
Published: (2025) -
Can Stationary Distributions of Scale-Invariant Neural Networks Be Described by the Thermodynamics of an Ideal Gas?
by: Sadrtdinov, Ildus, et al.
Published: (2025) -
Why Gaussian Diffusion Models Fail on Discrete Data and How to Prevent It?
by: Shabalin, Alexander, et al.
Published: (2026) -
Unsupervised Process Reward Models
by: Gadetsky, Artyom, et al.
Published: (2026)