Gespeichert in:
| Hauptverfasser: | Sadrtdinov, Ildus, Kodryan, Maxim, Pokonechny, Eduard, Lobacheva, Ekaterina, Vetrov, Dmitry |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2410.22113 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
To Stay or Not to Stay in the Pre-train Basin: Insights on Ensembling in Transfer Learning
von: Sadrtdinov, Ildus, et al.
Veröffentlicht: (2023)
von: Sadrtdinov, Ildus, et al.
Veröffentlicht: (2023)
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training
von: Sadrtdinov, Ildus, et al.
Veröffentlicht: (2025)
von: Sadrtdinov, Ildus, et al.
Veröffentlicht: (2025)
Can Stationary Distributions of Scale-Invariant Neural Networks Be Described by the Thermodynamics of an Ideal Gas?
von: Sadrtdinov, Ildus, et al.
Veröffentlicht: (2025)
von: Sadrtdinov, Ildus, et al.
Veröffentlicht: (2025)
Why Gaussian Diffusion Models Fail on Discrete Data and How to Prevent It?
von: Shabalin, Alexander, et al.
Veröffentlicht: (2026)
von: Shabalin, Alexander, et al.
Veröffentlicht: (2026)
Unsupervised Process Reward Models
von: Gadetsky, Artyom, et al.
Veröffentlicht: (2026)
von: Gadetsky, Artyom, et al.
Veröffentlicht: (2026)
SDE Matching: Scalable and Simulation-Free Training of Latent Stochastic Differential Equations
von: Bartosh, Grigory, et al.
Veröffentlicht: (2025)
von: Bartosh, Grigory, et al.
Veröffentlicht: (2025)
Neural Diffusion Models
von: Bartosh, Grigory, et al.
Veröffentlicht: (2023)
von: Bartosh, Grigory, et al.
Veröffentlicht: (2023)
Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates
von: Kovalev, Dmitry, et al.
Veröffentlicht: (2025)
von: Kovalev, Dmitry, et al.
Veröffentlicht: (2025)
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning
von: Mircea, Andrei, et al.
Veröffentlicht: (2025)
von: Mircea, Andrei, et al.
Veröffentlicht: (2025)
Regularized Distribution Matching Distillation for One-step Unpaired Image-to-Image Translation
von: Rakitin, Denis, et al.
Veröffentlicht: (2024)
von: Rakitin, Denis, et al.
Veröffentlicht: (2024)
Generative Flow Networks as Entropy-Regularized RL
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2023)
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2023)
Streaming Generation of Co-Speech Gestures via Accelerated Rolling Diffusion
von: Vu, Evgeniia, et al.
Veröffentlicht: (2025)
von: Vu, Evgeniia, et al.
Veröffentlicht: (2025)
Neural Flow Diffusion Models: Learnable Forward Process for Improved Diffusion Modelling
von: Bartosh, Grigory, et al.
Veröffentlicht: (2024)
von: Bartosh, Grigory, et al.
Veröffentlicht: (2024)
Improving GFlowNets with Monte Carlo Tree Search
von: Morozov, Nikita, et al.
Veröffentlicht: (2024)
von: Morozov, Nikita, et al.
Veröffentlicht: (2024)
Do Large Language Models Reason Causally Like Us? Even Better?
von: Dettki, Hanna M., et al.
Veröffentlicht: (2025)
von: Dettki, Hanna M., et al.
Veröffentlicht: (2025)
Tight and Efficient Upper Bound on Spectral Norm of Convolutional Layers
von: Grishina, Ekaterina, et al.
Veröffentlicht: (2024)
von: Grishina, Ekaterina, et al.
Veröffentlicht: (2024)
On Linear Convergence in Smooth Convex-Concave Bilinearly-Coupled Saddle-Point Optimization: Lower Bounds and Optimal Algorithms
von: Kovalev, Dmitry, et al.
Veröffentlicht: (2024)
von: Kovalev, Dmitry, et al.
Veröffentlicht: (2024)
Nesterov Finds GRAAL: Optimal and Adaptive Gradient Method for Convex Optimization
von: Borodich, Ekaterina, et al.
Veröffentlicht: (2025)
von: Borodich, Ekaterina, et al.
Veröffentlicht: (2025)
ProcrustesGPT: Compressing LLMs with Structured Matrices and Orthogonal Transformations
von: Grishina, Ekaterina, et al.
Veröffentlicht: (2025)
von: Grishina, Ekaterina, et al.
Veröffentlicht: (2025)
Adaptive Destruction Processes for Diffusion Samplers
von: Gritsaev, Timofei, et al.
Veröffentlicht: (2025)
von: Gritsaev, Timofei, et al.
Veröffentlicht: (2025)
Guided Star-Shaped Masked Diffusion
von: Meshchaninov, Viacheslav, et al.
Veröffentlicht: (2025)
von: Meshchaninov, Viacheslav, et al.
Veröffentlicht: (2025)
GLGENN: A Novel Parameter-Light Equivariant Neural Networks Architecture Based on Clifford Geometric Algebras
von: Filimoshina, Ekaterina, et al.
Veröffentlicht: (2025)
von: Filimoshina, Ekaterina, et al.
Veröffentlicht: (2025)
Generalization error of spectral algorithms
von: Velikanov, Maksim, et al.
Veröffentlicht: (2024)
von: Velikanov, Maksim, et al.
Veröffentlicht: (2024)
Diffusion on language model encodings for protein sequence generation
von: Meshchaninov, Viacheslav, et al.
Veröffentlicht: (2024)
von: Meshchaninov, Viacheslav, et al.
Veröffentlicht: (2024)
Where Do Reasoning Models Refuse?
von: Yamaguchi, Kureha, et al.
Veröffentlicht: (2025)
von: Yamaguchi, Kureha, et al.
Veröffentlicht: (2025)
Gradual Optimization Learning for Conformational Energy Minimization
von: Tsypin, Artem, et al.
Veröffentlicht: (2023)
von: Tsypin, Artem, et al.
Veröffentlicht: (2023)
What Can Grokking Teach Us About Learning Under Nonstationarity?
von: Lyle, Clare, et al.
Veröffentlicht: (2025)
von: Lyle, Clare, et al.
Veröffentlicht: (2025)
Lower Bounds and Optimal Algorithms for Non-Smooth Convex Decentralized Optimization over Time-Varying Networks
von: Kovalev, Dmitry, et al.
Veröffentlicht: (2024)
von: Kovalev, Dmitry, et al.
Veröffentlicht: (2024)
Do LLMs Understand Romanian Driving Laws? A Study on Multimodal and Fine-Tuned Question Answering
von: Barbu, Eduard, et al.
Veröffentlicht: (2025)
von: Barbu, Eduard, et al.
Veröffentlicht: (2025)
Where Do the Joules Go? Diagnosing Inference Energy Consumption
von: Chung, Jae-Won, et al.
Veröffentlicht: (2026)
von: Chung, Jae-Won, et al.
Veröffentlicht: (2026)
Motivating Next-Gen Accelerators with Flexible (N:M) Activation Sparsity via Benchmarking Lightweight Post-Training Sparsification Approaches
von: Alanova, Shirin, et al.
Veröffentlicht: (2025)
von: Alanova, Shirin, et al.
Veröffentlicht: (2025)
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks
von: Sharifloo, Amir Molzam, et al.
Veröffentlicht: (2025)
von: Sharifloo, Amir Molzam, et al.
Veröffentlicht: (2025)
How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability
von: Im, Shawn, et al.
Veröffentlicht: (2026)
von: Im, Shawn, et al.
Veröffentlicht: (2026)
Communication Compression for Byzantine Robust Learning: New Efficient Algorithms and Improved Rates
von: Rammal, Ahmad, et al.
Veröffentlicht: (2023)
von: Rammal, Ahmad, et al.
Veröffentlicht: (2023)
HOTA: Hamiltonian framework for Optimal Transport Advection
von: Buzun, Nazar, et al.
Veröffentlicht: (2025)
von: Buzun, Nazar, et al.
Veröffentlicht: (2025)
See Beyond a Single View: Multi-Attribution Learning Leads to Better Conversion Rate Prediction
von: Chen, Sishuo, et al.
Veröffentlicht: (2025)
von: Chen, Sishuo, et al.
Veröffentlicht: (2025)
A Neural Operator based on Dynamic Mode Decomposition
von: Sakovich, Nikita, et al.
Veröffentlicht: (2025)
von: Sakovich, Nikita, et al.
Veröffentlicht: (2025)
Assessing Library Instruction: Where It Has Been and Where Is It Taking Us?
von: Kenney, Donald J.
Veröffentlicht: (1987)
von: Kenney, Donald J.
Veröffentlicht: (1987)
Uncertainty Quantification for Data-Driven Machine Learning Models in Nuclear Engineering Applications: Where We Are and What Do We Need?
von: Wu, Xu, et al.
Veröffentlicht: (2025)
von: Wu, Xu, et al.
Veröffentlicht: (2025)
SANIA: Polyak-type Optimization Framework Leads to Scale Invariant Stochastic Algorithms
von: Abdukhakimov, Farshed, et al.
Veröffentlicht: (2023)
von: Abdukhakimov, Farshed, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
To Stay or Not to Stay in the Pre-train Basin: Insights on Ensembling in Transfer Learning
von: Sadrtdinov, Ildus, et al.
Veröffentlicht: (2023) -
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training
von: Sadrtdinov, Ildus, et al.
Veröffentlicht: (2025) -
Can Stationary Distributions of Scale-Invariant Neural Networks Be Described by the Thermodynamics of an Ideal Gas?
von: Sadrtdinov, Ildus, et al.
Veröffentlicht: (2025) -
Why Gaussian Diffusion Models Fail on Discrete Data and How to Prevent It?
von: Shabalin, Alexander, et al.
Veröffentlicht: (2026) -
Unsupervised Process Reward Models
von: Gadetsky, Artyom, et al.
Veröffentlicht: (2026)