Parameter Expanded Stochastic Gradient Markov Chain Monte Carlo
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Hyunsu, Nam, Giung, Yun, Chulhee, Yang, Hongseok, Lee, Juho |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Variational Partial Group Convolutions for Input-Aware Partial Equivariance of Rotations and Color-Shifts
von: Kim, Hyunsu, et al.
Veröffentlicht: (2024)
von: Kim, Hyunsu, et al.
Veröffentlicht: (2024)
Axial Neural Networks for Dimension-Free Foundation Models
von: Kim, Hyunsu, et al.
Veröffentlicht: (2025)
von: Kim, Hyunsu, et al.
Veröffentlicht: (2025)
Sparse Weight Averaging with Multiple Particles for Iterative Magnitude Pruning
von: Choi, Moonseok, et al.
Veröffentlicht: (2023)
von: Choi, Moonseok, et al.
Veröffentlicht: (2023)
Fast Ensembling with Diffusion Schrödinger Bridge
von: Kim, Hyunsu, et al.
Veröffentlicht: (2024)
von: Kim, Hyunsu, et al.
Veröffentlicht: (2024)
Learning Infinitesimal Generators of Continuous Symmetries from Data
von: Ko, Gyeonghoon, et al.
Veröffentlicht: (2024)
von: Ko, Gyeonghoon, et al.
Veröffentlicht: (2024)
Improving Constrained Language Generation via Self-Distilled Twisted Sequential Monte Carlo
von: Kim, Sooyeon, et al.
Veröffentlicht: (2025)
von: Kim, Sooyeon, et al.
Veröffentlicht: (2025)
The Cost of Robustness: Tighter Bounds on Parameter Complexity for Robust Memorization in ReLU Nets
von: Kim, Yujun, et al.
Veröffentlicht: (2025)
von: Kim, Yujun, et al.
Veröffentlicht: (2025)
Ex Uno Pluria: Insights on Ensembling in Low Precision Number Systems
von: Nam, Giung, et al.
Veröffentlicht: (2024)
von: Nam, Giung, et al.
Veröffentlicht: (2024)
Provable Benefit of Cutout and CutMix for Feature Learning
von: Oh, Junsoo, et al.
Veröffentlicht: (2024)
von: Oh, Junsoo, et al.
Veröffentlicht: (2024)
Learning to Explore for Stochastic Gradient MCMC
von: Kim, SeungHyun, et al.
Veröffentlicht: (2024)
von: Kim, SeungHyun, et al.
Veröffentlicht: (2024)
Policy Gradient Algorithms with Monte Carlo Tree Learning for Non-Markov Decision Processes
von: Morimura, Tetsuro, et al.
Veröffentlicht: (2022)
von: Morimura, Tetsuro, et al.
Veröffentlicht: (2022)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
von: Kim, Jihwan, et al.
Veröffentlicht: (2026)
von: Kim, Jihwan, et al.
Veröffentlicht: (2026)
Minor First, Major Last: A Depth-Induced Implicit Bias of Sharpness-Aware Minimization
von: Moon, Chaewon, et al.
Veröffentlicht: (2026)
von: Moon, Chaewon, et al.
Veröffentlicht: (2026)
Lipsum-FT: Robust Fine-Tuning of Zero-Shot Models Using Random Text Guidance
von: Nam, Giung, et al.
Veröffentlicht: (2024)
von: Nam, Giung, et al.
Veröffentlicht: (2024)
Continuous Monte Carlo Graph Search
von: Kujanpää, Kalle, et al.
Veröffentlicht: (2022)
von: Kujanpää, Kalle, et al.
Veröffentlicht: (2022)
Enhancing Transfer Learning with Flexible Nonparametric Posterior Sampling
von: Lee, Hyungi, et al.
Veröffentlicht: (2024)
von: Lee, Hyungi, et al.
Veröffentlicht: (2024)
Lightweight Dataset Pruning without Full Training via Example Difficulty and Prediction Uncertainty
von: Cho, Yeseul, et al.
Veröffentlicht: (2025)
von: Cho, Yeseul, et al.
Veröffentlicht: (2025)
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of Plasticity
von: Shin, Baekrok, et al.
Veröffentlicht: (2024)
von: Shin, Baekrok, et al.
Veröffentlicht: (2024)
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
von: Baek, Beomhan, et al.
Veröffentlicht: (2025)
von: Baek, Beomhan, et al.
Veröffentlicht: (2025)
MC$^2$A: Enabling Algorithm-Hardware Co-Design for Efficient Markov Chain Monte Carlo Acceleration
von: Zhao, Shirui, et al.
Veröffentlicht: (2025)
von: Zhao, Shirui, et al.
Veröffentlicht: (2025)
Stochastic Parameter Decomposition
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2025)
von: Bushnaq, Lucius, et al.
Veröffentlicht: (2025)
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
von: Song, Minhak, et al.
Veröffentlicht: (2025)
von: Song, Minhak, et al.
Veröffentlicht: (2025)
Efficient Monte Carlo Tree Search via On-the-Fly State-Conditioned Action Abstraction
von: Kwak, Yunhyeok, et al.
Veröffentlicht: (2024)
von: Kwak, Yunhyeok, et al.
Veröffentlicht: (2024)
Policy Gradient for Robust Markov Decision Processes
von: Wang, Qiuhao, et al.
Veröffentlicht: (2024)
von: Wang, Qiuhao, et al.
Veröffentlicht: (2024)
MemEIC: A Step Toward Continual and Compositional Knowledge Editing
von: Seong, Jin, et al.
Veröffentlicht: (2025)
von: Seong, Jin, et al.
Veröffentlicht: (2025)
Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thought
von: Huang, Jianhao, et al.
Veröffentlicht: (2025)
von: Huang, Jianhao, et al.
Veröffentlicht: (2025)
Monte Carlo Permutation Search
von: Cazenave, Tristan
Veröffentlicht: (2025)
von: Cazenave, Tristan
Veröffentlicht: (2025)
Clarifying Shampoo: Adapting Spectral Descent to Stochasticity and the Parameter Trajectory
von: Eschenhagen, Runa, et al.
Veröffentlicht: (2026)
von: Eschenhagen, Runa, et al.
Veröffentlicht: (2026)
Accurate Large-sample Uncertainty Quantification using Stochastic Gradient Markov Chain Monte Carlo
von: Wang, Yu, et al.
Veröffentlicht: (2026)
von: Wang, Yu, et al.
Veröffentlicht: (2026)
Distilled Thompson Sampling: Practical and Efficient Thompson Sampling via Imitation Learning
von: Namkoong, Hongseok, et al.
Veröffentlicht: (2020)
von: Namkoong, Hongseok, et al.
Veröffentlicht: (2020)
Epistemic Monte Carlo Tree Search
von: Oren, Yaniv, et al.
Veröffentlicht: (2022)
von: Oren, Yaniv, et al.
Veröffentlicht: (2022)
Linear attention is (maybe) all you need (to understand transformer optimization)
von: Ahn, Kwangjun, et al.
Veröffentlicht: (2023)
von: Ahn, Kwangjun, et al.
Veröffentlicht: (2023)
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
Active Learning with Selective Time-Step Acquisition for PDEs
von: Kim, Yegon, et al.
Veröffentlicht: (2025)
von: Kim, Yegon, et al.
Veröffentlicht: (2025)
A Scalable and Transferable Time Series Prediction Framework for Demand Forecasting
von: Park, Young-Jin, et al.
Veröffentlicht: (2024)
von: Park, Young-Jin, et al.
Veröffentlicht: (2024)
Principled Gradient-based Markov Chain Monte Carlo for Text Generation
von: Du, Li, et al.
Veröffentlicht: (2023)
von: Du, Li, et al.
Veröffentlicht: (2023)
Mitigating Gradient Overlap in Deep Residual Networks with Gradient Normalization for Improved Non-Convex Optimization
von: Yun, Juyoung
Veröffentlicht: (2024)
von: Yun, Juyoung
Veröffentlicht: (2024)
Twice Sequential Monte Carlo for Tree Search
von: Oren, Yaniv, et al.
Veröffentlicht: (2025)
von: Oren, Yaniv, et al.
Veröffentlicht: (2025)
Doubly Robust Monte Carlo Tree Search
von: Liu, Manqing, et al.
Veröffentlicht: (2025)
von: Liu, Manqing, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Variational Partial Group Convolutions for Input-Aware Partial Equivariance of Rotations and Color-Shifts
von: Kim, Hyunsu, et al.
Veröffentlicht: (2024) -
Axial Neural Networks for Dimension-Free Foundation Models
von: Kim, Hyunsu, et al.
Veröffentlicht: (2025) -
Sparse Weight Averaging with Multiple Particles for Iterative Magnitude Pruning
von: Choi, Moonseok, et al.
Veröffentlicht: (2023) -
Fast Ensembling with Diffusion Schrödinger Bridge
von: Kim, Hyunsu, et al.
Veröffentlicht: (2024) -
Learning Infinitesimal Generators of Continuous Symmetries from Data
von: Ko, Gyeonghoon, et al.
Veröffentlicht: (2024)