Direct Soft-Policy Sampling via Langevin Dynamics
Fuente:
arXiv
Saved in:
| Main Authors: | Ki, Donghyeon, Ahn, Hee-Jun, Kim, Kyungyoon, Lee, Byung-Jun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Score-Based One-step MeanFlow Policy Optimization
by: Kim, Kyungyoon, et al.
Published: (2026)
by: Kim, Kyungyoon, et al.
Published: (2026)
Actor-Critic without Actor
by: Ki, Donghyeon, et al.
Published: (2025)
by: Ki, Donghyeon, et al.
Published: (2025)
FlashMoE: Reducing SSD I/O Bottlenecks via ML-Based Cache Replacement for Mixture-of-Experts Inference on Edge Devices
by: Kim, Byeongju, et al.
Published: (2026)
by: Kim, Byeongju, et al.
Published: (2026)
Offline Reinforcement Learning with Penalized Action Noise Injection
by: Oh, JunHyeok, et al.
Published: (2025)
by: Oh, JunHyeok, et al.
Published: (2025)
FairDICE: Fairness-Driven Offline Multi-Objective Reinforcement Learning
by: Kim, Woosung, et al.
Published: (2025)
by: Kim, Woosung, et al.
Published: (2025)
Adaptive Non-uniform Timestep Sampling for Accelerating Diffusion Model Training
by: Kim, Myunsoo, et al.
Published: (2024)
by: Kim, Myunsoo, et al.
Published: (2024)
Aligning AI Agents via Information-Directed Sampling
by: Jeon, Hong Jun, et al.
Published: (2024)
by: Jeon, Hong Jun, et al.
Published: (2024)
NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations
by: Kim, Myunsoo, et al.
Published: (2025)
by: Kim, Myunsoo, et al.
Published: (2025)
TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning
by: Lee, Hayeong, et al.
Published: (2026)
by: Lee, Hayeong, et al.
Published: (2026)
DIAR: Diffusion-model-guided Implicit Q-learning with Adaptive Revaluation
by: Park, Jaehyun, et al.
Published: (2024)
by: Park, Jaehyun, et al.
Published: (2024)
ARCLE: The Abstraction and Reasoning Corpus Learning Environment for Reinforcement Learning
by: Lee, Hosung, et al.
Published: (2024)
by: Lee, Hosung, et al.
Published: (2024)
A Framework for Mining Collectively-Behaving Bots in MMORPGs
by: Kim, Hyunsoo, et al.
Published: (2025)
by: Kim, Hyunsoo, et al.
Published: (2025)
Prior-Guided Diffusion Planning for Offline Reinforcement Learning
by: Ki, Donghyeon, et al.
Published: (2025)
by: Ki, Donghyeon, et al.
Published: (2025)
Soft Deterministic Policy Gradient with Gaussian Smoothing
by: Na, Hyunjun, et al.
Published: (2026)
by: Na, Hyunjun, et al.
Published: (2026)
Multiple Invertible and Partial-Equivariant Function for Latent Vector Transformation to Enhance Disentanglement in VAEs
by: Jung, Hee-Jun, et al.
Published: (2025)
by: Jung, Hee-Jun, et al.
Published: (2025)
CFASL: Composite Factor-Aligned Symmetry Learning for Disentanglement in Variational AutoEncoder
by: Jung, Hee-Jun, et al.
Published: (2024)
by: Jung, Hee-Jun, et al.
Published: (2024)
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
by: Lee, Haanvid, et al.
Published: (2024)
by: Lee, Haanvid, et al.
Published: (2024)
PID-controlled Langevin Dynamics for Faster Sampling of Generative Models
by: Chen, Hongyi, et al.
Published: (2025)
by: Chen, Hongyi, et al.
Published: (2025)
Diffusion Fine-Tuning via Reparameterized Policy Gradient of the Soft Q-Function
by: Kang, Hyeongyu, et al.
Published: (2025)
by: Kang, Hyeongyu, et al.
Published: (2025)
RW-NSGCN: A Robust Approach to Structural Attacks via Negative Sampling
by: He, Shuqi, et al.
Published: (2024)
by: He, Shuqi, et al.
Published: (2024)
SOE: Sample-Efficient Robot Policy Self-Improvement via On-Manifold Exploration
by: Jin, Yang, et al.
Published: (2025)
by: Jin, Yang, et al.
Published: (2025)
Efficient Approximate Posterior Sampling with Annealed Langevin Monte Carlo
by: Parulekar, Advait, et al.
Published: (2025)
by: Parulekar, Advait, et al.
Published: (2025)
Accelerating Approximate Thompson Sampling with Underdamped Langevin Monte Carlo
by: Zheng, Haoyang, et al.
Published: (2024)
by: Zheng, Haoyang, et al.
Published: (2024)
Constrained Exploration via Reflected Replica Exchange Stochastic Gradient Langevin Dynamics
by: Zheng, Haoyang, et al.
Published: (2024)
by: Zheng, Haoyang, et al.
Published: (2024)
VPO: Leveraging the Number of Votes in Preference Optimization
by: Cho, Jae Hyeon, et al.
Published: (2024)
by: Cho, Jae Hyeon, et al.
Published: (2024)
World Models via Policy-Guided Trajectory Diffusion
by: Rigter, Marc, et al.
Published: (2023)
by: Rigter, Marc, et al.
Published: (2023)
PG-Rainbow: Using Distributional Reinforcement Learning in Policy Gradient Methods
by: Jeon, WooJae, et al.
Published: (2024)
by: Jeon, WooJae, et al.
Published: (2024)
Towards Scalable Handwriting Communication via EEG Decoding and Latent Embedding Integration
by: Kim, Jun-Young, et al.
Published: (2024)
by: Kim, Jun-Young, et al.
Published: (2024)
Imagined Speech State Classification for Robust Brain-Computer Interface
by: Ko, Byung-Kwan, et al.
Published: (2024)
by: Ko, Byung-Kwan, et al.
Published: (2024)
EPIC: Graph Augmentation with Edit Path Interpolation via Learnable Cost
by: Heo, Jaeseung, et al.
Published: (2023)
by: Heo, Jaeseung, et al.
Published: (2023)
Rethinking Langevin Thompson Sampling from A Stochastic Approximation Perspective
by: Wang, Weixin, et al.
Published: (2025)
by: Wang, Weixin, et al.
Published: (2025)
Label-based Graph Augmentation with Metapath for Graph Anomaly Detection
by: Kim, Hwan, et al.
Published: (2023)
by: Kim, Hwan, et al.
Published: (2023)
On Predicting Post-Click Conversion Rate via Counterfactual Inference
by: Ahn, Junhyung, et al.
Published: (2025)
by: Ahn, Junhyung, et al.
Published: (2025)
Offline Imitation Learning by Controlling the Effective Planning Horizon
by: Ahn, Hee-Jun, et al.
Published: (2024)
by: Ahn, Hee-Jun, et al.
Published: (2024)
Soft Sequence Policy Optimization
by: Glazyrina, Svetlana, et al.
Published: (2026)
by: Glazyrina, Svetlana, et al.
Published: (2026)
FinTexTS: Financial Text-Paired Time-Series Dataset via Semantic-Based and Multi-Level Pairing
by: Lee, Jaehoon, et al.
Published: (2026)
by: Lee, Jaehoon, et al.
Published: (2026)
Inconsistency-Aware Minimization: Improving Generalization with Unlabeled Data
by: Kim, Hee-Sung, et al.
Published: (2026)
by: Kim, Hee-Sung, et al.
Published: (2026)
Diffusion-Based Offline RL for Improved Decision-Making in Augmented ARC Task
by: Kim, Yunho, et al.
Published: (2024)
by: Kim, Yunho, et al.
Published: (2024)
IMLE Policy: Fast and Sample Efficient Visuomotor Policy Learning via Implicit Maximum Likelihood Estimation
by: Rana, Krishan, et al.
Published: (2025)
by: Rana, Krishan, et al.
Published: (2025)
SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy Optimization
by: Zheng, Zhi, et al.
Published: (2025)
by: Zheng, Zhi, et al.
Published: (2025)
Similar Items
-
Score-Based One-step MeanFlow Policy Optimization
by: Kim, Kyungyoon, et al.
Published: (2026) -
Actor-Critic without Actor
by: Ki, Donghyeon, et al.
Published: (2025) -
FlashMoE: Reducing SSD I/O Bottlenecks via ML-Based Cache Replacement for Mixture-of-Experts Inference on Edge Devices
by: Kim, Byeongju, et al.
Published: (2026) -
Offline Reinforcement Learning with Penalized Action Noise Injection
by: Oh, JunHyeok, et al.
Published: (2025) -
FairDICE: Fairness-Driven Offline Multi-Objective Reinforcement Learning
by: Kim, Woosung, et al.
Published: (2025)