Direct Soft-Policy Sampling via Langevin Dynamics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ki, Donghyeon, Ahn, Hee-Jun, Kim, Kyungyoon, Lee, Byung-Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Score-Based One-step MeanFlow Policy Optimization
von: Kim, Kyungyoon, et al.
Veröffentlicht: (2026)
von: Kim, Kyungyoon, et al.
Veröffentlicht: (2026)
Actor-Critic without Actor
von: Ki, Donghyeon, et al.
Veröffentlicht: (2025)
von: Ki, Donghyeon, et al.
Veröffentlicht: (2025)
FlashMoE: Reducing SSD I/O Bottlenecks via ML-Based Cache Replacement for Mixture-of-Experts Inference on Edge Devices
von: Kim, Byeongju, et al.
Veröffentlicht: (2026)
von: Kim, Byeongju, et al.
Veröffentlicht: (2026)
Offline Reinforcement Learning with Penalized Action Noise Injection
von: Oh, JunHyeok, et al.
Veröffentlicht: (2025)
von: Oh, JunHyeok, et al.
Veröffentlicht: (2025)
FairDICE: Fairness-Driven Offline Multi-Objective Reinforcement Learning
von: Kim, Woosung, et al.
Veröffentlicht: (2025)
von: Kim, Woosung, et al.
Veröffentlicht: (2025)
Adaptive Non-uniform Timestep Sampling for Accelerating Diffusion Model Training
von: Kim, Myunsoo, et al.
Veröffentlicht: (2024)
von: Kim, Myunsoo, et al.
Veröffentlicht: (2024)
Aligning AI Agents via Information-Directed Sampling
von: Jeon, Hong Jun, et al.
Veröffentlicht: (2024)
von: Jeon, Hong Jun, et al.
Veröffentlicht: (2024)
NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations
von: Kim, Myunsoo, et al.
Veröffentlicht: (2025)
von: Kim, Myunsoo, et al.
Veröffentlicht: (2025)
TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning
von: Lee, Hayeong, et al.
Veröffentlicht: (2026)
von: Lee, Hayeong, et al.
Veröffentlicht: (2026)
DIAR: Diffusion-model-guided Implicit Q-learning with Adaptive Revaluation
von: Park, Jaehyun, et al.
Veröffentlicht: (2024)
von: Park, Jaehyun, et al.
Veröffentlicht: (2024)
ARCLE: The Abstraction and Reasoning Corpus Learning Environment for Reinforcement Learning
von: Lee, Hosung, et al.
Veröffentlicht: (2024)
von: Lee, Hosung, et al.
Veröffentlicht: (2024)
A Framework for Mining Collectively-Behaving Bots in MMORPGs
von: Kim, Hyunsoo, et al.
Veröffentlicht: (2025)
von: Kim, Hyunsoo, et al.
Veröffentlicht: (2025)
Prior-Guided Diffusion Planning for Offline Reinforcement Learning
von: Ki, Donghyeon, et al.
Veröffentlicht: (2025)
von: Ki, Donghyeon, et al.
Veröffentlicht: (2025)
Soft Deterministic Policy Gradient with Gaussian Smoothing
von: Na, Hyunjun, et al.
Veröffentlicht: (2026)
von: Na, Hyunjun, et al.
Veröffentlicht: (2026)
Multiple Invertible and Partial-Equivariant Function for Latent Vector Transformation to Enhance Disentanglement in VAEs
von: Jung, Hee-Jun, et al.
Veröffentlicht: (2025)
von: Jung, Hee-Jun, et al.
Veröffentlicht: (2025)
CFASL: Composite Factor-Aligned Symmetry Learning for Disentanglement in Variational AutoEncoder
von: Jung, Hee-Jun, et al.
Veröffentlicht: (2024)
von: Jung, Hee-Jun, et al.
Veröffentlicht: (2024)
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
von: Lee, Haanvid, et al.
Veröffentlicht: (2024)
von: Lee, Haanvid, et al.
Veröffentlicht: (2024)
PID-controlled Langevin Dynamics for Faster Sampling of Generative Models
von: Chen, Hongyi, et al.
Veröffentlicht: (2025)
von: Chen, Hongyi, et al.
Veröffentlicht: (2025)
Diffusion Fine-Tuning via Reparameterized Policy Gradient of the Soft Q-Function
von: Kang, Hyeongyu, et al.
Veröffentlicht: (2025)
von: Kang, Hyeongyu, et al.
Veröffentlicht: (2025)
RW-NSGCN: A Robust Approach to Structural Attacks via Negative Sampling
von: He, Shuqi, et al.
Veröffentlicht: (2024)
von: He, Shuqi, et al.
Veröffentlicht: (2024)
SOE: Sample-Efficient Robot Policy Self-Improvement via On-Manifold Exploration
von: Jin, Yang, et al.
Veröffentlicht: (2025)
von: Jin, Yang, et al.
Veröffentlicht: (2025)
Efficient Approximate Posterior Sampling with Annealed Langevin Monte Carlo
von: Parulekar, Advait, et al.
Veröffentlicht: (2025)
von: Parulekar, Advait, et al.
Veröffentlicht: (2025)
Accelerating Approximate Thompson Sampling with Underdamped Langevin Monte Carlo
von: Zheng, Haoyang, et al.
Veröffentlicht: (2024)
von: Zheng, Haoyang, et al.
Veröffentlicht: (2024)
Constrained Exploration via Reflected Replica Exchange Stochastic Gradient Langevin Dynamics
von: Zheng, Haoyang, et al.
Veröffentlicht: (2024)
von: Zheng, Haoyang, et al.
Veröffentlicht: (2024)
VPO: Leveraging the Number of Votes in Preference Optimization
von: Cho, Jae Hyeon, et al.
Veröffentlicht: (2024)
von: Cho, Jae Hyeon, et al.
Veröffentlicht: (2024)
World Models via Policy-Guided Trajectory Diffusion
von: Rigter, Marc, et al.
Veröffentlicht: (2023)
von: Rigter, Marc, et al.
Veröffentlicht: (2023)
PG-Rainbow: Using Distributional Reinforcement Learning in Policy Gradient Methods
von: Jeon, WooJae, et al.
Veröffentlicht: (2024)
von: Jeon, WooJae, et al.
Veröffentlicht: (2024)
Towards Scalable Handwriting Communication via EEG Decoding and Latent Embedding Integration
von: Kim, Jun-Young, et al.
Veröffentlicht: (2024)
von: Kim, Jun-Young, et al.
Veröffentlicht: (2024)
Imagined Speech State Classification for Robust Brain-Computer Interface
von: Ko, Byung-Kwan, et al.
Veröffentlicht: (2024)
von: Ko, Byung-Kwan, et al.
Veröffentlicht: (2024)
EPIC: Graph Augmentation with Edit Path Interpolation via Learnable Cost
von: Heo, Jaeseung, et al.
Veröffentlicht: (2023)
von: Heo, Jaeseung, et al.
Veröffentlicht: (2023)
Rethinking Langevin Thompson Sampling from A Stochastic Approximation Perspective
von: Wang, Weixin, et al.
Veröffentlicht: (2025)
von: Wang, Weixin, et al.
Veröffentlicht: (2025)
Label-based Graph Augmentation with Metapath for Graph Anomaly Detection
von: Kim, Hwan, et al.
Veröffentlicht: (2023)
von: Kim, Hwan, et al.
Veröffentlicht: (2023)
On Predicting Post-Click Conversion Rate via Counterfactual Inference
von: Ahn, Junhyung, et al.
Veröffentlicht: (2025)
von: Ahn, Junhyung, et al.
Veröffentlicht: (2025)
Offline Imitation Learning by Controlling the Effective Planning Horizon
von: Ahn, Hee-Jun, et al.
Veröffentlicht: (2024)
von: Ahn, Hee-Jun, et al.
Veröffentlicht: (2024)
Soft Sequence Policy Optimization
von: Glazyrina, Svetlana, et al.
Veröffentlicht: (2026)
von: Glazyrina, Svetlana, et al.
Veröffentlicht: (2026)
FinTexTS: Financial Text-Paired Time-Series Dataset via Semantic-Based and Multi-Level Pairing
von: Lee, Jaehoon, et al.
Veröffentlicht: (2026)
von: Lee, Jaehoon, et al.
Veröffentlicht: (2026)
Inconsistency-Aware Minimization: Improving Generalization with Unlabeled Data
von: Kim, Hee-Sung, et al.
Veröffentlicht: (2026)
von: Kim, Hee-Sung, et al.
Veröffentlicht: (2026)
Diffusion-Based Offline RL for Improved Decision-Making in Augmented ARC Task
von: Kim, Yunho, et al.
Veröffentlicht: (2024)
von: Kim, Yunho, et al.
Veröffentlicht: (2024)
IMLE Policy: Fast and Sample Efficient Visuomotor Policy Learning via Implicit Maximum Likelihood Estimation
von: Rana, Krishan, et al.
Veröffentlicht: (2025)
von: Rana, Krishan, et al.
Veröffentlicht: (2025)
SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy Optimization
von: Zheng, Zhi, et al.
Veröffentlicht: (2025)
von: Zheng, Zhi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Score-Based One-step MeanFlow Policy Optimization
von: Kim, Kyungyoon, et al.
Veröffentlicht: (2026) -
Actor-Critic without Actor
von: Ki, Donghyeon, et al.
Veröffentlicht: (2025) -
FlashMoE: Reducing SSD I/O Bottlenecks via ML-Based Cache Replacement for Mixture-of-Experts Inference on Edge Devices
von: Kim, Byeongju, et al.
Veröffentlicht: (2026) -
Offline Reinforcement Learning with Penalized Action Noise Injection
von: Oh, JunHyeok, et al.
Veröffentlicht: (2025) -
FairDICE: Fairness-Driven Offline Multi-Objective Reinforcement Learning
von: Kim, Woosung, et al.
Veröffentlicht: (2025)