Q-Star Meets Scalable Posterior Sampling: Bridging Theory and Practice via HyperAgent
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yingru, Xu, Jiawei, Han, Lei, Luo, Zhi-Quan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scalable Exploration via Ensemble++
by: Li, Yingru, et al.
Published: (2024)
by: Li, Yingru, et al.
Published: (2024)
Prior-dependent analysis of posterior sampling reinforcement learning with function approximation
by: Li, Yingru, et al.
Published: (2024)
by: Li, Yingru, et al.
Published: (2024)
Optimistic Thompson Sampling for No-Regret Learning in Unknown Games
by: Li, Yingru, et al.
Published: (2024)
by: Li, Yingru, et al.
Published: (2024)
Bridging Distributional and Risk-sensitive Reinforcement Learning with Provable Regret Bounds
by: Liang, Hao, et al.
Published: (2022)
by: Liang, Hao, et al.
Published: (2022)
Logit Dynamics in Softmax Policy Gradient Methods
by: Li, Yingru
Published: (2025)
by: Li, Yingru
Published: (2025)
HyperAgent: Generalist Software Engineering Agents to Solve Coding Tasks at Scale
by: Phan, Huy Nhat, et al.
Published: (2024)
by: Phan, Huy Nhat, et al.
Published: (2024)
Scalable In-Context Q-Learning
by: Liu, Jinmei, et al.
Published: (2025)
by: Liu, Jinmei, et al.
Published: (2025)
Beyond Precision: Training-Inference Mismatch is an Optimization Problem and Simple LR Scheduling Fixes It
by: Zhang, Yaxiang, et al.
Published: (2026)
by: Zhang, Yaxiang, et al.
Published: (2026)
Bridging Theory and Practice in Link Representation with Graph Neural Networks
by: Lachi, Veronica, et al.
Published: (2025)
by: Lachi, Veronica, et al.
Published: (2025)
Scalable Differentiable Causal Discovery in the Presence of Latent Confounders with Skeleton Posterior (Extended Version)
by: Ma, Pingchuan, et al.
Published: (2024)
by: Ma, Pingchuan, et al.
Published: (2024)
Tackling Data Corruption in Offline Reinforcement Learning via Sequence Modeling
by: Xu, Jiawei, et al.
Published: (2024)
by: Xu, Jiawei, et al.
Published: (2024)
Language Agents Meet Causality -- Bridging LLMs and Causal World Models
by: Gkountouras, John, et al.
Published: (2024)
by: Gkountouras, John, et al.
Published: (2024)
Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
by: Xiong, Wei, et al.
Published: (2023)
by: Xiong, Wei, et al.
Published: (2023)
\textsc{MasFACT}: Continual Multi-Agent Topology Learning via Geometry-Aware Posterior Transfer
by: Wang, Xuefei, et al.
Published: (2026)
by: Wang, Xuefei, et al.
Published: (2026)
Identity Bridge: Enabling Implicit Reasoning via Shared Latent Memory
by: Lin, Pengxiao, et al.
Published: (2025)
by: Lin, Pengxiao, et al.
Published: (2025)
Solving Diffusion Inverse Problems with Restart Posterior Sampling
by: Ahmed, Bilal, et al.
Published: (2025)
by: Ahmed, Bilal, et al.
Published: (2025)
On Sample-Efficient Offline Reinforcement Learning: Data Diversity, Posterior Sampling, and Beyond
by: Nguyen-Tang, Thanh, et al.
Published: (2024)
by: Nguyen-Tang, Thanh, et al.
Published: (2024)
Large Language Model Agent for Hyper-Parameter Optimization
by: Liu, Siyi, et al.
Published: (2024)
by: Liu, Siyi, et al.
Published: (2024)
Flexible Bayesian Last Layer Models Using Implicit Priors and Diffusion Posterior Sampling
by: Xu, Jian, et al.
Published: (2024)
by: Xu, Jian, et al.
Published: (2024)
Dynamic Vocabulary Pruning: Stable LLM-RL by Taming the Tail
by: Li, Yingru, et al.
Published: (2025)
by: Li, Yingru, et al.
Published: (2025)
Efficient Approximate Posterior Sampling with Annealed Langevin Monte Carlo
by: Parulekar, Advait, et al.
Published: (2025)
by: Parulekar, Advait, et al.
Published: (2025)
FLAC: Maximum Entropy RL via Kinetic Energy Regularized Bridge Matching
by: Lv, Lei, et al.
Published: (2026)
by: Lv, Lei, et al.
Published: (2026)
Towards Scalable and Deep Graph Neural Networks via Noise Masking
by: Liang, Yuxuan, et al.
Published: (2024)
by: Liang, Yuxuan, et al.
Published: (2024)
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
Trust Region Masking for Long-Horizon LLM Reinforcement Learning
by: Li, Yingru, et al.
Published: (2025)
by: Li, Yingru, et al.
Published: (2025)
HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents
by: Li, Guankai, et al.
Published: (2026)
by: Li, Guankai, et al.
Published: (2026)
Divergence-Augmented Policy Optimization
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
Diffusion Posterior Sampling is Computationally Intractable
by: Gupta, Shivam, et al.
Published: (2024)
by: Gupta, Shivam, et al.
Published: (2024)
Coupled Data and Measurement Space Dynamics for Enhanced Diffusion Posterior Sampling
by: Hamidi, Shayan Mohajer, et al.
Published: (2025)
by: Hamidi, Shayan Mohajer, et al.
Published: (2025)
Distilled Thompson Sampling: Practical and Efficient Thompson Sampling via Imitation Learning
by: Namkoong, Hongseok, et al.
Published: (2020)
by: Namkoong, Hongseok, et al.
Published: (2020)
Agentic Unlearning: When LLM Agent Meets Machine Unlearning
by: Wang, Bin, et al.
Published: (2026)
by: Wang, Bin, et al.
Published: (2026)
SVRG and Beyond via Posterior Correction
by: Daheim, Nico, et al.
Published: (2025)
by: Daheim, Nico, et al.
Published: (2025)
Non-Linear Reinforcement Learning in Large Action Spaces: Structural Conditions and Sample-efficiency of Posterior Sampling
by: Agarwal, Alekh, et al.
Published: (2022)
by: Agarwal, Alekh, et al.
Published: (2022)
Adjoint Sampling: Highly Scalable Diffusion Samplers via Adjoint Matching
by: Havens, Aaron, et al.
Published: (2025)
by: Havens, Aaron, et al.
Published: (2025)
TensorHyper-VQC: A Tensor-Train-Guided Hypernetwork for Robust and Scalable Variational Quantum Computing
by: Qi, Jun, et al.
Published: (2025)
by: Qi, Jun, et al.
Published: (2025)
Fast and Robust Likelihood-Guided Diffusion Posterior Sampling with Amortized Variational Inference
by: Zheng, Léon, et al.
Published: (2026)
by: Zheng, Léon, et al.
Published: (2026)
Learning the Model While Learning Q: Finite-Time Sample Complexity of Online SyncMBQ
by: Lim, Han-Dong, et al.
Published: (2024)
by: Lim, Han-Dong, et al.
Published: (2024)
Fine-Tuning Diffusion Models for Molecular Generation via Reinforcement Learning and Fast Sampling
by: Lin, Guang, et al.
Published: (2026)
by: Lin, Guang, et al.
Published: (2026)
A Note on Hybrid Online Reinforcement and Imitation Learning for LLMs: Formulations and Algorithms
by: Li, Yingru, et al.
Published: (2025)
by: Li, Yingru, et al.
Published: (2025)
Bridging Geometric States via Geometric Diffusion Bridge
by: Luo, Shengjie, et al.
Published: (2024)
by: Luo, Shengjie, et al.
Published: (2024)
Similar Items
-
Scalable Exploration via Ensemble++
by: Li, Yingru, et al.
Published: (2024) -
Prior-dependent analysis of posterior sampling reinforcement learning with function approximation
by: Li, Yingru, et al.
Published: (2024) -
Optimistic Thompson Sampling for No-Regret Learning in Unknown Games
by: Li, Yingru, et al.
Published: (2024) -
Bridging Distributional and Risk-sensitive Reinforcement Learning with Provable Regret Bounds
by: Liang, Hao, et al.
Published: (2022) -
Logit Dynamics in Softmax Policy Gradient Methods
by: Li, Yingru
Published: (2025)