Fast Non-Episodic Finite-Horizon RL with K-Step Lookahead Thresholding
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Jiamin, Gan, Kyra |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Restless to Contextual: A Thresholding Bandit Reformulation For Finite-horizon Improvement
by: Xu, Jiamin, et al.
Published: (2025)
by: Xu, Jiamin, et al.
Published: (2025)
Integrating Causal DAGs in Deep RL: Activating Minimal Markovian States with Multi-Order Exposure
by: Xu, Jiamin, et al.
Published: (2026)
by: Xu, Jiamin, et al.
Published: (2026)
Reinforcement Learning with Multi-Step Lookahead Information Via Adaptive Batching
by: Merlis, Nadav
Published: (2026)
by: Merlis, Nadav
Published: (2026)
Prior-Aligned Meta-RL: Thompson Sampling with Learned Priors and Guarantees in Finite-Horizon MDPs
by: Zhou, Runlin, et al.
Published: (2025)
by: Zhou, Runlin, et al.
Published: (2025)
LoSAM: Local Search in Additive Noise Models with Mixed Mechanisms and General Noise for Global Causal Discovery
by: Hiremath, Sujai, et al.
Published: (2024)
by: Hiremath, Sujai, et al.
Published: (2024)
CSPI-MT: Calibrated Safe Policy Improvement with Multiple Testing for Threshold Policies
by: Cho, Brian M, et al.
Published: (2024)
by: Cho, Brian M, et al.
Published: (2024)
EARL-BO: Reinforcement Learning for Multi-Step Lookahead, High-Dimensional Bayesian Optimization
by: Cheon, Mujin, et al.
Published: (2024)
by: Cheon, Mujin, et al.
Published: (2024)
Next-Depth Lookahead Tree
by: Lee, Jaeho, et al.
Published: (2025)
by: Lee, Jaeho, et al.
Published: (2025)
Reinforcement Learning with Lookahead Information
by: Merlis, Nadav
Published: (2024)
by: Merlis, Nadav
Published: (2024)
Generalization and Optimization of SGD with Lookahead
by: Li, Kangcheng, et al.
Published: (2025)
by: Li, Kangcheng, et al.
Published: (2025)
Peeking with PEAK: Sequential, Nonparametric Composite Hypothesis Tests for Means of Multiple Data Streams
by: Cho, Brian, et al.
Published: (2024)
by: Cho, Brian, et al.
Published: (2024)
Data Augmentation for Continual RL via Adversarial Gradient Episodic Memory
by: Wu, Sihao, et al.
Published: (2024)
by: Wu, Sihao, et al.
Published: (2024)
Lookahead Counterfactual Fairness
by: Zuo, Zhiqun, et al.
Published: (2024)
by: Zuo, Zhiqun, et al.
Published: (2024)
Extending Differential Temporal Difference Methods for Episodic Problems
by: De Asis, Kris, et al.
Published: (2026)
by: De Asis, Kris, et al.
Published: (2026)
LEAP: Unlocking dLLM Parallelism via Lookahead Early-Convergence Token Detection
by: Zhang, Haohui, et al.
Published: (2026)
by: Zhang, Haohui, et al.
Published: (2026)
Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces
by: Yu, Shixing, et al.
Published: (2026)
by: Yu, Shixing, et al.
Published: (2026)
The Value of Reward Lookahead in Reinforcement Learning
by: Merlis, Nadav, et al.
Published: (2024)
by: Merlis, Nadav, et al.
Published: (2024)
Bridging RL Theory and Practice with the Effective Horizon
by: Laidlaw, Cassidy, et al.
Published: (2023)
by: Laidlaw, Cassidy, et al.
Published: (2023)
LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation
by: Ahn, Jinwoo, et al.
Published: (2026)
by: Ahn, Jinwoo, et al.
Published: (2026)
Federated Causal Inference in Healthcare: Methods, Challenges, and Applications
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
Deep Doubly Debiased Longitudinal Effect Estimation with ICE G-Computation
by: Chen, Wenxin, et al.
Published: (2026)
by: Chen, Wenxin, et al.
Published: (2026)
Smooth Multi-Policy Causal Effect Estimation in Longitudinal Settings
by: Chen, Wenxin, et al.
Published: (2026)
by: Chen, Wenxin, et al.
Published: (2026)
When Additive Noise Meets Unobserved Mediators: Bivariate Denoising Diffusion for Causal Discovery
by: Meier, Dominik, et al.
Published: (2025)
by: Meier, Dominik, et al.
Published: (2025)
MOSIC: Model-Agnostic Optimal Subgroup Identification with Multi-Constraint for Improved Reliability
by: Chen, Wenxin, et al.
Published: (2025)
by: Chen, Wenxin, et al.
Published: (2025)
Causal Attention with Lookahead Keys
by: Song, Zhuoqing, et al.
Published: (2025)
by: Song, Zhuoqing, et al.
Published: (2025)
Policy Mirror Descent with Lookahead
by: Protopapas, Kimon, et al.
Published: (2024)
by: Protopapas, Kimon, et al.
Published: (2024)
Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL
by: Wu, Ian, et al.
Published: (2026)
by: Wu, Ian, et al.
Published: (2026)
Horizon Reduction Makes RL Scalable
by: Park, Seohong, et al.
Published: (2025)
by: Park, Seohong, et al.
Published: (2025)
Lookahead Path Likelihood Optimization for Diffusion LLMs
by: Liu, Xuejie, et al.
Published: (2026)
by: Liu, Xuejie, et al.
Published: (2026)
The Horizon Threshold in Cooperative Multi-Agent Reward-Free Exploration
by: Barnea, Idan, et al.
Published: (2026)
by: Barnea, Idan, et al.
Published: (2026)
Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation
by: Kitamura, Toshinori, et al.
Published: (2025)
by: Kitamura, Toshinori, et al.
Published: (2025)
Kernel Debiased Plug-in Estimation: Simultaneous, Automated Debiasing without Influence Functions for Many Target Parameters
by: Cho, Brian, et al.
Published: (2023)
by: Cho, Brian, et al.
Published: (2023)
Reward Maximization for Pure Exploration: Minimax Optimal Good Arm Identification for Nonparametric Multi-Armed Bandits
by: Cho, Brian, et al.
Published: (2024)
by: Cho, Brian, et al.
Published: (2024)
Dual-Granularity Contrastive Reward via Generated Episodic Guidance for Efficient Embodied RL
by: Liu, Xin, et al.
Published: (2026)
by: Liu, Xin, et al.
Published: (2026)
Momentum Boosted Episodic Memory for Improving Learning in Long-Tailed RL Environments
by: Fernandes, Dolton, et al.
Published: (2025)
by: Fernandes, Dolton, et al.
Published: (2025)
Scaling Speculative Decoding with Lookahead Reasoning
by: Fu, Yichao, et al.
Published: (2025)
by: Fu, Yichao, et al.
Published: (2025)
Lookahead identification in adversarial bandits: accuracy and memory bounds
by: Brukhim, Nataly, et al.
Published: (2026)
by: Brukhim, Nataly, et al.
Published: (2026)
From Guess2Graph: When and How Can Unreliable Experts Safely Boost Causal Discovery in Finite Samples?
by: Hiremath, Sujai, et al.
Published: (2025)
by: Hiremath, Sujai, et al.
Published: (2025)
Lookahead Drifting Model
by: Zhang, Guoqiang, et al.
Published: (2026)
by: Zhang, Guoqiang, et al.
Published: (2026)
Contextual Bandits with Budgeted Information Reveal
by: Gan, Kyra, et al.
Published: (2023)
by: Gan, Kyra, et al.
Published: (2023)
Similar Items
-
From Restless to Contextual: A Thresholding Bandit Reformulation For Finite-horizon Improvement
by: Xu, Jiamin, et al.
Published: (2025) -
Integrating Causal DAGs in Deep RL: Activating Minimal Markovian States with Multi-Order Exposure
by: Xu, Jiamin, et al.
Published: (2026) -
Reinforcement Learning with Multi-Step Lookahead Information Via Adaptive Batching
by: Merlis, Nadav
Published: (2026) -
Prior-Aligned Meta-RL: Thompson Sampling with Learned Priors and Guarantees in Finite-Horizon MDPs
by: Zhou, Runlin, et al.
Published: (2025) -
LoSAM: Local Search in Additive Noise Models with Mixed Mechanisms and General Noise for Global Causal Discovery
by: Hiremath, Sujai, et al.
Published: (2024)