Counterfactual experience augmented off-policy reinforcement learning
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Sunbowen, Gong, Yicheng, Deng, Chao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MacLight: Multi-scene Aggregation Convolutional Learning for Traffic Signal Control
by: Lee, Sunbowen, et al.
Published: (2024)
by: Lee, Sunbowen, et al.
Published: (2024)
Soft $Q(λ)$: A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces
by: Mahajan, Pranav, et al.
Published: (2026)
by: Mahajan, Pranav, et al.
Published: (2026)
An advantage based policy transfer algorithm for reinforcement learning with measures of transferability
by: Alam, Md Ferdous, et al.
Published: (2023)
by: Alam, Md Ferdous, et al.
Published: (2023)
Delayed homomorphic reinforcement learning for environments with delayed feedback
by: Lee, Jongsoo, et al.
Published: (2026)
by: Lee, Jongsoo, et al.
Published: (2026)
Normalization and effective learning rates in reinforcement learning
by: Lyle, Clare, et al.
Published: (2024)
by: Lyle, Clare, et al.
Published: (2024)
Self-Distilled Disentangled Learning for Counterfactual Prediction
by: Li, Xinshu, et al.
Published: (2024)
by: Li, Xinshu, et al.
Published: (2024)
Multi-agent assignment via state augmented reinforcement learning
by: Agorio, Leopoldo, et al.
Published: (2024)
by: Agorio, Leopoldo, et al.
Published: (2024)
Curriculum reinforcement learning with measurable task representation learning
by: Wen, Yongyan, et al.
Published: (2026)
by: Wen, Yongyan, et al.
Published: (2026)
Rapid analysis of point-contact Andreev reflection spectra via machine learning with adaptive data augmentation
by: Lee, Dongik, et al.
Published: (2025)
by: Lee, Dongik, et al.
Published: (2025)
The geometry of invariant learning: an information-theoretic analysis of data augmentation and generalization
by: Bouyahia, Abdelali, et al.
Published: (2026)
by: Bouyahia, Abdelali, et al.
Published: (2026)
Bellman operator convergence enhancements in reinforcement learning algorithms
by: Kadurha, David Krame, et al.
Published: (2025)
by: Kadurha, David Krame, et al.
Published: (2025)
Causal prompting model-based offline reinforcement learning
by: Yu, Xuehui, et al.
Published: (2024)
by: Yu, Xuehui, et al.
Published: (2024)
Deep reinforcement learning with time-scale invariant memory
by: Kabir, Md Rysul, et al.
Published: (2024)
by: Kabir, Md Rysul, et al.
Published: (2024)
Offline reinforcement learning for job-shop scheduling problems
by: Echeverria, Imanol, et al.
Published: (2024)
by: Echeverria, Imanol, et al.
Published: (2024)
On Predicting Post-Click Conversion Rate via Counterfactual Inference
by: Ahn, Junhyung, et al.
Published: (2025)
by: Ahn, Junhyung, et al.
Published: (2025)
Leveraging weights signals -- Predicting and improving generalizability in reinforcement learning
by: Moulin, Olivier, et al.
Published: (2025)
by: Moulin, Olivier, et al.
Published: (2025)
Dynamic feature selection in medical predictive monitoring by reinforcement learning
by: Chen, Yutong, et al.
Published: (2024)
by: Chen, Yutong, et al.
Published: (2024)
Economic span selection of bridge based on deep reinforcement learning
by: Zhang, Leye, et al.
Published: (2024)
by: Zhang, Leye, et al.
Published: (2024)
Not all tokens are needed(NAT): token efficient reinforcement learning
by: Sang, Hejian, et al.
Published: (2026)
by: Sang, Hejian, et al.
Published: (2026)
Extreme value forecasting using relevance-based data augmentation with deep learning models
by: Hua, Junru, et al.
Published: (2025)
by: Hua, Junru, et al.
Published: (2025)
Deep deterministic policy gradient with symmetric data augmentation for lateral attitude tracking control of a fixed-wing aircraft
by: Li, Yifei, et al.
Published: (2024)
by: Li, Yifei, et al.
Published: (2024)
Automating proton PBS treatment planning for head and neck cancers using policy gradient-based deep reinforcement learning
by: Wang, Qingqing, et al.
Published: (2024)
by: Wang, Qingqing, et al.
Published: (2024)
An efficient deep reinforcement learning environment for flexible job-shop scheduling
by: Wu, Xinquan, et al.
Published: (2025)
by: Wu, Xinquan, et al.
Published: (2025)
Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning
by: Kobayashi, Seijin, et al.
Published: (2025)
by: Kobayashi, Seijin, et al.
Published: (2025)
On the consistency of hyper-parameter selection in value-based deep reinforcement learning
by: Obando-Ceron, Johan, et al.
Published: (2024)
by: Obando-Ceron, Johan, et al.
Published: (2024)
CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectives
by: Saghafian, Armin, et al.
Published: (2024)
by: Saghafian, Armin, et al.
Published: (2024)
Task diversity produces systematic transfer but inhibits continual reinforcement learning
by: Seth, Purab, et al.
Published: (2026)
by: Seth, Purab, et al.
Published: (2026)
Policy-shaped prediction: avoiding distractions in model-based reinforcement learning
by: Hutson, Miles, et al.
Published: (2024)
by: Hutson, Miles, et al.
Published: (2024)
Found-RL: foundation model-enhanced reinforcement learning for autonomous driving
by: Qu, Yansong, et al.
Published: (2026)
by: Qu, Yansong, et al.
Published: (2026)
Solving Bayesian inverse problems with diffusion priors and off-policy RL
by: Scimeca, Luca, et al.
Published: (2025)
by: Scimeca, Luca, et al.
Published: (2025)
Robust off-policy Reinforcement Learning via Soft Constrained Adversary
by: Nakanishi, Kosuke, et al.
Published: (2024)
by: Nakanishi, Kosuke, et al.
Published: (2024)
Survey on reinforcement learning for language processing
by: Uc-Cetina, Victor, et al.
Published: (2021)
by: Uc-Cetina, Victor, et al.
Published: (2021)
An approach of deep reinforcement learning for maximizing the net present value of stochastic projects
by: Xu, Wei, et al.
Published: (2025)
by: Xu, Wei, et al.
Published: (2025)
Designing an efficient and equitable humanitarian supply chain dynamically via reinforcement learning
by: Jin, Weijia
Published: (2025)
by: Jin, Weijia
Published: (2025)
Learning to summarize user information for personalized reinforcement learning from human feedback
by: Nam, Hyunji, et al.
Published: (2025)
by: Nam, Hyunji, et al.
Published: (2025)
Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs
by: Gan, Xingwei, et al.
Published: (2026)
by: Gan, Xingwei, et al.
Published: (2026)
Deep reinforcement learning for irrigation scheduling using high-dimensional sensor feedback
by: Saikai, Yuji, et al.
Published: (2023)
by: Saikai, Yuji, et al.
Published: (2023)
CHARME: A chain-based reinforcement learning approach for the minor embedding problem
by: Ngo, Hoang M., et al.
Published: (2024)
by: Ngo, Hoang M., et al.
Published: (2024)
Convergence of a model-free entropy-regularized inverse reinforcement learning algorithm
by: Renard, Titouan, et al.
Published: (2024)
by: Renard, Titouan, et al.
Published: (2024)
AFU: Actor-Free critic Updates in off-policy RL for continuous control
by: Perrin-Gilbert, Nicolas
Published: (2024)
by: Perrin-Gilbert, Nicolas
Published: (2024)
Similar Items
-
MacLight: Multi-scene Aggregation Convolutional Learning for Traffic Signal Control
by: Lee, Sunbowen, et al.
Published: (2024) -
Soft $Q(λ)$: A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces
by: Mahajan, Pranav, et al.
Published: (2026) -
An advantage based policy transfer algorithm for reinforcement learning with measures of transferability
by: Alam, Md Ferdous, et al.
Published: (2023) -
Delayed homomorphic reinforcement learning for environments with delayed feedback
by: Lee, Jongsoo, et al.
Published: (2026) -
Normalization and effective learning rates in reinforcement learning
by: Lyle, Clare, et al.
Published: (2024)