Predicting Long Term Sequential Policy Value Using Softer Surrogates
Fuente:
arXiv
Saved in:
| Main Authors: | Nam, Hyunji, Nie, Allen, Gao, Ge, Syrgkanis, Vasilis, Brunskill, Emma |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?
by: Hatgis-Kessell, Stephane, et al.
Published: (2026)
by: Hatgis-Kessell, Stephane, et al.
Published: (2026)
Short-Long Policy Evaluation with Novel Actions
by: Nam, Hyunji Alex, et al.
Published: (2024)
by: Nam, Hyunji Alex, et al.
Published: (2024)
Learning Causal Representations from General Environments: Identifiability and Intrinsic Ambiguity
by: Jin, Jikai, et al.
Published: (2023)
by: Jin, Jikai, et al.
Published: (2023)
Consistency of Neural Causal Partial Identification
by: Tan, Jiyuan, et al.
Published: (2024)
by: Tan, Jiyuan, et al.
Published: (2024)
Preference Learning with Response Time: Robust Losses and Guarantees
by: Sawarni, Ayush, et al.
Published: (2025)
by: Sawarni, Ayush, et al.
Published: (2025)
Towards efficient representation identification in supervised learning
by: Ahuja, Kartik, et al.
Published: (2022)
by: Ahuja, Kartik, et al.
Published: (2022)
Empirical Analysis of Model Selection for Heterogeneous Causal Effect Estimation
by: Mahajan, Divyat, et al.
Published: (2022)
by: Mahajan, Divyat, et al.
Published: (2022)
Prescriptive Scaling Reveals the Evolution of Language Model Capabilities
by: Zhang, Hanlin, et al.
Published: (2026)
by: Zhang, Hanlin, et al.
Published: (2026)
Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
by: Jin, Jikai, et al.
Published: (2025)
by: Jin, Jikai, et al.
Published: (2025)
Experiment Planning with Function Approximation
by: Pacchiano, Aldo, et al.
Published: (2024)
by: Pacchiano, Aldo, et al.
Published: (2024)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)
Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data
by: Nam, Hyunji, et al.
Published: (2026)
by: Nam, Hyunji, et al.
Published: (2026)
Adaptive Interventions with User-Defined Goals for Health Behavior Change
by: Mandyam, Aishwarya, et al.
Published: (2023)
by: Mandyam, Aishwarya, et al.
Published: (2023)
Data Augmentation Policy Search for Long-Term Forecasting
by: Nochumsohn, Liran, et al.
Published: (2024)
by: Nochumsohn, Liran, et al.
Published: (2024)
OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple Estimators
by: Nie, Allen, et al.
Published: (2024)
by: Nie, Allen, et al.
Published: (2024)
Predicting Long-Term Student Outcomes from Short-Term EdTech Log Data
by: Gao, Ge, et al.
Published: (2024)
by: Gao, Ge, et al.
Published: (2024)
Integrating Sequential and Relational Modeling for User Events: Datasets and Prediction Tasks
by: Fathony, Rizal, et al.
Published: (2025)
by: Fathony, Rizal, et al.
Published: (2025)
Learning to summarize user information for personalized reinforcement learning from human feedback
by: Nam, Hyunji, et al.
Published: (2025)
by: Nam, Hyunji, et al.
Published: (2025)
Long-Term Outlier Prediction Through Outlier Score Modeling
by: Aoki, Yuma, et al.
Published: (2026)
by: Aoki, Yuma, et al.
Published: (2026)
Efficient RL for optimizing conversation level outcomes with an LLM-based tutor
by: Nam, Hyunji, et al.
Published: (2025)
by: Nam, Hyunji, et al.
Published: (2025)
PolicyLong: Towards On-Policy Context Extension
by: Jia, Junlong, et al.
Published: (2026)
by: Jia, Junlong, et al.
Published: (2026)
Inference on Optimal Policy Values and Other Irregular Functionals via Softmax Smoothing
by: Whitehouse, Justin, et al.
Published: (2025)
by: Whitehouse, Justin, et al.
Published: (2025)
Adaptive Estimation and Inference in Conditional Moment Models via the Discrepancy Principle
by: Tan, Jiyuan, et al.
Published: (2026)
by: Tan, Jiyuan, et al.
Published: (2026)
SPO: Sequential Monte Carlo Policy Optimisation
by: Macfarlane, Matthew V, et al.
Published: (2024)
by: Macfarlane, Matthew V, et al.
Published: (2024)
Sequential Policy Gradient for Adaptive Hyperparameter Optimization
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
TN-SHAP-G: Graph-Structured Tensor Network Surrogates for Shapley Values and Interactions
by: Heidari, Farzaneh, et al.
Published: (2026)
by: Heidari, Farzaneh, et al.
Published: (2026)
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
by: Huang, Chenghua, et al.
Published: (2025)
by: Huang, Chenghua, et al.
Published: (2025)
Predicting Outcomes in Video Games with Long Short Term Memory Networks
by: Chulajata, Kittimate, et al.
Published: (2024)
by: Chulajata, Kittimate, et al.
Published: (2024)
Federated Learning for Early Prediction of EV Charging Demand
by: Perifanis, Vasilis, et al.
Published: (2026)
by: Perifanis, Vasilis, et al.
Published: (2026)
Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients
by: Thrampoulidis, Christos, et al.
Published: (2025)
by: Thrampoulidis, Christos, et al.
Published: (2025)
Enhancing Next Destination Prediction: A Novel Long Short-Term Memory Neural Network Approach Using Real-World Airline Data
by: Salihoglu, Salih, et al.
Published: (2024)
by: Salihoglu, Salih, et al.
Published: (2024)
ISMRNN: An Implicitly Segmented RNN Method with Mamba for Long-Term Time Series Forecasting
by: Zhao, GaoXiang, et al.
Published: (2024)
by: Zhao, GaoXiang, et al.
Published: (2024)
Surrogate Neural Networks Local Stability for Aircraft Predictive Maintenance
by: Ducoffe, Mélanie, et al.
Published: (2024)
by: Ducoffe, Mélanie, et al.
Published: (2024)
xMTrans: Temporal Attentive Cross-Modality Fusion Transformer for Long-Term Traffic Prediction
by: Ung, Huy Quang, et al.
Published: (2024)
by: Ung, Huy Quang, et al.
Published: (2024)
Hiformer: Hybrid Frequency Feature Enhancement Inverted Transformer for Long-Term Wind Power Prediction
by: Wan, Chongyang, et al.
Published: (2024)
by: Wan, Chongyang, et al.
Published: (2024)
Long-Term Prediction Accuracy Improvement of Data-Driven Medium-Range Global Weather Forecast
by: Hu, Yifan, et al.
Published: (2024)
by: Hu, Yifan, et al.
Published: (2024)
STM3: Mixture of Multiscale Mamba for Long-Term Spatio-Temporal Time-Series Prediction
by: Chen, Haolong, et al.
Published: (2025)
by: Chen, Haolong, et al.
Published: (2025)
Sequential Monte Carlo for Policy Optimization in Continuous POMDPs
by: Abdulsamad, Hany, et al.
Published: (2025)
by: Abdulsamad, Hany, et al.
Published: (2025)
Solar Active Regions Emergence Prediction Using Long Short-Term Memory Networks
by: Kasapis, Spiridon, et al.
Published: (2024)
by: Kasapis, Spiridon, et al.
Published: (2024)
Disentangled Parameter-Efficient Linear Model for Long-Term Time Series Forecasting
by: Zhao, Yuang, et al.
Published: (2024)
by: Zhao, Yuang, et al.
Published: (2024)
Similar Items
-
When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?
by: Hatgis-Kessell, Stephane, et al.
Published: (2026) -
Short-Long Policy Evaluation with Novel Actions
by: Nam, Hyunji Alex, et al.
Published: (2024) -
Learning Causal Representations from General Environments: Identifiability and Intrinsic Ambiguity
by: Jin, Jikai, et al.
Published: (2023) -
Consistency of Neural Causal Partial Identification
by: Tan, Jiyuan, et al.
Published: (2024) -
Preference Learning with Response Time: Robust Losses and Guarantees
by: Sawarni, Ayush, et al.
Published: (2025)