Saved in:
| Main Authors: | Takahashi, Tatsuki, Maru, Chihiro, Shoji, Hiroko |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2502.08993 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Off-Policy Evaluation of Ranking Policies via Embedding-Space User Behavior Modeling
by: Takahashi, Tatsuki, et al.
Published: (2025)
by: Takahashi, Tatsuki, et al.
Published: (2025)
RATFM: Retrieval-augmented Time Series Foundation Model for Anomaly Detection
by: Maru, Chihiro, et al.
Published: (2025)
by: Maru, Chihiro, et al.
Published: (2025)
Off-Policy Evaluation Under Nonignorable Missing Data
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
Off-Policy Evaluation and Learning for Survival Outcomes under Censoring
by: Kubota, Kohsuke, et al.
Published: (2026)
by: Kubota, Kohsuke, et al.
Published: (2026)
Abstract Reward Processes: Leveraging State Abstraction for Consistent Off-Policy Evaluation
by: Chaudhari, Shreyas, et al.
Published: (2024)
by: Chaudhari, Shreyas, et al.
Published: (2024)
Doubly Calibrated Estimator for Recommendation on Data Missing Not At Random
by: Kweon, Wonbin, et al.
Published: (2024)
by: Kweon, Wonbin, et al.
Published: (2024)
Coherent Off-Policy Improvement of Large Behavior Models with Learned Rewards
by: Scherer, Christian, et al.
Published: (2026)
by: Scherer, Christian, et al.
Published: (2026)
On (Normalised) Discounted Cumulative Gain as an Off-Policy Evaluation Metric for Top-$n$ Recommendation
by: Jeunen, Olivier, et al.
Published: (2023)
by: Jeunen, Olivier, et al.
Published: (2023)
Cross-Validated Off-Policy Evaluation
by: Cief, Matej, et al.
Published: (2024)
by: Cief, Matej, et al.
Published: (2024)
When is Off-Policy Evaluation (Reward Modeling) Useful in Contextual Bandits? A Data-Centric Perspective
by: Sun, Hao, et al.
Published: (2023)
by: Sun, Hao, et al.
Published: (2023)
Off-Policy Evaluation for Ranking Policies under Deterministic Logging Policies
by: Tanaka, Koichi, et al.
Published: (2026)
by: Tanaka, Koichi, et al.
Published: (2026)
Doubly-Robust Off-Policy Evaluation with Estimated Logging Policy
by: Lee, Kyungbok, et al.
Published: (2024)
by: Lee, Kyungbok, et al.
Published: (2024)
Long-term Off-Policy Evaluation and Learning
by: Saito, Yuta, et al.
Published: (2024)
by: Saito, Yuta, et al.
Published: (2024)
A Graph-Enhanced Deep-Reinforcement Learning Framework for the Aircraft Landing Problem
by: Maru, Vatsal
Published: (2025)
by: Maru, Vatsal
Published: (2025)
Clustering Context in Off-Policy Evaluation
by: Guzman-Olivares, Daniel, et al.
Published: (2025)
by: Guzman-Olivares, Daniel, et al.
Published: (2025)
Concept-driven Off Policy Evaluation
by: Majumdar, Ritam, et al.
Published: (2024)
by: Majumdar, Ritam, et al.
Published: (2024)
Off-Policy Evaluation of Slate Bandit Policies via Optimizing Abstraction
by: Kiyohara, Haruka, et al.
Published: (2024)
by: Kiyohara, Haruka, et al.
Published: (2024)
Off-Policy Evaluation from Logged Human Feedback
by: Bhargava, Aniruddha, et al.
Published: (2024)
by: Bhargava, Aniruddha, et al.
Published: (2024)
RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning
by: Hisaki, Yukinari, et al.
Published: (2024)
by: Hisaki, Yukinari, et al.
Published: (2024)
Off-Policy Evaluation and Learning for Matching Markets
by: Hayashi, Yudai, et al.
Published: (2025)
by: Hayashi, Yudai, et al.
Published: (2025)
Learning Action Embeddings for Off-Policy Evaluation
by: Cief, Matej, et al.
Published: (2023)
by: Cief, Matej, et al.
Published: (2023)
Breaking the Curse of Repulsion: Optimistic Distributionally Robust Policy Optimization for Off-Policy Generative Recommendation
by: Jiang, Jie, et al.
Published: (2026)
by: Jiang, Jie, et al.
Published: (2026)
Off-Policy Reinforcement Learning with High Dimensional Reward
by: Lee, Dong Neuck, et al.
Published: (2024)
by: Lee, Dong Neuck, et al.
Published: (2024)
Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction
by: Guan, Zhong, et al.
Published: (2026)
by: Guan, Zhong, et al.
Published: (2026)
Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
by: Ackermann, Johannes, et al.
Published: (2025)
by: Ackermann, Johannes, et al.
Published: (2025)
Log-Sum-Exponential Estimator for Off-Policy Evaluation and Learning
by: Behnamnia, Armin, et al.
Published: (2025)
by: Behnamnia, Armin, et al.
Published: (2025)
Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning
by: Sakhi, Otmane, et al.
Published: (2024)
by: Sakhi, Otmane, et al.
Published: (2024)
CANDOR: Counterfactual ANnotated DOubly Robust Off-Policy Evaluation
by: Mandyam, Aishwarya, et al.
Published: (2024)
by: Mandyam, Aishwarya, et al.
Published: (2024)
Missing Pattern Recognized Diffusion Imputation Model for Missing Not At Random
by: Sim, Gyuwon, et al.
Published: (2026)
by: Sim, Gyuwon, et al.
Published: (2026)
Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
by: He, Haoran, et al.
Published: (2025)
by: He, Haoran, et al.
Published: (2025)
IntOPE: Off-Policy Evaluation in the Presence of Interference
by: Bai, Yuqi, et al.
Published: (2024)
by: Bai, Yuqi, et al.
Published: (2024)
Data Poisoning Attacks on Off-Policy Policy Evaluation Methods
by: Lobo, Elita, et al.
Published: (2024)
by: Lobo, Elita, et al.
Published: (2024)
Distributional Off-Policy Evaluation with Deep Quantile Process Regression
by: Kuang, Qi, et al.
Published: (2026)
by: Kuang, Qi, et al.
Published: (2026)
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
by: Lee, Haanvid, et al.
Published: (2024)
by: Lee, Haanvid, et al.
Published: (2024)
Context-Action Embedding Learning for Off-Policy Evaluation in Contextual Bandits
by: Chandak, Kushagra, et al.
Published: (2025)
by: Chandak, Kushagra, et al.
Published: (2025)
DOLCE: Decomposing Off-Policy Evaluation/Learning into Lagged and Current Effects
by: Tamano, Shu
Published: (2025)
by: Tamano, Shu
Published: (2025)
From Weighting to Modeling: A Nonparametric Estimator for Off-Policy Evaluation
by: Zhu, Rong J. B.
Published: (2026)
by: Zhu, Rong J. B.
Published: (2026)
Off-Policy Evaluation Using Information Borrowing and Context-Based Switching
by: Dasgupta, Sutanoy, et al.
Published: (2021)
by: Dasgupta, Sutanoy, et al.
Published: (2021)
Off-Policy Evaluation and Learning for the Future under Non-Stationarity
by: Shimizu, Tatsuhiro, et al.
Published: (2025)
by: Shimizu, Tatsuhiro, et al.
Published: (2025)
Robustness of Refugee-Matching Gains to Off-Policy Evaluation Choices
by: Bansak, Kirk, et al.
Published: (2026)
by: Bansak, Kirk, et al.
Published: (2026)
Similar Items
-
Off-Policy Evaluation of Ranking Policies via Embedding-Space User Behavior Modeling
by: Takahashi, Tatsuki, et al.
Published: (2025) -
RATFM: Retrieval-augmented Time Series Foundation Model for Anomaly Detection
by: Maru, Chihiro, et al.
Published: (2025) -
Off-Policy Evaluation Under Nonignorable Missing Data
by: Wang, Han, et al.
Published: (2025) -
Off-Policy Evaluation and Learning for Survival Outcomes under Censoring
by: Kubota, Kohsuke, et al.
Published: (2026) -
Abstract Reward Processes: Leveraging State Abstraction for Consistent Off-Policy Evaluation
by: Chaudhari, Shreyas, et al.
Published: (2024)