Off-Policy Evaluation Using Information Borrowing and Context-Based Switching
Fuente:
arXiv
Saved in:
| Main Authors: | Dasgupta, Sutanoy, Niu, Yabo, Panaganti, Kishan, Kalathil, Dileep, Pati, Debdeep, Mallick, Bani |
|---|---|
| Format: | Preprint |
| Published: |
2021
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robust LLM Alignment via Distributionally Robust Direct Preference Optimization
by: Xu, Zaiyan, et al.
Published: (2025)
by: Xu, Zaiyan, et al.
Published: (2025)
On Quantification of Borrowing of Information in Hierarchical Bayesian Models
by: Ghosh, Prasenjit, et al.
Published: (2025)
by: Ghosh, Prasenjit, et al.
Published: (2025)
Factorized Fusion Shrinkage for Dynamic Relational Data
by: Zhao, Peng, et al.
Published: (2022)
by: Zhao, Peng, et al.
Published: (2022)
Hierarchical Bayesian Operator-induced Symbolic Regression Trees for Structural Learning of Scientific Expressions
by: Roy, Somjit, et al.
Published: (2025)
by: Roy, Somjit, et al.
Published: (2025)
Structured Optimal Variational Inference for Dynamic Latent Space Models
by: Zhao, Peng, et al.
Published: (2022)
by: Zhao, Peng, et al.
Published: (2022)
Tail-adaptive Bayesian shrinkage
by: Lee, Se Yoon, et al.
Published: (2020)
by: Lee, Se Yoon, et al.
Published: (2020)
Frequentist Regret Analysis of Gaussian Process Thompson Sampling via Fractional Posteriors
by: Roy, Somjit, et al.
Published: (2026)
by: Roy, Somjit, et al.
Published: (2026)
Model-Free Robust $ϕ$-Divergence Reinforcement Learning Using Both Offline and Online Data
by: Panaganti, Kishan, et al.
Published: (2024)
by: Panaganti, Kishan, et al.
Published: (2024)
Blocked Gibbs sampler for hierarchical Dirichlet processes
by: Das, Snigdha, et al.
Published: (2023)
by: Das, Snigdha, et al.
Published: (2023)
Stability of Sequential and Parallel Coordinate Ascent Variational Inference
by: Pati, Debdeep
Published: (2026)
by: Pati, Debdeep
Published: (2026)
Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity Guarantees
by: Ganguly, Sourav, et al.
Published: (2025)
by: Ganguly, Sourav, et al.
Published: (2025)
AttackGNN: Red-Teaming GNNs in Hardware Security Using Reinforcement Learning
by: Gohil, Vasudev, et al.
Published: (2024)
by: Gohil, Vasudev, et al.
Published: (2024)
Tractable Equilibrium Computation in Markov Games through Risk Aversion
by: Mazumdar, Eric, et al.
Published: (2024)
by: Mazumdar, Eric, et al.
Published: (2024)
Inference-Time Search Using Side Information for Diffusion-Based Image Reconstruction
by: Farahbakhsh, Mahdi, et al.
Published: (2025)
by: Farahbakhsh, Mahdi, et al.
Published: (2025)
MAVIS: Multi-Objective Alignment via Inference-Time Value-Guided Selection
by: Carleton, Jeremy, et al.
Published: (2025)
by: Carleton, Jeremy, et al.
Published: (2025)
Constrained Reweighting of Distributions: an Optimal Transport Approach
by: Chakraborty, Abhisek, et al.
Published: (2023)
by: Chakraborty, Abhisek, et al.
Published: (2023)
Adaptive Outer-Loop Control of Quadrotors via Reinforcement Learning
by: Saj, Vishnu, et al.
Published: (2026)
by: Saj, Vishnu, et al.
Published: (2026)
Global-Local Dirichlet Processes for Identifying Pan-Cancer Subpopulations Using Both Shared and Cancer-Specific Data
by: Chakrabarti, Arhit, et al.
Published: (2025)
by: Chakrabarti, Arhit, et al.
Published: (2025)
PowerMamba: A Deep State Space Model and Comprehensive Benchmark for Time Series Prediction in Electric Power Systems
by: Menati, Ali, et al.
Published: (2024)
by: Menati, Ali, et al.
Published: (2024)
Federated Ensemble-Directed Offline Reinforcement Learning
by: Rengarajan, Desik, et al.
Published: (2023)
by: Rengarajan, Desik, et al.
Published: (2023)
Hybrid Transfer Reinforcement Learning: Provable Sample Efficiency from Shifted-Dynamics Data
by: Qu, Chengrui, et al.
Published: (2024)
by: Qu, Chengrui, et al.
Published: (2024)
In-Context Learning for Gradient-Free Receiver Adaptation: Principles, Applications, and Theory
by: Zecchin, Matteo, et al.
Published: (2025)
by: Zecchin, Matteo, et al.
Published: (2025)
Adaptive finite element type decomposition of Gaussian processes
by: Kim, Jaehoan, et al.
Published: (2025)
by: Kim, Jaehoan, et al.
Published: (2025)
Risk-Averse Total-Reward Reinforcement Learning
by: Su, Xihong, et al.
Published: (2025)
by: Su, Xihong, et al.
Published: (2025)
Clustering Context in Off-Policy Evaluation
by: Guzman-Olivares, Daniel, et al.
Published: (2025)
by: Guzman-Olivares, Daniel, et al.
Published: (2025)
KL-regularization Itself is Differentially Private in Bandits and RLHF
by: Zhang, Yizhou, et al.
Published: (2025)
by: Zhang, Yizhou, et al.
Published: (2025)
FIT-GNN: Faster Inference Time for GNNs that 'FIT' in Memory Using Coarsening
by: Roy, Shubhajit, et al.
Published: (2024)
by: Roy, Shubhajit, et al.
Published: (2024)
A Generalized Tangent Approximation based Variational Inference Framework for Strongly Super-Gaussian Likelihoods
by: Roy, Somjit, et al.
Published: (2025)
by: Roy, Somjit, et al.
Published: (2025)
Distributionally Robust Constrained Reinforcement Learning under Strong Duality
by: Zhang, Zhengfei, et al.
Published: (2024)
by: Zhang, Zhengfei, et al.
Published: (2024)
Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning
by: Panaganti, Kishan, et al.
Published: (2026)
by: Panaganti, Kishan, et al.
Published: (2026)
Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values
by: Yu, Dian, et al.
Published: (2025)
by: Yu, Dian, et al.
Published: (2025)
Context-Action Embedding Learning for Off-Policy Evaluation in Contextual Bandits
by: Chandak, Kushagra, et al.
Published: (2025)
by: Chandak, Kushagra, et al.
Published: (2025)
GSpaRC: Gaussian Splatting for Real-time Reconstruction of RF Channels
by: Nukapotula, Bhavya Sai, et al.
Published: (2025)
by: Nukapotula, Bhavya Sai, et al.
Published: (2025)
Optimistic World Models: Efficient Exploration in Model-Based Deep Reinforcement Learning
by: Mete, Akshay, et al.
Published: (2026)
by: Mete, Akshay, et al.
Published: (2026)
Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning
by: Liang, Zhenwen, et al.
Published: (2025)
by: Liang, Zhenwen, et al.
Published: (2025)
VaSST: Variational Inference for Symbolic Regression using Soft Symbolic Trees
by: Roy, Somjit, et al.
Published: (2026)
by: Roy, Somjit, et al.
Published: (2026)
Cross-Validated Off-Policy Evaluation
by: Cief, Matej, et al.
Published: (2024)
by: Cief, Matej, et al.
Published: (2024)
Structured Reinforcement Learning for Media Streaming at the Wireless Edge
by: Bura, Archana, et al.
Published: (2024)
by: Bura, Archana, et al.
Published: (2024)
Off-Policy Evaluation for Ranking Policies under Deterministic Logging Policies
by: Tanaka, Koichi, et al.
Published: (2026)
by: Tanaka, Koichi, et al.
Published: (2026)
Guided Self-Evolving LLMs with Minimal Human Supervision
by: Yu, Wenhao, et al.
Published: (2025)
by: Yu, Wenhao, et al.
Published: (2025)
Similar Items
-
Robust LLM Alignment via Distributionally Robust Direct Preference Optimization
by: Xu, Zaiyan, et al.
Published: (2025) -
On Quantification of Borrowing of Information in Hierarchical Bayesian Models
by: Ghosh, Prasenjit, et al.
Published: (2025) -
Factorized Fusion Shrinkage for Dynamic Relational Data
by: Zhao, Peng, et al.
Published: (2022) -
Hierarchical Bayesian Operator-induced Symbolic Regression Trees for Structural Learning of Scientific Expressions
by: Roy, Somjit, et al.
Published: (2025) -
Structured Optimal Variational Inference for Dynamic Latent Space Models
by: Zhao, Peng, et al.
Published: (2022)