Gespeichert in:
| Hauptverfasser: | Mukherjee, Subhojyoti, Hanna, Josiah P., Nowak, Robert |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2406.02165 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SPEED: Experimental Design for Policy Evaluation in Linear Heteroscedastic Bandits
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2023)
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2023)
Pretraining Decision Transformers with Reward Prediction for In-Context Multi-task Structured Bandit Learning
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
On-Policy Policy Gradient Reinforcement Learning Without On-Policy Sampling
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2023)
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2023)
Off-Policy Evaluation from Logged Human Feedback
von: Bhargava, Aniruddha, et al.
Veröffentlicht: (2024)
von: Bhargava, Aniruddha, et al.
Veröffentlicht: (2024)
An Empirical Study on the Power of Future Prediction in Partially Observable Environments
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
Centralized Adaptive Sampling for Reliable Co-Training of Independent Multi-Agent Policies
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2025)
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2025)
Adaptive Exploration for Data-Efficient General Value Function Evaluations
von: Jain, Arushi, et al.
Veröffentlicht: (2024)
von: Jain, Arushi, et al.
Veröffentlicht: (2024)
Partial Policy Gradients for RL in LLMs
von: Mathur, Puneet, et al.
Veröffentlicht: (2026)
von: Mathur, Puneet, et al.
Veröffentlicht: (2026)
Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation
von: Zhou, Hongyi, et al.
Veröffentlicht: (2025)
von: Zhou, Hongyi, et al.
Veröffentlicht: (2025)
Efficient and Interpretable Bandit Algorithms
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2023)
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2023)
Understanding when Dynamics-Invariant Data Augmentations Benefit Model-Free Reinforcement Learning Updates
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2023)
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2023)
MDP Planning as Policy Inference
von: Tolpin, David
Veröffentlicht: (2026)
von: Tolpin, David
Veröffentlicht: (2026)
Optimal Design for Human Preference Elicitation
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
SaVe-TAG: LLM-based Interpolation for Long-Tailed Text-Attributed Graphs
von: Wang, Leyao, et al.
Veröffentlicht: (2024)
von: Wang, Leyao, et al.
Veröffentlicht: (2024)
Agentic Planning with Reasoning for Image Styling via Offline RL
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2026)
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2026)
Logits are All We Need to Adapt Closed Models
von: Hiranandani, Gaurush, et al.
Veröffentlicht: (2025)
von: Hiranandani, Gaurush, et al.
Veröffentlicht: (2025)
Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2026)
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2026)
Guided Data Augmentation for Offline Reinforcement Learning and Imitation Learning
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2023)
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2023)
Sparsely Multimodal Data Fusion
von: Bjorgaard, Josiah
Veröffentlicht: (2024)
von: Bjorgaard, Josiah
Veröffentlicht: (2024)
Optimal Posterior Sampling for Policy Identification in Tabular Markov Decision Processes
von: Kone, Cyrille, et al.
Veröffentlicht: (2026)
von: Kone, Cyrille, et al.
Veröffentlicht: (2026)
Stable Offline Value Function Learning with Bisimulation-based Representations
von: Pavse, Brahma S., et al.
Veröffentlicht: (2024)
von: Pavse, Brahma S., et al.
Veröffentlicht: (2024)
Structured Evaluation of Synthetic Tabular Data
von: Yang, Scott Cheng-Hsin, et al.
Veröffentlicht: (2024)
von: Yang, Scott Cheng-Hsin, et al.
Veröffentlicht: (2024)
Experimental Design for Active Transductive Inference in Large Language Models
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024)
MIDST Challenge at SaTML 2025: Membership Inference over Diffusion-models-based Synthetic Tabular data
von: Shafieinejad, Masoumeh, et al.
Veröffentlicht: (2026)
von: Shafieinejad, Masoumeh, et al.
Veröffentlicht: (2026)
Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment?
von: Markgraf, Hannah, et al.
Veröffentlicht: (2025)
von: Markgraf, Hannah, et al.
Veröffentlicht: (2025)
Learning to Stabilize Online Reinforcement Learning in Unbounded State Spaces
von: Pavse, Brahma S., et al.
Veröffentlicht: (2023)
von: Pavse, Brahma S., et al.
Veröffentlicht: (2023)
VeRO: An Evaluation Harness for Agents to Optimize Agents
von: Ursekar, Varun, et al.
Veröffentlicht: (2026)
von: Ursekar, Varun, et al.
Veröffentlicht: (2026)
ICU-Sepsis: A Benchmark MDP Built from Real Medical Data
von: Choudhary, Kartik, et al.
Veröffentlicht: (2024)
von: Choudhary, Kartik, et al.
Veröffentlicht: (2024)
Sarcasm Detection in Tweets with BERT and GloVe Embeddings
von: Khatri, Akshay, et al.
Veröffentlicht: (2020)
von: Khatri, Akshay, et al.
Veröffentlicht: (2020)
Offline Policy Evaluation for Reinforcement Learning with Adaptively Collected Data
von: Madhow, Sunil, et al.
Veröffentlicht: (2023)
von: Madhow, Sunil, et al.
Veröffentlicht: (2023)
CoVeR: Conformal Calibration for Versatile and Reliable Autoregressive Next-Token Prediction
von: Chen, Yuzhu, et al.
Veröffentlicht: (2025)
von: Chen, Yuzhu, et al.
Veröffentlicht: (2025)
AdvantageFlow: Advantage-Weighted Least Squares for RL in Flow Models
von: Kveton, Branislav, et al.
Veröffentlicht: (2026)
von: Kveton, Branislav, et al.
Veröffentlicht: (2026)
Geometric Re-Analysis of Classical MDP Solving Algorithms
von: Mustafin, Arsenii, et al.
Veröffentlicht: (2025)
von: Mustafin, Arsenii, et al.
Veröffentlicht: (2025)
Memisis: Orchestrating and Evaluating Synthetic Data for Tabular Health Datasets
von: Nagesh, Nitish, et al.
Veröffentlicht: (2026)
von: Nagesh, Nitish, et al.
Veröffentlicht: (2026)
Evaluating Generative Models for Tabular Data: Novel Metrics and Benchmarking
von: Herurkar, Dayananda, et al.
Veröffentlicht: (2025)
von: Herurkar, Dayananda, et al.
Veröffentlicht: (2025)
A Systematic Evaluation of Generative Models on Tabular Transportation Data
von: Wang, Chengen, et al.
Veröffentlicht: (2025)
von: Wang, Chengen, et al.
Veröffentlicht: (2025)
FEST: A Unified Framework for Evaluating Synthetic Tabular Data
von: Niu, Weijie, et al.
Veröffentlicht: (2025)
von: Niu, Weijie, et al.
Veröffentlicht: (2025)
MDP Geometry, Normalization and Reward Balancing Solvers
von: Mustafin, Arsenii, et al.
Veröffentlicht: (2024)
von: Mustafin, Arsenii, et al.
Veröffentlicht: (2024)
Learning to Reason in LLMs by Expectation Maximization
von: Lee, Junghyun, et al.
Veröffentlicht: (2025)
von: Lee, Junghyun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SPEED: Experimental Design for Policy Evaluation in Linear Heteroscedastic Bandits
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2023) -
Pretraining Decision Transformers with Reward Prediction for In-Context Multi-task Structured Bandit Learning
von: Mukherjee, Subhojyoti, et al.
Veröffentlicht: (2024) -
On-Policy Policy Gradient Reinforcement Learning Without On-Policy Sampling
von: Corrado, Nicholas E., et al.
Veröffentlicht: (2023) -
Off-Policy Evaluation from Logged Human Feedback
von: Bhargava, Aniruddha, et al.
Veröffentlicht: (2024) -
An Empirical Study on the Power of Future Prediction in Partially Observable Environments
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)