APRIL: Annotations for Policy evaluation with Reliable Inference from LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Mandyam, Aishwarya, Limaye, Kalyani, Engelhardt, Barbara E., Alsentzer, Emily |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Compositional Q-learning for electrolyte repletion with imbalanced patient sub-populations
by: Mandyam, Aishwarya, et al.
Published: (2021)
by: Mandyam, Aishwarya, et al.
Published: (2021)
Adaptive Interventions with User-Defined Goals for Health Behavior Change
by: Mandyam, Aishwarya, et al.
Published: (2023)
by: Mandyam, Aishwarya, et al.
Published: (2023)
Kernel Density Bayesian Inverse Reinforcement Learning
by: Mandyam, Aishwarya, et al.
Published: (2023)
by: Mandyam, Aishwarya, et al.
Published: (2023)
CANDOR: Counterfactual ANnotated DOubly Robust Off-Policy Evaluation
by: Mandyam, Aishwarya, et al.
Published: (2024)
by: Mandyam, Aishwarya, et al.
Published: (2024)
PERRY: Policy Evaluation with Confidence Intervals using Auxiliary Data
by: Mandyam, Aishwarya, et al.
Published: (2025)
by: Mandyam, Aishwarya, et al.
Published: (2025)
Preference-Guided Diffusion for Multi-Objective Offline Optimization
by: Annadani, Yashas, et al.
Published: (2025)
by: Annadani, Yashas, et al.
Published: (2025)
Sample Efficient Preference Alignment in LLMs via Active Exploration
by: Mehta, Viraj, et al.
Published: (2023)
by: Mehta, Viraj, et al.
Published: (2023)
Understanding Annotator Safety Policy with Interpretability
by: Oesterling, Alex, et al.
Published: (2026)
by: Oesterling, Alex, et al.
Published: (2026)
APRIL: Active Partial Rollouts in Reinforcement Learning to Tame Long-tail Generation
by: Zhou, Yuzhen, et al.
Published: (2025)
by: Zhou, Yuzhen, et al.
Published: (2025)
Non-Myopic Multi-Objective Bayesian Optimization
by: Belakaria, Syrine, et al.
Published: (2024)
by: Belakaria, Syrine, et al.
Published: (2024)
RadAnnotate: Large Language Models for Efficient and Reliable Radiology Report Annotation
by: Shetty, Saisha Pradeep, et al.
Published: (2026)
by: Shetty, Saisha Pradeep, et al.
Published: (2026)
Distance-informed Neural Processes
by: Venkataramanan, Aishwarya, et al.
Published: (2025)
by: Venkataramanan, Aishwarya, et al.
Published: (2025)
Multi-LLM Adaptive Conformal Inference for Reliable LLM Responses
by: Noh, Kangjun, et al.
Published: (2026)
by: Noh, Kangjun, et al.
Published: (2026)
Active Learning for Derivative-Based Global Sensitivity Analysis with Gaussian Processes
by: Belakaria, Syrine, et al.
Published: (2024)
by: Belakaria, Syrine, et al.
Published: (2024)
Geometric Scaling of Bayesian Inference in LLMs
by: Agarwal, Naman, et al.
Published: (2025)
by: Agarwal, Naman, et al.
Published: (2025)
Out of Context: Reliability in Multimodal Anomaly Detection Requires Contextual Inference
by: Wilkinghoff, Kevin, et al.
Published: (2026)
by: Wilkinghoff, Kevin, et al.
Published: (2026)
Partial Policy Gradients for RL in LLMs
by: Mathur, Puneet, et al.
Published: (2026)
by: Mathur, Puneet, et al.
Published: (2026)
Annotation-Assisted Learning of Treatment Policies From Multimodal Electronic Health Records
by: Arno, Henri, et al.
Published: (2025)
by: Arno, Henri, et al.
Published: (2025)
Sharpe Ratio-Guided Active Learning for Preference Optimization in RLHF
by: Belakaria, Syrine, et al.
Published: (2025)
by: Belakaria, Syrine, et al.
Published: (2025)
Reliable Policy Iteration: Performance Robustness Across Architecture and Environment Perturbations
by: Eshwar, S. R., et al.
Published: (2025)
by: Eshwar, S. R., et al.
Published: (2025)
Deep Reinforcement Learning for Inventory Networks: Toward Reliable Policy Optimization
by: Alvo, Matias, et al.
Published: (2023)
by: Alvo, Matias, et al.
Published: (2023)
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
by: Yang, Zhicheng, et al.
Published: (2026)
by: Yang, Zhicheng, et al.
Published: (2026)
Chemical Reaction Extraction from Long Patent Documents
by: Jadhav, Aishwarya, et al.
Published: (2024)
by: Jadhav, Aishwarya, et al.
Published: (2024)
Calibration-Aware Policy Optimization for Reasoning LLMs
by: Wang, Ziqi, et al.
Published: (2026)
by: Wang, Ziqi, et al.
Published: (2026)
HyPAC: Cost-Efficient LLMs-Human Hybrid Annotation with PAC Error Guarantees
by: Zeng, Hao, et al.
Published: (2026)
by: Zeng, Hao, et al.
Published: (2026)
A Real-time Anomaly Detection Using Convolutional Autoencoder with Dynamic Threshold
by: Maitra, Sarit, et al.
Published: (2024)
by: Maitra, Sarit, et al.
Published: (2024)
ImaginationPolicy: Towards Generalizable, Precise and Reliable End-to-End Policy for Robotic Manipulation
by: Lu, Dekun, et al.
Published: (2025)
by: Lu, Dekun, et al.
Published: (2025)
When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
by: Liu, Xiaogeng, et al.
Published: (2026)
by: Liu, Xiaogeng, et al.
Published: (2026)
EARL: Entropy-Aware RL Alignment of LLMs for Reliable RTL Code Generation
by: Shi, Jiahe, et al.
Published: (2025)
by: Shi, Jiahe, et al.
Published: (2025)
Extending Epistemic Uncertainty Beyond Parameters Would Assist in Designing Reliable LLMs
by: Nguyen-Hien, T. Duy, et al.
Published: (2025)
by: Nguyen-Hien, T. Duy, et al.
Published: (2025)
RAmBLA: A Framework for Evaluating the Reliability of LLMs as Assistants in the Biomedical Domain
by: Bolton, William James, et al.
Published: (2024)
by: Bolton, William James, et al.
Published: (2024)
From Imperfect Signals to Trustworthy Structure: Confidence-Aware Inference from Heterogeneous and Reliability-Varying Utility Data
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
Bootstrapping LLMs via Preference-Based Policy Optimization
by: Jia, Chen
Published: (2025)
by: Jia, Chen
Published: (2025)
One-Way Policy Optimization for Self-Evolving LLMs
by: Yang, Shuo, et al.
Published: (2026)
by: Yang, Shuo, et al.
Published: (2026)
APRIL: API Synthesis with Automatic Prompt Optimization and Reinforcement Learning
by: Zhong, Hua, et al.
Published: (2025)
by: Zhong, Hua, et al.
Published: (2025)
RL-Struct: A Lightweight Reinforcement Learning Framework for Reliable Structured Output in LLMs
by: Hu, Ruike, et al.
Published: (2025)
by: Hu, Ruike, et al.
Published: (2025)
Generalized Policy Gradient with History-Aware Decision Transformer for Reliable Routing over Graph Signals
by: Wei, Xing, et al.
Published: (2025)
by: Wei, Xing, et al.
Published: (2025)
When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?
by: Hatgis-Kessell, Stephane, et al.
Published: (2026)
by: Hatgis-Kessell, Stephane, et al.
Published: (2026)
Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs
by: Huang, Luke J., et al.
Published: (2026)
by: Huang, Luke J., et al.
Published: (2026)
PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training
by: Lv, Mingrui, et al.
Published: (2025)
by: Lv, Mingrui, et al.
Published: (2025)
Similar Items
-
Compositional Q-learning for electrolyte repletion with imbalanced patient sub-populations
by: Mandyam, Aishwarya, et al.
Published: (2021) -
Adaptive Interventions with User-Defined Goals for Health Behavior Change
by: Mandyam, Aishwarya, et al.
Published: (2023) -
Kernel Density Bayesian Inverse Reinforcement Learning
by: Mandyam, Aishwarya, et al.
Published: (2023) -
CANDOR: Counterfactual ANnotated DOubly Robust Off-Policy Evaluation
by: Mandyam, Aishwarya, et al.
Published: (2024) -
PERRY: Policy Evaluation with Confidence Intervals using Auxiliary Data
by: Mandyam, Aishwarya, et al.
Published: (2025)