Optimal Policy Minimum Bayesian Risk
Fuente:
arXiv
Saved in:
| Main Authors: | Astudillo, Ramón Fernandez, Sultan, Md Arafat, Trivedi, Aashka, El-Kurdi, Yousef, Naseem, Tahira, Florian, Radu, Roukos, Salim |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Refinement of Language Models from External Proxy Metrics Feedback
by: Ramji, Keshav, et al.
Published: (2024)
by: Ramji, Keshav, et al.
Published: (2024)
Latent Principle Discovery for Language Model Self-Improvement
by: Ramji, Keshav, et al.
Published: (2025)
by: Ramji, Keshav, et al.
Published: (2025)
Latency and Token-Aware Test-Time Compute
by: Huang, Jenny Y., et al.
Published: (2025)
by: Huang, Jenny Y., et al.
Published: (2025)
An Empirical Investigation into the Effect of Parameter Choices in Knowledge Distillation
by: Sultan, Md Arafat, et al.
Published: (2024)
by: Sultan, Md Arafat, et al.
Published: (2024)
BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback
by: Pandey, Gaurav, et al.
Published: (2024)
by: Pandey, Gaurav, et al.
Published: (2024)
DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math?
by: Lee, Young-Suk, et al.
Published: (2026)
by: Lee, Young-Suk, et al.
Published: (2026)
Graph-based Uncertainty Metrics for Long-form Language Model Outputs
by: Jiang, Mingjian, et al.
Published: (2024)
by: Jiang, Mingjian, et al.
Published: (2024)
Confidence-Weighted Token Set Cover for Early Hypothesis Pruning in Self-Consistency
by: Sultan, Md Arafat, et al.
Published: (2025)
by: Sultan, Md Arafat, et al.
Published: (2025)
Efficient Models for the Detection of Hate, Abuse and Profanity
by: Tillmann, Christoph, et al.
Published: (2024)
by: Tillmann, Christoph, et al.
Published: (2024)
Prompts as Auto-Optimized Training Hyperparameters: Training Best-in-Class IR Models from Scratch with 10 Gold Labels
by: Xian, Jasper, et al.
Published: (2024)
by: Xian, Jasper, et al.
Published: (2024)
Multi-Document Grounded Multi-Turn Synthetic Dialog Generation
by: Lee, Young-Suk, et al.
Published: (2024)
by: Lee, Young-Suk, et al.
Published: (2024)
Bayesian preference elicitation for decision support in multiobjective optimization
by: Huber, Felix, et al.
Published: (2025)
by: Huber, Felix, et al.
Published: (2025)
Thinking Without Words: Efficient Latent Reasoning with Abstract Chain-of-Thought
by: Ramji, Keshav, et al.
Published: (2026)
by: Ramji, Keshav, et al.
Published: (2026)
Learning Conformal Abstention Policies for Adaptive Risk Management in Large Language and Vision-Language Models
by: Tayebati, Sina, et al.
Published: (2025)
by: Tayebati, Sina, et al.
Published: (2025)
Formally Specifying the High-Level Behavior of LLM-Based Agents
by: Crouse, Maxwell, et al.
Published: (2023)
by: Crouse, Maxwell, et al.
Published: (2023)
Structured Chain-of-Thought Prompting for Few-Shot Generation of Content-Grounded QA Conversations
by: Sultan, Md Arafat, et al.
Published: (2024)
by: Sultan, Md Arafat, et al.
Published: (2024)
Generalizing Scaling Laws for Dense and Sparse Large Language Models
by: Hossain, Md Arafat, et al.
Published: (2025)
by: Hossain, Md Arafat, et al.
Published: (2025)
BadImplant: Injection-based Multi-Targeted Graph Backdoor Attack
by: Khan, Md Nabi Newaz, et al.
Published: (2026)
by: Khan, Md Nabi Newaz, et al.
Published: (2026)
Uncertainty-Aware Decoding with Minimum Bayes Risk
by: Daheim, Nico, et al.
Published: (2025)
by: Daheim, Nico, et al.
Published: (2025)
Insertion Based Sequence Generation with Learnable Order Dynamics
by: Patel, Dhruvesh, et al.
Published: (2026)
by: Patel, Dhruvesh, et al.
Published: (2026)
A Simulated Annealing-Based Multiobjective Optimization Algorithm for Minimum Weight Minimum Connected Dominating Set Problem
by: Dahmri, Hayet, et al.
Published: (2023)
by: Dahmri, Hayet, et al.
Published: (2023)
SAFE-KD: Risk-Controlled Early-Exit Distillation for Vision Backbones
by: Khazem, Salim
Published: (2026)
by: Khazem, Salim
Published: (2026)
Optimistic Exploration for Risk-Averse Constrained Reinforcement Learning
by: McCarthy, James, et al.
Published: (2025)
by: McCarthy, James, et al.
Published: (2025)
An Offline Risk-aware Policy Selection Method for Bayesian Markov Decision Processes
by: Angelotti, Giorgio, et al.
Published: (2021)
by: Angelotti, Giorgio, et al.
Published: (2021)
Agreement-Constrained Probabilistic Minimum Bayes Risk Decoding
by: Natsumi, Koki, et al.
Published: (2025)
by: Natsumi, Koki, et al.
Published: (2025)
Optimal Policy Learning with Observational Data in Multi-Action Scenarios: Estimation, Risk Preference, and Potential Failures
by: Cerulli, Giovanni
Published: (2024)
by: Cerulli, Giovanni
Published: (2024)
HAEPO: History-Aggregated Exploratory Policy Optimization
by: Trivedi, Gaurish, et al.
Published: (2025)
by: Trivedi, Gaurish, et al.
Published: (2025)
Sufficient Conditions for Stability of Minimum-Norm Interpolating Deep ReLU Networks
by: Harzli, Ouns El, et al.
Published: (2026)
by: Harzli, Ouns El, et al.
Published: (2026)
Improved Sample Complexity For Diffusion Model Training Without Empirical Risk Minimizer Access
by: Gaur, Mudit, et al.
Published: (2025)
by: Gaur, Mudit, et al.
Published: (2025)
Bayesian Modeling for Uncertainty Management in Financial Risk Forecasting and Compliance
by: Mamun, Sharif Al, et al.
Published: (2025)
by: Mamun, Sharif Al, et al.
Published: (2025)
Information as Structural Alignment: A Dynamical Theory of Continual Learning
by: Negulescu, Radu
Published: (2026)
by: Negulescu, Radu
Published: (2026)
Optimal partition of feature using Bayesian classifier
by: Vishwakarma, Sanjay, et al.
Published: (2023)
by: Vishwakarma, Sanjay, et al.
Published: (2023)
Bayesian Regression for Predicting Subscription to Bank Term Deposits in Direct Marketing Campaigns
by: Tanvir, Muhammad Farhan, et al.
Published: (2024)
by: Tanvir, Muhammad Farhan, et al.
Published: (2024)
Rank-1 Approximation of Inverse Fisher for Natural Policy Gradients in Deep Reinforcement Learning
by: Huo, Yingxiao, et al.
Published: (2026)
by: Huo, Yingxiao, et al.
Published: (2026)
BEACON: Bayesian Optimal Stopping for Efficient LLM Sampling
by: Wan, Guangya, et al.
Published: (2025)
by: Wan, Guangya, et al.
Published: (2025)
PAC-Bayesian Reinforcement Learning Trains Generalizable Policies
by: Zitouni, Abdelkrim, et al.
Published: (2025)
by: Zitouni, Abdelkrim, et al.
Published: (2025)
Large Language Models for Drug Overdose Prediction from Longitudinal Medical Records
by: Nahian, Md Sultan Al, et al.
Published: (2025)
by: Nahian, Md Sultan Al, et al.
Published: (2025)
Hotel Booking Cancellation Prediction Using Applied Bayesian Models
by: Jishan, Md Asifuzzaman, et al.
Published: (2024)
by: Jishan, Md Asifuzzaman, et al.
Published: (2024)
Off-OAB: Off-Policy Policy Gradient Method with Optimal Action-Dependent Baseline
by: Meng, Wenjia, et al.
Published: (2024)
by: Meng, Wenjia, et al.
Published: (2024)
Measures of Variability for Risk-averse Policy Gradient
by: Luo, Yudong, et al.
Published: (2025)
by: Luo, Yudong, et al.
Published: (2025)
Similar Items
-
Self-Refinement of Language Models from External Proxy Metrics Feedback
by: Ramji, Keshav, et al.
Published: (2024) -
Latent Principle Discovery for Language Model Self-Improvement
by: Ramji, Keshav, et al.
Published: (2025) -
Latency and Token-Aware Test-Time Compute
by: Huang, Jenny Y., et al.
Published: (2025) -
An Empirical Investigation into the Effect of Parameter Choices in Knowledge Distillation
by: Sultan, Md Arafat, et al.
Published: (2024) -
BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedback
by: Pandey, Gaurav, et al.
Published: (2024)