Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
Fuente:
arXiv
Saved in:
| Main Authors: | Chittepu, Yaswanth, Metevier, Blossom, Schwarzer, Will, Hoag, Austin, Niekum, Scott, Thomas, Philip S. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Margin RLHF via Preference over Preferences
by: Chittepu, Yaswanth, et al.
Published: (2025)
by: Chittepu, Yaswanth, et al.
Published: (2025)
Safe RLHF Beyond Expectation: Stochastic Dominance for Universal Spectral Risk Control
by: Chittepu, Yaswanth, et al.
Published: (2026)
by: Chittepu, Yaswanth, et al.
Published: (2026)
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
by: Rafailov, Rafael, et al.
Published: (2024)
by: Rafailov, Rafael, et al.
Published: (2024)
Evaluation-Aware Reinforcement Learning
by: Deshmukh, Shripad Vilasrao, et al.
Published: (2025)
by: Deshmukh, Shripad Vilasrao, et al.
Published: (2025)
Confidence Adjusted Surprise Measure for Active Resourceful Trials (CA-SMART): A Data-driven Active Learning Framework for Accelerating Material Discovery under Resource Constraints
by: Raihan, Ahmed Shoyeb, et al.
Published: (2025)
by: Raihan, Ahmed Shoyeb, et al.
Published: (2025)
Conformal Safety Monitoring for Flight Testing: A Case Study in Data-Driven Safety Learning
by: Feldman, Aaron O., et al.
Published: (2025)
by: Feldman, Aaron O., et al.
Published: (2025)
Unified Representation of Genomic and Biomedical Concepts through Multi-Task, Multi-Source Contrastive Learning
by: Yuan, Hongyi, et al.
Published: (2024)
by: Yuan, Hongyi, et al.
Published: (2024)
ML-Tool-Bench: Tool-Augmented Planning for ML Tasks
by: Chittepu, Yaswanth, et al.
Published: (2025)
by: Chittepu, Yaswanth, et al.
Published: (2025)
The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety
by: Springer, Max, et al.
Published: (2026)
by: Springer, Max, et al.
Published: (2026)
Beyond the Hype: Embeddings vs. Prompting for Multiclass Classification Tasks
by: Kokkodis, Marios, et al.
Published: (2025)
by: Kokkodis, Marios, et al.
Published: (2025)
ICE-ID: A Novel Historical Census Dataset for Longitudinal Identity Resolution
by: de Carvalho, Gonçalo Hora, et al.
Published: (2025)
by: de Carvalho, Gonçalo Hora, et al.
Published: (2025)
Domain-Shift-Aware Conformal Prediction for Large Language Models
by: Lin, Zhexiao, et al.
Published: (2025)
by: Lin, Zhexiao, et al.
Published: (2025)
Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning
by: Zhou, Cai, et al.
Published: (2026)
by: Zhou, Cai, et al.
Published: (2026)
Reliable and Efficient Amortized Model-based Evaluation
by: Truong, Sang, et al.
Published: (2025)
by: Truong, Sang, et al.
Published: (2025)
Uncertainty-Aware Adaptation of Large Language Models for Protein-Protein Interaction Analysis
by: Jantre, Sanket, et al.
Published: (2025)
by: Jantre, Sanket, et al.
Published: (2025)
Beyond Words: How Large Language Models Perform in Quantitative Management Problem-Solving
by: Kuzmanko, Jonathan
Published: (2025)
by: Kuzmanko, Jonathan
Published: (2025)
Augmented Risk Prediction for the Onset of Alzheimer's Disease from Electronic Health Records with Large Language Models
by: Wang, Jiankun, et al.
Published: (2024)
by: Wang, Jiankun, et al.
Published: (2024)
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
by: Lee, Harrison, et al.
Published: (2023)
by: Lee, Harrison, et al.
Published: (2023)
Distributionally Robust Reinforcement Learning with Human Feedback
by: Mandal, Debmalya, et al.
Published: (2025)
by: Mandal, Debmalya, et al.
Published: (2025)
Contrastive Preference Learning: Learning from Human Feedback without RL
by: Hejna, Joey, et al.
Published: (2023)
by: Hejna, Joey, et al.
Published: (2023)
Parameter Efficient Reinforcement Learning from Human Feedback
by: Sidahmed, Hakim, et al.
Published: (2024)
by: Sidahmed, Hakim, et al.
Published: (2024)
Conditional diffusions for amortized neural posterior estimation
by: Chen, Tianyu, et al.
Published: (2024)
by: Chen, Tianyu, et al.
Published: (2024)
Removing Spurious Correlation from Neural Network Interpretations
by: Fotouhi, Milad, et al.
Published: (2024)
by: Fotouhi, Milad, et al.
Published: (2024)
Language Models as Causal Effect Generators
by: Bynum, Lucius E. J., et al.
Published: (2024)
by: Bynum, Lucius E. J., et al.
Published: (2024)
Cross-Domain Energy-Guided Diffusion Generation for Off-Dynamics Reinforcement Learning
by: Yang, Yu, et al.
Published: (2026)
by: Yang, Yu, et al.
Published: (2026)
Training ML Models with Predictable Failures
by: Schwarzer, Will, et al.
Published: (2026)
by: Schwarzer, Will, et al.
Published: (2026)
Dual RL: Unification and New Methods for Reinforcement and Imitation Learning
by: Sikchi, Harshit, et al.
Published: (2023)
by: Sikchi, Harshit, et al.
Published: (2023)
An Optimal Discriminator Weighted Imitation Perspective for Reinforcement Learning
by: Xu, Haoran, et al.
Published: (2025)
by: Xu, Haoran, et al.
Published: (2025)
Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
by: Ackermann, Johannes, et al.
Published: (2025)
by: Ackermann, Johannes, et al.
Published: (2025)
Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble
by: Zhang, Shun, et al.
Published: (2024)
by: Zhang, Shun, et al.
Published: (2024)
Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
by: Poddar, Sriyash, et al.
Published: (2024)
by: Poddar, Sriyash, et al.
Published: (2024)
FactsR: A Safer Method for Producing High Quality Healthcare Documentation
by: Hansen, Victor Petrén Bach, et al.
Published: (2025)
by: Hansen, Victor Petrén Bach, et al.
Published: (2025)
Reinforcement Learning with Backtracking Feedback
by: Sel, Bilgehan, et al.
Published: (2026)
by: Sel, Bilgehan, et al.
Published: (2026)
Chitchat with AI: Understand the supply chain carbon disclosure of companies worldwide through Large Language Model
by: Hang, Haotian, et al.
Published: (2025)
by: Hang, Haotian, et al.
Published: (2025)
Contextual Phenotyping of Pediatric Sepsis Cohort Using Large Language Models
by: Nagori, Aditya, et al.
Published: (2025)
by: Nagori, Aditya, et al.
Published: (2025)
FreDF: Learning to Forecast in the Frequency Domain
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
A Statistical Theory of Regularization-Based Continual Learning
by: Zhao, Xuyang, et al.
Published: (2024)
by: Zhao, Xuyang, et al.
Published: (2024)
RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
by: Chaudhari, Shreyas, et al.
Published: (2024)
by: Chaudhari, Shreyas, et al.
Published: (2024)
Curious Causality-Seeking Agents Learn Meta Causal World
by: Zhao, Zhiyu, et al.
Published: (2025)
by: Zhao, Zhiyu, et al.
Published: (2025)
Reinforcement Learning from Human Feedback with Active Queries
by: Ji, Kaixuan, et al.
Published: (2024)
by: Ji, Kaixuan, et al.
Published: (2024)
Similar Items
-
Adaptive Margin RLHF via Preference over Preferences
by: Chittepu, Yaswanth, et al.
Published: (2025) -
Safe RLHF Beyond Expectation: Stochastic Dominance for Universal Spectral Risk Control
by: Chittepu, Yaswanth, et al.
Published: (2026) -
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
by: Rafailov, Rafael, et al.
Published: (2024) -
Evaluation-Aware Reinforcement Learning
by: Deshmukh, Shripad Vilasrao, et al.
Published: (2025) -
Confidence Adjusted Surprise Measure for Active Resourceful Trials (CA-SMART): A Data-driven Active Learning Framework for Accelerating Material Discovery under Resource Constraints
by: Raihan, Ahmed Shoyeb, et al.
Published: (2025)