BaNEL: Exploration Posteriors for Generative Modeling Using Only Negative Rewards
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Sangyun, Amos, Brandon, Fanti, Giulia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving the Training of Rectified Flows
von: Lee, Sangyun, et al.
Veröffentlicht: (2024)
von: Lee, Sangyun, et al.
Veröffentlicht: (2024)
Truncated Consistency Models
von: Lee, Sangyun, et al.
Veröffentlicht: (2024)
von: Lee, Sangyun, et al.
Veröffentlicht: (2024)
Value-Based Pre-Training with Downstream Feedback
von: Ke, Shuqi, et al.
Veröffentlicht: (2026)
von: Ke, Shuqi, et al.
Veröffentlicht: (2026)
MaxShapley: Towards Incentive-compatible Generative Search with Fair Context Attribution
von: Patel, Sara, et al.
Veröffentlicht: (2025)
von: Patel, Sara, et al.
Veröffentlicht: (2025)
On amortizing convex conjugates for optimal transport
von: Amos, Brandon
Veröffentlicht: (2022)
von: Amos, Brandon
Veröffentlicht: (2022)
Tutorial on amortized optimization
von: Amos, Brandon
Veröffentlicht: (2022)
von: Amos, Brandon
Veröffentlicht: (2022)
Mixture-of-Linear-Experts for Long-term Time Series Forecasting
von: Ni, Ronghao, et al.
Veröffentlicht: (2023)
von: Ni, Ronghao, et al.
Veröffentlicht: (2023)
Notes on the Reward Representation of Posterior Updates
von: Ortega, Pedro A.
Veröffentlicht: (2026)
von: Ortega, Pedro A.
Veröffentlicht: (2026)
Online Intrinsic Rewards for Decision Making Agents from Large Language Model Feedback
von: Zheng, Qinqing, et al.
Veröffentlicht: (2024)
von: Zheng, Qinqing, et al.
Veröffentlicht: (2024)
Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference
von: Lee, Sangyun, et al.
Veröffentlicht: (2026)
von: Lee, Sangyun, et al.
Veröffentlicht: (2026)
Calibrating Large Language Models Using Their Generations Only
von: Ulmer, Dennis, et al.
Veröffentlicht: (2024)
von: Ulmer, Dennis, et al.
Veröffentlicht: (2024)
Not Only Rewards But Also Constraints: Applications on Legged Robot Locomotion
von: Kim, Yunho, et al.
Veröffentlicht: (2023)
von: Kim, Yunho, et al.
Veröffentlicht: (2023)
Uncertainty-Aware Reward-Free Exploration with General Function Approximation
von: Zhang, Junkai, et al.
Veröffentlicht: (2024)
von: Zhang, Junkai, et al.
Veröffentlicht: (2024)
Exploration by Random Reward Perturbation
von: Ma, Haozhe, et al.
Veröffentlicht: (2025)
von: Ma, Haozhe, et al.
Veröffentlicht: (2025)
Exploration Through Introspection: A Self-Aware Reward Model
von: Petrowski, Michael, et al.
Veröffentlicht: (2026)
von: Petrowski, Michael, et al.
Veröffentlicht: (2026)
Grounding LTL Tasks in Sub-Symbolic RL Environments for Zero-Shot Generalization
von: Pannacci, Matteo, et al.
Veröffentlicht: (2026)
von: Pannacci, Matteo, et al.
Veröffentlicht: (2026)
Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
von: Zhang, Qingru, et al.
Veröffentlicht: (2025)
von: Zhang, Qingru, et al.
Veröffentlicht: (2025)
Reward Generation via Large Vision-Language Model in Offline Reinforcement Learning
von: Lee, Younghwan, et al.
Veröffentlicht: (2025)
von: Lee, Younghwan, et al.
Veröffentlicht: (2025)
Pretrained deep models outperform GBDTs in Learning-To-Rank under label scarcity
von: Hou, Charlie, et al.
Veröffentlicht: (2023)
von: Hou, Charlie, et al.
Veröffentlicht: (2023)
Adaptive Exploration for Multi-Reward Multi-Policy Evaluation
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
Provably Efficient Exploration in Reward Machines with Low Regret
von: Bourel, Hippolyte, et al.
Veröffentlicht: (2024)
von: Bourel, Hippolyte, et al.
Veröffentlicht: (2024)
Characterizing the Training Dynamics of Private Fine-tuning with Langevin diffusion
von: Ke, Shuqi, et al.
Veröffentlicht: (2024)
von: Ke, Shuqi, et al.
Veröffentlicht: (2024)
TaskMet: Task-Driven Metric Learning for Model Learning
von: Bansal, Dishank, et al.
Veröffentlicht: (2023)
von: Bansal, Dishank, et al.
Veröffentlicht: (2023)
RIFT: Repurposing Negative Samples via Reward-Informed Fine-Tuning
von: Liu, Zehua, et al.
Veröffentlicht: (2026)
von: Liu, Zehua, et al.
Veröffentlicht: (2026)
SemiReward: A General Reward Model for Semi-supervised Learning
von: Li, Siyuan, et al.
Veröffentlicht: (2023)
von: Li, Siyuan, et al.
Veröffentlicht: (2023)
Cross-Domain Imitation Learning via Optimal Transport
von: Fickinger, Arnaud, et al.
Veröffentlicht: (2021)
von: Fickinger, Arnaud, et al.
Veröffentlicht: (2021)
Posterior Mean Matching: Generative Modeling through Online Bayesian Inference
von: Salazar, Sebastian, et al.
Veröffentlicht: (2024)
von: Salazar, Sebastian, et al.
Veröffentlicht: (2024)
Progressive Inference: Explaining Decoder-Only Sequence Classification Models Using Intermediate Predictions
von: Kariyappa, Sanjay, et al.
Veröffentlicht: (2024)
von: Kariyappa, Sanjay, et al.
Veröffentlicht: (2024)
Flexible Bayesian Last Layer Models Using Implicit Priors and Diffusion Posterior Sampling
von: Xu, Jian, et al.
Veröffentlicht: (2024)
von: Xu, Jian, et al.
Veröffentlicht: (2024)
Temporal Representations for Exploration: Learning Complex Exploratory Behavior without Extrinsic Rewards
von: Mohamed, Faisal, et al.
Veröffentlicht: (2026)
von: Mohamed, Faisal, et al.
Veröffentlicht: (2026)
DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management
von: Su, Xuerui, et al.
Veröffentlicht: (2025)
von: Su, Xuerui, et al.
Veröffentlicht: (2025)
IRPM: Intergroup Relative Preference Modeling for Pointwise Generative Reward Models
von: Song, Haonan, et al.
Veröffentlicht: (2026)
von: Song, Haonan, et al.
Veröffentlicht: (2026)
Mitigating Label Shift in Tabular In-Context Learning via Test-Time Posterior Adjustment
von: Lee, Seunghan
Veröffentlicht: (2026)
von: Lee, Seunghan
Veröffentlicht: (2026)
OMG-RL:Offline Model-based Guided Reward Learning for Heparin Treatment
von: Lim, Yooseok, et al.
Veröffentlicht: (2024)
von: Lim, Yooseok, et al.
Veröffentlicht: (2024)
MIR: Efficient Exploration in Episodic Multi-Agent Reinforcement Learning via Mutual Intrinsic Reward
von: Chen, Kesheng, et al.
Veröffentlicht: (2025)
von: Chen, Kesheng, et al.
Veröffentlicht: (2025)
Process Reward Models That Think
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2025)
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2025)
CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward Modeling
von: Liu, Dengcan, et al.
Veröffentlicht: (2026)
von: Liu, Dengcan, et al.
Veröffentlicht: (2026)
Confidence-Controlled Exploration: Efficient Sparse-Reward Policy Learning for Robot Navigation
von: Patel, Bhrij, et al.
Veröffentlicht: (2023)
von: Patel, Bhrij, et al.
Veröffentlicht: (2023)
Safe Exploration Using Bayesian World Models and Log-Barrier Optimization
von: As, Yarden, et al.
Veröffentlicht: (2024)
von: As, Yarden, et al.
Veröffentlicht: (2024)
DreamSmooth: Improving Model-based Reinforcement Learning via Reward Smoothing
von: Lee, Vint, et al.
Veröffentlicht: (2023)
von: Lee, Vint, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Improving the Training of Rectified Flows
von: Lee, Sangyun, et al.
Veröffentlicht: (2024) -
Truncated Consistency Models
von: Lee, Sangyun, et al.
Veröffentlicht: (2024) -
Value-Based Pre-Training with Downstream Feedback
von: Ke, Shuqi, et al.
Veröffentlicht: (2026) -
MaxShapley: Towards Incentive-compatible Generative Search with Fair Context Attribution
von: Patel, Sara, et al.
Veröffentlicht: (2025) -
On amortizing convex conjugates for optimal transport
von: Amos, Brandon
Veröffentlicht: (2022)