Guardado en:
| Autores principales: | Nazir, Mohammad Saif, Banerjee, Chayan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2503.22723 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Mining--Gym: A Configurable RL Benchmarking Environment for Truck Dispatch Scheduling
por: Banerjee, Chayan, et al.
Publicado: (2025)
por: Banerjee, Chayan, et al.
Publicado: (2025)
Assessing the Zero-Shot Capabilities of LLMs for Action Evaluation in RL
por: Pignatelli, Eduardo, et al.
Publicado: (2024)
por: Pignatelli, Eduardo, et al.
Publicado: (2024)
Online Human Action Detection during Escorting
por: Mondal, Siddhartha, et al.
Publicado: (2025)
por: Mondal, Siddhartha, et al.
Publicado: (2025)
Adaptive Querying for Reward Learning from Human Feedback
por: Anand, Yashwanthi, et al.
Publicado: (2024)
por: Anand, Yashwanthi, et al.
Publicado: (2024)
Contrastive Preference Learning: Learning from Human Feedback without RL
por: Hejna, Joey, et al.
Publicado: (2023)
por: Hejna, Joey, et al.
Publicado: (2023)
Automatic Reward Shaping from Multi-Objective Human Heuristics
por: Xie, Yuqing, et al.
Publicado: (2025)
por: Xie, Yuqing, et al.
Publicado: (2025)
Boosting LLM Reasoning via Human-Inspired Reward Shaping
por: Lin, Wenze, et al.
Publicado: (2026)
por: Lin, Wenze, et al.
Publicado: (2026)
RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards
por: Wang, Zhilin, et al.
Publicado: (2025)
por: Wang, Zhilin, et al.
Publicado: (2025)
Can Differentiable Decision Trees Enable Interpretable Reward Learning from Human Feedback?
por: Kalra, Akansha, et al.
Publicado: (2023)
por: Kalra, Akansha, et al.
Publicado: (2023)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
por: Ackermann, Johannes, et al.
Publicado: (2026)
por: Ackermann, Johannes, et al.
Publicado: (2026)
Zero-Shot Instruction Following in RL via Structured LTL Representations
por: Giuri, Mattia, et al.
Publicado: (2025)
por: Giuri, Mattia, et al.
Publicado: (2025)
Zero-Shot Instruction Following in RL via Structured LTL Representations
por: Jackermeier, Mathias, et al.
Publicado: (2026)
por: Jackermeier, Mathias, et al.
Publicado: (2026)
Verifier-Free RL for LLMs via Intrinsic Gradient-Norm Reward
por: Wen, Xuexiang, et al.
Publicado: (2026)
por: Wen, Xuexiang, et al.
Publicado: (2026)
CogGuide: Human-Like Guidance for Zero-Shot Omni-Modal Reasoning
por: Shou, Zhou-Peng, et al.
Publicado: (2025)
por: Shou, Zhou-Peng, et al.
Publicado: (2025)
Feedback Loops With Language Models Drive In-Context Reward Hacking
por: Pan, Alexander, et al.
Publicado: (2024)
por: Pan, Alexander, et al.
Publicado: (2024)
A Unified Linear Programming Framework for Offline Reward Learning from Human Demonstrations and Feedback
por: Kim, Kihyun, et al.
Publicado: (2024)
por: Kim, Kihyun, et al.
Publicado: (2024)
IPBC: An Interactive Projection-Based Framework for Human-in-the-Loop Semi-Supervised Clustering of High-Dimensional Data
por: Zare, Mohammad
Publicado: (2026)
por: Zare, Mohammad
Publicado: (2026)
Unsupervised Zero-Shot Reinforcement Learning via Functional Reward Encodings
por: Frans, Kevin, et al.
Publicado: (2024)
por: Frans, Kevin, et al.
Publicado: (2024)
Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning
por: Rocamonde, Juan, et al.
Publicado: (2023)
por: Rocamonde, Juan, et al.
Publicado: (2023)
What Can You Do When You Have Zero Rewards During RL?
por: Prakash, Jatin, et al.
Publicado: (2025)
por: Prakash, Jatin, et al.
Publicado: (2025)
Grounding LTL Tasks in Sub-Symbolic RL Environments for Zero-Shot Generalization
por: Pannacci, Matteo, et al.
Publicado: (2026)
por: Pannacci, Matteo, et al.
Publicado: (2026)
Debate as Reward: A Multi-Agent Reward System for Scientific Ideation via RL Post-Training
por: Salimi, Moein, et al.
Publicado: (2026)
por: Salimi, Moein, et al.
Publicado: (2026)
Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
por: Ackermann, Johannes, et al.
Publicado: (2025)
por: Ackermann, Johannes, et al.
Publicado: (2025)
Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble
por: Zhang, Shun, et al.
Publicado: (2024)
por: Zhang, Shun, et al.
Publicado: (2024)
TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics
por: Chen, Shirui, et al.
Publicado: (2026)
por: Chen, Shirui, et al.
Publicado: (2026)
Towards Improving Reward Design in RL: A Reward Alignment Metric for RL Practitioners
por: Muslimani, Calarina, et al.
Publicado: (2025)
por: Muslimani, Calarina, et al.
Publicado: (2025)
Enhancing Financial Fraud Detection with Human-in-the-Loop Feedback and Feedback Propagation
por: Kadam, Prashank
Publicado: (2024)
por: Kadam, Prashank
Publicado: (2024)
Fusing Reward and Dueling Feedback in Stochastic Bandits
por: Wang, Xuchuang, et al.
Publicado: (2025)
por: Wang, Xuchuang, et al.
Publicado: (2025)
Prompt Optimization with Human Feedback
por: Lin, Xiaoqiang, et al.
Publicado: (2024)
por: Lin, Xiaoqiang, et al.
Publicado: (2024)
Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
por: Liang, Zhiyuan, et al.
Publicado: (2025)
por: Liang, Zhiyuan, et al.
Publicado: (2025)
Anomaly-Aware Vision-Language Adapters for Zero-Shot Anomaly Detection
por: Aqeel, Muhammad, et al.
Publicado: (2026)
por: Aqeel, Muhammad, et al.
Publicado: (2026)
Omni-Thinker: Scaling Multi-Task RL in LLMs with Hybrid Reward and Task Scheduling
por: Li, Derek, et al.
Publicado: (2025)
por: Li, Derek, et al.
Publicado: (2025)
Enhancing LLMs for Physics Problem-Solving using Reinforcement Learning with Human-AI Feedback
por: Anand, Avinash, et al.
Publicado: (2024)
por: Anand, Avinash, et al.
Publicado: (2024)
Dynamics-Aligned Latent Imagination in Contextual World Models for Zero-Shot Generalization
por: Röder, Frank, et al.
Publicado: (2025)
por: Röder, Frank, et al.
Publicado: (2025)
Zero-Shot Robustification of Zero-Shot Models
por: Adila, Dyah, et al.
Publicado: (2023)
por: Adila, Dyah, et al.
Publicado: (2023)
Shaping Zero-Shot Coordination via State Blocking
por: Kang, Mingu, et al.
Publicado: (2026)
por: Kang, Mingu, et al.
Publicado: (2026)
Capturing Individual Human Preferences with Reward Features
por: Barreto, André, et al.
Publicado: (2025)
por: Barreto, André, et al.
Publicado: (2025)
Bootstrapped Reward Shaping
por: Adamczyk, Jacob, et al.
Publicado: (2025)
por: Adamczyk, Jacob, et al.
Publicado: (2025)
Human in the Latent Loop (HILL): Interactively Guiding Model Training Through Human Intuition
por: Geissler, Daniel, et al.
Publicado: (2025)
por: Geissler, Daniel, et al.
Publicado: (2025)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
por: Hatgis-Kessell, Stephane, et al.
Publicado: (2025)
por: Hatgis-Kessell, Stephane, et al.
Publicado: (2025)
Ejemplares similares
-
Mining--Gym: A Configurable RL Benchmarking Environment for Truck Dispatch Scheduling
por: Banerjee, Chayan, et al.
Publicado: (2025) -
Assessing the Zero-Shot Capabilities of LLMs for Action Evaluation in RL
por: Pignatelli, Eduardo, et al.
Publicado: (2024) -
Online Human Action Detection during Escorting
por: Mondal, Siddhartha, et al.
Publicado: (2025) -
Adaptive Querying for Reward Learning from Human Feedback
por: Anand, Yashwanthi, et al.
Publicado: (2024) -
Contrastive Preference Learning: Learning from Human Feedback without RL
por: Hejna, Joey, et al.
Publicado: (2023)