Reducing Oracle Feedback with Vision-Language Embeddings for Preference-Based RL
Fuente:
arXiv
Saved in:
| Main Authors: | Ghosh, Udita, Raychaudhuri, Dripta S., Li, Jiachen, Karydis, Konstantinos, Roy-Chowdhury, Amit |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning
by: Ghosh, Udita, et al.
Published: (2025)
by: Ghosh, Udita, et al.
Published: (2025)
Robust Offline Imitation Learning from Diverse Auxiliary Data
by: Ghosh, Udita, et al.
Published: (2024)
by: Ghosh, Udita, et al.
Published: (2024)
Toward Autonomous Laboratory Safety Monitoring with Vision Language Models: Learning to See Hazards Through Scene Structure
by: Chakraborty, Trishna, et al.
Published: (2026)
by: Chakraborty, Trishna, et al.
Published: (2026)
CONTRAST: Continual Multi-source Adaptation to Dynamic Distributions
by: Ahmed, Sk Miraj, et al.
Published: (2024)
by: Ahmed, Sk Miraj, et al.
Published: (2024)
HEAL: An Empirical Study on Hallucinations in Embodied Agents Driven by Large Language Models
by: Chakraborty, Trishna, et al.
Published: (2025)
by: Chakraborty, Trishna, et al.
Published: (2025)
Towards Source-Free Machine Unlearning
by: Ahmed, Sk Miraj, et al.
Published: (2025)
by: Ahmed, Sk Miraj, et al.
Published: (2025)
Offline Constrained RLHF with Multiple Preference Oracles
by: Latham, Brenden, et al.
Published: (2026)
by: Latham, Brenden, et al.
Published: (2026)
Contrastive Preference Learning: Learning from Human Feedback without RL
by: Hejna, Joey, et al.
Published: (2023)
by: Hejna, Joey, et al.
Published: (2023)
Making RL with Preference-based Feedback Efficient via Randomization
by: Wu, Runzhe, et al.
Published: (2023)
by: Wu, Runzhe, et al.
Published: (2023)
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
by: Wang, Yufei, et al.
Published: (2024)
by: Wang, Yufei, et al.
Published: (2024)
TractOracle: towards an anatomically-informed reward function for RL-based tractography
by: Théberge, Antoine, et al.
Published: (2024)
by: Théberge, Antoine, et al.
Published: (2024)
A data balancing approach towards design of an expert system for Heart Disease Prediction
by: Karmakar, Rahul, et al.
Published: (2024)
by: Karmakar, Rahul, et al.
Published: (2024)
RL Token: Bootstrapping Online RL with Vision-Language-Action Models
by: Xu, Charles, et al.
Published: (2026)
by: Xu, Charles, et al.
Published: (2026)
Exploring the robustness of TractOracle methods in RL-based tractography
by: Levesque, Jeremi, et al.
Published: (2025)
by: Levesque, Jeremi, et al.
Published: (2025)
$π_\texttt{RL}$: Online RL Fine-tuning for Flow-based Vision-Language-Action Models
by: Chen, Kang, et al.
Published: (2025)
by: Chen, Kang, et al.
Published: (2025)
Visual Alignment of Medical Vision-Language Models for Grounded Radiology Report Generation
by: Bose, Sarosij, et al.
Published: (2025)
by: Bose, Sarosij, et al.
Published: (2025)
Oracle-Robust Online Alignment for Large Language Models
by: Li, Zimeng, et al.
Published: (2026)
by: Li, Zimeng, et al.
Published: (2026)
Vision-based Xylem Wetness Classification in Stem Water Potential Determination
by: Peiris, Pamodya, et al.
Published: (2024)
by: Peiris, Pamodya, et al.
Published: (2024)
ComPO: Preference Alignment via Comparison Oracles
by: Chen, Peter, et al.
Published: (2025)
by: Chen, Peter, et al.
Published: (2025)
Drive My Way: Preference Alignment of Vision-Language-Action Model for Personalized Driving
by: Wang, Zehao, et al.
Published: (2026)
by: Wang, Zehao, et al.
Published: (2026)
RLLaVA: An RL-central Framework for Language and Vision Assistants
by: Zhao, Lei, et al.
Published: (2025)
by: Zhao, Lei, et al.
Published: (2025)
SeGuE: Semantic Guided Exploration for Mobile Robots
by: Simons, Cody, et al.
Published: (2025)
by: Simons, Cody, et al.
Published: (2025)
Language-guided Robust Navigation for Mobile Robots in Dynamically-changing Environments
by: Simons, Cody, et al.
Published: (2024)
by: Simons, Cody, et al.
Published: (2024)
VLP: Vision-Language Preference Learning for Embodied Manipulation
by: Liu, Runze, et al.
Published: (2025)
by: Liu, Runze, et al.
Published: (2025)
RA-PbRL: Provably Efficient Risk-Aware Preference-Based Reinforcement Learning
by: Zhao, Yujie, et al.
Published: (2024)
by: Zhao, Yujie, et al.
Published: (2024)
Conformal Prediction and MLLM aided Uncertainty Quantification in Scene Graph Generation
by: Nag, Sayak, et al.
Published: (2025)
by: Nag, Sayak, et al.
Published: (2025)
VARP: Reinforcement Learning from Vision-Language Model Feedback with Agent Regularized Preferences
by: Singh, Anukriti, et al.
Published: (2025)
by: Singh, Anukriti, et al.
Published: (2025)
Provably Robust DPO: Aligning Language Models with Noisy Feedback
by: Chowdhury, Sayak Ray, et al.
Published: (2024)
by: Chowdhury, Sayak Ray, et al.
Published: (2024)
Combinatorial Reinforcement Learning with Preference Feedback
by: Lee, Joongkyu, et al.
Published: (2025)
by: Lee, Joongkyu, et al.
Published: (2025)
Queueing Matching Bandits with Preference Feedback
by: Kim, Jung-hun, et al.
Published: (2024)
by: Kim, Jung-hun, et al.
Published: (2024)
RL with Learnable Textual Feedback: A Bilevel Approach
by: Singh, Utsav, et al.
Published: (2026)
by: Singh, Utsav, et al.
Published: (2026)
Towards Generalizable Safety in Crowd Navigation via Conformal Uncertainty Handling
by: Yao, Jianpeng, et al.
Published: (2025)
by: Yao, Jianpeng, et al.
Published: (2025)
SoNIC: Safe Social Navigation with Adaptive Conformal Inference and Constrained Reinforcement Learning
by: Yao, Jianpeng, et al.
Published: (2024)
by: Yao, Jianpeng, et al.
Published: (2024)
DOPL: Direct Online Preference Learning for Restless Bandits with Preference Feedback
by: Xiong, Guojun, et al.
Published: (2024)
by: Xiong, Guojun, et al.
Published: (2024)
Span-Agnostic Optimal Sample Complexity and Oracle Inequalities for Average-Reward RL
by: Zurek, Matthew, et al.
Published: (2025)
by: Zurek, Matthew, et al.
Published: (2025)
Approximating Safety Feedback Without a Safety Oracle via Model Predictive Control
by: Pflueger, Jeff, et al.
Published: (2025)
by: Pflueger, Jeff, et al.
Published: (2025)
Neural Dueling Bandits: Preference-Based Optimization with Human Feedback
by: Verma, Arun, et al.
Published: (2024)
by: Verma, Arun, et al.
Published: (2024)
Design Considerations in Offline Preference-based RL
by: Agarwal, Alekh, et al.
Published: (2025)
by: Agarwal, Alekh, et al.
Published: (2025)
reBandit: Random Effects based Online RL algorithm for Reducing Cannabis Use
by: Ghosh, Susobhan, et al.
Published: (2024)
by: Ghosh, Susobhan, et al.
Published: (2024)
How Well Can Preference Optimization Generalize Under Noisy Feedback?
by: Im, Shawn, et al.
Published: (2025)
by: Im, Shawn, et al.
Published: (2025)
Similar Items
-
Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning
by: Ghosh, Udita, et al.
Published: (2025) -
Robust Offline Imitation Learning from Diverse Auxiliary Data
by: Ghosh, Udita, et al.
Published: (2024) -
Toward Autonomous Laboratory Safety Monitoring with Vision Language Models: Learning to See Hazards Through Scene Structure
by: Chakraborty, Trishna, et al.
Published: (2026) -
CONTRAST: Continual Multi-source Adaptation to Dynamic Distributions
by: Ahmed, Sk Miraj, et al.
Published: (2024) -
HEAL: An Empirical Study on Hallucinations in Embodied Agents Driven by Large Language Models
by: Chakraborty, Trishna, et al.
Published: (2025)