Making RL with Preference-based Feedback Efficient via Randomization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Runzhe, Sun, Wen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback
von: Zhang, Rongtao, et al.
Veröffentlicht: (2026)
von: Zhang, Rongtao, et al.
Veröffentlicht: (2026)
Introducing User Feedback-based Counterfactual Explanations (UFCE)
von: Suffian, Muhammad, et al.
Veröffentlicht: (2024)
von: Suffian, Muhammad, et al.
Veröffentlicht: (2024)
Understanding Impact of Human Feedback via Influence Functions
von: Min, Taywon, et al.
Veröffentlicht: (2025)
von: Min, Taywon, et al.
Veröffentlicht: (2025)
VARP: Reinforcement Learning from Vision-Language Model Feedback with Agent Regularized Preferences
von: Singh, Anukriti, et al.
Veröffentlicht: (2025)
von: Singh, Anukriti, et al.
Veröffentlicht: (2025)
Enhancing Preference-based Linear Bandits via Human Response Time
von: Li, Shen, et al.
Veröffentlicht: (2024)
von: Li, Shen, et al.
Veröffentlicht: (2024)
Tell me why: Training preferences-based RL with human preferences and step-level explanations
von: Karalus, Jakob
Veröffentlicht: (2024)
von: Karalus, Jakob
Veröffentlicht: (2024)
TimeSRL: Generalizable Time-Series Behavioral Modeling via Semantic RL-Tuned LLMs -- A Case Study in Mental Health
von: Fan, Yuang, et al.
Veröffentlicht: (2026)
von: Fan, Yuang, et al.
Veröffentlicht: (2026)
Influencing Humans to Conform to Preference Models for RLHF
von: Hatgis-Kessell, Stephane, et al.
Veröffentlicht: (2025)
von: Hatgis-Kessell, Stephane, et al.
Veröffentlicht: (2025)
Towards Interactive Reinforcement Learning with Intrinsic Feedback
von: Poole, Benjamin, et al.
Veröffentlicht: (2021)
von: Poole, Benjamin, et al.
Veröffentlicht: (2021)
Towards Green Wearable Computing: A Physics-Aware Spiking Neural Network for Energy-Efficient IMU-based Human Activity Recognition
von: Zheng, Naichuan, et al.
Veröffentlicht: (2026)
von: Zheng, Naichuan, et al.
Veröffentlicht: (2026)
Interactive Example-based Explanations to Improve Health Professionals' Onboarding with AI for Human-AI Collaborative Decision Making
von: Lee, Min Hun, et al.
Veröffentlicht: (2024)
von: Lee, Min Hun, et al.
Veröffentlicht: (2024)
Hindsight PRIORs for Reward Learning from Human Preferences
von: Verma, Mudit, et al.
Veröffentlicht: (2024)
von: Verma, Mudit, et al.
Veröffentlicht: (2024)
Learning a Canonical Basis of Human Preferences from Binary Ratings
von: Vodrahalli, Kailas, et al.
Veröffentlicht: (2025)
von: Vodrahalli, Kailas, et al.
Veröffentlicht: (2025)
Sentiment Analysis in Learning Management Systems Understanding Student Feedback at Scale
von: Almutairi, Mohammed
Veröffentlicht: (2025)
von: Almutairi, Mohammed
Veröffentlicht: (2025)
Optimizing Data Delivery: Insights from User Preferences on Visuals, Tables, and Text
von: Luera, Reuben, et al.
Veröffentlicht: (2024)
von: Luera, Reuben, et al.
Veröffentlicht: (2024)
MENTOR: Guiding Hierarchical Reinforcement Learning with Human Feedback and Dynamic Distance Constraint
von: Zhou, Xinglin, et al.
Veröffentlicht: (2024)
von: Zhou, Xinglin, et al.
Veröffentlicht: (2024)
Cognitive Exoskeleton: Augmenting Human Cognition with an AI-Mediated Intelligent Visual Feedback
von: Xu, Songlin, et al.
Veröffentlicht: (2025)
von: Xu, Songlin, et al.
Veröffentlicht: (2025)
Feedback-Aware Monte Carlo Tree Search for Efficient Information Seeking in Goal-Oriented Conversations
von: Chopra, Harshita, et al.
Veröffentlicht: (2025)
von: Chopra, Harshita, et al.
Veröffentlicht: (2025)
Making Language Models Better Tool Learners with Execution Feedback
von: Qiao, Shuofei, et al.
Veröffentlicht: (2023)
von: Qiao, Shuofei, et al.
Veröffentlicht: (2023)
From Accuracy to Readiness: Metrics and Benchmarks for Human-AI Decision-Making
von: Lee, Min Hun
Veröffentlicht: (2026)
von: Lee, Min Hun
Veröffentlicht: (2026)
An Approach to Joint Hybrid Decision Making between Humans and Artificial Intelligence
von: Rockbach, Jonas D., et al.
Veröffentlicht: (2025)
von: Rockbach, Jonas D., et al.
Veröffentlicht: (2025)
AI, Meet Human: Learning Paradigms for Hybrid Decision Making Systems
von: Punzi, Clara, et al.
Veröffentlicht: (2024)
von: Punzi, Clara, et al.
Veröffentlicht: (2024)
Towards Uncertainty Aware Task Delegation and Human-AI Collaborative Decision-Making
von: Lee, Min Hun, et al.
Veröffentlicht: (2025)
von: Lee, Min Hun, et al.
Veröffentlicht: (2025)
Efficient Speech Command Recognition Leveraging Spiking Neural Network and Curriculum Learning-based Knowledge Distillation
von: Wang, Jiaqi, et al.
Veröffentlicht: (2024)
von: Wang, Jiaqi, et al.
Veröffentlicht: (2024)
Improving Health Professionals' Onboarding with AI and XAI for Trustworthy Human-AI Collaborative Decision Making
von: Lee, Min Hun, et al.
Veröffentlicht: (2024)
von: Lee, Min Hun, et al.
Veröffentlicht: (2024)
Protecting Multiple Types of Privacy Simultaneously in EEG-based Brain-Computer Interfaces
von: Meng, Lubin, et al.
Veröffentlicht: (2024)
von: Meng, Lubin, et al.
Veröffentlicht: (2024)
Spatial Distillation based Distribution Alignment (SDDA) for Cross-Headset EEG Classification
von: Liu, Dingkun, et al.
Veröffentlicht: (2025)
von: Liu, Dingkun, et al.
Veröffentlicht: (2025)
Reciprocal Learning of Intent Inferral with Augmented Visual Feedback for Stroke
von: Xu, Jingxi, et al.
Veröffentlicht: (2024)
von: Xu, Jingxi, et al.
Veröffentlicht: (2024)
RAICL: Retrieval-Augmented In-Context Learning for Vision-Language-Model Based EEG Seizure Detection
von: Li, Siyang, et al.
Veröffentlicht: (2026)
von: Li, Siyang, et al.
Veröffentlicht: (2026)
HERO: Human-Feedback Efficient Reinforcement Learning for Online Diffusion Model Finetuning
von: Hiranaka, Ayano, et al.
Veröffentlicht: (2024)
von: Hiranaka, Ayano, et al.
Veröffentlicht: (2024)
Conformalized Interactive Imitation Learning: Handling Expert Shift and Intermittent Feedback
von: Zhao, Michelle, et al.
Veröffentlicht: (2024)
von: Zhao, Michelle, et al.
Veröffentlicht: (2024)
Quantifying the Effect of Feedback Frequency in Interactive Reinforcement Learning for Robotic Tasks
von: Harnack, Daniel, et al.
Veröffentlicht: (2022)
von: Harnack, Daniel, et al.
Veröffentlicht: (2022)
Contrastive Learning from Exploratory Actions: Leveraging Natural Interactions for Preference Elicitation
von: Dennler, Nathaniel, et al.
Veröffentlicht: (2025)
von: Dennler, Nathaniel, et al.
Veröffentlicht: (2025)
Improving User Experience in Preference-Based Optimization of Reward Functions for Assistive Robots
von: Dennler, Nathaniel, et al.
Veröffentlicht: (2024)
von: Dennler, Nathaniel, et al.
Veröffentlicht: (2024)
SpiroActive: Active Learning for Efficient Data Acquisition for Spirometry
von: Jain, Ankita Kumari, et al.
Veröffentlicht: (2024)
von: Jain, Ankita Kumari, et al.
Veröffentlicht: (2024)
Talaria: Interactively Optimizing Machine Learning Models for Efficient Inference
von: Hohman, Fred, et al.
Veröffentlicht: (2024)
von: Hohman, Fred, et al.
Veröffentlicht: (2024)
Uni-RLHF: Universal Platform and Benchmark Suite for Reinforcement Learning with Diverse Human Feedback
von: Yuan, Yifu, et al.
Veröffentlicht: (2024)
von: Yuan, Yifu, et al.
Veröffentlicht: (2024)
Investigating Thematic Patterns and User Preferences in LLM Interactions using BERTopic
von: Bhandarkar, Abhay, et al.
Veröffentlicht: (2025)
von: Bhandarkar, Abhay, et al.
Veröffentlicht: (2025)
FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users
von: Singh, Anikait, et al.
Veröffentlicht: (2025)
von: Singh, Anikait, et al.
Veröffentlicht: (2025)
Human-like Bots for Tactical Shooters Using Compute-Efficient Sensors
von: Justesen, Niels, et al.
Veröffentlicht: (2024)
von: Justesen, Niels, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Vibrotactile Preference Learning: Uncertainty-Aware Preference Learning for Personalized Vibration Feedback
von: Zhang, Rongtao, et al.
Veröffentlicht: (2026) -
Introducing User Feedback-based Counterfactual Explanations (UFCE)
von: Suffian, Muhammad, et al.
Veröffentlicht: (2024) -
Understanding Impact of Human Feedback via Influence Functions
von: Min, Taywon, et al.
Veröffentlicht: (2025) -
VARP: Reinforcement Learning from Vision-Language Model Feedback with Agent Regularized Preferences
von: Singh, Anukriti, et al.
Veröffentlicht: (2025) -
Enhancing Preference-based Linear Bandits via Human Response Time
von: Li, Shen, et al.
Veröffentlicht: (2024)