Contextual bandits with entropy-based human feedback
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Seraj, Raihan, Meng, Lili, Sylvain, Tristan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AutoCast++: Enhancing World Event Prediction with Zero-shot Ranking-based Context Retrieval
von: Yan, Qi, et al.
Veröffentlicht: (2023)
von: Yan, Qi, et al.
Veröffentlicht: (2023)
Generalizing Multi-Step Inverse Models for Representation Learning to Finite-Memory POMDPs
von: Wu, Lili, et al.
Veröffentlicht: (2024)
von: Wu, Lili, et al.
Veröffentlicht: (2024)
Robust Reinforcement Learning Objectives for Sequential Recommender Systems
von: Mozifian, Melissa, et al.
Veröffentlicht: (2023)
von: Mozifian, Melissa, et al.
Veröffentlicht: (2023)
Heterogeneous Decentralized Diffusion Models
von: Jiang, Zhiying, et al.
Veröffentlicht: (2026)
von: Jiang, Zhiying, et al.
Veröffentlicht: (2026)
Spectral bandits
von: Kocák, Tomáš, et al.
Veröffentlicht: (2026)
von: Kocák, Tomáš, et al.
Veröffentlicht: (2026)
Linear bandits with polylogarithmic minimax regret
von: Lumbreras, Josep, et al.
Veröffentlicht: (2024)
von: Lumbreras, Josep, et al.
Veröffentlicht: (2024)
Adversarial bandit optimization for approximately linear functions
von: Cheng, Zhuoyu, et al.
Veröffentlicht: (2025)
von: Cheng, Zhuoyu, et al.
Veröffentlicht: (2025)
Reinforcement learning with combinatorial actions for coupled restless bandits
von: Xu, Lily, et al.
Veröffentlicht: (2025)
von: Xu, Lily, et al.
Veröffentlicht: (2025)
Can MLLMs generate human-like feedback in grading multimodal short answers?
von: Sil, Pritam, et al.
Veröffentlicht: (2024)
von: Sil, Pritam, et al.
Veröffentlicht: (2024)
Rejecting Hallucinated State Targets during Planning
von: Zhao, Mingde, et al.
Veröffentlicht: (2024)
von: Zhao, Mingde, et al.
Veröffentlicht: (2024)
PcLast: Discovering Plannable Continuous Latent States
von: Koul, Anurag, et al.
Veröffentlicht: (2023)
von: Koul, Anurag, et al.
Veröffentlicht: (2023)
Functional multi-armed bandit and the best function identification problems
von: Dorn, Yuriy, et al.
Veröffentlicht: (2025)
von: Dorn, Yuriy, et al.
Veröffentlicht: (2025)
Learning to summarize user information for personalized reinforcement learning from human feedback
von: Nam, Hyunji, et al.
Veröffentlicht: (2025)
von: Nam, Hyunji, et al.
Veröffentlicht: (2025)
Maximizing the efficiency of human feedback in AI alignment: a comparative analysis
von: Chouliaras, Andreas, et al.
Veröffentlicht: (2025)
von: Chouliaras, Andreas, et al.
Veröffentlicht: (2025)
Towards In-Vehicle Multi-Task Facial Attribute Recognition: Investigating Synthetic Data and Vision Foundation Models
von: Seraj, Esmaeil, et al.
Veröffentlicht: (2024)
von: Seraj, Esmaeil, et al.
Veröffentlicht: (2024)
PsyAgent: Constructing Human-like Agents Based on Psychological Modeling and Contextual Interaction
von: Meng, Zibin, et al.
Veröffentlicht: (2026)
von: Meng, Zibin, et al.
Veröffentlicht: (2026)
Softmax gradient policy for variance minimization and risk-averse multi armed bandits
von: Turinici, Gabriel
Veröffentlicht: (2026)
von: Turinici, Gabriel
Veröffentlicht: (2026)
Improving Reliable Navigation under Uncertainty via Predictions Informed by Non-Local Information
von: Arnob, Raihan Islam, et al.
Veröffentlicht: (2023)
von: Arnob, Raihan Islam, et al.
Veröffentlicht: (2023)
Silencing the Guardrails: Inference-Time Jailbreaking via Dynamic Contextual Representation Ablation
von: Xing, Wenpeng, et al.
Veröffentlicht: (2026)
von: Xing, Wenpeng, et al.
Veröffentlicht: (2026)
Eterna is Solved
von: Cazenave, Tristan
Veröffentlicht: (2025)
von: Cazenave, Tristan
Veröffentlicht: (2025)
Monte Carlo Search Algorithms Discovering Monte Carlo Tree Search Exploration Terms
von: Cazenave, Tristan
Veröffentlicht: (2024)
von: Cazenave, Tristan
Veröffentlicht: (2024)
Generalized Nested Rollout Policy Adaptation with Limited Repetitions
von: Cazenave, Tristan
Veröffentlicht: (2024)
von: Cazenave, Tristan
Veröffentlicht: (2024)
Learning a Prior for Monte Carlo Search by Replaying Solutions to Combinatorial Problems
von: Cazenave, Tristan
Veröffentlicht: (2024)
von: Cazenave, Tristan
Veröffentlicht: (2024)
False Data Injection Attack Detection in Edge-based Smart Metering Networks with Federated Learning
von: Uddin, Md Raihan, et al.
Veröffentlicht: (2024)
von: Uddin, Md Raihan, et al.
Veröffentlicht: (2024)
Prior-informed optimization of treatment recommendation via bandit algorithms trained on large language model-processed historical records
von: Nessari, Saman, et al.
Veröffentlicht: (2025)
von: Nessari, Saman, et al.
Veröffentlicht: (2025)
Use Your INSTINCT: INSTruction optimization for LLMs usIng Neural bandits Coupled with Transformers
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2023)
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2023)
AIoT-based Continuous, Contextualized, and Explainable Driving Assessment for Older Adults
von: Liu, Yimeng, et al.
Veröffentlicht: (2026)
von: Liu, Yimeng, et al.
Veröffentlicht: (2026)
MBExplainer: Multilevel bandit-based explanations for downstream models with augmented graph embeddings
von: Golgoon, Ashkan, et al.
Veröffentlicht: (2024)
von: Golgoon, Ashkan, et al.
Veröffentlicht: (2024)
Optimizing Alignment with Less: Leveraging Data Augmentation for Personalized Evaluation
von: Seraj, Javad, et al.
Veröffentlicht: (2024)
von: Seraj, Javad, et al.
Veröffentlicht: (2024)
Interpretable experiential learning based on state history and global feedback
von: Kolonin, Anton
Veröffentlicht: (2026)
von: Kolonin, Anton
Veröffentlicht: (2026)
Safety through feedback in Constrained RL
von: Chirra, Shashank Reddy, et al.
Veröffentlicht: (2024)
von: Chirra, Shashank Reddy, et al.
Veröffentlicht: (2024)
Sample-based Dynamic Hierarchical Transformer with Layer and Head Flexibility via Contextual Bandit
von: Meng, Fanfei, et al.
Veröffentlicht: (2023)
von: Meng, Fanfei, et al.
Veröffentlicht: (2023)
Realtime Dynamic Gaze Target Tracking and Depth-Level Estimation
von: Seraj, Esmaeil, et al.
Veröffentlicht: (2024)
von: Seraj, Esmaeil, et al.
Veröffentlicht: (2024)
QD-VMR: Query Debiasing with Contextual Understanding Enhancement for Video Moment Retrieval
von: Gao, Chenghua, et al.
Veröffentlicht: (2024)
von: Gao, Chenghua, et al.
Veröffentlicht: (2024)
BeeRNA: tertiary structure-based RNA inverse folding using Artificial Bee Colony
von: Mlaweh, Mehyar, et al.
Veröffentlicht: (2025)
von: Mlaweh, Mehyar, et al.
Veröffentlicht: (2025)
Contextual Confidence and Generative AI
von: Jain, Shrey, et al.
Veröffentlicht: (2023)
von: Jain, Shrey, et al.
Veröffentlicht: (2023)
Enhancing AI-based Generation of Software Exploits with Contextual Information
von: Liguori, Pietro, et al.
Veröffentlicht: (2024)
von: Liguori, Pietro, et al.
Veröffentlicht: (2024)
Emergence of Physical Intelligence via Controllable Information Production
von: Shah, Tristan, et al.
Veröffentlicht: (2026)
von: Shah, Tristan, et al.
Veröffentlicht: (2026)
Monte Carlo Permutation Search
von: Cazenave, Tristan
Veröffentlicht: (2025)
von: Cazenave, Tristan
Veröffentlicht: (2025)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AutoCast++: Enhancing World Event Prediction with Zero-shot Ranking-based Context Retrieval
von: Yan, Qi, et al.
Veröffentlicht: (2023) -
Generalizing Multi-Step Inverse Models for Representation Learning to Finite-Memory POMDPs
von: Wu, Lili, et al.
Veröffentlicht: (2024) -
Robust Reinforcement Learning Objectives for Sequential Recommender Systems
von: Mozifian, Melissa, et al.
Veröffentlicht: (2023) -
Heterogeneous Decentralized Diffusion Models
von: Jiang, Zhiying, et al.
Veröffentlicht: (2026) -
Spectral bandits
von: Kocák, Tomáš, et al.
Veröffentlicht: (2026)