Aligning Medical Conversational AI through Online Reinforcement Learning with Information-Theoretic Rewards
Fuente:
arXiv
Salvato in:
| Autori principali: | Verma, Tanvi, Zhou, Yang, Goh, Rick Siow Mong, Liu, Yong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AiRacleX: Automated Detection of Price Oracle Manipulations via LLM-Driven Knowledge Mining and Prompt Generation
di: Gao, Bo, et al.
Pubblicazione: (2025)
di: Gao, Bo, et al.
Pubblicazione: (2025)
Enabling Energy-Efficient Deployment of Large Language Models on Memristor Crossbar: A Synergy of Large and Small
di: Wang, Zhehui, et al.
Pubblicazione: (2024)
di: Wang, Zhehui, et al.
Pubblicazione: (2024)
Secure and Explainable Fraud Detection in Finance via Hierarchical Multi-source Dataset Distillation
di: Qian, Yiming, et al.
Pubblicazione: (2025)
di: Qian, Yiming, et al.
Pubblicazione: (2025)
UrFound: Towards Universal Retinal Foundation Models via Knowledge-Guided Masked Modeling
di: Yu, Kai, et al.
Pubblicazione: (2024)
di: Yu, Kai, et al.
Pubblicazione: (2024)
AdvMIM: Adversarial Masked Image Modeling for Semi-Supervised Medical Image Segmentation
di: Zhu, Lei, et al.
Pubblicazione: (2025)
di: Zhu, Lei, et al.
Pubblicazione: (2025)
MaskedCLIP: Bridging the Masked and CLIP Space for Semi-Supervised Medical Vision-Language Pre-training
di: Zhu, Lei, et al.
Pubblicazione: (2025)
di: Zhu, Lei, et al.
Pubblicazione: (2025)
History-Aware and Dynamic Client Contribution in Federated Learning
di: Ghosh, Bishwamittra, et al.
Pubblicazione: (2024)
di: Ghosh, Bishwamittra, et al.
Pubblicazione: (2024)
Optimizing Neural Networks with Learnable Non-Linear Activation Functions via Lookup-Based FPGA Acceleration
di: Yin, Mengyuan, et al.
Pubblicazione: (2025)
di: Yin, Mengyuan, et al.
Pubblicazione: (2025)
Aligning Crowd Feedback via Distributional Preference Reward Modeling
di: Li, Dexun, et al.
Pubblicazione: (2024)
di: Li, Dexun, et al.
Pubblicazione: (2024)
Enhancing Community Vision Screening -- AI Driven Retinal Photography for Early Disease Detection and Patient Trust
di: Lei, Xiaofeng, et al.
Pubblicazione: (2024)
di: Lei, Xiaofeng, et al.
Pubblicazione: (2024)
Is Quantum Optimization Ready? An Effort Towards Neural Network Compression using Adiabatic Quantum Computing
di: Wang, Zhehui, et al.
Pubblicazione: (2025)
di: Wang, Zhehui, et al.
Pubblicazione: (2025)
RLPeri: Accelerating Visual Perimetry Test with Reinforcement Learning and Convolutional Feature Extraction
di: Verma, Tanvi, et al.
Pubblicazione: (2024)
di: Verma, Tanvi, et al.
Pubblicazione: (2024)
Direct Advantage Regression: Aligning LLMs with Online AI Reward
di: He, Li, et al.
Pubblicazione: (2025)
di: He, Li, et al.
Pubblicazione: (2025)
Safe Reinforcement Learning with Learned Non-Markovian Safety Constraints
di: Low, Siow Meng, et al.
Pubblicazione: (2024)
di: Low, Siow Meng, et al.
Pubblicazione: (2024)
Partially Supervised Unpaired Multi-Modal Learning for Label-Efficient Medical Image Segmentation
di: Zhu, Lei, et al.
Pubblicazione: (2025)
di: Zhu, Lei, et al.
Pubblicazione: (2025)
Rewarding Creativity: A Human-Aligned Generative Reward Model for Reinforcement Learning in Storytelling
di: Li, Zhaoyan, et al.
Pubblicazione: (2026)
di: Li, Zhaoyan, et al.
Pubblicazione: (2026)
Towards Open-Ended Emotional Support Conversations in LLMs via Reinforcement Learning with Future-Oriented Rewards
di: Yang, Ting, et al.
Pubblicazione: (2025)
di: Yang, Ting, et al.
Pubblicazione: (2025)
Enhancing Inverse Reinforcement Learning through Encoding Dynamic Information in Reward Shaping
di: Zhan, Simon Sinong, et al.
Pubblicazione: (2024)
di: Zhan, Simon Sinong, et al.
Pubblicazione: (2024)
Secure Code Generation via Online Reinforcement Learning with Vulnerability Reward Model
di: Wu, Tianyi, et al.
Pubblicazione: (2026)
di: Wu, Tianyi, et al.
Pubblicazione: (2026)
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations
di: Hoang, Huy, et al.
Pubblicazione: (2025)
di: Hoang, Huy, et al.
Pubblicazione: (2025)
Incorporating Human Flexibility through Reward Preferences in Human-AI Teaming
di: Bhambri, Siddhant, et al.
Pubblicazione: (2023)
di: Bhambri, Siddhant, et al.
Pubblicazione: (2023)
Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable Rewards
di: Nguyen, Hieu Trung, et al.
Pubblicazione: (2026)
di: Nguyen, Hieu Trung, et al.
Pubblicazione: (2026)
Optimas: Optimizing Compound AI Systems with Globally Aligned Local Rewards
di: Wu, Shirley, et al.
Pubblicazione: (2025)
di: Wu, Shirley, et al.
Pubblicazione: (2025)
Reward Hacking in Rubric-Based Reinforcement Learning
di: Mahmoud, Anas, et al.
Pubblicazione: (2026)
di: Mahmoud, Anas, et al.
Pubblicazione: (2026)
HAVA: Hybrid Approach to Value-Alignment through Reward Weighing for Reinforcement Learning
di: Varys, Kryspin, et al.
Pubblicazione: (2025)
di: Varys, Kryspin, et al.
Pubblicazione: (2025)
MRG-R1: Reinforcement Learning for Clinically Aligned Medical Report Generation
di: Wang, Pengyu, et al.
Pubblicazione: (2025)
di: Wang, Pengyu, et al.
Pubblicazione: (2025)
CoT-Space: A Theoretical Framework for Internal Slow-Thinking via Reinforcement Learning
di: Gan, Zeyu, et al.
Pubblicazione: (2025)
di: Gan, Zeyu, et al.
Pubblicazione: (2025)
Process Reinforcement through Implicit Rewards
di: Cui, Ganqu, et al.
Pubblicazione: (2025)
di: Cui, Ganqu, et al.
Pubblicazione: (2025)
CAPSULE: Control-Theoretic Action Perturbations for Safe Uncertainty-Aware Reinforcement Learning
di: Narava, Rahul, et al.
Pubblicazione: (2026)
di: Narava, Rahul, et al.
Pubblicazione: (2026)
SALP-CG: Standard-Aligned LLM Pipeline for Classifying and Grading Large Volumes of Online Conversational Health Data
di: Yan, Yiwei, et al.
Pubblicazione: (2025)
di: Yan, Yiwei, et al.
Pubblicazione: (2025)
VORTEX: Aligning Task Utility and Human Preferences through LLM-Guided Reward Shaping
di: Xiong, Guojun, et al.
Pubblicazione: (2025)
di: Xiong, Guojun, et al.
Pubblicazione: (2025)
EDGE: A Theoretical Framework for Misconception-Aware Adaptive Learning
di: Verma, Ananda Prakash
Pubblicazione: (2025)
di: Verma, Ananda Prakash
Pubblicazione: (2025)
BenchX: A Unified Benchmark Framework for Medical Vision-Language Pretraining on Chest X-Rays
di: Zhou, Yang, et al.
Pubblicazione: (2024)
di: Zhou, Yang, et al.
Pubblicazione: (2024)
An Aggregation-Free Federated Learning for Tackling Data Heterogeneity
di: Wang, Yuan, et al.
Pubblicazione: (2024)
di: Wang, Yuan, et al.
Pubblicazione: (2024)
A Practical Approach to using Supervised Machine Learning Models to Classify Aviation Safety Occurrences
di: Siow, Bryan Y.
Pubblicazione: (2025)
di: Siow, Bryan Y.
Pubblicazione: (2025)
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective
di: He, Shenghua, et al.
Pubblicazione: (2025)
di: He, Shenghua, et al.
Pubblicazione: (2025)
DialogXpert: Driving Intelligent and Emotion-Aware Conversations through Online Value-Based Reinforcement Learning with LLM Priors
di: Rakib, Tazeek Bin Abdur, et al.
Pubblicazione: (2025)
di: Rakib, Tazeek Bin Abdur, et al.
Pubblicazione: (2025)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
di: Miao, Yuchun, et al.
Pubblicazione: (2024)
di: Miao, Yuchun, et al.
Pubblicazione: (2024)
AI-Slop to AI-Polish? Aligning Language Models through Edit-Based Writing Rewards and Test-time Computation
di: Chakrabarty, Tuhin, et al.
Pubblicazione: (2025)
di: Chakrabarty, Tuhin, et al.
Pubblicazione: (2025)
Multi-Agent Collaborative Reward Design for Enhancing Reasoning in Reinforcement Learning
di: Yang, Pei, et al.
Pubblicazione: (2025)
di: Yang, Pei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
AiRacleX: Automated Detection of Price Oracle Manipulations via LLM-Driven Knowledge Mining and Prompt Generation
di: Gao, Bo, et al.
Pubblicazione: (2025) -
Enabling Energy-Efficient Deployment of Large Language Models on Memristor Crossbar: A Synergy of Large and Small
di: Wang, Zhehui, et al.
Pubblicazione: (2024) -
Secure and Explainable Fraud Detection in Finance via Hierarchical Multi-source Dataset Distillation
di: Qian, Yiming, et al.
Pubblicazione: (2025) -
UrFound: Towards Universal Retinal Foundation Models via Knowledge-Guided Masked Modeling
di: Yu, Kai, et al.
Pubblicazione: (2024) -
AdvMIM: Adversarial Masked Image Modeling for Semi-Supervised Medical Image Segmentation
di: Zhu, Lei, et al.
Pubblicazione: (2025)