POLO: Preference-Guided Multi-Turn Reinforcement Learning for Lead Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Ziqing, Wen, Yibo, Pattie, William, Luo, Xiao, Wu, Weimin, Hu, Jerry Yao-Chieh, Pandey, Abhishek, Liu, Han, Ding, Kaize |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MolMem: Memory-Augmented Agentic Reinforcement Learning for Sample-Efficient Molecular Optimization
by: Wang, Ziqing, et al.
Published: (2026)
by: Wang, Ziqing, et al.
Published: (2026)
A Survey of Large Language Models for Text-Guided Molecular Discovery: from Molecule Generation to Optimization
by: Wang, Ziqing, et al.
Published: (2025)
by: Wang, Ziqing, et al.
Published: (2025)
Pareto-Optimal Energy Alignment for Designing Nature-Like Antibodies
by: Wen, Yibo, et al.
Published: (2024)
by: Wen, Yibo, et al.
Published: (2024)
HopRank: Self-Supervised LLM Preference-Tuning on Graphs for Few-Shot Node Classification
by: Wang, Ziqing, et al.
Published: (2026)
by: Wang, Ziqing, et al.
Published: (2026)
GlassMol: Interpretable Molecular Property Prediction with Concept Bottleneck Models
by: Rivera, Oscar, et al.
Published: (2026)
by: Rivera, Oscar, et al.
Published: (2026)
On Statistical Rates and Provably Efficient Criteria of Latent Diffusion Transformers (DiTs)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
AMANDA: Agentic Medical Knowledge Augmentation for Data-Efficient Medical Visual Question Answering
by: Wang, Ziqing, et al.
Published: (2025)
by: Wang, Ziqing, et al.
Published: (2025)
Genome-Factory: A Library for Tuning, Deploying, and Interpreting Genomic Foundation Models
by: Wu, Weimin, et al.
Published: (2025)
by: Wu, Weimin, et al.
Published: (2025)
In-Context Deep Learning via Transformer Models
by: Wu, Weimin, et al.
Published: (2024)
by: Wu, Weimin, et al.
Published: (2024)
Discrete Flow Matching Policy Optimization
by: Su, Maojiang, et al.
Published: (2026)
by: Su, Maojiang, et al.
Published: (2026)
On Structured State-Space Duality
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Universal Approximation with Softmax Attention
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Beyond Sharp Minima: Robust LLM Unlearning via Feedback-Guided Multi-Point Optimization
by: Wu, Wenhan, et al.
Published: (2025)
by: Wu, Wenhan, et al.
Published: (2025)
Provably Optimal Memory Capacity for Modern Hopfield Models: Transformer-Compatible Dense Associative Memories as Spherical Codes
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
LEMON: Learning Executable Multi-Agent Orchestration via Counterfactual Reinforcement Learning
by: Chen, Xudong, et al.
Published: (2026)
by: Chen, Xudong, et al.
Published: (2026)
User Simulator-Guided Multi-Turn Preference Optimization for Reasoning LLM-based Conversational Recommendation
by: Xiang, Xingyuan, et al.
Published: (2026)
by: Xiang, Xingyuan, et al.
Published: (2026)
Direct Multi-Turn Preference Optimization for Language Agents
by: Shi, Wentao, et al.
Published: (2024)
by: Shi, Wentao, et al.
Published: (2024)
On Statistical Rates of Conditional Diffusion Transformers: Approximation, Estimation and Minimax Optimality
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
Token-Importance Guided Direct Preference Optimization
by: Yang, Ning, et al.
Published: (2025)
by: Yang, Ning, et al.
Published: (2025)
Dynamic Context Tuning for Retrieval-Augmented Generation: Enhancing Multi-Turn Planning and Tool Adaptation
by: Soni, Jubin Abhishek, et al.
Published: (2025)
by: Soni, Jubin Abhishek, et al.
Published: (2025)
Reinforcing Diffusion Models by Direct Group Preference Optimization
by: Luo, Yihong, et al.
Published: (2025)
by: Luo, Yihong, et al.
Published: (2025)
RPRO: Ranked Preference Reinforcement Optimization for Enhancing Medical QA and Diagnostic Reasoning
by: Hsu, Chia-Hsuan, et al.
Published: (2025)
by: Hsu, Chia-Hsuan, et al.
Published: (2025)
Towards Acyclic Preference Evaluation of Language Models via Multiple Evaluators
by: Hu, Zhengyu, et al.
Published: (2024)
by: Hu, Zhengyu, et al.
Published: (2024)
CoAct: Co-Active LLM Preference Learning with Human-AI Synergy
by: Xu, Ruiyao, et al.
Published: (2026)
by: Xu, Ruiyao, et al.
Published: (2026)
On Flow Matching KL Divergence
by: Su, Maojiang, et al.
Published: (2025)
by: Su, Maojiang, et al.
Published: (2025)
Attention Mechanism, Max-Affine Partition, and Universal Approximation
by: Liu, Hude, et al.
Published: (2025)
by: Liu, Hude, et al.
Published: (2025)
On Computational Limits of Modern Hopfield Models: A Fine-Grained Complexity Analysis
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2024)
Expectation Confirmation Preference Optimization for Multi-Turn Conversational Recommendation Agent
by: Feng, Xueyang, et al.
Published: (2025)
by: Feng, Xueyang, et al.
Published: (2025)
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
by: Ding, Yifeng, et al.
Published: (2025)
by: Ding, Yifeng, et al.
Published: (2025)
RobustFT: Robust Supervised Fine-tuning for Large Language Models under Noisy Response
by: Luo, Junyu, et al.
Published: (2024)
by: Luo, Junyu, et al.
Published: (2024)
Human-in-the-Loop Policy Optimization for Preference-Based Multi-Objective Reinforcement Learning
by: Li, Ke, et al.
Published: (2024)
by: Li, Ke, et al.
Published: (2024)
Uniform Memory Retrieval with Larger Capacity for Modern Hopfield Models
by: Wu, Dennis, et al.
Published: (2024)
by: Wu, Dennis, et al.
Published: (2024)
In-Context Algorithm Emulation in Fixed-Weight Transformers
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Transformer Approximations from ReLUs
by: Hu, Jerry Yao-Chieh, et al.
Published: (2026)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2026)
Mind the Inconspicuous: Revealing the Hidden Weakness in Aligned LLMs' Refusal Boundaries
by: Yu, Jiahao, et al.
Published: (2024)
by: Yu, Jiahao, et al.
Published: (2024)
Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization
by: Yu, Jiahao, et al.
Published: (2025)
by: Yu, Jiahao, et al.
Published: (2025)
GNN-as-Judge: Unleashing the Power of LLMs for Graph Learning with GNN Feedback
by: Xu, Ruiyao, et al.
Published: (2026)
by: Xu, Ruiyao, et al.
Published: (2026)
Large Language Models for Anomaly and Out-of-Distribution Detection: A Survey
by: Xu, Ruiyao, et al.
Published: (2024)
by: Xu, Ruiyao, et al.
Published: (2024)
Beyond Generalization: A Survey of Out-Of-Distribution Adaptation on Graphs
by: Liu, Shuhan, et al.
Published: (2024)
by: Liu, Shuhan, et al.
Published: (2024)
Residual Reinforcement Learning for Robot Teleoperation under Stochastic Delays
by: Deng, Kaize, et al.
Published: (2026)
by: Deng, Kaize, et al.
Published: (2026)
Similar Items
-
MolMem: Memory-Augmented Agentic Reinforcement Learning for Sample-Efficient Molecular Optimization
by: Wang, Ziqing, et al.
Published: (2026) -
A Survey of Large Language Models for Text-Guided Molecular Discovery: from Molecule Generation to Optimization
by: Wang, Ziqing, et al.
Published: (2025) -
Pareto-Optimal Energy Alignment for Designing Nature-Like Antibodies
by: Wen, Yibo, et al.
Published: (2024) -
HopRank: Self-Supervised LLM Preference-Tuning on Graphs for Few-Shot Node Classification
by: Wang, Ziqing, et al.
Published: (2026) -
GlassMol: Interpretable Molecular Property Prediction with Concept Bottleneck Models
by: Rivera, Oscar, et al.
Published: (2026)