In-context Ranking Preference Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Junda, Surana, Rohan, Xie, Zhouhang, Shen, Yiran, Xia, Yu, Yu, Tong, Rossi, Ryan A., Ammanabrolu, Prithviraj, McAuley, Julian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization
by: Surana, Rohan, et al.
Published: (2026)
by: Surana, Rohan, et al.
Published: (2026)
Evaluation on Entity Matching in Recommender Systems
by: Huang, Zihan, et al.
Published: (2026)
by: Huang, Zihan, et al.
Published: (2026)
Importance Sampling for Multi-Negative Multimodal Direct Preference Optimization
by: Li, Xintong, et al.
Published: (2025)
by: Li, Xintong, et al.
Published: (2025)
Active Learning for Direct Preference Optimization
by: Kveton, Branislav, et al.
Published: (2025)
by: Kveton, Branislav, et al.
Published: (2025)
AMPS: Adaptive Modality Preference Steering via Functional Entropy
by: Huang, Zihan, et al.
Published: (2026)
by: Huang, Zihan, et al.
Published: (2026)
A Survey on Personalized and Pluralistic Preference Alignment in Large Language Models
by: Xie, Zhouhang, et al.
Published: (2025)
by: Xie, Zhouhang, et al.
Published: (2025)
Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards
by: Shen, Yiran, et al.
Published: (2025)
by: Shen, Yiran, et al.
Published: (2025)
F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking
by: Surana, Rohan, et al.
Published: (2026)
by: Surana, Rohan, et al.
Published: (2026)
Preference-Based Learning in Audio Applications: A Systematic Analysis
by: Broukhim, Aaron, et al.
Published: (2025)
by: Broukhim, Aaron, et al.
Published: (2025)
Skill-CMIB: Multimodal Agent Skill for Consistent Action via Conditional Multimodal Information Bottleneck
by: Huang, Zihan, et al.
Published: (2026)
by: Huang, Zihan, et al.
Published: (2026)
WS-GRPO: Weakly-Supervised Group-Relative Policy Optimization for Rollout-Efficient Reasoning
by: Mundada, Gagan, et al.
Published: (2026)
by: Mundada, Gagan, et al.
Published: (2026)
Skill-R1: Agent Skill Evolution via Reinforcement Learning
by: Vishe, Yash, et al.
Published: (2026)
by: Vishe, Yash, et al.
Published: (2026)
From Reviews to Dialogues: Active Synthesis for Zero-Shot LLM-based Conversational Recommender System
by: Surana, Rohan, et al.
Published: (2025)
by: Surana, Rohan, et al.
Published: (2025)
Beyond Needle(s) in the Embodied Haystack: Environment, Architecture, and Training Considerations for Long Context Reasoning
by: Kim, Bosung, et al.
Published: (2025)
by: Kim, Bosung, et al.
Published: (2025)
A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning
by: Wang, Ruiyi, et al.
Published: (2025)
by: Wang, Ruiyi, et al.
Published: (2025)
Futga: Towards Fine-grained Music Understanding through Temporally-enhanced Generative Augmentation
by: Wu, Junda, et al.
Published: (2024)
by: Wu, Junda, et al.
Published: (2024)
Federated Large Language Models: Current Progress and Future Directions
by: Yao, Yuhang, et al.
Published: (2024)
by: Yao, Yuhang, et al.
Published: (2024)
Mitigating Visual Knowledge Forgetting in MLLM Instruction-tuning via Modality-decoupled Gradient Descent
by: Wu, Junda, et al.
Published: (2025)
by: Wu, Junda, et al.
Published: (2025)
How Reasoning Evolves from Post-Training Data: An Empirical Study Using Chess
by: Dionisopoulos, Lucas, et al.
Published: (2026)
by: Dionisopoulos, Lucas, et al.
Published: (2026)
SAND: Boosting LLM Agents with Self-Taught Action Deliberation
by: Xia, Yu, et al.
Published: (2025)
by: Xia, Yu, et al.
Published: (2025)
CoMMIT: Coordinated Multimodal Instruction Tuning
by: Li, Xintong, et al.
Published: (2024)
by: Li, Xintong, et al.
Published: (2024)
FINEST: Stabilizing Recommendations by Rank-Preserving Fine-Tuning
by: Oh, Sejoon, et al.
Published: (2024)
by: Oh, Sejoon, et al.
Published: (2024)
Listwise Preference Diffusion Optimization for User Behavior Trajectories Prediction
by: Huang, Hongtao, et al.
Published: (2025)
by: Huang, Hongtao, et al.
Published: (2025)
CTRLS: Chain-of-Thought Reasoning via Latent State-Transition
by: Wu, Junda, et al.
Published: (2025)
by: Wu, Junda, et al.
Published: (2025)
FERA: Uncertainty-Aware Federated Reasoning for Large Language Models
by: Wang, Ruhan, et al.
Published: (2026)
by: Wang, Ruhan, et al.
Published: (2026)
Learning to Hint for Reinforcement Learning
by: Xia, Yu, et al.
Published: (2026)
by: Xia, Yu, et al.
Published: (2026)
Extending Input Contexts of Language Models through Training on Segmented Sequences
by: Karypis, Petros, et al.
Published: (2023)
by: Karypis, Petros, et al.
Published: (2023)
OCEAN: Offline Chain-of-thought Evaluation and Alignment in Large Language Models
by: Wu, Junda, et al.
Published: (2024)
by: Wu, Junda, et al.
Published: (2024)
Pluralistic Off-policy Evaluation and Alignment
by: Huang, Chengkai, et al.
Published: (2025)
by: Huang, Chengkai, et al.
Published: (2025)
DICE: Dynamic In-Context Example Selection in LLM Agents via Efficient Knowledge Transfer
by: Wang, Ruoyu, et al.
Published: (2025)
by: Wang, Ruoyu, et al.
Published: (2025)
Knowledge-Aware Query Expansion with Large Language Models for Textual and Relational Retrieval
by: Xia, Yu, et al.
Published: (2024)
by: Xia, Yu, et al.
Published: (2024)
DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation
by: Novack, Zachary, et al.
Published: (2024)
by: Novack, Zachary, et al.
Published: (2024)
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning
by: Surana, Rohan, et al.
Published: (2026)
by: Surana, Rohan, et al.
Published: (2026)
Visual Prompting in Multimodal Large Language Models: A Survey
by: Wu, Junda, et al.
Published: (2024)
by: Wu, Junda, et al.
Published: (2024)
ThinkRouter: Efficient Reasoning via Routing Thinking between Latent and Discrete Spaces
by: Xu, Xin, et al.
Published: (2026)
by: Xu, Xin, et al.
Published: (2026)
Bridging Conversational and Collaborative Signals for Conversational Recommendation
by: Rabiah, Ahmad Bin, et al.
Published: (2024)
by: Rabiah, Ahmad Bin, et al.
Published: (2024)
MusiCRS: Benchmarking Audio-Centric Conversational Recommendation
by: Surana, Rohan, et al.
Published: (2025)
by: Surana, Rohan, et al.
Published: (2025)
CSyMR: Benchmarking Compositional Music Information Retrieval in Symbolic Music Reasoning
by: Wang, Boyang, et al.
Published: (2025)
by: Wang, Boyang, et al.
Published: (2025)
Critique-out-Loud Reward Models
by: Ankner, Zachary, et al.
Published: (2024)
by: Ankner, Zachary, et al.
Published: (2024)
Multi-Agent Collaborative Filtering: Orchestrating Users and Items for Agentic Recommendations
by: Xia, Yu, et al.
Published: (2025)
by: Xia, Yu, et al.
Published: (2025)
Similar Items
-
MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization
by: Surana, Rohan, et al.
Published: (2026) -
Evaluation on Entity Matching in Recommender Systems
by: Huang, Zihan, et al.
Published: (2026) -
Importance Sampling for Multi-Negative Multimodal Direct Preference Optimization
by: Li, Xintong, et al.
Published: (2025) -
Active Learning for Direct Preference Optimization
by: Kveton, Branislav, et al.
Published: (2025) -
AMPS: Adaptive Modality Preference Steering via Functional Entropy
by: Huang, Zihan, et al.
Published: (2026)