F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Surana, Rohan, Mundada, Gagan, Wu, Junda, Li, Xintong, Jiao, Yizhu, Jin, Bowen, Zhou, Sizhe, Yu, Tong, Sinha, Ritwik, Han, Jiawei, Shang, Jingbo, McAuley, Julian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WS-GRPO: Weakly-Supervised Group-Relative Policy Optimization for Rollout-Efficient Reasoning
von: Mundada, Gagan, et al.
Veröffentlicht: (2026)
von: Mundada, Gagan, et al.
Veröffentlicht: (2026)
Importance Sampling for Multi-Negative Multimodal Direct Preference Optimization
von: Li, Xintong, et al.
Veröffentlicht: (2025)
von: Li, Xintong, et al.
Veröffentlicht: (2025)
MusiCRS: Benchmarking Audio-Centric Conversational Recommendation
von: Surana, Rohan, et al.
Veröffentlicht: (2025)
von: Surana, Rohan, et al.
Veröffentlicht: (2025)
AMPS: Adaptive Modality Preference Steering via Functional Entropy
von: Huang, Zihan, et al.
Veröffentlicht: (2026)
von: Huang, Zihan, et al.
Veröffentlicht: (2026)
Explainable Chain-of-Thought Reasoning: An Empirical Analysis on State-Aware Reasoning Dynamics
von: Yu, Sheldon, et al.
Veröffentlicht: (2025)
von: Yu, Sheldon, et al.
Veröffentlicht: (2025)
OLIVIA: Online Learning via Inference-time Action Adaptation for Decision Making in LLM ReAct Agents
von: Yu, Sheldon, et al.
Veröffentlicht: (2026)
von: Yu, Sheldon, et al.
Veröffentlicht: (2026)
MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization
von: Surana, Rohan, et al.
Veröffentlicht: (2026)
von: Surana, Rohan, et al.
Veröffentlicht: (2026)
Active Learning for Direct Preference Optimization
von: Kveton, Branislav, et al.
Veröffentlicht: (2025)
von: Kveton, Branislav, et al.
Veröffentlicht: (2025)
WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music Reasoning
von: Mundada, Gagan, et al.
Veröffentlicht: (2025)
von: Mundada, Gagan, et al.
Veröffentlicht: (2025)
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning
von: Surana, Rohan, et al.
Veröffentlicht: (2026)
von: Surana, Rohan, et al.
Veröffentlicht: (2026)
In-context Ranking Preference Optimization
von: Wu, Junda, et al.
Veröffentlicht: (2025)
von: Wu, Junda, et al.
Veröffentlicht: (2025)
Evaluation on Entity Matching in Recommender Systems
von: Huang, Zihan, et al.
Veröffentlicht: (2026)
von: Huang, Zihan, et al.
Veröffentlicht: (2026)
Skill-R1: Agent Skill Evolution via Reinforcement Learning
von: Vishe, Yash, et al.
Veröffentlicht: (2026)
von: Vishe, Yash, et al.
Veröffentlicht: (2026)
CoMMIT: Coordinated Multimodal Instruction Tuning
von: Li, Xintong, et al.
Veröffentlicht: (2024)
von: Li, Xintong, et al.
Veröffentlicht: (2024)
SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes
von: Wang, Chuhan, et al.
Veröffentlicht: (2026)
von: Wang, Chuhan, et al.
Veröffentlicht: (2026)
From Reviews to Dialogues: Active Synthesis for Zero-Shot LLM-based Conversational Recommender System
von: Surana, Rohan, et al.
Veröffentlicht: (2025)
von: Surana, Rohan, et al.
Veröffentlicht: (2025)
CTRLS: Chain-of-Thought Reasoning via Latent State-Transition
von: Wu, Junda, et al.
Veröffentlicht: (2025)
von: Wu, Junda, et al.
Veröffentlicht: (2025)
Skill-CMIB: Multimodal Agent Skill for Consistent Action via Conditional Multimodal Information Bottleneck
von: Huang, Zihan, et al.
Veröffentlicht: (2026)
von: Huang, Zihan, et al.
Veröffentlicht: (2026)
OCEAN: Offline Chain-of-thought Evaluation and Alignment in Large Language Models
von: Wu, Junda, et al.
Veröffentlicht: (2024)
von: Wu, Junda, et al.
Veröffentlicht: (2024)
Mitigating Visual Knowledge Forgetting in MLLM Instruction-tuning via Modality-decoupled Gradient Descent
von: Wu, Junda, et al.
Veröffentlicht: (2025)
von: Wu, Junda, et al.
Veröffentlicht: (2025)
Listwise Preference Diffusion Optimization for User Behavior Trajectories Prediction
von: Huang, Hongtao, et al.
Veröffentlicht: (2025)
von: Huang, Hongtao, et al.
Veröffentlicht: (2025)
Weakly-supervised VLM-guided Partial Contrastive Learning for Visual Language Navigation
von: Wang, Ruoyu, et al.
Veröffentlicht: (2025)
von: Wang, Ruoyu, et al.
Veröffentlicht: (2025)
FINEST: Stabilizing Recommendations by Rank-Preserving Fine-Tuning
von: Oh, Sejoon, et al.
Veröffentlicht: (2024)
von: Oh, Sejoon, et al.
Veröffentlicht: (2024)
Reddit2Deezer: A Scalable Dataset for Real-World Grounded Conversational Music Recommendation
von: Kim, Haven, et al.
Veröffentlicht: (2026)
von: Kim, Haven, et al.
Veröffentlicht: (2026)
TEXT2DB: Integration-Aware Information Extraction with Large Language Model Agents
von: Jiao, Yizhu, et al.
Veröffentlicht: (2025)
von: Jiao, Yizhu, et al.
Veröffentlicht: (2025)
A Survey on Personalized and Pluralistic Preference Alignment in Large Language Models
von: Xie, Zhouhang, et al.
Veröffentlicht: (2025)
von: Xie, Zhouhang, et al.
Veröffentlicht: (2025)
Three central limit theorems for the unbounded excursion component of a Gaussian field
von: McAuley, Michael
Veröffentlicht: (2024)
von: McAuley, Michael
Veröffentlicht: (2024)
Children in Custody
von: McAuley, Mary
Veröffentlicht: (2022)
von: McAuley, Mary
Veröffentlicht: (2022)
Politics and the Soviet Union / Mary McAuley
von: McAuley, Mary
von: McAuley, Mary
Limit theorems for non-local functionals of smooth Gaussian fields via quasi-association
von: McAuley, Michael
Veröffentlicht: (2026)
von: McAuley, Michael
Veröffentlicht: (2026)
GSPRec: Temporal-Aware Graph Spectral Filtering for Recommendation
von: Rabiah, Ahmad Bin, et al.
Veröffentlicht: (2025)
von: Rabiah, Ahmad Bin, et al.
Veröffentlicht: (2025)
InstructGraph: Boosting Large Language Models via Graph-centric Instruction Tuning and Preference Alignment
von: Wang, Jianing, et al.
Veröffentlicht: (2024)
von: Wang, Jianing, et al.
Veröffentlicht: (2024)
CoLLAP: Contrastive Long-form Language-Audio Pretraining with Musical Temporal Structure Augmentation
von: Wu, Junda, et al.
Veröffentlicht: (2024)
von: Wu, Junda, et al.
Veröffentlicht: (2024)
PDB-Eval: An Evaluation of Large Multimodal Models for Description and Explanation of Personalized Driving Behavior
von: Wu, Junda, et al.
Veröffentlicht: (2025)
von: Wu, Junda, et al.
Veröffentlicht: (2025)
CoRAL: Collaborative Retrieval-Augmented Large Language Models Improve Long-tail Recommendation
von: Wu, Junda, et al.
Veröffentlicht: (2024)
von: Wu, Junda, et al.
Veröffentlicht: (2024)
DICE: Dynamic In-Context Example Selection in LLM Agents via Efficient Knowledge Transfer
von: Wang, Ruoyu, et al.
Veröffentlicht: (2025)
von: Wang, Ruoyu, et al.
Veröffentlicht: (2025)
Knowledge-Aware Query Expansion with Large Language Models for Textual and Relational Retrieval
von: Xia, Yu, et al.
Veröffentlicht: (2024)
von: Xia, Yu, et al.
Veröffentlicht: (2024)
Establishing Knowledge Preference in Language Models
von: Zhou, Sizhe, et al.
Veröffentlicht: (2024)
von: Zhou, Sizhe, et al.
Veröffentlicht: (2024)
Self-Updatable Large Language Models by Integrating Context into Model Parameters
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
Extending Input Contexts of Language Models through Training on Segmented Sequences
von: Karypis, Petros, et al.
Veröffentlicht: (2023)
von: Karypis, Petros, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
WS-GRPO: Weakly-Supervised Group-Relative Policy Optimization for Rollout-Efficient Reasoning
von: Mundada, Gagan, et al.
Veröffentlicht: (2026) -
Importance Sampling for Multi-Negative Multimodal Direct Preference Optimization
von: Li, Xintong, et al.
Veröffentlicht: (2025) -
MusiCRS: Benchmarking Audio-Centric Conversational Recommendation
von: Surana, Rohan, et al.
Veröffentlicht: (2025) -
AMPS: Adaptive Modality Preference Steering via Functional Entropy
von: Huang, Zihan, et al.
Veröffentlicht: (2026) -
Explainable Chain-of-Thought Reasoning: An Empirical Analysis on State-Aware Reasoning Dynamics
von: Yu, Sheldon, et al.
Veröffentlicht: (2025)