Policy-Gradient Training of Language Models for Ranking
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Ge, Chang, Jonathan D., Cardie, Claire, Brantley, Kianté, Joachim, Thorsten |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WildVis: Open Source Visualizer for Million-Scale Chat Logs in the Wild
by: Deng, Yuntian, et al.
Published: (2024)
by: Deng, Yuntian, et al.
Published: (2024)
Ranking with Long-Term Constraints
by: Brantley, Kianté, et al.
Published: (2023)
by: Brantley, Kianté, et al.
Published: (2023)
Personas within Parameters: Fine-Tuning Small Language Models with Low-Rank Adapters to Mimic User Behaviors
by: Thakur, Himanshu, et al.
Published: (2025)
by: Thakur, Himanshu, et al.
Published: (2025)
NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
by: Lee, Chankyu, et al.
Published: (2024)
by: Lee, Chankyu, et al.
Published: (2024)
UniGLM: Training One Unified Language Model for Text-Attributed Graph Embedding
by: Fang, Yi, et al.
Published: (2024)
by: Fang, Yi, et al.
Published: (2024)
ReasonRank: Empowering Passage Ranking with Strong Reasoning Ability
by: Liu, Wenhan, et al.
Published: (2025)
by: Liu, Wenhan, et al.
Published: (2025)
ACER: Automatic Language Model Context Extension via Retrieval
by: Gao, Luyu, et al.
Published: (2024)
by: Gao, Luyu, et al.
Published: (2024)
HAPO: Training Language Models to Reason Concisely via History-Aware Policy Optimization
by: Huang, Chengyu, et al.
Published: (2025)
by: Huang, Chengyu, et al.
Published: (2025)
StealthRank: LLM Ranking Manipulation via Stealthy Prompt Optimization
by: Tang, Yiming, et al.
Published: (2025)
by: Tang, Yiming, et al.
Published: (2025)
RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs
by: Yu, Yue, et al.
Published: (2024)
by: Yu, Yue, et al.
Published: (2024)
Personalized Product Search Ranking: A Multi-Task Learning Approach with Tabular and Non-Tabular Data
by: Morishetti, Lalitesh, et al.
Published: (2025)
by: Morishetti, Lalitesh, et al.
Published: (2025)
Large Language Models are Learnable Planners for Long-Term Recommendation
by: Shi, Wentao, et al.
Published: (2024)
by: Shi, Wentao, et al.
Published: (2024)
Ranked List Truncation for Large Language Model-based Re-Ranking
by: Meng, Chuan, et al.
Published: (2024)
by: Meng, Chuan, et al.
Published: (2024)
Aligning LLM Agents by Learning Latent Preference from User Edits
by: Gao, Ge, et al.
Published: (2024)
by: Gao, Ge, et al.
Published: (2024)
RankPO: Preference Optimization for Job-Talent Matching
by: Zhang, Yafei, et al.
Published: (2025)
by: Zhang, Yafei, et al.
Published: (2025)
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF
by: Gao, Zhaolin, et al.
Published: (2024)
by: Gao, Zhaolin, et al.
Published: (2024)
Pessimistic Off-Policy Optimization for Learning to Rank
by: Cief, Matej, et al.
Published: (2022)
by: Cief, Matej, et al.
Published: (2022)
AcuRank: Uncertainty-Aware Adaptive Computation for Listwise Reranking
by: Yoon, Soyoung, et al.
Published: (2025)
by: Yoon, Soyoung, et al.
Published: (2025)
UP5: Unbiased Foundation Model for Fairness-aware Recommendation
by: Hua, Wenyue, et al.
Published: (2023)
by: Hua, Wenyue, et al.
Published: (2023)
DIVE: Embedding Compression via Self-Limiting Gradient Updates
by: Zhao, Dongfang
Published: (2026)
by: Zhao, Dongfang
Published: (2026)
Large Language Model Augmented Exercise Retrieval for Personalized Language Learning
by: Xu, Austin, et al.
Published: (2024)
by: Xu, Austin, et al.
Published: (2024)
Automated Query-Product Relevance Labeling using Large Language Models for E-commerce Search
by: Sachdev, Jayant, et al.
Published: (2025)
by: Sachdev, Jayant, et al.
Published: (2025)
MedSlice: Fine-Tuned Large Language Models for Secure Clinical Note Sectioning
by: Davis, Joshua, et al.
Published: (2025)
by: Davis, Joshua, et al.
Published: (2025)
Agentic Entropy-Balanced Policy Optimization
by: Dong, Guanting, et al.
Published: (2025)
by: Dong, Guanting, et al.
Published: (2025)
MixLM: High-Throughput and Effective LLM Ranking via Text-Embedding Mix-Interaction
by: Li, Guoyao, et al.
Published: (2025)
by: Li, Guoyao, et al.
Published: (2025)
On Bilingual Lexicon Induction with Large Language Models
by: Li, Yaoyiran, et al.
Published: (2023)
by: Li, Yaoyiran, et al.
Published: (2023)
Dataset Reset Policy Optimization for RLHF
by: Chang, Jonathan D., et al.
Published: (2024)
by: Chang, Jonathan D., et al.
Published: (2024)
The Factuality of Large Language Models in the Legal Domain
by: Hamdani, Rajaa El, et al.
Published: (2024)
by: Hamdani, Rajaa El, et al.
Published: (2024)
User Embedding Model for Personalized Language Prompting
by: Doddapaneni, Sumanth, et al.
Published: (2024)
by: Doddapaneni, Sumanth, et al.
Published: (2024)
Approaching Human-Level Forecasting with Language Models
by: Halawi, Danny, et al.
Published: (2024)
by: Halawi, Danny, et al.
Published: (2024)
AutoTask: Task Aware Multi-Faceted Single Model for Multi-Task Ads Relevance
by: Guo, Shouchang, et al.
Published: (2024)
by: Guo, Shouchang, et al.
Published: (2024)
Retrieval meets Long Context Large Language Models
by: Xu, Peng, et al.
Published: (2023)
by: Xu, Peng, et al.
Published: (2023)
Graph-based Confidence Calibration for Large Language Models
by: Li, Yukun, et al.
Published: (2024)
by: Li, Yukun, et al.
Published: (2024)
Uncertainty-Aware Explainable Recommendation with Large Language Models
by: Peng, Yicui, et al.
Published: (2024)
by: Peng, Yicui, et al.
Published: (2024)
Offline RL for Adaptive Policy Retrieval in Prior Authorization
by: Sharifullin, Ruslan, et al.
Published: (2026)
by: Sharifullin, Ruslan, et al.
Published: (2026)
A Surprising Failure? Multimodal LLMs and the NLVR Challenge
by: Wu, Anne, et al.
Published: (2024)
by: Wu, Anne, et al.
Published: (2024)
CrackSQL: A Hybrid SQL Dialect Translation System Powered by Large Language Models
by: Zhou, Wei, et al.
Published: (2025)
by: Zhou, Wei, et al.
Published: (2025)
PaCE: Parsimonious Concept Engineering for Large Language Models
by: Luo, Jinqi, et al.
Published: (2024)
by: Luo, Jinqi, et al.
Published: (2024)
News Recommendation with Category Description by a Large Language Model
by: Yada, Yuki, et al.
Published: (2024)
by: Yada, Yuki, et al.
Published: (2024)
TableRAG: Million-Token Table Understanding with Language Models
by: Chen, Si-An, et al.
Published: (2024)
by: Chen, Si-An, et al.
Published: (2024)
Similar Items
-
WildVis: Open Source Visualizer for Million-Scale Chat Logs in the Wild
by: Deng, Yuntian, et al.
Published: (2024) -
Ranking with Long-Term Constraints
by: Brantley, Kianté, et al.
Published: (2023) -
Personas within Parameters: Fine-Tuning Small Language Models with Low-Rank Adapters to Mimic User Behaviors
by: Thakur, Himanshu, et al.
Published: (2025) -
NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
by: Lee, Chankyu, et al.
Published: (2024) -
UniGLM: Training One Unified Language Model for Text-Attributed Graph Embedding
by: Fang, Yi, et al.
Published: (2024)