Fast Slate Policy Optimization: Going Beyond Plackett-Luce
Fuente:
arXiv
Saved in:
| Main Authors: | Sakhi, Otmane, Rohde, David, Chopin, Nicolas |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Probabilistic Rank and Reward: A Scalable Model for Slate Recommendation
by: Aouali, Imad, et al.
Published: (2022)
by: Aouali, Imad, et al.
Published: (2022)
Prompt-to-Slate: Diffusion Models for Prompt-Conditioned Slate Generation
by: Tomasi, Federico, et al.
Published: (2024)
by: Tomasi, Federico, et al.
Published: (2024)
RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents
by: Aouali, Imad, et al.
Published: (2026)
by: Aouali, Imad, et al.
Published: (2026)
Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning
by: Sakhi, Otmane, et al.
Published: (2024)
by: Sakhi, Otmane, et al.
Published: (2024)
Self-Consistency via Marginal Sharpening
by: Arzhantsev, Aleksei, et al.
Published: (2026)
by: Arzhantsev, Aleksei, et al.
Published: (2026)
Exploiting Similarities in A/B Testing with Off-Policy Estimation
by: Sakhi, Otmane, et al.
Published: (2025)
by: Sakhi, Otmane, et al.
Published: (2025)
Off-Policy Learning in Large Action Spaces: Optimization Matters More Than Estimation
by: Aouali, Imad, et al.
Published: (2025)
by: Aouali, Imad, et al.
Published: (2025)
Position Paper: Why the Shooting in the Dark Method Dominates Recommender Systems Practice; A Call to Abandon Anti-Utopian Thinking
by: Rohde, David
Published: (2024)
by: Rohde, David
Published: (2024)
Sequential Off-Policy Learning with Logarithmic Smoothing
by: Haddouche, Maxime, et al.
Published: (2025)
by: Haddouche, Maxime, et al.
Published: (2025)
Non-Linear Counterfactual Aggregate Optimization
by: Heymann, Benjamin, et al.
Published: (2025)
by: Heymann, Benjamin, et al.
Published: (2025)
Proximal Ranking Policy Optimization for Practical Safety in Counterfactual Learning to Rank
by: Gupta, Shashank, et al.
Published: (2024)
by: Gupta, Shashank, et al.
Published: (2024)
Confounding is a Pervasive Problem in Real World Recommender Systems
by: Merkov, Alexander, et al.
Published: (2025)
by: Merkov, Alexander, et al.
Published: (2025)
$Δ\text{-}{\rm OPE}$: Off-Policy Estimation with Pairs of Policies
by: Jeunen, Olivier, et al.
Published: (2024)
by: Jeunen, Olivier, et al.
Published: (2024)
Pessimistic Off-Policy Optimization for Learning to Rank
by: Cief, Matej, et al.
Published: (2022)
by: Cief, Matej, et al.
Published: (2022)
Ultra Fast Warm Start Solution for Graph Recommendations
by: Yusupov, Viacheslav, et al.
Published: (2025)
by: Yusupov, Viacheslav, et al.
Published: (2025)
Accelerating Matrix Factorization by Dynamic Pruning for Fast Recommendation
by: Wu, Yining, et al.
Published: (2024)
by: Wu, Yining, et al.
Published: (2024)
ContextGNN: Beyond Two-Tower Recommendation Systems
by: Yuan, Yiwen, et al.
Published: (2024)
by: Yuan, Yiwen, et al.
Published: (2024)
Aligning GPTRec with Beyond-Accuracy Goals with Reinforcement Learning
by: Petrov, Aleksandr, et al.
Published: (2024)
by: Petrov, Aleksandr, et al.
Published: (2024)
Beyond Item Dissimilarities: Diversifying by Intent in Recommender Systems
by: Wang, Yuyan, et al.
Published: (2024)
by: Wang, Yuyan, et al.
Published: (2024)
Off-Policy Evaluation and Learning for Matching Markets
by: Hayashi, Yudai, et al.
Published: (2025)
by: Hayashi, Yudai, et al.
Published: (2025)
Multi-Objective Recommendation via Multivariate Policy Learning
by: Jeunen, Olivier, et al.
Published: (2024)
by: Jeunen, Olivier, et al.
Published: (2024)
IntOPE: Off-Policy Evaluation in the Presence of Interference
by: Bai, Yuqi, et al.
Published: (2024)
by: Bai, Yuqi, et al.
Published: (2024)
Optimal Baseline Corrections for Off-Policy Contextual Bandits
by: Gupta, Shashank, et al.
Published: (2024)
by: Gupta, Shashank, et al.
Published: (2024)
Scaling Down, Serving Fast: Compressing and Deploying Efficient LLMs for Recommendation Systems
by: Behdin, Kayhan, et al.
Published: (2025)
by: Behdin, Kayhan, et al.
Published: (2025)
Cross-Sensory Brain Passage Retrieval: Scaling Beyond Visual to Audio
by: McGuire, Niall, et al.
Published: (2026)
by: McGuire, Niall, et al.
Published: (2026)
Beyond Static Evaluation: Rethinking the Assessment of Personalized Agent Adaptability in Information Retrieval
by: Kaur, Kirandeep, et al.
Published: (2025)
by: Kaur, Kirandeep, et al.
Published: (2025)
Policy-Guided Causal State Representation for Offline Reinforcement Learning Recommendation
by: Wang, Siyu, et al.
Published: (2025)
by: Wang, Siyu, et al.
Published: (2025)
Additive Control Variates Dominate Self-Normalisation in Off-Policy Evaluation
by: Jeunen, Olivier, et al.
Published: (2026)
by: Jeunen, Olivier, et al.
Published: (2026)
Understanding the Effects of the Baidu-ULTR Logging Policy on Two-Tower Models
by: de Haan, Morris, et al.
Published: (2024)
by: de Haan, Morris, et al.
Published: (2024)
Learning to Bid in Repeated Second-Price Auctions with Dynamic Values and Aggregated Feedback
by: Heymann, Benjamin, et al.
Published: (2026)
by: Heymann, Benjamin, et al.
Published: (2026)
Content Moderation in TV Search: Balancing Policy Compliance, Relevance, and User Experience
by: Hande, Adeep, et al.
Published: (2025)
by: Hande, Adeep, et al.
Published: (2025)
CASP: Support-Aware Offline Policy Selection for Two-Stage Recommender Systems
by: Chapagain, Nilson
Published: (2026)
by: Chapagain, Nilson
Published: (2026)
Beauty Beyond Words: Explainable Beauty Product Recommendations Using Ingredient-Based Product Attributes
by: Liu, Siliang, et al.
Published: (2024)
by: Liu, Siliang, et al.
Published: (2024)
On (Normalised) Discounted Cumulative Gain as an Off-Policy Evaluation Metric for Top-$n$ Recommendation
by: Jeunen, Olivier, et al.
Published: (2023)
by: Jeunen, Olivier, et al.
Published: (2023)
Minimizing Live Experiments in Recommender Systems: User Simulation to Evaluate Preference Elicitation Policies
by: Hsu, Chih-Wei, et al.
Published: (2024)
by: Hsu, Chih-Wei, et al.
Published: (2024)
GEO: Generative Engine Optimization
by: Aggarwal, Pranjal, et al.
Published: (2023)
by: Aggarwal, Pranjal, et al.
Published: (2023)
See Beyond a Single View: Multi-Attribution Learning Leads to Better Conversion Rate Prediction
by: Chen, Sishuo, et al.
Published: (2025)
by: Chen, Sishuo, et al.
Published: (2025)
From Entity Reliability to Clean Feedback: An Entity-Aware Denoising Framework Beyond Interaction-Level Signals
by: Liu, Ze, et al.
Published: (2025)
by: Liu, Ze, et al.
Published: (2025)
Meta Off-Policy Estimation
by: Jeunen, Olivier
Published: (2025)
by: Jeunen, Olivier
Published: (2025)
Make It Long, Keep It Fast: End-to-End 10K Long User Behavior Sequence Modeling for Billion-Scale Douyin Recommendation
by: Guan, Lin, et al.
Published: (2025)
by: Guan, Lin, et al.
Published: (2025)
Similar Items
-
Probabilistic Rank and Reward: A Scalable Model for Slate Recommendation
by: Aouali, Imad, et al.
Published: (2022) -
Prompt-to-Slate: Diffusion Models for Prompt-Conditioned Slate Generation
by: Tomasi, Federico, et al.
Published: (2024) -
RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents
by: Aouali, Imad, et al.
Published: (2026) -
Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning
by: Sakhi, Otmane, et al.
Published: (2024) -
Self-Consistency via Marginal Sharpening
by: Arzhantsev, Aleksei, et al.
Published: (2026)