Safely Exploring Novel Actions in Recommender Systems via Deployment-Efficient Policy Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Kiyohara, Haruka, Narita, Yusuke, Saito, Yuta, Tateno, Kei, Udagawa, Takuma |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Off-Policy Evaluation and Learning for the Future under Non-Stationarity
por: Shimizu, Tatsuhiro, et al.
Publicado: (2025)
por: Shimizu, Tatsuhiro, et al.
Publicado: (2025)
Counterfactual Reciprocal Recommender Systems for User-to-User Matching
por: Kawamura, Kazuki, et al.
Publicado: (2025)
por: Kawamura, Kazuki, et al.
Publicado: (2025)
Effective Off-Policy Evaluation and Learning in Contextual Combinatorial Bandits
por: Shimizu, Tatsuhiro, et al.
Publicado: (2024)
por: Shimizu, Tatsuhiro, et al.
Publicado: (2024)
SCOPE-RL: A Python Library for Offline Reinforcement Learning and Off-Policy Evaluation
por: Kiyohara, Haruka, et al.
Publicado: (2023)
por: Kiyohara, Haruka, et al.
Publicado: (2023)
Offline Contextual Bandits in the Presence of New Actions
por: Kishimoto, Ren, et al.
Publicado: (2026)
por: Kishimoto, Ren, et al.
Publicado: (2026)
Towards Assessing and Benchmarking Risk-Return Tradeoff of Off-Policy Evaluation
por: Kiyohara, Haruka, et al.
Publicado: (2023)
por: Kiyohara, Haruka, et al.
Publicado: (2023)
Off-Policy Evaluation for Ranking Policies under Deterministic Logging Policies
por: Tanaka, Koichi, et al.
Publicado: (2026)
por: Tanaka, Koichi, et al.
Publicado: (2026)
Not Just What, But When: Integrating Irregular Intervals to LLM for Sequential Recommendation
por: Du, Wei-Wei, et al.
Publicado: (2025)
por: Du, Wei-Wei, et al.
Publicado: (2025)
Prompt Optimization with Logged Bandit Data
por: Kiyohara, Haruka, et al.
Publicado: (2025)
por: Kiyohara, Haruka, et al.
Publicado: (2025)
Off-Policy Evaluation of Slate Bandit Policies via Optimizing Abstraction
por: Kiyohara, Haruka, et al.
Publicado: (2024)
por: Kiyohara, Haruka, et al.
Publicado: (2024)
Safe Deployment of Offline Reinforcement Learning via Input Convex Action Correction
por: Durkin, Alex, et al.
Publicado: (2025)
por: Durkin, Alex, et al.
Publicado: (2025)
Off-Policy Evaluation and Learning for Survival Outcomes under Censoring
por: Kubota, Kohsuke, et al.
Publicado: (2026)
por: Kubota, Kohsuke, et al.
Publicado: (2026)
Gradients as an Action: Towards Communication-Efficient Federated Recommender Systems via Adaptive Action Sharing
por: Lu, Zhufeng, et al.
Publicado: (2025)
por: Lu, Zhufeng, et al.
Publicado: (2025)
Peer Learning: Learning Complex Policies in Groups from Scratch via Action Recommendations
por: Derstroff, Cedric, et al.
Publicado: (2023)
por: Derstroff, Cedric, et al.
Publicado: (2023)
Safe Planning and Policy Optimization via World Model Learning
por: Latyshev, Artem, et al.
Publicado: (2025)
por: Latyshev, Artem, et al.
Publicado: (2025)
Q-realign: Piggybacking Realignment on Quantization for Safe and Efficient LLM Deployment
por: Tan, Qitao, et al.
Publicado: (2026)
por: Tan, Qitao, et al.
Publicado: (2026)
Data-Efficient Safe Policy Improvement Using Parametric Structure
por: Engelen, Kasper, et al.
Publicado: (2025)
por: Engelen, Kasper, et al.
Publicado: (2025)
From Algorithm to Hardware: A Survey on Efficient and Safe Deployment of Deep Neural Networks
por: Geng, Xue, et al.
Publicado: (2024)
por: Geng, Xue, et al.
Publicado: (2024)
PROTEA: Offline Evaluation and Iterative Refinement for Multi-Agent LLM Workflows
por: Kawamura, Kazuki, et al.
Publicado: (2026)
por: Kawamura, Kazuki, et al.
Publicado: (2026)
Learning Safe Numeric Planning Action Models
por: Mordoch, Argaman, et al.
Publicado: (2023)
por: Mordoch, Argaman, et al.
Publicado: (2023)
SymLight: Exploring Interpretable and Deployable Symbolic Policies for Traffic Signal Control
por: Liao, Xiao-Cheng, et al.
Publicado: (2025)
por: Liao, Xiao-Cheng, et al.
Publicado: (2025)
Towards Fast Safe Online Reinforcement Learning via Policy Finetuning
por: Chen, Keru, et al.
Publicado: (2024)
por: Chen, Keru, et al.
Publicado: (2024)
SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning
por: Anisimov, Maksim, et al.
Publicado: (2026)
por: Anisimov, Maksim, et al.
Publicado: (2026)
SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning
por: Zhang, Borong, et al.
Publicado: (2025)
por: Zhang, Borong, et al.
Publicado: (2025)
Safe Exploration via Policy Priors
por: Wendl, Manuel, et al.
Publicado: (2026)
por: Wendl, Manuel, et al.
Publicado: (2026)
A Constraint Programming Approach to Fair High School Course Scheduling
por: Kiyohara, Mitsuka, et al.
Publicado: (2024)
por: Kiyohara, Mitsuka, et al.
Publicado: (2024)
Safe Explicable Policy Search
por: Hanni, Akkamahadevi, et al.
Publicado: (2025)
por: Hanni, Akkamahadevi, et al.
Publicado: (2025)
Prediction with Action: Visual Policy Learning via Joint Denoising Process
por: Guo, Yanjiang, et al.
Publicado: (2024)
por: Guo, Yanjiang, et al.
Publicado: (2024)
Enhancing Policy Learning with World-Action Model
por: Han, Yuci, et al.
Publicado: (2026)
por: Han, Yuci, et al.
Publicado: (2026)
SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment
por: Li, Hao, et al.
Publicado: (2026)
por: Li, Hao, et al.
Publicado: (2026)
Safe Flow Q-Learning: Offline Safe Reinforcement Learning with Reachability-Based Flow Policies
por: Tayal, Mumuksh, et al.
Publicado: (2026)
por: Tayal, Mumuksh, et al.
Publicado: (2026)
STARec: An Efficient Agent Framework for Recommender Systems via Autonomous Deliberate Reasoning
por: Wu, Chenghao, et al.
Publicado: (2025)
por: Wu, Chenghao, et al.
Publicado: (2025)
SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred Trajectories
por: Burnwal, Returaj, et al.
Publicado: (2025)
por: Burnwal, Returaj, et al.
Publicado: (2025)
Iterative Batch Reinforcement Learning via Safe Diversified Model-based Policy Search
por: Najib, Amna, et al.
Publicado: (2024)
por: Najib, Amna, et al.
Publicado: (2024)
Efficient Multi-Task Learning via Generalist Recommender
por: Wang, Luyang, et al.
Publicado: (2025)
por: Wang, Luyang, et al.
Publicado: (2025)
NRR-Core: Non-Resolution Reasoning as a Computational Framework for Contextual Identity and Ambiguity Preservation
por: Saito, Kei
Publicado: (2025)
por: Saito, Kei
Publicado: (2025)
NRR-Phi: Text-to-State Mapping for Ambiguity Preservation in LLM Inference
por: Saito, Kei
Publicado: (2026)
por: Saito, Kei
Publicado: (2026)
Completion at the Boundary (CaB): Deployable Switching with Completion-Aware Control under Limited Calibration
por: Sano, Yusuke, et al.
Publicado: (2026)
por: Sano, Yusuke, et al.
Publicado: (2026)
Transferable Reinforcement Learning via Probabilistic Latent Embeddings and Dynamic Policy Adaptation for Sim-to-Real Deployment
por: Han, Gengyue, et al.
Publicado: (2026)
por: Han, Gengyue, et al.
Publicado: (2026)
Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization
por: Hua, Xingyuan, et al.
Publicado: (2026)
por: Hua, Xingyuan, et al.
Publicado: (2026)
Ejemplares similares
-
Off-Policy Evaluation and Learning for the Future under Non-Stationarity
por: Shimizu, Tatsuhiro, et al.
Publicado: (2025) -
Counterfactual Reciprocal Recommender Systems for User-to-User Matching
por: Kawamura, Kazuki, et al.
Publicado: (2025) -
Effective Off-Policy Evaluation and Learning in Contextual Combinatorial Bandits
por: Shimizu, Tatsuhiro, et al.
Publicado: (2024) -
SCOPE-RL: A Python Library for Offline Reinforcement Learning and Off-Policy Evaluation
por: Kiyohara, Haruka, et al.
Publicado: (2023) -
Offline Contextual Bandits in the Presence of New Actions
por: Kishimoto, Ren, et al.
Publicado: (2026)