Don't Let Bandit Feedback Pull Continual LLM-Recommender Updates Off Target
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Taesan, Yun, Hyeongjun, Choo, Jaegul, Park, Chung |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do's and Don'ts: Learning Desirable Skills with Instruction Videos
by: Kim, Hyunseung, et al.
Published: (2024)
by: Kim, Hyunseung, et al.
Published: (2024)
When Model Meets New Normals: Test-time Adaptation for Unsupervised Time-series Anomaly Detection
by: Kim, Dongmin, et al.
Published: (2023)
by: Kim, Dongmin, et al.
Published: (2023)
Benchmarking is Broken -- Don't Let AI be its Own Judge
by: Cheng, Zerui, et al.
Published: (2025)
by: Cheng, Zerui, et al.
Published: (2025)
EPIC: Effective Prompting for Imbalanced-Class Data Synthesis in Tabular Data Classification via Large Language Models
by: Kim, Jinhee, et al.
Published: (2024)
by: Kim, Jinhee, et al.
Published: (2024)
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
by: Plyusov, Daniil, et al.
Published: (2026)
by: Plyusov, Daniil, et al.
Published: (2026)
Pacer and Runner: Cooperative Learning Framework between Single- and Cross-Domain Sequential Recommendation
by: Park, Chung, et al.
Published: (2024)
by: Park, Chung, et al.
Published: (2024)
Self-Supervised Contrastive Learning for Long-term Forecasting
by: Park, Junwoo, et al.
Published: (2024)
by: Park, Junwoo, et al.
Published: (2024)
Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess
by: Hwang, Dongyoon, et al.
Published: (2025)
by: Hwang, Dongyoon, et al.
Published: (2025)
Towards Trustworthy LLM-Based Recommendation via Rationale Integration
by: Park, Chung, et al.
Published: (2025)
by: Park, Chung, et al.
Published: (2025)
Don't Retrain, Just Reuse: Recovering Dual-Target Molecules from Single-Target Diffusion Models
by: Zeng, Qingyuan, et al.
Published: (2026)
by: Zeng, Qingyuan, et al.
Published: (2026)
Influential Bandits: Pulling an Arm May Change the Environment
by: Sato, Ryoma, et al.
Published: (2025)
by: Sato, Ryoma, et al.
Published: (2025)
Trust, Don't Trust, or Flip: Robust Preference-Based Reinforcement Learning with Multi-Expert Feedback
by: Hosseini, Seyed Amir, et al.
Published: (2026)
by: Hosseini, Seyed Amir, et al.
Published: (2026)
Don't Freeze, Don't Crash: Extending the Safe Operating Range of Neural Navigation in Dense Crowds
by: Zhang, Jiefu, et al.
Published: (2026)
by: Zhang, Jiefu, et al.
Published: (2026)
Know What You Don't Know: Uncertainty Calibration of Process Reward Models
by: Park, Young-Jin, et al.
Published: (2025)
by: Park, Young-Jin, et al.
Published: (2025)
Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization
by: Zhou, Jin Peng, et al.
Published: (2024)
by: Zhou, Jin Peng, et al.
Published: (2024)
Don't Forget the Critic: Value-Based Data Rehearsal for Multi-Cyclic Continual Reinforcement Learning
by: Poole, Benjamin, et al.
Published: (2026)
by: Poole, Benjamin, et al.
Published: (2026)
Slow and Steady Wins the Race: Maintaining Plasticity with Hare and Tortoise Networks
by: Lee, Hojoon, et al.
Published: (2024)
by: Lee, Hojoon, et al.
Published: (2024)
Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints
by: Bowyer, Sam, et al.
Published: (2025)
by: Bowyer, Sam, et al.
Published: (2025)
Transformers Don't In-Context Learn Least Squares Regression
by: Hill, Joshua, et al.
Published: (2025)
by: Hill, Joshua, et al.
Published: (2025)
Hybrid Federated Learning for Noise-Robust Training
by: Kim, Yongjun, et al.
Published: (2026)
by: Kim, Yongjun, et al.
Published: (2026)
Don't Forget It! Conditional Sparse Autoencoder Clamping Works for Unlearning
by: Khoriaty, Matthew, et al.
Published: (2025)
by: Khoriaty, Matthew, et al.
Published: (2025)
Don't Waste Your Time: Early Stopping Cross-Validation
by: Bergman, Edward, et al.
Published: (2024)
by: Bergman, Edward, et al.
Published: (2024)
LLM Cyber Evaluations Don't Capture Real-World Risk
by: Lukošiūtė, Kamilė, et al.
Published: (2025)
by: Lukošiūtė, Kamilė, et al.
Published: (2025)
Don't Shoot The Breeze: Topic Continuity Model Using Nonlinear Naive Bayes With Attention
by: Pi, Shu-Ting, et al.
Published: (2026)
by: Pi, Shu-Ting, et al.
Published: (2026)
Investigating Pre-Training Objectives for Generalization in Vision-Based Reinforcement Learning
by: Kim, Donghu, et al.
Published: (2024)
by: Kim, Donghu, et al.
Published: (2024)
xAI-Drop: Don't Use What You Cannot Explain
by: De Luca, Vincenzo Marco, et al.
Published: (2024)
by: De Luca, Vincenzo Marco, et al.
Published: (2024)
Don't Lag, RAG: Training-Free Adversarial Detection Using RAG
by: Kazoom, Roie, et al.
Published: (2025)
by: Kazoom, Roie, et al.
Published: (2025)
Don't be lazy: CompleteP enables compute-efficient deep transformers
by: Dey, Nolan, et al.
Published: (2025)
by: Dey, Nolan, et al.
Published: (2025)
Don't Judge by the Look: Towards Motion Coherent Video Representation
by: Zhang, Yitian, et al.
Published: (2024)
by: Zhang, Yitian, et al.
Published: (2024)
Prediction Bottlenecks Don't Discover Causal Structure (But Here's What They Actually Do)
by: Lade, Ankit Hemant, et al.
Published: (2026)
by: Lade, Ankit Hemant, et al.
Published: (2026)
mHC-lite: You Don't Need 20 Sinkhorn-Knopp Iterations
by: Yang, Yongyi, et al.
Published: (2026)
by: Yang, Yongyi, et al.
Published: (2026)
Don't stop me now: Rethinking Validation Criteria for Model Parameter Selection
by: Apicella, Andrea, et al.
Published: (2026)
by: Apicella, Andrea, et al.
Published: (2026)
Linear and Neural Dueling Bandits with Delayed Feedback
by: Wang, Xiangyi, et al.
Published: (2026)
by: Wang, Xiangyi, et al.
Published: (2026)
Know What You Don't Know: Selective Prediction for Early Exit DNNs
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
Don't throw the baby out with the bathwater: How and why deep learning for ARC
by: Cole, Jack, et al.
Published: (2025)
by: Cole, Jack, et al.
Published: (2025)
Trust, or Don't Predict: Introducing the CWSA Family for Confidence-Aware Model Evaluation
by: Shahnazari, Kourosh, et al.
Published: (2025)
by: Shahnazari, Kourosh, et al.
Published: (2025)
Fusing Reward and Dueling Feedback in Stochastic Bandits
by: Wang, Xuchuang, et al.
Published: (2025)
by: Wang, Xuchuang, et al.
Published: (2025)
Effective Off-Policy Evaluation and Learning in Contextual Combinatorial Bandits
by: Shimizu, Tatsuhiro, et al.
Published: (2024)
by: Shimizu, Tatsuhiro, et al.
Published: (2024)
Reasoning Models Don't Always Say What They Think
by: Chen, Yanda, et al.
Published: (2025)
by: Chen, Yanda, et al.
Published: (2025)
Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
by: Kachaev, Nikita, et al.
Published: (2025)
by: Kachaev, Nikita, et al.
Published: (2025)
Similar Items
-
Do's and Don'ts: Learning Desirable Skills with Instruction Videos
by: Kim, Hyunseung, et al.
Published: (2024) -
When Model Meets New Normals: Test-time Adaptation for Unsupervised Time-series Anomaly Detection
by: Kim, Dongmin, et al.
Published: (2023) -
Benchmarking is Broken -- Don't Let AI be its Own Judge
by: Cheng, Zerui, et al.
Published: (2025) -
EPIC: Effective Prompting for Imbalanced-Class Data Synthesis in Tabular Data Classification via Large Language Models
by: Kim, Jinhee, et al.
Published: (2024) -
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
by: Plyusov, Daniil, et al.
Published: (2026)