Scouting By Reward: VLM-TO-IRL-Driven Player Selection For Esports
Fuente:
arXiv
Guardado en:
| Autores principales: | Yan, Qing, Yang, Wenyu, Wang, Yufei, Ma, Wenhao, Hu, Linchong, Jin, Yifei, Dahbura, Anton |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning
por: Wan, Zhenglin, et al.
Publicado: (2025)
por: Wan, Zhenglin, et al.
Publicado: (2025)
PandaSkill - Player Performance and Skill Rating in Esports: Application to League of Legends
por: De Bois, Maxime, et al.
Publicado: (2025)
por: De Bois, Maxime, et al.
Publicado: (2025)
The Dark Side of Rich Rewards: Understanding and Mitigating Noise in VLM Rewards
por: Huang, Sukai, et al.
Publicado: (2024)
por: Huang, Sukai, et al.
Publicado: (2024)
ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation
por: Feng, Youhe, et al.
Publicado: (2026)
por: Feng, Youhe, et al.
Publicado: (2026)
CoMI-IRL: Contrastive Multi-Intention Inverse Reinforcement Learning
por: Mone, Antonio, et al.
Publicado: (2026)
por: Mone, Antonio, et al.
Publicado: (2026)
From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models
por: Gumbsch, Christian, et al.
Publicado: (2026)
por: Gumbsch, Christian, et al.
Publicado: (2026)
Model-Free Inference of Investor Preferences: A Relative Entropy IRL Approach
por: Xu, Chen
Publicado: (2026)
por: Xu, Chen
Publicado: (2026)
VLM-C4L: Continual Core Dataset Learning with Corner Case Optimization via Vision-Language Models for Autonomous Driving
por: Hu, Haibo, et al.
Publicado: (2025)
por: Hu, Haibo, et al.
Publicado: (2025)
Goal-Driven Reward by Video Diffusion Models for Reinforcement Learning
por: Wang, Qi, et al.
Publicado: (2025)
por: Wang, Qi, et al.
Publicado: (2025)
IRL for Restless Multi-Armed Bandits with Applications in Maternal and Child Health
por: Jain, Gauri, et al.
Publicado: (2024)
por: Jain, Gauri, et al.
Publicado: (2024)
TreeIRL: Safe Urban Driving with Tree Search and Inverse Reinforcement Learning
por: Tomov, Momchil S., et al.
Publicado: (2025)
por: Tomov, Momchil S., et al.
Publicado: (2025)
SurgIRL: Towards Life-Long Learning for Surgical Automation by Incremental Reinforcement Learning
por: Ho, Yun-Jie, et al.
Publicado: (2024)
por: Ho, Yun-Jie, et al.
Publicado: (2024)
DuoGuard: A Two-Player RL-Driven Framework for Multilingual LLM Guardrails
por: Deng, Yihe, et al.
Publicado: (2025)
por: Deng, Yihe, et al.
Publicado: (2025)
AutoScout: Structured Optimization for Automating ML System Configuration
por: Shong, Jimmy, et al.
Publicado: (2026)
por: Shong, Jimmy, et al.
Publicado: (2026)
Reinforcement Learning From Imperfect Corrective Actions And Proxy Rewards
por: Jiang, Zhaohui, et al.
Publicado: (2024)
por: Jiang, Zhaohui, et al.
Publicado: (2024)
Semi-Supervised Reward Modeling via Iterative Self-Training
por: He, Yifei, et al.
Publicado: (2024)
por: He, Yifei, et al.
Publicado: (2024)
CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward Modeling
por: Liu, Dengcan, et al.
Publicado: (2026)
por: Liu, Dengcan, et al.
Publicado: (2026)
Maximum Causal Entropy IRL in Mean-Field Games and GNEP Framework for Forward RL
por: Anahtarci, Berkay, et al.
Publicado: (2024)
por: Anahtarci, Berkay, et al.
Publicado: (2024)
Learning Dynamics of VLM Finetuning
por: Zhang, Jusheng, et al.
Publicado: (2025)
por: Zhang, Jusheng, et al.
Publicado: (2025)
Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
por: Li, Muyang, et al.
Publicado: (2026)
por: Li, Muyang, et al.
Publicado: (2026)
SafeCoT: Improving VLM Safety with Minimal Reasoning
por: Ma, Jiachen, et al.
Publicado: (2025)
por: Ma, Jiachen, et al.
Publicado: (2025)
Scout Before You Attend: Sketch-and-Walk Sparse Attention for Efficient LLM Inference
por: Le, Hoang Anh Duy, et al.
Publicado: (2026)
por: Le, Hoang Anh Duy, et al.
Publicado: (2026)
Reward-Zero: Language Embedding Driven Implicit Reward Mechanisms for Reinforcement Learning
por: Zhang, Heng, et al.
Publicado: (2026)
por: Zhang, Heng, et al.
Publicado: (2026)
CoIRL-AD: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous Driving
por: Zheng, Xiaoji, et al.
Publicado: (2025)
por: Zheng, Xiaoji, et al.
Publicado: (2025)
Rethinking Evaluation Metric for Probability Estimation Models Using Esports Data
por: Choi, Euihyeon, et al.
Publicado: (2023)
por: Choi, Euihyeon, et al.
Publicado: (2023)
ScoutAttention: Efficient KV Cache Offloading via Layer-Ahead CPU Pre-computation for LLM Inference
por: Zhang, Qiuyang, et al.
Publicado: (2026)
por: Zhang, Qiuyang, et al.
Publicado: (2026)
Reward-RAG: Enhancing RAG with Reward Driven Supervision
por: Nguyen, Thang, et al.
Publicado: (2024)
por: Nguyen, Thang, et al.
Publicado: (2024)
RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedback
por: Wang, Yufei, et al.
Publicado: (2024)
por: Wang, Yufei, et al.
Publicado: (2024)
A Lecture Note on Offline RL and IRL, Part II: Foundations of Inverse Reinforcement Learning and Dynamic Discrete Choice Models
por: Kang, Enoch Hyunwook
Publicado: (2026)
por: Kang, Enoch Hyunwook
Publicado: (2026)
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
por: Duan, Kaiwen, et al.
Publicado: (2025)
por: Duan, Kaiwen, et al.
Publicado: (2025)
Diffusion Classifier-Driven Reward for Offline Preference-based Reinforcement Learning
por: Pang, Teng, et al.
Publicado: (2025)
por: Pang, Teng, et al.
Publicado: (2025)
Enhancing Financial Market Predictions: Causality-Driven Feature Selection
por: Liang, Wenhao, et al.
Publicado: (2024)
por: Liang, Wenhao, et al.
Publicado: (2024)
Automatic Reward Shaping from Multi-Objective Human Heuristics
por: Xie, Yuqing, et al.
Publicado: (2025)
por: Xie, Yuqing, et al.
Publicado: (2025)
Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown
por: Lou, Xingzhou, et al.
Publicado: (2024)
por: Lou, Xingzhou, et al.
Publicado: (2024)
Mean-Field Diffuser: Scaling Offline MARL to Thousands of Agents
por: Li, Wenhao, et al.
Publicado: (2026)
por: Li, Wenhao, et al.
Publicado: (2026)
Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning
por: Ma, Weiyu, et al.
Publicado: (2026)
por: Ma, Weiyu, et al.
Publicado: (2026)
DavIR: Data Selection via Implicit Reward for Large Language Models
por: Zhou, Haotian, et al.
Publicado: (2023)
por: Zhou, Haotian, et al.
Publicado: (2023)
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
por: Qi, Xuan, et al.
Publicado: (2025)
por: Qi, Xuan, et al.
Publicado: (2025)
Laplacian Canonization: A Minimalist Approach to Sign and Basis Invariant Spectral Embedding
por: Ma, Jiangyan, et al.
Publicado: (2023)
por: Ma, Jiangyan, et al.
Publicado: (2023)
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
por: Beigi, Mohammad, et al.
Publicado: (2026)
por: Beigi, Mohammad, et al.
Publicado: (2026)
Ejemplares similares
-
FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning
por: Wan, Zhenglin, et al.
Publicado: (2025) -
PandaSkill - Player Performance and Skill Rating in Esports: Application to League of Legends
por: De Bois, Maxime, et al.
Publicado: (2025) -
The Dark Side of Rich Rewards: Understanding and Mitigating Noise in VLM Rewards
por: Huang, Sukai, et al.
Publicado: (2024) -
ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation
por: Feng, Youhe, et al.
Publicado: (2026) -
CoMI-IRL: Contrastive Multi-Intention Inverse Reinforcement Learning
por: Mone, Antonio, et al.
Publicado: (2026)