SOLAR-RL: Semi-Online Long-horizon Assignment Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Jichao, Bian, Liuyang, Zhou, Yufeng, Xiao, Han, Pan, Yue, Wang, Guozhi, Wang, Hao, Wang, Zhaoxiong, Wen, Yafei, Chen, Xiaoxin, Ren, Shuai, Zeng, Lingfang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UI-Mem: Self-Evolving Experience Memory for Online Reinforcement Learning in Mobile GUI Agents
by: Xiao, Han, et al.
Published: (2026)
by: Xiao, Han, et al.
Published: (2026)
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
by: Xiao, Han, et al.
Published: (2025)
by: Xiao, Han, et al.
Published: (2025)
BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices
by: Lu, Xudong, et al.
Published: (2024)
by: Lu, Xudong, et al.
Published: (2024)
Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement Learning
by: Wang, Qi, et al.
Published: (2023)
by: Wang, Qi, et al.
Published: (2023)
SALT: Step-level Advantage Assignment for Long-horizon Agents via Trajectory Graph
by: Li, Jiazheng, et al.
Published: (2025)
by: Li, Jiazheng, et al.
Published: (2025)
Credit Assignment and Efficient Exploration based on Influence Scope in Multi-agent Reinforcement Learning
by: Han, Shuai, et al.
Published: (2025)
by: Han, Shuai, et al.
Published: (2025)
Asynchronous Credit Assignment for Multi-Agent Reinforcement Learning
by: Liang, Yongheng, et al.
Published: (2024)
by: Liang, Yongheng, et al.
Published: (2024)
Decoupling Ego-Motion from Target Dynamics via Dual-Interval Motion Cues for UAV Detection
by: Wang, Liuyang, et al.
Published: (2026)
by: Wang, Liuyang, et al.
Published: (2026)
VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning
by: Chen, Boyu, et al.
Published: (2025)
by: Chen, Boyu, et al.
Published: (2025)
Role-RL: Online Long-Context Processing with Role Reinforcement Learning for Distinct LLMs in Their Optimal Roles
by: He, Lewei, et al.
Published: (2024)
by: He, Lewei, et al.
Published: (2024)
GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment
by: Lv, Minxuan, et al.
Published: (2026)
by: Lv, Minxuan, et al.
Published: (2026)
Eosinophilic Granulomatous Polyangiitis Presenting With Finger Swelling as the Main Manifestation: A Case Report and Analysis
by: Lingfang Zhou
Published: (2026)
by: Lingfang Zhou
Published: (2026)
Tunable phase transitions from semimetals to Chern insulators in two-dimensional quadratic-band-crossing materials
by: Bian, Wen-Hao, et al.
Published: (2025)
by: Bian, Wen-Hao, et al.
Published: (2025)
Critical behavior around the fixed points driven by fermion-fermion interactions and disorders in the nodal-line superconductors
by: Bian, Wen-Hao, et al.
Published: (2024)
by: Bian, Wen-Hao, et al.
Published: (2024)
UNEX-RL: Reinforcing Long-Term Rewards in Multi-Stage Recommender Systems with UNidirectional EXecution
by: Zhang, Gengrui, et al.
Published: (2024)
by: Zhang, Gengrui, et al.
Published: (2024)
Toward Optimal Statistical Inference in Noisy Linear Quadratic Reinforcement Learning over a Finite Horizon
by: Pan, Bo, et al.
Published: (2025)
by: Pan, Bo, et al.
Published: (2025)
MoRL: Reinforced Reasoning for Unified Motion Understanding and Generation
by: Wang, Hongpeng, et al.
Published: (2026)
by: Wang, Hongpeng, et al.
Published: (2026)
Predicting the Potential Spread of Invasive Reptiles From Hong Kong and Taiwan to Other Regions of China
by: Chaosheng Mu, et al.
Published: (2026)
by: Chaosheng Mu, et al.
Published: (2026)
Search-Based Credit Assignment for Offline Preference-Based Reinforcement Learning
by: Gao, Xiancheng, et al.
Published: (2025)
by: Gao, Xiancheng, et al.
Published: (2025)
Optimistic Critic Reconstruction and Constrained Fine-Tuning for General Offline-to-Online RL
by: Luo, Qin-Wen, et al.
Published: (2024)
by: Luo, Qin-Wen, et al.
Published: (2024)
AdamO: A Collapse-Suppressed Optimizer for Offline RL
by: Qiao, Nan, et al.
Published: (2026)
by: Qiao, Nan, et al.
Published: (2026)
A Quantum Technology for Reinforcement Learning on Channel Assignment
by: Zengjing Chen, et al.
Published: (2024)
by: Zengjing Chen, et al.
Published: (2024)
AutoGMap: Learning to Map Large-scale Sparse Graphs on Memristive Crossbars
by: Lyu, Bo, et al.
Published: (2021)
by: Lyu, Bo, et al.
Published: (2021)
DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents
by: Wang, Taiyi, et al.
Published: (2024)
by: Wang, Taiyi, et al.
Published: (2024)
MobileRL: Online Agentic Reinforcement Learning for Mobile GUI Agents
by: Xu, Yifan, et al.
Published: (2025)
by: Xu, Yifan, et al.
Published: (2025)
LoongRL: Reinforcement Learning for Advanced Reasoning over Long Contexts
by: Wang, Siyuan, et al.
Published: (2025)
by: Wang, Siyuan, et al.
Published: (2025)
NurseSchedRL: Attention-Guided Reinforcement Learning for Nurse-Patient Assignment
by: Koduri, Harsha
Published: (2025)
by: Koduri, Harsha
Published: (2025)
Harmonizing Dense and Sparse Signals in Multi-turn RL: Dual-Horizon Credit Assignment for Industrial Sales Agents
by: Yang, Haojin, et al.
Published: (2026)
by: Yang, Haojin, et al.
Published: (2026)
SoPo: Text-to-Motion Generation Using Semi-Online Preference Optimization
by: Tan, Xiaofeng, et al.
Published: (2024)
by: Tan, Xiaofeng, et al.
Published: (2024)
Dual-Path Coupled Image Deraining Network via Spatial-Frequency Interaction
by: He, Yuhong, et al.
Published: (2024)
by: He, Yuhong, et al.
Published: (2024)
Scalable and Reliable Multi-agent Reinforcement Learning for Traffic Assignment
by: Wang, Leizhen, et al.
Published: (2025)
by: Wang, Leizhen, et al.
Published: (2025)
Three-Octave Supercontinuum Generation Spanning from Ultraviolet in Lithium Tantalate Waveguides
by: Wang, Lingfang, et al.
Published: (2025)
by: Wang, Lingfang, et al.
Published: (2025)
Online Robust Reinforcement Learning with General Function Approximation
by: Ghosh, Debamita, et al.
Published: (2025)
by: Ghosh, Debamita, et al.
Published: (2025)
The role of acetylation and deacetylation in cancer metabolism
by: Cuicui Wang, et al.
Published: (2025)
by: Cuicui Wang, et al.
Published: (2025)
Strain‐Amplified Interfacial Electric Field in an S‐Scheme Cu‐TCPP/CdSe Heterojunction for Efficient Photocatalytic H 2 O 2 Production in Pure Water
by: Yuqin Bian, et al.
Published: (2025)
by: Yuqin Bian, et al.
Published: (2025)
Evolution of Cooperation in Innovation Networks: A Network Public Goods Game Approach
by: Zilong Wang, et al.
Published: (2026)
by: Zilong Wang, et al.
Published: (2026)
Multiplicity and concentration of nontrivial solutions for Kirchhoff–Schrödinger–Poisson system with steep potential well
by: Liuyang Shao, et al.
Published: (2024)
by: Liuyang Shao, et al.
Published: (2024)
SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
by: Wang, Hanlin, et al.
Published: (2025)
by: Wang, Hanlin, et al.
Published: (2025)
DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts
by: Gai, Jiading, et al.
Published: (2026)
by: Gai, Jiading, et al.
Published: (2026)
Similar Items
-
UI-Mem: Self-Evolving Experience Memory for Online Reinforcement Learning in Mobile GUI Agents
by: Xiao, Han, et al.
Published: (2026) -
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
by: Wang, Hao, et al.
Published: (2026) -
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
by: Xiao, Han, et al.
Published: (2025) -
BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices
by: Lu, Xudong, et al.
Published: (2024) -
Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement Learning
by: Wang, Qi, et al.
Published: (2023)