Omni-Thinker: Scaling Multi-Task RL in LLMs with Hybrid Reward and Task Scheduling
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Derek, Zhou, Jiaming, Brunswic, Leo Maxime, Ghaddar, Abbas, Sun, Qianyi, Ma, Liheng, Luo, Yu, Li, Dong, Coates, Mark, Hao, Jianye, Zhang, Yingxue |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Logical Reasoning in Large Language Models through Graph-based Synthetic Data
by: Zhou, Jiaming, et al.
Published: (2024)
by: Zhou, Jiaming, et al.
Published: (2024)
A Theory of Multi-Agent Generative Flow Networks
by: Brunswic, Leo Maxime, et al.
Published: (2025)
by: Brunswic, Leo Maxime, et al.
Published: (2025)
CKGConv: General Graph Convolution with Continuous Kernels
by: Ma, Liheng, et al.
Published: (2024)
by: Ma, Liheng, et al.
Published: (2024)
Alexandrov Theorem for 2+1 flat radiant spacetimes
by: Brunswic, Léo
Published: (2020)
by: Brunswic, Léo
Published: (2020)
On branched coverings of singular $(G,X)$-manifolds
by: Brunswic, Léo
Published: (2020)
by: Brunswic, Léo
Published: (2020)
One Demo Is All It Takes: Planning Domain Derivation with LLMs from A Single Demonstration
by: Huang, Jinbang, et al.
Published: (2025)
by: Huang, Jinbang, et al.
Published: (2025)
Multi-resolution Time-Series Transformer for Long-term Forecasting
by: Zhang, Yitian, et al.
Published: (2023)
by: Zhang, Yitian, et al.
Published: (2023)
GraphPPD: Posterior Predictive Modelling for Graph-Level Inference
by: Pal, Soumyasundar, et al.
Published: (2025)
by: Pal, Soumyasundar, et al.
Published: (2025)
Plain Transformers Can be Powerful Graph Learners
by: Ma, Liheng, et al.
Published: (2025)
by: Ma, Liheng, et al.
Published: (2025)
A Theory of Non-Acyclic Generative Flow Networks
by: Brunswic, Leo Maxime, et al.
Published: (2023)
by: Brunswic, Leo Maxime, et al.
Published: (2023)
Sparse Decomposition of Graph Neural Networks
by: Hu, Yaochen, et al.
Published: (2024)
by: Hu, Yaochen, et al.
Published: (2024)
OmniEVA: Embodied Versatile Planner via Task-Adaptive 3D-Grounded and Embodiment-aware Reasoning
by: Liu, Yuecheng, et al.
Published: (2025)
by: Liu, Yuecheng, et al.
Published: (2025)
On the importance of Data Scale in Pretraining Arabic Language Models
by: Ghaddar, Abbas, et al.
Published: (2024)
by: Ghaddar, Abbas, et al.
Published: (2024)
BOSCH: Black-Box Binary Optimization for Short-Context Attention-Head Selection in LLMs
by: Ghaddar, Abbas, et al.
Published: (2026)
by: Ghaddar, Abbas, et al.
Published: (2026)
Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs
by: Alomrani, Mohammad Ali, et al.
Published: (2025)
by: Alomrani, Mohammad Ali, et al.
Published: (2025)
PosterOmni: Generalized Artistic Poster Creation via Task Distillation and Unified Reward Feedback
by: Chen, Sixiang, et al.
Published: (2026)
by: Chen, Sixiang, et al.
Published: (2026)
Beyond Distribution Sharpening: The Importance of Task Rewards
by: Mittal, Sarthak, et al.
Published: (2026)
by: Mittal, Sarthak, et al.
Published: (2026)
JPmHC Dynamical Isometry via Orthogonal Hyper-Connections
by: Sengupta, Biswa, et al.
Published: (2026)
by: Sengupta, Biswa, et al.
Published: (2026)
Extracting and Following Paths for Robust Relational Reasoning with Large Language Models
by: Zhang, Ge, et al.
Published: (2024)
by: Zhang, Ge, et al.
Published: (2024)
Generative Models in Decision Making: A Survey
by: Shao, Xinyu, et al.
Published: (2025)
by: Shao, Xinyu, et al.
Published: (2025)
Hybrid Reward Normalization for Process-supervised Non-verifiable Agentic Tasks
by: Xu, Peiran, et al.
Published: (2025)
by: Xu, Peiran, et al.
Published: (2025)
FEval-TTC: Fair Evaluation Protocol for Test-Time Compute
by: Rumiantsev, Pavel, et al.
Published: (2025)
by: Rumiantsev, Pavel, et al.
Published: (2025)
Refining Answer Distributions for Improved Large Language Model Reasoning
by: Pal, Soumyasundar, et al.
Published: (2024)
by: Pal, Soumyasundar, et al.
Published: (2024)
A Balanced Neuro-Symbolic Approach for Commonsense Abductive Logic
by: Cotnareanu, Joseph, et al.
Published: (2026)
by: Cotnareanu, Joseph, et al.
Published: (2026)
Personalized Negative Reservoir for Incremental Learning in Recommender Systems
by: Valkanas, Antonios, et al.
Published: (2024)
by: Valkanas, Antonios, et al.
Published: (2024)
Learning Reward Machines in Cooperative Multi-Agent Tasks
by: Ardon, Leo, et al.
Published: (2023)
by: Ardon, Leo, et al.
Published: (2023)
H-WM: Robotic Task and Motion Planning Guided by Hierarchical World Model
by: Huang, Jinbang, et al.
Published: (2026)
by: Huang, Jinbang, et al.
Published: (2026)
Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning
by: Mahrooghi, Ilia, et al.
Published: (2026)
by: Mahrooghi, Ilia, et al.
Published: (2026)
uTeBC-NLP at SemEval-2024 Task 9: Can LLMs be Lateral Thinkers?
by: Sadeghi, Pouya, et al.
Published: (2024)
by: Sadeghi, Pouya, et al.
Published: (2024)
Ergodic Generative Flows
by: Brunswic, Leo Maxime, et al.
Published: (2025)
by: Brunswic, Leo Maxime, et al.
Published: (2025)
SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling
by: Zhang, Yiqi, et al.
Published: (2026)
by: Zhang, Yiqi, et al.
Published: (2026)
Advancing Multi-Agent RAG Systems with Minimalist Reinforcement Learning
by: Wu, Yihong, et al.
Published: (2025)
by: Wu, Yihong, et al.
Published: (2025)
DyG2Vec: Efficient Representation Learning for Dynamic Graphs
by: Alomrani, Mohammad Ali, et al.
Published: (2022)
by: Alomrani, Mohammad Ali, et al.
Published: (2022)
Exploring RL-based LLM Training for Formal Language Tasks with Programmed Rewards
by: Padula, Alexander G., et al.
Published: (2024)
by: Padula, Alexander G., et al.
Published: (2024)
Parametric Feature Transfer: One-shot Federated Learning with Foundation Models
by: Beitollahi, Mahdi, et al.
Published: (2024)
by: Beitollahi, Mahdi, et al.
Published: (2024)
SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards
by: Batra, Hunar, et al.
Published: (2025)
by: Batra, Hunar, et al.
Published: (2025)
MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
by: Wang, Chengyao, et al.
Published: (2025)
by: Wang, Chengyao, et al.
Published: (2025)
Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
by: Wu, Runzhe, et al.
Published: (2025)
by: Wu, Runzhe, et al.
Published: (2025)
When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?
by: Hatgis-Kessell, Stephane, et al.
Published: (2026)
by: Hatgis-Kessell, Stephane, et al.
Published: (2026)
OmniOVCD: Streamlining Open-Vocabulary Change Detection with SAM 3
by: Zhang, Xu, et al.
Published: (2026)
by: Zhang, Xu, et al.
Published: (2026)
Similar Items
-
Enhancing Logical Reasoning in Large Language Models through Graph-based Synthetic Data
by: Zhou, Jiaming, et al.
Published: (2024) -
A Theory of Multi-Agent Generative Flow Networks
by: Brunswic, Leo Maxime, et al.
Published: (2025) -
CKGConv: General Graph Convolution with Continuous Kernels
by: Ma, Liheng, et al.
Published: (2024) -
Alexandrov Theorem for 2+1 flat radiant spacetimes
by: Brunswic, Léo
Published: (2020) -
On branched coverings of singular $(G,X)$-manifolds
by: Brunswic, Léo
Published: (2020)