COSMO-RL: Towards Trustworthy LMRMs via Joint Safety and Stability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Yizhuo, Chen, Mingkang, Liu, Qiuhua, Weng, Fenghua, Qu, Wanying, Yang, Yue, Jiang, Yugang, Wu, Zuxuan, Fu, Yanwei, Shao, Wenqi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908577248051200
author Ding, Yizhuo
Chen, Mingkang
Liu, Qiuhua
Weng, Fenghua
Qu, Wanying
Yang, Yue
Jiang, Yugang
Wu, Zuxuan
Fu, Yanwei
Shao, Wenqi
author_facet Ding, Yizhuo
Chen, Mingkang
Liu, Qiuhua
Weng, Fenghua
Qu, Wanying
Yang, Yue
Jiang, Yugang
Wu, Zuxuan
Fu, Yanwei
Shao, Wenqi
contents Large Multimodal Reasoning Models (LMRMs) are moving into real applications, where they must be both useful and safe. Safety is especially challenging in multimodal settings: images and text can be combined to bypass guardrails, and single objective training can cause policy drift that yields over-refusal on benign inputs or unsafe compliance on risky ones. We present COSMO-RL, a mixed reinforcement learning framework that trains reasoning oriented LMRMs under multimodal, multitask, and multiobjective signals, and we release the resulting model, COSMO-R1. Our approach aims to let safety and capability grow together in one stable pipeline rather than competing during alignment. In experiments, COSMO-R1 improves safety while maintaining-and often improving multimodal reasoning and instruction following, shows stronger robustness to multimodal jailbreaks, and reduces unnecessary refusals. The framework also transfers across backbones with consistent gains. Ablations support the design choices, indicating a simple path to advancing safety and general capability together in LMRMs.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04196
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle COSMO-RL: Towards Trustworthy LMRMs via Joint Safety and Stability
Ding, Yizhuo
Chen, Mingkang
Liu, Qiuhua
Weng, Fenghua
Qu, Wanying
Yang, Yue
Jiang, Yugang
Wu, Zuxuan
Fu, Yanwei
Shao, Wenqi
Artificial Intelligence
Machine Learning
Large Multimodal Reasoning Models (LMRMs) are moving into real applications, where they must be both useful and safe. Safety is especially challenging in multimodal settings: images and text can be combined to bypass guardrails, and single objective training can cause policy drift that yields over-refusal on benign inputs or unsafe compliance on risky ones. We present COSMO-RL, a mixed reinforcement learning framework that trains reasoning oriented LMRMs under multimodal, multitask, and multiobjective signals, and we release the resulting model, COSMO-R1. Our approach aims to let safety and capability grow together in one stable pipeline rather than competing during alignment. In experiments, COSMO-R1 improves safety while maintaining-and often improving multimodal reasoning and instruction following, shows stronger robustness to multimodal jailbreaks, and reduces unnecessary refusals. The framework also transfers across backbones with consistent gains. Ablations support the design choices, indicating a simple path to advancing safety and general capability together in LMRMs.
title COSMO-RL: Towards Trustworthy LMRMs via Joint Safety and Stability
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.04196