Xiaomi MiMo-VL-Miloco Technical Report

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Jiaze, Chen, Jingyang, Qu, Yuxun, Xu, Shijie, Lin, Zhenru, Zhu, Junyou, Xu, Boshen, Tan, Wenhui, Fu, Pei, Ju, Jianzhong, Luo, Zhenbo, Luan, Jian
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911331324526592
author Li, Jiaze
Chen, Jingyang
Qu, Yuxun
Xu, Shijie
Lin, Zhenru
Zhu, Junyou
Xu, Boshen
Tan, Wenhui
Fu, Pei
Ju, Jianzhong
Luo, Zhenbo
Luan, Jian
author_facet Li, Jiaze
Chen, Jingyang
Qu, Yuxun
Xu, Shijie
Lin, Zhenru
Zhu, Junyou
Xu, Boshen
Tan, Wenhui
Fu, Pei
Ju, Jianzhong
Luo, Zhenbo
Luan, Jian
contents We open-source MiMo-VL-Miloco-7B and its quantized variant MiMo-VL-Miloco-7B-GGUF, a pair of home-centric vision-language models that achieve strong performance on both home-scenario understanding and general multimodal reasoning. Built on the MiMo-VL-7B backbone, MiMo-VL-Miloco-7B is specialized for smart-home environments, attaining leading F1 scores on gesture recognition and common home-scenario understanding, while also delivering consistent gains across video benchmarks such as Video-MME, Video-MMMU, and Charades-STA, as well as language understanding benchmarks including MMMU-Pro and MMLU-Pro. In our experiments, MiMo-VL-Miloco-7B outperforms strong closed-source and open-source baselines on home-scenario understanding and several multimodal reasoning benchmarks. To balance specialization and generality, we design a two-stage training pipeline that combines supervised fine-tuning with reinforcement learning based on Group Relative Policy Optimization, leveraging efficient multi-domain data. We further incorporate chain-of-thought supervision and token-budget-aware reasoning, enabling the model to learn knowledge in a data-efficient manner while also performing reasoning efficiently. Our analysis shows that targeted home-scenario training not only enhances activity and gesture understanding, but also improves text-only reasoning with only modest trade-offs on document-centric tasks. Model checkpoints, quantized GGUF weights, and our home-scenario evaluation toolkit are publicly available at https://github.com/XiaoMi/xiaomi-mimo-vl-miloco to support research and deployment in real-world smart-home applications.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17436
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Xiaomi MiMo-VL-Miloco Technical Report
Li, Jiaze
Chen, Jingyang
Qu, Yuxun
Xu, Shijie
Lin, Zhenru
Zhu, Junyou
Xu, Boshen
Tan, Wenhui
Fu, Pei
Ju, Jianzhong
Luo, Zhenbo
Luan, Jian
Computer Vision and Pattern Recognition
We open-source MiMo-VL-Miloco-7B and its quantized variant MiMo-VL-Miloco-7B-GGUF, a pair of home-centric vision-language models that achieve strong performance on both home-scenario understanding and general multimodal reasoning. Built on the MiMo-VL-7B backbone, MiMo-VL-Miloco-7B is specialized for smart-home environments, attaining leading F1 scores on gesture recognition and common home-scenario understanding, while also delivering consistent gains across video benchmarks such as Video-MME, Video-MMMU, and Charades-STA, as well as language understanding benchmarks including MMMU-Pro and MMLU-Pro. In our experiments, MiMo-VL-Miloco-7B outperforms strong closed-source and open-source baselines on home-scenario understanding and several multimodal reasoning benchmarks. To balance specialization and generality, we design a two-stage training pipeline that combines supervised fine-tuning with reinforcement learning based on Group Relative Policy Optimization, leveraging efficient multi-domain data. We further incorporate chain-of-thought supervision and token-budget-aware reasoning, enabling the model to learn knowledge in a data-efficient manner while also performing reasoning efficiently. Our analysis shows that targeted home-scenario training not only enhances activity and gesture understanding, but also improves text-only reasoning with only modest trade-offs on document-centric tasks. Model checkpoints, quantized GGUF weights, and our home-scenario evaluation toolkit are publicly available at https://github.com/XiaoMi/xiaomi-mimo-vl-miloco to support research and deployment in real-world smart-home applications.
title Xiaomi MiMo-VL-Miloco Technical Report
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.17436