LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yuan, Yuqian, Zhang, Wenqiao, Lin, Juekai, Zhong, Yu, Gao, Mingjian, Yu, Binhe, Cao, Yunqi, Li, Wentong, Zhuang, Yueting, Ooi, Beng Chin
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913046995140608
author Yuan, Yuqian
Zhang, Wenqiao
Lin, Juekai
Zhong, Yu
Gao, Mingjian
Yu, Binhe
Cao, Yunqi
Li, Wentong
Zhuang, Yueting
Ooi, Beng Chin
author_facet Yuan, Yuqian
Zhang, Wenqiao
Lin, Juekai
Zhong, Yu
Gao, Mingjian
Yu, Binhe
Cao, Yunqi
Li, Wentong
Zhuang, Yueting
Ooi, Beng Chin
contents Large Multimodal Models (LMMs) have achieved remarkable progress in general-purpose vision--language understanding, yet they remain limited in tasks requiring precise object-level grounding, fine-grained spatial reasoning, and controllable visual manipulation. In particular, existing systems often struggle to identify the correct instance, preserve object identity across interactions, and localize or modify designated regions with high precision. Object-centric vision provides a principled framework for addressing these challenges by promoting explicit representations and operations over visual entities, thereby extending multimodal systems from global scene understanding to object-level understanding, segmentation, editing, and generation. This paper presents a comprehensive review of recent advances at the convergence of LMMs and object-centric vision. We organize the literature into four major themes: object-centric visual understanding, object-centric referring segmentation, object-centric visual editing, and object-centric visual generation. We further summarize the key modeling paradigms, learning strategies, and evaluation protocols that support these capabilities. Finally, we discuss open challenges and future directions, including robust instance permanence, fine-grained spatial control, consistent multi-step interaction, unified cross-task modeling, and reliable benchmarking under distribution shift. We hope this paper provides a structured perspective on the development of scalable, precise, and trustworthy object-centric multimodal systems.
format Preprint
id arxiv_https___arxiv_org_abs_2604_11789
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation
Yuan, Yuqian
Zhang, Wenqiao
Lin, Juekai
Zhong, Yu
Gao, Mingjian
Yu, Binhe
Cao, Yunqi
Li, Wentong
Zhuang, Yueting
Ooi, Beng Chin
Computer Vision and Pattern Recognition
Large Multimodal Models (LMMs) have achieved remarkable progress in general-purpose vision--language understanding, yet they remain limited in tasks requiring precise object-level grounding, fine-grained spatial reasoning, and controllable visual manipulation. In particular, existing systems often struggle to identify the correct instance, preserve object identity across interactions, and localize or modify designated regions with high precision. Object-centric vision provides a principled framework for addressing these challenges by promoting explicit representations and operations over visual entities, thereby extending multimodal systems from global scene understanding to object-level understanding, segmentation, editing, and generation. This paper presents a comprehensive review of recent advances at the convergence of LMMs and object-centric vision. We organize the literature into four major themes: object-centric visual understanding, object-centric referring segmentation, object-centric visual editing, and object-centric visual generation. We further summarize the key modeling paradigms, learning strategies, and evaluation protocols that support these capabilities. Finally, we discuss open challenges and future directions, including robust instance permanence, fine-grained spatial control, consistent multi-step interaction, unified cross-task modeling, and reliable benchmarking under distribution shift. We hope this paper provides a structured perspective on the development of scalable, precise, and trustworthy object-centric multimodal systems.
title LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.11789