SegLLM: Multi-round Reasoning Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, XuDong, Zhang, Shaolun, Li, Shufan, Kallidromitis, Konstantinos, Li, Kehan, Kato, Yusuke, Kozuka, Kazuki, Darrell, Trevor
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912099392815104
author Wang, XuDong
Zhang, Shaolun
Li, Shufan
Kallidromitis, Konstantinos
Li, Kehan
Kato, Yusuke
Kozuka, Kazuki
Darrell, Trevor
author_facet Wang, XuDong
Zhang, Shaolun
Li, Shufan
Kallidromitis, Konstantinos
Li, Kehan
Kato, Yusuke
Kozuka, Kazuki
Darrell, Trevor
contents We present SegLLM, a novel multi-round interactive reasoning segmentation model that enhances LLM-based segmentation by exploiting conversational memory of both visual and textual outputs. By leveraging a mask-aware multimodal LLM, SegLLM re-integrates previous segmentation results into its input stream, enabling it to reason about complex user intentions and segment objects in relation to previously identified entities, including positional, interactional, and hierarchical relationships, across multiple interactions. This capability allows SegLLM to respond to visual and text queries in a chat-like manner. Evaluated on the newly curated MRSeg benchmark, SegLLM outperforms existing methods in multi-round interactive reasoning segmentation by over 20%. Additionally, we observed that training on multi-round reasoning segmentation data enhances performance on standard single-round referring segmentation and localization tasks, resulting in a 5.5% increase in cIoU for referring expression segmentation and a 4.5% improvement in Acc@0.5 for referring expression localization.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18923
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SegLLM: Multi-round Reasoning Segmentation
Wang, XuDong
Zhang, Shaolun
Li, Shufan
Kallidromitis, Konstantinos
Li, Kehan
Kato, Yusuke
Kozuka, Kazuki
Darrell, Trevor
Computer Vision and Pattern Recognition
Artificial Intelligence
We present SegLLM, a novel multi-round interactive reasoning segmentation model that enhances LLM-based segmentation by exploiting conversational memory of both visual and textual outputs. By leveraging a mask-aware multimodal LLM, SegLLM re-integrates previous segmentation results into its input stream, enabling it to reason about complex user intentions and segment objects in relation to previously identified entities, including positional, interactional, and hierarchical relationships, across multiple interactions. This capability allows SegLLM to respond to visual and text queries in a chat-like manner. Evaluated on the newly curated MRSeg benchmark, SegLLM outperforms existing methods in multi-round interactive reasoning segmentation by over 20%. Additionally, we observed that training on multi-round reasoning segmentation data enhances performance on standard single-round referring segmentation and localization tasks, resulting in a 5.5% increase in cIoU for referring expression segmentation and a 4.5% improvement in Acc@0.5 for referring expression localization.
title SegLLM: Multi-round Reasoning Segmentation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2410.18923