Autoregressive Universal Video Segmentation Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Heo, Miran, Hwang, Sukjun, Chen, Min-Hung, Wang, Yu-Chiang Frank, Gu, Albert, Kim, Seon Joo, Hachiuma, Ryo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918130935136256
author Heo, Miran
Hwang, Sukjun
Chen, Min-Hung
Wang, Yu-Chiang Frank
Gu, Albert
Kim, Seon Joo
Hachiuma, Ryo
author_facet Heo, Miran
Hwang, Sukjun
Chen, Min-Hung
Wang, Yu-Chiang Frank
Gu, Albert
Kim, Seon Joo
Hachiuma, Ryo
contents Recent video foundation models such as SAM2 excel at prompted video segmentation by treating masks as a general-purpose primitive. However, many real-world settings require unprompted segmentation that aims to detect and track all objects in a video without external cues, leaving today's landscape fragmented across task-specific models and pipelines. We recast streaming video segmentation as sequential mask prediction, analogous to language modeling, and introduce the Autoregressive Universal Segmentation Model (AUSM), a single architecture that unifies both prompted and unprompted video segmentation. Built on recent state-space models, AUSM maintains a fixed-size spatial state and scales to video streams of arbitrary length. Furthermore, all components of AUSM are designed for parallel training across frames, yielding substantial speedups over iterative training. On standard benchmarks (DAVIS17, YouTube-VOS 2018 & 2019, MOSE, YouTube-VIS 2019 & 2021, and OVIS) AUSM outperforms prior universal streaming video segmentation methods and achieves up to 2.5x faster training on 16-frame sequences.
format Preprint
id arxiv_https___arxiv_org_abs_2508_19242
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Autoregressive Universal Video Segmentation Model
Heo, Miran
Hwang, Sukjun
Chen, Min-Hung
Wang, Yu-Chiang Frank
Gu, Albert
Kim, Seon Joo
Hachiuma, Ryo
Computer Vision and Pattern Recognition
Recent video foundation models such as SAM2 excel at prompted video segmentation by treating masks as a general-purpose primitive. However, many real-world settings require unprompted segmentation that aims to detect and track all objects in a video without external cues, leaving today's landscape fragmented across task-specific models and pipelines. We recast streaming video segmentation as sequential mask prediction, analogous to language modeling, and introduce the Autoregressive Universal Segmentation Model (AUSM), a single architecture that unifies both prompted and unprompted video segmentation. Built on recent state-space models, AUSM maintains a fixed-size spatial state and scales to video streams of arbitrary length. Furthermore, all components of AUSM are designed for parallel training across frames, yielding substantial speedups over iterative training. On standard benchmarks (DAVIS17, YouTube-VOS 2018 & 2019, MOSE, YouTube-VIS 2019 & 2021, and OVIS) AUSM outperforms prior universal streaming video segmentation methods and achieves up to 2.5x faster training on 16-frame sequences.
title Autoregressive Universal Video Segmentation Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.19242