Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Su Ho, Hyun, Jeongseok, Lee, Pilhyeon, Shim, Minho, Wee, Dongyoon, Kim, Seon Joo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917431519215616
author Han, Su Ho
Hyun, Jeongseok
Lee, Pilhyeon
Shim, Minho
Wee, Dongyoon
Kim, Seon Joo
author_facet Han, Su Ho
Hyun, Jeongseok
Lee, Pilhyeon
Shim, Minho
Wee, Dongyoon
Kim, Seon Joo
contents Multimodal large language models (MLLMs) demonstrate strong video understanding by attending to visual tokens relevant to textual queries. To directly adapt this for localization in a training-free manner, we cast video reasoning segmentation as a video QA task and extract attention maps via rollout mechanism. However, raw attention maps are noisy and poorly aligned with object regions. We propose Decomposed Attention Fusion (DecAF), which refines these maps through two mechanisms: (1) contrastive object-background fusion and (2) complementary video-frame fusion. This method suppresses irrelevant activations and enhances object-focused cues, enabling direct conversion of attention maps into coarse segmentation masks. In addition, we introduce attention-guided SAM2 prompting for obtaining fine-grained masks. Unlike existing methods that jointly train MLLMs with SAM, our method operates entirely without retraining. DecAF outperforms training-free methods and achieves performance comparable to training-based methods on both referring and reasoning VOS benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2510_19592
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
Han, Su Ho
Hyun, Jeongseok
Lee, Pilhyeon
Shim, Minho
Wee, Dongyoon
Kim, Seon Joo
Computer Vision and Pattern Recognition
Multimodal large language models (MLLMs) demonstrate strong video understanding by attending to visual tokens relevant to textual queries. To directly adapt this for localization in a training-free manner, we cast video reasoning segmentation as a video QA task and extract attention maps via rollout mechanism. However, raw attention maps are noisy and poorly aligned with object regions. We propose Decomposed Attention Fusion (DecAF), which refines these maps through two mechanisms: (1) contrastive object-background fusion and (2) complementary video-frame fusion. This method suppresses irrelevant activations and enhances object-focused cues, enabling direct conversion of attention maps into coarse segmentation masks. In addition, we introduce attention-guided SAM2 prompting for obtaining fine-grained masks. Unlike existing methods that jointly train MLLMs with SAM, our method operates entirely without retraining. DecAF outperforms training-free methods and achieves performance comparable to training-based methods on both referring and reasoning VOS benchmarks.
title Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.19592