Comp-Attn: Present-and-Align Attention for Compositional Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Hongyu, Deng, Yufan, Yuan, Shenghai, Zhao, Yian, Jin, Peng, Hou, Xuehan, Liu, Chang, Chen, Jie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912777984016384
author Zhang, Hongyu
Deng, Yufan
Yuan, Shenghai
Zhao, Yian
Jin, Peng
Hou, Xuehan
Liu, Chang
Chen, Jie
author_facet Zhang, Hongyu
Deng, Yufan
Yuan, Shenghai
Zhao, Yian
Jin, Peng
Hou, Xuehan
Liu, Chang
Chen, Jie
contents In the domain of text-to-video (T2V) generation, reliably synthesizing compositional content involving multiple subjects with intricate relations is still underexplored. The main challenges are twofold: 1) Subject presence, where not all subjects can be presented in the video; 2) Inter-subject relations, where the interaction and spatial relationship between subjects are misaligned. Existing methods adopt techniques, such as inference-time latent optimization or layout control, which fail to address both issues simultaneously. To tackle these problems, we propose Comp-Attn, a composition-aware cross-attention variant that follows a Present-and-Align paradigm: it decouples the two challenges by enforcing subject presence at the condition level and achieving relational alignment at the attention-distribution level. Specifically, 1) We introduce Subject-aware Condition Interpolation (SCI) to reinforce subject-specific conditions and ensure each subject's presence; 2) We propose Layout-forcing Attention Modulation (LAM), which dynamically enforces the attention distribution to align with the relational layout of multiple subjects. Comp-Attn can be seamlessly integrated into various T2V baselines in a training-free manner, boosting T2V-CompBench scores by 15.7\% and 11.7\% on Wan2.1-T2V-14B and Wan2.2-T2V-A14B with only a 5\% increase in inference time. Meanwhile, it also achieves strong performance on VBench and T2I-CompBench, demonstrating its scalability in general video generation and compositional text-to-image (T2I) tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_14428
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Comp-Attn: Present-and-Align Attention for Compositional Video Generation
Zhang, Hongyu
Deng, Yufan
Yuan, Shenghai
Zhao, Yian
Jin, Peng
Hou, Xuehan
Liu, Chang
Chen, Jie
Computer Vision and Pattern Recognition
Artificial Intelligence
In the domain of text-to-video (T2V) generation, reliably synthesizing compositional content involving multiple subjects with intricate relations is still underexplored. The main challenges are twofold: 1) Subject presence, where not all subjects can be presented in the video; 2) Inter-subject relations, where the interaction and spatial relationship between subjects are misaligned. Existing methods adopt techniques, such as inference-time latent optimization or layout control, which fail to address both issues simultaneously. To tackle these problems, we propose Comp-Attn, a composition-aware cross-attention variant that follows a Present-and-Align paradigm: it decouples the two challenges by enforcing subject presence at the condition level and achieving relational alignment at the attention-distribution level. Specifically, 1) We introduce Subject-aware Condition Interpolation (SCI) to reinforce subject-specific conditions and ensure each subject's presence; 2) We propose Layout-forcing Attention Modulation (LAM), which dynamically enforces the attention distribution to align with the relational layout of multiple subjects. Comp-Attn can be seamlessly integrated into various T2V baselines in a training-free manner, boosting T2V-CompBench scores by 15.7\% and 11.7\% on Wan2.1-T2V-14B and Wan2.2-T2V-A14B with only a 5\% increase in inference time. Meanwhile, it also achieves strong performance on VBench and T2I-CompBench, demonstrating its scalability in general video generation and compositional text-to-image (T2I) tasks.
title Comp-Attn: Present-and-Align Attention for Compositional Video Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2503.14428