Comp-Attn: Present-and-Align Attention for Compositional Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Hongyu, Deng, Yufan, Yuan, Shenghai, Zhao, Yian, Jin, Peng, Hou, Xuehan, Liu, Chang, Chen, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation
by: Zhang, Hongyu, et al.
Published: (2026)
by: Zhang, Hongyu, et al.
Published: (2026)
Focal Guidance: Unlocking Controllability from Semantic-Weak Layers in Video Diffusion Models
by: Yin, Yuanyang, et al.
Published: (2026)
by: Yin, Yuanyang, et al.
Published: (2026)
Rectified SpaAttn: Revisiting Attention Sparsity for Efficient Video Generation
by: Liu, Xuewen, et al.
Published: (2025)
by: Liu, Xuewen, et al.
Published: (2025)
AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image Generation
by: Pang, Lianyu, et al.
Published: (2024)
by: Pang, Lianyu, et al.
Published: (2024)
DiTFastAttn: Attention Compression for Diffusion Transformer Models
by: Yuan, Zhihang, et al.
Published: (2024)
by: Yuan, Zhihang, et al.
Published: (2024)
Video Generation with Predictive Latents
by: Zhao, Yian, et al.
Published: (2026)
by: Zhao, Yian, et al.
Published: (2026)
OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
by: Yuan, Shenghai, et al.
Published: (2025)
by: Yuan, Shenghai, et al.
Published: (2025)
D-Attn: Decomposed Attention for Large Vision-and-Language Models
by: Kuo, Chia-Wen, et al.
Published: (2025)
by: Kuo, Chia-Wen, et al.
Published: (2025)
VideoTetris: Towards Compositional Text-to-Video Generation
by: Tian, Ye, et al.
Published: (2024)
by: Tian, Ye, et al.
Published: (2024)
ColFigPhotoAttnNet: Reliable Finger Photo Presentation Attack Detection Leveraging Window-Attention on Color Spaces
by: Vurity, Anudeep, et al.
Published: (2025)
by: Vurity, Anudeep, et al.
Published: (2025)
MaskAttn-SDXL: Controllable Region-Level Text-To-Image Generation
by: Chang, Yu, et al.
Published: (2025)
by: Chang, Yu, et al.
Published: (2025)
CompAlign: Improving Compositional Text-to-Image Generation with a Complex Benchmark and Fine-Grained Feedback
by: Wan, Yixin, et al.
Published: (2025)
by: Wan, Yixin, et al.
Published: (2025)
AttnMod: Attention-Based New Art Styles
by: Su, Shih-Chieh
Published: (2024)
by: Su, Shih-Chieh
Published: (2024)
CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance
by: Deng, Yufan, et al.
Published: (2025)
by: Deng, Yufan, et al.
Published: (2025)
Helios: Real Real-Time Long Video Generation Model
by: Yuan, Shenghai, et al.
Published: (2026)
by: Yuan, Shenghai, et al.
Published: (2026)
GradAttn: Replacing Fixed Residual Connections with Task-Modulated Attention Pathways
by: Ghoshal, Soudeep, et al.
Published: (2026)
by: Ghoshal, Soudeep, et al.
Published: (2026)
CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation
by: Zhu, Zixin, et al.
Published: (2025)
by: Zhu, Zixin, et al.
Published: (2025)
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
PresentAgent: Multimodal Agent for Presentation Video Generation
by: Shi, Jingwei, et al.
Published: (2025)
by: Shi, Jingwei, et al.
Published: (2025)
I2V-Adapter: A General Image-to-Video Adapter for Diffusion Models
by: Guo, Xun, et al.
Published: (2023)
by: Guo, Xun, et al.
Published: (2023)
EF-VI: Enhancing End-Frame Injection for Video Inbetweening
by: Chen, Liuhan, et al.
Published: (2025)
by: Chen, Liuhan, et al.
Published: (2025)
MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
by: Deng, Yufan, et al.
Published: (2025)
by: Deng, Yufan, et al.
Published: (2025)
AttnRouter: Per-Category Attention Routing for Training-Free Image Editing on MMDiT
by: Li, Guandong, et al.
Published: (2026)
by: Li, Guandong, et al.
Published: (2026)
UniComp: Rethinking Video Compression Through Informational Uniqueness
by: Yuan, Chao, et al.
Published: (2025)
by: Yuan, Chao, et al.
Published: (2025)
DiTFastAttnV2: Head-wise Attention Compression for Multi-Modality Diffusion Transformers
by: Zhang, Hanling, et al.
Published: (2025)
by: Zhang, Hanling, et al.
Published: (2025)
Identity-Preserving Text-to-Video Generation by Frequency Decomposition
by: Yuan, Shenghai, et al.
Published: (2024)
by: Yuan, Shenghai, et al.
Published: (2024)
HumanNet: Scaling Human-centric Video Learning to One Million Hours
by: Deng, Yufan, et al.
Published: (2026)
by: Deng, Yufan, et al.
Published: (2026)
DFSAttn: Dynamic Fine-grained Sparse Attention for Efficient Video Generation
by: Hu, Jie, et al.
Published: (2026)
by: Hu, Jie, et al.
Published: (2026)
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
by: Liu, Yexin, et al.
Published: (2025)
by: Liu, Yexin, et al.
Published: (2025)
FlashI2V: Fourier-Guided Latent Shifting Prevents Conditional Image Leakage in Image-to-Video Generation
by: Ge, Yunyang, et al.
Published: (2025)
by: Ge, Yunyang, et al.
Published: (2025)
Video-Bench: Human-Aligned Video Generation Benchmark
by: Han, Hui, et al.
Published: (2025)
by: Han, Hui, et al.
Published: (2025)
Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model
by: Bui, Phuoc-Nguyen, et al.
Published: (2025)
by: Bui, Phuoc-Nguyen, et al.
Published: (2025)
DATransNet: Dynamic Attention Transformer Network for Infrared Small Target Detection
by: Hu, Chen, et al.
Published: (2024)
by: Hu, Chen, et al.
Published: (2024)
iSegMan: Interactive Segment-and-Manipulate 3D Gaussians
by: Zhao, Yian, et al.
Published: (2025)
by: Zhao, Yian, et al.
Published: (2025)
MANGO:Natural Multi-speaker 3D Talking Head Generation via 2D-Lifted Enhancement
by: Zhu, Lei, et al.
Published: (2026)
by: Zhu, Lei, et al.
Published: (2026)
Generative Model-Based Feature Attention Module for Video Action Analysis
by: Wang, Guiqin, et al.
Published: (2025)
by: Wang, Guiqin, et al.
Published: (2025)
MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators
by: Yuan, Shenghai, et al.
Published: (2024)
by: Yuan, Shenghai, et al.
Published: (2024)
Matten: Video Generation with Mamba-Attention
by: Gao, Yu, et al.
Published: (2024)
by: Gao, Yu, et al.
Published: (2024)
Rethinking Video Generation Model for the Embodied World
by: Deng, Yufan, et al.
Published: (2026)
by: Deng, Yufan, et al.
Published: (2026)
Comp4D: LLM-Guided Compositional 4D Scene Generation
by: Xu, Dejia, et al.
Published: (2024)
by: Xu, Dejia, et al.
Published: (2024)
Similar Items
-
TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation
by: Zhang, Hongyu, et al.
Published: (2026) -
Focal Guidance: Unlocking Controllability from Semantic-Weak Layers in Video Diffusion Models
by: Yin, Yuanyang, et al.
Published: (2026) -
Rectified SpaAttn: Revisiting Attention Sparsity for Efficient Video Generation
by: Liu, Xuewen, et al.
Published: (2025) -
AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image Generation
by: Pang, Lianyu, et al.
Published: (2024) -
DiTFastAttn: Attention Compression for Diffusion Transformer Models
by: Yuan, Zhihang, et al.
Published: (2024)