Rethinking Multi-Condition DiTs: Eliminating Redundant Attention via Position-Alignment and Keyword-Scoping
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Chao, Wei, Tianyi, Chen, Yiling, Zhou, Wenbo, Yu, Nenghai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SD-DiT: Unleashing the Power of Self-supervised Discrimination in Diffusion Transformer
von: Zhu, Rui, et al.
Veröffentlicht: (2024)
von: Zhu, Rui, et al.
Veröffentlicht: (2024)
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
von: Wang, Xiang, et al.
Veröffentlicht: (2025)
von: Wang, Xiang, et al.
Veröffentlicht: (2025)
Unveiling Redundancy in Diffusion Transformers (DiTs): A Systematic Study
von: Sun, Xibo, et al.
Veröffentlicht: (2024)
von: Sun, Xibo, et al.
Veröffentlicht: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
Rethinking Multi-view Representation Learning via Distilled Disentangling
von: Ke, Guanzhou, et al.
Veröffentlicht: (2024)
von: Ke, Guanzhou, et al.
Veröffentlicht: (2024)
Untwisting RoPE: Frequency Control for Shared Attention in DiTs
von: Mikaeili, Aryan, et al.
Veröffentlicht: (2026)
von: Mikaeili, Aryan, et al.
Veröffentlicht: (2026)
Cross Modal Fine-Grained Alignment via Granularity-Aware and Region-Uncertain Modeling
von: Liu, Jiale, et al.
Veröffentlicht: (2025)
von: Liu, Jiale, et al.
Veröffentlicht: (2025)
MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and Retention
von: Wang, Tianyi, et al.
Veröffentlicht: (2025)
von: Wang, Tianyi, et al.
Veröffentlicht: (2025)
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
von: Cai, Lingling, et al.
Veröffentlicht: (2024)
von: Cai, Lingling, et al.
Veröffentlicht: (2024)
Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
von: Zhou, Chao, et al.
Veröffentlicht: (2025)
von: Zhou, Chao, et al.
Veröffentlicht: (2025)
DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation
von: Cai, Minghong, et al.
Veröffentlicht: (2024)
von: Cai, Minghong, et al.
Veröffentlicht: (2024)
Clean Image May be Dangerous: Data Poisoning Attacks Against Deep Hashing
von: Li, Shuai, et al.
Veröffentlicht: (2025)
von: Li, Shuai, et al.
Veröffentlicht: (2025)
U-DiTs: Downsample Tokens in U-Shaped Diffusion Transformers
von: Tian, Yuchuan, et al.
Veröffentlicht: (2024)
von: Tian, Yuchuan, et al.
Veröffentlicht: (2024)
SpecFLASH: A Latent-Guided Semi-autoregressive Speculative Decoding Framework for Efficient Multimodal Generation
von: Wang, Zihua, et al.
Veröffentlicht: (2025)
von: Wang, Zihua, et al.
Veröffentlicht: (2025)
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
CLIP-PCQA: Exploring Subjective-Aligned Vision-Language Modeling for Point Cloud Quality Assessment
von: Liu, Yating, et al.
Veröffentlicht: (2025)
von: Liu, Yating, et al.
Veröffentlicht: (2025)
Riemann-based Multi-scale Attention Reasoning Network for Text-3D Retrieval
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
Multi-scale Attention Guided Pose Transfer
von: Roy, Prasun, et al.
Veröffentlicht: (2022)
von: Roy, Prasun, et al.
Veröffentlicht: (2022)
Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning
von: Song, Zijie, et al.
Veröffentlicht: (2023)
von: Song, Zijie, et al.
Veröffentlicht: (2023)
Art2Mus: Artwork-to-Music Generation via Visual Conditioning and Large-Scale Cross-Modal Alignment
von: Rinaldi, Ivan, et al.
Veröffentlicht: (2026)
von: Rinaldi, Ivan, et al.
Veröffentlicht: (2026)
Contrastive Pre-Training with Multi-View Fusion for No-Reference Point Cloud Quality Assessment
von: Shan, Ziyu, et al.
Veröffentlicht: (2024)
von: Shan, Ziyu, et al.
Veröffentlicht: (2024)
Regularizing Subspace Redundancy of Low-Rank Adaptation
von: Zhu, Yue, et al.
Veröffentlicht: (2025)
von: Zhu, Yue, et al.
Veröffentlicht: (2025)
DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization
von: Nguyen, Ngoc-Son, et al.
Veröffentlicht: (2026)
von: Nguyen, Ngoc-Son, et al.
Veröffentlicht: (2026)
Visual Grounding with Multi-modal Conditional Adaptation
von: Yao, Ruilin, et al.
Veröffentlicht: (2024)
von: Yao, Ruilin, et al.
Veröffentlicht: (2024)
Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
von: Song, Shezheng, et al.
Veröffentlicht: (2026)
von: Song, Shezheng, et al.
Veröffentlicht: (2026)
Noisy-Correspondence Learning for Text-to-Image Person Re-identification
von: Qin, Yang, et al.
Veröffentlicht: (2023)
von: Qin, Yang, et al.
Veröffentlicht: (2023)
A Hierarchical Compression Technique for 3D Gaussian Splatting Compression
von: Huang, He, et al.
Veröffentlicht: (2024)
von: Huang, He, et al.
Veröffentlicht: (2024)
Inclusion 2024 Global Multimedia Deepfake Detection Challenge: Towards Multi-dimensional Face Forgery Detection
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
DepthGait: Multi-Scale Cross-Level Feature Fusion of RGB-Derived Depth and Silhouette Sequences for Robust Gait Recognition
von: Li, Xinzhu, et al.
Veröffentlicht: (2025)
von: Li, Xinzhu, et al.
Veröffentlicht: (2025)
ByteNet: Rethinking Multimedia File Fragment Classification through Visual Perspectives
von: Liu, Wenyang, et al.
Veröffentlicht: (2024)
von: Liu, Wenyang, et al.
Veröffentlicht: (2024)
Signal: Selective Interaction and Global-local Alignment for Multi-Modal Object Re-Identification
von: Liu, Yangyang, et al.
Veröffentlicht: (2025)
von: Liu, Yangyang, et al.
Veröffentlicht: (2025)
PixCLIP: Achieving Fine-grained Visual Language Understanding via Any-granularity Pixel-Text Alignment Learning
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query
von: Chow, Wei, et al.
Veröffentlicht: (2025)
von: Chow, Wei, et al.
Veröffentlicht: (2025)
Textured mesh Quality Assessment using Geometry and Color Field Similarity
von: Yang, Kaifa, et al.
Veröffentlicht: (2025)
von: Yang, Kaifa, et al.
Veröffentlicht: (2025)
PAME: Self-Supervised Masked Autoencoder for No-Reference Point Cloud Quality Assessment
von: Shan, Ziyu, et al.
Veröffentlicht: (2024)
von: Shan, Ziyu, et al.
Veröffentlicht: (2024)
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces
von: E, Shaojun, et al.
Veröffentlicht: (2025)
von: E, Shaojun, et al.
Veröffentlicht: (2025)
Transcending Fusion: A Multi-Scale Alignment Method for Remote Sensing Image-Text Retrieval
von: Yang, Rui, et al.
Veröffentlicht: (2024)
von: Yang, Rui, et al.
Veröffentlicht: (2024)
Wills Aligner: Multi-Subject Collaborative Brain Visual Decoding
von: Bao, Guangyin, et al.
Veröffentlicht: (2024)
von: Bao, Guangyin, et al.
Veröffentlicht: (2024)
ContextBLIP: Doubly Contextual Alignment for Contrastive Image Retrieval from Linguistically Complex Descriptions
von: Lin, Honglin, et al.
Veröffentlicht: (2024)
von: Lin, Honglin, et al.
Veröffentlicht: (2024)
POINTS1.5: Building a Vision-Language Model towards Real World Applications
von: Liu, Yuan, et al.
Veröffentlicht: (2024)
von: Liu, Yuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SD-DiT: Unleashing the Power of Self-supervised Discrimination in Diffusion Transformer
von: Zhu, Rui, et al.
Veröffentlicht: (2024) -
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
von: Wang, Xiang, et al.
Veröffentlicht: (2025) -
Unveiling Redundancy in Diffusion Transformers (DiTs): A Systematic Study
von: Sun, Xibo, et al.
Veröffentlicht: (2024) -
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024) -
Rethinking Multi-view Representation Learning via Distilled Disentangling
von: Ke, Guanzhou, et al.
Veröffentlicht: (2024)