LLM-guided Instance-level Image Manipulation with Diffusion U-Net Cross-Attention Maps
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Palaev, Andrey, Khan, Adil, Kazmi, Syed M. Ahsan |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
AD-Net: Attention-based dilated convolutional residual network with guided decoder for robust skin lesion segmentation
par: Naveed, Asim, et autres
Publié: (2024)
par: Naveed, Asim, et autres
Publié: (2024)
Attention-Augmented YOLOv8 with Ghost Convolution for Real-Time Vehicle Detection in Intelligent Transportation Systems
par: Ullah, Syed Sajid, et autres
Publié: (2026)
par: Ullah, Syed Sajid, et autres
Publié: (2026)
LSSF-Net: Lightweight Segmentation with Self-Awareness, Spatial Attention, and Focal Modulation
par: Farooq, Hamza, et autres
Publié: (2024)
par: Farooq, Hamza, et autres
Publié: (2024)
Efficient Video Object Segmentation via Modulated Cross-Attention Memory
par: Shaker, Abdelrahman, et autres
Publié: (2024)
par: Shaker, Abdelrahman, et autres
Publié: (2024)
InstanceAssemble: Layout-Aware Image Generation via Instance Assembling Attention
par: Xiang, Qiang, et autres
Publié: (2025)
par: Xiang, Qiang, et autres
Publié: (2025)
InstanceGen: Image Generation with Instance-level Instructions
par: Sella, Etai, et autres
Publié: (2025)
par: Sella, Etai, et autres
Publié: (2025)
Geometry-guided Cross-view Diffusion for One-to-many Cross-view Image Synthesis
par: Lin, Tao Jun, et autres
Publié: (2024)
par: Lin, Tao Jun, et autres
Publié: (2024)
InstanceDiffusion: Instance-level Control for Image Generation
par: Wang, Xudong, et autres
Publié: (2024)
par: Wang, Xudong, et autres
Publié: (2024)
Dynamic Importance in Diffusion U-Net for Enhanced Image Synthesis
par: Wang, Xi, et autres
Publié: (2025)
par: Wang, Xi, et autres
Publié: (2025)
BookNet: Book Image Rectification via Cross-Page Attention Network
par: Liu, Shaokai, et autres
Publié: (2026)
par: Liu, Shaokai, et autres
Publié: (2026)
A Unified Attention U-Net Framework for Cross-Modality Tumor Segmentation in MRI and CT
par: Rai, Nishan, et autres
Publié: (2026)
par: Rai, Nishan, et autres
Publié: (2026)
Transformer-Driven Active Transfer Learning for Cross-Hyperspectral Image Classification
par: Ahmad, Muhammad, et autres
Publié: (2024)
par: Ahmad, Muhammad, et autres
Publié: (2024)
Layout-to-Image Generation with Localized Descriptions using ControlNet with Cross-Attention Control
par: Lukovnikov, Denis, et autres
Publié: (2024)
par: Lukovnikov, Denis, et autres
Publié: (2024)
AttentionBender: Manipulating Cross-Attention in Video Diffusion Transformers as a Creative Probe
par: Cole, Adam, et autres
Publié: (2026)
par: Cole, Adam, et autres
Publié: (2026)
Handwritten Digit Recognition: An Ensemble-Based Approach for Superior Performance
par: Ullah, Syed Sajid, et autres
Publié: (2025)
par: Ullah, Syed Sajid, et autres
Publié: (2025)
LMBF-Net: A Lightweight Multipath Bidirectional Focal Attention Network for Multifeatures Segmentation
par: Khan, Tariq M, et autres
Publié: (2024)
par: Khan, Tariq M, et autres
Publié: (2024)
Soft Knowledge Distillation with Multi-Dimensional Cross-Net Attention for Image Restoration Models Compression
par: Zhang, Yongheng, et autres
Publié: (2025)
par: Zhang, Yongheng, et autres
Publié: (2025)
DCAU-Net: Differential Cross Attention and Channel-Spatial Feature Fusion for Medical Image Segmentation
par: Li, Yanxin, et autres
Publié: (2026)
par: Li, Yanxin, et autres
Publié: (2026)
Rethinking Attention-Based Multiple Instance Learning for Whole-Slide Pathological Image Classification: An Instance Attribute Viewpoint
par: Cai, Linghan, et autres
Publié: (2024)
par: Cai, Linghan, et autres
Publié: (2024)
Attention-Challenging Multiple Instance Learning for Whole Slide Image Classification
par: Zhang, Yunlong, et autres
Publié: (2023)
par: Zhang, Yunlong, et autres
Publié: (2023)
ASMIL: Attention-Stabilized Multiple Instance Learning for Whole Slide Imaging
par: Ye, Linfeng, et autres
Publié: (2026)
par: Ye, Linfeng, et autres
Publié: (2026)
Single-Reference Text-to-Image Manipulation with Dual Contrastive Denoising Score
par: Israr, Syed Muhmmad, et autres
Publié: (2025)
par: Israr, Syed Muhmmad, et autres
Publié: (2025)
Latents of latents to delineate pixels: hybrid Matryoshka autoencoder-to-U-Net pairing for segmenting large medical images in GPU-poor and low-data regimes
par: Syed, Tahir, et autres
Publié: (2025)
par: Syed, Tahir, et autres
Publié: (2025)
Soft-Hard Attention U-Net Model and Benchmark Dataset for Multiscale Image Shadow Removal
par: Cholopoulou, Eirini, et autres
Publié: (2024)
par: Cholopoulou, Eirini, et autres
Publié: (2024)
Speedrunning ImageNet Diffusion
par: Bhanded, Swayam
Publié: (2025)
par: Bhanded, Swayam
Publié: (2025)
Faithful Counterfactual Visual Explanations (FCVE)
par: Khan, Bismillah, et autres
Publié: (2025)
par: Khan, Bismillah, et autres
Publié: (2025)
Garment Attribute Manipulation with Multi-level Attention
par: Casula, Vittorio, et autres
Publié: (2024)
par: Casula, Vittorio, et autres
Publié: (2024)
Prompt Augmentation for Self-supervised Text-guided Image Manipulation
par: Bodur, Rumeysa, et autres
Publié: (2024)
par: Bodur, Rumeysa, et autres
Publié: (2024)
VMix: Improving Text-to-Image Diffusion Model with Cross-Attention Mixing Control
par: Wu, Shaojin, et autres
Publié: (2024)
par: Wu, Shaojin, et autres
Publié: (2024)
Object-Conditioned Energy-Based Attention Map Alignment in Text-to-Image Diffusion Models
par: Zhang, Yasi, et autres
Publié: (2024)
par: Zhang, Yasi, et autres
Publié: (2024)
CalibNet: Dual-branch Cross-modal Calibration for RGB-D Salient Instance Segmentation
par: Pei, Jialun, et autres
Publié: (2023)
par: Pei, Jialun, et autres
Publié: (2023)
AttMetNet: Attention-Enhanced Deep Neural Network for Methane Plume Detection in Sentinel-2 Satellite Imagery
par: Ahsan, Rakib, et autres
Publié: (2025)
par: Ahsan, Rakib, et autres
Publié: (2025)
nnY-Net: Swin-NeXt with Cross-Attention for 3D Medical Images Segmentation
par: Liu, Haixu, et autres
Publié: (2025)
par: Liu, Haixu, et autres
Publié: (2025)
U-REPA: Aligning Diffusion U-Nets to ViTs
par: Tian, Yuchuan, et autres
Publié: (2025)
par: Tian, Yuchuan, et autres
Publié: (2025)
SalFAU-Net: Saliency Fusion Attention U-Net for Salient Object Detection
par: Mulat, Kassaw Abraham, et autres
Publié: (2024)
par: Mulat, Kassaw Abraham, et autres
Publié: (2024)
Joint Image-Instance Spatial-Temporal Attention for Few-shot Action Recognition
par: Qian, Zefeng, et autres
Publié: (2025)
par: Qian, Zefeng, et autres
Publié: (2025)
Region Guided Attention Network for Retinal Vessel Segmentation
par: Javed, Syed, et autres
Publié: (2024)
par: Javed, Syed, et autres
Publié: (2024)
Attention-based U-Net Method for Autonomous Lane Detection
par: Tangestanizadeh, Mohammadhamed, et autres
Publié: (2024)
par: Tangestanizadeh, Mohammadhamed, et autres
Publié: (2024)
Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image Editing
par: Liu, Bingyan, et autres
Publié: (2024)
par: Liu, Bingyan, et autres
Publié: (2024)
Hierarchical and Step-Layer-Wise Tuning of Attention Specialty for Multi-Instance Synthesis in Diffusion Transformers
par: Zhang, Chunyang, et autres
Publié: (2025)
par: Zhang, Chunyang, et autres
Publié: (2025)
Documents similaires
-
AD-Net: Attention-based dilated convolutional residual network with guided decoder for robust skin lesion segmentation
par: Naveed, Asim, et autres
Publié: (2024) -
Attention-Augmented YOLOv8 with Ghost Convolution for Real-Time Vehicle Detection in Intelligent Transportation Systems
par: Ullah, Syed Sajid, et autres
Publié: (2026) -
LSSF-Net: Lightweight Segmentation with Self-Awareness, Spatial Attention, and Focal Modulation
par: Farooq, Hamza, et autres
Publié: (2024) -
Efficient Video Object Segmentation via Modulated Cross-Attention Memory
par: Shaker, Abdelrahman, et autres
Publié: (2024) -
InstanceAssemble: Layout-Aware Image Generation via Instance Assembling Attention
par: Xiang, Qiang, et autres
Publié: (2025)