CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Gaoyang, Fu, Bingtao, Fan, Qingnan, Zhang, Qi, Liu, Runxing, Gu, Hong, Zhang, Huaqi, Liu, Xinguo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AuthFace: Towards Authentic Blind Face Restoration with Face-oriented Generative Diffusion Prior
by: Liang, Guoqiang, et al.
Published: (2024)
by: Liang, Guoqiang, et al.
Published: (2024)
BokehDiff: Neural Lens Blur with One-Step Diffusion
by: Zhu, Chengxuan, et al.
Published: (2025)
by: Zhu, Chengxuan, et al.
Published: (2025)
One-Shot Refiner: Boosting Feed-forward Novel View Synthesis via One-Step Diffusion
by: Dong, Yitong, et al.
Published: (2026)
by: Dong, Yitong, et al.
Published: (2026)
GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion Priors
by: Yin, Xingyilang, et al.
Published: (2025)
by: Yin, Xingyilang, et al.
Published: (2025)
LIPE: Learning Personalized Identity Prior for Non-rigid Image Editing
by: Liu, Aoyang, et al.
Published: (2024)
by: Liu, Aoyang, et al.
Published: (2024)
Agentar-Fin-OCR
by: Qian, Siyi, et al.
Published: (2026)
by: Qian, Siyi, et al.
Published: (2026)
RAP-SR: RestorAtion Prior Enhancement in Diffusion Models for Realistic Image Super-Resolution
by: Wang, Jiangang, et al.
Published: (2024)
by: Wang, Jiangang, et al.
Published: (2024)
FreeDiff: Progressive Frequency Truncation for Image Editing with Diffusion Models
by: Wu, Wei, et al.
Published: (2024)
by: Wu, Wei, et al.
Published: (2024)
MT-PCR: Hybrid Mamba-Transformer Network with Spatial Serialization for Point Cloud Registration
by: Liu, Bingxi, et al.
Published: (2025)
by: Liu, Bingxi, et al.
Published: (2025)
TinySR: Pruning Diffusion for Real-World Image Super-Resolution
by: Dong, Linwei, et al.
Published: (2025)
by: Dong, Linwei, et al.
Published: (2025)
Text-Aware Real-World Image Super-Resolution via Diffusion Model with Joint Segmentation Decoders
by: Hu, Qiming, et al.
Published: (2025)
by: Hu, Qiming, et al.
Published: (2025)
TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
by: Luan, Bozhi, et al.
Published: (2024)
by: Luan, Bozhi, et al.
Published: (2024)
Hero-SR: One-Step Diffusion for Super-Resolution with Human Perception Priors
by: Wang, Jiangang, et al.
Published: (2024)
by: Wang, Jiangang, et al.
Published: (2024)
TextInPlace: Indoor Visual Place Recognition in Repetitive Structures with Scene Text Spotting and Verification
by: Tao, Huaqi, et al.
Published: (2025)
by: Tao, Huaqi, et al.
Published: (2025)
Efficient RGB-D Scene Understanding via Multi-task Adaptive Learning and Cross-dimensional Feature Guidance
by: Sun, Guodong, et al.
Published: (2026)
by: Sun, Guodong, et al.
Published: (2026)
TSD-SR: One-Step Diffusion with Target Score Distillation for Real-World Image Super-Resolution
by: Dong, Linwei, et al.
Published: (2024)
by: Dong, Linwei, et al.
Published: (2024)
OMGSR: You Only Need One Mid-timestep Guidance for Real-World Image Super-Resolution
by: Wu, Zhiqiang, et al.
Published: (2025)
by: Wu, Zhiqiang, et al.
Published: (2025)
Textualize Visual Prompt for Image Editing via Diffusion Bridge
by: Xu, Pengcheng, et al.
Published: (2025)
by: Xu, Pengcheng, et al.
Published: (2025)
Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure Guidance
by: Luo, Minxing, et al.
Published: (2025)
by: Luo, Minxing, et al.
Published: (2025)
Real-Time Glottis Detection Framework via Spatial-decoupled Feature Learning for Nasal Transnasal Intubation
by: Liu, Jinyu, et al.
Published: (2026)
by: Liu, Jinyu, et al.
Published: (2026)
Enhancing Text-to-Image Diffusion Transformer via Split-Text Conditioning
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
by: Zhang, Yanzhe, et al.
Published: (2023)
by: Zhang, Yanzhe, et al.
Published: (2023)
Tuning-Free Adaptive Style Incorporation for Structure-Consistent Text-Driven Style Transfer
by: Ge, Yanqi, et al.
Published: (2024)
by: Ge, Yanqi, et al.
Published: (2024)
Boosting Text-to-Image Diffusion Models via Core Token Attention-Based Seed Selection
by: Zhang, Yunzhe, et al.
Published: (2026)
by: Zhang, Yunzhe, et al.
Published: (2026)
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
by: Jiang, Dongzhi, et al.
Published: (2024)
by: Jiang, Dongzhi, et al.
Published: (2024)
Ultra-High-Definition Reference-Based Landmark Image Super-Resolution with Generative Diffusion Prior
by: Shi, Zhenning, et al.
Published: (2025)
by: Shi, Zhenning, et al.
Published: (2025)
UnlearnCanvas: Stylized Image Dataset for Enhanced Machine Unlearning Evaluation in Diffusion Models
by: Zhang, Yihua, et al.
Published: (2024)
by: Zhang, Yihua, et al.
Published: (2024)
Hierarchical Visual Relocalization with Nearest View Synthesis from Feature Gaussian Splatting
by: Tao, Huaqi, et al.
Published: (2026)
by: Tao, Huaqi, et al.
Published: (2026)
From Text to Mask: Localizing Entities Using the Attention of Text-to-Image Diffusion Models
by: Xiao, Changming, et al.
Published: (2023)
by: Xiao, Changming, et al.
Published: (2023)
Decoding with Structured Awareness: Integrating Directional, Frequency-Spatial, and Structural Attention for Medical Image Segmentation
by: Zhang, Fan, et al.
Published: (2025)
by: Zhang, Fan, et al.
Published: (2025)
LiveMoments: Reselected Key Photo Restoration in Live Photos via Reference-guided Diffusion
by: Xue, Clara, et al.
Published: (2026)
by: Xue, Clara, et al.
Published: (2026)
Disciplined Diffusion: Text-to-Image Diffusion Model against NSFW Generation
by: Zhang, Chi, et al.
Published: (2026)
by: Zhang, Chi, et al.
Published: (2026)
Debiasing Text-to-Image Diffusion Models
by: He, Ruifei, et al.
Published: (2024)
by: He, Ruifei, et al.
Published: (2024)
Multi-Scale Diffusion: Enhancing Spatial Layout in High-Resolution Panoramic Image Generation
by: Zhang, Xiaoyu, et al.
Published: (2024)
by: Zhang, Xiaoyu, et al.
Published: (2024)
Image Super-Resolution with Text Prompt Diffusion
by: Chen, Zheng, et al.
Published: (2023)
by: Chen, Zheng, et al.
Published: (2023)
DiffRIS: Enhancing Referring Remote Sensing Image Segmentation with Pre-trained Text-to-Image Diffusion Models
by: Dong, Zhe, et al.
Published: (2025)
by: Dong, Zhe, et al.
Published: (2025)
CtxMIM: Context-Enhanced Masked Image Modeling for Remote Sensing Image Understanding
by: Zhang, Mingming, et al.
Published: (2023)
by: Zhang, Mingming, et al.
Published: (2023)
Learning from Mistakes: Iterative Prompt Relabeling for Text-to-Image Diffusion Model Training
by: Chen, Xinyan, et al.
Published: (2023)
by: Chen, Xinyan, et al.
Published: (2023)
Improving GFlowNets for Text-to-Image Diffusion Alignment
by: Zhang, Dinghuai, et al.
Published: (2024)
by: Zhang, Dinghuai, et al.
Published: (2024)
Noise Diffusion for Enhancing Semantic Faithfulness in Text-to-Image Synthesis
by: Miao, Boming, et al.
Published: (2024)
by: Miao, Boming, et al.
Published: (2024)
Similar Items
-
AuthFace: Towards Authentic Blind Face Restoration with Face-oriented Generative Diffusion Prior
by: Liang, Guoqiang, et al.
Published: (2024) -
BokehDiff: Neural Lens Blur with One-Step Diffusion
by: Zhu, Chengxuan, et al.
Published: (2025) -
One-Shot Refiner: Boosting Feed-forward Novel View Synthesis via One-Step Diffusion
by: Dong, Yitong, et al.
Published: (2026) -
GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion Priors
by: Yin, Xingyilang, et al.
Published: (2025) -
LIPE: Learning Personalized Identity Prior for Non-rigid Image Editing
by: Liu, Aoyang, et al.
Published: (2024)