Saved in:
| Main Authors: | Zhang, Jiexuan, Du, Yiheng, Wang, Qian, Li, Weiqi, Gu, Yu, Zhang, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2509.17088 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AlignGen: Boosting Personalized Image Generation with Cross-Modality Prior Alignment
by: Lin, Yiheng, et al.
Published: (2025)
by: Lin, Yiheng, et al.
Published: (2025)
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
by: Wu, Yi, et al.
Published: (2025)
by: Wu, Yi, et al.
Published: (2025)
ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned Representation
by: Huang, Qing, et al.
Published: (2026)
by: Huang, Qing, et al.
Published: (2026)
Style Aligned Image Generation via Shared Attention
by: Hertz, Amir, et al.
Published: (2023)
by: Hertz, Amir, et al.
Published: (2023)
Style-Aligned Image Composition for Robust Detection of Abnormal Cells in Cytopathology
by: Qi, Qiuyi, et al.
Published: (2025)
by: Qi, Qiuyi, et al.
Published: (2025)
OmniGen2: Towards Instruction-Aligned Multimodal Generation
by: Wu, Chenyuan, et al.
Published: (2025)
by: Wu, Chenyuan, et al.
Published: (2025)
AlignCVC: Aligning Cross-View Consistency for Single-Image-to-3D Generation
by: Liang, Xinyue, et al.
Published: (2025)
by: Liang, Xinyue, et al.
Published: (2025)
RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation
by: Zhang, Ruoxuan, et al.
Published: (2025)
by: Zhang, Ruoxuan, et al.
Published: (2025)
Re-Align: Structured Reasoning-guided Alignment for In-Context Image Generation and Editing
by: He, Runze, et al.
Published: (2026)
by: He, Runze, et al.
Published: (2026)
GenArena: How Can We Achieve Human-Aligned Evaluation for Visual Generation Tasks?
by: Li, Ruihang, et al.
Published: (2026)
by: Li, Ruihang, et al.
Published: (2026)
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
by: Wu, Yixuan, et al.
Published: (2025)
by: Wu, Yixuan, et al.
Published: (2025)
OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation
by: Li, Weiqi, et al.
Published: (2024)
by: Li, Weiqi, et al.
Published: (2024)
AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image Generation
by: Pang, Lianyu, et al.
Published: (2024)
by: Pang, Lianyu, et al.
Published: (2024)
ArtCrafter: Text-Image Aligning Style Transfer via Embedding Reframing
by: Huang, Nisha, et al.
Published: (2025)
by: Huang, Nisha, et al.
Published: (2025)
CAST: Component-Aligned 3D Scene Reconstruction from an RGB Image
by: Yao, Kaixin, et al.
Published: (2025)
by: Yao, Kaixin, et al.
Published: (2025)
360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion Model
by: Wang, Qian, et al.
Published: (2024)
by: Wang, Qian, et al.
Published: (2024)
AlignTok: Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models
by: Chen, Bowei, et al.
Published: (2025)
by: Chen, Bowei, et al.
Published: (2025)
Learning to Align Generative Appearance Priors for Fine-grained Image Retrieval
by: Wang, Shijie, et al.
Published: (2026)
by: Wang, Shijie, et al.
Published: (2026)
StyleTextGen: Style-Conditioned Multilingual Scene Text Generation
by: Chen, Zeyu, et al.
Published: (2026)
by: Chen, Zeyu, et al.
Published: (2026)
TAMISeg: Text-Aligned Multi-scale Medical Image Segmentation with Semantic Encoder Distillation
by: Gao, Qiang, et al.
Published: (2026)
by: Gao, Qiang, et al.
Published: (2026)
EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement
by: Xu, Zitong, et al.
Published: (2026)
by: Xu, Zitong, et al.
Published: (2026)
SeedEdit: Align Image Re-Generation to Image Editing
by: Shi, Yichun, et al.
Published: (2024)
by: Shi, Yichun, et al.
Published: (2024)
Asynchronous Denoising Diffusion Models for Aligning Text-to-Image Generation
by: Hu, Zijing, et al.
Published: (2025)
by: Hu, Zijing, et al.
Published: (2025)
Skill-Aligned Annotation for Reliable Evaluation in Text-to-Image Generation
by: Eldesokey, Abdelrahman, et al.
Published: (2026)
by: Eldesokey, Abdelrahman, et al.
Published: (2026)
UniTriGen: Unified Triplet Generation of Aligned Visible-Infrared-Label for Few-Shot RGB-T Semantic Segmentation
by: Zhou, Ping, et al.
Published: (2026)
by: Zhou, Ping, et al.
Published: (2026)
Style-NeRF2NeRF: 3D Style Transfer From Style-Aligned Multi-View Images
by: Fujiwara, Haruo, et al.
Published: (2024)
by: Fujiwara, Haruo, et al.
Published: (2024)
Driving-Video Dehazing with Non-Aligned Regularization for Safety Assistance
by: Fan, Junkai, et al.
Published: (2024)
by: Fan, Junkai, et al.
Published: (2024)
Pixal3D: Pixel-Aligned 3D Generation from Images
by: Li, Dong-Yang, et al.
Published: (2026)
by: Li, Dong-Yang, et al.
Published: (2026)
PhaSR: Generalized Image Shadow Removal with Physically Aligned Priors
by: Lee, Chia-Ming, et al.
Published: (2026)
by: Lee, Chia-Ming, et al.
Published: (2026)
SAQ-SAM: Semantically-Aligned Quantization for Segment Anything Model
by: Zhang, Jing, et al.
Published: (2025)
by: Zhang, Jing, et al.
Published: (2025)
Aligning Medical Images with General Knowledge from Large Language Models
by: Fang, Xiao, et al.
Published: (2024)
by: Fang, Xiao, et al.
Published: (2024)
AGHI-QA: A Subjective-Aligned Dataset and Metric for AI-Generated Human Images
by: Li, Yunhao, et al.
Published: (2025)
by: Li, Yunhao, et al.
Published: (2025)
Parallax to Align Them All: An OmniParallax Attention Mechanism for Distributed Multi-View Image Compression
by: Zhang, Haotian, et al.
Published: (2026)
by: Zhang, Haotian, et al.
Published: (2026)
Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation
by: Zhang, Wenchao, et al.
Published: (2025)
by: Zhang, Wenchao, et al.
Published: (2025)
Align-DETR: Enhancing End-to-end Object Detection with Aligned Loss
by: Cai, Zhi, et al.
Published: (2023)
by: Cai, Zhi, et al.
Published: (2023)
Fusion in Your Way: Aligning Image Fusion with Heterogeneous Demands via Direct Preference Optimization
by: Su, Weijian, et al.
Published: (2026)
by: Su, Weijian, et al.
Published: (2026)
Disentangle-then-Align: Non-Iterative Hybrid Multimodal Image Registration via Cross-Scale Feature Disentanglement
by: Zhang, Chunlei, et al.
Published: (2026)
by: Zhang, Chunlei, et al.
Published: (2026)
ChatGen: Automatic Text-to-Image Generation From FreeStyle Chatting
by: Jia, Chengyou, et al.
Published: (2024)
by: Jia, Chengyou, et al.
Published: (2024)
Video-Bench: Human-Aligned Video Generation Benchmark
by: Han, Hui, et al.
Published: (2025)
by: Han, Hui, et al.
Published: (2025)
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection
by: Wang, Shuyu, et al.
Published: (2025)
by: Wang, Shuyu, et al.
Published: (2025)
Similar Items
-
AlignGen: Boosting Personalized Image Generation with Cross-Modality Prior Alignment
by: Lin, Yiheng, et al.
Published: (2025) -
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
by: Wu, Yi, et al.
Published: (2025) -
ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned Representation
by: Huang, Qing, et al.
Published: (2026) -
Style Aligned Image Generation via Shared Attention
by: Hertz, Amir, et al.
Published: (2023) -
Style-Aligned Image Composition for Robust Detection of Abnormal Cells in Cytopathology
by: Qi, Qiuyi, et al.
Published: (2025)