ID-Aligner: Enhancing Identity-Preserving Text-to-Image Generation with Reward Feedback Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Weifeng, Zhang, Jiacheng, Wu, Jie, Wu, Hefeng, Xiao, Xuefeng, Lin, Liang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
by: Chen, Weifeng, et al.
Published: (2023)
by: Chen, Weifeng, et al.
Published: (2023)
ConsistentID: Portrait Generation with Multimodal Fine-Grained Identity Preserving
by: Huang, Jiehui, et al.
Published: (2024)
by: Huang, Jiehui, et al.
Published: (2024)
DiffusionAgent: Navigating Expert Models for Agentic Image Generation
by: Qin, Jie, et al.
Published: (2024)
by: Qin, Jie, et al.
Published: (2024)
IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model
by: Ji, Yatai, et al.
Published: (2024)
by: Ji, Yatai, et al.
Published: (2024)
InstantID: Zero-shot Identity-Preserving Generation in Seconds
by: Wang, Qixun, et al.
Published: (2024)
by: Wang, Qixun, et al.
Published: (2024)
DiffusionReward: Enhancing Blind Face Restoration through Reward Feedback Learning
by: Wu, Bin, et al.
Published: (2025)
by: Wu, Bin, et al.
Published: (2025)
ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video Generation
by: Wu, Mingyang, et al.
Published: (2026)
by: Wu, Mingyang, et al.
Published: (2026)
Learning Joint ID-Textual Representation for ID-Preserving Image Synthesis
by: Liu, Zichuan, et al.
Published: (2025)
by: Liu, Zichuan, et al.
Published: (2025)
FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation
by: Wang, Yuanzhi, et al.
Published: (2026)
by: Wang, Yuanzhi, et al.
Published: (2026)
IFAdapter: Instance Feature Control for Grounded Text-to-Image Generation
by: Wu, Yinwei, et al.
Published: (2024)
by: Wu, Yinwei, et al.
Published: (2024)
Concat-ID: Towards Universal Identity-Preserving Video Synthesis
by: Zhong, Yong, et al.
Published: (2025)
by: Zhong, Yong, et al.
Published: (2025)
Slot-ID: Identity-Preserving Video Generation from Reference Videos via Slot-Based Temporal Identity Encoding
by: Lai, Yixuan, et al.
Published: (2026)
by: Lai, Yixuan, et al.
Published: (2026)
OnlineVPO: Align Video Diffusion Model with Online Video-Centric Preference Optimization
by: Zhang, Jiacheng, et al.
Published: (2024)
by: Zhang, Jiacheng, et al.
Published: (2024)
RoBridge: A Hierarchical Architecture Bridging Cognition and Execution for General Robotic Manipulation
by: Zhang, Kaidong, et al.
Published: (2025)
by: Zhang, Kaidong, et al.
Published: (2025)
Personalized Reward Modeling for Text-to-Image Generation
by: Lee, Jeongeun, et al.
Published: (2025)
by: Lee, Jeongeun, et al.
Published: (2025)
ControlNet++: Improving Conditional Controls with Efficient Consistency Feedback
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
Anonymization Prompt Learning for Facial Privacy-Preserving Text-to-Image Generation
by: Shi, Liang, et al.
Published: (2024)
by: Shi, Liang, et al.
Published: (2024)
DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation
by: Chen, Hong, et al.
Published: (2023)
by: Chen, Hong, et al.
Published: (2023)
EmoFeedback$^2$: Reinforcement of Continuous Emotional Image Generation via LVLM-based Reward and Textual Feedback
by: Jia, Jingyang, et al.
Published: (2025)
by: Jia, Jingyang, et al.
Published: (2025)
StableAnimator: High-Quality Identity-Preserving Human Image Animation
by: Tu, Shuyuan, et al.
Published: (2024)
by: Tu, Shuyuan, et al.
Published: (2024)
Identity-Preserving Image-to-Video Generation via Reward-Guided Optimization
by: Shen, Liao, et al.
Published: (2025)
by: Shen, Liao, et al.
Published: (2025)
RestorerID: Towards Tuning-Free Face Restoration with ID Preservation
by: Ying, Jiacheng, et al.
Published: (2024)
by: Ying, Jiacheng, et al.
Published: (2024)
Deep Reward Supervisions for Tuning Text-to-Image Diffusion Models
by: Wu, Xiaoshi, et al.
Published: (2024)
by: Wu, Xiaoshi, et al.
Published: (2024)
ID-Animator: Zero-Shot Identity-Preserving Human Video Generation
by: He, Xuanhua, et al.
Published: (2024)
by: He, Xuanhua, et al.
Published: (2024)
Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards
by: Kim, Seungwook, et al.
Published: (2026)
by: Kim, Seungwook, et al.
Published: (2026)
FITA: Fine-grained Image-Text Aligner for Radiology Report Generation
by: Yang, Honglong, et al.
Published: (2024)
by: Yang, Honglong, et al.
Published: (2024)
Enhancing Image Caption Generation Using Reinforcement Learning with Human Feedback
by: L, Adarsh N, et al.
Published: (2024)
by: L, Adarsh N, et al.
Published: (2024)
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation
by: Zhou, Sashuai, et al.
Published: (2026)
by: Zhou, Sashuai, et al.
Published: (2026)
Enhancing Spatial Understanding in Image Generation via Reward Modeling
by: Tang, Zhenyu, et al.
Published: (2026)
by: Tang, Zhenyu, et al.
Published: (2026)
RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards
by: Ye, Junyan, et al.
Published: (2025)
by: Ye, Junyan, et al.
Published: (2025)
ID-EA: Identity-driven Text Enhancement and Adaptation with Textual Inversion for Personalized Text-to-Image Generation
by: Jin, Hyun-Jun, et al.
Published: (2025)
by: Jin, Hyun-Jun, et al.
Published: (2025)
Enhancing Reward Models for High-quality Image Generation: Beyond Text-Image Alignment
by: Ba, Ying, et al.
Published: (2025)
by: Ba, Ying, et al.
Published: (2025)
TextMatch: Enhancing Image-Text Consistency Through Multimodal Optimization
by: Luo, Yucong, et al.
Published: (2024)
by: Luo, Yucong, et al.
Published: (2024)
WithAnyone: Towards Controllable and ID Consistent Image Generation
by: Xu, Hengyuan, et al.
Published: (2025)
by: Xu, Hengyuan, et al.
Published: (2025)
Dual-Perspective Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial Labels
by: Pu, Tao, et al.
Published: (2022)
by: Pu, Tao, et al.
Published: (2022)
Multi-path Exploration and Feedback Adjustment for Text-to-Image Person Retrieval
by: Kang, Bin, et al.
Published: (2024)
by: Kang, Bin, et al.
Published: (2024)
PhyCustom: Towards Realistic Physical Customization in Text-to-Image Generation
by: Wu, Fan, et al.
Published: (2025)
by: Wu, Fan, et al.
Published: (2025)
AnyPhoto: Multi-Person Identity Preserving Image Generation with ID Adaptive Modulation on Location Canvas
by: Yuan, Longhui
Published: (2026)
by: Yuan, Longhui
Published: (2026)
ArtAug: Enhancing Text-to-Image Generation through Synthesis-Understanding Interaction
by: Duan, Zhongjie, et al.
Published: (2024)
by: Duan, Zhongjie, et al.
Published: (2024)
Similar Items
-
Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
by: Chen, Weifeng, et al.
Published: (2023) -
ConsistentID: Portrait Generation with Multimodal Fine-Grained Identity Preserving
by: Huang, Jiehui, et al.
Published: (2024) -
DiffusionAgent: Navigating Expert Models for Agentic Image Generation
by: Qin, Jie, et al.
Published: (2024) -
IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model
by: Ji, Yatai, et al.
Published: (2024) -
InstantID: Zero-shot Identity-Preserving Generation in Seconds
by: Wang, Qixun, et al.
Published: (2024)