Pareto-Enhanced Portrait Generation: Vision-Aligned Text Supervision for Alignment, Realism, and Aesthetics
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yunlong, Shi, Jinjin, Gao, Wenbin, Xu, Xuran, Shi, Runyu, Huang, Ying |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SIGMA: Bridging Structural and Distributional Gaps for Vision Foundation Model Adaptation
by: Xiong, Lingyu, et al.
Published: (2026)
by: Xiong, Lingyu, et al.
Published: (2026)
Optimizing Negative Prompts for Enhanced Aesthetics and Fidelity in Text-To-Image Generation
by: Ogezi, Michael, et al.
Published: (2024)
by: Ogezi, Michael, et al.
Published: (2024)
Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and Algorithms
by: Zhang, Miaosen, et al.
Published: (2024)
by: Zhang, Miaosen, et al.
Published: (2024)
BodyMetric: Evaluating the Realism of Human Bodies in Text-to-Image Generation
by: Andreou, Nefeli, et al.
Published: (2024)
by: Andreou, Nefeli, et al.
Published: (2024)
X-Portrait: Expressive Portrait Animation with Hierarchical Motion Attention
by: Xie, You, et al.
Published: (2024)
by: Xie, You, et al.
Published: (2024)
Personalized Safety Alignment for Text-to-Image Diffusion Models
by: Lei, Yu, et al.
Published: (2025)
by: Lei, Yu, et al.
Published: (2025)
Domain Generalizable Portrait Style Transfer
by: Wang, Xinbo, et al.
Published: (2025)
by: Wang, Xinbo, et al.
Published: (2025)
RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realism
by: Peruzzo, Elia, et al.
Published: (2025)
by: Peruzzo, Elia, et al.
Published: (2025)
Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation
by: Li, Daiqing, et al.
Published: (2024)
by: Li, Daiqing, et al.
Published: (2024)
POCA: Pareto-Optimal Curriculum Alignment for Visual Text Generation
by: Fan, Yaohou, et al.
Published: (2026)
by: Fan, Yaohou, et al.
Published: (2026)
RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment
by: Gu, Difei, et al.
Published: (2025)
by: Gu, Difei, et al.
Published: (2025)
Advancing Comprehensive Aesthetic Insight with Multi-Scale Text-Guided Self-Supervised Learning
by: Liu, Yuti, et al.
Published: (2024)
by: Liu, Yuti, et al.
Published: (2024)
Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs Reasoning
by: Liu, Junming, et al.
Published: (2025)
by: Liu, Junming, et al.
Published: (2025)
Enhance Vision-Language Alignment with Noise
by: Huang, Sida, et al.
Published: (2024)
by: Huang, Sida, et al.
Published: (2024)
Before the Shutter: Aesthetic and Actionable Portrait Photography Planning in 3D Scenes
by: Jiang, Ruixiang, et al.
Published: (2026)
by: Jiang, Ruixiang, et al.
Published: (2026)
CiQi-Agent: Aligning Vision, Tools and Aesthetics in Multimodal Agent for Cultural Reasoning on Chinese Porcelains
by: Wang, Wenhan, et al.
Published: (2026)
by: Wang, Wenhan, et al.
Published: (2026)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
Disentangling Foreground and Background Motion for Enhanced Realism in Human Video Generation
by: Liu, Jinlin, et al.
Published: (2024)
by: Liu, Jinlin, et al.
Published: (2024)
GEA: Generation-Enhanced Alignment for Text-to-Image Person Retrieval
by: Zou, Hao, et al.
Published: (2025)
by: Zou, Hao, et al.
Published: (2025)
TagAlign: Improving Vision-Language Alignment with Multi-Tag Classification
by: Liu, Qinying, et al.
Published: (2023)
by: Liu, Qinying, et al.
Published: (2023)
Parametric Shadow Control for Portrait Generation in Text-to-Image Diffusion Models
by: Cai, Haoming, et al.
Published: (2025)
by: Cai, Haoming, et al.
Published: (2025)
AlignGemini: Generalizable AI-Generated Image Detection Through Task-Model Alignment
by: Chen, Ruoxin, et al.
Published: (2025)
by: Chen, Ruoxin, et al.
Published: (2025)
On the Diversity and Realism of Distilled Dataset: An Efficient Dataset Distillation Paradigm
by: Sun, Peng, et al.
Published: (2023)
by: Sun, Peng, et al.
Published: (2023)
TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards
by: Cui, Mingxuan, et al.
Published: (2026)
by: Cui, Mingxuan, et al.
Published: (2026)
Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model
by: Feng, Qianhan, et al.
Published: (2024)
by: Feng, Qianhan, et al.
Published: (2024)
CLIP Brings Better Features to Visual Aesthetics Learners
by: Xu, Liwu, et al.
Published: (2023)
by: Xu, Liwu, et al.
Published: (2023)
Unlocking the Essence of Beauty: Advanced Aesthetic Reasoning with Relative-Absolute Policy Optimization
by: Liu, Boyang, et al.
Published: (2025)
by: Liu, Boyang, et al.
Published: (2025)
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data
by: Ogezi, Michael, et al.
Published: (2025)
by: Ogezi, Michael, et al.
Published: (2025)
"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
by: Gu, Jing, et al.
Published: (2025)
by: Gu, Jing, et al.
Published: (2025)
Multi-Cache Enhanced Prototype Learning for Test-Time Generalization of Vision-Language Models
by: Chen, Xinyu, et al.
Published: (2025)
by: Chen, Xinyu, et al.
Published: (2025)
AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
PASTA: Part-Aware Sketch-to-3D Shape Generation with Text-Aligned Prior
by: Lee, Seunggwan, et al.
Published: (2025)
by: Lee, Seunggwan, et al.
Published: (2025)
Reflective Human-Machine Co-adaptation for Enhanced Text-to-Image Generation Dialogue System
by: Feng, Yuheng, et al.
Published: (2024)
by: Feng, Yuheng, et al.
Published: (2024)
AlignIT: Enhancing Prompt Alignment in Customization of Text-to-Image Models
by: Agarwal, Aishwarya, et al.
Published: (2024)
by: Agarwal, Aishwarya, et al.
Published: (2024)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
by: Jin, Zhe, et al.
Published: (2025)
by: Jin, Zhe, et al.
Published: (2025)
RealCompo: Balancing Realism and Compositionality Improves Text-to-Image Diffusion Models
by: Zhang, Xinchen, et al.
Published: (2024)
by: Zhang, Xinchen, et al.
Published: (2024)
Enhancing Alignment for Unified Multimodal Models via Semantically-Grounded Supervision
by: Kim, Jiyeong, et al.
Published: (2026)
by: Kim, Jiyeong, et al.
Published: (2026)
Position: Universal Aesthetic Alignment Narrows Artistic Expression
by: Guo, Wenqi Marshall, et al.
Published: (2025)
by: Guo, Wenqi Marshall, et al.
Published: (2025)
ConsistentID: Portrait Generation with Multimodal Fine-Grained Identity Preserving
by: Huang, Jiehui, et al.
Published: (2024)
by: Huang, Jiehui, et al.
Published: (2024)
FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video Generation
by: Tan, Weiting, et al.
Published: (2026)
by: Tan, Weiting, et al.
Published: (2026)
Similar Items
-
SIGMA: Bridging Structural and Distributional Gaps for Vision Foundation Model Adaptation
by: Xiong, Lingyu, et al.
Published: (2026) -
Optimizing Negative Prompts for Enhanced Aesthetics and Fidelity in Text-To-Image Generation
by: Ogezi, Michael, et al.
Published: (2024) -
Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and Algorithms
by: Zhang, Miaosen, et al.
Published: (2024) -
BodyMetric: Evaluating the Realism of Human Bodies in Text-to-Image Generation
by: Andreou, Nefeli, et al.
Published: (2024) -
X-Portrait: Expressive Portrait Animation with Hierarchical Motion Attention
by: Xie, You, et al.
Published: (2024)