CosmicMan: A Text-to-Image Foundation Model for Humans
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Shikai, Fu, Jianglin, Liu, Kaiyuan, Wang, Wentao, Lin, Kwan-Yee, Wu, Wayne |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HOI-Dyn: Learning Interaction Dynamics for Human-Object Motion Diffusion
by: Wu, Lin, et al.
Published: (2025)
by: Wu, Lin, et al.
Published: (2025)
Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
by: Zhao, Shihao, et al.
Published: (2024)
by: Zhao, Shihao, et al.
Published: (2024)
Visually-grounded Humanoid Agents
by: Ye, Hang, et al.
Published: (2026)
by: Ye, Hang, et al.
Published: (2026)
A Survey on 3D Human Avatar Modeling -- From Reconstruction to Generation
by: Wang, Ruihe, et al.
Published: (2024)
by: Wang, Ruihe, et al.
Published: (2024)
Modeling Thousands of Human Annotators for Generalizable Text-to-Image Person Re-identification
by: Jiang, Jiayu, et al.
Published: (2025)
by: Jiang, Jiayu, et al.
Published: (2025)
Parameterization-driven Neural Surface Reconstruction for Object-oriented Editing in Neural Rendering
by: Xu, Baixin, et al.
Published: (2023)
by: Xu, Baixin, et al.
Published: (2023)
Text-guided Foundation Model Adaptation for Long-Tailed Medical Image Classification
by: Li, Sirui, et al.
Published: (2024)
by: Li, Sirui, et al.
Published: (2024)
Rank-Aware Agglomeration of Foundation Models for Immunohistochemistry Image Cell Counting
by: Huang, Zuqi, et al.
Published: (2025)
by: Huang, Zuqi, et al.
Published: (2025)
TimeWalker: Personalized Neural Space for Lifelong Head Avatars
by: Pan, Dongwei, et al.
Published: (2024)
by: Pan, Dongwei, et al.
Published: (2024)
A Token-level Text Image Foundation Model for Document Understanding
by: Guan, Tongkun, et al.
Published: (2025)
by: Guan, Tongkun, et al.
Published: (2025)
Mind the Gap: Aligning Vision Foundation Models to Image Feature Matching
by: Liu, Yuhan, et al.
Published: (2025)
by: Liu, Yuhan, et al.
Published: (2025)
Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models
by: Chen, Dong, et al.
Published: (2026)
by: Chen, Dong, et al.
Published: (2026)
MM-Retinal: Knowledge-Enhanced Foundational Pretraining with Fundus Image-Text Expertise
by: Wu, Ruiqi, et al.
Published: (2024)
by: Wu, Ruiqi, et al.
Published: (2024)
Joint Optimization for 4D Human-Scene Reconstruction in the Wild
by: Liu, Zhizheng, et al.
Published: (2025)
by: Liu, Zhizheng, et al.
Published: (2025)
Restore-R1: Efficient Image Restoration Agents via Reinforcement Learning with Multimodal LLM Perceptual Feedback
by: Lu, Jianglin, et al.
Published: (2025)
by: Lu, Jianglin, et al.
Published: (2025)
AvatarGO: Zero-shot 4D Human-Object Interaction Generation and Animation
by: Cao, Yukang, et al.
Published: (2024)
by: Cao, Yukang, et al.
Published: (2024)
PLACE: Adaptive Layout-Semantic Fusion for Semantic Image Synthesis
by: Lv, Zhengyao, et al.
Published: (2024)
by: Lv, Zhengyao, et al.
Published: (2024)
Learning Multi-dimensional Human Preference for Text-to-Image Generation
by: Zhang, Sixian, et al.
Published: (2024)
by: Zhang, Sixian, et al.
Published: (2024)
Detecting Human Artifacts from Text-to-Image Models
by: Wang, Kaihong, et al.
Published: (2024)
by: Wang, Kaihong, et al.
Published: (2024)
ArtiFade: Learning to Generate High-quality Subject from Blemished Images
by: Yang, Shuya, et al.
Published: (2024)
by: Yang, Shuya, et al.
Published: (2024)
VipDiff: Towards Coherent and Diverse Video Inpainting via Training-free Denoising Diffusion Models
by: Xie, Chaohao, et al.
Published: (2025)
by: Xie, Chaohao, et al.
Published: (2025)
TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models
by: Xiao, Yao, et al.
Published: (2025)
by: Xiao, Yao, et al.
Published: (2025)
LooC: Effective Low-Dimensional Codebook for Compositional Vector Quantization
by: Li, Jie, et al.
Published: (2026)
by: Li, Jie, et al.
Published: (2026)
Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation
by: Cao, Chenjie, et al.
Published: (2025)
by: Cao, Chenjie, et al.
Published: (2025)
Text4Seg: Reimagining Image Segmentation as Text Generation
by: Lan, Mengcheng, et al.
Published: (2024)
by: Lan, Mengcheng, et al.
Published: (2024)
Let Humanoids Hike! Integrative Skill Development on Complex Trails
by: Lin, Kwan-Yee, et al.
Published: (2025)
by: Lin, Kwan-Yee, et al.
Published: (2025)
DeltaEdit: Exploring Text-free Training for Text-Driven Image Manipulation
by: Lyu, Yueming, et al.
Published: (2023)
by: Lyu, Yueming, et al.
Published: (2023)
AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image Generation
by: Pang, Lianyu, et al.
Published: (2024)
by: Pang, Lianyu, et al.
Published: (2024)
S-VGGT: Structure-Aware Subscene Decomposition for Scalable 3D Foundation Models
by: Li, Xinze, et al.
Published: (2026)
by: Li, Xinze, et al.
Published: (2026)
Scaling Down Text Encoders of Text-to-Image Diffusion Models
by: Wang, Lifu, et al.
Published: (2025)
by: Wang, Lifu, et al.
Published: (2025)
Evaluating and Predicting Distorted Human Body Parts for Generated Images
by: Ma, Lu, et al.
Published: (2025)
by: Ma, Lu, et al.
Published: (2025)
Modular Energy Steering for Safe Text-to-Image Generation with Foundation Models
by: Tan, Yaoteng, et al.
Published: (2026)
by: Tan, Yaoteng, et al.
Published: (2026)
Continual Unlearning for Foundational Text-to-Image Models without Generalization Erosion
by: Thakral, Kartik, et al.
Published: (2025)
by: Thakral, Kartik, et al.
Published: (2025)
Fine-Grained Erasure in Text-to-Image Diffusion-based Foundation Models
by: Thakral, Kartik, et al.
Published: (2025)
by: Thakral, Kartik, et al.
Published: (2025)
GeoBridge: A Semantic-Anchored Multi-View Foundation Model Bridging Images and Text for Geo-Localization
by: Song, Zixuan, et al.
Published: (2025)
by: Song, Zixuan, et al.
Published: (2025)
From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning
by: He, Honglin, et al.
Published: (2025)
by: He, Honglin, et al.
Published: (2025)
UniHuman: A Unified Model for Editing Human Images in the Wild
by: Li, Nannan, et al.
Published: (2023)
by: Li, Nannan, et al.
Published: (2023)
Tuning-Free Image Customization with Image and Text Guidance
by: Li, Pengzhi, et al.
Published: (2024)
by: Li, Pengzhi, et al.
Published: (2024)
Interactive Visual Assessment for Text-to-Image Generation Models
by: Mi, Xiaoyue, et al.
Published: (2024)
by: Mi, Xiaoyue, et al.
Published: (2024)
Unveiling Structural Memorization: Structural Membership Inference Attack for Text-to-Image Diffusion Models
by: Li, Qiao, et al.
Published: (2024)
by: Li, Qiao, et al.
Published: (2024)
Similar Items
-
HOI-Dyn: Learning Interaction Dynamics for Human-Object Motion Diffusion
by: Wu, Lin, et al.
Published: (2025) -
Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
by: Zhao, Shihao, et al.
Published: (2024) -
Visually-grounded Humanoid Agents
by: Ye, Hang, et al.
Published: (2026) -
A Survey on 3D Human Avatar Modeling -- From Reconstruction to Generation
by: Wang, Ruihe, et al.
Published: (2024) -
Modeling Thousands of Human Annotators for Generalizable Text-to-Image Person Re-identification
by: Jiang, Jiayu, et al.
Published: (2025)