MEAT: Multiview Diffusion Model for Human Generation on Megapixels with Mesh Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yuhan, Hong, Fangzhou, Yang, Shuai, Jiang, Liming, Wu, Wayne, Loy, Chen Change
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912270065336320
author Wang, Yuhan
Hong, Fangzhou
Yang, Shuai
Jiang, Liming
Wu, Wayne
Loy, Chen Change
author_facet Wang, Yuhan
Hong, Fangzhou
Yang, Shuai
Jiang, Liming
Wu, Wayne
Loy, Chen Change
contents Multiview diffusion models have shown considerable success in image-to-3D generation for general objects. However, when applied to human data, existing methods have yet to deliver promising results, largely due to the challenges of scaling multiview attention to higher resolutions. In this paper, we explore human multiview diffusion models at the megapixel level and introduce a solution called mesh attention to enable training at 1024x1024 resolution. Using a clothed human mesh as a central coarse geometric representation, the proposed mesh attention leverages rasterization and projection to establish direct cross-view coordinate correspondences. This approach significantly reduces the complexity of multiview attention while maintaining cross-view consistency. Building on this foundation, we devise a mesh attention block and combine it with keypoint conditioning to create our human-specific multiview diffusion model, MEAT. In addition, we present valuable insights into applying multiview human motion videos for diffusion training, addressing the longstanding issue of data scarcity. Extensive experiments show that MEAT effectively generates dense, consistent multiview human images at the megapixel level, outperforming existing multiview diffusion methods.
format Preprint
id arxiv_https___arxiv_org_abs_2503_08664
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MEAT: Multiview Diffusion Model for Human Generation on Megapixels with Mesh Attention
Wang, Yuhan
Hong, Fangzhou
Yang, Shuai
Jiang, Liming
Wu, Wayne
Loy, Chen Change
Computer Vision and Pattern Recognition
Artificial Intelligence
Multiview diffusion models have shown considerable success in image-to-3D generation for general objects. However, when applied to human data, existing methods have yet to deliver promising results, largely due to the challenges of scaling multiview attention to higher resolutions. In this paper, we explore human multiview diffusion models at the megapixel level and introduce a solution called mesh attention to enable training at 1024x1024 resolution. Using a clothed human mesh as a central coarse geometric representation, the proposed mesh attention leverages rasterization and projection to establish direct cross-view coordinate correspondences. This approach significantly reduces the complexity of multiview attention while maintaining cross-view consistency. Building on this foundation, we devise a mesh attention block and combine it with keypoint conditioning to create our human-specific multiview diffusion model, MEAT. In addition, we present valuable insights into applying multiview human motion videos for diffusion training, addressing the longstanding issue of data scarcity. Extensive experiments show that MEAT effectively generates dense, consistent multiview human images at the megapixel level, outperforming existing multiview diffusion methods.
title MEAT: Multiview Diffusion Model for Human Generation on Megapixels with Mesh Attention
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2503.08664