FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lyu, Weijie, Yang, Ming-Hsuan, Shu, Zhixin
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912945777147904
author Lyu, Weijie
Yang, Ming-Hsuan
Shu, Zhixin
author_facet Lyu, Weijie
Yang, Ming-Hsuan
Shu, Zhixin
contents We introduce FaceCam, a system that generates video under customizable camera trajectories for monocular human portrait video input. Recent camera control approaches based on large video-generation models have shown promising progress but often exhibit geometric distortions and visual artifacts on portrait videos due to scale-ambiguous camera representations or 3D reconstruction errors. To overcome these limitations, we propose a face-tailored scale-aware representation for camera transformations that provides deterministic conditioning without relying on 3D priors. We train a video generation model on both multi-view studio captures and in-the-wild monocular videos, and introduce two camera-control data generation strategies: synthetic camera motion and multi-shot stitching, to exploit stationary training cameras while generalizing to dynamic, continuous camera trajectories at inference time. Experiments on Ava-256 dataset and diverse in-the-wild videos demonstrate that FaceCam achieves superior performance in camera controllability, visual quality, identity and motion preservation.
format Preprint
id arxiv_https___arxiv_org_abs_2603_05506
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning
Lyu, Weijie
Yang, Ming-Hsuan
Shu, Zhixin
Computer Vision and Pattern Recognition
We introduce FaceCam, a system that generates video under customizable camera trajectories for monocular human portrait video input. Recent camera control approaches based on large video-generation models have shown promising progress but often exhibit geometric distortions and visual artifacts on portrait videos due to scale-ambiguous camera representations or 3D reconstruction errors. To overcome these limitations, we propose a face-tailored scale-aware representation for camera transformations that provides deterministic conditioning without relying on 3D priors. We train a video generation model on both multi-view studio captures and in-the-wild monocular videos, and introduce two camera-control data generation strategies: synthetic camera motion and multi-shot stitching, to exploit stationary training cameras while generalizing to dynamic, continuous camera trajectories at inference time. Experiments on Ava-256 dataset and diverse in-the-wild videos demonstrate that FaceCam achieves superior performance in camera controllability, visual quality, identity and motion preservation.
title FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.05506