Salvato in:
Dettagli Bibliografici
Autori principali: Khalifi, Omar El, Rossi, Thomas, Fossey, Oscar, Fouque, Thibault, Mizrahi, Ulysse, Torr, Philip, Laptev, Ivan, Pizzati, Fabio, Bellot-Gurlet, Baptiste
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:https://arxiv.org/abs/2605.06667
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915989772304384
author Khalifi, Omar El
Rossi, Thomas
Fossey, Oscar
Fouque, Thibault
Mizrahi, Ulysse
Torr, Philip
Laptev, Ivan
Pizzati, Fabio
Bellot-Gurlet, Baptiste
author_facet Khalifi, Omar El
Rossi, Thomas
Fossey, Oscar
Fouque, Thibault
Mizrahi, Ulysse
Torr, Philip
Laptev, Ivan
Pizzati, Fabio
Bellot-Gurlet, Baptiste
contents For artistic applications, video generation requires fine-grained control over both performance and cinematography, i.e., the actor's motion and the camera trajectory. We present ActCam, a zero-shot method for video generation that jointly transfers character motion from a driving video into a new scene and enables per-frame control of intrinsic and extrinsic camera parameters. ActCam builds on any pretrained image-to-video diffusion model that accepts conditioning in terms of scene depth and character pose. Given a source video with a moving character and a target camera motion, ActCam generates pose and depth conditions that remain geometrically consistent across frames. We then run a single sampling process with a two-phase conditioning schedule: early denoising steps condition on both pose and sparse depth to enforce scene structure, after which depth is dropped and pose-only guidance refines high-frequency details without over-constraining the generation. We evaluate ActCam on multiple benchmarks spanning diverse character motions and challenging viewpoint changes. We find that, compared to pose-only control and other pose and camera methods, ActCam improves camera adherence and motion fidelity, and is preferred in human evaluations, especially under large viewpoint changes. Our results highlight that careful camera-consistent conditioning and staged guidance can enable strong joint camera and motion control without training. Project page: https://elkhomar.github.io/actcam/.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06667
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation
Khalifi, Omar El
Rossi, Thomas
Fossey, Oscar
Fouque, Thibault
Mizrahi, Ulysse
Torr, Philip
Laptev, Ivan
Pizzati, Fabio
Bellot-Gurlet, Baptiste
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
For artistic applications, video generation requires fine-grained control over both performance and cinematography, i.e., the actor's motion and the camera trajectory. We present ActCam, a zero-shot method for video generation that jointly transfers character motion from a driving video into a new scene and enables per-frame control of intrinsic and extrinsic camera parameters. ActCam builds on any pretrained image-to-video diffusion model that accepts conditioning in terms of scene depth and character pose. Given a source video with a moving character and a target camera motion, ActCam generates pose and depth conditions that remain geometrically consistent across frames. We then run a single sampling process with a two-phase conditioning schedule: early denoising steps condition on both pose and sparse depth to enforce scene structure, after which depth is dropped and pose-only guidance refines high-frequency details without over-constraining the generation. We evaluate ActCam on multiple benchmarks spanning diverse character motions and challenging viewpoint changes. We find that, compared to pose-only control and other pose and camera methods, ActCam improves camera adherence and motion fidelity, and is preferred in human evaluations, especially under large viewpoint changes. Our results highlight that careful camera-consistent conditioning and staged guidance can enable strong joint camera and motion control without training. Project page: https://elkhomar.github.io/actcam/.
title ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.06667