Avatar Fingerprinting for Authorized Use of Synthetic Talking-Head Videos

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Prashnani, Ekta, Nagano, Koki, De Mello, Shalini, Luebke, David, Gallo, Orazio
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911975578009600
author Prashnani, Ekta
Nagano, Koki
De Mello, Shalini
Luebke, David
Gallo, Orazio
author_facet Prashnani, Ekta
Nagano, Koki
De Mello, Shalini
Luebke, David
Gallo, Orazio
contents Modern avatar generators allow anyone to synthesize photorealistic real-time talking avatars, ushering in a new era of avatar-based human communication, such as with immersive AR/VR interactions or videoconferencing with limited bandwidths. Their safe adoption, however, requires a mechanism to verify if the rendered avatar is trustworthy: does it use the appearance of an individual without their consent? We term this task avatar fingerprinting. To tackle it, we first introduce a large-scale dataset of real and synthetic videos of people interacting on a video call, where the synthetic videos are generated using the facial appearance of one person and the expressions of another. We verify the identity driving the expressions in a synthetic video, by learning motion signatures that are independent of the facial appearance shown. Our solution, the first in this space, achieves an average AUC of 0.85. Critical to its practical use, it also generalizes to new generators never seen in training (average AUC of 0.83). The proposed dataset and other resources can be found at: https://research.nvidia.com/labs/nxp/avatar-fingerprinting/.
format Preprint
id arxiv_https___arxiv_org_abs_2305_03713
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Avatar Fingerprinting for Authorized Use of Synthetic Talking-Head Videos
Prashnani, Ekta
Nagano, Koki
De Mello, Shalini
Luebke, David
Gallo, Orazio
Computer Vision and Pattern Recognition
Modern avatar generators allow anyone to synthesize photorealistic real-time talking avatars, ushering in a new era of avatar-based human communication, such as with immersive AR/VR interactions or videoconferencing with limited bandwidths. Their safe adoption, however, requires a mechanism to verify if the rendered avatar is trustworthy: does it use the appearance of an individual without their consent? We term this task avatar fingerprinting. To tackle it, we first introduce a large-scale dataset of real and synthetic videos of people interacting on a video call, where the synthetic videos are generated using the facial appearance of one person and the expressions of another. We verify the identity driving the expressions in a synthetic video, by learning motion signatures that are independent of the facial appearance shown. Our solution, the first in this space, achieves an average AUC of 0.85. Critical to its practical use, it also generalizes to new generators never seen in training (average AUC of 0.83). The proposed dataset and other resources can be found at: https://research.nvidia.com/labs/nxp/avatar-fingerprinting/.
title Avatar Fingerprinting for Authorized Use of Synthetic Talking-Head Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2305.03713