<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>Controllable Radiance Fields for Dynamic Face Synthesis</title></titleStmt>
			<publicationStmt>
				<publisher></publisher>
				<date>2022</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10387045</idno>
					<idno type="doi"></idno>
					<title level='j'>3DV</title>
<idno>0219-6921</idno>
<biblScope unit="volume"></biblScope>
<biblScope unit="issue"></biblScope>					

					<author>Peiye Zhuang</author><author>Liqian Ma</author><author>Oluwasanmi Koyejo</author><author>Alexander Schwing</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[Recent work on 3D-aware image synthesis has achieved compelling results using advances in neural rendering. However, 3D-aware synthesis of face dynamics hasn't received much attention. Here, we study how to explicitly control generative model synthesis of face dynamics exhibiting non-rigid motion (e.g., facial expression change), while simultaneously ensuring 3D-awareness. For this we propose a Controllable Radiance Field (CoRF): 1) Motion control is achieved by embedding motion features within the layered latent motion space of a style-based generator; 2) To ensure consistency of background, motion features and subject-specific attributes such as lighting, texture, shapes, albedo, and identity, a face parsing net, a head regressor and an identity encoder are incorporated. On head image/video data we show that CoRFs are 3D-aware while enabling editing of identity, viewing directions, and motion.]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.">Introduction</head><p>Face synthesis is an important task with applications in digital content creation, film-making and Virtual Reality (VR). Generative Adversarial Nets (GANs), as a powerful generative model, have demonstrated remarkable success in generating high-quality faces. However, despite the impressive performance of GANs in both 2D <ref type="bibr">[15,</ref><ref type="bibr">22,</ref><ref type="bibr">2,</ref><ref type="bibr">24,</ref><ref type="bibr">25,</ref><ref type="bibr">23]</ref> and 3D image synthesis <ref type="bibr">[35,</ref><ref type="bibr">46,</ref><ref type="bibr">7,</ref><ref type="bibr">16]</ref>, its extension to both 3D-aware and motion-controllable image synthesis has not been fully explored. This is largely due to the complexity of physical motion and the challenging yet consistent appearance changes of an identity. For example, dynamic head synthesis requires maintaining some 3D consistency over time to preserve facial identity while permitting other 3D deformations due to expression changes. It is even more challenging to enable interactive manipulation of identity, motion, and viewing directions.</p><p>Addressing this task of generalizable, 3D-aware, and motion-controllable synthesis of face dynamics enables to generate never-before-seen faces including their dynamic motion, while controlling the viewpoint, as shown in Fig. <ref type="figure">1</ref>.</p><p>Most related to this task are motion transfer methods such as face reenactment <ref type="bibr">[69,</ref><ref type="bibr">64,</ref><ref type="bibr">47,</ref><ref type="bibr">61,</ref><ref type="bibr">70,</ref><ref type="bibr">67,</ref><ref type="bibr">49,</ref><ref type="bibr">48]</ref>. Specifically, prior work animates a subject shown in a source image given target motion from a driving video. For this, prior methods commonly use 2D GANs to produce dynamic subject images with additional guidance such as reference keypoints <ref type="bibr">[69,</ref><ref type="bibr">64,</ref><ref type="bibr">47,</ref><ref type="bibr">61,</ref><ref type="bibr">70,</ref><ref type="bibr">67]</ref> and self-learned motion representations <ref type="bibr">[48,</ref><ref type="bibr">49]</ref>. However, due to intrinsic limitations of 2D convolutions, these approaches lack 3D-awareness. This leads to two main failure modes: 1) spurious identity changes -the subject, head pose or facial expression of the source image is distinct from the target when rotating a source head by a large angle; 2) serious distortions when transferring motion across identity based on keypoints which contain identity-specific information such as face shapes. To address these failure modes, recent face reenactment works <ref type="bibr">[40,</ref><ref type="bibr">13]</ref> estimate parameters of 3D morphable face models, e.g., poses and expression, as additional guidance. Notably, Guy et al. <ref type="bibr">[13]</ref> explicitly represent heads in a 3D radiance field showing impressive 3Dconsistency. However, one model per face identity needs to be trained <ref type="bibr">[13]</ref>.</p><p>Different from this direction, we develop a generalizable method of motion-controllable synthesis for novel, neverbefore-seen identity generation. This differs from prior works which either consider motion control but drop 3Dawareness <ref type="bibr">[43,</ref><ref type="bibr">56,</ref><ref type="bibr">44,</ref><ref type="bibr">63,</ref><ref type="bibr">69,</ref><ref type="bibr">64,</ref><ref type="bibr">48,</ref><ref type="bibr">47,</ref><ref type="bibr">61,</ref><ref type="bibr">70,</ref><ref type="bibr">67]</ref>, or are 3D-aware but don't consider dynamics <ref type="bibr">[46,</ref><ref type="bibr">7,</ref><ref type="bibr">16]</ref> or can't generate never-before-seen identities <ref type="bibr">[40,</ref><ref type="bibr">13]</ref>.</p><p>To achieve our goal, we propose controllable radiance fields (CoRFs). We learn CoRFs from RGB images/videos with unknown camera poses. This requires to address two main challenges: 1) how to effectively represent and control identity and motion in 3D; and 2) how to ensure spatiotemporal consistency across views and time.</p><p>To control identity and motion, CoRFs use a stylebased radiance field generator which takes low-dimensional identity and motion representations as input, as shown in Fig. <ref type="figure">1</ref>. Unlike prior head reconstruction work <ref type="bibr">[40,</ref><ref type="bibr">13]</ref> that uses ground-truth images for self-supervision, there is no ground-truth target image for a generated image of a neverbefore-seen person. Thus, we propose an additional motion reconstruction loss to supervise motion control.</p><p>To ensure spatio-temporal consistency, we propose three consistency constraints on face attributes, identities, and background. Specifically, at each training step, we generate two images given the same identity representation yet different motion representations. We encourage that the two images share identical environment and subject-specific attributes. For this, we apply a head regressor and an identity encoder. A head regressor decomposes the images into representations of a statistic face model, including lighting, shape, texture and albedo. Then, we compute a consistency loss which compares the predicted attribute parameters of the paired synthetic images. Moreover, we use an identity encoder to ensure the paired synthetic images share the same identity. Further, to encourage a consistent background across time, we recognize background regions using a face parsing net and employ a background consistency loss between paired synthetic images that share the same identity representation.</p><p>As there is no direct baseline, we compare CoRFs to two types of work that are most related: 1) face reenactment work <ref type="bibr">[48,</ref><ref type="bibr">49]</ref> that can control facial expression transfer; and 2) video synthesis work <ref type="bibr">[56,</ref><ref type="bibr">44,</ref><ref type="bibr">63,</ref><ref type="bibr">55]</ref> that can produce videos for novel identities. We evaluate the improvements of our method using the Fr&#233;chet Inception Distance (FID) <ref type="bibr">[19]</ref> for videos, motion control scores, and 3Daware consistency metrics on FFHQ <ref type="bibr">[22]</ref> and two face video benchmarks <ref type="bibr">[42,</ref><ref type="bibr">9]</ref> at a resolution of 256 &#215; 256. Moreover, we show additional applications such as novel view synthesis and motion interpolation in the latent motion space. Contributions. 1) We study the new task of generalizable, 3D-aware, and motion-controllable face generation. 2) We develop CoRFs which enable editing of motion, and viewpoints of never-before-seen subjects during synthesis. For this, we study techniques that aid motion control and spatiotemporal consistency. 3) CoRFs improve visual quality and spatio-temporal consistency of free-view motion synthesis compared to multiple baselines <ref type="bibr">[56,</ref><ref type="bibr">44,</ref><ref type="bibr">63,</ref><ref type="bibr">48,</ref><ref type="bibr">69,</ref><ref type="bibr">49,</ref><ref type="bibr">55]</ref> on popular face benchmarks <ref type="bibr">[22,</ref><ref type="bibr">42,</ref><ref type="bibr">9]</ref> using multiple metrics for image quality, temporal consistency, identity preservation, and expression transfer.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.">Related work</head><p>3D-aware image synthesis. GANs <ref type="bibr">[15]</ref> have significantly advanced 2D image synthesis capabilities in recent years, addressing early concerns regarding diversity, resolution, and photo-realism <ref type="bibr">[22,</ref><ref type="bibr">2,</ref><ref type="bibr">24,</ref><ref type="bibr">25,</ref><ref type="bibr">23,</ref><ref type="bibr">17]</ref>. More recent GANs aim to extend 2D image synthesis to 3D-aware image generation <ref type="bibr">[37,</ref><ref type="bibr">53,</ref><ref type="bibr">34,</ref><ref type="bibr">35,</ref><ref type="bibr">11]</ref>, while permitting explicit camera control. For this, methods <ref type="bibr">[37,</ref><ref type="bibr">53,</ref><ref type="bibr">34,</ref><ref type="bibr">35,</ref><ref type="bibr">11]</ref> use an implicit neural representation.</p><p>Implicit neural representations <ref type="bibr">[32,</ref><ref type="bibr">8,</ref><ref type="bibr">31,</ref><ref type="bibr">66,</ref><ref type="bibr">28,</ref><ref type="bibr">51,</ref><ref type="bibr">50,</ref><ref type="bibr">38,</ref><ref type="bibr">5]</ref> have been introduced to encode 3D objects (i.e., its shape and/or appearance) via a parameterized neural net. Compared to voxel-based <ref type="bibr">[21,</ref><ref type="bibr">41,</ref><ref type="bibr">65]</ref> and meshbased <ref type="bibr">[36,</ref><ref type="bibr">59]</ref> methods, implicit neural functions represent 3D objects in continuous space and are not restricted to a particular object topology. A recent variation, Neural Radiance Fields (NeRFs) <ref type="bibr">[32]</ref>, represents the appearance and geometry of a static real-world scene with a multi-layer perceptron (MLP), enabling impressive novel-view synthesis with multi-view consistency via volume rendering. Different from the studied method, those classical approaches can't generate novel scenes.</p><p>To address this, recent work <ref type="bibr">[46,</ref><ref type="bibr">7,</ref><ref type="bibr">16,</ref><ref type="bibr">72,</ref><ref type="bibr">71]</ref> studies 3D-aware generative models using NeRFs as a gener-  ator. For instance, GRAF <ref type="bibr">[46]</ref> and &#960;-GAN <ref type="bibr">[7]</ref> operate on a randomly sampled camera pose and latent vectors from prior distributions to produce a radiance field via an MLP. StyleNeRF and CIPS-3D <ref type="bibr">[16,</ref><ref type="bibr">72]</ref> combine a shallow NeRF network to provide low-resolution radiance fields and a 2D rendering network to produce high-resolution images with fine details. LolNeRF <ref type="bibr">[39]</ref> learns 3D objects by optimizing foreground and background NeRFs together with a learnable per-image table of latent codes. Zhao et al. <ref type="bibr">[71]</ref> develop a generative multi-plane image (GMPI) representation to ensure view-consistency. These works only generate multi-view images for static scenes. In contrast, the proposed CoRF targets generation of face dynamics while enabling to control motion. We also notice concurrent NeRFbased generative models <ref type="bibr">[52,</ref><ref type="bibr">20]</ref> aiming for 3D-aware semantic appearance or expression editing.</p><p>Dynamic object synthesis. Dynamic object synthesis has been studied using 2D-and 3D-based methods. Some 2Dbased methods unconditionally generate dynamic instances via convolutional neural nets such as GANs <ref type="bibr">[58,</ref><ref type="bibr">43,</ref><ref type="bibr">44,</ref><ref type="bibr">56,</ref><ref type="bibr">63]</ref>, but struggle with motion control. Follow-up work incorporates additional reference information during synthesis <ref type="bibr">[69,</ref><ref type="bibr">64,</ref><ref type="bibr">48,</ref><ref type="bibr">47,</ref><ref type="bibr">61,</ref><ref type="bibr">70,</ref><ref type="bibr">62,</ref><ref type="bibr">40]</ref>, e.g., for tasks like face reenactment. Our method is related but differs in that CoRF synthesizes images and controls motion for never-beforeseen identities from free viewpoints. 3D-based methods recover face geometry and control the expression by changing geometry parameters <ref type="bibr">[54,</ref><ref type="bibr">26,</ref><ref type="bibr">14,</ref><ref type="bibr">33]</ref>. While they can control facial motion effectively, they struggle to produce real-istic hair, teeth, and accessories <ref type="bibr">[14,</ref><ref type="bibr">33]</ref>. Recent work <ref type="bibr">[13]</ref> adapts neural implicit methods, e.g., NeRFs to learn head reconstruction from images. However, one model is trained per identity, i.e., these methods don't generalize across different identities. Unlike prior work <ref type="bibr">[13]</ref>, CoRF generalizes across identities. Motion conditioning. We draw inspiration from GANbased image inversion and editing work <ref type="bibr">[1,</ref><ref type="bibr">73]</ref>. To be concrete, Abdal et al. <ref type="bibr">[1]</ref> proposed to embed images back to an extended latent identity space of a StyleGAN <ref type="bibr">[24]</ref>, i.e., the W + space. This extended latent identity space was shown to be remarkably useful for image editing and expression transfer in follow-up work <ref type="bibr">[73]</ref>. In our case, we embed a motion representation into a layered latent motion space, which we then broadcast to every synthesis layer in the generator. In Section 4, we illustrate that such a motion embedding strategy benefits motion control. 3D consistency preservation. To ensure 3D consistency over time, a 3D convolutional discriminator is commonly used to distinguish the source of input video clips <ref type="bibr">[55]</ref>. Recent work <ref type="bibr">[68]</ref> finds that when employing an implicit neural rendering net as a generator, a 2D discriminator for 2 synthetic frames in a video is sufficient. Inspired by this result, we propose consistency losses on paired synthetic frames without using a 3D convolutional discriminator.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.">Method</head><p>We aim for generalizable, 3D-aware and motioncontrollable synthesis of face dynamics. Concretely, we seek to generate never-before-seen faces while controlling dynamic motions and viewpoint arbitrarily. For this, we propose controllable radiance fields (CoRF), illustrated in Fig. <ref type="figure">2</ref>, for which we provide an overview next. Overview. CoRFs are composed of a generator G, a discriminator D, a regressor R and an identity encoder E. The CoRF generator G renders an RGB image I based on a noise vector z and a motion representation m. Concretely, the generator G first applies two mapping modules f z and f m , before rendering the image via a style-based synthesis module G s , i.e., G = {f z , f m , G s }. Specifically, to ensure effective motion control, we embed motion representations m into a layered latent motion space using a mapping net f m , which is then combined with a style vector w := f z (z) via a summation. The discriminator D is applied to ensure photo-realism. The regressor R and the encoder E ensure spatio-temporal consistency of synthetic faces over dynamic changes.</p><p>During training, different from prior work, we generate two images given the same identity yet different motion representations. To preserve consistency over views and poses, the regressor R and the identity encoder E extract features for the paired images which are encouraged to be similar via a consistency loss L consist and an identity loss L id . To supervise motion control, we estimate the motion representation from images and incorporate a motion reconstruction loss L motion , to minimize the distance between the estimated motion and the conditioned one. Finally, we apply a background loss, L bg , to preserve background consistency over dynamic changes.</p><p>In the following, we briefly describe the generator G and the discriminator D (Section 3.1). We then describe our contributions: 1) how to control motion (Section 3.2) and 2) how to preserve spatio-temporal consistency (Section 3.3).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.1.">Generator and Discriminator</head><p>Generative radiance fields. We adopt the generative strategy of StyleNeRF <ref type="bibr">[16]</ref> where a style-based radiance field generator G renders a synthetic image I . Specifically, we render colors for each pixel from an integral feature F r . To obtain an integral feature F r , we integrate features F of 3D points along a camera ray r. We present the generator G and its workflow in Fig. <ref type="figure">2</ref>.</p><p>Formally, we cast rays from the camera origin o through the pixels of the image plane, and accumulate the sampled density &#963; and feature values F along the rays with near and far planes t n and t f . The aggregated feature F r is computed using the rendering equation <ref type="bibr">[29]</ref> as follows:</p><p>where</p><p>Here, d refers to the ray direction and T (t) corresponds to the accumulated transmittance along the ray r from t n to t. The functions &#963;(r(t), m, z) and F (r(t), d, m, z) are the density and the feature at a 3D location, respectively.</p><p>Different from StyleNeRF <ref type="bibr">[16]</ref>, the generator G in CoRF is conditioned on a motion representation m &#8712; R 50 for motion control. To this end, a ReLU mapping network f m embeds the motion representation m into a latent space, which is then injected to the synthesis module G s via weight modulation <ref type="bibr">[24]</ref>. We discuss the details in Section 3.2. Also different from StyleNeRF <ref type="bibr">[16]</ref>, the generator G does not contain upsampling blocks. Instead, we increase the image resolution by sampling more rays. Discriminator. We use the estimated camera pose &#958; and the motion feature m of a training image I &#8764; p real to control generation of new images I &#8764; p syn . The discriminator D is conditioned on a camera pose &#958; and a motion feature m, following the conditional strategy of StyleGAN2-ADA <ref type="bibr">[23,</ref><ref type="bibr">6]</ref>. We use a non-saturating adversarial loss with a gradient regularizer <ref type="bibr">[30]</ref> applied on the discriminator D. We refer to this loss as L adv , i.e.,</p><p>where f (u)</p><p>-log(1 + exp(-u)). Further, p m and p &#958; refer to distributions over motion vectors and pose respectively. In practice, we extract motion features and pose from training data samples using <ref type="bibr">[12]</ref>. During training, we randomly sample a training image with its corresponding estimated motion feature and pose to control image synthesis.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.2.">Motion control</head><p>A key goal of CoRF is to control synthesis of 3D-aware face dynamics. For this, we condition the generator G on low-dimensional motion representations m. In this section, we discuss our conditioning method, motion extraction from images, and the supervision which we use during training to encourage motion control. Motion conditioning. We embed the motion representation in a layered latent motion space using the mapping module f m . The layered latent motion space is a concatenation of N latent motion vectors, denoted as [d 1 , . . . , d N ] T := f m (m), where d i &#8712; R 512 for i &#8712; {1, . . . , N }. Here, N corresponds to the number of layers in the synthesis module G s , which we find to lead to feature embeddings with superior performance. The latent motion vectors are then combined with the latent style vector w. Formally, we combine style vector and motion representation via a summation, i.e.,</p><p>The conditioned latent vectors are then given to each layer of the synthesis module G s via modulation <ref type="bibr">[24]</ref>. Instead of summation, other conditioning strategies for control exist. However, we find that not every strategy works equally well in our task. For example, recent head reconstruction work <ref type="bibr">[13]</ref> concatenates a motion representation with a noise vector z which is then provided as input to the first layer of the model. We compare our conditioning strategy with this prior method <ref type="bibr">[13]</ref> in Section 4 and find that our motion conditioning strategy accelerates convergence of the motion reconstruction loss L motion . Efficient motion extraction. To obtain a low-dimensional motion representation m for face dynamics, we extract coefficients of a blendshape basis of a statistical head model using a pre-trained regressor R with an accurate statistical head model <ref type="bibr">[27]</ref>. For this we use a publicly available package. *  Note that our method is also compatible with other pretrained head models. Compared with existing work <ref type="bibr">[13]</ref> that uses an optimization-based approach for expression estimation, use of a pre-trained regressor R enjoys two advantages. First, using the regressor R is computationally more efficient as coefficient prediction requires only a single forward-pass. Second, we can use the regressor R as a differentiable plug-in to regress and supervise the expression representation of synthetically generated frames. Motion learning. A motion reconstruction loss L motion is applied to supervise motion control. It assesses whether the synthetically generated image exhibits the motion for which the representation was provided as input. For this we use a simple quadratic loss, i.e.,</p><p>where m is estimated by the regressor R, i.e., m = R(I ).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.3.">Spatio-temporal consistency</head><p>The aforementioned losses are insufficient to ensure spatio-temporal consistency. To address this, we propose * Code available from <ref type="url">https://github.com/YadiraF/DECA</ref> two losses which focus on the generated face and one loss which focuses on background regions. All three losses are simple yet effective at preserving spatio-temporal consistency. In summary, to optimize CoRF we minimize the weighted objective</p><p>where the hyper-parameters &#955; 1 , &#955; 2 , . . . , &#955; 5 adjust the strength of the different loss terms. In practice, the regressor R and the identity encoder E are pre-trained offline on a large-scale dataset <ref type="bibr">[12,</ref><ref type="bibr">4]</ref>. We describe each of the three additional losses L consist , L id and L bg next. Consistency of attributes. The motion conditioning occasionally changes lighting and person-specific attributes such as albedo and shape, resulting in flickering textures in rendered images. To alleviate this issue, we propose a consistency loss to supervise attributes independent from motion (see Fig. <ref type="figure">3</ref>). Specifically, we extract lighting, texture, shape, and albedo representations for paired artificially generated images which have an identical noise vector z, yet a different motion representation m. These extracted attributes are independent from motion and thus should be identical for paired images. To encourage their consistency via a loss, we subsume representations of these attributes (lighting, albedo, texture, shape) for both paired images in the set A = {(l 1 , l 2 ), (a 1 , a 2 ), (t 1 , t 2 ), (s 1 , s 2 )}. We then compute the consistency loss</p><p>by comparing pairs of estimated attributes x 1 and x 2 . Here, l, a, t, and s are the attributes of lighting, albedo, texture, and shape, respectively. &#946; x is the loss weight. In practice, these attributes are predicted by the regressor R. Consistency of identity. Beyond attribute consistency, we use an identity loss on paired images, following prior face reconstruction work <ref type="bibr">[12]</ref> (see Fig. <ref type="figure">3</ref>). To be concrete, an encoder E outputs identity embeddings of paired generated images which have an identical noise vector z yet a different motion representation. Let I 1 and I 2 refer to the two generated images. The identity loss is proportional to the negative cosine similarity between embeddings of both images, i.e., it is computed via</p><p>Consistency of background. Beyond the constraints on faces, we also want to encourage that the background remains mostly static over time. For this, we recognize background regions of paired synthetic images I 1 and I 2 , using Source Driving FOMM <ref type="bibr">[48]</ref> MRAA <ref type="bibr">[49]</ref> Ours Source Driving FOMM <ref type="bibr">[48]</ref> MRAA <ref type="bibr">[49]</ref> Ours Figure <ref type="figure">4</ref>: Visual comparison of face reenactment on FFHQ. We transfer expression and pose from driving images to the source images. We compare our results with results of FOMM <ref type="bibr">[48]</ref> and MRAA <ref type="bibr">[49]</ref> at a resolution of 256 &#215; 256 in challenging cases. We show face reenactment across identity, gender, pose, or with a partially occluded driving image.</p><p>a pre-trained face parsing net. &#8224; The corresponding binary background masks, which we refer to via s 1 and s 2 , use a value of 1 to indicate the background regions and a value of 0 to indicate the face regions. Using both masks, we define a background consistency loss on the overlapping background regions of paired images via</p><p>where &#8226; denotes an element-wise product. Due to the limited space, we illustrate details of the background loss L bg in the supplementary material.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.">Experiments 4.1. Experimental settings</head><p>Datasets and preprocessing. We evaluate the proposed approach on FFHQ <ref type="bibr">[22]</ref> and two face video datasets, including Faceforensics++ <ref type="bibr">[42]</ref> and VoxCeleb2 <ref type="bibr">[9]</ref>. (i) FFHQ <ref type="bibr">[22]</ref>: We show CoRF learning motion dynamics on FFHQ, which is a single image dataset. (ii) The Faceforen-sics++ dataset <ref type="bibr">[42]</ref> contains 1,000 raw video sequences.</p><p>(iii) The VoxCeleb2 dataset <ref type="bibr">[9]</ref> contains over 1 million videos for 6,112 identities. As videos in these two face datasets <ref type="bibr">[42,</ref><ref type="bibr">9]</ref> contain low-quality frames with small or occluded faces, a pre-processing step is applied to clean the data. Specifically, we detect the faces in the videos using a pretrained landmark detection model <ref type="bibr">[3]</ref>, sharpen faces, filter out frames with small, blurry, or occluded faces, and align the remaining faces. We finally use 990 and 4,804 videos in each dataset <ref type="bibr">[42,</ref><ref type="bibr">9]</ref> for CoRF training.</p><p>Baselines and evaluation metrics. We evaluate CoRF from two perspectives. First, we quantitatively compare our CoRF with face reenactment work <ref type="bibr">[48,</ref><ref type="bibr">55]</ref>. To this end, we use the Fr&#233;chet Inception Distance (FID) <ref type="bibr">[19]</ref> and 3 cosine similarity scores for identity preservation and expression transfer, namely ID, EXP &#8224; and EXP &#8225; . Regarding identity &#8224; Code available from <ref type="url">https://github.com/zllrunning/face-</ref>parsing.PyTorch Source Driving Bi-layer <ref type="bibr">[69]</ref> Ours Figure <ref type="figure">5</ref>: Visual comparison to Bi-layer <ref type="bibr">[69]</ref>.</p><p>preservation, we extract face features using a popular image identity recognition mode pretrained on VGGface2 <ref type="bibr">[4]</ref>. Regarding expression transfer, for a fair evaluation, we use another pretrained facial expression estimation network <ref type="bibr">[10]</ref> (EXP &#8225; ), which differs from the encoder <ref type="bibr">[12]</ref> (EXP &#8224; ) used in our model. Second, as CoRF generates never-before-seen dynamic faces from a prior noise distribution, similar to unconditional video synthesis approaches <ref type="bibr">[56,</ref><ref type="bibr">44,</ref><ref type="bibr">63,</ref><ref type="bibr">55]</ref>, we compare CoRF with these baselines using the Fr&#233;chet Video Distance (FVD) <ref type="bibr">[57]</ref> score. The FVD score is a video-level FID score commonly used to examine both visual quality and temporal coherence of synthesized dynamics <ref type="bibr">[63,</ref><ref type="bibr">18]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.2.">Results</head><p>Comparison to face reenactment methods. We present a visual comparison to FOMM <ref type="bibr">[48]</ref>, MRAA <ref type="bibr">[49]</ref>, and Bilayer <ref type="bibr">[69]</ref> in Fig. <ref type="figure">4</ref> and Fig. <ref type="figure">5</ref>. For this, we transfer expressions and poses from real 'driving' images to synthetic source images in four challenging cases, including crossidentity, cross-gender, cross-pose, and driving image occlusion. We observe that the baseline methods struggle in these cases, while CoRF better preserves face geometry and identity. Additionally, we show targeted video generation on the VoxCeleb2 dataset <ref type="bibr">[9]</ref> at a 256 &#215; 256 resolution in Fig. <ref type="figure">6</ref>. Different from the face reenactment baselines <ref type="bibr">[48,</ref><ref type="bibr">49]</ref>, CoRF is capable of controlling more factors such as novel views and field-of-view (FOV) without the requirement of a reference image, e.g., in Fig. <ref type="figure">7</ref>. We show additional results in the supplementary material. We generate head videos (row 2 and 4) conditioned on facial expression information of the driving frames (row 1 and 3) at a 256 &#215; 256 resolution. Note, the synthetic portraits have a similar expression as the driving frame (see highlighted eyes in row 2, and the mouth in row 4).</p><p>In Tab. 1, we quantitatively compare face reenactment methods on 10,000 pairs of source and driving images (variances in the supplementary material). We find CoRF to significantly improve upon the baselines <ref type="bibr">[48,</ref><ref type="bibr">49]</ref> with regards to identity preservation and expression transformation. Complex expression animation. We show complex expression animation in Fig. <ref type="figure">8</ref>. Among all randomly generated images, we observe the results of the proposed method to follow the expressions in the driving image, highlighting the method's capability. Motion interpolation in the latent space. To analyze the latent motion space, we render from multiple views a trajectory of synthetic images conditioned on linear interpolations between two motion representations, as shown in Fig. <ref type="figure">9</ref>. The smooth interpolation between two expressions over viewpoint changes demonstrates that the latent motion space captures semantics and CoRF enables to generate dynamic faces from multiple views with 3D-awareness. Comparison to unconditional video synthesis methods. We present a quantitative temporal coherence comparison in Tab. 2 where we use 10,000 images/videos for numerical evaluation. Tab. 2 shows that CoRF significantly improves video-level (FVD) quality compared to the baselines <ref type="bibr">[63,</ref><ref type="bibr">FaceF [42]</ref> Vox2 We compare our method with FOMM <ref type="bibr">[48]</ref> and MRAA <ref type="bibr">[47]</ref> at a 256 &#215; 256 resolution. We compute the FID score <ref type="bibr">[19]</ref>, the face identity preservation score (ID), and expression similarities (two expression extractors <ref type="bibr">[12]</ref> and <ref type="bibr">[10]</ref> are used, denoted as EXP &#8224; and EXP &#8225; , respectively).</p><p>Driving Result Driving Result Driving Result 55]. In Fig. <ref type="figure">10</ref>, we also show a comparison of unconditional video synthesis methods <ref type="bibr">[63,</ref><ref type="bibr">55]</ref> at a 64 &#215; 64 resolution (size is chosen to be compatible with the baselines <ref type="bibr">[63,</ref><ref type="bibr">55]</ref>). We observe that CoRF produces sharper faces and persistent head geometry over time and space (row 3). Ablation study. In Fig. <ref type="figure">11</ref> we show the motion reconstruction loss curves during training with different motion conditioning strategies. The baseline method concatenates a motion representation m with a noise vector z, as used in prior work <ref type="bibr">[13]</ref> (blue curve). Compared to the baseline method, our motion embedding strategy (green and red curves) accelerates convergence of the motion reconstruction loss L motion . Moreover, we evaluate the effectiveness Table <ref type="table">2</ref>: Quantitative evaluation for video synthesis. We evaluate our method using FVD <ref type="bibr">[57]</ref> scores (&#8595;) and compare to MoCoGAN <ref type="bibr">[56]</ref>, G 3 AN [63], TGANv2 <ref type="bibr">[44]</ref> at the 64 &#215; 64 resolution, and with MOCOGAN-HD <ref type="bibr">[55]</ref> at the 256 &#215; 256 resolution.</p><p>[63]</p><p>[55]</p><p>Ours Figure <ref type="figure">10</ref>: Visual comparison of video synthesis. We compare our method with popular unconditional video synthesis approaches G 3 AN <ref type="bibr">[63]</ref> and MoCoGAN-HD <ref type="bibr">[55]</ref> on VoxCeleb2 <ref type="bibr">[9]</ref>. The appearance of our method's results remains consistent over view and expression changes, while baselines struggle (see frames highlighted with red border).</p><p>of conditioning the discriminator on motion. We find that motion conditioning of the discriminator (green curve) is slightly better than omission of motion conditioning (red curve). Thus, we use a conditional discriminator in all the experiments. Fig. <ref type="figure">12</ref> provides a visual comparison for the ablation settings of the proposed consistency losses, illustrated below the figure (column 1-3). For each setting, we render two synthetic images applying an identical identity representation yet a different motion representation (top and bottom). We find that the identity loss L id , the background loss L bg , and the consistency loss L consist help to preserve the face identity and a stable background over expression and pose changes. The table on the right shows the averaged loss values with or without optimizing the corresponding loss terms, which supports the visual observation.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.">Conclusion</head><p>We propose controllable radiance fields (CoRFs) to address the task of generalizable, 3D-aware, and motioncontrollable synthesis of face dynamics. CoRFs allow not only explicit 3D-aware motion control and arbitrary view-  Figure <ref type="figure">12</ref>: Ablation study for losses. We render paired synthetic images conditioned on the same identity representation but different motion representations (row 1-2) using different ablation settings for identity, background, and consistency losses. In the table, we show averaged values of the losses L id and L bg (column 2-3), without (or with) optimizing them.</p><p>points but also diverse identity generation in photo-realistic quality. Extensive experiments highlight that CoRFs improve performance w.r.t. 3D consistency and identity preservation over pose and motion changes, compared to state-of-the-art methods on benchmark video datasets. We hope this work inspires future work on the challenging task of 3D-aware dynamic face synthesis.</p></div></body>
		</text>
</TEI>
