View Translation: Geometry-Guided Latent Diffusion for Cross-View Image Synthesis on Ground Vehicles
Synthesizing novel camera views is important for autonomous ground vehicles, with applications in surround-view monitoring, occlusion recovery, and training data augmentation. We present {View Translation}, a geometry-guided latent diffusion framework that generates a target camera view from a source image, relative camera pose, and an available target-view depth prior. The method combines three components: a Vector Quantized Variational Autoencoder for compact latent encoding, a depth-based warping module that projects the source image into the target view to provide geometric guidance, and a ControlNet-augmented denoising UNet conditioned on source appearance, relative pose, and an auxiliary Image-Depth fusion network. Evaluated on KITTI and a simulated off-road dataset, our method achieves competitive FID while improving LPIPS and PSNR over baseline approaches, supporting cross-view synthesis for ground vehicle perception.
more »
« less
An official website of the United States government

