NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

WALT3D: Generating Realistic Training Data from Time-Lapse Imagery for Reconstructing Dynamic Objects Under Occlusion

https://doi.org/10.1109/CVPR52733.2024.00909

Vuong, Khiem; Reddy, N Dinesh; Tamburo, Robert; Narasimhan, Srinivasa G (June 2024, IEEE)

Full Text Available
Learned Two-Plane Perspective Prior based Image Resampling for Efficient Object Detection

https://doi.org/10.1109/CVPR52729.2023.01284

Ghosh, Anurag; Reddy, N. Dinesh; Mertz, Christoph; Narasimhan, Srinivasa G. (June 2023, IEEE Conference on Computer Vision and Pattern Recognition)

Full Text Available
Traffic4D: Single View Longitudinal 4D Reconstruction of Repetitious Activity using Self-Supervised Experts

Li, Fangyu; Reddy, N. Dinesh; Chen, Xudong; Narasimhan, Srinivasa G. (July 2021, IEEE Intelligent Vehicles Symposium)

Reconstructing 4D vehicular activity (3D space and time) from cameras is useful for autonomous vehicles, commuters and local authorities to plan for smarter and safer cities. Traffic is inherently repetitious over long periods, yet current deep learning-based 3D reconstruction methods have not considered such repetitions and have difficulty generalizing to new intersection-installed cameras. We present a novel approach exploiting longitudinal (long-term) repetitious motion as self-supervision to reconstruct 3D vehicular activity from a video captured by a single fixed camera. Starting from off-the-shelf 2D keypoint detections, our algorithm optimizes 3D vehicle shapes and poses, and then clusters their trajectories in 3D space. The 2D keypoints and trajectory clusters accumulated over long-term are later used to improve the 2D and 3D keypoints via self-supervision without any human annotation. Our method improves reconstruction accuracy over state of the art on scenes with a significant visual difference from the keypoint detector’s training data, and has many applications including velocity estimation, anomaly detection and vehicle counting. We demonstrate results on traffic videos captured at multiple city intersections, collected using our smartphones, YouTube, and other public datasets.
more » « less
Full Text Available
TesseTrack: End-to-End Learnable Multi-Person Articulated 3D Pose Tracking

Reddy, N Dinesh; Guigues, Laurent; Pischulini, Leonid; Eledath, Jayan; Narasimhan, Srinivasa (June 2021, International Conference on Computer Vision and Pattern Recognition (CVPR))

We consider the task of 3D pose estimation and tracking of multiple people seen in an arbitrary number of camera feeds. We propose TesseTrack, a novel top-down approach that simultaneously reasons about multiple individuals’ 3D body joint reconstructions and associations in space and time in a single end-to-end learnable framework. At the core of our approach is a novel spatio-temporal formulation that operates in a common voxelized feature space aggregated from single- or multiple camera views. After a person detection step, a 4D CNN produces short-term persons pecific representations which are then linked across time by a differentiable matcher. The linked descriptions are then merged and deconvolved into 3D poses. This joint spatio-temporal formulation contrasts with previous piecewise strategies that treat 2D pose estimation, 2D-to-3D lifting, and 3D pose tracking as independent sub-problems that are error-prone when solved in isolation. Furthermore, unlike previous methods, TesseTrack is robust to changes in the number of camera views and achieves very good results even if a single view is available at inference time. Quantitative evaluation of 3D pose reconstruction accuracy on standard benchmarks shows significant improvements over the state of the art. Evaluation of multi-person articulated 3D pose tracking in our novel evaluation framework demonstrates the superiority of TesseTrack over strong baselines.
more » « less
Full Text Available

Search for: All records