This paper presents an innovative solution to the challenge of part obsolescence in microelectronics, focusing on the semantic segmentation of PCB X-ray images using deep learning. Addressing the scarcity of annotated datasets, we developed a novel method to synthesize X-ray images of PCBs, employing virtual images with predefined geometries and inherent labeling to eliminate the need for manual annotation. Our approach involves creating realistic synthetic images that mimic actual X-ray projections, enhanced by incorporating noise profiles derived from real X-ray images. Two deep learning networks, based on the U-Net architecture with a VGG-16 backbone, were trained exclusively on these synthetic datasets to segment PCB junctions and traces. The results demonstrate the effectiveness of this synthetic data-driven approach, with the networks achieving high Jaccard indices on real PCB X-ray images. This study not only offers a scalable and cost-effective alternative for dataset generation in microelectronics but also highlights the potential of synthetic data in training models for complex image analysis tasks, suggesting broad applications in various domains where data scarcity is a concern.
more »
« less
This content will become publicly available on July 1, 2027
Synthetic crack generation using dynamic programming and elastic deformation to enhance segmentation of concrete and pavement defects
Abstract Accurate crack detection in concrete and pavement images is critical for infrastructure assessment but is limited by the scarcity of large, consistently annotated datasets. Supervised learning methods are particularly sensitive to data scarcity, often overfitting and generalizing poorly across crack types and imaging conditions. This study proposes a synthetic crack generation framework to augment or partially replace real datasets while reducing annotation effort. Synthetic cracks are generated by tracing minimal and maximal cumulative energy paths on random noise fields using dynamic programming, producing realistic one-pixel-wide crack strands. These are expanded via variable-width morphological dilation and deformed through geometric transformations and elastic deformation to model variations in width, tortuosity, and boundary irregularities across longitudinal, transverse, and shear cracks. The synthetic data trains a filter-based segmentation and connected component classification system rather than an end-to-end model. Over 2.25 million unique samples are generated across diverse scales and geometries. Elastic deformation increases geometric diversity, raising the mean normalized pairwise feature distance from approximately 0.17 to 0.31. Evaluation on Cracks-200, CDLN, and DeepCrack datasets shows performance comparable to human-annotated training data, with F1-scores up to 0.79 and mIoU exceeding 0.80. These results demonstrate that synthetic crack data can effectively supplement or substitute real annotated datasets, reducing annotation effort while preserving segmentation performance.
more »
« less
- Award ID(s):
- 2118061
- PAR ID:
- 10698339
- Publisher / Repository:
- Springer Nature
- Date Published:
- Journal Name:
- Innovative Infrastructure Solutions
- Volume:
- 11
- Issue:
- 7
- ISSN:
- 2364-4176
- Format(s):
- Medium: X
- Sponsoring Org:
- National Science Foundation
More Like this
-
-
Semantic segmentation of medical images is pivotal in applications like disease diagnosis and treatment planning. While deep learning automates this task effectively, it struggles in ultra low-data regimes for the scarcity of annotated segmentation masks. To address this, we propose a generative deep learning framework that produces high-quality image-mask pairs as auxiliary training data. Unlike traditional generative models that separate data generation from model training, ours uses multi-level optimization for end-to-end data generation. This allows segmentation performance to guide the generation process, producing data tailored to improve segmentation outcomes. Our method demonstrates strong generalization across 11 medical image segmentation tasks and 19 datasets, covering various diseases, organs, and modalities. It improves performance by 10–20% (absolute) in both same- and out-of-domain settings and requires 8–20 times less training data than existing approaches. This greatly enhances the feasibility and cost-effectiveness of deep learning in data-limited medical imaging scenarios.more » « less
-
Deep learning models for computer vision require large amounts of annotated data, which is particularly challenging to acquire in animal behavioral research due to resource constraints and ethical considerations. This paper addresses data scarcity in bee monitoring applications by developing a parametric 3D honeybee model for synthetic dataset generation using game engine technology. Our methodology combines a pre-built model with a novel parametrization system, creating a hybrid approach that enables morphological parametrization, paint code customization, and pollen modeling through a unified data structure compatible with the replicANT platform in Unreal Engine. Evaluation of synthetic datasets for bee detection using YOLOv8 demonstrated that the addition of synthetic data to real bases (50-100 images) improved detection accuracy to 0.97 mAP50, significantly outperforming models trained on equivalent real data alone. For bee re-identification tasks using paint codes, synthetic data enabled color estimation with an average color distance of 0.29 in [0,1] normalized RGB space, showing promising transfer from synthetic to real domains despite remaining domain gaps. Our results demonstrate that high-quality 3D models can serve as effective pre-training foundations, reducing real data requirements while maintaining competitive performance, advancing ecological monitoring through scalable synthetic data generation.more » « less
-
Abstract Advances in electron microscopy, image segmentation and computational infrastructure have given rise to large-scale and richly annotated connectomic datasets, which are increasingly shared across communities. To enable collaboration, users need to be able to concurrently create annotations and correct errors in the automated segmentation by proofreading. In large datasets, every proofreading edit relabels cell identities of millions of voxels and thousands of annotations like synapses. For analysis, users require immediate and reproducible access to this changing and expanding data landscape. Here we present the Connectome Annotation Versioning Engine (CAVE), a computational infrastructure that provides scalable solutions for proofreading and flexible annotation support for fast analysis queries at arbitrary time points. Deployed as a suite of web services, CAVE empowers distributed communities to perform reproducible connectome analysis in up to petascale datasets (~1 mm3) while proofreading and annotating is ongoing.more » « less
-
Abstract Advances in Electron Microscopy, image segmentation and computational infrastructure have given rise to large-scale and richly annotated connectomic datasets which are increasingly shared across communities. To enable collaboration, users need to be able to concurrently create new annotations and correct errors in the automated segmentation by proofreading. In large datasets, every proofreading edit relabels cell identities of millions of voxels and thousands of annotations like synapses. For analysis, users require immediate and reproducible access to this constantly changing and expanding data landscape. Here, we present the Connectome Annotation Versioning Engine (CAVE), a computational infrastructure for immediate and reproducible connectome analysis in up-to petascale datasets (∼1mm3) while proofreading and annotating is ongoing. For segmentation, CAVE provides a distributed proofreading infrastructure for continuous versioning of large reconstructions. Annotations in CAVE are defined by locations such that they can be quickly assigned to the underlying segment which enables fast analysis queries of CAVE’s data for arbitrary time points. CAVE supports schematized, extensible annotations, so that researchers can readily design novel annotation types. CAVE is already used for many connectomics datasets, including the largest datasets available to date.more » « less
An official website of the United States government
