<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>Concrete spalling detection system based on semantic segmentation using deep architectures</title></titleStmt>
			<publicationStmt>
				<publisher>Elsevier</publisher>
				<date>08/01/2024</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10509067</idno>
					<idno type="doi">10.1016/j.compstruc.2024.107398</idno>
					<title level='j'>Computers &amp; Structures</title>
<idno>0045-7949</idno>
<biblScope unit="volume">300</biblScope>
<biblScope unit="issue">C</biblScope>					

					<author>Tamanna Yasmin</author><author>Duc La</author><author>Kien La</author><author>Minh Tuan Nguyen</author><author>Hung Manh La</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[This paper presents a method for detecting the location of spalling and assessing the severity level of the spalling in concrete surfaces. The proposed method is constructed based on deep learning architectures and multi-class semantic segmentation. The proposed method can detect each pixel as a non-spalling, a deepspalling, or a shallow-spalling. The proposed method consists of three dierent deep learning architectures with several encoders as backbone networks. Both qualitative and quantitative analyses show that the deep learning architecture with a certain encoder network can detect spalling with dierent severity levels very well. Additionally, the paper proposes a method to analyze the deep spalling areas of concrete to show their severity levels. The performance analysis shows that this approach provides very convincing results with respect to the actual aected spalling areas. The results convey that this paper achieved a higher level of performance for detecting spalling and assessing the severity of the spalling.]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.">Introduction</head><p>Concrete distress, such as spalling, poses life-threatening risks, necessitating regular maintenance to prevent hazardous incidents <ref type="bibr">[1]</ref> <ref type="bibr">[2]</ref>. Spalling, a concrete abnormality caused by heavy loads and surface erosion, compromises structural integrity in bridges and buildings. It is required to conduct regular inspections to ensure structural integrity <ref type="bibr">[3]</ref><ref type="bibr">[4]</ref><ref type="bibr">[5]</ref>. To address this issue, autonomous detection systems are now mandatory.</p><p>To manage spalling e&#57344;ectively, precise measurements and categorization are essential. Depending on spalling conditions, post-inspection actions vary. Priority is given to large or deep spalling, while smaller or shallower instances may follow. Detecting spalling alone is insu&#57345;cient; assessing severity (based on size <ref type="bibr">[6]</ref> or depth) is crucial to prioritize repairs. Spalling severity ranges from deep (high risk) to shallow or non-existent. Deep spalling signi&#57436;cantly impacts structural health.</p><p>Therefore, this paper introduces spalling detection and categorization: deep, shallow, and non-spalling. We also propose a method for ranking the severity among deep spalling areas. The severity levels of spalling are shown in Fig. <ref type="figure">1</ref>.</p><p>The rest of the paper is organized as follows: Section 1 also presents the related works and our contribution. Section 2 outlines the research methodology and our proposed work. Section 3 analyzes the results. Finally, the conclusion and future work of this paper are given in section 4.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.1.">Literature review</head><p>In recent years, there has been an increasing interest in spalling detection techniques for various infrastructure applications such as metro tunnels, subway networks, bridges, and railway surfaces. These methods often utilize machine vision, laser scanning, deep learning, and infrared thermography to detect and evaluate spalling.</p><p>One approach proposed for detecting spalling in subway networks is based on image processing <ref type="bibr">[7]</ref>. The color image of the spalled area is processed to remove noise and extract surface features. A 3D visualization is created from the extracted features, and spalling severity and depth are detected using a projection of the spalling intensity curve and regression analysis.  Another approach suggested for metro tunnels utilizes surface roughness analysis based on a 3D mobile laser scanning system <ref type="bibr">[8]</ref>. Point cloud data obtained from the scanning system is used to analyze the surface roughness, which is then used to detect concrete spalling.</p><p>A machine learning and vision-based approach has been developed for subways to detect and quantify spalling <ref type="bibr">[9]</ref>. This approach involves extracting important features from images, removing noise, and detecting surface distresses in subways.</p><p>To detect concrete spalling automatically in multiple spots (within a single structural element or in multiple structural elements), a deep learning-based method has been developed <ref type="bibr">[10]</ref>. In this work, an inexpensive depth camera is integrated with a faster region-based convolutional neural network (Faster R-CNN) to automatically detect, localize, and quantify the spalling damage.</p><p>Another deep learning-based real-time multi-drone approach has been proposed to detect surface defects in high-rise civil structures <ref type="bibr">[11]</ref>. To detect &#57436;ve types of concrete surface defect images of two classes (crack and spalling): vertical crack, horizontal crack, diagonal crack, branch crack, and spalling, the authors have used the deep learning model YOLO-v3 (You Look Only Once-version3) and the edge computing principle.</p><p>A morphological attention ensemble learning method for surface defect detection at the bounding box level is proposed to detect three types of defects (crack, e&#57346;orescence, and spalling) <ref type="bibr">[12,</ref><ref type="bibr">13]</ref>. The authors propose a specialized loss for each defect to demonstrate improved defect-recognition accuracy. Moreover, the deformable convolutional network (DCN) <ref type="bibr">[14]</ref> and multi-task ensemble learning techniques have been exploited to adaptively extract defect features according to the feature shapes and to apply the loss of each defect, respectively.</p><p>Another approach based on Faster RCNN is proposed to detect four di&#57344;erent damage types: surface cracks, spalling (including fa&#231;ade spalling and concrete spalling), and severe damage with exposed rebars and severely buckled rebars <ref type="bibr">[15]</ref>. Since the purpose of their study is to present a timely assessment of defects and damages that occurred due to an earthquake to buildings, the authors manually evaluated the approach using annotated image data collected from damaged concrete buildings during several past earthquakes.</p><p>For rail surface spalling detection, a real-time visual inspection system has been proposed that utilizes image acquisition and image processing sub-systems <ref type="bibr">[16]</ref>. Images captured by a camera are segmented, and spalling on the rail surface is detected using histogram curve information in the longitudinal direction of the track image.</p><p>An optical detection algorithm based on visual salience has been proposed for rail surface spalling detection <ref type="bibr">[17]</ref>. This algorithm uses a threshold value to detect the di&#57344;erence between spalled and non-spalled areas after removing unnecessary noises from the neighborhood area of spalling.</p><p>A novel automated 3D spalling defects inspection system for railway tunnel linings has been proposed that uses laser intensity and depth information for accurate spalling detection <ref type="bibr">[18]</ref>. A spalling intensity depurator network is also proposed for automatic feature extraction, and the system produces 3D inspection results with quantitative analysis of the spalled area.</p><p>Deep learning approaches have also been developed for automatic detection of cracks and spalling in buildings and bridges <ref type="bibr">[19]</ref>. These approaches utilize Mask R-CNNs for continuous segmentation of damaged areas in bridges and buildings, and the deep CNN architectures can be extended for surface damage detection and evaluation.</p><p>To automatically detect concrete spalling, image texture and piecewise linear stochastic gradient descent logistic regression are used for pattern recognition <ref type="bibr">[20]</ref>. Image textures are extracted from images, and statistical properties are used to categorize the condition of the concrete surface into non-spalled and spalled classes.</p><p>A terrestrial laser scanner was utilized in this study to simultaneously localize and quantify spalling defects on concrete surfaces, as reported by <ref type="bibr">Kim et al. in 2015 [21]</ref>. The proposed method combines features with complementary properties to enhance the localization and quanti&#57436;cation of spalling defects. To extract relevant information, such as the condition and size of the damaged portion of the concrete surface, a defect classi&#57436;er was developed. The concrete structure was scanned using a terrestrial laser scanner, and a region of interest was selected for analysis. The scanner captured 3D coordinate information of the scanned points within the selected region. Once the raw scanned data was ready, the proposed method initiated the defect detection process.</p><p>Spalling in concrete structures can happen during &#57436;re conditions, as reported by <ref type="bibr">Kodur et al. in 2021 [22]</ref>. The proposed approach considers factors like pore pressure, thermal gradients, and structural loading as contributors to spalling. Comparing the approach's predictions to experimental data from full-scale &#57436;re resistance tests on concrete beams of di&#57344;erent strengths, the analysis reveals that pore pressure-induced stresses are the primary cause of spalling. However, thermal and mechanical stress levels also play a role in spalling. The extent of spalling signi&#57436;cantly a&#57344;ects the &#57436;re resistance of concrete beams in severe &#57436;re scenarios.</p><p>The method proposed by Naser et al. in 2019 <ref type="bibr">[23]</ref> is based on Machine Cognition (MC) to obtain expression in order to detect defects in concrete structures due to &#57436;re conditions. These expressions consider the geometric attributes, material composition, and distinct characteristics of reinforced concrete (RC) columns. Their purpose is to predict the occurrence and intensity of &#57436;re-induced spalling and assess the &#57436;re resistance of these structural components.</p><p>Another approach to identifying essential factors that in&#57437;uence the occurrence of &#57436;re-induced spalling in RC columns o&#57344;ers an exploration of how data science and machine learning techniques can be employed <ref type="bibr">[24]</ref>. Nine distinct algorithms (naive Bayes, generalized linear model, logistic regression, fast large margin, deep learning, decision tree, random forest, gradient boosted trees, and support vector machine) have been used for this study to examine data collected from 185 &#57436;re experiments. These algorithms were similarly employed to pinpoint the essential attributes in&#57437;uencing the likelihood of &#57436;re-induced spalling in RC columns and to create tools for immediate spalling prediction.</p><p>The study focused on investigating the e&#57344;ects of di&#57344;erent types and sizes of specimens on concrete spalling when exposed to a hydrocarbon &#57436;re, as reported by Mohd et al. in 2018 <ref type="bibr">[25]</ref>. Four di&#57344;erent types of specimen sizes, including cylinders, columns, and panels, were analyzed to isolate the variables that a&#57344;ect concrete spalling. Additionally, three aggregate sizes were used in the concrete mixes to determine the impact of aggregate size on concrete spalling. The investigation also included analyzing the e&#57344;ect of aggregate type on concrete spalling.</p><p>Concrete spalling detection can be achieved using active infrared thermography, as reported by <ref type="bibr">Tanaka et al. in 2006 [26]</ref>. Various irradiation devices, such as halogen lamps, xenon arc lamps, and farinfrared irradiation devices, can be used to heat the concrete surface and create a temperature gradient for detecting spalling. Active infrared thermography was chosen in this study due to its ability to provide measurements independent of meteorological conditions. Photogrammetry, laser scanning, and Light Detection and Ranging (LiDAR) are technologies commonly used for surface damage detection, including spalling, as reported by Zhang et al. in 2022 <ref type="bibr">[27]</ref>. In the proposed method, point cloud data from the laser scanner was utilized to detect spalling and quantify its key properties in reinforced concrete columns. The &#57436;rst phase of the method involved removing noise points and calibrating the coordinate system of the captured point cloud data, which was then sliced into thin layers for analysis of the damaged areas. The second phase included detecting points corresponding to distressed and nondistressed areas. Finally, linear interpolation was used to calculate the spalling area and lost concrete volume.</p><p>A computer application has been developed to automatically evaluate spalling and detect spalling severity in concrete bridges <ref type="bibr">[28]</ref>. The proposed approach utilizes a single-objective particle swarm optimization model based on the Tsallis entropy function to detect spalling. In the second phase, the severity of spalling is evaluated by generating a comprehensive analysis of the bridge deck image using the Daubechies discrete wavelet transform feature description algorithm. A hybrid arti-&#57436;cial neural network-particle swarm optimization model is used in the second phase to accurately predict the spalling area and overcome the limitations of the gradient descent algorithm.</p><p>The timely and accurate detection of spalling and its severity is critical, and computer vision plays a vital role in this context by extracting numerical information from various sources such as depth images, digital images, videos, and 3D point clouds, processing the data, and taking appropriate actions <ref type="bibr">[29]</ref>. A computer vision-based approach for classifying concrete spalling severity has been developed <ref type="bibr">[30]</ref>. This method utilizes concrete images and categorizes the severity levels as shallow spall or deep spall. Features of the concrete surface, including statistical measurements of color channels, gray-level run length, and centersymmetric local binary pattern, are used to optimize the support vector machine classi&#57436;er using the jelly&#57436;sh search metaheuristic to divide the data into shallow spalling and deep spalling based on a decision boundary.</p><p>An entropy-based automated method has been proposed, which consists of three signi&#57436;cant parts and utilizes computer vision technologies for spalling detection <ref type="bibr">[31]</ref>. The spalling detection phase employs a segmentation model that integrates a multi-objective invasive weed optimization and information theory-based formalism of images. The feature extraction phase combines singular value decomposition and discrete wavelet transform to obtain e&#57345;cient image information. The third phase involves developing a rating system for spalling severity based on its area and depth.</p><p>The study introduces a method for identifying spalling damage using point cloud data and incorporating the damaged elements into a Building Information Model (BIM) by enhancing the as-built Industry Foundation Classes (IFC) model with semantic information <ref type="bibr">[32]</ref>. The authors present a methodology for creating the as-built BIM, recon-structing the geometric properties of identi&#57436;ed damage point clusters, and enhancing the associated IFC model with semantic information.</p><p>A deep neural network called MaDnet (material-and-damagenetwork), which is designed to perform dual tasks, it can identify both the material type (concrete, steel, asphalt) and distinguish between &#57436;ne (cracks, exposed rebar) and coarse (spalling, corrosion) structural damage simultaneously <ref type="bibr">[33]</ref>. The authors utilize semantic segmentation, which involves assigning material and damage labels to individual pixels in the image. The connection between material and damage is integrated by training shared &#57436;lters using a multi-objective optimization approach.</p><p>Computer vision approaches o&#57344;er solutions for detecting spalling severity on concrete surfaces <ref type="bibr">[34]</ref>. The proposed work utilizes Extreme Gradient Boosting Machine and Deep Convolutional Neural Network (DCNN) to classify image data into shallow spall and deep spall. Feature extraction methods such as local binary pattern, center symmetric local binary pattern, local ternary pattern, and attractive repulsive center symmetric local binary pattern are used to extract properties of spalling from the concrete surface image. The prediction performance of the Extreme Gradient Boosting Machine is enhanced using the Aquila optimizer metaheuristic.</p><p>In summary, various approaches have been proposed for detecting spalling, or surface damage, in subway networks, metro tunnels, rail surfaces, buildings, and bridges. These approaches utilize image processing, laser scanning, machine learning, and other techniques for extracting features, removing noise, and detecting spalling severity and depth. Deep learning approaches, such as Mask R-CNNs, have been used for continuous segmentation of damaged areas in bridges and buildings. Terrestrial laser scanners have been utilized for localization and quanti&#57436;cation of spalling defects on concrete surfaces. Active infrared thermography has been used for concrete spalling detection, and photogrammetry, laser scanning, and LiDAR have been used for surface damage detection. Computer applications and optimization models have also been developed for automated evaluation of spalling severity. These approaches aim to provide accurate and e&#57345;cient detection of spalling in various structures, aiding in maintenance and repair e&#57344;orts. However, there is a limited number of methods that focus on identifying and categorizing the severity of concrete spalling. We have provided a contrasting analysis between our proposed work and previous works in Table <ref type="table">1</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.2.">Contributions</head><p>Ensuring the structural health of concrete is crucial for maintaining the wellness of civil infrastructure. Therefore, detecting spalling and classifying its severity level has a signi&#57436;cant impact on achieving this goal. Although there are several methods for detecting spalling, there are few approaches for classifying its severity level. This is important because it helps prioritize spalling maintenance, especially in critical areas.</p><p>Previous approaches for spalling detection have some limitations, such as that "non-spalling" areas were not categorized as a level of severity. To properly identify distressed surfaces and non-a&#57344;ected areas, these "non-spalling" areas should be included in the severity level. Additionally, spalling classes should be discretely segmented with proper visual mapping based on severity level. Thus, the classi&#57436;cation of severity can be measured by how deep or shallow the spalling is, or whether there is no spalling at all. As a result of the bene&#57436;ts of image segmentation techniques in various &#57436;elds, we considered this problem to be one of semantic segmentation.</p><p>To overcome these limitations, we propose a method for detecting and classifying spalling severity levels using deep architecture and encoder-decoder networks. Our approach uses pixel-by-pixel multiclass semantic segmentation to categorize spalling as non-spalling, shallow, or deep. We conducted a comparative analysis to determine the best combination of deep architecture and encoder-decoder networks. Ac-T. Yasmin, D. La, K. La et al. Hong et al. <ref type="bibr">[12]</ref> Ghosh et al. <ref type="bibr">[15]</ref> Kim et al. <ref type="bibr">[21]</ref> Tanaka et al. <ref type="bibr">[26]</ref> Zhang et al. <ref type="bibr">[27]</ref> Concrete surface</p><p>Kumar et al. <ref type="bibr">[11]</ref> Bai et al. <ref type="bibr">[19]</ref> Abdelkader et al. <ref type="bibr">[28]</ref> Isailovic et al. <ref type="bibr">[32]</ref> High-rise civil structure</p><p>Pham et al. <ref type="bibr">[16]</ref> Rail Surface</p><p>Naser et al. <ref type="bibr">[23]</ref> Naser et al. <ref type="bibr">[24]</ref> Mohd et al. <ref type="bibr">[24]</ref> Concrete surface under &#57436;re condition</p><p>Hoang et al. <ref type="bibr">[30]</ref> Nguyen et al. cording to the severity level, deep spallings are more crucial since they a&#57344;ect the concrete surfaces more alarmingly. They are crucial and need further analysis. The a&#57344;ected area of a deep spalling may vary according to di&#57344;erent sizes. We also proposed a pixel-wise severity ranking method to calculate the ranking of severity for deep spalling areas.</p><p>Our proposed approach is designed to identify distressed surfaces and non-a&#57344;ected areas accurately and to classify the severity of spalling more precisely. It o&#57344;ers several contributions, including a deep architecture-based method with di&#57344;erent backbone networks, multi-class segmentation using pixel-by-pixel categorization, a pixelwise severity ranking method, and qualitative and quantitative analysis to obtain the best results.</p><p>Overall, our proposed method provides an e&#57344;ective solution for detecting and classifying spalling severity levels. This will help engineers and maintenance personnel to prioritize and plan spalling maintenance activities more e&#57345;ciently, resulting in improved infrastructure wellness and durability.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.">Research methodology</head><p>In this section, we have presented a comprehensive analysis of the di&#57344;erent aspects of our proposed method for detecting spalling and severity levels. The approach for detecting spalling and spalling severity levels is based on Deep encoder-decoder networks. In recent years, several encoder-decoder-based deep convolutional networks have been proposed; SegNet <ref type="bibr">[35]</ref>, UNet <ref type="bibr">[36]</ref>, PSPNet <ref type="bibr">[37]</ref>, FCN <ref type="bibr">[38]</ref>, DeepLab <ref type="bibr">[39]</ref>, DeepCrack <ref type="bibr">[40]</ref>. We have selected SegNet, PSPNet, and UNet for our proposed architecture. Our proposed approach delineated the use of SegNet, PSPNet, and UNet along with variations in the backbone networks to predict the best deep architecture-backbone network pair. A comparative analysis of the performance achieved from the three deep architectures in terms of di&#57344;erent performance metrics has been discussed in the results and discussion section. We have employed di&#57344;erent encoder modules leveraged within the context of the di&#57344;erent deep architectures including ResNet-50 <ref type="bibr">[41]</ref>[42], VGG-19 <ref type="bibr">[43]</ref>, Xception <ref type="bibr">[44]</ref>, and MobileNet <ref type="bibr">[45]</ref>.</p><p>We have discussed several concepts related to our proposed approach. First, several deep encoder-decoder-based architectures will be discussed. The preparation of datasets and the data augmentation process will be included in this section. Along with these discussions, our proposed methodology for detecting spalling and severity levels using deep encoder-decoder networks will be outlined.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.1.">Deep encoder-decoder architecture</head><p>SegNet: The SegNet is an encoder-decoder network based architecture <ref type="bibr">[35]</ref>. SegNet architecture-based image segmentation process has been used to extract abnormal skin lesions from dermoscopy image <ref type="bibr">[46]</ref>, for gland segmentation from colon cancer histology images <ref type="bibr">[47]</ref>, to detect dark spots in oil spill areas <ref type="bibr">[48]</ref>, for automated brain tumor segmentation on multi-modal MR image <ref type="bibr">[49]</ref>, to detect pixel level crack detection <ref type="bibr">[50]</ref>, and for the inspection and evaluation of bridge decks <ref type="bibr">[51]</ref>.</p><p>This architecture was proposed for pixel-wise semantic segmentation. The architecture for SegNet with encoder-decoder block is shown in Fig. <ref type="figure">2</ref>. The encoder block of SegNet architecture contains 13 convolutional layers for feature maps which leads to object classi&#57436;cation. The dense convolutions, ReLU non-linearity, and a non-overlapping max-pooling are performed by encoder block <ref type="bibr">[52]</ref>. The max pooling is performed with a (2 &#215; 2) window. The SegNet architecture avoids the fully connected layers to gain higher-resolution feature maps at the deepest encoder output. The down-sampling is the &#57436;nal step of the encoder. In the decoder block up-sampling and convolutions are performed <ref type="bibr">[35]</ref>. The decoder conducts the up-sampling and calls the max pooling indices of the corresponding encoder layer. There is a K-class softmax classi&#57436;er at the end to predict the class for each pixel.   UNet: Several works used the UNet architecture for image segmentation; brain tumor image segmentation using UNet <ref type="bibr">[53]</ref> and UNet-VGG16 <ref type="bibr">[54]</ref>, COVID-19 lung CT image segmentation <ref type="bibr">[55]</ref>, crack detection model <ref type="bibr">[56]</ref>, and dental panoramic image segmentation <ref type="bibr">[57]</ref>.</p><p>UNet is an encoder-decoder-based architecture consisting of four encoder and four decoder blocks. Fig. <ref type="figure">3</ref> shows the overview of UNet architecture. The encoder block contains two 3 &#215; 3 convolutions <ref type="bibr">[36]</ref>.</p><p>A ReLU activation function comes after each convolution. The encoder component of the UNet architecture functions as an image feature extractor and gathers the image's features. Each encoder block's number of feature channels is doubled and its spatial dimensions are cut in half by the encoder network. A link connects the encoder blocks and decoder blocks together. The resulting output of the encoder blocks' ReLU activation function connects to the matching decoder blocks. Two (3 &#215; 3) convolutions are used in the connection between the encoder and decoder blocks, and each convolution is followed by a ReLU activation function. By providing supplementary information, this connection enables the decoder to build stronger semantic features. The decoder network has half the number of feature channels and doubles the spatial dimensions. A (2 &#215; 2) transpose convolution is present in the decoder's initial stage. Using a concatenation process of convolution and connection, the feature maps are transferred through the connection between the encoder and decoder. A segmentation mask is created in the decoder section. A (1 &#215; 1) convolution with sigmoid activation is applied to the output generated by the &#57436;nal decoder. Using an activation function, the segmentation mask is transformed into pixel-wise categorization.</p><p>PSPNet: The architectural overview of PSPNet is shown in Fig. <ref type="figure">4</ref>. PSPNet is one of the most well-recognized image segmentation models. PSPNet-based semantic segmentation process used in image semantic segmentation <ref type="bibr">[59]</ref>, pavement distress detection <ref type="bibr">[60]</ref> and crack detection <ref type="bibr">[61]</ref>[62], arms and hands segmentation for egocentric perspective using image segmentation <ref type="bibr">[63]</ref>, and image segmentation for coronary angiography <ref type="bibr">[64]</ref>. This architecture has two blocks like most semantic segmentation models: PSPNet encoder and PSPNet decoder. The PSPNet encoder consists of the CNN backbone with dilated convolutions <ref type="bibr">[65]</ref> and the pyramid pooling module. Dilated convolution layers are used in place of the typical convolutional layers in the backbone's last layers, which helps to increase the receptive &#57436;eld. The last two blocks of the backbone contain these dilated convolution layers. As a result, the feature that is added at the end of the backbone has more features. During convolution, the value of dilation indicates the sparsity. In comparison to standard convolution, dilated convolution has a broader receptive &#57436;eld. The size of the used context information is found from the size of the receptive &#57436;eld. The pyramid pooling module is the primary component of this model since it enables the model to recognize the global context in the image and classify the pixels according to that context. The backbone's feature map is pooled at di&#57344;erent sizes, passed through a convolution layer, and then upsampled to bring the pooled features up to the size of the original feature map. The original feature map and the upsampled maps are &#57436;nally concatenated before being sent to the decoder. This method aggregates the overall context by fusing the information at di&#57344;erent scales. The decoder will then take those features and turn them into predictions by feeding them into its layers once the encoder has extracted the image's features. The decoder is another network that processes inputted characteristics to provide predictions.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.2.">Backbone network</head><p>Several CNN's (convolutional neural network) backbone networks have made signi&#57436;cant advancements with the highest quality performances over the past few years. These network architectures may effectively extract an image's feature mapping, providing a strong base network for semantic segmentation <ref type="bibr">[66]</ref>. We have used ResNet-50, VGG-19, MobileNet, and Xception as feature extractors in our deep architectures. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.3.">Preparation of dataset</head><p>Large volumes of data are required to train, validate, and test the models in deep network architectures <ref type="bibr">[76]</ref>. As a result, managing a well-balanced dataset is a crucial step. For our proposed method, we have collected images of di&#57344;erent buildings and bridges. For bridge data collection, we mainly used our developed robots integrated with nondestructive evaluation sensors and cameras to collect the bridge deck surface images <ref type="bibr">[77]</ref><ref type="bibr">[78]</ref><ref type="bibr">[79]</ref><ref type="bibr">[80]</ref><ref type="bibr">[81]</ref>. We collected images at di&#57344;erent times of the day to maintain the non-uniformity of the environment. We have employed a data augmentation procedure for our proposed architecture to help with the data management issue. We assigned the labels of nonspalling, deep spalling, and shallow spalling, along with the labels of severity, to each image in our collection. As a result, during the training process, pixel mapping is automatically generated from the labeling of the image.</p><p>The non-spalling area and the severity levels of spalling are annotated with RGB combinations because the labeling of images follows the RGB range. The process of annotating images is an arduous and time-consuming task <ref type="bibr">[82]</ref>. We have annotated each image according to the spalling area; Deep spalling area, shallow spalling area, and nonspalling area. The example of the annotation process of an image is shown in Fig. <ref type="figure">5</ref>.</p><p>We have prepared a collection of images for each original image and respective annotated image using the data augmentation procedure. The augmentation procedure chooses a random picture for each image as well as a random pixel point for the tagged image. A selected image and pixel map of the relevant original image is made in accordance with that. The augmentation method generates a number of sub-images at random from the pixel locations by &#57437;ipping or rotating the pixel map. The augmentation process of a sample image is shown in Fig. <ref type="figure">6</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.4.">Proposed architecture</head><p>We have proposed the method using three di&#57344;erent types of deep architectures with di&#57344;erent backbone networks for detecting the spalling  <ref type="figure">4</ref>), UNet (Fig. <ref type="figure">3</ref>), and SegNet (Fig. <ref type="figure">2</ref>).</p><p>For the encoder part, we have employed ResNet-50, VGG-19, Xception, and MobileNet. The encoder block has convolution and pooling layers. A set of down-sampled feature maps are produced by each part of the encoder using an input picture or feature map. The pooling layers help the encoder to form integrated feature points after the feature is extracted from the convolution layers.</p><p>The decoder is essentially a mirrored encoder. The di&#57344;erence between the decoder block and encoder block is the up-sampling layer instead of the pooling layers. It gradually upsamples the encoder's output and semantically projects into high-resolution pixel space from the low-resolution identi&#57436;able feature maps.</p><p>The advantage of employing deep learning-based image segmentation architecture is, the segmentation model di&#57344;erentiate each pixel at the pixel level as well as projects the features with the di&#57344;erent category at various stages into the pixel space learned by the encoder to fully segment the target region <ref type="bibr">[83]</ref>. Moreover, using the concatenation process the decoder connects to the corresponding encoder and helps to reduce the loss that happened during the down-sampling process. Therefore, due to the advantage and performance of deep learning-based image segmentation architecture in several &#57436;elds <ref type="bibr">[84]</ref>[85] <ref type="bibr">[86]</ref>, we have proposed the use of deep architectures with encoder-decoder networks to detect spalling and severity level. We have considered the spalling and severity detection process as multi-class image segmentation. Therefore, as the output or semantic segmentation of the given data from the encoder-decoder network, we get the segmented area of spalling; non-   spalling, deep spalling, or shallow spalling. Fig. <ref type="figure">7</ref> shows the overview of the proposed methodology to detect the spalling severity levels.</p><p>For the deep architecture-based proposed method, we have annotated the images of deep spalling based on the exposed reinforcing steel bars. The annotated images are used to detect the spalling severity level. These deep spalling areas based on the reinforcing steel bars can be large, very large, or small. Along with the depth, the ratio of the affected deep spall areas helps provide more insight into the severity. For that reason, we have proposed a method to calculate the ranking of severity for the deep spalling area. This proposed method determines the a&#57344;ected deep spalling area using pixel-wise calculation. Afterwards, the ranking of deep spalling areas is determined according to the ratio of the a&#57344;ected area (number of a&#57344;ected pixels) with respect to the overall area (number of total pixels). We have determined the ratio using Equation <ref type="bibr">(1)</ref>, where &#57344;_&#57345;&#57346;&#57347;&#57348;&#57349; provides the number of total pixels of deep spalling area, and &#57350; _&#57345;&#57346;&#57347;&#57348;&#57349; counts the total number of pixels for the entire image. The ranking of deep spalling areas is categorized as "very severe," "medium severe," and "less severe" based on the value of the ratio using a prede&#57436;ned threshold.</p><p>Hence, we &#57436;rst detected spalling and its severity. Moreover, we discussed the comparative analysis of the performances achieved by the deep architectures. Using pixel-wise calculation, we determined the severity ranking for deep spalling areas. We have provided a comparative analysis of severity ranking for three selected image categories in the Result and Discussion section.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.">Result and discussion</head><p>In this section, we are going to present the performance analysis of the proposed deep architecture with di&#57344;erent encoder-decoder networks. The results and experimental analysis part includes dataset preparations, experimental setup, and qualitative and quantitative analysis of proposed architectures.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.1.">Experimental setup</head><p>The dataset contains images of spalling in buildings and bridges. These images have di&#57344;erent types of noises, namely, oil spills, faded colors, and stones. It is di&#57345;cult to detect any abnormality in the crucial corners of bridges, for example, the intersection of pillars, due to  the di&#57344;erence in light. The images were taken at di&#57344;erent times of the day to avoid any impact of light and shadow on the result of spalling detection.</p><p>We have used GIMP (GNU Image Manipulation Program) to annotate our images in the pixel-by-pixel map. GIMP is one of the most popular illustration and image editing programs available <ref type="bibr">[87]</ref>. We have annotated each image according to the spall class. The reason behind the image size (256 &#215; 256) is to focus on the speci&#57436;c spall class with any noises or light di&#57344;erences. Our dataset contains di&#57344;erent categories of images: only deep spalling, deep spalling with non-spalling area, only shallow spalling, shallow spalling with non-spalling area, and non-spalling. The spalling area is categorized as deep spalling when the reinforcing steel bars are exposed. The shallow spalling areas are the ones whose condition lies between deep spalling and non-spalling. Moreover, the spalling without exposed steel bars was considered shallow spalling in this paper.</p><p>The method was trained and tested on a system with a GTX 1080 GPU. The size of the image was (256 &#215; 256). For the multi-class classi&#57436;cation problem, we used categorical cross-entropy (CCE) loss, which is also known as Softmax loss. The Adam optimizer was used to optimize the architecture with a learning rate of 0.001. We have used an augmentation process (described in Fig. <ref type="figure">6</ref>) during the training and validation phases. The use of the augmentation process during the training and validation phases has the advantage of avoiding over&#57436;tting problems <ref type="bibr">[82]</ref>. The dataset contains 10000 images for training and another 2000 images for validation. The CCE loss curve of the training and validation for PSPNet, SegNet, and UNet are shown in Fig. <ref type="figure">8,</ref><ref type="figure">9</ref>, and 10, respectively. We recorded the loss for all the deep encoder-decoder combinations. In the loss curve, the training loss and validation loss show how the model &#57436;ts the training data and the new data, respectively. The loss is measured by the error between its predicted output and the true output. Our goal is to get the loss value as close as 0. During the starting phase, Fig. <ref type="figure">8,</ref><ref type="figure">9</ref>, and 10 show gaps between the training and validation curves. The use of the augmentation process during the training and validation phases helped to reduce the gaps gradually. Since the gaps were reducing and both of the loss curves were getting close to 0, the model was &#57436;tting well. The training and validation curves started converging approximately after 90 epochs which indicates a desirable characteristic. We maintained 100 epochs for all the architectures to avoid over&#57436;tting.</p><p>The testing phase was conducted on 300 images. First, we tested the proposed approach with 100 images to determine the di&#57344;erence in performance achieved based on the number of test images. The di&#57344;erence in performance for these two datasets is negligible (approximately 0.01% for all metrics), which justi&#57436;es a consistent performance. Hence, we have presented the performance analysis only for the dataset with 300 images. The proposed approach trained for 100 epochs. Therefore, on each epoch, it was trained on a di&#57344;erent dataset because of the augmentation process.</p><p>Because spalling is detected based on its severity levels, we have presented non-statistical qualitative analysis as well as quantitative analysis with statistical measurements. In the quantitative analysis sub-  section, we evaluated the performance of three deep architectures with di&#57344;erent encoder-decoder networks using di&#57344;erent metrics. The qualitative analysis subsection describes the performance comparison based on the results of spalling detection and severity level.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.2.">Quantitative analysis</head><p>This section presents a performance-based statistical analysis and deep learning-based image segmentation architectures with di&#57344;erent encoder-decoder networks. The overall performance for spalling detection with severity level is shown in Table <ref type="table">2</ref>.</p><p>We have used Dice loss, mIoU, Precision, Recall, and Accuracy metrics for the performance analysis. The dice loss referred to the loss level for the combination architecture with di&#57344;erent encoder-decoder networks. We have performed the calculation on Equation ( <ref type="formula">2</ref>), (3), and (4) to &#57436;nd out the Accuracy, Precision, and Recall respectively. From Equation (5), we get the calculation for IoU for each class which helps to calculate the mean value of IoU for all classes.</p><p>Table <ref type="table">3</ref> describes the quantitative measures used for evaluating the performance of deep network architectures. The lower Dice Loss values are more appropriate since they re&#57437;ect the degree of loss incurred by the di&#57344;erent combinations of network frameworks employed in the proposed system for spalling and severity detection. The higher values for all other performance measures re&#57437;ect the proposed spalling and severity detection system's improved performance. The suggested system performs well for spalling and severity detection as the mIoU, precision, recall, and accuracy values increase.</p><p>The statistical performance for PSPNet is shown in Table <ref type="table">2</ref>; the result for PSPNet architecture with encoders namely, Xception, ResNet-50, MobileNet, and VGG-19, respectively. According to the metrics discussed above, PSPNet architecture with Xception gives the best result among all the combinations (e.g., Dice Loss: 5.94%, mIoU: 88.78%, Precision: 94.67%, Recall: 98.43%, Accuracy: 96.06%). The result for ResNet-50 is pretty close to Xception. PSPNet with default encoderdecoder network gives comparatively poor results than with the other encoder-decoder networks (e.g., Dice Loss: 8.33%, mIoU: 84.62%, Precision: 92.53%, Recall: 90.82%, Accuracy: 92.40%). The performance of VGG-19 with the PSPNet architecture shows that it closely follows the performance of PSPNet with the default encoder-decoder network (e.g., Dice Loss: 7.82%, mIoU: 85.49%, Precision: 94.44%, Recall: 92.95%, Accuracy: 92.97%). For PSPNet, the decreasing CNN layers (during employing Xception, ResNet-50, MobileNet, VGG-19) have a negative impact on the performance for spalling and severity level detection.     The above discussion and performance evaluation shown in Table <ref type="table">2</ref> infer that PSPNet gives comparatively good performance for detecting spalling and severity levels among the three deep architectures. For all three deep architectures, Xception gives the best result. The VGG-19 provides comparatively poor performance compared to other encoderdecoder networks for detecting spalling and severity levels with SegNet and UNet architectures. For PSPNet architecture, the VGG-19 encoder and PSPNet with the default encoder-decoder network both provide poorer performance than other encoder-decoder networks.</p><p>Table <ref type="table">4</ref> shows the results for the severity ranking of deep spalling for three image categories.</p><p>Several ranking methods are proposed for civil infrastructure <ref type="bibr">[88] [89]</ref>. Moreover, we have analyzed our dataset and observed that we should categorize the severity ranking for deep spalling areas. We have 300 images for testing the deep architectures for spalling severity detection. Among the 300 images, there are around 100 images of deep spalling areas. Based on our observation, we have de&#57436;ned the ranking for deep spalling areas as very severe, medium severe, and less severe. Fig. <ref type="figure">11</ref> shows the spalling ratio of 100 test images for deep spalling ar-eas. According to the spalling ratio and from our observations of the dataset, we have de&#57436;ned thresholds for the severity ranking. The pre-de&#57436;ned thresholds for the severity ranking are de&#57436;ned as: less severe when Ratio &#8804; 0.39, medium severe when, 0.4 &#8804; Ratio &#8804; 0.69, and very severe when Ratio &#8805; 0.7. Table <ref type="table">4</ref> displays the ratio, which is expressed in terms of 100%. We have selected three di&#57344;erent categories of images from the 100 deep spalling images. In Table <ref type="table">4</ref>, we presented the results of the severity ranking for each deep architecture with di&#57344;erent backbone networks based on each selected image category. The Input Size displays the total number of pixels of the image, which is the same for all the images, Spalling Size refers to the number of pixels a&#57344;ected by deep spalling, Ratio is calculated using Equation (1) and shown in 100%, and Severity Ranking determines the ranking of severity based on the ratio and the prede&#57436;ned threshold value. We compared the severity ranking to ground truth for each image category. For most of the models, the severity ranking follows the ranking of ground truth very closely.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.3.">Qualitative analysis</head><p>This section presents the qualitative analysis of the proposed approach to show the non-statistical performance of the deep architectures with di&#57344;erent encoder-decoder networks. The performance evaluation of di&#57344;erent deep architectures for detecting spalling and severity level segmentation has been shown in Fig. <ref type="figure">12</ref>, Fig. <ref type="figure">13</ref>, and Fig. <ref type="figure">14</ref>. The results highlight the overall performance of spalling and severity detection based on deep spalling, shallow spalling, and non-spalling images.</p><p>In Fig. <ref type="figure">12</ref>, the results are shown for the PSPNet framework with di&#57344;erent encoder-decoder networks. We have mentioned earlier that the images in our dataset are categorized as only deep spalling, deep spalling with non-spalling area, only shallow spalling, shallow spalling with non-spalling area, and non-spalling. For the PSPNet framework, only deep spalling, shallow spalling with non-spalling, and non-spalling areas were chosen as input to present the performance evaluation. In Fig. <ref type="figure">12</ref>, we have original image, ground truth which is pixel-by-pixel The PSPNet architecture with the default encoder and VGG-19 network gives comparatively poor results compared to the other encoderdecoder networks. Among the three severity classes of spalling, nonspalling areas are predicted to be more accurate for all the architecture combinations.</p><p>The results are shown for the UNet framework with di&#57344;erent encoder-decoder networks in Fig. <ref type="figure">13</ref>. For the UNet framework, we have chosen here deep spalling with a non-spalling area, shallow spalling with a non-spalling area, and non-spalling area as input for performance evaluation. A sticker serves as noise in the shallow spalling image. We have the original image, ground truth, which is a pixel-by-pixel mapping of the original image for each spalling class, the result for UNet architecture, and the result for UNet architecture with encoders namely, Xception, ResNet-50, MobileNet, and VGG-19, as shown in Fig. <ref type="figure">13</ref>, respectively. Fig. <ref type="figure">13</ref> for Xception shows that the UNet architecture gives the best result among all the combinations. The comparative analysis shows that MobileNet follows the results of Xception. ResNet-50 provides pretty low performance compared to Xception and MobileNet, unlike PSPNet. In comparison to ground truth, the results provided by Unet architecture, UNet architecture with ResNet-50, and VGG-19 MobileNet have some inaccurate predictions for deep spalling, shallow spalling, and non-spalling images. Fig. <ref type="figure">13</ref> shows that the VGG-19 encoder performs poorly in comparison to the other encoder-decoder networks.</p><p>The comparative analysis of SegNet architecture with the default encoder-decoder and with Xception, ResNet-50, MobileNet, and VGG-19 is shown in Fig. <ref type="figure">14</ref>. The categorization of images for performance evaluation of the SegNet framework was chosen here as deep spalling with the non-spalling area, shallow spalling with non-spalling, and nonspalling. In Fig. <ref type="figure">14</ref>, we have the original image of deep spalling, shallow spalling, and non-spalling area, ground truth which is pixel-by-pixel mapping of the original image for each spalling class, result for Seg-Net architecture, result for SegNet architecture with encoders namely, Xception, ResNet-50, MobileNet, and VGG-19, respectively. The Seg-Net architecture with the VGG-19 encoder gives comparatively poor results compared to other encoder-decoder networks like the UNet architecture. VGG-19 shows poor performance, especially for non-spalling and shallow spalling classes. The non-spalling is predicted pretty accurately by most of the architecture combinations, except for the Seg-Net architecture with VGG-19. In Fig. <ref type="figure">14</ref>, the comparative analysis presents that the SegNet architecture with Xception gives the best result among all the combinations for all the spalling severity classes. The results for ResNet-50 show that it matches the result for Xception pretty closely. SegNet architecture with the default encoder-decoder network gives poor results compared to ground truth, especially for shallow spalling. MobileNet gives some incorrect predictions for deep and shallow spalling areas.</p><p>Based on the discussion above and the performance shown in Fig. <ref type="figure">12</ref>, 13, and 14, it can be concluded that most deep architectures with encoder-decoder networks provide comparatively good results for non-spalling areas. The performance evaluation for predicting deep spalling and shallow spalling closely follows the performance evaluation for predicting non-spalling areas. The performance evaluation shows that, among the three deep architectures, PSPNet shows the best performance for detecting spalling and severity classi&#57436;cation. The Xception gives the best results for detecting deep spalling, shallow spalling, and non-spalling with SegNet, UNet, and PSPNet deep architectures. Comparatively, VGG-19 shows poor performance in detecting spalling and severity levels with UNet and SegNet architectures. For PSPNet, the VGG-19 encoder closely follows the performance of PSPNet with the default encoder-decoder network.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.">Conclusions and future work</head><p>This paper presents an innovative deep learning-based approach to detect spalling and its severity levels in civil infrastructure using encoder-decoder networks. The proposed method &#57436;lls the gap in the literature, where very few methods exist for detecting the severity level of spalling accurately. Our study shows that deep learning-based architectures with encoder-decoder networks o&#57344;er high performance in detecting spalling severity levels in di&#57344;erent &#57436;elds, including civil infrastructure.  We incorporated three di&#57344;erent deep architectures and four backbone networks in our proposed methodology to achieve the best performance. Our results indicate that the PSPNet-based deep architecture with the Xception encoder o&#57344;ers the best performance. We have also conducted statistical and non-statistical analyses to demonstrate the proposed method's high performance.</p><p>Our study has several potential future directions, including improving the proposed deep architecture's e&#57345;ciency by reducing power consumption and memory requirements while achieving better performance in detecting spalling and severity levels. Additionally, this approach's adaptability to detect various concrete distresses using a deep architecture-based combined detection process is worth exploring. Overall, our proposed method provides a promising solution for detecting and classifying spalling severity levels in civil infrastructure, which is crucial for ensuring the structural health of concrete.  </p></div></body>
		</text>
</TEI>
