<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>High Definition, Inexpensive, Underwater Mapping</title></titleStmt>
			<publicationStmt>
				<publisher></publisher>
				<date>2022</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10339319</idno>
					<idno type="doi">10.1109/ICRA46639.2022.9811695</idno>
					<title level='j'>IEEE International Conference on Robotics and Automation  (ICRA)</title>
<idno></idno>
<biblScope unit="volume"></biblScope>
<biblScope unit="issue"></biblScope>					

					<author>Bharat Joshi</author><author>Marios Xanthidis</author><author>Sharmin Rahman</author><author>Ioannis Rekleitis</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[In this paper we present a complete framework for Underwater SLAM utilizing a single inexpensive sensor. Over the recent years, imaging technology of action cameras is producing stunning results even under the challenging conditions of the underwater domain. The GoPro 9 camera provides high definition video in synchronization with an Inertial Measurement Unit (IMU) data stream encoded in a single mp4 file. The visual inertial SLAM framework is augmented to adjust the map after each loop closure. Data collected at an artificial wreck of the coast of South Carolina and in caverns and caves in Florida demonstrate the robustness of the proposed approach in a variety of conditions.]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head>I. INTRODUCTION</head><p>The underwater domain presents a special allure since the early days of exploration <ref type="bibr">[1]</ref>; coral reefs, shipwrecks, and underwater caves all present unique views like nothing most people see above water. Underwater exploration using acoustic sensors is well studied, however, the resulting representations convey only limited information; in contrast vision based mapping presents the most familiar representations <ref type="bibr">[2]</ref>- <ref type="bibr">[5]</ref>. Underwater is a very challenging environment for cameras. The visibility is limited, sometimes objects after a few meters disappear; color attenuation, colors disappear with depth starting with red <ref type="bibr">[6]</ref>, <ref type="bibr">[7]</ref>; floating particulates generate blurriness; there is varying illumination resulting from caustic patterns due to waves up to complete lack of ambient light inside caves; and the reduced number of features makes localization challenging.</p><p>Autonomous Underwater Vehicles (AUVs) and Remotely Operated Vehicles (ROVs) range in cost from a few thousand to hundreds of thousand of dollars. Furthermore, camera technologies for these vehicles, unless at the higher end of the spectrum provide images not of the highest quality; one major challenge is the light has to pass from water to the AUVs window, through air, then the lens of the camera generating additional distortions. In the recent years so-called action cameras and in particular the GoPro cameras have produced exceptional imagery for a fraction of the cost. The improvements in image quality though, were limited by the single camera view which made estimation of scale Fig. <ref type="figure">1</ref>: Collecting data in an underwater cavern, Ginnie Spring, FL, USA. GoPro 9 camera is attached to the stereo-rig <ref type="bibr">[15]</ref>, lighting from two Keldan lights <ref type="bibr">[16]</ref>. near impossible. As shown in Joshi et al. <ref type="bibr">[8]</ref> and Quattrini Li et al. <ref type="bibr">[9]</ref> monocular vision without inertial data has very low accuracy. From the GoPro 5 black, the video contains embedded inertial data at a rate of 200 Hz, without synchronization information. Starting from GoPro 8, the video contains inertial information along with necessary timing information for camera IMU synchronization thus making underwater state estimation feasible. In this paper we tested some of the most promising open-source Visual Inertial SLAM packages <ref type="bibr">[10]</ref>- <ref type="bibr">[14]</ref>, in a variety of environments with very accurate results. Furthermore, the SVIn2 framework <ref type="bibr">[10]</ref> is augmented with updating the 3D pose of the detected visual features after loop closure producing a consistent global map. The code is publicly available <ref type="foot">1</ref> .</p><p>Experiments conducted over an artificial reef (refuelling barge wreck) off the coast of South Carolina; inside the Devil System cave, FL; at the spring waters (open water) of Troy Springs State Park, FL; and in the cavern of Ginnie Springs, FL. In particular inside the Ginnie Springs cavern, five fiducial markers were placed in different locations in order to estimate the accuracy of the estimated trajectory. These datasets are made publicly available together with calibration data and basic scripts for evaluating ground truth <ref type="foot">2</ref> .</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>II. RELATED WORK</head><p>A stereo GoPro setup was used to map underwater caves <ref type="bibr">[17]</ref>, however, this technology is no longer available, and it is only recently that the IMU of the GoPro 9 allows scale-accurate results from a single camera. GoPro cameras have been studied in underwater settings <ref type="bibr">[18]</ref>, <ref type="bibr">[19]</ref>, and used, due to the high quality imagery in a variety of underwater tasks, for coral monitoring <ref type="bibr">[19]</ref>- <ref type="bibr">[22]</ref>, underwater archaeology <ref type="bibr">[23]</ref>, and seafloor reconstruction <ref type="bibr">[24]</ref>, <ref type="bibr">[25]</ref>.</p><p>Wreck mapping has been studied using a variety of techniques all around the world. Photogrammetry of manually obtained images resulted in mosaics in Demesticha et al. <ref type="bibr">[3]</ref>, or from an ROV, see Nornes et al. <ref type="bibr">[26]</ref>. While the Arrows EU project provides an overview of robotic technology used <ref type="bibr">[27]</ref>. Menna et al. <ref type="bibr">[28]</ref> provide a comprehensive review of techniques used. Mapping projects extend from Italy <ref type="bibr">[29]</ref>, Spain <ref type="bibr">[30]</ref>, Canada <ref type="bibr">[31]</ref>, Qatar <ref type="bibr">[32]</ref>, up to the arctic <ref type="bibr">[33]</ref>. With the most famous wreck explorations of the Titanic <ref type="bibr">[2]</ref> and the Antikythera <ref type="bibr">[34]</ref> shipwrecks.</p><p>Coral reef mapping also utilizes vision. By creating specialized sensors <ref type="bibr">[15]</ref>, <ref type="bibr">[35]</ref> or utilizing UAVs <ref type="bibr">[36]</ref> there is a need for Underwater SLAM <ref type="bibr">[37]</ref>. Due to the deteriorating health of the coral reefs, it is important to document the state of the different reefs and to measure the rate of deterioration. Of particular interest is to identify resilient species to assist re-population efforts.</p><p>There are few datasets from underwater experiments <ref type="bibr">[8]</ref>, <ref type="bibr">[9]</ref>, <ref type="bibr">[38]</ref>, however, obtaining ground truth is extremely challenging. The Aqualoc dataset <ref type="bibr">[38]</ref> used the trajectory estimated using global optimization package Colmap <ref type="bibr">[14]</ref> as ground truth. This work will contribute a novel collection of datasets, and in select cases a set of permanent landmarks to act as ground truth. This dataset contains high definition/resolution images and inertial data at 200 Hz. The image quality is far superior than existing datasets.</p><p>Mapping underwater caves is extremely challenging due to the total lack of ambient light. Wakulla Springs cave is one of the most well known and efforts to map it include the Wakulla 2 project <ref type="bibr">[39]</ref>, <ref type="bibr">[40]</ref> utilizing mainly acoustic sensors, as was mapping a cenote <ref type="bibr">[41]</ref>. Nocerino et al. <ref type="bibr">[42]</ref> proposed the use of multiple cameras on a ROV for mapping caves, then use it to map caves in Sicily <ref type="bibr">[43]</ref>. Malios et al. <ref type="bibr">[44]</ref> proposed also a SLAM framework for confined spaces. The works of Rahman et al. <ref type="bibr">[10]</ref>, <ref type="bibr">[45]</ref> has demonstrated accurate results over long trajectories in a variety of settings. On this work we utilize a subset of this work SVIn2 utilizing the initialization and loop closure extension over OKVIS <ref type="bibr">[46]</ref>. Furthermore, a framework for denser reconstructions used the plethora of shadows in the cave environment <ref type="bibr">[47]</ref>. We have augmented this framework to produce a consistent map, by updating the triangulated features after every loop closure.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>III. PROPOSED APPROACH</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>A. Sensor Setup</head><p>The GoPro 9 consists of a color camera, an IMU, and a GPS. GPS does not work underwater; thus it was not used in this work. However, GPS information can be fused with Visual Inertial Navigation Systems (VINS) during above-water operations. For calibrating the camera intrinsic parameters and the extrinsic parameters of the sensor setup, we use a grid of AprilTags <ref type="bibr">[48]</ref>.  sensor has an inbuilt 12-bit A/D converter to shoot highspeed and high-definition videos using horizontal and vertical binning and subsampling readout. The sensor has on-chip R, G, and B primary color mosaic filters for better color capture. GoPro 9 has multiple settings for recording the video, while many can be used there are certain modes which are prohibitive to VIO operations due to the non-linear transformation of the image as detailed in <ref type="bibr">[49]</ref>. We found that the SuperView mode generates non-linear distortions that thwart calibration of the camera intrinsic parameters during data collection. The videos were recorded at full High Definition (HD) resolution of 1960&#215;1080 with wide lens setting: horizontal field-of-view (FOV) 118 &#8226; , vertical FOV 69 &#8226; , and hypersmooth level set to off. Hypersmooth levels control the electronic image stabilization that predicts camera motion and compensates for it by cropping the viewable image. Hypersmoothing can effectively crop up to 10% of the image frame and the amount of cropping depends on the amount of motion, rendering this mode extremely challenging for VIO applications. GoPro 9 also includes a Bosch BMI260 IMU equipped with 16-bit 3-axis MEMS accelerometer and gyroscope. GoPro 9 inherently records IMU data at 200Hz. The timestamps of IMU and camera are synchronized using the timing information from metadata encoded inside the MP4 video.</p><p>GoPro 9 also includes a UBlox UBX-M8030 GNSS chip capable of concurrent reception of up to 3 GNSS (GPS, Galileo, GLONASS, BeiDou) and accuracy of 2 m horizontal circular error probable, meaning 50% of measurements fall inside circle of 2 m.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>B. GoPro Telemetry Extraction</head><p>The GoPro 9 MP4 video file is divided into multiple streams namely video encoded with H.265 encoder, audio encoded in advanced audio coding (AAC) format, timecode (audio-video synchronization information), GoPro fdsc data stream for file repair and GoPro telemetry stream in GoPro Metadata Format referred as GPMF <ref type="bibr">[50]</ref>. GPMF -is a modified Key, Length, Value solution, with a 32-bit aligned payload, that is both compact, fully extensible, and somewhat human readable in a hex editor. Please refer to <ref type="bibr">[50]</ref> for more details on GPMF, here we focus on the camera-IMU synchronization.</p><p>GPMF is divided into payloads, extracted using gpmfparser <ref type="bibr">[50]</ref>, with each payload containing sensor measurements for 1.01 seconds while recording at frame rate of 29.97 Hz as shown in Fig. <ref type="figure">2</ref>. A particular sensor information is obtained from payload using FourCC-7-bit 4 character ASCII key, for instance 'ACCL' for accelerometer, 'GYRO' for gyroscope, and 'SHUT' for shutter exposure times. The payload also contains the starting time of each payload in microseconds relative to the start of the video capture. Since images are encoded in the video stream, we use the start of 'SHUT' payload to find relative timing with the accelerometer and gyroscope measurements. Using the start and end of payloads along with the number of measurements in that payload, we interpolate the timing of all measurements. We decode the video stream and extract images from MP4 file using the FFmpeg library and combine them with the IMU measurements using timing information from the GPMF payload. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>C. Calibration</head><p>Firstly, we calibrate the camera intrinsic parameters. We use the camera-calib sequences where we move the GoPro 9 camera in front of a calibration pattern placed in an indoor swimming pool from different viewing angles.</p><p>The intrinsic noise parameters of the IMU are required for the probabilistic modeling of the IMU measurements used in state estimation algorithms and the camera-IMU extrinsic parameters. We assume that IMU measurements (both linear accelerations and angular velocities) are perturbed by zeromean uncorrelated white noise with standard deviation &#963; w and random walk bias, which is the integration of white noise with standard deviation &#963; b . To determine the characteristics of the IMU noise, the Allan deviation plot &#963; Allan (&#964; ) as a function of averaging time needs to be plotted. In a log-log plot white noise appears on Allan Deviation plot as a slope with gradient -1 2 and &#963; w can be found at &#964; = 1. Similarly, the bias random walk can be found by fitting a straight line with slope 1  2 at &#964; = 3; for details see <ref type="bibr">[52]</ref>. Fig. <ref type="figure">3</ref> shows the Allan deviation plot for GoPro 9 camera along with the noise parameters.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>D. Global Map</head><p>Most VIO packages ( <ref type="bibr">[10]</ref>, <ref type="bibr">[11]</ref>, <ref type="bibr">[13]</ref>), when applying loop closure they update only the pose graph, leaving the triangulated features in their original estimate. In contrast COLMAP <ref type="bibr">[14]</ref> is a global optimization package and the final result optimizes both poses and feature 3D locations. We run COLMAP by using 2 images per second resulting in 2000 to 3000 images in cavern and cave sequences, which takes on average 7 to 10 hours. ORB-SLAM3 <ref type="bibr">[12]</ref> performs global bundle adjustment after loop closure, however often diverges over large dataset as the local mapping is stopped during global optimization.</p><p>An enhancement is proposed for the SVIn2 framework to update the 3D pose of the tracked features every time loop closure occurs in order the 3D features to be consistent with the pose-graph optimization results. We maintain the pose graph with the keyframes as vertices for loop closure and the edges indicate the relative pose constraint between keyframes. For each keyframe f , the VIO module passes the following information to the loop closure module:</p><p>&#8226; T wf pose of keyframe in world coordinate system &#8226; 3D points visible in keyframe l i with each point l = [P w , F, Q, I] has the following attributes: position in world frame P w &#8712; R 3 , index of keyframe F &#8712; N, landmark quality Q &#8712; R, and position of keypoint in image I l &#8712; N 2 . SVIn2, provides a quality measure of a 3D point as the ratio of the square root of minimum and maximum eigenvalues of the Hessian block matrix associated with the 3D point. For each 3D point, we calculate its local position in keyframe as P f = T -1 wf P w , color C as RGB value at pixel location I. We collect all the observations of 3D landmark l from multiple keyframes as O f = [P f , F, Q f , C f ] hashed using keyframe index resulting in O(1) lookup time for each observation. In the event of loop closure, we deform the global map so that the relative pose between each point and its attached keyframes remains unchanged. We fuse the multiple landmark observations from keyframe f = 1 to N to obtain global position P w , color C and quality Q as shown in Eq. <ref type="bibr">(1)</ref>. It should be noted that P f is expressed as the relative position with respect to the keyframe and whenever the pose of keyframe T wf changes due to loop closure updates, the location of the 3D landmarks in the global frame also changes accordingly; producing a globally consistent map as shown in Fig. <ref type="figure">4</ref>. Realistically, the number of observations of a 3D point is quite small compared to the total 3D points in the scene; thus global map building scales linearly O(n) with the number of points in the scene.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>IV. DATASETS</head><p>The GoPro9 Underwater VIO dataset consists of calibration sequences in addition to odometry evaluation sequences recorded using 2 GoPro 9 cameras. The two cameras are referred as g i , i &#8712; <ref type="bibr">[1,</ref><ref type="bibr">2]</ref>. Although we provide calibration results for all the datasets, the calibration sequences are made available to facilitate users wanting to preform their own calibration. Datasets can be categorized as:</p><p>&#8226; camera-calib: for calibrating the camera intrinsic parameters underwater. We provide calibration sequences using two calibration patterns: a grid of AprilTags and a checkerboard for each camera. The calibration patterns are recorded with slow camera motions at a swimming pool; making sure the calibration patterns are viewed from varying distances and orientations.</p><p>&#8226; camera-imu-calib: for calibrating the camera-IMU extrinsic parameters in order to determine the relative pose between the IMU and the camera. The camera is moved in front of the Apriltag grid exciting all 6 degrees of freedom. This sequence is recorded indoor just to expedite the calibration process as the extrinsic parameters are the same above and below water. Moreover, we also provide a camera-calib indoor sequence to assist with the camera-IMU calibration.</p><p>&#8226; imu-static: contains IMU data to estimate white noise and random walk bias parameters. These sequences are recorded with the camera stationary for at least 4 hours for each cameras.</p><p>&#8226; shipwreck: two sequences collected by handheld GoPro 9 on an artificial reef (refueling barge wreck) 55 Km outside of Charleston, SC, USA; see Fig. <ref type="figure">5(a)</ref>.</p><p>&#8226; spring open water one sequence collected by handheld GoPro 9 at the basin of Troy Springs State Park FL, USA; see Fig. <ref type="figure">5(b</ref>). The settings had image stabilization on (hypersmooth) resulting into arbitrary cropping of the field of view.</p><p>&#8226; cavern: three trajectories traversed inside the ballroom cavern at Ginnie Springs, FL; see Fig. <ref type="figure">5(c</ref>). Inside the cavern five markers (AR single tags <ref type="bibr">[53]</ref>) were placed to establish ground truth measurements. Each trajectory consisted of several loops each observing slightly different parts of the cavern but all ensuring the tag of that part of the cavern was visible.</p><p>&#8226; cave: two sequences were collected at the Devil's system, FL; see Fig. <ref type="figure">5(d)</ref>. For the first sequence the GoPro was mounted on the stereo rig and second sequence was using only the GoPro.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>V. EXPERIMENTAL RESULTS</head><p>Due to absence of GPS in underwater environments or motion capture systems, we use COLMAP <ref type="bibr">[14]</ref> to generate baseline trajectories. COLMAP is a structure-from-motion (SfM) pipeline equipped with global bundle adjustment and loop closure capabilities; thus producing consistent camera trajectory and 3D reconstruction. Even though COLMAP provides good estimation of shape of trajectories, they can not be considered as ground truth. As monocular SfM inherently suffers from scale observability constraints and global optimization does not converge over large trajectories, we consider the estimated trajectories as accurate up to scale. We found that relative scale between COLMAP and all VIO trajectories was almost equal. Hence, COLMAP trajectories are scaled by scaling factor calculated from the average of all VIO trajectories in subsequent sections unless otherwise specified.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>A. Tracking Evaluation Metrics</head><p>As COLMAP does not provide accurate scale, we evaluate the accuracy of the various tracking algorithms using the absolute trajectory error (ATE) metric after Sim(3) alignment <ref type="bibr">[54]</ref>. ATE is calculated as the root mean squared difference between ground truth 3D positions obtained from COLMAP p i and corresponding estimated 3D positions pi aligned using optimal Sim(3) rotation matrix R, translation t, and scaling factor s:</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>B. Tracking Results</head><p>We compare the performance of various open source visual-inertial odometry (VIO) methods on the above described datasets. Since, most of the VIO algorithms have parameters that are tuned at VGA resolution; their performance was found to be better at quarter resolution (960&#215;540 pixels). However, we provide datasets at full high definition resolution (1920&#215;1080 pixels) to further the research in underwater VIO and SfM. Unless specified otherwise, all the VIO algorithms use quarter resolution dataset.</p><p>We evaluate the performance of VINS-Mono <ref type="bibr">[11]</ref>, ORB-SLAM3 <ref type="bibr">[12]</ref>, SVIn2 <ref type="bibr">[10]</ref> and OpenVINS <ref type="bibr">[13]</ref> based on the RMSE of the ATE as shown in Table <ref type="table">II</ref>  All system are able to track most of the sequences until the end. VINS-Mono and SVIn2 were able to track the complete trajectory consistently with good accuracy. In the cavern sequences, ORB-SLAM3 took too long during global bundle adjustment after loop closure and lost track as it disabled local mapping during global optimization. However, ORB-SLAM3 is equipped with map merging and was able to relocalize and merge the disjoint maps. This produced slightly inferior performance in the cavern sequences. OpenVINS required smooth motion for initialization as it only relies on IMU measurement for gravity alignment and orientation initialization. Hence, OpenVINS might diverge unless data collection is started from a static position or good initialization is found. Finally, it is worth noting that the open water sequence was recorded with the hypersmooth option activated resulting in failures in ORB-SLAM3 and OpenVINS.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>C. AR-tag based Validation</head><p>As there is no continuous tracking for absolute ground truth, 3D landmark-based validation with AR tags is used to quantify the accuracy of the evaluated methods. As a part of the experimental setup in the cavern sequences with multiple loops, we placed 5 different AR-tags printed on waterproof paper at different locations inside the cavern. We observe the variance of the position of the AR-tags from their mean position over the whole length of trajectory. If the trajectories do not drift over time, the markers must be observed at the same location during multiple visits. Among the cavern sequences, we detected most AR-tags in the g1 cavern2 sequence; therefore, this sequence is used as reference for further analysis.</p><p>We determined the relative position between the camera and the tags in g1 cavern2 sequence using ar track alvar <ref type="foot">3</ref> .</p><p>By projecting a 3D cube over the tags using the pose estimate, we observed higher noise in the orientation estimate; hence, only the position of the AR-tags is used for the error analysis. Once the relative position from ar track alvar is found, the global position can be found as T k W M = T k W C * P CM where T k W M is the marker position in world coordinate frame W at time k, T k W C is the pose of the camera C in W at time k (produced by SLAM/odometry system), and P CM is the relative position of marker M from camera. Fig. <ref type="figure">7</ref> shows boxplots of the displacement from the mean position of the markers over the whole length of the trajectory for all the different methods, including COLMAP. Table <ref type="table">III</ref> shows the summary of the standard deviation in translation and the average distance error. All the algorithms performed well with slightly inferior performance of ORB-SLAM3 due to tracking issues. Fig. <ref type="figure">8</ref> shows the position of the tags in g1 cavern2 sequence observed by different packages along with the COLMAP trajectory as reference.  </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>D. Global Mapping</head><p>We evaluate the global map produced by enhancing SVIn2 to update the triangulated feature positions after loop closures by comparing with COLMAP's sparse pointcloud. The pointclouds come from different sources and differ in size, so we align the pointclouds from COLMAP and SVIn2 as follows:</p><p>1) Perform voxel downsampling with voxel size of 10cm.</p><p>2) Compute FPFH <ref type="bibr">[55]</ref> feature descriptor describing local geometric signature for each point. 3) Find correspondence between pointclouds by computing similarity score between FPFH descriptors. 4) Feed all putative correspondences to TEASER++ <ref type="bibr">[56]</ref> to perform global registration finding transformation to align corresponding points. 5) Fine tune registration by running ICP over original point cloud with TEASER++ solution as initial guess. The reconstruction results are compared based on registration accuracy using fitness and inlier rmse metrics. More specifically, fitness is the ratio of number of inlier correspondences (distance less than voxel size) and number of points in SVIn2 pointcloud. Whereas, inlier rmse is the root mean squared error of all inlier correspondences. Table <ref type="table">IV</ref> shows the similarity between SVIn2 and COLMAP reconstruction based on fitness and inlier rmse metrics. Fig. <ref type="figure">10</ref> shows aligned sparse reconstruction obtained from COLMAP and SVIn2 in g1 shipwreck1 and g1 cavern2 sequence.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>VI. CONCLUSION</head><p>In this work we presented a complete pipeline for underwater SLAM utilizing a commonly available, inexpensive, x (m)       SVIn2 framework was augmented to correct the 3D features according to the updated pose graph after successful loop closure calculations. The resulting map demonstrated accuracy similar to the much slower global optimization COLMAP package. The experimental results verify that the specific camera is capable of producing accurate estimates of the trajectory together with consistent sparse representations of the environment. Future uses of the proposed framework would be in recording and documenting the surroundings and the trajectory of AUVs operating autonomously in challenging underwater environments such as caves and shipwrecks <ref type="bibr">[57]</ref>, <ref type="bibr">[58]</ref>.</p><p>Currently we are investigating synchronization methods between the GoPro camera and other devices such as Autonomous Underwater Vehicle and sensor suites. By introducing additional data streams such as water depth and magnetometer data, both providing absolute values, we expect to reduce the drift accumulating over long trajectories without loops and increase the overall accuracy.</p></div><note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="1" xml:id="foot_0"><p>https://github.com/AutonomousFieldRoboticsLab/ gopro_ros</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="2" xml:id="foot_1"><p>https://afrl.cse.sc.edu/afrl/resources/datasets/</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="3" xml:id="foot_2"><p>http://wiki.ros.org/ar_track_alvar</p></note>
		</body>
		</text>
</TEI>
