<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>Remote referencing strategy for high-resolution coded ptychographic imaging</title></titleStmt>
			<publicationStmt>
				<publisher></publisher>
				<date>01/01/2023</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10438527</idno>
					<idno type="doi">10.1364/OL.481395</idno>
					<title level='j'>Optics Letters</title>
<idno>0146-9592</idno>
<biblScope unit="volume">48</biblScope>
<biblScope unit="issue">2</biblScope>					

					<author>Tianbo Wang</author><author>Pengming Song</author><author>Shaowei Jiang</author><author>Ruihai Wang</author><author>Liming Yang</author><author>Chengfei Guo</author><author>Zibang Zhang</author><author>Guoan Zheng</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[The applications of conventional ptychography are limited by its relatively low resolution and throughput in the visible light regime. The new development of coded ptychography (CP) has addressed these issues and achieved the highest numerical aperture for large-area optical imaging in a lensless configuration. A high-quality reconstruction of CP relies on precise tracking of the coded sensor’s positional shifts. The coded layer on the sensor, however, prevents the use of cross correlation analysis for motion tracking. Here we derive and analyze the motion tracking model of CP. A novel, to the best of our knowledge, remote referencing scheme and its subsequent refinement pipeline are developed for blind image acquisition. By using this approach, we can suppress the correlation peak caused by the coded surface and recover the positional shifts with deep sub-pixel accuracy. In contrast with common positional refinement methods, the reported approach can be disentangled from the iterative phase retrieval process and is computationally efficient. It allows blind image acquisition without motion feedback from the scanning process. It also provides a robust and reliable solution for implementing ptychography with high imaging throughput. We validate this approach by performing high-resolution whole slide imaging of bio-specimens.]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><p>Ptychography was first developed for solving the phase problem in electron crystallography <ref type="bibr">[1]</ref>. It has grown rapidly in recent years and attracted attention from different research communities. In a typical implementation, the object is translated through a spatially confined probe beam and the corresponding diffraction patterns are recorded in the far-field <ref type="bibr">[2]</ref>. The reconstruction process iteratively imposes two sets of constraints. In the spatial domain, the spatially confined probe beam serves as the support constraint to limit the physical extent of the object for each measurement. In the Fourier domain, the diffraction measurements serve as the Fourier magnitude constraints for the estimated solution. Ptychography does not require a reference beam as in holography. It also lifts the isolated-object requirement of conventional coherent diffraction imaging approaches <ref type="bibr">[3]</ref>. In the past few years, it has become an indispensable imaging tool in most synchrotron and national laboratories worldwide <ref type="bibr">[4]</ref>.</p><p>In the visible light regime, the new development of coded ptychography (CP) enables high-resolution, high-throughput optical imaging in a lensless configuration <ref type="bibr">[5,</ref><ref type="bibr">6]</ref>. Figure <ref type="figure">1(a)</ref> shows the schematic of a typical implementation of CP, where the light waves propagate from the object plane to the coded surface plane, and then to the detector plane. As shown in Fig. <ref type="figure">1</ref>(b), the coded surface, denoted as cs(x &#8242; , y &#8242; ), can be formed by coating a layer of microbeads <ref type="bibr">[7,</ref><ref type="bibr">8]</ref> or by smearing a monolayer of blood cells on the sensor's coverglass <ref type="bibr">[6,</ref><ref type="bibr">9]</ref>. Planes (x, y), (x &#8242; , y &#8242; ), and (x &#8242;&#8242; , y &#8242;&#8242; ) in Fig. <ref type="figure">1</ref> denotes different defocus planes. In the following, we use them interchangeably for representing the lateral coordinates. The coded surface serves as a high-resolution scattering lens with a theoretically unlimited field of view. We note that the sensitivity of the pixel array underneath the coded surface falls off quickly for light waves with large incident angles <ref type="bibr">[5]</ref>. With CP, the coded surface can redirect the large-angle diffracted waves into smaller angles for detection. As such, the otherwise inaccessible high-resolution details can be acquired using the pixel array underneath. Previous demonstrations of CP rely on a clear region on the sensor surface for positional tracking <ref type="bibr">[5,</ref><ref type="bibr">6,</ref><ref type="bibr">8,</ref><ref type="bibr">9]</ref>, thereby sacrificing the valuable imaging area and throughput. Furthermore, the tracking process becomes unreliable when imaging a sparse sample or the edge part of a dense sample. In these cases, there will be no object at the clear region for positional tracking. Here we aim to resolve the positional tracking problem caused by the coded surface in CP and provide a reliable solution for implementing CP by researchers in different fields.</p><p>In the image acquisition process of CP, the object (or the coded image sensor) is translated to different positions (x i , y i ) and the corresponding diffraction measurements I i (x, y) are captured for reconstruction. The imaging model can be expressed as where w(x, y) is the object exit wavefront at the coded surface plane, psf d2 presents the free-space propagation kernel for a distance of d 2 , and "*" represents convolution.</p><p>In our experiment, we used the Sony IMX 226 sensor for image acquisition and the ASI MS-2000 stage for sample translation. The scanning step size is 1-3 microns between adjacent acquisitions with d 2 = 840 &#181;m. A 10-mW collimated laser beam is used for sample illumination and the exposure time is &#8764;1 ms. By acquiring images at 30 frames per second, the continuous scanning process generates a negligible motion blur of &#8764;40 nm. In contrast, conventional ptychography has a large scanning step size. One needs to fully stop the stage for image acquisition, preventing its operation at the full camera frame rate.</p><p>A high-quality reconstruction of CP relies on the precise tracking of the object's positional shifts (x i , y i )s. However, the coded surface cs(x, y) modulates the object's light waves and prevents the use of cross correlation analysis for motion tracking. To derive the closed-form correlation map between the first image I 1 (x, y) and the ith image I i (x, y), we make a first-order approximation to Eq. ( <ref type="formula">1</ref>): where &#8710;w, &#8710;cs are the first-order expansions of the object exit wavefront and the coded surface profile, and &#8710;w &#8242; (xx i , y -</p><p>Here "Re" represents the real part of the complex expression, and &#955; is the wavelength. With Eq. ( <ref type="formula">2</ref>), we can calculate the correlation map R I 1 I i (x, y) between I 1 (x, y) and I i (x, y) as follows:</p><p>(3) where "conj" represents the complex conjugate of the expression, "F " represents the Fourier transform operation, and R &#8710;w &#8242; and R &#8710;cs &#8242; represent the autocorrelation maps of &#8710;w &#8242; (x, y) and &#8710;cs &#8242; (x, y), respectively. The goal of the motion tracking process is to recover the positional shift (x i , y i ) from the correlation map R I 1 I i (x, y) in Eq. ( <ref type="formula">3</ref>). We can see that the map contains two peaks, one at (x i , y i ) and the other at (0, 0). The strong peak caused by the coded surface prevents us to precisely locate the peak at the position (x i , y i ).</p><p>To address this issue, we developed the following pipeline with two main steps: (1) reduce the impact of the coded surface profile via a remote referencing strategy; and (2) enhance the object profile by smearing out the coded surface profile. The unique combination of these two steps allows us to minimize the impact caused by the coded surface in Eq. (3). Figure <ref type="figure">2</ref> shows the overview of the proposed pipeline. In step 1, we capture a reference image I ref (x, y) by moving the object to a remote location (typically &gt;2 mm away from other scanning positions). Similar to Eq. ( <ref type="formula">2</ref>), this reference image can be approximated as</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>y). (4)</head><p>We can then calculate the correlation map between I ref (x, y) and I i (x, y) as follows:</p><p>where R &#8710;cs &#8242; remote &#8710;cs &#8242; represents the cross correlation map between &#8710;cs &#8242; (x, y) and &#8710;cs &#8242; remote (x, y). In contrast with Eq. ( <ref type="formula">3</ref>), this term is a constant because different regions of the blood-coded surface are not correlated with each other. Therefore, Eq. ( <ref type="formula">5</ref>) only contains one peak at the position (x i , y i ). In Figs. <ref type="figure">2(a</ref>) and 2(b), we can then recover the initial positional shift by locating the peak of R I ref I i (x, y) as follows:</p><p>where x ref and y ref are constants and can be subtracted to obtain x i and y i . The second step in Fig. <ref type="figure">2(c</ref>) is to enhance the object  <ref type="formula">7</ref>) and ( <ref type="formula">8</ref>).</p><p>profile by smearing out the coded surface profile. We use the following equation to generate the updated reference image:</p><p>where measurements are shifted back based on the initial positional shift obtained from Eq. ( <ref type="formula">6</ref>). The positional shift can be updated accordingly:</p><p>We typically repeat Eqs. ( <ref type="formula">7</ref>) and ( <ref type="formula">8</ref>) three times to obtain the final positional shift in Fig. <ref type="figure">2(d)</ref>. Table 1. Summary of the Motion Tracking Performance Reference Image Mean Error Standard Deviation First raw image 14.11 &#181;m 6.57 &#181;m Divide the sum 0.75 &#181;m 0.55 &#181;m Remote referencing 0.22 &#181;m 0.1 &#181;m With refinement 0.2 &#181;m 0.08 &#181;m tracking in Fig. <ref type="figure">3(a</ref>). An efficient sub-pixel registration approach with deep sub-pixel accuracy is adopted in our analysis <ref type="bibr">[10]</ref>. Figure <ref type="figure">3</ref>(b1) shows the correlation map of two captured raw images with the coded layer presented. From this map, we can see one peak at (x i , y i ) and the other at (0, 0). The strong peak caused by the coded surface prevents us to locate the peak at the position (x i , y i ), especially when the shifts are small. Figure <ref type="figure">3</ref>(b2) shows the recovered positional shifts and the error map calculated based on the ground-truth positions in Fig. <ref type="figure">3</ref>(a2). In Fig. <ref type="figure">3</ref>(c1), we use the first raw image as the reference and minimize the impact of the coded surface by dividing it by the sum of all images <ref type="bibr">[11]</ref>. This strategy is similar to dividing an image captured in the absence of the object. A better tracking performance can be achieved but the coded surface peak still causes problems for small positional shifts. Figure <ref type="figure">3(d)</ref> shows the performance of the proposed remote referencing strategy, where the peak of the coded surface has been removed. With the refinement step in Eqs. ( <ref type="formula">7</ref>) and ( <ref type="formula">8</ref>), we can obtain deep sub-pixel accuracy in Fig. <ref type="figure">3(e)</ref>. Table <ref type="table">1</ref> summarizes the performances of different cases.</p><p>We validate the imaging performance using a mouse kidney slide Fig. <ref type="figure">4</ref>. The raw image is shown in Fig. <ref type="figure">4</ref> subsequent refinement pipeline, we do not need any positional feedback from the motion stage. The empty space on the coded surface is also no longer needed for positional tracking, enabling the use of all sensor pixels for diffraction data acquisition. The common iterative motion refinement strategy in ptychography is also not needed in the reported platform <ref type="bibr">[12]</ref>. To acquire the large field-of-view images of bio-specimens, we can blindly translate the object to different positions and continuously acquire the corresponding diffraction data for CP reconstruction. In summary, we have analyzed the motion tracking model of CP and reported a novel positional recovery pipeline for blind image acquisition. By using the reported approach, we can suppress the correlation peak caused by the coded surface and recover the positional shifts with deep sub-pixel accuracy. In contrast with common positional refinement methods, the reported approach can be disentangled from the iterative phase retrieval process and is computationally efficient. It allows blind image acquisition without motion feedback from the scanning process. It can also be applied in diffuser-based electron and light microscopy <ref type="bibr">[13,</ref><ref type="bibr">14]</ref>.</p><p>We note that for weak-phase samples, the raw image can be divided by the sum of all images before performing crossrelation analysis, i.e., replacing I i (x, y) with I i (x, y)/ T &#8721;&#65025; i=1 I i (x, y) in Figs. <ref type="figure">4(d</ref>) and 4(e). Doing so effectively suppresses the impact of the coded surface. If the obtained cross-relation map is not uniform, we can also subtract the ith correlation map with the average of the adjacent maps to better remove the background. Lastly, there are many ways to implement the reported approach. For example, one may not need to move the object to a distant location to capture a reference image. Instead, we can divide the captured images into two sets. The first image can be used as the reference image for the second set. The last image can be used as the reference image for the first set.</p></div><note xmlns="http://www.tei-c.org/ns/1.0" place="foot" xml:id="foot_0"><p>0146-9592/23/020485-04 Journal &#169; 2023 Optica Publishing Group</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" xml:id="foot_1"><p>Vol. 48, No. 2 / 15 January 2023 / Optics Letters</p></note>
		</body>
		</text>
</TEI>
