<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>Informed Chemical Classification of Organophosphorus Compounds via Unsupervised Machine Learning of X-ray Absorption Spectroscopy and X-ray Emission Spectroscopy</title></titleStmt>
			<publicationStmt>
				<publisher></publisher>
				<date>07/28/2022</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10353624</idno>
					<idno type="doi">10.1021/acs.jpca.2c03635</idno>
					<title level='j'>The Journal of Physical Chemistry A</title>
<idno>1089-5639</idno>
<biblScope unit="volume">126</biblScope>
<biblScope unit="issue">29</biblScope>					

					<author>Samantha Tetef</author><author>Vikram Kashyap</author><author>William M. Holden</author><author>Alexandra Velian</author><author>Niranjan Govind</author><author>Gerald T. Seidler</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[We analyze an ensemble of organophosphorus compounds to form an unbiased characterization of the information encoded in their X-ray absorption near-edge structure (XANES) and valence-to-core X-ray emission spectra (VtC-XES). Data-driven emergence of chemical classes via unsupervised machine learning, specifically cluster analysis in the Uniform Manifold Approximation and Projection (UMAP) embedding, finds spectral sensitivity to coordination, oxidation, aromaticity, intramolecular hydrogen bonding, and ligand identity. Subsequently, we implement supervised machine learning via Gaussian process classifiers to identify confidence in predictions that match our initial qualitative assessments of clustering. The results further support the benefit of utilizing unsupervised machine learning as a precursor to supervised machine learning, which we term Unsupervised Validation of Classes (UVC), a result that goes beyond the present case of X-ray spectroscopies.]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head>I. INTRODUCTION</head><p>The information content in any spectroscopy method is constrained by the lossiness of the underlying quantum mechanics that connects an atomic-scale structure and dynamics to experimental observables. Further limitations to the sensitivity of spectroscopy techniques often include the inherently nonlinear or stochastic responses of the experimental probe. These facts constrain our ability to correlate physical measurements, e.g., spectral features, to desired microscopic properties. Thus, the emergence of data science and machine learning (ML) in spectroscopy, with applications in all fields in physical sciences, has exploded. <ref type="bibr">[1]</ref><ref type="bibr">[2]</ref><ref type="bibr">[3]</ref><ref type="bibr">[4]</ref><ref type="bibr">[5]</ref> These datadriven models can frequently disentangle and infer patterns from inherently lossy observables as well as provide insight into the information encoded in spectra.</p><p>In general, supervised ML studies across a wide range of spectroscopies target either predicting properties from spectra or correlating specific properties of interest to spectral features. <ref type="bibr">6</ref> This necessarily assumes that sufficient information is, in fact, encoded in spectra; otherwise, supervised ML models will correlate spurious features to requested properties. This detail of encoded information is often addressed by handselecting a targeted training domain, an approach that is deeply contingent on the accuracy and completeness of prior knowledge. <ref type="bibr">7</ref> Clearly, issues will arise if the training domain is too small or biased. First, if the training domain is too small, the model will be unable to generalize well beyond its specialized scope, which violates the essential assumption that the training and test data are sampled from the same distribution. Second, although some bias is essential for any machine learning model, <ref type="bibr">8</ref> unwanted bias, especially from unrepresentative data, blindly undermines reliability of inferences and has led to contemporary ethical concerns. <ref type="bibr">[9]</ref><ref type="bibr">[10]</ref><ref type="bibr">[11]</ref><ref type="bibr">[12]</ref> In an effort to combat unwanted bias as well as provide generalizability to complex datasets, this study demonstrates the value of the discovery cycle exemplified in Figure <ref type="figure">1</ref>. This process validates encoded information via unsupervised machine learning, i.e., cluster analysis on a reduced-dimensional embedding of the spectra, before passing either the embedding or the original spectra&#65533;selected as an unbiased training (sub)set&#65533;to a supervised ML model. This approach decreases the risk of implicit biases and spurious correlations introduced by supervised ML by adding steps (3) and ( <ref type="formula">4</ref>) to validate spectral sensitivity of the training dataset to properties requested during supervised predictions. The continuation of inferences from supervised ML back to experimental design and (primarily) data simulation is obviously informed by the resulting errors achieved by the supervised ML model. This cycle touches on other related ways people have used unsupervised ML as a precursor to or in a cycle with supervised ML. These approaches have included semisupervised machine learning, <ref type="bibr">13</ref> pretraining a neural network, <ref type="bibr">14</ref> feature selection or generation, <ref type="bibr">15</ref> and human-in-the-loop learning, <ref type="bibr">16</ref> which have been used in a multitude of fields, such as gene expression <ref type="bibr">17</ref> and marketing. <ref type="bibr">18</ref> Given the ubiquity of concerns about the scope and bias when constructing training datasets in supervised ML, we propose that our approach, which we term Unsupervised Validation of Classes (UVC), has relevance beyond the present case of X-ray spectroscopies as well as contributes to efforts to close the loop between artificial intelligence and scientific understanding. <ref type="bibr">19</ref> Here, we apply the framework of Figure <ref type="figure">1</ref> to both X-ray absorption spectroscopy (XAS) and X-ray emission spectroscopy (XES). XAS has seen an explosion of ML applications.  XAS is most commonly used in chemistry, biology, and materials science to investigate the element-specific local coordination environment and electronic structure, with applications including energy storage, <ref type="bibr">42,</ref><ref type="bibr">43</ref> catalysis, <ref type="bibr">44</ref> and photochemical dynamics. <ref type="bibr">45</ref> XAS, which includes both the Xray absorption near-edge structure (XANES) and extended Xray absorption fine structure (EXAFS), probes the unoccupied electronic states of the excited state of a chosen atomic species.</p><p>Conversely, relaxation to fill the core hole results in either nonradiative (Auger) or radiative processes. The latter results in the emission of X-ray fluorescence that can be finely characterized by XES for insight into the occupied electronic states. <ref type="bibr">[46]</ref><ref type="bibr">[47]</ref><ref type="bibr">[48]</ref> Often discussed as complementary to XANES in information content, valence-to-core XES (VtC-XES) is produced when electrons deexcite from the valence shell to fill the core hole, giving direct information about the occupied electronic states involved in bonding. <ref type="bibr">49,</ref><ref type="bibr">50</ref> While XAS and XES have traditionally been synchrotron-based methods, we note that their access, including for VtC-XES, is now being steadily augmented with a renaissance of laboratory-based spectrometers, <ref type="bibr">[51]</ref><ref type="bibr">[52]</ref><ref type="bibr">[53]</ref> including in studies of sufficient scale for data science methods. <ref type="bibr">54</ref> In the first study to use supervised ML in XAS, Timoshenko et al. <ref type="bibr">20</ref> successfully inferred coordination numbers of Pt nanoclusters from XANES spectra using a neural network, a result that would otherwise require (human) expert analysis of EXAFS. Zheng et al. <ref type="bibr">24</ref> also predicted coordination, except using a random forest model. Notably, Torrisi et al. <ref type="bibr">36</ref> likewise used a random forest model to correlate polynomial fitting parameters of spectra to properties like bond distance. Other works utilizing both supervised and unsupervised machine learning in XAS include a XANES matching algorithm, 25 hierarchical clustering on spectra, <ref type="bibr">26</ref> and use of an autoencoder to correlate coordination to a reduced-dimensional representation of spectra. <ref type="bibr">27</ref> Most of these studies assumed that desired information was in fact encoded in spectra, largely because of hand-crafting relevant training datasets. However, our approach (Figure <ref type="figure">1</ref>), via the unsupervised machine learning precursor, allows for explorative and unbiased refinement of chemical descriptors&#65533;a step that we propose is necessary, and likely sufficient, when addressing much more complex datasets.</p><p>The present study is prompted by our recent work <ref type="bibr">55</ref> that compared the variance and information content of sulfur Kedge XANES to VtC-XES K&#946; spectra for sulforganics. We found that nonlinear dimensionality reduction algorithms, a subset of unsupervised ML, provided an effective way to extract spectral features and thus important chemical information encoded in spectra. Moreover, our results exemplified the benefits of utilizing unsupervised ML to mold and understand the full potential of supervised ML analysis. <ref type="bibr">56</ref> Here, we investigate the information content and sensitivity of phosphorus K-edge XANES and VtC-XES K&#946; in a more complex chemical system, organophosphorus compounds, and indeed find sensitivity to a wider range of chemical properties, including coordination, oxidation, aromaticity, intramolecular hydrogen bonding, and ligand identity. The proximity of phosphorus to sulfur in the periodic table allows for the same theoretical parameters to generate spectra (and thus obtain similar experimental agreement) as our previous study and also leverages the more diverse bonding environment of phosphorus. The dataset of spectra is calculated from molecular structures gathered from the PubChem 57 database using moldl, a new open-source tool that we have developed for this purpose. <ref type="bibr">58</ref> For the rest of this paper, we will refer to the phosphorus K-edge XANES and VtC-XES K&#946; as just XANES and VtC-XES, respectively, for brevity.</p><p>Organophosphorus compounds have much higher total variance than sulforganics, as well as higher variance within the same bonding geometry. We can therefore tune the input domain to account for these highly variant structures, allowing us to understand the sensitivity of these spectra to a wider range of properties. In addition, we can find, in an unbiased way, the extent of the chemically relevant information that may be extracted using dimensionality reduction algorithms, especially when confined to very limited dimensions. These explorations allow for full utilization of real spectral information during supervised ML predictions.</p><p>To this end, we use the Uniform Manifold Approximation and Projection (UMAP) <ref type="bibr">59</ref> for dimensionality reduction, which allows us to develop chemical classes by examining clustering of spectra in a two-dimensional embedding. UMAP is a nonlinear embedding similar to t-distributed Stochastic Neighbor Embedding (t-SNE), <ref type="bibr">60</ref> which was used in our recent work <ref type="bibr">55</ref> to extract chemical classes. Like t-SNE, UMAP constructs a graph-based representation of the data in the highdimensional space to generate a similarity comparison, and then it tries to match the similarity comparison in a lowdimensional representation of the data. However, UMAP utilizes a different cost function, namely, cross-entropy instead of KL divergence, which further enables the global structure to be preserved, albeit at the cost of the "crowding problem". <ref type="bibr">60</ref> Moreover, given the proper choice in hyperparameters, UMAP can retain global similarity such that distances between clusters can be interpreted (given the manifold remains connected). This contrasts t-SNE, where its cost function, Flowchart of an analysis framework that uses unsupervised machine learning (such as cluster analysis) as a precursor to predictions on spectra via supervised machine learning, which can then inform experimental design and data creation.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>The Journal of Physical Chemistry A</head><p>the KL divergence, goes to zero at large distances. The result is that t-SNE is not penalized for putting unalike data either far or very far away, and thus interpretation of similarity is only valid on a relatively local (intracluster) scale. These properties of UMAP allow it to generate a mapping function that can then be used to map subsequent data, which is why UMAP is called a "parametric embedding" and contrasts t-SNE's requirement that the entire training dataset must be used to predict new data. Thus, UMAP can be used for future data compression and has the potential for better interpretation of overall global similarities. These advantages have led to its recent popularity, such as in single-cell RNA sequencing (scRNA-seq) data analysis, <ref type="bibr">61</ref> but UMAP has not yet seen use in XAS analysis.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>II. METHODS</head><p>Our methods for the electronic structure calculations closely follow that of Tetef et al. <ref type="bibr">55</ref> Molecular structures were downloaded from the PubChem database using our opensource Python module called moldl 58 that allows for users to easily write scripts that can search the PubChem database and store the resulting structures, with metadata, in a local database indexed by PubChem Compound IDentification (CID) numbers. The downloaded structures can then be sorted using customizable filters, and selected molecules can be exported in multiple formats (SDF, MOL, and XYZ). This tool is accessible to any researcher for use in projects that require the collection and management of molecular structure datasets. A total of 1196 compounds were downloaded and managed in this study, while 756 of them were structurally viable for our desired analysis.</p><p>Both XANES and VtC-XES were calculated with the opensource NWChem computational chemistry software package 62,63 via the same pipeline as specified in Tetef et al. <ref type="bibr">55</ref> To summarize, both spectra were computed using the Sapporo QZP-2012 basis set <ref type="bibr">64</ref> for P, while the remaining atoms were represented using the 6-31G* basis set and the PBE0 exchange correlation functional. <ref type="bibr">65</ref> Additionally, the Stuttgart RLC ECP <ref type="bibr">66</ref> was substituted for atoms heavier than phosphorus. As in Tetef et al., <ref type="bibr">55</ref> a post-processing energy-dependent linear broadening scheme was applied to XANES transitions, starting with a full width half-maximum (FWHM) Lorentz broadening of 0.5 eV at the whiteline, and then linearly increasing to 4.0 eV FWHM at 20 eV past the whiteline. An energy shift of 50 eV was applied to all XANES transitions to align with experimental data. <ref type="bibr">67</ref> For VtC-XES, the calculated transitions were all shifted by -19 eV to align to the experiment. <ref type="bibr">68</ref> A FWHM Lorentz broadening of 0.5 eV and a FWHM Gaussian broadening of 1.5 eV were added to each transition to agree with experimental data. <ref type="bibr">68</ref> Because NWChem calculates a selfconsistent field density functional theory (DFT) solution for both XANES <ref type="bibr">69</ref> and VtC-XES, <ref type="bibr">70</ref> this solution serves as a reference for the time-dependent DFT (TDDFT)-based X-ray spectroscopy calculations and thus only one internally consistent energy shift is required for each system.</p><p>Finally, both XANES and VtC-XES were individually normalized by their total K&#945; intensities. The K&#945; transitions scale in intensity proportional to the compound size (like VtC-XES and XANES calculations) but are very nearly independent of all environmental effects, thus providing an absolute scale to maintain relative intensities across the entire ensemble.</p><p>The sulforganics study of Holden et al. <ref type="bibr">54</ref> for the experimental VtC-XES and NWChem calculations showed excellent agreement, as did additional calculations and comparison to XANES in Tetef et al. <ref type="bibr">55</ref> Here, in Figures <ref type="figure">S1</ref> and<ref type="figure">S2</ref>, we more modestly validate the performance of NWChem against several VtC-XES taken with the same instrument and methodology as Holden et al., <ref type="bibr">54</ref> and also </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>The Journal of Physical Chemistry A</head><p>validate the performance against several XANES spectra from Persson et al. <ref type="bibr">67</ref> Briefly, Principal Component Analysis (PCA) was implemented using the scikit-learn 71 package in Python and was applied to the original spectra before UMAP to speed up computation and decrease noise. The number of principal components kept from the PCA was the number of components necessary to explain at least 95% of the variance in the dataset. For example, the number of retained principal components was 7 and 14 components for VtC-XES and XANES spectra, respectively, for the dataset consisting of all tricoordinate and tetracoordinate compounds, as shown in Figure <ref type="figure">S3</ref>. Some reconstructed spectra using the 95% variance cutoff are shown in Figures <ref type="figure">S4</ref> and<ref type="figure">S5</ref>. The difference in the number of principal components required for VtC-XES and XANES suggests that XANES spectra have more variation and thus more nonlinear features, which is unsurprising.</p><p>UMAP was implemented using the umap-learn module <ref type="bibr">72</ref> with default hyperparameters. Again, as mentioned above, to accelerate computing, the UMAP algorithm was applied to the PCA coefficients at the 95% variance level, thus decreasing the dimensionality of the training space from 1000 to either 7 or 14. For Figures <ref type="figure">2</ref><ref type="figure">3</ref><ref type="figure">4</ref><ref type="figure">5</ref><ref type="figure">6</ref>, the number of UMAP output components was constrained to two for visualization purposes, while for Figure <ref type="figure">7</ref>, the output dimensionality was set to five (found through the hyperparameter optimization discussed below).</p><p>Finally, to help illustrate the value of unsupervised ML as a precursor to supervised ML, we applied supervised ML in the form of a Gaussian process <ref type="bibr">73</ref> classifier to the UMAP representation for all five classification schemes determined by the two-dimensional cluster analysis. The Gaussian processes were implemented using scikit-learn, <ref type="bibr">71</ref> which utilizes the Laplace approximation as detailed by Rasmussen and Williams. <ref type="bibr">73</ref> A separate classifier was trained for each of the five classification schemes, shown in Table <ref type="table">S1</ref>, for both VtC-XES and XANES data.</p><p>A test set was specified for each classifier, which comprised of a random selection of 15% within each class, with the rest of the data specified as training. A validation set was then randomly selected within that training set to optimize model hyperparameters. These hyperparameters were found to be five dimensions for the UMAP embedding, with the optimal kernels for the Gaussian Process selected as Rational Quadratic for both VtC-XES and XANES spectra. All data and analysis code for this study is publicly available. 74</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>III. RESULTS AND DISCUSSION</head><p>The first two sections below follow the general approach in Tetef et al. <ref type="bibr">55</ref> wherein we investigate the heuristically expected chemical sensitivities in VtC-XES and XANES (Section III.I) and then, when subclusters are observed within an expectedly dominant chemical classification, we investigate unexpected sensitivities to further structure or electronic refinements (Section III.II). This has several important results, including delineation of both similar and different sensitivities of VtC-XES and XANES to chemical classifications, as well as the emergence of spectral sensitivity to second-shell coordination for phosphates.</p><p>The final section (Section III.III), on the other hand, seeks to address the motivating hypothesis illustrated in Figure <ref type="figure">1</ref>, i.e., unsupervised ML can usefully inform supervised ML. We demonstrate that confidence of predictions directly correlates to our qualitative cluster analysis, thus validating that the strength of information encoded in VtC-XES and XANES can vary between spectroscopies, depending on the system and property of interest.</p><p>III.I. Unbiased Verification of Heuristic Classes. To begin, heuristically, one expects phosphorus coordination to yield the strongest distinguishing features between spectra, specifically the distinction between tricoordinate phosphorus The Journal of Physical Chemistry A and tetracoordinate phosphorus. Not only do these coordination geometries have different hybridized orbital characters but they are also often a proxy for the oxidation state. In organophosphorus compounds with tricoordinate phosphorus centers, phosphorus is typically in a 3+ oxidation state, whereas compounds with tetracoordinate phosphorus centers usually have phosphorus in a 5+ oxidation state.</p><p>We chose compounds with a diverse number of oxygens bonded to phosphorus within these two coordination configurations (with all other bonding atoms as carbon) to further vary the effective charge on phosphorus. We then applied UMAP to VtC-XES and XANES spectra to create a two-dimensional embedding of the ensemble. The results are color-coded based on whether the compound includes tricoordinate phosphorus or tetracoordinate phosphorus, as shown in Figure <ref type="figure">2</ref>. All R groups bonded to phosphorus (or bonded to the oxygens bonded to phosphorus) are constrained to be exclusively carbons (e.g., alkyl or aryl chains), and sometimes hydrogens (when bound to the oxygen) to achieve hydroxyl groups, but only for phosphates (which we will explore later).</p><p>As expected, coordination distinguishes most of the groupings of the compounds, with a handful of outliers. We have further labeled some example compounds a-h (right panels) in each cluster with their corresponding VtC-XES and XANES spectra. (The identity of compounds a-h is defined in Table <ref type="table">S2</ref> in the supplementary section, but it is sufficient to say that they span a wide range of local coordination and oxidation states.) Note how some compounds that are in the same cluster in the VtC-XES embedding are in different clusters in the XANES embedding, and vice versa. For example, compounds a, b, and c are together in the XANES embedding, but they are in three different clusters in the VtC-XES embedding, which we will discuss later as being due to the number of oxygen ligands. These observations clearly indicate that VtC-XES and XANES encode information differently and that there are chemically relevant subgroupings within each coordination.</p><p>Seeking to elucidate the chemical subgroupings, Figure <ref type="figure">3</ref> shows the embedding color-coded within each of the tri-and tetracoordinate classes based on the number of oxygens bonded to phosphorus and the corresponding named chemical classifications. The spectral averages for both VtC-XES and XANES spectra for each class are shown in Figure <ref type="figure">S6</ref>, while the spectral averages for each cluster are shown in Figure <ref type="figure">S7</ref>. Figure <ref type="figure">3</ref> shows very clear retention of chemically relevant information, with some similarities and differences between VtC-XES and XANES. We will now discuss the expected change in spectra based on the chemical signatures in this ensemble, the resulting successes in information encoding, the differences between the two spectroscopies, and (importantly) the occurrence of outliers in the UMAP embedding, specifically, if the outliers correspond to molecules whose electronic structure is somehow strongly anomalous with respect to their general chemical class.</p><p>First, we expect effective charge of phosphorus to have the biggest impact on both VtC-XES and XANES spectra. For VtC-XES, the ligand peaks (the small low-energy peak in Figure <ref type="figure">S6</ref>) will increase in both energy and intensity with an increase in phosphorus oxidation. From a molecular orbital perspective, this trend is from both a larger overlap between the ligand valence orbital and the phosphorus 3p orbital (valence shell) and the increased number of oxygen ligands. In general, this feature (which also changes with different ligand symmetries and orientation) is why VtC-XES is strongly sensitive to ligand identity. <ref type="bibr">75</ref> For XANES spectra, an increase in the oxidation of phosphorus, i.e., the number of oxygen ligands within a coordination, will cause a blue shift of the absorption edge, also demonstrated again by the average spectra in Figure <ref type="figure">S6</ref>.</p><p>Second, in terms of successful information encoding, we see that the number of oxygen ligands supplies much more information to explain the groupings in the UMAP representation than just coordination. For example, the highest oxidation compounds&#65533;the phosphates (blue)&#65533;are separated from all other compounds in both VtC-XES and XANES embeddings and are even subdivided into two clusters for both (this is due to a combination of chemical properties, which we will explore later in (Section III.II) and is the reason compounds e and d are separated in the XANES embedding but not the VtC-XES embedding).</p><p>Third, we consider the similarities and differences of information encoding by XANES and VtC-XES in Figure <ref type="figure">3</ref>. In terms of differences, VtC-XES segregated the tetracoordinate phosphonates (yellow) from other compounds, whereas XANES segregated the tricoordinate trialkyl phosphines (orange) from the rest of the ensemble. Additionally, VtC-XES separated the phosphine oxides (pink) into two subclusters not seen in the XANES embedding, while the tricoordinate phosphite esters (red) get their own cluster in XANES but not in the VtC-XES embedding.</p><p>A closer look at these differences in the UMAP embeddings for VtC-XES and XANES is exemplified by the example compounds a-c. In this case, although compound b (tetracoordinate, phosphine oxide, one oxygen ligand) has a more reduced P atom compared to compounds a and c (both tetracoordinate, phosphates, four oxygen ligands), it is in a different cluster in the VtC-XES embedding but in the same cluster, albeit at the opposite end, as compounds a and c in the XANES embedding. We see that VtC-XES for compound b is in fact vastly different than the spectra of a and c, but its XANES counterpart is more similar to the others. This difference is grouping is likely indicative of the variation within the two spectroscopies. Because UMAP compares both local and global similarities between spectra, this trend might indicate that VtC-XES have more discrete spectral features (especially regarding the charge on phosphorus) compared to a continuous variation in XANES spectral features (for example, a continuous shift in the absorption edge).</p><p>Finally, moving to apparent outliers, one clear example is the location of compound a, diethyl (chloromethyl) phosphonate in the VtC-XES embedding. Compound a is a phosphonate but has a chlorinated carbon ligand, which effectively pulls more charge from phosphorus, thus making the carbon act more like an oxygen and the phosphorus having a higher oxidation. Likewise, both phosphonates, like compound a, in the phosphate cluster having a chlorinated R 1 ligand are thus grouped with the nominally "higher oxidation" compounds instead of the cluster with compound c (diethyl methanephosphonate).</p><p>As for further outliers, note that although compound f is a phosphinite with nominal P(III) oxidation from its tricoordinate P, it has a distinct number of oxygen ligands (one) compared to g (trialkyl phosphine, tricoordinate, no O ligands) and h (phosphite ester, tricoordinate, three O ligands). In these UMAP embeddings, compound f is more</p><p>The Journal of Physical Chemistry A similar to the higher oxidation compounds in VtC-XES compared to XANES spectra. Upon further examination, the other trialkyl phosphines in the cluster with compound f in the VtC-XES embedding are anomalous&#65533;they all have nitrile functional groups bonded to the phosphorus atom. Thus, in this case, VtC-XES seems to determine outliers more definitively than XANES, where the distinction falls on the second nearest-neighbor identity.</p><p>These observations bring us to our next hypothesis that VtC-XES and XANES are both sensitive to ligand identity. As stated earlier, VtC-XES is highly sensitive to ligand identity via changes in the ligand peak feature. <ref type="bibr">47</ref> Again, because the absorption edge of a XANES spectrum shifts with oxidation, the electronegativity of ligands will cause the biggest spectral change. However, even for ligands with approximately the same electronegativity, different phase shifts and cross sections cause finer changes to XANES spectra.</p><p>To systematically probe the effect of ligand identity, a series of tetracoordinate phosphorus compounds (phosphates) were evaluated, in which the oxygen substituents were replaced with one or two sulfur atoms with the local bonding environment around the phosphorus otherwise unchanged. Compared to oxygen, sulfur is significantly less electronegative, with a Pauling electronegativity value near that of carbon and phosphorus. <ref type="bibr">76</ref> Thus, while differences in photoelectron scattering can influence the XANES, we generally expect that these oxygen-to-sulfur ligand substitutions cause the biggest spectral change by adjusting the effective charge on the phosphorous. The resulting clusters are shown in Figure <ref type="figure">4</ref>. Note that the phosphates are the same compounds that were used in the ensemble appearing in Figures <ref type="figure">2</ref> and<ref type="figure">3</ref>, but that we have added additional chemical classes&#65533;phosphorothioates and dithiophosphates&#65533;to create the ensemble appearing in Figure <ref type="figure">4</ref>.  </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>The Journal of Physical Chemistry A</head><p>The different ligand identities drive cluster separations in Figure <ref type="figure">4</ref>, but do not exhaust the refinement of chemical classification; we return below to the question of further classification within phosphates. However, in Figure <ref type="figure">4</ref>, VtC-XES has a clear outlier&#65533;the phosphorothioate (green) in the dithiophosphate cluster (red) in the second inset of that figure. Chemically, this compound (PubChem CID 104781, tertbutylbicyclophosphorothionate) is structurally different from others because the oxygens form one edge of a carbon tetrahedrane. Thus, a clear chemical outlier, in terms of electronic structure, is also flagged as an outlier in the UMAP embedding because UMAP grouped this compound with dithiophosphates instead of with phosphorothioates.</p><p>We next analyze whether the spectra would be sensitive to substitutions of R groups (if bonded to an oxygen) with a hydrogen atom, thus forming hydroxyl groups, as shown in Figure <ref type="figure">5</ref>. Here, we have taken phosphinate and phosphonate as starting points, and consecutively replaced O-R groups with OH groups. Note that the phosphinates and phosphonates are the same compounds that were used in the ensemble appearing in Figures <ref type="figure">2</ref> and<ref type="figure">3</ref>, but that we have added additional chemical classes&#65533;phosphenic acids, half phosphonic acids, and phosphonic acids&#65533;to create the ensemble appearing in Figure <ref type="figure">5</ref>.</p><p>In general, this distinction seems to be better illuminated by VtC-XES than the XANES (which is shown in Figure <ref type="figure">S8</ref>), as the clustering in VtC-XES is suggestive of a sensitivity to hydroxyl groups. However, Figure <ref type="figure">5</ref> also exemplifies that first nearest neighbors, e.g., the oxygen ligands directly bonded to phosphorus, likely cause the biggest spectral changes and thus are the biggest contributing factor to clustering, which is consistent with our earlier observations. III.II. Emergent Chemical Fingerprints from Clusters. Above, we motivated our classes by important chemical properties that we heuristically expected to yield the biggest spectral differences. However, even within this chemically driven framework, there are subclusters within our heuristic chemical classes that are instead emergent from UMAP. For example, we found that subclustering of the phosphate chemical class (exemplified by multiple separate subclusters in Figures <ref type="figure">3</ref> and<ref type="figure">4</ref>) was caused by unexpected variations in the secondary substituent (atoms bound to oxygens, not directly to phosphorus), indicating that XANES spectra are sensitive to even more subtle details than anticipated.</p><p>Let us examine this subdivision of the phosphates, specifically in the UMAP embedding of their XANES spectra. For just phosphates, we achieve the embedding shown in Figure <ref type="figure">6</ref>, which has labeled the phosphates into four clusters determined by the dbscan 77 clustering algorithm: I, II, III, and IV. The average spectrum for each cluster is shown at the bottom and the common structural motifs for each cluster are shown to the right.</p><p>77% of Cluster I is comprised of compounds with two alkyl R groups and the third group either alkyl or aryl rings. This distinction is different from Clusters II to IV as they instead typically have two R groups as H atoms instead of carbonbased groups. Cluster II is the largest subcluster and 94% of the compounds have two hydroxyl groups bonded to phosphorus and the last R group an alkyl chain. These two clusters are the most distinct.</p><p>On the other hand, Clusters III and IV are similar in composition. Cluster III is comprised of compounds with the third R group as: (a) alkyl rings or cycloalkanes (36%), (b) aromatic rings (23%), or (c) take part in intramolecular hydrogen bonding with one of the hydroxyl groups bonding to phosphorus. Cluster IV compounds are structurally very similar to Cluster III compounds, even though their spectra are distinct. However, 54% of Cluster IV compounds have their third R group as aromatic rings. For some example compounds in each cluster along with their spectra and structure, see Figures <ref type="figure">S9-S12</ref>. All compounds in Clusters I-IV can also be viewed in Figures <ref type="figure">S13-S16</ref>. Additionally, given the linear nature of Clusters I, III, and IV in the UMAP embedding, we tested the correlation between the embedding location and the energy of the absorption edge, as shown in Figure <ref type="figure">S17</ref>, and found no strong correlation.</p><p>Furthermore, color-coding the phosphates based on a 10dimensional clustering and then visualizing them in two dimensions yields very nearly the same classifications, as shown in Figure <ref type="figure">S18</ref>. Thus, the two-dimensional embedding is retaining enough information to categorize the phosphates appropriately. Even expanding the embedding space to three dimensions instead of two for all previous embeddings yields very nearly the same clustering, as shown in Figure <ref type="figure">S19</ref>. This retention in information&#65533;yet complex clustering of com-pounds&#65533;further supports the nonlinear nature of spectra and the idea that properties are complexly encoded in spectra and conversely, spectral features do not correlate solely to a single, high-variant attribute but rather a combination of electronic or chemical properties.</p><p>Taken en masse, these results show the extent to which chemically relevant information is, or is not, encoded by the quantum mechanics involved in XANES and VtC-XES. As to the specific algorithm, UMAP can be used iteratively as more data is collected. Thus, it has the potential to shown evolutions through the domain space, similar to the latent space of a variational autoencoder (VAE), <ref type="bibr">78</ref> given proper tuning of its two hyperparameters: the number of expected neighbors in a cluster and the minimum distance between points. For an overview of the effect of those two hyperparameters on the </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>The Journal of Physical Chemistry A</head><p>UMAP embedding, see Figures <ref type="figure">S20</ref> and<ref type="figure">S21</ref>. Finally, and of key importance here, UMAP can generate embeddings of spectra that can be used for unbiased refinement of the training dataset in addition to a preprocessing step before supervised ML predictions.</p><p>III.III. Validation of Chemical Fingerprints from Cluster Analysis. In our prior work on sulforganics <ref type="bibr">55</ref> and in the present above work on the more complex case of organophosphorus compounds, we have demonstrated a convincing utility of advanced, nonlinear unsupervised ML tools for evaluating the chemically relevant information in VtC-XES and XANES spectra. We now return to our hypothesis presented in the introduction and illustrated in Figure <ref type="figure">1</ref>, where we propose that such an unsupervised ML method can productively inform the use of supervised ML tasks.</p><p>The most common use of supervised ML in X-ray spectroscopy is to predict numerical properties, such as bond length or coordination, from XANES spectra. <ref type="bibr">[20]</ref><ref type="bibr">[21]</ref><ref type="bibr">[22]</ref><ref type="bibr">24,</ref><ref type="bibr">31</ref> Here, we instead predict chemical classes from both VtC-XES and XANES spectra. Moreover, we predict these classes from a fivedimensional UMAP representation of the spectra instead of from the original spectra themselves. Such preprocessing through dimensionality reduction can help separate inherently correlated and nonlinear spectral features <ref type="bibr">56</ref> as well as greatly reduce both the computational cost and the effect of spectral noise.</p><p>Furthermore, we use a Gaussian process (GP) to incorporate prior knowledge into our models and generate an informed predictor. <ref type="bibr">73</ref> A GP is a nonparametric kernel method that formally incorporates Bayes rule into the model, which not only allows for priors to be specified training but also allows for a probabilistic interpretation of the results. This probability gives uncertainty estimates, or conversely confidence, of the predictions. We note that one of the biggest downsides of a GP is that it scales poorly, which is another reason why applying a nonlinear dimensionality reduction routine like UMAP beforehand can transform this problem into a computationally tractable one.</p><p>The results of training a GP on each of the five classification schemes (see Table <ref type="table">S1</ref>) we developed&#65533;coordination, number of oxygen ligands, phosphate subcluster, number of sulfur ligands, and number of hydroxyl ligands&#65533;are shown in Figure <ref type="figure">7</ref>, with the average accuracy score on the test set as well as the probability of that prediction, i.e., the confidence score, shown.</p><p>There is a clear correlation between the average accuracy and confidence, indicating that the GP is, in fact, properly modeling uncertainty of predictions.</p><p>Finally, the accuracies and confidence of each prediction across VtC-XES and XANES data match what we observed in our two-dimensional UMAP figures. This correlation is clearly demonstrated in the hydroxyl ligand and phosphate subcluster classification schemes, where the XANES and VtC-XES, respectively, poorly cluster by these schemes, and the low corresponding GP confidence reflects this. Overall, these results further validate that visualizing data via a dimensionality reduction algorithm like UMAP correlates to extractable information content and can properly inform classes to be used for supervised ML.</p><p>However, we note that care must be taken to ensure transferability when training any supervised ML model on theoretical spectra to then make predictions on experimental data, the obvious next step of our GPs. Ensuring transferability might mean appropriately modeling for noise, the spectral line shape, or any systematic errors in the theoretical model.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>IV. CONCLUSIONS</head><p>By utilizing Uniform Manifold Approximation and Projection (UMAP) and analyzing the resulting clustering in a twodimensional embedding of VtC-XES and XANES spectra of an ensemble of organophosphorus compounds, we find sensitivity to coordination and ligand identity (specifically by distinguishing the number of oxygen ligands, sulfur ligands, and hydroxyl groups). Additionally, the XANES was clearly more sensitive to phosphate subgroupings due to an unexpected, unintuitive fingerprint that emerged from the clustering in the unsupervised machine learning tool, UMAP.</p><p>These results culminate in a valuable analysis framework: (1) applying nonlinear dimensionality reduction routines and cluster analysis to check for both heuristic chemical sensitivities and emergent ones present in the spectra, (2) applying dimensionality reduction methods like UMAP before querying supervised ML models, and (3) utilizing models that incorporate prior knowledge, such as a Gaussian process, to estimate uncertainty or confidence of these predictions on the clustering-informed classes. Furthermore, this framework, which we call Unsupervised Validation of Classes (UVC) and illustrate in Figure <ref type="figure">1</ref>, is broadly applicable&#65533;it can easily be expanded to both other systems and other spectroscopies&#65533; The Supporting Information is available free of charge at <ref type="url">https://pubs.acs.org/doi/10.1021/acs.jpca.2c03635</ref>.</p><p>Theory versus experiment: VtC-XES (Figure <ref type="figure">S1</ref>); theory versus experiment: XANES (Figure <ref type="figure">S2</ref>); scree plot of VtC-XES and XANES data (Figure <ref type="figure">S3</ref>); PCA reconstruction of VtC-XES (Figure <ref type="figure">S4</ref>); PCA reconstruction of XANES spectra (Figure <ref type="figure">S5</ref>); class averages of spectra with different coordinations (Figure <ref type="figure">S6</ref>); cluster averages of spectra with different coordinations (Figure <ref type="figure">S7</ref>); UMAP representation of XANES with H atom substitutions (Figure <ref type="figure">S8</ref>); phosphate subcluster I example spectra (Figure <ref type="figure">S9</ref>); phosphate subcluster II example spectra (Figure <ref type="figure">S10</ref>); phosphate subcluster III example spectra (Figure <ref type="figure">S11</ref>); phosphate subcluster IV example spectra (Figure <ref type="figure">S12</ref>); phosphate subcluster I structures (Figure <ref type="figure">S13</ref>); phosphate subcluster II structures (Figure <ref type="figure">S14</ref>); phosphate subcluster III structures (Figure <ref type="figure">S15</ref>); phosphate subcluster IV structures (Figure <ref type="figure">S16</ref>); phosphate subcluster correlations (Figure <ref type="figure">S17</ref>); phosphate subclusters: 10-dimensional clustering (Figure <ref type="figure">S18</ref>); 3D UMAP visualizations (Figure <ref type="figure">S19</ref>); changing UMAP hyperparameters: number of neighbors (Figure <ref type="figure">S20</ref>); changing UMAP hyperparameters: minimum distance (Figure <ref type="figure">S21</ref>); classification table (Table <ref type="table">S1</ref>); and table for compounds a-h (Table <ref type="table">S2</ref>) (PDF)</p><p>&#9632; </p></div><note xmlns="http://www.tei-c.org/ns/1.0" place="foot" xml:id="foot_0"><p>https://doi.org/10.1021/acs.jpca.2c03635 J. Phys. Chem. A 2022, 126, 4862-4872 Downloaded via UNIV OF WASHINGTON on September 1, 2022 at 19:30:19 (UTC).See https://pubs.acs.org/sharingguidelines for options on how to legitimately share published articles.</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" xml:id="foot_1"><p>https://doi.org/10.1021/acs.jpca.2c03635 J. Phys. Chem. A 2022, 126, 4862-4872</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" xml:id="foot_2"><p>The Journal of Physical Chemistry A</p></note>
		</body>
		</text>
</TEI>
