<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>The eROSITA Final Equatorial-Depth Survey (eFEDS): Identification and characterization of the counterparts to point-like sources</title></titleStmt>
			<publicationStmt>
				<publisher></publisher>
				<date>05/01/2022</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10332760</idno>
					<idno type="doi">10.1051/0004-6361/202141631</idno>
					<title level='j'>Astronomy &amp; Astrophysics</title>
<idno>0004-6361</idno>
<biblScope unit="volume">661</biblScope>
<biblScope unit="issue"></biblScope>					

					<author>M. Salvato</author><author>J. Wolf</author><author>T. Dwelly</author><author>A. Georgakakis</author><author>M. Brusa</author><author>A. Merloni</author><author>T. Liu</author><author>Y. Toba</author><author>K. Nandra</author><author>G. Lamer</author><author>J. Buchner</author><author>C. Schneider</author><author>S. Freund</author><author>A. Rau</author><author>A. Schwope</author><author>A. Nishizawa</author><author>M. Klein</author><author>R. Arcodia</author><author>J. Comparat</author><author>B. Musiimenta</author><author>T. Nagao</author><author>H. Brunner</author><author>A. Malyali</author><author>A. Finoguenov</author><author>S. Anderson</author><author>Y. Shen</author><author>H. Ibarra-Medel</author><author>J. Trump</author><author>W. N. Brandt</author><author>C. M. Urry</author><author>C. Rivera</author><author>M. Krumpe</author><author>T. Urrutia</author><author>T. Miyaji</author><author>K. Ichikawa</author><author>D. P. Schneider</author><author>A. Fresco</author><author>T. Boller</author><author>J. Haase</author><author>J. Brownstein</author><author>R. R. Lane</author><author>D. Bizyaev</author><author>C. Nitschelm</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[Context.              In November 2019, eROSITA on board of the Spektrum-Roentgen-Gamma (SRG) observatory started to map the entire sky in X-rays. After the four-year survey program, it will reach a flux limit that is about 25 times deeper than ROSAT. During the SRG performance verification phase, eROSITA observed a contiguous 140 deg              2              area of the sky down to the final depth of the eROSITA all-sky survey (eROSITA Final Equatorial-Depth Survey; eFEDS), with the goal of obtaining a census of the X-ray emitting populations (stars, compact objects, galaxies, clusters of galaxies, and active galactic nuclei) that will be discovered over the entire sky.                                      Aims.              This paper presents the identification of the counterparts to the point sources detected in eFEDS in the main and hard samples and their multi-wavelength properties, including redshift.                                      Methods.              To identifyy the counterparts, we combined the results from two independent methods (              NWAY              and              ASTROMATCH              ), trained on the multi-wavelength properties of a sample of 23k              XMM-Newton              sources detected in the DESI Legacy Imaging Survey DR8. Then spectroscopic redshifts and photometry from ancillary surveys were collated to compute photometric redshifts.                                      Results.              Of the eFEDS sources, 24 774 of 27 369 have reliable counterparts (90.5%) in the main sample and 231 of 246 sourcess (93.9%) have counterparts in the hard sample, including 2514 (3) sources for which a second counterpart is equally likely. By means of reliable spectra,              Gaia              parallaxes, and/or multi-wavelength properties, we have classified the reliable counterparts in both samples into Galactic (2695) and extragalactic sources (22 079). For about 340 of the extragalactic sources, we cannot rule out the possibility that they are unresolved clusters or belong to clusters. Inspection of the distributions of the X-ray sources in various optical/IR colour-magnitude spaces reveal a rich variety of diverse classes of objects. The photometric redshifts are most reliable within the KiDS/VIKING area, where deep near-infrared data are also available.                                      Conclusions.              This paper accompanies the eROSITA early data release of all the observations performed during the performance and verification phase. Together with the catalogues of primary and secondary counterparts to the main and hard samples of the eFEDS survey, this paper releases their multi-wavelength properties and redshifts.]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.">Introduction</head><p>Across the electromagnetic spectrum, sensitive wide-area surveys serve multiple purposes. First and foremost, they help astronomers to draw a map of our cosmic neighbourhood. In doing so, they reveal the inner workings of the Milky Way, the local group, and the filamentary large-scale structure underpinning the distribution of matter. Secondly, by observing and cataloguing large numbers of stars, galaxies, groups, clusters, and superclusters of galaxies that are the main visible tracers of this large-scale structure, wide area surveys also provide new statistical tools for the study of classes and populations of astronomical objects, thus helping astronomers to better understand their life cycles, interactions, and ultimately, their physical properties.</p><p>The data are only available at the CDS via anonymous ftp to cdsarc.u-strasbg.fr <ref type="bibr">(130.79.128.5)</ref> or via <ref type="url">http://cdsarc. u-strasbg.fr/viz-bin/cat/J/A+A/661/A3</ref> X-ray surveys in particular reveal fundamental physical processes that are invisible at other wavelengths. Examples are the hot, diffuse plasma that virialises and thermalises within massive dark matter knots; accretion of matter onto compact objects, both Galactic and extragalactic; and the magnetic coronae of mostly young, fast-rotating stars. These are all phenomena that are accessible by X-ray sensitive instruments.</p><p>extended ROentgen Survey with an Imaging Telescope Array (eROSITA, <ref type="bibr">Predehl et al. 2021</ref>) on board the Spektrum-R&#214;ntgen-Gamma (SRG) mission <ref type="bibr">(Sunyaev et al. 2021)</ref>, was designed to provide sensitive X-ray imaging and spectroscopy over a large field of view, thus unlocking unprecedented capabilities for surveying large areas of the sky to deep flux levels. Moreover, the SRG mission plan includes a long (four years), uninterrupted all-sky survey program (the eROSITA All-Sky Survey: eRASS; <ref type="bibr">Predehl et al. 2021</ref>) capable of detecting millions of X-ray sources for the first time.</p><p>In order to demonstrate these ground-breaking survey capabilities and prepare for the science exploitation of the A&amp;A 661, A3 (2022) Fig. <ref type="figure">1</ref>. eFEDS X-ray and multi-wavelength coverage. The thick blue line shows the outer bound of the region that was searched for X-ray sources. The thin beaded blue line shows the region with at least 500 seconds of effective X-ray exposure depth. We indicate the approximate coverage of several selected surveys that are particularly important for this work: Subaru HSC-Wide (shaded green region), KiDS/VIKING (dashed magenta box), GAMA09-DR3 (hatched red box). The eFEDS field is also covered in several other important surveys that completely (or almost completely) enclose the displayed region: e.g., the Galex all-sky surveys (in the UV), Gaia (in optical), Legacy Survey DR8 (optical combined with Gaia and WISE), VHS, and UKIDSS (in the near-infrared), WISE/NEOWISE-R (in the mid-infrared), and the SDSS (optical imaging and spectroscopy). upcoming all-sky survey, the contiguous 140 square degrees of the eROSITA Final Equatorial-Depth survey (eFEDS; <ref type="bibr">Brunner et al. 2022</ref>) were observed during the SRG calibration and performance verification phase, between 3 and 7 November 2019. The entire field, centred at RA 136 and Dec +2 (see Fig. <ref type="figure">1</ref>), was observed to an approximate depth of &#8764;2.2 ks (&#8764;1.2 ks after correcting for telescope vignetting), corresponding to a limiting flux of F 0.5-2 keV &#8764; 6.5 &#215; 10 -15 erg s -1 cm -2 . The eFEDS field was chosen from among the extragalactic areas with the richest multi-wavelength coverage visible by eROSITA in November 2019. The observations are just about 50% deeper than anticipated for eRASS:8 at the end of the planned four-year program in the ecliptic equatorial region (&#8764;1.1 &#215; 10 -14 erg cm -2 s -1 ; Predehl et al. 2021). In other words, the eFEDS exposure corresponds to roughly the 80th percentile of the expected eRASS:8 exposure distribution over the whole sky. eFEDS therefore is a fair representation of what the final eROSITA all-sky survey will be, enabling scientists to face and solve the challenges that will accompany their work for the duration of the survey.</p><p>As discussed in detail in <ref type="bibr">Brunner et al. (2022)</ref>, the X-ray catalogues generated by the analysis of the eFEDS eROSITA data comprise a main catalogue, with 27 910 sources detected above a detection likelihood of 6 in the most sensitive 0.2-2.3 keV band, and a Hard catalogue, containing 246 sources detected above a detection likelihood of 10 in the less sensitive 2.3-5 keV band.</p><p>In this paper, we focus on the point-like (i.e. with an extension likelihood EXT_LIKE = 0 1 ) X-ray sources in these catalogs (27 369 and 246 for the main and hard sample, respectively) 1 This parameter is obtained from the task srctool of the eSASS software <ref type="bibr">(Brunner et al. 2022)</ref>. and describe the procedure of (i) reliably identifying multiwavelength counterparts to the eROSITA sources, (ii) classifying and characterising their properties, and (iii) providing reliable redshift measurements (spectroscopic when available and photometric otherwise). The identification and determination of the reliability of the counterparts, the computation of the photometric redshifts (photo-z), and the characterisation of the sample follow the same procedure for the main and hard samples. For simplicity, we discuss here specifically only the main sample: the two catalogs overlap for 226 of the 246 hard sources, respectively. While we provide the catalogue of counterparts for all the sources in both samples, the properties of the sources in the hard sample are presented and discussed in Nandra et al. (in prep.). The papers about the X-ray spectral analysis <ref type="bibr">(Liu et</ref>  The structure of the paper is as follows: in Sect. 2 we summarise the availability of ancillary data that were used to identify the X-ray counterparts and photo-z estimates. Sections 3 and 4 describe the methods we used to identify the counterparts, while in Sect. 5, the counterparts are finally assigned. In Sect. 6 the counterparts are separated into Galactic and extragalactic sources using morphological, photometric, and proper motion information.  Notes. For LS8, the required depth for DESI is listed. For Ls8/WISE, the listed depth is taken from Meisner et al. ( <ref type="formula">2019</ref>) and is computed using WISE detected sources. Because the photometry we used is forced photometry at the position of optically detected sources, the depth is higher.</p><p>Fig. <ref type="figure">2</ref>. Illustration of the bandpass and relative transmission curves of the UV, optical, and near-infrared photometry we used to compute photo-z.</p><p>For clarity, we do not show the WISE bandpasses.</p><p>Because of its size, the field is well populated by stars, AGNs, clusters, and nearby galaxies. Each eROSITA working group has developed independent methods for the identification of sources of interest. In Sects. 5.3 and 7.5, a comparison is made with two main source classes: stellar coronal emitters, and clusters of galaxies. The ultimate goal is to consolidate the counterparts and classify them at the same time. Section 6 presents the multi-wavelength properties of the counterparts and characterises their Galactic or extragalactic nature. Section 7 presents and discusses the photo-z computed with Le PHARE <ref type="bibr">(Ilbert et al. 2006;</ref><ref type="bibr">Arnouts et al. 1999)</ref>, including a comparison with DNNz (Nishizawa et al., in prep.), an independent method based on machine-learning. Section 8 describes the released data. The basic properties of the point-source eFEDS population based on redshift, photometry, and X-ray flux are presented in Sect. 9. The conclusions in Sect. 10 close the paper, with a forecast of the results and challenges that we will face when working with data from the eROSITA all-sky survey.</p><p>The description of the catalogs that we release is provided in the appendix, together with the list of templates we used to compute the photo-z. Throughout the paper, we assume AB magnitudes unless stated otherwise. In order to allow direct comparison with existing works from the literature of X-ray surveys, we adopt a flat &#923;CDM cosmology with h = H 0 /[100 km s -1 Mpc -1 ] = 0.7, &#8486; M = 0.3, and &#8486; &#923; = 0.7.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.">Supporting data</head><p>Fo studies of X-ray sources (taken singularly or as a population), the entire spectral energy distribution (SED) needs to be constructed and the redshift needs to be determined. Redshift can only rarely be obtained directly from X-ray spectra. It is instead routinely obtained either via optical or near-infrared spectroscopy or via photometric techniques. However, for this to work, the counterparts to the X-ray sources need to be determined first. Deep and homogeneous multi-wavelength data are therefore a prerequisite for any complete population study of an X-ray survey.</p><p>The main challenge is that in survey mode, eROSITA has a half-energy width (HEW) of 26 arcsec <ref type="foot">2</ref>  . By construction, the eFEDS field is placed in an area that fully encompasses the GAMA09 equatorial field <ref type="bibr">(Driver et al. 2009</ref>) and is rich in ancillary photometric and spectroscopic data <ref type="bibr">(Merloni et al., in prep.)</ref>. We list and describe the surveys we used in more detail below. Table <ref type="table">1</ref> summarises the depth in each filter, and Fig. <ref type="figure">2</ref> shows the coverage of the ancillary data in wavelength.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.1.">Supporting the associations</head><p>The counterparts were identified using the DESI Legacy Imaging Survey DR8 (LS8; <ref type="bibr">Dey et al. 2019</ref>) for various reasons.</p><p>First of all, LS8 covers the field homogeneously and has sufficient depth, based on the expected optical properties of the X-ray population <ref type="bibr">(Merloni et al. 2012</ref>). In addition, the survey together with Gaia also provides the AllWISE tractor (Lang 2014) photometry extracted at the position of the optical sources. Finally, the survey covers 14 000 square degrees of sky, thus providing a sufficient number of sources external to eFEDS that can be used as training and validation samples to test the association (see Sect. 3). The absolute astrometry of the LS8 catalogue is registered to the Gaia DR2 astrometric frame, with residuals typically smaller than 30 milliarcseconds <ref type="foot">3</ref> . While the catalogued positional uncertainties of individual LS8 sources are often much larger than this systematic limit (especially toward fainter magnitudes), they are still extremely small relative to those of the eFEDS X-ray sources and are set to 0.1 for the entire survey. Photometry and parallax measures from Gaia, which are optimised for point-like sources, are ideal for the identification of the stars in our sample. For this purpose, EDR3 <ref type="foot">4</ref> has been used (Gaia Collaboration 2020) instead of the Gaia DR 2 provided by LS8.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.2.">Supporting photometric redshifts</head><p>To compute the photo-z, we used the following data sets:</p><p>GALEX. The NASA satellite GALEX has mapped the entire sky in the far-and near-UV (FUV and NUV) between 2003 and 2012, with a typical depth of 19.9 and 20.8 AB magnitude in FUV and NUV, respectively. We used the catalogue catalogue GR6/7 presented in <ref type="bibr">Bianchi (2014)</ref> that is available via Vizier.</p><p>Kilo-degree Survey (KiDS). <ref type="foot">5</ref> The survey mapped 1350 deg 2 in u, g, r, i bands using VST/OmegaCAM. The same area was also covered by the VISTA Kilo-Degree Infrared Galaxy Survey (VIKING; <ref type="bibr">Edge et al. 2013</ref>) in Z, Y, J, H, K. We used the catalogue presented in <ref type="bibr">Kuijken et al. (2019)</ref>; it has ZYJHK aperture-matched forced photometry to the ugri source positions. About 65 deg 2 of sky are shared between KiDS/VIKING and eFEDS.</p><p>HSC S19A. The Hyper Suprime-Cam (HSC; <ref type="bibr">Miyazaki et al. 2018</ref>) Subaru Strategic Program survey (HSC-SSP; <ref type="bibr">Aihara et al. 2018a</ref>) is an ongoing optical imaging survey with five broadband filters (g-, r-, i-, z-, and y-band) and four narrow-band filters (see <ref type="bibr">Aihara et al. 2018b</ref>). We used S19A wide data obtained from March 2014 to April 2019, which provide forced photometry for the five bands, with the 5&#963; limiting magnitudes as listed in Table <ref type="table">1</ref> (see <ref type="bibr">Aihara et al. 2018b</ref><ref type="bibr">Aihara et al. , 2019))</ref>. The astrometric uncertainty is approximately 40 milliarcseconds in rms.</p><p>VISTA/VHS. The entire southern hemisphere has been observed by VISTA in the near-infrared, and at least for J and Ks, the depth is 30 times the depth of 2MASS <ref type="bibr">(McMahon et al. 2013)</ref>. We used the DR4 data that are available via Vizier.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>WISE.</head><p>The Wide-field Infrared Survey Explorer (WISE; <ref type="bibr">Wright et al. 2010</ref>) scanned the entire sky in the 3.4, 4.6, 12, and 22 &#181;m bands over the course of one year (hereafter W1, W2, W3, and W4). Afterwards, the survey continued with observations in W1 and W2 only. The photometry in W1, W2, W3, and W4 from LS8 includes all five years of publicly available WISE and NEOWISE reactivation <ref type="bibr">(Meisner et al. 2019)</ref>. It was measured using the TRACTOR algorithm <ref type="bibr">(Lang 2014</ref>) at the position of grz detected sources.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.3.">Optical spectroscopy</head><p>The eFEDS field has previously been observed by several spectroscopic surveys, most notably GAMA, SDSS, WiggleZ, 2SLAQ, and LAMOST. Many of the existing spectra are of high enough quality for us to use them for science applications, in particular, when we just need redshift and basic classification (i.e. deciding whether the source is a star, QSO, or galaxy). However, a careful collation and homogenisation of the existing spectroscopy catalogues was first needed to provide a reliable compendium of these data.</p><p>The largest body of spectroscopic redshift information comes from the SDSS survey <ref type="bibr">(York et al. 2000;</ref><ref type="bibr">Gunn et al. 2006;</ref><ref type="bibr">Smee et al. 2013;</ref><ref type="bibr">Abdurro'uf et al. 2022)</ref>, totalling more than 68k spectra of 61k science targets within the outer bounds of the eFEDS field. We collected archival public data from SDSS phases I-IV <ref type="bibr">(Ahumada et al. 2020)</ref>, as well as the results of the recent dedicated SPIDERS (Spectroscopic identifications of eROSITA sources) campaign <ref type="bibr">(Comparat et al. 2020, Merloni et al., in prep.)</ref>, within SDSS-IV <ref type="bibr">(Blanton et al. 2017</ref>) following up eFEDS X-ray sources. A small team of the authors visually inspected all of the SDSS 1D spectra that lie in the vicinity of eFEDS X-ray sources, correcting occasional pipeline failures, and grading the spec-z onto a common normalised quality (NORMQ) scale between 3 and -1. NORMQ can be interpreted as follows: spec-z with NORMQ = 3 are those with 'secure' spectroscopic redshifts, those with NORMQ = 2 are 'not secure' (although a large fraction are expected to be at the correct redshift), spec-z with NORMQ = 1 are 'bad' (e.g. low Signal to Noise. S/N, problematic extraction, dropped fibres), and those with NORMQ = -1 are 'blazar candidates'. For completeness, we also retained SDSS-DR16 spectroscopic redshifts in the eFEDS field that do not lie near eFEDS X-ray detections, but only when they satisfied all the following criteria: SN_MEDIAN_ALL&gt;2.0, ZWARNING = 0, SPECPRIMARY = 1, and 0 &lt; Z_ERR &lt; 0.002. An exhaustive description of the SDSS dataset within the eFEDS field will be presented separately by <ref type="bibr">Merloni et al. (in prep.)</ref>.</p><p>We also gathered published spectroscopic redshifts and classifications (hereafter 'spec-z') from the literature where they overlap the eFEDS footprint. The detailed breakdown is presented in Table <ref type="table">2</ref>. In order to gather spec-z from smaller surveys that might only contribute a few redshifts each, we also queried the Simbad database (as of 5 March 2021; <ref type="bibr">Wenger et al. 2000)</ref> in the vicinity of the eFEDS X-ray counterpart positions.</p><p>For the purposes of this work, we placed greater weight on purity than on completeness. Therefore, where the parent survey catalogues included some metric of quality/reliability, we applied strict criteria to retain only the most secure spec-z information. The filtering criteria applied to the original catalogues and the number of spec-z considered from each catalogue are listed in Table <ref type="table">2</ref>. We assumed that after these quality filtering steps, all the archival spec-z are 'secure' (i.e. NORMQ = 3), except for Simbad, for which we adopted NORMQ = 2. This is meant in this case to be interpreted as 'not yet proven to be secure'.</p><p>All these spec-z were progressively collated into a single catalogue, with a single redshift and classification per sky position, using a match in coordinates between the coordinates listed for the nine input spectroscopic catalogs. An appropriate search radius (in the range 1-3 arcsec) was chosen according to the A3, page 4 of 32 Total unique objects 143 637</p><p>Notes. N specz is the number of spectroscopic redshift measurements that pass the quality threshold (applied to columns provided in the originating catalogue). For Simbad, the number of entries is limited to objects lying within 3 arcsec of the optical coordinates of counterparts to eFEDS sources. Some astrophysical objects appear in two or more redshift catalogues.</p><p>expected positional fidelity and/or fibre sizes associated with each input spectroscopic catalogue. After this de-duplication step, we were left with 143 637 unique entries over the eFEDS field, 108 834 of which are secure (i.e. NORMQ = 3).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.">Counterpart identification: Method</head><p>Because of the large PSF of the eROSITA telescopes and the small number of photons associated with typical X-ray detections, the 1&#963; rms positional uncertainties of individual X-ray sources can be several arcseconds. Specifically, in eFEDS, the mean positional error is 4. 7 and extends above 20 arcsec only for a handful of sources <ref type="bibr">(Brunner et al. 2022</ref>). For the expected optical/infrared magnitude distribution of X-ray sources at the depth of the eFEDS (see e.g. <ref type="bibr">Merloni et al. 2012;</ref><ref type="bibr">Menzel et al. 2016)</ref>, the sky density of the relevant astrophysical source populations is often high. For this reason, the identification of the true associations cannot be determined solely by closestneighbour searches, as there will be several potential counterparts within the error circle of any given X-ray source. Taking this into account, the identification of the counterparts of eFEDS point-like sources has been performed using two independent methods. NWAY <ref type="bibr">(Salvato et al. 2019</ref>), based on Bayesian statistics, and ASTROMATCH <ref type="bibr">(Ruiz et al. 2018</ref>), based on the maximum likelihood ratio (MLR; Sutherland &amp; Saunders 1992), have been specifically developed to identify the correct counterparts to X-ray sources, independently of their Galactic or extragalactic nature. In order to assess the probability (or likelihood) of an object to be the correct counterpart to an eFEDS sources, the two methods first take the separation between the sources, their positional accuracy and the number density of the sources in the ancillary data into account. The difference between the methods resides then in the adoption of specific features that are able to distinguish an X-ray emitter (regardless of its Galactic or extragalactic nature) from a random source in the field. Both methods determined the features (priors) using a representative training sample constructed using secure counterparts to X-ray sources detected in 3XMM. In the case of NWAY, the prior was also determined by comparing the features of the sources in the training sample with the features of the field sources present within 30 arcsec from the 3XMM sources. The disentangling power of the priors was then tested on a blind validation sample of 3415 Chandra sources with secure counterparts, where the accuracy of the Chandra position was made eFEDS-like (see Sect. 4.1). The detailed description of the construction of training, validation, and respective field sample is presented in Appendix A.</p><p>Here we provide a short description of NWAY and ASTROMATCH and how their respective priors were determined.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.1.">NWAY enhanced with photometric priors defined via machine-learning</head><p>In addition to astrometry, that is, (a) the separation between an X-ray source and a candidate counterpart; (b) the associated positional uncertainties and (c) the number densities of the sources in the two catalogs, the photometry of potential counterparts is valuable information to determine whether they are associated with a given X-ray detection. Traditionally, the likelihood ratio associated with angular distance was multiplied by a factor accounting for the magnitude distributions and the sky density of a population of X-ray sources and background objects (e.g. </p><p>, where D &#966; and D m refer to the astrometric and photometric information, respectively. For any possible association, the modifying factor P(D m | H) is computed from the feature (e.g. magnitude or colour) m of the counterpart candidate and from the expected distribution of this observable for X-ray sources and field (non-X-ray) sources. We call such factors "priors" to NWAY, as they enter as a priori information in the ultimate matching process. These priors are posteriors previously learned from other data. For further details of the formalism, we refer to Salvato et al. (2019) and the NWAY documentation <ref type="foot">6</ref> .</p><p>In order to take full advantage of the LS8 ancillary catalogue, we extended this approach for the eFEDS counterpart identification. Instead of using a subset of magnitudes, colours, and their associated distributions, we trained a random forest By construction, the training sample is highly imbalanced, since the field objects strongly outnumber the X-ray sources. We therefore opted for a weighting scheme that automatically adjusted weights of training examples for the class imbalance.</p><p>The trained model was evaluated on the test set, resulting in the confusion matrix presented in Fig. <ref type="figure">3</ref>. We note that the cut in the class prediction for the presented confusion matrix is made at p X-ray = 0.50, where p X-ray is the predicted probability that a counterpart candidate is X-ray emitting. Since NWAY uses the continuous predicted probability as modifying factor for the likelihood P(D | H), real counterparts with rare or untypical photometric features, that is, with p X-ray 0.50, may still be selected by the algorithm if the astrometric configuration favours them. We obtained a high recall fraction of 2585/(2585 + 457) = 85%, while the fractional leakage of contaminating field objects remained low: 738/(738 + 58041) = 1%. We note that p X-ray was only computed from the photometric and proper motion properties of the LS8 sources. In particular, it does not depend on coordinates and positional uncertainties. As discussed in the previous section, this allows us to split the likelihood of a match into independent astrometric and photometric terms:</p><p>the photometric term, is directly related to p X-ray .</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.1.2.">NWAY association run</head><p>Using the trained model, we predicted p X-ray for all LS8 sources in the eFEDS field. We then ran the NWAY matching procedure Fig. <ref type="figure">3</ref>. Confusion matrix resulting from the random forest prediction on an independent test set. X-ray sources are labelled "real X-ray", and field objects are labelled "field". Numbers on the right downward diagonal correspond to correctly predicted classes. At this step, the separation between the sources and the X-ray position is not considered. using the ratio p X-ray /(1 -p X-ray ) for P(D m | H). This was done by adding p X-ray as a column to the LS8 catalogue and activating it as a prior column in NWAY with the mag option. We set a radius of 30 arcsec from each eFEDS X-ray source, considering all LS8 sources within this radius. This may appear to be a relatively large maximal separation, given the mean eFEDS positional error of 4.5 arcsec; however, we wish to account also for the largest positional uncertainties of a few objects in the eFEDS source catalogue (see <ref type="bibr">Brunner et al. 2022</ref>) and the use of a large search radius minimises the probability of missing counterparts that are widely separated from the X-ray centroid position. The sky coverage of eFEDS and LS8 is 140 deg 2 and N eFEDS &#215; &#960; &#215; (30 ) 2 -A overlap , respectively, where N eFEDS is the number of eFEDS sources (point-like or extended) and A overlap is the overlap area of neighbouring search windows around the Xray sources. As described in the appendix of Salvato et al. ( <ref type="formula">2019</ref>), the area coverages were used to compute the number densities, which in turn were used to compute the probability for an eFEDS source to have a counterpart (p_any) and the probability for each source in LS8 to be the correct counterpart (p_i). These two quantities were then used to assign a counterpart to an eFEDS source: while the LS8 source with the highest p_i was considered to be the best available counterpart, we used the p_any value to decide whether the identification of the counterpart was reliable (see Sect. 4.2 for details).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.2.">MLR approach</head><p>The maximum likelihood ratio (MLR) statistic for the correct pairing of sources from multiple catalogues was introduced in the seminal work of <ref type="bibr">Sutherland &amp; Saunders (1992)</ref> and is widely used, although mostly to pair only two catalogs. For sources detected at two different wavebands and separated by angular distance r on the plane of the sky, the likelihood ratio provides a measure of the probability that the two sources are true counterparts normalised by the probability that they are random alignments. Quantitatively, this is estimated as</p><p>where q( -&#8594; m) is the prior knowledge about the properties of the true associations, such as the distribution of their apparent A3, page 6 of 32 magnitudes at given spectral window, their colours, and/or the spatial extent of the observed light in a given waveband. The collection of all possible source properties for which a prior probability can be estimated is represented by the vector -&#8594; m.</p><p>The quantity n( -&#8594; m) is the sky density of all known source populations in the parameter space of -&#8594; m. It measures the expected contamination rate from background/foreground sources that are randomly projected on the sky within a distance r off a given position. The probability that the true associations are separated by the distance r is measured by the quantity f (r). This depends on the positional uncertainties of the matched catalogues.</p><p>For the MLR applied to the eFEDS X-ray sources, a multidimensional prior was used that combines knowledge of the optical and mid-infrared colours/magnitudes of X-ray sources as well as their optical extent, that is, point-like versus extended.</p><p>The version of MLR we applied to the eFEDS work is based on the ASTROMATCH<ref type="foot">foot_6</ref> implementation. This tool has been specifically designed to deal with the complexity of widearea surveys that contain a very large number of sources. The HEALPix multi-order coverage map (MOC<ref type="foot">foot_7</ref> ) technology was used to describe the footprint of a catalogue of astrophysical sources. The KD-tree library as implemented in the ASTROPY package (Astropy Collaboration 2013, 2018) was used to accelerate spatial searches of potential counterparts within a radius r of a given sky position. The core ASTROMATCH functionality was expanded to enable the use of multi-dimensional priors. The version of ASTROMATCH adopted in this work is therefore a fork (github.com/ageorgakakis/astromatch) of the main development branch.</p><p>Like for NWAY, the optical counterparts were investigated out to a maximum radius of 30 arcsec. We assumed that the positional uncertainties of the X-ray and optical catalogues follow a normal distribution. The quantity f (r) is therefore represented by a Gaussian with &#963; parameter estimated as the sum in quadrature of the X-ray and optical positional uncertainties.</p><p>The priors were generated using the training sample defined in Appendix A.1. The LS8 photometric properties of the sources in that sample were explored to identify parameter spaces in which they are separate from the general LS8 field population. After some experimentation, we opted for the following three independent priors:</p><p>(1) A space that includes the WISE colour W1 -W2, the WISE magnitude W2, and the optical extent of a source. For the latter, we used the LS8 parameter TYPE, which provides information about the optical morphology of sources. In our application we only differentiated between optically unresolved (TYPE = "PSF") and optically extended (TYPE "PSF") populations;</p><p>(2) A space that includes the optical/WISE colour r -W2, the optical magnitude g, and the optical extent of a source. For the latter, we used the Legacy-DR8 parameter TYPE, as explained above;</p><p>(3) The distribution of the Gaia G magnitudes listed in the LS8 catalogues. This is to identify X-ray sources associated with very bright counterparts;</p><p>The distribution of the training sample sources in the parameter spaces above was used to define two three-dimensional priors and one one-dimensional independent prior. These were provided as input to the ASTROMATCH code, together with the distribution of the sources in the field population when the association for eFEDS was computed. For a given eFEDS source, all the potential associations within the search radius of 30 arcsec were identified. Each of them was assigned one LR value for each of the three priors using Eq. (1).The LS8 source with the highest value of LR from one of the three priors was considered to be the counterpart.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.">Comparing NWAY and ASTROMATCH on a validation sample</head><p>In order to compare completeness and purity of NWAY and ASTROMATCH, the same setting as we adopted to identify the counterparts to eFEDS was used to determine the best counterparts to a blind validation sample of 3415 counterparts to Chandra sources (see Appendix A). This validation sample was used as a truth table to test the performance of NWAY and ASTROMATCH in finding counterparts and to define the p_any and LR_BEST thresholds above which a counterpart is considered secure.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.1.">eROSITA-like validation sample</head><p>The Chandra sources were assigned eROSITA positional errors by randomly sampling from the astrometric uncertainties listed in the core eFEDS source catalogue. We accounted for the flux dependence of these uncertainties by matching any given Chandra source with a certain flux from 0.5-2 keV to only those eFEDS sources with a similar 0.6-2.3 keV flux within a margin of 0.5 dex. The flux transformation between the Chandra and eFEDS spectral bands is small, about 2% for a powerlaw spectral energy distribution with &#915; = 1.9, for instance, and was ignored. The positional uncertainty, &#963;, assigned to each of the Chandra sources, can be split into a right-ascension and a declination component. It was assumed that these two uncertainties are equal, and therefore, &#948;RA = &#948;Dec = &#963;/ &#8730; 2. Under the assumption that both the &#948;RA and &#948;Dec are normally distributed, the total radial positional uncertainty follows the Rayleigh distribution with a scale parameter &#963;.</p><p>Instead of directly using the assigned &#963; as the astrometric error to be applied to the Chandra positions to make them resemble the eFEDS astrometric accuracy, we preferred to add further randomness to the experiment. For each Chandra source, the assigned &#963; was treated as the scale factor of the Rayleigh distribution and a deviate was drawn that represented the positional error. This was applied to the sky coordinates of the optical counterpart of the Chandra source, and the new offset position was taken as the centroid of the X-ray source in the case of an eFEDS-like observation.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.2.">Probability thresholds definition</head><p>The LS8 counterparts to the Chandra eFEDS-like sources were identified by NWAY and ASTROMATCH using the same setup as we adopted for the real eFEDS observation. The resulting catalogue of best counterparts was matched with the true associations, providing a direct comparison between the methods, and, at the same time, providing a measure of the false-positive identification rate of the eFEDS counterpart catalogue.</p><p>First, we compared the primary identifications returned by NWAY and ASTROMATCH to true identifications stored in the validation sample. NWAY and ASTROMATCH correctly identified 3216 out of 3394 (95%) and 3024 out of 3394 (89%) sources, respectively. NWAY has a higher success rate. Additionally, NWAY has a smaller fraction of sources with a second  possible counterpart (115 sources against 367). Another way to look at the results is to compare purity and completeness for the two methods. At any given value of p_any/LR_BEST, we defined as purity the fraction of sources with a correct identification. In addition, we defined as completeness the fraction of sources for which we were able to assign a counterpart (see Fig. <ref type="figure">4</ref>). Both methods have very high purity and completeness. NWAY provides a sample that is purer, consistent with the fact that very few sources have a second possible counterpart, in addition to the correct one. Combined with the success rate, this makes NWAY the more robust method for determining the counterparts. Its strength comes first of all from the capability to account for complicated priors involving multiple features (essentially resembling an entire SED, together with other physical properties), from different catalogs at the same time. Furthermore, the Bayesian statistics upon which NWAY is based also allows accounting for sources that are lacking one or more of the features.</p><p>Similarly to what is traditionally done in maximum likelihood (see e.g. <ref type="bibr">Brusa et al. 2007</ref>), the intersection between the completeness and purity can be used to define a threshold above/below which the counterparts is considered reliable. This corresponds to 0.035 for p_any and 0.45 for LR_BEST.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.">Determination of the counterparts to eFEDS sources</head><p>While for the large majority of the cases the two methods select the same counterparts, there are cases where they do not agree Notes. In the last column, the fraction of the "different ctps" with respect to the whole sample is also reported. or where they identify multiple likely associations. In the following we describe the procedure that we adopted for the final assignment of the counterparts. Then, after the consolidation of the counterpart, we describe a further test for consistency that was done by comparing the results of the association with an independent method, Ham-Star <ref type="bibr">(Schneider et al. 2022)</ref>, which is tuned to identify Galactic coronal X-ray emitters (Sect. 5.3). The same process was then repeated for the 246 sources in the eFEDS hard point-source catalog. From now on, all numbers and descriptions are given for the main sample unless specified otherwise.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.1.">Comparison of counterparts from NWAY and ASTROMATCH</head><p>For 24 193 out of 27 369 (88.4%) eFEDS point-like sources in the main sample, NWAY and ASTROMATCH point at the same counterpart. They disagree for 3176 (11.6%) of the cases. The numbers are quoted at this stage regardless of the p_any or LR_BEST thresholds, which are used instead below in order to assign a flag for the quality of the proposed counterpart.</p><p>Table <ref type="table">4</ref> summarises the number of eFEDS sources with the agreement/disagreement between the two methods as a function of detection likelihood of the X-ray source. Sources with low detection likelihood values have larger X-ray positional errors on average, and a larger number of spurious sources is expected from simulations <ref type="bibr">(Brunner et al. 2022;</ref><ref type="bibr">Liu et al. 2022c</ref>). It is therefore not surprising that the largest discrepancies are observed at the lowest detection likelihoods (Fig. <ref type="figure">5</ref>). The disagreement drops from 11.6 to 6.5% when we consider only eFEDS sources with DET_LIKE grater than 10, suggesting that at low detection likelihood, a fraction of eFEDS sources might be spurious detections where NWAY and ASTROMATCH assign a different field source. The notion that these are field sources is also supported by the fact that for about 50% of eFEDS sources with DET_LIKE below 10 and with different counterparts, both p_any and LR_BEST are below the threshold.</p><p>A3, page 8 of 32 Table <ref type="table">5</ref> summarises the comparison between the two methods and also takes the reliability of the associations into account. In this table we further split the sample with the same counterparts ("same ctps" for brevity) into two subsamples: one for which the proposed counterparts are the only associations suggested by both methods ("single solutions"; 86.3% of the entire sample), and one for which, although both methods point to the same associations, at least an additional counterpart at lower significance exists from at least one method ("multiple solutions"; 2.1% of the entire sample).</p><p>The different priors and the different methods we used to assign the counterparts explain the selection of different counterparts in the different ctps sample. ASTROMATCH uses three priors, but they are each used independently, and for any given eFEDS source, the counterpart is assigned by the prior with the higher probability. Instead, NWAY uses all the features at the same time, and the best counterpart is the one that mimics best the training sample in a multidimensional space. We consider this second method more reliable, and for this reason, we decided to always list as primary the counterpart suggested by NWAY, unless LR_BEST is above the threshold and p_any is not.</p><p>Interestingly, we note that in the different ctps sample, the primary counterpart assigned by one method is the secondary counterpart assigned by the other for about 25% of the cases.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.2.">Assigning a quality to the proposed counterparts</head><p>As a consequence of the discussion above, each counterpart in the catalogue was flagged as follows ([number] refers to the number of sources in the category):</p><p>-CTP_quality = 4: when NWAY and ASTROMATCH agree on the counterpart, and both p_any and LR_BEST are above threshold [20 873 sources; black in <ref type="bibr">Table 5]</ref>;</p><p>-CTP_quality = 3: when NWAY and ASTROMATCH agree on the counterpart, but only one of the methods assigns the counterpart with a probability above the threshold [1379 sources; blue in <ref type="bibr">Table 5]</ref>;</p><p>-CTP_quality = 2: when there is more than one possible reliable counterpart. This includes a) all the sources in the different ctps sample with at least one probability above the threshold, and b) the sources in the same ctps sample with possible secondary solutions [2522 sources in total; cyan in <ref type="bibr">Table 5]</ref>. Because of the low spatial resolution of eROSITA, this last case implies that both sources contribute to the X-ray flux. A supplementary catalogue with the properties of the secondary counterparts for these 2522 sources is also released (Sect. 8).</p><p>-CTP_quality = 1: when NWAY and ASTROMATCH agree on the counterpart, but both p_any and LR_BEST are below the threshold [1370 sources; purple in <ref type="bibr">Table 5]</ref>; a probability below the threshold does not necessarily imply an incorrect counterpart. It might also indicate that the counterpart is correct, but its features do not sufficiently mimic those in the training sample.</p><p>-CTP_quality = 0: when NWAY and ASTROMATCH indicate different counterparts and both p_any and LR_BEST are below the threshold [1225 sources; red in <ref type="bibr">Table 5]</ref>.</p><p>Counterparts with quality 4, 3, and 2 are considered reliable (90.5% of the main sample and 93.9% of the hard sample), while sources with quality 1 or 0 are considered unreliable (9.5% of the main sample and 6.1% of the hard sample).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.3.">Comparison with an independent association method tuned to stars: HamStar</head><p>The content of the eFEDS point-source catalogue was also analysed in order to specifically identify stellar coronal X-ray emitters with sufficiently well-defined properties. This method, called HamStar in the following, is based on the properties expected for this type of star; the details are presented in <ref type="bibr">Schneider et al. (2022)</ref>. In short, HamStar performs a binary classification between stellar coronal emitters and other objects. This classification is based on the concept of eligible stellar counterparts, that is, the match catalogue contained only stellar objects that may reasonably be responsible for the X-ray sources. Specifically, the parent sample that HamStar used included only sources from Gaia EDR3 that (a) are brighter than 19th magnitude in G band (implied by the stellar saturation limit of L X /L bol 10 -3 and the depth of eFEDS); (b) have accurate magnitudes in all three Gaia photometric bands (to apply colourdependent corrections); (c) have a parallax value at least three times higher than the parallax error (to select only genuine stars).</p><p>Then, a positional match between sources in eFEDS and the eligible stellar candidates was made, considering all sources within 5&#963; of the positional uncertainty of the eFEDS source as possible stellar counterparts. Finally, the matching probabilities of all possible counterparts were adjusted based on the value of the two-dimensional Bayes map at the counterpart Bp-Rp colour and ratio of X-ray to G-band flux. Based on the Ham-Star algorithm, 2060 eFEDS sources are expected to be stellar <ref type="bibr">(Schneider et al. 2022</ref>). The vast majority of them have a unique Gaia counterpart, and only 83 eFEDS sources have two possible counterparts.</p><p>Of the 2060 eFEDS sources with a counterpart from Ham-Star, a counterpart for 1883 is identified here that is less than 2 arcsec from the counterpart proposed by Hamstar. We assume A3, page 9 of 32 A&amp;A 661, <ref type="bibr">A3 (2022)</ref> that this is the same source. We visually inspected the cutouts of the 29 sources for which the separation between the counterpart proposed by Hamstar and this work is between 2 and 3 arcsec, and concluded that the counterparts are the same for 9 sources, but the sources are heavily saturated in LS8 so that the coordinates are not sufficiently precise. This corresponds to an 92% agreement; incidentally, this value corresponds almost exactly to the expected reliability and completeness of HamStar <ref type="bibr">(Schneider et al. 2022)</ref>. All these sources are then classified as "secure Galactic" in Sect. <ref type="bibr">6</ref>.</p><p>HamStar applies well-understood X-ray-to-optical properties of stars to a well-defined subsample of Gaia sources. On the other hand, the training samples used by NWAY and ASTRO-MATCH include various classes of X-ray emitters: stars and compact objects, AGN, and galaxies, including the bright ones at the centre of clusters (BCG). We considered the prior defined by NWAY and ASTROMATCH to be more representative of the population of X-ray emitters in general and decided to keep the counterpart assigned in the previous section rather than changing counterparts for the 177 sources for which the methods indicate different counterparts. However, we degraded the CTP_quality because an alternative solution might apply. Interestingly, only 4 out of 211 sources were considered secure, with CTP_quality = = 3. All other counterparts had an already low CTP_quality.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.4.">Separation and magnitude distribution of the counterparts</head><p>For 24 427 out of 24 774 (98.5%) of the sources with CTP_quality&#8805;2, the separation between the X-ray position and the assigned LS8 counterpart is smaller than 15 arcsec, with a mean of 4.3 arcsec. As might be expected, there is a trend for larger average X-ray-optical separations at lower values of DET_LIKE; lower detection likelihood sources typically have larger X-ray positional uncertainty (see <ref type="bibr">Brunner et al. 2022</ref>). The distribution of the observed X-OIR separations normalised by the X-ray positional uncertainty is shown in Fig. <ref type="figure">6</ref> as a function of the r magnitude of the counterpart. The distribution is broadly comparable to the expectation of a Rayleigh distribution with a scale factor = 1. In Fig. <ref type="figure">7</ref> we show the distribution of the sample in X-ray flux versus optical magnitude space. The sample is subdivided into objects with more secure counterparts (CTP_quality &#8805; 2) and objects with less reliable counterparts (CTP_quality &#8804; 1). The less reliable counterparts tend to have fainter optical magnitudes for a given X-ray flux than the more secure counterparts.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="6.">Source characterisation and classification</head><p>After the identification of the counterparts, the different classes of objects need to be separated to understand physical processes and populations. The most important separation is between extragalactic sources (galaxies, AGN, and QSOs) and galactic sources (stars, compact objects, etc.). In the following, we describe how we classifed the sources and how the validation tests were performed.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="6.1.">Galactic and extragalactic sources</head><p>In order to classify sources in the most reliable way, we used a combination of methods and various information: spectroscopy, parallax measurements from Gaia, colours, and morphology from imaging surveys. None of the methods is infallible because they all depend on the quality of the data (e.g. S/N for spectra, depth, and image resolution) and because of the degeneracy in colour-redshift space for many of the sources. We therefore adopted a multi-step approach: at each step, we extracted from the pool of sources those that can be classified with high reliability as either extragalactic or Galactic. Figure <ref type="figure">8</ref> shows an illustration of the decision tree we adopted for the classification, together with the number of sources in each the classes. The procedure is described below in detail.</p><p>We first applied the classification based on spectroscopy or high parallax. These can be considered primary methods as they are highly pure, but certainly not complete. The sources thus classified were defined as "secure Galactic" or "secure extragalactic". Briefly, we defined as secure extragalactic all sources with spectroscopic redshift &gt;0.002 and NORMQ = 3 (step 1 in Fig. <ref type="figure">8</ref>) and as secure Galactic all sources satisfying at least one of the criteria spectroscopic redshift &lt;0.002 and NORMQ = 3, significant parallax from Gaia EDR3 (above 3&#963;), or agreement with HamStar counterparts (step 2).</p><p>Next, we extracted those of the sources that were still in the pool that appeared extended in the optical images. Depending on whether photometry from HSC was available, a source was defined as extended (EXT) if it satisfied &#8710;mag = mag Kron -mag psf &gt; 0.1 <ref type="bibr">(2)</ref> simultaneously in g,r,i,z from HSC imaging data (e.g. Palanque-Delabrouille et al. 2011), or, when no photometry from HSC was available (either because the source is outside the field or A3, page 10 of 32 because of saturated photometry),</p><p>The EXT sources were then flagged as "likely extragalactic" (step 3). This is considered a secondary classifier because in poor seeing conditions, for example, point-like sources (or stellar binary systems) would also be misclassified as extended (see the discussion presented in <ref type="bibr">Hsu et al. 2014</ref>).</p><p>The sources classified as "secure" were then projected in the LS8 z-W1 versus g-r plane (see inset in the left panel of Fig. <ref type="figure">9</ref>), following <ref type="bibr">Ruiz et al. (2018)</ref>. There, we empirically defined a line separator, described as</p><p>which provides a sharp separation between secure Galactic and extragalactic sources; a negligible fraction of secure extragalactic sources lies below the separator (left panel of Fig. <ref type="figure">9</ref>). Then, for all the sources still in the pool and with available photometry from LS8 (step 4), we classified the sources below the line as "likely Galactic" (step 5). The remaining sources in the pool with available LS8 photometry were classified as "likely Galactic/extragalactic" (step 6) depending on whether they fell below or above the line in the W1 versus X-ray flux plane (see inset in the right panel of Fig. <ref type="figure">9</ref>), as defined in Salvato et al. (2019), W1 + 1.625 * log(F 0.5-2keV ) + 6.101 = 0, <ref type="bibr">(5)</ref> with W1 in the Vega system and X-ray flux in cgs. Originally, a similar line separator was introduced by <ref type="bibr">Maccacaro et</ref>   <ref type="figure">9</ref>). It has the advantage of generality, as the W1 photometry and the X-ray fluxes are available virtually for all the eFEDS sources. Finally, we assumed for the sources without complete information from LS8 that they are extragalactic (step 7), unless they are below the W1-X line defined in Eq. ( <ref type="formula">5</ref>) (step 8).</p><p>In this manner, a simple but reliable four-way classification scheme (secure/likely Galactic/extragalactic) was achieved. The final distribution of the four classes of sources in the g-r-z-W1 vs. W1-X planes is shown in Fig. <ref type="figure">10</ref>. The two line separators identify four wedges, two of which can be used to define almost 100% pure subsamples of Galactic/extragalactic X-ray selected sources. The four wedges are described below.</p><p>-Top left: 724 sources, out of which 637 (87.9%) are Galactic (463 and 428, respectively, only considering sources with reliable counterparts, CTP_quality&gt; = 2).</p><p>-Top right: 23 874 sources, out of which 23809 (99.7%) are extragalactic (21 711 and 21647 for CTP_quality&gt; = 2).</p><p>-Bottom left: 1391 sources, out of which 1373 (98.7%) are Galactic (1337 and 1319 for CTP_quality&gt; = 2).</p><p>-Bottom right: 1380 sources, out of which 479 (34.7%) are extragalactic ( 1263 and 379 for CTP_quality&gt; = 2).</p><p>It is important to recall that the order of the steps taken in the decision tree is crucial to limit the misclassification of the sources as much as possible. For example, the use of spectroscopic redshift in the first step allowed us to identify the bright and nearby extragalactic sources that would have been misclassified as Galactic in the W1-X plane. Similarly, the adoption of the high parallax from Gaia allowed us to identify secure galactic sources that would have been misclassified as extragalactic in the z-W1 versus g-r plane.</p><p>In summary, the eFEDS main sample comprises 24 393 sources classified as extragalactic (5377 secure and 19 016 likely) A3, page 11 of 32 A&amp;A 661, A3 (2022) Fig. <ref type="figure">8</ref>. Decision tree we adopted to assign each eFEDS point source to the Galactic or extragalactic class. First we classified the sources on the basis of the most secure methods (e.g. high-confidence redshift) and then proceeded with less reliable methods (e.g. based on colours) on the sources remaining in the pool, creating less pure samples. The numbers listed at each step include all sources, i.e., they also include those with an insecure counterpart (CTP_quality &lt; 2). and 2976 classified as Galactic (2566 secure and 410 likely). All these numbers are reported in Table <ref type="table">6</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="6.2.">Validation of the classification using external samples</head><p>We carried out sanity checks of the classification framework against two external catalogues whose members are expected to   <ref type="figure">9</ref>. Distribution of sources flagged as "secure Galactic" (red) and "secure extragalactic" (blue) in the g-r vs z-W1 (left) and W1 vs. X-ray (right) planes that were used to determine a line separator (black line) to classify sources in steps 5, 6, and 8 of the flowchart presented in Fig. <ref type="figure">8</ref>. The line separator on the right has fewer Galactic sources that fall into the extragalactic locus. However, with the line separator defined on the left, only a handful of extragalactic sources fall into the Galactic locus. This makes this classifier more efficient when the four photometric points are available. Fig. <ref type="figure">10</ref>. Four eFEDS X-ray source classes "secure Galactic" (red), "likely Galactic" (orange), "secure extragalactic" (blue), and "likely extragalactic" (cyan) defined in the flowchart presented in Fig. <ref type="figure">8</ref>, distributed according to their distance from the two lines defined in Fig. <ref type="figure">9</ref>. Three of the four wedges thus defined contain extragalactic or Galactic samples that are up to 99% pure (see text for details). positional matches were made against our best matching optical counterpart positions, with a search radius of 3 arcsec for the FIRST radio component catalogue <ref type="foot">9</ref> and 1 arscec for GUA (GUA objects were considered when they had PROB_RF &gt; 0.8). We examined the rate at which sources we classified as Galactic or extragalactic (both secure and likely) were matched to objects in these external catalogues (see Table <ref type="table">6</ref>). There is a very low Notes. The low rate at which our classification logic classifies both radio sources and AGN candidates as secure Galactic or likely Galactic suggests that our classifications are robust.</p><p>rate of apparent disparities between our classifications and those that may be derived by matches to the external catalogues. For example, only 0.19% of secure Galactic sources have a radio counterpart in FIRST, compared to 7.0% of the secure extraglactic sample. Likewise, only 0.19% of the secure Galactic sources are matched to candidate AGN from Shu et al. ( <ref type="formula">2019</ref>), compared to 54% of the secure extragalactic subsample.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="6.3.">Very nearby galaxies</head><p>Unlike what happens in pencil-beam surveys, there are numerous very nearby and thus resolved galaxies within eFEDS. the HECATE galaxies, but do not coincide with the centre of the galaxy, but rather with a source that could be either an ULX in the galaxy or an extragalactic source in the background. For these 7 sources, dedicated studies will be needed to identify the exact origin of the X-ray emission.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="7.">Photometric redshifts</head><p>Photo-z of AGN and X-ray selected sources in general have developed dramatically in the past ten years, bringing the redshift accuracy and the fraction of outliers (usual quantities measured to assess the quality of the photo-z) comparable to those measured for normal galaxies. Regardless of whether photo-z are computed via SED fitting or via machine-learning, accurate photo-z for AGN are less straightforward to obtain than those for non-active galaxies (see <ref type="bibr">Salvato et al. 2018</ref>, for a review of the topic), mainly because for each multi-wavelength data point, the relative contribution of host and nuclear emission is unknown and redshift dependent. Redshift, however, is the parameter that we are trying to determine. To add to the difficulty, the impact of dust extinction and variability must not be fogotten. Variability is an intrinsic property of AGN. Especially for wide-area surveys, where data are taken over many years, this can noticeably affect the accuracy of photo-z if it is not accounted for (e.g. <ref type="bibr">Simm et al. 2015)</ref>, as was possible to do in COSMOS, for instance <ref type="bibr">(Salvato et al. 2009</ref><ref type="bibr">(Salvato et al. , 2011;;</ref><ref type="bibr">Marchesi et al. 2016)</ref>. In eFEDS, we also have to face the issue that the photometry is not homogenised, and different surveys cover different parts of the field at different depth and there are different ways of computing the photometry (Kron, Petrosian, apertures, model, etc). In the following we describe the procedure we adopted to compute photo-z using LePHARE <ref type="bibr">(Arnouts et al. 1999;</ref><ref type="bibr">Ilbert et al. 2006</ref>). We then proceed with an estimate of the reliability of the photo-z and a comparison with DNNZ, an independent computation of photo-z using machine-learning (Nischizawa et al., in prep.).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="7.1.">Photo-z computation</head><p>We computed the photo-z for the sources classified as extragalactic. In order to minimise systematic effects, we used different types of photometry, depending on the survey; in particular, we tried to avoid photometry derived from models for the extended and nearby sources because usual models are good representation of point-like, disk-like, and bulge-like sources, but are unable to represent a local Seyfert galaxy, for example, in which nuclear and host components both contribute to the total flux. For this reason, we used total fluxes from GALEX; Kron and cmodel photometry from HSC, depending on whether the source was extended (see below); and GAAP (Gaussian Aperture and Photometry) from KiDS+VIKING. From VHS, we adopted Petrosian photometry as it appears to agree better with the VISTA/VIKING photometry. All the photometry was corrected for Galactic extinction using E(B-V) from LS8. Depending on whether the source was in the area covered by KiDS+VIKING, within HSC but outside KiDS and outside HSC, different bands were available 10 . 10 In particular, for HSC, in the S19A release available to us at the time of this work, photometry in r2 and i2 filters is provided. However, the filters have changed during the survey, and depending on the coordinates of the sources, the fraction of data obtained with the original or the new filters changes. In order to account for this at any location, we adopted the filter that was used to obtain at least 50% of the data. This solution is not optimal and will affect the quality of the photo-z in some areas.</p><p>The computation of the photo-z followed the procedure already outlined in previous works <ref type="bibr">(Salvato et al. 2009</ref><ref type="bibr">(Salvato et al. , 2011;;</ref><ref type="bibr">Fotopoulou et al. 2012;</ref><ref type="bibr">Hsu et al. 2014;</ref><ref type="bibr">Marchesi et al. 2016;</ref><ref type="bibr">Ananna et al. 2017)</ref>, where sources were treated differently, depending on whether the optical images indicate them being extended (EXT) or a point-like/unresolved (PLIKE), following Sect. 6. This step is particularly important, as sources in the two samples are treated differently, using different priors and templates.</p><p>In addition, the fitting templates were selected on the basis of the X-ray depth and coverage of the surveys, keeping in mind that bright AGN, for instance, will mostly be absent in a deep pencil-beam survey. These surveys are characterised instead by host-galaxy dominated sources. Given the similar X-ray depth, the libraries used in <ref type="bibr">Ananna et al. (2017)</ref> for the Stripe-82X survey were a good starting point for our work on eFEDS. However, a new library of templates for AGN and hybrids (AGN and host) was recently presented in <ref type="bibr">Brown et al. (2019)</ref>. The authors used photometry and archival spectroscopy of 41 AGN to create an additional set of 75 new hybrid templates. With respect to previous AGN templates, they have the advantage that they are empirical for the entire wavelength coverage and that the contribution from the host and AGN components is fully taken into account when the final SED is created, including dust attenuation and emission lines.</p><p>eFEDS is particularly rich in sources with reliable spectroscopy (see Sect. 2.3), allowing for a better tuning of the templates to be used to compute the photo-z. To optimise the template choice, the colours of all the sources with reliable spectroscopy were plotted as a function of redshift, together with the theoretical colours from all the templates available (Fig. <ref type="figure">11</ref> illustrates this for i-z and W1-W2 for the EXT and PLIKE samples, respectively, for all the templates that we ultimately adopted). In selecting the templates, we tried to limit their number (to control degeneracy in the redshift solution), while at the same time compiling a list representative of the entire population. The sources that in Fig. <ref type="figure">11</ref> are outside the parameter space covered by the templates can be interpreted in various ways, from problems in the photometry of the specific objects due to blending with neighbours or variability or lack of certain features in the templates. While we will further investigate this latter possibility for the future eROSITA surveys, here we recall that the figures are representative of only two colours, while in selecting the templates, we study all the colours that our photometric set allows.</p><p>An important point to keep in mind is the fact that despite being rich, the available spectroscopic sample is not representative of the entire eFEDS population, as is shown in Fig. <ref type="figure">12</ref>. For this reason, the final library should also include some templates for types of sources that are expected to be present in eFEDS, but have not necessarily been identified so far. In particular, we created a set of templates using the archetype of type 1 AGN from the counterparts of ROSAT/2RXS <ref type="bibr">(Salvato et al. 2019</ref>) observed within SDSS-IV/SPIDERS presented in <ref type="bibr">Comparat et al. (2020)</ref>, extended in UV and mid-infrared with various slopes. For the non-empirical templates, reddening was also considered, using the extinction law of Prevot <ref type="bibr">(Prevot et al. 1984</ref>) with E(B-V) values from 0 to 0.4 in steps of 0.1. The selected templates are presented in Appendix B.</p><p>As output, LE PHARE provides the best value for the best photo-z together with the upper and lower 1, 2, and 3&#963; error, the best combination of template, extinction law, and extinction value, the quality of the fit, and the pdz, the latter being the redshift probability distribution defined as pdz = F(z) dz A3, page 14 of 32   <ref type="formula">2019</ref>), this is not the case when the photometric set is not rich, and pdz can be high also for a poor fit just because there are no sufficient constraints.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="7.2.">Reliability of photo-z</head><p>The final comparison between photo-z and spec-z, considering EXT and PLIKE sources together, for the area within KiDS+VIKING and within HSC, but outside KiDS+VIKING, is shown in Fig. <ref type="figure">13</ref>. We used the standard metrics to measure the quality of photo-z (see <ref type="bibr">Salvato et</ref>   Notes. In each row, the difference between the numerators in the N spec /N tot columns provides the number of sources with spectroscopy for which DNNZ could not provide a photo-z, mostly because the HSC photometry is saturated for these sources. The results are listed in Table <ref type="table">7</ref>. Figure <ref type="figure">14</ref> shows the same results, but split as a function of z-band magnitude from LS8, X-ray flux, and spectroscopic redshift.</p><p>Ideally, for the best computation of photo-z, in particular, for sources dominated by emission lines such as AGN, photometry from broad-band filters across the entire spectral range should be complemented by narrow-band and near-infrared photometry and should be homogenised (e.g. <ref type="bibr">Salvato et al. 2009</ref><ref type="bibr">Salvato et al. , 2018))</ref>. While narrow-band photometry is not available, at least some of the surveys provide homogenised photometry. For niear-infrared photometry, the VISTA/VHS data are not sufficiently deep. The effect on the photo-z is clearly visible in all the panels of Fig. <ref type="figure">14</ref>, where the fraction of outliers is usually higher and the accuracy lower (high value of &#963; NMAD in the area without VIKING coverage; dotted lines). Not only are the near-infrared data shallow outside the KiDS+VIKING area, they are also just a collection of photometric points computed in different ways, simply matched in coordinates. For this reason, based on the footprints shown in Fig. <ref type="figure">1</ref>, we can think of the photo-z in eFEDS as divided into three regions that reflect the quality of the available photometry: the inner area is covered by deep forced photometry in KiDS+VIKING; the area that is within HSC, but outside KiDS+VIKING, for which some near-infrared information is provided by the shallow VISTA/VHS; and the area outside HSC for which the optical photometry is provided by LS8 alone.</p><p>The lack of deep near-infrared data also creates an unusual number of sources at high-z (z &gt; 3), most of which are most likely incorrect. For example, the number of sources with A3, page 16 of 32 photo-z &gt; 3 is 188 within KiDS and 819 in the HSC area outside KiDS, although the area is about the same size. Within KiDS+VIKING, LE PHARE correctly estimates the redshift for 40 of the 55 (72.7%) sources that are spectroscopically confirmed to be at a redshift higher than 3. Most of these high-z sources in excess can be easily identified and flagged by noticing that they are characterised by having high pdz even though they are in the area outside the HSC, that is, with a very limited number of photometric points to be fitted (see Sect. 7.4).</p><p>Figure <ref type="figure">14</ref> also shows how the accuracy degrades and the fraction of outliers increases for the PLIKE that are X-ray bright (top, second panels from the left). These sources are dominated by the AGN component with an SED close to a power law, for which the lack of narrow-band photometry that would identify the emission lines does not allow breaking the degeneracy in the redshift solutions. However, in eFEDS, there are only 47 extragalactic sources with an X-ray flux above 5 &#215; 10 -13 erg cm -2 s -1 , and a reliable spectroscopic redshift is available for 39 of them, so that the low quality of the photo-z for these sources has only a limited effect. Finally, Fig. <ref type="figure">14</ref> shows an undesired high fraction of outliers at low redshift, where photo-z values for normal galaxies are usually extremely accurate. The problem for AGN probably originates from the fact that both KiDS and HSC photometry are based on fitted models and not on total fluxes. Models are not able to account properly for the contribution of the nuclear component that is comparable to the one from the host. Photometry from models can represent sources at very low redshift well, where the AGN contribution is negligible with respect to that from the host, and at high redshift, where the flux is dominated by the AGN component.</p><p>As already highlighted in the past, it is always easier to obtain a reliable photo-z for galaxy-dominated sources with characteristic breaks in the SED. AGN-dominated sources are degenerate in the redshift solution, especially when little photometry is available, even within the KiDS+VIKING area (compare the dashed lines for EXT and PLIKE).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="7.3.">Comparison with DNNz</head><p>Within the HSC collaboration, the computation of photo-z is available in many flavours. The method that performs better on AGN is DNNZ (Nishizawa et al., in prep.). It is based on machine-learning and exclusively uses HSC photometry, trained on the rich spectroscopic sample available for both AGN and normal galaxies within the entire HSC region (beyond the area in common with eFEDS). The DNNZ is based on the multi-layer perceptron (MLP) that takes the cmodel flux, PSF-matched aperture flux, and the second-order moment size measured at five HSC filter bands as inputs, and takes posterior probability as an output. In total, 3 &#215; 5 inputs and output PDF were binned in A3, page 17 of 32 A&amp;A 661, A3 (2022) Fig. <ref type="figure">15</ref>. Direct comparison between photo-z computed in this work with LE PHARE and DNNZ, within the HSC area for all the EXT (left panel) and PLIKE (central panel) sources. By construction, true EXT sources should not have spectroscopic redshift exceeding z &#8776; 1. It is not possible to decide a priori whether the photo-z are incorrect or if the sources were placed erroneously in the EXT sample due to some issue of the photometry. In the middle panel, we highlight the sources that have spectroscopic redshift higher than 3 in red (see text for details). Right panel: Comparison between photo-z from LE PHARE and spec-z for the sources for which LE PHARE and DNNZ agree. Notes. The table clearly indicates that when LE PHARE and DNNZ disagree, DNNZ has a higher fraction of outliers among the spectroscopic sample, while when the two codes agree, the difference in fraction of outliers is marginal. To define agreement, we used 1+mean (LEPHARE,DNNZ). The small difference in the fraction of outliers for the two methods when they agree depends on how close they are to the real spectroscopic value. 100 bins from z = 0 to z = 7. We have five hidden layers, and each layer has 100 nodes that are fully connected to the nodes in the neighbouring layers. With a 50k spectroscopic sample, it takes almost a whole day to train this machine with NVIDIA GeForce RTX 2080Ti GPU.</p><p>One interesting feature of DNNZ is that it was trained for any type of extragalactic source, without any particular tuning for AGN. In Table <ref type="table">7</ref>, the performances of DNNZ are directly compared with the output from LE PHARE. Remarkably, the accuracy of DNNZ is higher in general than for LE PHARE, although with a higher fraction of outliers.</p><p>Interestingly, although only HSC photometry was used, DNNZ also shows a remarkable difference in the quality of the photo-z for the sources within or outside the area covered by KiDS+VIKING. This is probably due to the combined photometry from the filters r and r2 and i and i2 that were changed during the survey. Most of the KiDS+VIKING area has been homogeneously observed only in i and r band, while the rest of the area has a mixture of observations. Taking this into account, we can compare LE PHARE and DNNZ in the area within KiDS+VIKING and split by TYPE. Figure <ref type="figure">15</ref> shows that both sets of photo-z have some systematics (vertical and horizontal substructures) that are due on the one hand to the imbalance between galaxies and AGN in the training of DNNZ, and on the other hand, to the degeneracies in the solution for power-lawdominated AGN and limited availability in photometry for LE PHARE.</p><p>However, when the photometry is sufficient and of good quality, SED fitting can correctly predict the redshift of AGN also when it is higher than 3 (middle panel of Fig. <ref type="figure">15</ref>; sources in red). This is a current limitation for photo-z computed via machine-learning because the sample of this type of source available for training is small (see Nishizawa et al., in prep.).</p><p>When the photo-z from DNNZ is available, we can measure the mean photo-z between the values proposed by the two methods for each source. Assuming this value is the right one, we have that for 60.6% of the extragalactic sources with CTP_quality &#8805; 2 DNNZ and LE PHARE agree (|zp LePHAREzp DNNz | &lt; 0.15 &#215; (1 + mean(zp LePHARE , zp DNNz )). The comparison between LE PHARE and the spectroscopic redshift for 3919 sources with spec-z is shown in the third panel of Fig. <ref type="figure">15</ref>; the fraction of outliers with respect to the spectroscopic sample is extremely small and the accuracy is very high, comparable to the accuracy that is routinely obtained for normal galaxies, using purely broad-band photometry.</p><p>Table <ref type="table">8</ref> summarises the result for DNNZ and LE PHARE separately, within and outside KiDS+VIKING. For the about 7500 sources for which the two methods provide differing A3, page 18 of 32 results, the spectroscopic sample does not help to distinguish the best photo-z because the spectroscopic sample is very small (756 and 553 sources in the two areas, respectively) and not representative of the magnitude distribution in the sample (mean r value of the spectroscopic sample 20; the mean r value of the sample for which DNNZ and LE PHARE disagree is 21.5. See also the next section and Fig. <ref type="figure">16</ref>). Photo-z derived via machine-learning Notes. The lower the pdz, the lower the quality of the fitting.</p><p>are well known to be very reliable only within the parameter space represented by the training sample and have little predictive power outside this space (e.g. <ref type="bibr">Brescia et al. 2019)</ref>. Keeping this in mind, we decided to rely on the prediction power of SED fitting and to rely on the results from LE PHARE. However, we also report the results from DNNZ and flag the sources for which LE PHARE and DNNZ agree or disagree (see Sect. 7.4).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="7.4.">CTP_REDSHIFT and CTP_REDSHIFT_GRADE in the final catalogue</head><p>In the final catalogue we report the spectroscopic redshifts (regardless of their reliability) and the photo-z from both LE PHARE and DNNZ. In addition, for each source we summarise in the two columns CTP_REDSHIFT and CTP_REDSHIFT_GRADE our best knowledge of redshift and its reliability. The column CTP_REDSHIFT lists original spectroscopic redshift when it is available and reliable (NORMQ = 3). The redshift is set to 0 for all the sources that are classified as GALACTIC (either SECURE or LIKELY) or for which no reliable redshift is available. To the remaining sources we assign the photo-z from LE PHARE.</p><p>Then in the column CTP_REDSHIFT_GRADE we provide a grade of confidence to the redshifts. The grades are listed below.</p><p>-CTP_REDSHIFT_GRADE = 5: this is a higher grade, assigned to the sources with reliable spectroscopic redshift. Of the 6591 sources in this category, 5377 are extragalactic sources and 1214 are Galactic (6465/6591 with CTP_quality &#8805; 2).</p><p>-CTP_REDSHIFT_GRADE = 4: this is assigned to the sources for which the photo-z from LE PHARE and DNNZ agree (10 949 in total, 9643 of which have CTP_quality &#8805; 2), because in the previous section we demonstrated that for this subsample, the fraction of outliers is very small and the accuracy very high. By construction, all the Galactic sources without spectroscopic redshift have CTP_REDSHIFT_GRADE = 4 because DNNZ and LE PHARE are set to zero and belong to this subsample (2995 sources).</p><p>-CTP_REDSHIFT_GRADE = 3: this is assigned to the sources for which LE PHARE and DNNZ disagree and pdz &gt; 40 (6741 in total, 6057 of the sources with CTP_quality &#8805; 2). The threshold at pdz &gt; 40 was set by considering the fraction of outliers as a function of pdz in the sample with spectroscopic redshift (see Table <ref type="table">9</ref>). At the same time, we searched for the value of pdz that minimised the number of outliers and maximised the number of sources with z_phot &gt; 4. The latter is suspiciously too high. This is due to the lack of deep photometry not only in the UV, but also in near-infrared: the large majority of these high-z sources are concentrated in the area outside KiDS. -CTP_REDSHIFT_GRADE = 2: this is assigned to the remaining sources for which LE PHARE and DNNZ disagree and pdz &lt; 40, for which we are less confident about the photometric redshifts, This group includes only 1326 sources, 1092 of which have a CTP_quality &#8805; 2.</p><p>Figure <ref type="figure">16</ref> shows the distribution of the sources for each of the REDSHIFT_GRADE in the magnitude redshift plane. The mean value of the redshift and magnitude for each of the subsamples is also indicated. 7.5. Flagging sources likely associated with clusters of galaxies in the point-like sample By construction, the eFEDS X-ray point-source catalogue is expected to be very little contaminated by clusters of galaxies; still, a low probability that a source is actually a cluster remains, as was shown in the simulations we performed for eFEDS <ref type="bibr">(Liu et al. 2022c</ref>). Clusters end up in the point-like sample for many reasons <ref type="bibr">(Willis et al. 2021)</ref>. Most obviously, clusters with a small apparent size or at low detection likelihood can fall below the thresholds that are used to define the extension of the X-ray source. In addition, clusters could leak into the point-like sample because of source splitting and superimposition of a bright point source and a cluster. With this in mind, we ran the multi-component matched filter cluster confirmation tool (MCMF; <ref type="bibr">Klein et al. 2018</ref><ref type="bibr">Klein et al. , 2019) )</ref> on the eFEDS point source catalog. We ran MCMF as in the eFEDS extended sources catalogue <ref type="bibr">(Klein et al. 2022</ref>) after adjusting some of the parameters (e.g. limiting the area search from the X-ray position). As in <ref type="bibr">Klein et al. (2022)</ref>, we defined a "contamination fraction", f cont , which expresses the probability for an optical concentration of red galaxies to be a chance alignment along the line of sight to the X-ray source. This is the key selection criterion for selecting cluster candidates, and it immediately provides an estimate of the catalogue contamination. A catalogue created by selecting f cont &lt; a is expected to have a contamination fraction of a, assuming the input catalogue is highly contaminated.</p><p>Because of the high number density of sources in the pointlike sample and/or the possibility that the emission from an actual cluster is split into many point sources, it can happen that many close X-ray sources point to the same optical cluster. A simple cut in f cont will therefore yield a much larger sample of sources than real clusters in that catalog, causing the contamination fraction to be much higher than expected. To compensate for this, for each eFEDS point source that is close to an optical overdensity, an environmental flag is set to true for the source that is closest to the overdensity and that is at least 0.75 Mpc away from a cluster detected in the extent-selected sample (Liu et al. 2022a,c; Klein et al. 2022) at similar redshift. Only when the flag is set to true is the point-source further considered as a candidate for being a cluster.</p><p>For this latter subgroup of sources, following Klein et al. (2022), MCMF assigns a redshift to the cluster (via the red sequence). In addition, the photo-z of the counterpart to the point-like sources is recomputed assuming they are passive galaxies. The two redshifts are then compared with the redshift computed by LEPHARE, as described in the previous section (Sect <ref type="bibr">. 7)</ref>.</p><p>Combining all the information described above, we define a new flag, Cluster_Class, which indicates the possibility that an eFEDS X-ray (point-like) source is actually a cluster or belongs to a cluster.</p><p>-Cluster_class = 5: CTP_QUALITY &#8804; 1 &amp; f cont &lt; 0.2 and the environmental flag set to true: the counterpart NWAY or ASTROMATCH is considered unreliable and the X-ray emission is more likely associated with a cluster (top left panel of Fig. <ref type="figure">17</ref>; 120 cases).</p><p>-Cluster_class = 4: CTP_QUALITY &#8805; 2 &amp; f cont &lt; 0.2 with the environmental flag set to true, the optical colours of the sources are typical of passive galaxies, and the redshift computed with LEPHARE coincides with the redshift of the optical cluster: the counterpart is reliable, and the point source is a galaxy member (possibly the BCG) of the optically detected cluster (top right panel in Fig. <ref type="figure">17</ref>; 63 cases).</p><p>-Cluster_class = 3: CTP_QUALITY &#8805; 2 &amp; f cont &lt; 0.2 and the environmental flag set to true and the redshift computed assuming an AGN template is consistent with the redshift of the optical cluster, but the optical colours of the counterpart are not typical of a passive galaxy: the counterpart is correct and the source is a cluster member (bottom left panel from the left of Fig. <ref type="figure">17</ref>; 96 cases).</p><p>-Cluster_class = 2: CTP_QUALITY &#8805; 2 &amp; f cont &lt; 0.01 and the environmental flag set to true, while the photo-z computed by the three methods disagree: the counterpart is reliable, and the source (AGN) is just projected on a likely cluster (bottom right panel of Fig. <ref type="figure">17</ref> In all the other cases, the X-ray emission is from a genuine point-source and likely not from the extended, hot intercluster medium. A dedicated effort is currently ongoing to confirm the secure clusters in the point-source catalogue and to characterise and measure the X-ray and radio properties of the confirmed clusters (e.g. <ref type="bibr">Bulbul et al. 2022</ref>).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="8.">Data release</head><p>The catalogs listing the properties of the counterparts to eFEDS point-like sources in the main and hard samples samples <ref type="bibr">(Brunner et al. 2022</ref>) associated with this paper are available via CDS/Vizier and via the web page at MPE dedicated to the eROSITA data release<ref type="foot">foot_9</ref> . The list of the columns and their description for the two samples is available in Appendix D. Only the basic X-ray properties are listed here (columns 1-9). For the complete list, we refer to the catalogs released by <ref type="bibr">Brunner et al. (2022)</ref>. After the columns reporting the key X-ray properties of the sources, Cols. 10-36 report the results of the counterpart (CTP) association, followed by the key parameters from NWAY, ASTROMATCH, and HamStar. Next (Cols. 36-49) we present the photometry from the recent Gaia EDR3 release in the original photometric system, followed by all the collected photometry, corrected for extinction (Cols. . We recall that the HSC photometry from S19A in i and r bands was split into i, i2, and r, r2, and that Kron is listed for EXT sources, while cmodel is listed for PLIKE (see Sect.  </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="9.">Discussion</head><p>The size and depth of the eFEDS X-ray survey, combined with ancillary data both in photometry and spectroscopy, allows us to paint a comprehensive picture of the average population of X-ray sources that contribute the bulk of the cosmic X-ray background (CXB) flux at energies &lt;10 keV (see e.g. <ref type="bibr">Gilli et al. 2007</ref>) in its Galactic and extragalactic content. The identification of the optical/IR counterparts, to a high degree of completeness and reliability, as discussed here, will facilitate detailed population studies of X-ray active stars, Galactic compact objects, and AGN. Here we briefly outline the main properties of our sample by examining the distributions of the X-ray sources in various colour/redshift spaces in detail.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="9.1.">Population studies</head><p>Figure <ref type="figure">18</ref> shows the distribution of all the eFEDS sources with a secure counterpart (CTP_quality &#8805; 2; see Sect. 5) in four different multi-band photometric spaces, chosen for their wide applicability to large areas of the sky.</p><p>The top panel shows sources in the z-W1 versus g-r space, colour-coded by their redshift. A few representative tracks of  various classes of extragalactic objects are overlaid. In addition to the clear separation between Galactic and extragalactic objects we discussed in Sect. 6, the X-ray points identify clear sequences of unobscured QSO, obscured Seyferts, and inactive galaxies. The inactive galaxies are best represented by the S0 and elliptical tracks, suggesting that some of them are the sources that are associated (or confused) with a cluster. The sources indicated by a yellow circle in the top left panel have CLUSTER_CLASS = 3,4 indicating that they belong to a cluster. In most cases, they are the BCG (see Sect. 7.5). These sources are best fit by the template of a passive galaxy, as the spectra for those available also suggest (e.g. lack of emission lines from star formation, strong  HK lines). However, some of the spectra together with the clear features from a non-star-forming galaxy also reveal the presence of broad emission lines typical of AGN (see <ref type="bibr">Bulbul et al. 2022)</ref>.</p><p>The top right panel of Fig. <ref type="figure">18</ref> shows the distribution of points in the mid-infrared (WISE W1) versus soft X-ray (0.5-2 keV) plane (same as Fig. <ref type="figure">9</ref>), originally introduced in Salvato et al. (2018). X-ray bright objects above the dashed line are typically AGN, while most of the IR bright objects below the line are Galactic X-ray emitting stars, with some contamination from nearby extragalactic objects. These sources are rare, but given the size of eFEDS, their number is non-negligible. Thus, when using this plot for other surveys, the size of the survey must be accounted for. The larger the surveys, the less efficient the line separator.</p><p>The bottom right panel shows the distribution of the sources in the Wise-only W1-W2 versus W2 colour-magnitude plane <ref type="foot">12</ref> . This is widely used to classify point sources, as it easily separates stars, with W1-W2 &#8776; 0, from QSOs, with W1-W2 &gt; 0.5 (see e. Finally, the bottom left panel shows the distribution of the eFEDS sources in the optical/mid-infrared diagram defined by the "all-sky available" G-W1 versus W1-W2, which is frequently used to separate QSO from stars in the Gaia catalog. As already pointed out in Sect. 6, 10% of the Galaxtic sources are too faint to be detected by Gaia. This is even more true for the extragalactic sources: the plot shows only 57% of the entire eFEDS sample. However, the plot shows insights into the population that the first eROSITA All-Sky Survey (eRASS1) will uncover. As expected, the X-ray selected eFEDS sources contain beyond stars and (unobscured) QSOs a tail at high G-W1 (i.e. bright mid-infrared, faint optical magnitudes) typical of inactive galaxies and/or mildly obscured AGN. This is indeed confirmed by comparing the location of the extragalactic eFEDS sources in the grzW1 plane with the X-ray hardness ratio measured from the X-ray counts in the bands in which eROSITA is most sensitive. The right panel of Fig. <ref type="figure">19</ref> shows the distribution of the sources in this plane, colour-coded by their average hardness ratio (defined as (H-S)/(H+S), where H and S are the counts in the ranges 1.0-2.0 keV 0.2-1.0 keV 13 , respectively), while the left panel highlights the loci of the most common classes of sources based on the distribution of template tracks. The hardest sources in the eROSITA band populate the optical/mid-infrared colour-space of Seyfert 2 galaxies and/or reddened QSOs. A detailed discussion of the X-ray spectral properties of the AGN in the main eFEDS sample will be presented in Liu et al. (2022b).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="9.2.">eFEDS stellar content</head><p>The eFEDS field spans a wide range of Galactic latitudes (from about +20 to about +40). Reassuringly, the fraction of the Xray sources that are classified as Galactic (see Sect. eFEDS. About 22.3% of the Galactic sources identified by NWAY/ASTROMATCH are fainter than the 19th magnitude (10% are not detected by Gaia).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="10.">Conclusions</head><p>We have presented the identification of the counterparts to the point sources in eFEDS listed in the main and hard catalogues <ref type="bibr">(Brunner et al. 2022</ref>), together with the study of their multiwavelength properties. eFEDS has a limiting flux of F 0.5-2 keV &#8764; 6.5 &#215; 10 -15 erg s -1 cm -2 and is a factor of &#8764;50% deeper than the final eROSITA all-sky survey. It can therefore also be used as a forecast for eRASS:8, not only for the population that eRASS:8 will reveal, but also for the challenges that are ahead of us with respect to counterpart identification and redshift determination.</p><p>-Counterpart identification: We used NWAY <ref type="bibr">(Salvato et al. 2019</ref>) and ASTROMATCH. In addition to spatial information from eFEDS and these codes use a prior based on the properties of a training sample of 3XMM sources that was tested on a validation sample of Chandra sources, made eFEDS-like in terms of positional accuracy. Each method has identified its own priors in a different way. For the validation sample, NWAY correctly identified 95% of the sources; only 2% of the sources have a possible second counterpart, compared with 89% and 10% for ASTRO-MATCH, but at the threshold adopted for p_any and LR_BEST, both methods have very high completeness and purity (above 95%). These remarkable results, well above the predicted completeness and purity mentioned in <ref type="bibr">Merloni et al. (2012)</ref>, are due to three important factors: the development of new methods for identifying the correct counterparts, large samples of X-ray detected sources with known counterparts, and the availability of sufficiently deep, homogenised, multi-wavelength photometry from optical to mid-infrared over very wide areas from which the SED of these sources that were to be used as training was constructed. In the next two years, by the time eROSITA will have completed the final all-sky survey, the methods will continue to improve and the training/validation samples will increase in size. Most importantly, the coverage of the multi-wavelength catalogs that are used to identify the counterparts will be larger. While the DESI Legacy Imaging Survey DR9 (LS9; <ref type="bibr">Dey et al. 2019</ref>) just became publicly available, the work on DR10 has started. The survey will cover virtually all of the eROSITA-DE area of the sky at sufficient depth, which is possible because the DECam data taken via the DeROSITAS survey are included (PI A. Zenteno). We predict that the identification of the counterparts for the entire eRASS will be of at least the same quality as eFEDS, also in the Galactic plane because the recently released Gaia EDR3 is included.</p><p>-Redshift determination: Given the lack of sufficiently deep near-infrared data outside the DES area (Sevilla-Noarbe et al. 2021), the possibility of obtaining reliable photometric redshifts via SED fitting will be low, at least until data from SpherEx <ref type="bibr">(Dor&#233; et al. 2018</ref>) will be made available (launch planned for Summer 2024). However, as demonstrated in Nishizawa et al. (in prep.) and <ref type="bibr">Borisov et al. (2021)</ref>, the increasing size and completeness of the spectroscopic sample that can be used for the training will enable reliable photometric redshifts for any type of X-ray extragalactic source. For example, the spectroscopic follow-up of the eROSITA-DE sources planned via Vista/4MOST and SDSS-V/BHM will allow us to obtain redshifts for 80% of the sources detected by eRASS:3, thus limiting the need of photo-z, and at the same time, ensuring a high quality of photo-z that will use these spectroscopically confirmed sources as training. 1. S0 2. S0_10_QSO2_90 3. S0_20_QSO2_80</p></div><note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="2" xml:id="foot_0"><p>i.e. comparable to the XMM Slew Survey: https://www.cosmos. esa.int/web/xmm-newton/xmmsl2-ug A3, page</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="3" xml:id="foot_1"><p>of 32 A&amp;A 661,A3 (2022)   </p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="3" xml:id="foot_2"><p>https://www.legacysurvey.org/dr8/description/ #astrometry</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="4" xml:id="foot_3"><p>https://www.cosmos.esa.int/web/gaia/earlydr3</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="5" xml:id="foot_4"><p>http://kids.strw.leidenuniv.nl/</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="6" xml:id="foot_5"><p>https://github.com/JohannesBuchner/NWAY A3, page 5 of 32</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="7" xml:id="foot_6"><p>https://github.com/ruizca/astromatch</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="8" xml:id="foot_7"><p>https://www.ivoa.net/documents/MOC/</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="9" xml:id="foot_8"><p>We only considered the radio components and made no attempt to handle complex sources appropriately.</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="11" xml:id="foot_9"><p>https://erosita.mpe.mpg.de/edr/eROSITAObservations/ Catalogues/ A3, page 20 of 32</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="12" xml:id="foot_10"><p>For this plot, we use the Vega System, so that the user can compare the figure with similar ones prepared using AllWISE all sky.</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="15" xml:id="foot_11"><p>in NWAY, p_any is the probability for each source in the primary catalogue (eFEDS in this case) to have a counterpart in the secondary catalogs; then, for each source in the secondary catalogues, p_i gives the probability to be the correct counterpart to the source in the primary catalogue (see more in the NWAY manual and Salvato et al.(2019)   </p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="16" xml:id="foot_12"><p>https://cxc.cfa.harvard.edu/csc2/index.html</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="17" xml:id="foot_13"><p>http://classic.sdss.org/dr5/algorithms/spectemplates</p></note>
			<note xmlns="http://www.tei-c.org/ns/1.0" place="foot" n="18" xml:id="foot_14"><p>The templates are slightly different than in Ananna et al. in the UV part.</p></note>
		</body>
		</text>
</TEI>
