skip to main content
US FlagAn official website of the United States government
dot gov icon
Official websites use .gov
A .gov website belongs to an official government organization in the United States.
https lock icon
Secure .gov websites use HTTPS
A lock ( lock ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites.


Title: The CAMELS Project: Expanding the Galaxy Formation Model Space with New ASTRID and 28-parameter TNG and SIMBA Suites
Abstract We present CAMELS-ASTRID, the third suite of hydrodynamical simulations in the Cosmology and Astrophysics with MachinE Learning (CAMELS) project, along with new simulation sets that extend the model parameter space based on the previous frameworks of CAMELS-TNG and CAMELS-SIMBA, to provide broader training sets and testing grounds for machine-learning algorithms designed for cosmological studies. CAMELS-ASTRID employs the galaxy formation model following the ASTRID simulation and contains 2124 hydrodynamic simulation runs that vary three cosmological parameters (Ωm8, Ωb) and four parameters controlling stellar and active galactic nucleus (AGN) feedback. Compared to the existing TNG and SIMBA simulation suites in CAMELS, the fiducial model of ASTRID features the mildest AGN feedback and predicts the least baryonic effect on the matter power spectrum. The training set of ASTRID covers a broader variation in the galaxy populations and the baryonic impact on the matter power spectrum compared to its TNG and SIMBA counterparts, which can make machine-learning models trained on the ASTRID suite exhibit better extrapolation performance when tested on other hydrodynamic simulation sets. We also introduce extension simulation sets in CAMELS that widely explore 28 parameters in the TNG and SIMBA models, demonstrating the enormity of the overall galaxy formation model parameter space and the complex nonlinear interplay between cosmology and astrophysical processes. With the new simulation suites, we show that building robust machine-learning models favors training and testing on the largest possible diversity of galaxy formation models. We also demonstrate that it is possible to train accurate neural networks to infer cosmological parameters using the high-dimensional TNG-SB28 simulation set.  more » « less
Award ID(s):
2108678 2108944
PAR ID:
10479791
Author(s) / Creator(s):
; ; ; ; ; ; ; ; ; ; ; ; ; ;
Publisher / Repository:
DOI PREFIX: 10.3847
Date Published:
Journal Name:
The Astrophysical Journal
Volume:
959
Issue:
2
ISSN:
0004-637X
Format(s):
Medium: X Size: Article No. 136
Size(s):
Article No. 136
Sponsoring Org:
National Science Foundation
More Like this
  1. ABSTRACT We quantify the cosmological spread of baryons relative to their initial neighbouring dark matter distribution using thousands of state-of-the-art simulations from the Cosmology and Astrophysics with MachinE Learning Simulations (CAMELS) project. We show that dark matter particles spread relative to their initial neighbouring distribution owing to chaotic gravitational dynamics on spatial scales comparable to their host dark matter halo. In contrast, gas in hydrodynamic simulations spreads much further from the initial neighbouring dark matter owing to feedback from supernovae (SNe) and active galactic nuclei (AGN). We show that large-scale baryon spread is very sensitive to model implementation details, with the fiducial simba model spreading ∼40 per cent of baryons >1 Mpc away compared to ∼10 per cent for the IllustrisTNG and astrid models. Increasing the efficiency of AGN-driven outflows greatly increases baryon spread while increasing the strength of SNe-driven winds can decrease spreading due to non-linear coupling of stellar and AGN feedback. We compare total matter power spectra between hydrodynamic and paired N-body simulations and demonstrate that the baryonic spread metric broadly captures the global impact of feedback on matter clustering over variations of cosmological and astrophysical parameters, initial conditions, and (to a lesser extent) galaxy formation models. Using symbolic regression, we find a function that reproduces the suppression of power by feedback as a function of wave number (k) and baryonic spread up to $$k \sim 10\, h$$ Mpc−1 in SIMBA while highlighting the challenge of developing models robust to variations in galaxy formation physics implementation. 
    more » « less
  2. Abstract Most diffuse baryons, including the circumgalactic medium (CGM) surrounding galaxies and the intergalactic medium (IGM) in the cosmic web, remain unmeasured and unconstrained. Fast radio bursts (FRBs) offer an unparalleled method to measure the electron dispersion measures (DMs) of ionized baryons. Their distribution can resolve the missing baryon problem and constrain the history of feedback theorized to impart significant energy to the CGM and IGM. We analyze the Cosmology and Astrophysics with Machine Learning Simulations using three suites, IllustrisTNG, SIMBA, and Astrid, each varying six parameters (two cosmological and four astrophysical feedback), for a total of 183 distinct simulation models. We find significantly different predictions between the fiducial models of the suites owing to their different implementations of feedback. SIMBA exhibits the strongest feedback, leading to the smoothest distribution of baryons and reducing the sight-line-to-sight-line variance in DMs betweenz= 0 and 1. Astrid has the weakest feedback and the largest variance. We calculate FRB CGM measurements as a function of galaxy impact parameter, with SIMBA showing the weakest DMs due to aggressive active galactic nucleus (AGN) feedback and Astrid the strongest. Within each suite, the largest differences are due to varying AGN feedback. IllustrisTNG shows the most sensitivity to supernova feedback, but this is due to the change in the AGN feedback strengths, demonstrating that black holes, not stars, are most capable of redistributing baryons in the IGM and CGM. We compare our statistics directly to recent observations, paving the way for the use of FRBs to constrain the physics of galaxy formation and evolution. 
    more » « less
  3. Most diffuse baryons, including the circumgalactic medium (CGM) surrounding galaxies and the intergalactic medium (IGM) in the cosmic web, remain unmeasured and unconstrained. Fast radio bursts (FRBs) offer an unparalleled method to measure the electron dispersion measures (DMs) of ionized baryons. Their distribution can resolve the missing baryon problem and constrain the history of feedback theorized to impart significant energy to the CGM and IGM. We analyze the Cosmology and Astrophysics with Machine Learning Simulations using three suites, IllustrisTNG, SIMBA, and Astrid, each varying six parameters (two cosmological and four astrophysical feedback), for a total of 183 distinct simulation models. We find significantly different predictions between the fiducial models of the suites owing to their different implementations of feedback. SIMBA exhibits the strongest feedback, leading to the smoothest distribution of baryons and reducing the sight-line-to-sight-line variance in DMs between z = 0 and 1. Astrid has the weakest feedback and the largest variance. We calculate FRB CGM measurements as a function of galaxy impact parameter, with SIMBA showing the weakest DMs due to aggressive active galactic nucleus (AGN) feedback and Astrid the strongest. Within each suite, the largest differences are due to varying AGN feedback. IllustrisTNG shows the most sensitivity to supernova feedback, but this is due to the change in the AGN feedback strengths, demonstrating that black holes, not stars, are most capable of redistributing baryons in the IGM and CGM. We compare our statistics directly to recent observations, paving the way for the use of FRBs to constrain the physics of galaxy formation and evolution. 
    more » « less
  4. Abstract As the next generation of large galaxy surveys come online, it is becoming increasingly important to develop and understand the machine-learning tools that analyze big astronomical data. Neural networks are powerful and capable of probing deep patterns in data, but they must be trained carefully on large and representative data sets. We present a new “hump” of the Cosmology and Astrophysics with MachinE Learning Simulations (CAMELS) project: CAMELS-SAM, encompassing one thousand dark-matter-only simulations of (100h−1cMpc)3with different cosmological parameters (Ωmandσ8) and run through the Santa Cruz semi-analytic model for galaxy formation over a broad range of astrophysical parameters. As a proof of concept for the power of this vast suite of simulated galaxies in a large volume and broad parameter space, we probe the power of simple clustering summary statistics to marginalize over astrophysics and constrain cosmology using neural networks. We use the two-point correlation, count-in-cells, and void probability functions, and we probe nonlinear and linear scales across 0.68 <R<27h−1cMpc. We find our neural networks can both marginalize over the uncertainties in astrophysics to constrain cosmology to 3%–8% error across various types of galaxy selections, while simultaneously learning about the SC-SAM astrophysical parameters. This work encompasses vital first steps toward creating algorithms able to marginalize over the uncertainties in our galaxy formation models and measure the underlying cosmology of our Universe. CAMELS-SAM has been publicly released alongside the rest of CAMELS, and it offers great potential to many applications of machine learning in astrophysics:https://camels-sam.readthedocs.io. 
    more » « less
  5. The baryonic physics shaping galaxy formation and evolution are complex, spanning a vast range of scales and making them challenging to model. Cosmological simulations rely on subgrid models that produce significantly different predictions. Understanding how models of stellar and active galactic nucleus (AGN) feedback affect baryon behavior across different halo masses and redshifts is essential. Using the SIMBA and IllustrisTNG suites from the Cosmology and Astrophysics with MachinE Learning Simulations (CAMELS) project, we explore the effect of parameters governing the subgrid implementation of stellar and AGN feedback. We find that while IllustrisTNG shows higher cumulative feedback energy across all halos, SIMBA demonstrates a greater spread of baryons, quantified by the closure radius and circumgalactic medium (CGM) gas fraction. This suggests that feedback in SIMBA couples more effectively to baryons and drives them more efficiently within the host halo. There is evidence that the different feedback modes are highly interrelated in these subgrid models. The parameters controlling the stellar feedback efficiency significantly impact AGN feedback, as seen in the suppression of black hole mass growth and delayed activation of AGN feedback to higher-mass halos with increasing stellar feedback efficiency in both simulations. Additionally, the AGN feedback efficiency parameters affect the CGM gas fraction at low halo masses in SIMBA, hinting at complex, nonlinear interactions between the AGN and supernova feedback modes. Overall, we demonstrate that stellar and AGN feedback are intimately interwoven, especially at low redshift, due to subgrid implementation, resulting in halo property effects that might initially seem counterintuitive. 
    more » « less