Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher.
Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?
Some links on this page may take you to non-federal websites. Their policies may differ from this site.
-
Physics intelligence and digital twins often require rapid and repeated performance evaluation of various engineering systems (e.g. robots, autonomous vehicles, semiconductor chips) to enable (almost) real-time actions or decision making. This has motivated the development of accelerated partial differential equation (PDE) solvers, in resource-constrained scenarios if the PDE solvers are to be deployed on the edge. Physics-informed neural networks (PINNs) have shown promise in solving high-dimensional PDEs, but the training time on state-of-the-art digital hardware (e.g., GPUs) is still orders-of-magnitude longer than the latency required for enabling real-time decision making. Photonic computing offers a potential solution to address this huge latency gap because of its ultra-high operation speed. However, the lack of photonic memory and the large device sizes prevent training real-size PINNs on photonic chips. This paper proposes a completely back-propagation-free (BP-free) and highly scalable framework for training real-size PINNs on silicon photonic platforms. Our approach involves three key innovations: (1) a sparse-grid Stein derivative estimator to avoid the BP in the loss evaluation of a PINN, (2) a dimension-reduced zeroth-order optimization via tensor-train decomposition to achieve better scalability and convergence in BP-free training, and (3) a scalable on-chip photonic PINN training accelerator design using photonic tensor cores. We validate our numerical methods on both low- and high-dimensional PDE benchmarks. Through pre-silicon simulation based on real device parameters, we further demonstrate the significant performance benefit (e.g., real-time training, huge chip area reduction) of our photonic accelerator. Our code is available at https://github.com/olokevin/scalable_bpfree_onn_training.more » « lessFree, publicly-accessible full text available August 5, 2027
-
Free, publicly-accessible full text available April 28, 2027
-
Free, publicly-accessible full text available November 26, 2026
-
Free, publicly-accessible full text available February 1, 2027
-
Abstract Mutations in isocitrate dehydrogenase 1 (IDH1) and 2 (IDH2) are common in multiple types of human cancer and cause accumulation of the oncometaboliteD-2-hydroxyglutarate (D2HG) instead of α-ketoglutarate, driving cancers like gliomas and acute myeloid leukaemia by blocking cell differentiation and promoting tumour growth. Here we discovered proteinO-2-hydroxyglutarylation by D2HG using chemical proteomics and further revealed distinct chiral preferences for D2HG andL-2-hydroxyglutarate (L2HG) modifications. D2HG modifications are upregulated in IDH-mutant cells or upon D2HG treatment, while L2HG modifications increase under hypoxic conditions or following L2HG treatment. Notably, two kinases MRCKA and SLK are modified by D2HG and L2HG, respectively, and confirmed by synthetic peptide standards. Phosphoproteomics revealed reduced phosphorylation of MRCKA and SLK substrates, suggesting crosstalk between D/L-2HG modification and kinase activity. These findings highlight distinctive roles of D/L-2HG modifications in cancer progression and suggest potential avenues for therapeutic targeting of oncometabolite-induced post-translational modifications.more » « lessFree, publicly-accessible full text available June 1, 2027
-
Free, publicly-accessible full text available June 3, 2027
-
Back propagation (BP) is the default solution for gradient computation in neural network training. However, implementing BP-based training on various edge devices such as FPGA, microcontrollers (MCUs), and analog computing platforms faces multiple major challenges, such as the lack of hardware resources, long time-to-market, and dramatic errors in a low-precision setting. This article presents a simple BP-free training scheme on an MCU, which makes edge training hardware design as easy as inference hardware design. We adopt a quantized zeroth-order method to estimate the gradients of quantized model parameters, which can overcome the error of a straight-through estimator in a low-precision BP scheme. We further employ a few dimension reduction methods (e.g., node perturbation, sparse training) to improve the convergence of zeroth-order training. Experiment results show that our BP-free training achieves comparable performance as BP-based training on adapting a pre-trained image classifier to various corrupted data on resource-constrained edge devices (e.g., an MCU with 1024-KB SRAM for dense full-model training, or an MCU with 256-KB SRAM for sparse training). This method is most suitable for application scenarios where memory cost and time-to-market are the major concerns, but longer latency can be tolerated.more » « lessFree, publicly-accessible full text available September 30, 2026
An official website of the United States government

Full Text Available