NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

PIM GPT a hybrid process in memory accelerator for autoregressive transformers

https://doi.org/10.1038/s44335-024-00004-2

Wu, Yuting; Wang, Ziyu; Lu, Wei_D (July 2024, npj Unconventional Computing)

Abstract Decoder-only Transformer models such as Generative Pre-trained Transformers (GPT) have demonstrated exceptional performance in text generation by autoregressively predicting the next token. However, the efficiency of running GPT on current hardware systems is bounded by low compute-to-memory-ratio and high memory access. In this work, we propose a Process-in-memory (PIM) GPT accelerator, PIM-GPT, which achieves end-to-end acceleration of GPT inference with high performance and high energy efficiency. PIM-GPT leverages DRAM-based PIM designs for executing multiply-accumulate (MAC) operations directly in the DRAM chips, eliminating the need to move matrix data off-chip. Non-linear functions and data communication are supported by an application specific integrated chip (ASIC). At the software level, mapping schemes are designed to maximize data locality and computation parallelism. Overall, PIM-GPT achieves 41 − 137 × , 631 − 1074 × speedup and 123 − 383 × , 320 − 602 × energy efficiency over GPU and CPU baseline on 8 GPT models with up to 1.4 billion parameters.
more » « less
Bulk‐Switching Memristor‐Based Compute‐In‐Memory Module for Deep Neural Network Training

https://doi.org/10.1002/adma.202305465

Wu, Yuting; Wang, Qiwen; Wang, Ziyu; Wang, Xinxin; Ayyagari, Buvna; Krishnan, Siddarth; Chudzik, Michael; Lu, Wei_D (October 2023, Advanced Materials)

Abstract The constant drive to achieve higher performance in deep neural networks (DNNs) has led to the proliferation of very large models. Model training, however, requires intensive computation time and energy. Memristor‐based compute‐in‐memory (CIM) modules can perform vector‐matrix multiplication (VMM) in place and in parallel, and have shown great promises in DNN inference applications. However, CIM‐based model training faces challenges due to non‐linear weight updates, device variations, and low‐precision. In this work, a mixed‐precision training scheme is experimentally implemented to mitigate these effects using a bulk‐switching memristor‐based CIM module. Low‐precision CIM modules are used to accelerate the expensive VMM operations, with high‐precision weight updates accumulated in digital units. Memristor devices are only changed when the accumulated weight update value exceeds a pre‐defined threshold. The proposed scheme is implemented with a system‐onchip of fully integrated analog CIM modules and digital sub‐systems, showing fast convergence of LeNet training to 97.73%. The efficacy of training larger models is evaluated using realistic hardware parameters and verifies that CIM modules can enable efficient mix‐precision DNN training with accuracy comparable to full‐precision software‐trained models. Additionally, models trained on chip are inherently robust to hardware variations, allowing direct mapping to CIM inference chips without additional re‐training.
more » « less
RN‐Net: Reservoir Nodes‐Enabled Neuromorphic Vision Sensing Network

https://doi.org/10.1002/aisy.202400265

Yoo, Sangmnin; Lee, Eric_Yeu‐Jer; Wang, Ziyu; Wang, Xinxin; Lu, Wei_D (May 2024, Advanced Intelligent Systems)

Neuromorphic computing systems promise high energy efficiency and low latency. In particular, when integrated with neuromorphic sensors, they can be used to produce intelligent systems for a broad range of applications. An event‐based camera is such a neuromorphic sensor, inspired by the sparse and asynchronous spike representation of the biological visual system. However, processing the event data requires either using expensive feature descriptors to transform spikes into frames, or using spiking neural networks (SNNs) that are expensive to train. In this work, a neural network architecture is proposed, reservoir nodes‐enabled neuromorphic vision sensing network (RN‐Net), based on dynamic temporal encoding by on‐sensor reservoirs and simple deep neural network (DNN) blocks. The reservoir nodes enable efficient temporal processing of asynchronous events by leveraging the native dynamics of the node devices, while the DNN blocks enable spatial feature processing. Combining these blocks in a hierarchical structure, the RN‐Net offers efficient processing for both local and global spatiotemporal features. RN‐Net executes dynamic vision tasks created by event‐based cameras at the highest accuracy reported to date at one order of magnitude smaller network size. The use of simple DNN and standard backpropagation‐based training rules further reduces implementation and training costs.
more » « less
Efficient data processing using tunable entropy-stabilized oxide memristors

https://doi.org/10.1038/s41928-024-01169-1

Yoo, Sangmin; Chae, Sieun; Chiang, Tony; Webb, Matthew; Ma, Tao; Paik, Hanjong; Park, Yongmo; Williams, Logan; Nomoto, Kazuki; Xing, Huili G; et al (May 2024, Nature Electronics)

Full Text Available
TT-CIM: Tensor Train Decomposition for Neural Network in RRAM-Based Compute-in-Memory Systems

https://doi.org/10.1109/TCSI.2023.3344550

Meng, Fan-Hsuan; Wu, Yuting; Zhang, Zhengya; Lu, Wei D (March 2024, IEEE Transactions on Circuits and Systems I: Regular Papers)

Full Text Available
Compute-In-Memory Technologies for Deep Learning Acceleration

https://doi.org/10.1109/MNANO.2023.3340321

Meng, Fan-husan; Lu, Wei D (February 2024, IEEE Nanotechnology Magazine)

Full Text Available
PowerGAN: A Machine Learning Approach for Power Side‐Channel Attack on Compute‐in‐Memory Accelerators

https://doi.org/10.1002/aisy.202300313

Wang, Ziyu; Wu, Yuting; Park, Yongmo; Yoo, Sangmin; Wang, Xinxin; Eshraghian, Jason_K; Lu, Wei_D (September 2023, Advanced Intelligent Systems)

Analog compute‐in‐memory (CIM) systems are promising candidates for deep neural network (DNN) inference acceleration. However, as the use of DNNs expands, protecting user input privacy has become increasingly important. Herein, a potential security vulnerability is identified wherein an adversary can reconstruct the user's private input data from a power side‐channel attack even without knowledge of the stored DNN model. An attack approach using a generative adversarial network is developed to achieve high‐quality data reconstruction from power leakage measurements. The analyses show that the attack methodology is effective in reconstructing user input data from power leakage of the analog CIM accelerator, even at large noise levels and after countermeasures. To demonstrate the efficacy of the proposed approach, an example of CIM inference of U‐Net for brain tumor detection is attacked, and the original magnetic resonance imaging medical images can be successfully reconstructed even at a noise level of 20% standard deviation of the maximum power signal value. This study highlights a potential security vulnerability in emerging analog CIM accelerators and raises awareness of needed safety features to protect user privacy in such systems.
more » « less
AR-PIM: An Adaptive-Range Processing-in-Memory Architecture

https://doi.org/10.1109/ISLPED58423.2023.10244186

Chou, Teyuh; Garcia-Redondo, Fernando; Whatmough, Paul; Zhang, Zhengya (August 2023, IEEE)

Full Text Available
Exploring Compute-in-Memory Architecture Granularity for Structured Pruning of Neural Networks

https://doi.org/10.1109/JETCAS.2022.3227471

Meng, Fan-Hsuan; Wang, Xinxin; Wang, Ziyu; Lee, Eric Yeu-Jer; Lu, Wei D. (December 2022, IEEE Journal on Emerging and Selected Topics in Circuits and Systems)

Full Text Available
RM-NTT: An RRAM-Based Compute-in-Memory Number Theoretic Transform Accelerator

https://doi.org/10.1109/JXCDC.2022.3202517

Park, Yongmo; Wang, Ziyu; Yoo, Sangmin; Lu, Wei D. (December 2022, IEEE Journal on Exploratory Solid-State Computational Devices and Circuits)

« Prev Next »

Search for: All records