High-Performance, Energy-Efficient, and Scalable Machine Learning Accelerator Design with Adaptive and Approximate On-Chip Interconnection
Open AccessMachine learning has been widely used in many applications, ranging from healthcare diagnostics and image classification to natural language processing and content generation. To efficiently run these applications, accelerators are developed for machine learning model inference or training. Machine learning accelerators are designed with multi-core or multi-chiplet systems to meet the demand for computation when handling inference or training workloads. However, existing accelerators consume large amounts of time and power transmitting data across on-chip interconnects, causing low utilization of arithmetic-logic units, extensive training or inference time, and high power consumption. As future machine learning workloads integrate more functionalities by scaling models with deeper layers, larger training datasets, and more complex model structures, the on-chip interconnect becomes the primary bottleneck for existing accelerator designs.This dissertation research explores the error tolerance and communication pattern of machine learning workloads and proposes accelerator designs with approximate and adaptive on-chip interconnection networks to reduce the power consumption and latency of data movement for efficient model training or inference. We observed that most of the machine learning models are implemented with neuron networks, which can tolerate modest errors during training and inference. Thus, we first develop the approximate communication framework for general approximate computing applications, which incorporates a quality control method and a data approximation mechanism to reduce the packet size for lower network power consumption and latency. Then, the framework is extended to support the approximation of the inference of the deep convolution neural networks. Specifically, the approximate communication framework incorporates quantization and contrast reduction techniques for fewer flits in each data packet and reduces traffic in network-on-chips (NoCs) while guaranteeing classification accuracy during the model inference. Finally, we propose two accelerator designs for efficient and scalable machine learning training with adaptive on-chip interconnection design. For efficient training of sparse models, we develop an adaptive compression technique for on-chip interconnect with a two-level load balance algorithm to address the workload imbalance problem as the model is pruned during training. We further enhance the scalability and efficiency of the chiplet-based system with adaptive network topology to handle the intensive and complex communication pattern during generative adversarial network (GAN) training.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.