Silicon Photonics Enabled High-Performance, Energy-Efficient, Flexible, and Scalable Deep Neural Network Accelerator Design
Open AccessDeep neural networks (DNNs) have achieved unprecedented success in a wide range of applications such as image classification and speech recognition. The superior inference accuracy of DNNs comes at the cost of high volumes of data and computation. Consequently, there has been an increased emphasis on accelerating DNNs by exploiting parallelism and specialization. The communication network is at the heart of a DNN accelerator and responsible for connecting numerous processing units and orchestrating data movement incurred from strategic scheduling of computation. As the intrinsic limitations of metallic-based interconnects have become increasingly severe along with the scaling of DNN accelerators, we propose to explore other disruptive interconnect technologies such as silicon photonics. However, the adoption of new technologies can lead to a paradigm shift in both network architecture design and dataflow optimization. This dissertation reexamines the basic properties of DNN applications in the context of silicon photonics and explores the network and dataflow co-design approach to improve performance, energy efficiency, flexibility, and scalability.We first develop a general chiplet-based platform where a variety of existing chip-scale DNN accelerators can be assembled to enable efficient scaling. This platform includes (1) a photonic network for inter-chiplet communication that can be dynamically tuned to adapt to different sets of communication patterns incurred from assembled accelerator chips, and (2) a system-level dataflow that leverages the broadcast capability of silicon photonics and minimizes the costly electrical-to-optical and optical-to-electrical signal conversions.We then extend the photonic network design to support seamless inter-chiplet and intra-chiplet communications and reveal the unique considerations for dataflow optimization when silicon photonic interconnects are assumed. Specifically, we propose a photonic network that enables single-chiplet and cross-chiplet multicast communications by establishing single-write-multiple-read channels within different sets of processing units. We propose a dataflow that maximizes spatial parallelism and minimizes unnecessary movement of intermediate data.We also address the communication challenges of the simultaneous implementation of multiple inference tasks on the same DNN accelerator in the cloud. In this case, hardware resources are partitioned and then regrouped for the execution of concurrent inference tasks. We propose three techniques: (1) a photonic network that can be divided into sub-networks at runtime, each providing seamless one-hop communication between partitions allocated to a specific inference task despite their physical locations, by leveraging the distance-independent latency nature of silicon photonics; (2) an extension of the previous output-stationary multicast-promoted dataflow so that it can be performed across physically disconnected partitions; (3) an algorithm that optimally allocates hardware resources based on computation demand and task priority. The combined effects of the above three techniques improve performance and energy efficiency, as well as Quality-of-Service and fairness.We lastly develop a highly flexible accelerator architecture that can accommodate spatially co-located inference tasks with diverse dataflows. The need for diverse dataflows is from the well-known observation that no dataflow is universally optimal due to the variations in DNN models and available hardware resources. We first carry out a two-step exhaustive dataflow exploration (on-chip and off-chip) and enumerate the set of communication patterns observed in each dataflow setup. Based on the information obtained from dataflow exploration, we design a photonic network that can not only be dynamically divided into sub-networks, but also adapt to the communication patterns of explored dataflows. In addition to the algorithm that allocates hardware resources to each concurrent inference task, we develop another algorithm that selects the optimal dataflow for an inference task based on the DNN model of this task and the allocated hardware resources, with the aim of minimizing the number of accesses to off-chip memory and on-chip global buffer.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.