Machine Learning Enabled Network-on-Chip Design for Low-power, High-Performance, and Flexible Manycore Architectures
Open AccessThe proliferation of core counts in chip-multiprocessors (CMP) enables the extraordinary growth in computing power to harness parallelism. The most critical barrier toward exploiting the parallelism lies in the underlying communication fabrics. This signals the advent of a paradigm shift from computation-centric to communication-centric systems. Consequently, a high-performance, power-efficient, and flexible Network-on-Chip (NoC) architecture is of great importance to current and future computing systems. This dissertation is to overcome multiple communication challenges facing future manycore computing systems. Specifically, we focus on four research tasks namely, (1) designing power and performance efficient NoC architectures, (2) applying machine learning techniques to NoC power management, and (3) enabling flexibility in heterogeneous manycores, and (4) optimizing the data movement for chiplet-based computing systems. First, the end of voltage scaling has made power consumption as the primary focus in current chip designs. To reduce the power consumption, a variety of power-saving techniques have been deployed at both circuit and architectural levels, but, without a careful design, these techniques can conversely jeopardize the system performance. We target managing the power consumption and performance together for on-chip communication. Specifically, we built simulation models to characterize the application behavior and identify the bottlenecks of deploying power-saving techniques. Leveraging the findings from our simulation studies, we designed a bypass routing mechanism to continue the packet transmission during sporadic traffic. The proposed design exploits the full benefits of power-gating while circumventing its performance overhead. Second, the design space for optimized NoCs has expanded significantly comprising a plethora of techniques. The simultaneous application of these techniques requires monitoring of a large number of system parameters, prediction of runtime application features, and decision to select the right mix of solutions. Manually designing rules and strategies to handle such a large design space is likely to require a substantial engineering effort and enormous resources. Machine learning techniques can be applied to automatically infer complex problems to meet the target objective. Given that, we developed a reinforcement learning (RL) approach in NoC architectures to self-learn a control policy with the goal of maximizing the power reductions and performance improvements. The proposed RL automatically explores the dynamic interactions among power gating, dynamic voltage and frequency scaling, and system parameters, learns the critical system parameters contained in the router and cache, and eventually evolves the optimal power management policy. This study shows that machine learning techniques could guide us to new and better solutions in NoC power management. Despite the many benefits of machine learning, it is costly to implement the RL model in the hardware. we thus proposed a simple artificial neural network to approximate the state-action table required by the RL model.Third, modern heterogeneous manycore architectures are comprised of a large collection of computing resources such as CPUs, GPUs, and accelerators. While the increased computational capability and diversity facilitate the concurrent execution of multiple applications, it puts a large burden on the on-chip communication fabric. To tackle the mentioned problem, we proposed a flexible NoC architecture that can provide efficient communication support for concurrent application execution. Specifically, the proposed design can dynamically allocate several disjoint regions of the NoC, called subNoCs, with different sizes and locations for various running applications. Each of the dynamically-allocated subNoC is capable of supporting a given topology such as a mesh, cmesh, torus, or tree, and thus satisfying different types of communication in terms of latency, bandwidth, power, and traffic patterns. Lastly, advanced packaging technology, such as silicon interposer, has emerged as an alternative to sustain the performance scaling by integrating multiple smaller chiplets within a single package. The tightly integrated a variety of chiplets compose a large-scale computing system, offering substantial computing capability to harness both heterogeneity and parallelism. However, such computing capability can only be unleashed if the underlying interconnection network can provide the required bandwidth and latency. We explored the communication needs existed in chiplet-based computing systems such as heterogeneous manycore architectures and deep neural network accelerators. On the top of this understanding, we designed a flexible interconnection network that can be tailored in response to myriad communication needs. The proposed interconnection network takes advantage of the wiring resources in the silicon interposer, composing a set of networks to support both inter- and intra-chiplet communications.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.