High-Performance, Energy-Efficient, and Flexible Accelerator Design for Deep Neural Networks, Graph Neural Networks, and Dynamic Graph Neural Networks
Open Access DepositedThe exponential growth of artificial intelligence (AI) applications has led to a dramatic surge in computational demands, far exceeding the capabilities of conventional CPUs and GPUs. This paradigm shift has necessitated the development of specialized hardware accelerators to deliver the required performance and energy efficiency for modern neural network workloads. Despite significant progress, designing efficient neural network accelerators presents several critical challenges. Communication overhead, both off-chip and on-chip, has become a dominant bottleneck due to its high energy cost and limited bandwidth scalability. Traditional accelerator designs often rely on centralized global buffers, which suffer from bandwidth congestion and poor scalability under increasing data movement demands. Furthermore, most existing architectures are fixed-function and lack the flexibility needed to accommodate the heterogeneous computation and communication patterns exhibited by diverse workloads such as deep neural networks (DNNs), graph neural networks (GNNs), and dynamic graph neural networks (DGNNs). To address these challenges, this dissertation proposes a cohesive set of accelerator architectures that share common design principles—reconfigurability, distributed tile-based organization, and communication-aware execution—while being tailored to the distinct characteristics of DNNs, GNNs, and DGNNs
A DGNN accelerator featuring a novel one-pass computation model that eliminates redundant inter-snapshot operations and significantly reduces communication costs. (4) DiTile-DGNN
A scalable accelerator for large-scale DGNN workloads based on a distributed-tile architecture, which supports redundancy-free parallelism, dynamic workload balancing, and reconfigurable interconnects for complex temporal graph processing. Collectively, the proposed designs deliver substantial improvements in performance, energy efficiency, and scalability across a wide range of AI applications. These solutions offer practical and generalizable architectural insights for next-generation neural network acceleration, particularly in domains such as social network analysis, recommendation systems, and traffic prediction.
A reconfigurable GNN accelerator that incorporates degree-aware mapping, workload-aware partitioning, and a unified NoC to efficiently support various GNN models with irregular sparsity patterns. (3) I-DGNN
(1) Venus
A versatile DNN accelerator that integrates distributed buffers and a flexible network-on-chip (NoC), enabling dynamic dataflow adaptation and alleviating the limitations of centralized memory architectures. (2) Aurora
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.