Exploration, Collaboration, and Applications in Multi-Agent Reinforcement Learning
Open Access DepositedMany decision-making problems in the real world naturally involve the participation of multiple decision-making agents (e.g., vehicles in a fleet and routers in a network) and thus need to be modeled as multi-agent problems. As a dominant approach, multi-agent reinforcement learning (MARL) continues to make headway in increasingly complex sequential decision-making problems. However, we also start to encounter `the curse of multi-agent' – i.e., the exponential growth of state/action spaces caused by an increasing number of agents significantly hinders the development of collaborative exploration strategies as well as the learning of joint decision-making policies. Substantial gaps remain in developing a solid foundation for multi-agent exploration and collaboration. To this end, this dissertation aims to break `the curse of multi-agent' by developing a holistic framework to enable highly scalable MARL algorithms with collaborative exploration and explainable communication. Furthermore, trained with millions or billions of parameters, deep reinforcement learning has achieved tremendous success in many sophisticated sequential decision-making problems. However, some of the key roadblocks preventing widespread adoption of such reinforcement learning algorithms in the real-world applications are the lack of fairness guarantees and explainability, as well as the need for novel hardware-software implementations that are orders of magnitude faster and more energy efficient. Existing notions/tools of fairness do not generalize well to settings where multiple intelligent agents - acting competitively, cooperatively, or neutrally - dynamically interact with an environment that is subjected to the actions of all agents. Further, existing work on Explainable AI (XAI) primarily focuses on the explainability of deep neural networks with very few exceptions. There remain several open problems in explainable reinforcement learning: `how to relate long-term objectives to an agent’s near-term policy and actions, what different behaviors exist in exploration and learning, and how to interpret agents’ influence on each other’s decision-making in multi-agent settings. This dissertation also aims to translate the innovations into real-world applications and investigate a wide range of emerging applications to amplify the impact of our research on data-driven decision-making. The interaction between theory and practice is mutual. As this dissertation motivates the development of new data-driven solutions in practice, %these exciting applications in turn provide valuable feedback (e.g., validating/falsifying key assumptions and imposing new constraints) to further improve our learning models and algorithms. it investigated the use of multi-agent decision-making in Real-time intelligent Control (RiC) of 5G O-RAN network slicing resource management. This dissertation delves into the intricate dynamics of exploration and collaboration within multi-agent decision-making problems, areas that are pivotal due to their inherent need for rapid processing capabilities and adept handling of distributional rewards. Central to our investigation are two key elements: the development of sophisticated exploration strategies and the enhancement of collaborative mechanisms in MARL.In the realm of exploration, we focus on the innovation of scaling techniques for option discovery, aiming to significantly improve the efficiency and effectiveness of decision-making processes in complex environments. For collaboration, we propose the creation of a discrete communication algorithm with a performance guarantee tailored for MARL, designed to facilitate more effective and explainable coordination among agents. The practical applications of these advancements are diverse and impactful. Firstly, we address the challenge of fairness in MARL, introducing a novel approach to reward engineering that seeks to balance and optimize the distribution of rewards among agents, thereby promoting equity within the system. Additionally, we apply our findings to the field of 5G network slicing, employing a Distributional MARL framework to effectively tackle the complexities of resource allocation decisions, a critical aspect in the optimization of 5G networks. In the domain of cybersecurity, we explore the integration of a software-hardware co-design with an explainable AI algorithm, targeting the acceleration of detection speeds in network intrusion scenarios. This approach not only enhances security measures but also brings clarity and understanding to the decision-making process.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.