Electronic Thesis/Dissertation
 

Why Architectures that Ignore Humans Fail to Satisfy Policy Objectives

Open Access Deposited

Centering Human Operators in Traditional & AI-Enabled Systems

Policies often fail to achieve their intended objectives due to architectural norms in engineering design that do not adequately include the role of human operators in the design, test, and evaluation process. Traditional complex engineered systems experience protracted design, development, and deployment cycles, meaning expected operating environments may change between system design and system fielding. This potential mismatch has created a need to design systems to be changeable, meaning they can respond to changes in the operating environment. Literature and practice have primarily focused on modularity, which allows a module to be easily swapped on or off a system to improve or change capabilities, as the key design mechanism for changeability. A core implicit assumption in many studies of changeability is that future changes can be anticipated, but often, unanticipated changes are needed during operations. Two complementary case studies, one of the C-130, where one platform was modified to perform many missions, and close air support in Desert Storm, where many platforms were modified to perform one mission, are used to examine how unanticipated changes occur in the field. These contrasting case studies reveal that excess, which is margin that is left after all design changes are locked in, and operational changes in which operators changed how their system was used can be more valuable in enabling unexpected changes than modularity. This is not to suggest that modularity has no use or that excess and operational changes are always the answer

classification (is the road clear or unclear) and routing (which alternative road should I take if the road is not clear). Through this modeling, I found that human-AI architectures alone created significant differences in performance and risk, even when the technical and environmental conditions did not change. Additionally, I found that some architectures were sensitive to changes in trust between humans and AI, suggesting a need to test explicitly for such variables, and that changes in the operating environment led to changes in the performance-risk profile of many architectures, suggesting that the expected operating environment and potential shifts in that environment affect which human-AI architecture is the most suitable. Additionally, the work on this testbed, which resulted from a Systems Engineering Research Center challenge sponsored by the United States Army DEVCOM Armaments Center, showed the variables needed to begin to model human-AI architecture for test and evaluation

oversight, delegation, and teaming. In oversight, the overseeing party can either directly control the system they are overseeing or can indirectly affect operations by providing warnings, information, or changing set points. Within direct oversight, approval or selection of a plan may be required for the system to function or oversight might be passive instead. In delegation, the ‘delegated to’ party can either directly perform tasks, take over control of the system, or indirectly support by collecting and providing information. The key differentiator is that the primary actor must actively delegate to the other party for them to act while in oversight, and the overseeing party chooses when to intervene. Both humans and AI can oversee each other and can delegate to one another. In teaming, both humans and AI must act, with task delegation being either fixed or dynamic. In addition to defining human-AI interaction on an architecture level and describing the various modalities of how humans and AI can be partnered together in a system, I also synthesize a wide range of literature on technology ethics, international law, and automation and labor studies to analyze what makes human-AI partnership difficult. Numerous factors like time pressure and over-trust in AI emerged from both theoretical research and through empirical cases. I map the challenges identified to the specific architectures where they are the most relevant. I also present a four step human-AI architecture process, consisting of selecting a level of analysis, allocating tasks and relationships at that level, assessing the potential pitfalls to successful human-AI partnership for the human-AI architectures selected, and assessing the system in its expected operating environment to understand whether issues like time pressure are likely to be present. While this proposed process is crucial for understanding how to begin to codify human-AI architecture as an explicit part of the system design process, it does not address how to test and evaluate the system more formally. Most work in AI evaluation is focused on the technical system, with emphasis on data quality, mechanistic interpretability, and benchmark testing, but a growing number of studies seek to evaluate how humans and AI perform together. Most human-AI system evaluation studies are focused on AI advice being presented to a human decision maker in different ways. Those that do evaluate distinctly different types of human-AI systems typically focus only on ‘in-the-loop’ participatory control vs ‘on-the-loop’ supervisory control. This emphasis on just a few types of human-AI interaction has meant that the literature is unclear on the tradeoffs of different architectures and how to test them. To address this gap, I created a simulation testbed to analyze how different human-AI architectures performed in the same environment. The simulation setting of the testbed is the traversal of an unmanned ground vehicle (UGV) across a field with improvised explosive devices (IEDs) or mines. If an IED is present on a road, it extends traversal time as the mine must be removed. We do not allow backtracking and assume that if a mine or IED is encountered, it must be removed. This problem resembles a Canadian Traveler Problem since there is some probability that a road the UGV hopes to traverse is blocked by an IED and that the IED can only be discovered once the UGV begins to traverse the road. Since the IED can be dismantled and could be predicted in advance, this is a case of Canadian Traveler Problem with remote sensing and neutralizations. Unlike a traditional Canadian Traveler Problem, the status of the road is not fully known until it is traversed. The key behavior I hoped to understand through this study is the tradeoffs between performance, measured as mean traversal time, and risk, measured as the number of IEDs encountered, for each architecture. I modeled various architectures for the two key tasks in this problem

rather, I argue that these two mechanisms have not been adequately addressed by the systems engineering design literature despite the ample evidence of their importance in practice. To be able to consider operator change effectively, changeability studies need to expand the system boundary beyond just the technical artifact to include the larger socio-technical system they are deployed in. The case studies also highlight how changeability mechanisms are often used in conjunction with each other as part of a larger changeability strategy that is not often discussed in the literature. In AI-enabled systems, policy often mandates human control over and collaboration with AI to ensure that systems are safe and trustworthy. The core ideas behind these policies are that human control can prevent AI from making mistakes and that human collaboration with AI can leverage the unique strengths of both to create superior human-AI systems. Literature has primarily focused on just a few types of human-AI interaction, leaving a great deal of uncertainty around how humans and AI should be partnered together to achieve these goals. To address this uncertainty, I synthesize a wide range of literature spanning human-computer interaction, human-robot interaction, and human factors studies along with an examination of many common AI systems in use today to create a taxonomy of human-AI interaction. I define my level of analysis on the architectural level, and I define human-AI architecture as the task allocation and relationship between humans and AI. This human-AI architecture framing is not dependent on the system, task, or function being defined as I define three key modalities of human-AI architecture

a schematic workflow of how the system functions and which tasks could be allocated between humans and AI, estimated performance for humans and AI in the tasks they might be expected to perform, parameters of expected operating environments, and a shared mental model of how humans and AI might interact. Such testbeds are intended to be used as trade space exploration tools, enabling system designers to consider which architectures should warrant additional testing with humans and to understand which variables seem to drive performance. Expanding the system boundary to include human operators and recognizing the conditions that make them successful can help designers architect systems in ways that may actually satisfy the high level policy objectives designers are supposed to meet.

Author Language Keyword Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Singh_gwu_0075A_17769.pdf File 2026-06-24 Embargo