Electronic Thesis/Dissertation
 

Optimized Routing and Scheduling in Home Health Care Using Deep Reinforcement Learning

Open Access Deposited

The growing elderly population in the U.S. has increased demand for home care services. Even though the primary objective of home health agencies is to provide patient care, home health care regulations, labor laws, contracts, and clinical standards require them to follow additional guidelines that may compromise their ability to do so. These hard constraints include matching clinician skills against patient care needs, scheduling clinicians during designated time windows, observing work time limits, and so on. Additionally, home health agencies must balance other secondary objectives, such as travel efficiency, equitable work distribution, patient preferences, and continuity of care. Classical scheduling methods typically handle these soft objectives sequentially. They first maximize the primary objective, such as visit fulfillment, before addressing secondary objectives. Similarly, metaheuristic scheduling algorithms often optimize a single primary objective rather than balancing conflicting objectives at the same time. To address the limitations of single objective optimization techniques shown in classical, heuristics and metaheuristic methods, we use deep reinforcement learning(DRL) to train an agent to account for competing secondary objectives. We start by formulating the problem as a Markov decision process. We define a reward function as an aggregator function composed of weighted signals of each secondary objective we wish to prioritize. We then train the agent using Proximal Policy Optimization algorithm to maximize the reward. The agent interacts with a simulation environment using samples of varying instance sizes (small

80 patients, 225 visits, 10 clinicians

medium

large

150 patients, 430 visits, 20 clinicians). Action masking enforces hard constraints so that clinicians are never assigned outside of their work hours, visits occur within specified time windows, and skill requirementsare satisfied. We configure the weights of the reward function to reflect priorities typically seen in home health. Visit fulfillment receives a higher weight than secondary objectives to reflect its position as the primary objective in home health care. To account for different home health operational priorities, we also implement three aggregator variants, Weighted Sum, Chebyshev, and Policy-Aware. We evaluated the DRL model against exact optimization method, metaheuristic method, and greedy approaches to address three questions. Does the DRL model achieve competitive Visit Fulfillment rates on small and medium problem instances? Does the DRL model explore the Pareto frontiers more effectively along visit fulfillment, idle rate, and utilization objectives? Does the DRL model perform comparably in the General Composite Score composed of ten secondary metrics? The results reveal that performance varies by instance size. On small size instances, the DRL model scored 96% on Visit Fulfillment compared to 81% for exact optimization methods. On medium size instances, the DRL model achieved a Visit Fulfillment rate of of 79-80% , showing no significant difference from the benchmark methods. Across small and medium scale, the DRL model achieved 6-140 times greater Pareto hypervolume than all benchmark approaches. These results show that DRL model has greater ability to explore trade-offs in small and medium instances. The General Composite Score shows 23% improvement over exact optimization method at small scale and comparative at medium scale. At large scale, the DRL model scored 53% on Visit Fulfillment, compared to 90% with exact optimization method. This indicates that either substantially more training, or architectural changes would be necessary for the DRL model to be comparable at larger instances. This praxis demonstrates that deep reinforcement learning is a viable routing and scheduling solution for small and medium size home health care agencies. The novelty of this work lies in dynamic action masking to enforce constraints, a decomposed reward function that uses weights to assign priority to secondary objectives and a regularization strategy to provide generalization over unseen instances.

40 patients, 120 visits, 5 clinicians

Author Language Keyword Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Omollo_gwu_0075A_17623.pdf Omollo_gwu_0075A_17623.pdf 2025-12-12 Open Access