Electronic Thesis/Dissertation
 

Predicting Long-Term Vehicle Trajectories Combining Contrastive Language - Image Pretraining (CLIP) Into a Diffusion Model

Open Access Deposited

CLIP, Diffusion, Multi-modal, Vision Models, Transformers, Trajectory predictions

1) How does the proposed model perform against other state of the art models (SOTA)? 2) How do adjustments in CLIP architecture and training impact the final average displacement error (ADE)? The research showed that the model performed within one standard deviation of similar models and that the more robust the CLIP labeling, the greater impact on the performance of the model. Keywords

From January to June 2025, 17,140 vehicular crash-related deaths occurred within theUnited States. Traffic growth within the nation is expected to increase by 1.5% from 2025 to 2027, with an anticipated 26% growth in the U.S. population leading to a 72% increase in passenger miles per year. While autonomous driving vehicles and intelligent driving systems reduce congestion and accidents, many car manufacturers cannot afford the cost of entry due to high-value systems such as LiDAR and on-board cameras. This limited presence of autonomous vehicles raises the risk of vehicular accidents. Moreover, most autonomous vehicles only predict 3–5 seconds out to mitigate immediate accident avoidance, leading to growing traffic congestion, increased accidents on roadways, and impacts on other fields such as urban planning and emergency services. The Georgia Department of Vehicle Services (2025) stated that most good drivers make driving decisions 12–15 seconds ahead of action. However, to date, research has been solely focused on near-term, 3–5-second predictions for vehicle trajectories, with almost no research extending to long-term predictions of 6–10 seconds. In this paper, by combining contrastive language image pretraining (CLIP) with diffusion models, which typically analyze multiple possible futures, we holistically examine long-term vehicle trajectories and introduce human-like intuition into multi-modal models to discern if visual and semantic information improve trajectory predictions over time. This praxis examens two research questions

Author Language Keyword Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of SEXTON_gwu_0075A_17649.pdf SEXTON_gwu_0075A_17649.pdf 2025-12-12 Open Access