Predicting Long-Term Vehicle Trajectories Combining Contrastive Language - Image Pretraining (CLIP) Into a Diffusion Model
Open Access DepositedCLIP, Diffusion, Multi-modal, Vision Models, Transformers, Trajectory predictions
1) How does the proposed model perform against other state of the art models (SOTA)? 2) How do adjustments in CLIP architecture and training impact the final average displacement error (ADE)? The research showed that the model performed within one standard deviation of similar models and that the more robust the CLIP labeling, the greater impact on the performance of the model. Keywords
From January to June 2025, 17,140 vehicular crash-related deaths occurred within theUnited States. Traffic growth within the nation is expected to increase by 1.5% from 2025 to 2027, with an anticipated 26% growth in the U.S. population leading to a 72% increase in passenger miles per year. While autonomous driving vehicles and intelligent driving systems reduce congestion and accidents, many car manufacturers cannot afford the cost of entry due to high-value systems such as LiDAR and on-board cameras. This limited presence of autonomous vehicles raises the risk of vehicular accidents. Moreover, most autonomous vehicles only predict 3–5 seconds out to mitigate immediate accident avoidance, leading to growing traffic congestion, increased accidents on roadways, and impacts on other fields such as urban planning and emergency services. The Georgia Department of Vehicle Services (2025) stated that most good drivers make driving decisions 12–15 seconds ahead of action. However, to date, research has been solely focused on near-term, 3–5-second predictions for vehicle trajectories, with almost no research extending to long-term predictions of 6–10 seconds. In this paper, by combining contrastive language image pretraining (CLIP) with diffusion models, which typically analyze multiple possible futures, we holistically examine long-term vehicle trajectories and introduce human-like intuition into multi-modal models to discern if visual and semantic information improve trajectory predictions over time. This praxis examens two research questions
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.