Electronic Thesis/Dissertation
 

Improving 3D Semantic Segmentation for Autonomous Vehicles Using Multimodal Vision Transformers

Open Access Deposited

LiDAR-based semantic segmentation is an important part of the perception system for autonomous vehicles (AV). It enables AVs to have better understanding of their surrounding environments. Traditional segmentation methods that use convolutional neural networks (CNNs) achieve strong performance on 2D images but limited results on LiDAR point clouds. Since the convolutions in CNNs operate locally, they struggle to integrate information across distant regions of a scene. This shortcoming limits their effectiveness to model long-range dependencies. Moreover, LiDAR data is inherently sparse, in which small or distant objects have only few points to represent them. These challenges make it difficult to classify point cloud using LiDAR information alone. This praxis develops a multimodal framework based on Vision Transformers (ViTs) to address the above shortcomings. The proposed method fuses LiDAR’s point cloud and camera’s RGB image before feeding them into a transformer-based architecture. The proposed method also integrates a custom patch-embedding and a scan-unfolding projection to optimize data representation for ViT’s input. The experimental results on the SemanticKITTI and Waymo Open datasets show that our multimodal ViT achieves a mean Intersection over Union (mIoU) of 0.71 on SemanticKITTI, representing a 15 percentage points improvement over CNN-based methods (mIoU=0.56). The improvements show that our proposed model can enhance the semantic segmentation accuracy and robustness for AV’s perception system.

Author Language Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Dinh_gwu_0075A_17888.pdf Dinh_gwu_0075A_17888.pdf 2026-06-24 Open Access