Geospatial Analysis and Object Detection for Military Vehicles Using Multimodal Large Language Models
Open Access DepositedRarePlanes, High-Resolution Ship Collection 2016 Multi-Scale (HRSC2016-MS), Dataset for Object Detection in Aerial Images version 1.5 (DOTA-v1.5), Fine-grAined object recognItion in high-Resolution one Million (FAIR1M), and Detection in Optical Remote Sensing Images (DIOR).The results support the potential of language-guided detection in data-limited environments, suggesting that the framework requires less retraining. In the supervised setting, the framework achieved performance within a 20% margin of the single-modal baselines on the RarePlanes dataset. However, it trailed behind the other baseline models on DOTA and HRSC2016-MS. In the few-shot setting, the framework scored 66 AP50 on only 10% of the FAIR1M training set, which is within 15 AP50 points of a fully supervised Mask R-CNN baseline. In the zero-shot setting, the framework performed well on DIOR with an AP50 of 44, while the DETR baseline reached 40 AP50. Overall, single-modal models perform better in supervised conditions, while the multimodal framework shows stronger performance in few-shot and zero-shot geospatial scenarios.
The National Geospatial-Intelligence Agency (NGA) heavily relies on classical single-modal computer vision models for geospatial intelligence (GEOINT), which leads to high training costs and domain-adaptation challenges. A multimodal detection framework that uses a multimodal large language model (MLLM) with a multimodal object detector can address these challenges. The framework combines GeoChat, an MLLM fine-tuned for grounded captioning, with the Modulated Detection Transformer-Geospatial (MDETR-G). MDETR-G incorporates a SatlasPretrain Shifted Window (Swin) Transformer backbone, deformable attention, and learnable temperatures for contrastive alignment and classification logits. The framework comparison evaluation includes the models You Only Look Once version 8 (YOLOv8), Faster Region-based Convolutional Neural Network (Faster R-CNN), Mask Region-based Convolutional Neural Network (Mask R-CNN), and Detection Transformer (DETR). The comparison is done across five benchmark datasets
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.