Synthetic Aperture Radar (SAR) Object Detection via Spectral Diffusion
Open Access DepositedDownloadable Content
Satellite-based SAR is an integral component of geospatial imaging, but recent computer vision advancements have been less successful with SAR imagery. Geospatial imagery is a key information source for natural disaster response, and wildfires represent a challenging problem that SAR is uniquely well-suited to address. SAR imagery is fundamentally different from optical imagery, and existing object detection (OD) technologies are less effective with geospatial imagery. Rather than fine-tuning OD algorithms to SAR imagery, this praxis uses diffusion models to transform SAR imagery into optical representations. A Visual Language Model (VLM) captions the SAR image. The image and caption are input to a Low Range Adaptation (LoRA) Stable Diffusion model to create a pseudo-optical output. Pseudo-optical output assessed by standard OD measures overall image fidelity and compatibility with existing OD tools. SAR-Aircraft 1.0 and Dataset for Object Detection in Aerial Images (DOTA) were selected for model training. Airports and aircraft were used since they provided generally uniform terrain with movable objects on a scale equivalent to the structures of interest during disaster response. To adapt the datasets to diffusion training, specific tools were developed to further clean and label the data. Models were trained and tested against multiple caption strategies to evaluate the impact of guide text on image output. The best performing VLM achieved an F1-score of 0.83, and precision of 0.79, the highest recall achieved was 0.89. The OD performance output was more nuanced. The highest F1-score against all aircraft classes was 0.51, with improved performance at higher detection confidence levels. On a per-class basis, models achieved as high as 1.0 recall at the expense of degraded performance against other classes. The VLM diffusion model demonstrated high performance in both object detection and image captioning. The stable diffusion pseudo-optical image generator created output compatible with existing optical image detectors, with reduced detection accuracy compared to the VLM. The workflow designed for this praxis demonstrates the potential for enhanced SAR analysis with diffusion models trained on a carefully curated image collection. Combined with the evidence of fine detail transfer, these results support further training against images collected during wildfire events.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.