Multi-Scaled MAMBA Transformer Mixer For Medical Image Classification
Open Access DepositedMS-MAMBATM
In AI-driven disease diagnosis, as well as pathology studies, medical image classification in two dimensions remains a task of utmost importance. The overarching objective in this community revolves around designing a lightweight yet extremely efficient, portable classification backbone. Latest contributions point towards a novel application of Mamba, a state-space model, for a specific task of sequence prediction. Its visual analogue, Vision Mamba, or ViM, has been identified as having a potential application replacing both classic CNN models, as well as Vision Transformer models. In this application, MS-MambaTM, a novel 2D image classification backbone with a multi-scale architecture, including a computationally efficient specialized sparse self-attention mechanism with Vision Mamba, has emerged, outperforming existing CNN, VisionTransformer, as well as Vision Mamba models on all scoring metrics for experiments carried out on RetinaMNIST, as well as ChestMNIST datasets. Specifically, On RetinaMNIST, MS-MambaTM achieves an ROC-AUC of 0.80, which corresponds to 8.1% relative improvement over CNN (0.74), 3.9% over ViT (0.77), and 6.7% over ViM (0.75). On ChestMNIST, MS-MambaTM achieves a Macro AUC of 0.818, compared to 0.773 for CNN and 0.779 for ViT and ViM, resulting in up to 5.8% relative gain. This research validates the effectiveness of multiscale attention-enhanced Mamba design in 2D medical image classification and paves the way for future applications in lightweight and clinically meaningful diagnostic systems.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.