Gender Bias Mitigation in Image Captioning
Open Access DepositedRich, Fair, or Rich and Fair
Image captioning models describe visual content using natural language but perpetuate and amplify gender biases from training data. Current bias mitigation techniques are computationally expensive, limiting practical deployment. This Praxis introduces a modular framework operating at inference time at lower cost. Bias Mitigation Modules (BMMs) extract demographic features from images and fuse them with visual features through weighted addition, guiding models toward accurate demographic predictions across multiple architectures. Three variants were evaluated
Rich BMM (trained on Academic dataset, seven attributes), Fair BMM (trained on FairFace, three attributes), Rich + Fair BMM (combined both datasets using curriculum learning). The framework was tested on UpDown, Transformer, and ClipCap. Bias reduction varied by model architecture and BMM variant. All configurations achieved measurable reductions in gender misclassification rates and bias amplification while maintaining caption quality. The modular approach offers a practical, low-cost solution without model retraining.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.