Will CLIP Zero-Shot, Using Text to Predict Zero-Shot Image Classification.
Open Access DepositedCLIP aligns visual and textual data into a shared embedding space, enabling zero-shot image classification. However, its performance is inconsistent across different classes, with some concepts underrepresented. This thesis explores whether the intrinsic properties of CLIP's text encoder can predict and enhance zero-shot classification accuracy based solely on text embeddings. We develop methods to analyze CLIP's text embedding space to predict zero-shot performance. we demonstrate that text embeddings alone can effectively predict and improve zero-shot accuracy. Additionally, We show how these characteristics can be applied directly to generate more effective text prompts, boosting zero-shot classification performance on multiple datasets and models.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.