Electronic Thesis/Dissertation
 

Will CLIP Zero-Shot, Using Text to Predict Zero-Shot Image Classification.

Open Access Deposited

CLIP aligns visual and textual data into a shared embedding space, enabling zero-shot image classification. However, its performance is inconsistent across different classes, with some concepts underrepresented. This thesis explores whether the intrinsic properties of CLIP's text encoder can predict and enhance zero-shot classification accuracy based solely on text embeddings. We develop methods to analyze CLIP's text embedding space to predict zero-shot performance. we demonstrate that text embeddings alone can effectively predict and improve zero-shot accuracy. Additionally, We show how these characteristics can be applied directly to generate more effective text prompts, boosting zero-shot classification performance on multiple datasets and models.

Author Language Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Wu_gwu_0075M_17151.pdf Wu_gwu_0075M_17151.pdf 2025-04-09 Open Access