Electronic Thesis/Dissertation
 

Context Matters: Employing Word Embeddings to Improve Text Classifier Performance on Peer-Reviewed Academic Journal Abstracts, a Test Case

Open Access

Bag-of-words is a commonly used text representation method for many text classification applications. However, bag-of-words representation fails to consider the context of the text because it only examines text documents based on the presence of individual words and explores relationships between texts with similar word choices (Bengfort, 2018). Understanding the context of a text is important to classify closely related texts correctly. In recent years, research has emerged to classify large corpora using a subset of the text specifically abstracts and metadata. However, this research almost exclusively focuses on medical and biomedical data sets derived from MEDLINE including the 2014 BioASQ Challenge data set for biomedical semantic indexing. This research aimed to show the benefit, in terms of increased text classification performance, of employing semantic analysis in data preprocessing to classify peer-reviewed journal abstracts by subject. The results of this research showed that semantic analysis preprocessing did not significantly improve classification performance for the research dataset. However, text classification is a viable option to automate some requirements elicitation activities and reduce the amount of manual intervention required to review and distribute requests for new Department of Defense Information Technology projects.

Author Language Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Gaither_gwu_0075A_15712.pdf Gaither_gwu_0075A_15712.pdf 2022-03-06 Open Access