Context Matters: Employing Word Embeddings to Improve Text Classifier Performance on Peer-Reviewed Academic Journal Abstracts, a Test Case
Open AccessBag-of-words is a commonly used text representation method for many text classification applications. However, bag-of-words representation fails to consider the context of the text because it only examines text documents based on the presence of individual words and explores relationships between texts with similar word choices (Bengfort, 2018). Understanding the context of a text is important to classify closely related texts correctly. In recent years, research has emerged to classify large corpora using a subset of the text specifically abstracts and metadata. However, this research almost exclusively focuses on medical and biomedical data sets derived from MEDLINE including the 2014 BioASQ Challenge data set for biomedical semantic indexing. This research aimed to show the benefit, in terms of increased text classification performance, of employing semantic analysis in data preprocessing to classify peer-reviewed journal abstracts by subject. The results of this research showed that semantic analysis preprocessing did not significantly improve classification performance for the research dataset. However, text classification is a viable option to automate some requirements elicitation activities and reduce the amount of manual intervention required to review and distribute requests for new Department of Defense Information Technology projects.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.