Electronic Thesis/Dissertation
 

Optimizing Accuracy and Efficiency Using Office of Scientific and Technical Information Records

Open Access Deposited

Enhancing Multi-Label Document Classification with AI

Repositories containing scientific and technical information are currently producing tens of thousands of new records per year, and thus, assigning subject terms to each record is becoming nearly impossible to do manually. In addition to the overwhelming volume of documents that are being produced daily by scientists around the world, many organizations have limited personnel and time to devote to maintaining accurate subject classification systems. One potential method for assisting with the challenge of providing timely and cost-effective subject classification for these records is through the application of machine learning techniques.

This paper evaluates the effectiveness of machine learning models in automating subject classification of records in the U.S. Department of Energy’s Office of Scientific and Technical Information (OSTI) repository using publicly accessible descriptive metadata.

Hierarchical Multi-Label Classification was the focus of this research since most documents in scientific and technical repositories will be assigned multiple subject terms at multiple levels within the same taxonomy. A reproducible machine learning pipeline was established and utilized for this project to predict subject terms from short descriptive metadata fields, i.e., Title and Abstract fields, as the primary input to the models. The classical baseline models used in this study were TF-IDF feature and logistic regression, while the transformer-based models were the state-of-the-art alternatives. All experiments used a fixed division of the available data into training, validation, and testing sets to allow for consistent and reproducible evaluations of the models. The decision threshold values for all models were determined only on the validation set to prevent data leakage when determining the threshold values and to avoid tuning the models on the test data.

Three supervised classification conditions were used in the study to evaluate the performance of the models: (1) the operational subject labels provided in the OSTI repository; (2) the externally curated subject labels from the PubMed database and the MeSH vocabulary; and (3) an OSTI evaluation protocol where the models were trained on the operational subject labels and tested on a subset of the test data that had been manually verified. The results of this study indicate that the use of metadata alone for predicting subject terms can provide significant support for Level-1 subject routing; however, predicting subject terms at the Level-2 of the taxonomy continues to be a challenging problem due to the large number of rare subject labels present in the OSTI repository, the extensive lexical variation that exists in the descriptive metadata fields, and the variability in the curation process of the subject labels among the different sources.

In summary, this study provides a reproducible framework for evaluating the effectiveness of machine learning models for performing hierarchical metadata classification using repository metadata. Additionally, this study demonstrates how both the organization and governance of taxonomies affect the performance of classification models. These results suggest that improvements in metadata governance and taxonomy design may be as important as advances in model architecture for improving automated subject classification in large scientific repositories.

Author Language Keyword Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Vasquez_gwu_0075A_17891.pdf Vasquez_gwu_0075A_17891.pdf 2026-06-24 Open Access