Electronic Thesis/Dissertation
 

Applying Machine Learning and Neural Language Models to Math Linguistics Tasks

Open Access Deposited

The mathematical language presents unique challenges for math-related tasks, including information retrieval (IR) systems and indexing math content in online libraries, such as creating an archive for math content. Traditional IR systems face challenges with the com- plexities of mathematical notation and terminology. Therefore, integrating math language processing (MLP) in IR systems can enhance the precision and recall of mathematical infor- mation retrieval. However, math language has different characteristics than other natural languages, requiring adaptive and special techniques to handle math language content.This dissertation explores the possibilities and challenges of adopting existing NLP techniques for processing mathematical language. We investigate methods for representing mathematical features, such as Bag of Words (BOW) and word embeddings. We examine various traditional machine learning approaches, including SVM, Decision Tree, Random Forest, and deep learning models like BiLSTM. Our findings emphasize the importance of mapping mathematical tokens to their corresponding meanings to enhance math language knowledge representation and extraction. To address these challenges, we investigate the possible implementation of Named Entity Recognition (NER) and chunking/segmentation algorithms in mathematical language processing. We develop specialized models for math formula chunking and Named-Function Entity Recognition (NFER). These models break down a mathematical formula into syn- tactically correlated parts or segments. Although both models are applied using a subset of well-known math entities, the positive results demonstrate the potential for adapting the models to math language. Our other contribution in this dissertation follows the recent trend of utilizing pre-trained language models (such as Roberta, DistilBERT, and BERT) in natural language processing. The aim is to assess the adaptability of these general models to tasks involving math language. We apply these pre-trained models to our math chunking and NFRM tasks, as well as to other math-related tasks. Our evaluation of these models highlights the need to develop a model that is trained specifically on math language or takes into account math content during training. To this end, we develop a mixed-domain model using a large math corpus and present the MBERT model. We evaluate the model on the same set of tasks and conduct a performance comparison. Although the MBERT model is not a final version, meaning it is not fully trained due to resource limitations, the results suggest promising outcomes with a fully pre-trained model. Our research results indicate that the models presented in this work have the potential to serve as a foundational framework for math-related tasks, particularly those where annotating math tokens is highly beneficial, especially in challenging areas such ad IR and STEM applications, where the complexities of mathematical notation and terminology require adaptive and specialized techniques.

Author Language Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Alshamari_gwu_0075A_17038.pdf File 2025-04-09 Embargo