A Deep Learning Approach for Classifying Pathology of Cancer Cell Lines using Vector Representations of Gene Expression and Attribute Data
Open AccessMachine Learning and Big Data are two burning topics which have not avoided the contemporary surge in accession of large-scale biomedical data. A variety of technologies and research platforms, such as the efficient employment of next-generation sequencing technologies, ChIP sequencing, or mass spectrometry pipelines, have allowed for the gathering of biochemical information capable of forming genetic, epigenetic, or proteomic answers to crucial research questions. The challenge, however, is with the continued and effective use of these technologies there will come the need to better make sense of and manage the data that is produced. The concept of machine learning aims to address this challenge, and its applications in biomedicine will no doubt be formulated as these advanced computational methods are refined by the wave of data scientists moving into related fields such as Bioinformatics or Genomics. In this paper, a use-case of deep learning - a subset of machine learning - is explored onto cancer cell line gene expression and other genetic data to perform a simple classification task: is the sample ‘primary’ or ‘metastatic’ in pathology? The genetic data is transformed into a word-embedding using the popular word2vec set of algorithms which is subsequently used to convert sample information into images recognized by a Convolutional Neural Network (CNN). A classification accuracy of 98% is reported and a forward-thinking discussion held on the importance and advantages that deep learning contains when compared to traditional machine learning and statistical applications in Bioinformatics.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.