Electronic Thesis/Dissertation
 

Classification of Alternate Reads in Expressed Tumor Variants (CLARET)

Open Access Deposited

A Deep Learning Pipeline for Four-Class Variant Classification from Tumor-Only RNA Sequencing

Somatic mutation detection in RNA sequencing data is a fundamentally harder problem than in DNA sequencing. Expressed variants in cancer transcriptomes arise from at least four distinct processes

somatic mutations, inherited germline polymorphisms, post-transcriptional RNA editing, and technical artifacts from alignment and sequencing errors. The read-level signatures of these classes overlap extensively, and the most common mitigation, matched normal sequencing, is unavailable for the large majority of published RNA-seq datasets and for most retrospective clinical cohorts. Existing RNA-seq variant callers either require a matched normal, perform binary variant detection without class resolution, or apply hard database filters that collapse in the biologically ambiguous cases that matter most for cancer research.This thesis presents CLARET (Classification of Alternate Reads in Expressed Tumor variants), a deep learning pipeline that performs four-class variant classification (somatic, germline, RNA-edited, artifact) directly from tumor-only bulk RNA sequencing. Each candidate variant is encoded as a 201 base pair by 17 channel tensor capturing local sequence context, allele frequency structure, read quality, and genomic annotations. A channel-grouped convolutional neural network (ChannelGroupNet) processes these tensors through separate streams for sequence, allele frequency, quality, and biological channels before fusion, and is trained under a three-phase curriculum that progresses from balanced sampling to class-imbalanced fine-tuning to hard-negative mining. A concordance cascade then integrates the CNN ensemble with an XGBoost meta-classifier that consumes 259-dimensional CNN embeddings, six biological features, and 38 tensor-derived spatial statistics, routing variants through a gnomAD population filter, an RNA-editing gate, and a final somatic-versus-non-somatic decision stage. CLARET was trained on 179 CCLE cancer cell line samples and evaluated on 65 independent DepMap test samples spanning 25 tissue types unseen during training. Under leave-one-sample-out cross-validation, the cascade achieved a somatic F1 of 0.703, macro F1 of 0.870, overall accuracy of 94.1%, and somatic AUPRC of 0.763, a 2.1-fold improvement over the CNN ensemble alone. Validation against DepMap whole-genome sequencing recovered 84.8% of RNA-visible somatic mutations. Analysis of the exonic confusion zone (exonic variants with VAF 0.4 to 0.6 and gnomAD absent) showed that read-level classification alone is insufficient in this regime, as missed somatic calls were largely indistinguishable from heterozygous germline at the read level, leaving population-database absence as the decisive discriminating signal. Zero-shot transfer to single-cell Oxford Nanopore long-read data produced 97.3% germline concordance with matched short-read labels without any retraining, and 99.1% of NOVEL RNA-editing predictions were A-to-G or T-to-C substitutions consistent with the ADAR mutational signature. Together these results demonstrate that learned four-class classification is a tractable framing for RNA-seq variant analysis, that population-database evidence is most effective as one feature among many rather than a hard pre-filter, and that CLARET is directly applicable to the large and growing body of tumor-only RNA sequencing data for which matched DNA is not available.

Author Language Keyword Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Sajjad_gwu_0075M_17938.pdf Sajjad_gwu_0075M_17938.pdf 2026-06-24 Open Access