Leveraging Machine Learning and Knowledge Graphs to Discover Novel Relationships in Biomarker Data
Open Access DepositedDownloadable Content
Biomarkers are essential to modern biomedical research and clinical practice, serving as measurable indicators of biological states, disease processes, and therapeutic responses. Despite their importance, biomarker knowledge remains fragmented across heterogeneous databases, inconsistently defined, and difficult to integrate computationally. This fragmentation limits reproducibility, interoperability, and large-scale analytical exploration.This dissertation presents a biomarker-centric computational framework for harmonizing, integrating, and analyzing biomarker knowledge at scale. Grounded in the FDA–NIH BEST (Biomarkers, EndpointS, and other Tools) definition, a formal data model was developed to represent biomarkers as measurable changes in biological entities associated with conditions, exposures, or interventions. The model distinguishes core definitional elements from contextual metadata and incorporates controlled vocabularies, ontology alignment, and identifier normalization to ensure semantic consistency and interoperability. This work is culminated and presented as BiomarkerKB, a functional biomarker knowledgebase. To support large-scale integration, a multi-layer quality assessment framework was implemented. Rule-based validation enforced structural constraints and generated weak supervision labels (Gold, Silver, Noisy), which were subsequently used to train interpretable machine learning models that assign continuous quality scores to biomarker records. This approach separates structural validation, probabilistic confidence estimation, and governance thresholds, enabling scalable quality control while preserving transparency. Harmonized biomarker data were transformed into the BiomarkerKB Knowledge Graph, implemented in Neo4j and aligned with the Biolink data model for interoperability with the Common Fund Data Ecosystem Unified Biomedical Knowledge Graph. Graph-based analyses demonstrate recovery of known biomarker–disease–drug relationships, identification of cross-disease molecular overlap, and systematic exploration of therapeutic connectivity. These results validate the knowledge graph as an analytical substrate for biomarker-centric network reasoning and hypothesis generation. Together, this work establishes a scalable infrastructure that integrates conceptual rigor, standardized data representation, embedded quality assessment, and graph-based analytics. The resulting framework transforms fragmented biomarker entities along with their annotations into a Findable, Accessible, Interoperable, and Reusable (FAIR)-compliant and computationally actionable knowledge resource, supporting future biomarker discovery, translational research, and precision medicine applications.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.