Author name disambiguation using graph contrastive learning
Open Access DepositedAuthor name ambiguity remains a constant challenging problem in academic databases, often resulting in significant misattribution and inflated metrics that affect assessments of research impact. This study proposed a novel graph-based author name disambiguation (AND) model that employs contrastive learning to enhance the differentiation between authors with similar or identical names. By incorporating co-authorship patterns and academic site associations, the synergistic graph-based author disambiguation (SGAD) model refines semantic and relational embeddings within the graph structure. The result show that co-authorship patterns emerge as the most influential feature for accurate disambiguation. This approach constructs a heterogeneous graph where nodes represent individual author instances from publications, linked by their collaborative and institutional relationships. A graph neural network (GNN) is trained to generate context-aware embeddings for each node in the graph. These embeddings are optimized through a contrastive loss function, which systematically maximizes the similarity for publication pairs belonging to the same true author while minimizing it for pairs from different authors. When evaluate the model on the large AMiner benchmark dataset, the proposed model reaches a F1 score of 0.8070. This represents a clear improvement over both traditional approaches and other recent deep learning models. These results show the value of combining graph-based relational learning with contrastive training and point to a framework that is both robust and scalable for assigning scholarly credit accurately.
- All rights reserved
Notice to Authors
If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.