Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Identifying disease-associated genes through experimental approaches remains costly and time-consuming. Computational approaches based on biological networks provide an effective alternative for disease gene prioritization and disease module identification. However, existing machine learning methods often require reliable negative examples, whereas in disease gene prioritization only a limited number of confirmed disease genes are available and the remaining genes cannot be confidently considered negative samples. This positive--unlabeled (PU) learning setting introduces substantial label uncertainty and remains a major challenge for computational disease gene discovery. In this study, we propose XG-PUL, a machine learning framework designed for disease gene prioritization under a PU learning scenario. XG-PUL employs a two-stage strategy in which label-independent node representations are first learned from the protein--protein interaction (PPI) network using Node2Vec. These representations are then integrated with complementary graph-topological features to capture both latent network patterns and explicit connectivity characteristics of genes. Subsequently, a bagging-based PU classifier is applied to aggregate multiple classifiers trained on different bootstrap samples of unlabeled genes, improving robustness against uncertainty in the unlabeled data. The proposed framework was evaluated on ten complex disease datasets and compared with representative network-based methods and graph learning approaches. XG-PUL demonstrated improved ranking performance, achieving higher AUPR and F1-score values across the evaluated diseases and ranking thresholds, while maintaining competitive AUC-ROC performance. Functional enrichment analysis and literature-based validation further confirmed the biological relevance of prioritized candidates, showing that XG-PUL can recover genes involved in disease-related pathways and molecular mechanisms beyond known disease annotations. These results demonstrate that integrating label-independent network representation learning with PU learning provides a robust framework for disease gene prioritization in incomplete biological knowledge scenarios.</p>

Show More

Keywords

disease learning genes gene prioritization

Related Articles

PORE

About

Connect