Deprecated: Function curl_close() is deprecated since 8.5, as it has no effect since PHP 8.0 in /home/u483256323/domains/poorvam.com/public_html/subdomains/pore/includes/api.php on line 184
Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Benchmarks of gene regulatory network (GRN) inference score predicted edges against curated reference databases that are known to be radically incomplete. Recent work has quantified the consequence as an upper bound: incomplete labels cap the area under the precision–recall curve (AUPRC) that any method can attain. A ceiling tells us that reported scores are too low; it does not tell us by how much, and it does not tell us whether the incompleteness distorts the ranking of methods. We take the next step and construct a de-biased point estimate. Treating STRINGphysical, BioGRID multi-validated and TRRUST as three partially overlapping “captures” of one latent interactome, we estimate per-database completeness π by capture–recapture, first with Lincoln–Petersen and then with log-linear models that carry explicit dependence terms, because these databases share curation practice and literature bias and are emphatically not independent. The estimated π then feeds an Elkan–Noto-style positive–unlabelled correction that yields bias-corrected AUPRC with bootstrap intervals for eight GRN inference methods, including three scGPT-derived variants, two co-expression baselines, GENIE3, GRNBoost2 and a random control. Two results matter. First, the absolute level of π is only weakly identified: across dependence models that all fit the observed overlaps, the estimated size of the latent interactome moves over an order of magnitude, from 2,579 to 46,864 edges in a 359,700-pair candidate universe, so completeness estimates range from 5.4% to 54.4%. Second, and more useful, this ambiguity is largely irrelevant to the question benchmarks actually ask. We prove that any constant-propensity correction — which is what a ceiling or a global π delivers — rescales every method’s AUPRC by the same factor and therefore cannot change the ranking, and that a propensity-weighted correction depends on the estimated completeness only through its shape across strata, not its level. The shape is far better identified than the level: three different dependence models that disagree about π by a factor of three agree about the relative completeness profile across literature-attention strata to within a few percent. Applying the propensity-weighted correction rescales AUPRC upward by an order of magnitude but leaves the ordering of the eight methods exactly unchanged in the primary tissue, under four different propensity models, and moves only one adjacent, statistically indistinguishable pair in the replication tissue. A controlled simulation in which the truth is known shows that the machinery does have the power to repair a mis-ordering when the labelling mechanism is strongly covariate-dependent; the real reference databases are not biased strongly enough, relative to how similar current GRN methods are to one another, for that to happen. We conclude that the published ordering of GRN methods is more robust to reference incompleteness than the size of the incompleteness suggests, while the absolute performance numbers are not interpretable at all.</p>

Show More

Keywords

methods auprc three completeness models

Related Articles


Deprecated: Function curl_close() is deprecated since 8.5, as it has no effect since PHP 8.0 in /home/u483256323/domains/poorvam.com/public_html/subdomains/pore/includes/api.php on line 76
PORE

About

Connect