Deprecated: Function curl_close() is deprecated since 8.5, as it has no effect since PHP 8.0 in /home/u483256323/domains/poorvam.com/public_html/subdomains/pore/includes/api.php on line 184
Abstract
<jats:p>\textbf{Background:} Biomolecular dataset construction often requires coordinated retrieval from heterogeneous repositories, identifier mapping, cross-reference enrichment, source-specific parsing, and provenance recording. These operations are frequently implemented through project-specific scripts, making acquisition procedures difficult to inspect, reproduce, or adapt across studies. We present SilkRoute, an open-source Python framework that formalizes biomolecular data acquisition as descriptor-defined, source-aware, and provenance-tracked workflows, providing a reproducible foundation for multi-source biomolecular dataset construction. \textbf{Results:} SilkRoute uses machine-readable YAML descriptors to specify dataset intent, biomolecular modality, workflow mode, query logic, enrichment resources, execution parameters, and export settings. These descriptors drive a common execution model that coordinates primary retrieval and downstream enrichment while preserving source-specific outputs, interaction evidence when available, the original workflow configuration, metadata, and run summaries. We evaluated this model through three representative acquisition scenarios spanning proteins, compounds, and molecular interactions. In the protein-centered workflow, SilkRoute retrieved 2,444 reviewed antimicrobial protein records from UniProt and generated complementary outputs from AlphaFold DB, InterPro, Pathway Commons, and the Protein Data Bank. In the compound-centered workflow, a ChEMBL IC$_{50}$ query produced 1,445,939 activity records organized into query-defined potency ranges. In the interaction-centered workflow, 2,253 UniProt protein records were expanded with 902,713 BioGRID interaction records and 5,702 STRING interaction-partner records. Across these scenarios, the framework successfully applied the same descriptor-defined acquisition model to distinct biomolecular entity types, retrieval strategies, enrichment paths, and output structures. \textbf{Conclusions:} SilkRoute extends beyond sequence retrieval by providing a reusable acquisition layer for constructing multi-source biomolecular datasets. By separating primary retrieval from enrichment and preserving source-aware outputs together with workflow descriptors and execution metadata, the framework makes acquisition procedures easier to inspect, reproduce, archive, and adapt. SilkRoute does not replace biological curation, label validation, deduplication, partitioning, or benchmarking, but provides structured and traceable acquisition packages that support these downstream processes.</jats:p>