Back to Search View Original Cite This Article

Abstract

<jats:p>Artificial intelligence for domain science relies on large, well-structured training corpora that integrate experimental data with explicit contextual knowledge and the tacit expertise underlying scientific decisions. However, scientific information remains fragmented across publications. In particular, the information regarding practical reasoning, assumptions, failed attempts, parameter choices, and interpretive judgment is often held by authors and unpublished, hindering its use for model training. Here, we propose PaperBot, a large language model (LLM)-enabled, author-curated scientific publishing framework that organizes these resources into a customed ChatBot, enabling the integration of connected, AI-ready data, explicit knowledge, and tacit scientific expertise. Within PaperBot, the LLM serves as an interactive reasoning and orchestration layer that connects knowledge content with raw and processed data, figures, methodological details, metadata, code, and author-provided interpretation. We introduce a standardized, platform-independent workflow for constructing PaperBots, supported by a reusable SKILL, defined schemas, and provenance requirements that enable the export of structured scientific records for downstream model development. Case studies in materials synthesis and characterization demonstrate how the LLM-enabled framework supports questioning, data visualization and inspection, numerical verification, and preparation of AI-ready datasets. By combining author-based organization of scientific data and contextual knowledge, PaperBot provides a scalable pathway toward provenance-rich training corpora and more reliable AI-enabled scientific workflows.</jats:p>

Show More

Keywords

scientific data knowledge training model

Related Articles

PORE

About

Connect