Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Chemical literature tables are an important source for chemical knowledge extraction and structured data. They contain heterogeneous text, numerical values, chemical formulas, footnotes, and molecular-structure depictions. Existing optical character recognition (OCR) models, general table recognition models, and vision-language models (VLMs) can recover part of row-column structure and text; for molecular-structure cells, however, they often return placeholders, empty contents, or weak descriptions, making it difficult to obtain computable Simplified Molecular Input Line Entry System (SMILES) representations. This study addresses row-column recognition, cell-text recovery, and molecular-structure depiction recognition by developing a heterogeneous cell recovery framework under table-coordinate constraints. We propose MolTabReco, a multi-stage framework that localizes the table region, uses VLM-predicted row-column structure and text results, and combines them with a structure detection model to generate cell locations and a unified table coordinate system. A molecular-structure cell discriminator routes each cell to a text recovery branch or an optical chemical structure recognition (OCSR)-SMILES branch. Non-structure cells use conservative VLM-first fusion; structure cells use a MolScribe-led OCSR-SMILES strategy with crop expansion, image enhancement, foreground segmentation, RDKit validity checking, canonicalization, and fragment filtering. Large language model (LLM)-based correction is restricted to controlled review for low-confidence or abnormal cells. On ChemTable under a shared VLM-predicted grid, MolTabReco achieved 44.05% canonical exact match (EM) on 1,682 molecular-structure cells. On dimension-matched pages, 1,159 structure cells reached 54.79% canonical EM, 60.31% no-stereochemistry EM, and 94.48% valid SMILES recognition. Among 1,095 parseable predictions, 65.8% had Tanimoto similarity ≥ 0.9 and 77.5% had similarity ≥ 0.5. The VLM baseline could not convert depictions into SMILES; its 8.09% score corresponded to placeholder or weak-content matching. Because structure cells are rare, Overall Value Accuracy increased from 82.53% to 83.29%, and Tree Edit Distance based Similarity from 0.8537 to 0.8595. MolTabReco converts structure cells that would otherwise remain at the placeholder level into canonical SMILES linked to row-column positions. Scientific Contribution MolTabReco extends chemical table recognition beyond layout and text extraction by jointly recovering molecular-structure cells as canonical SMILES under explicit table-coordinate constraints. The framework integrates table-region localization, VLM-predicted row-column grounding, structure-cell routing, and OCSR-based SMILES recovery so that chemical structures are linked to their original row-column context. This provides a reproducible pathway for converting molecular-structure cells that would otherwise remain empty or as image placeholders into computable chemical data.</p>

Show More

Keywords

cells structure chemical molecularstructure recognition

Related Articles

PORE

About

Connect