Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p> <bold>Purpose:</bold> Automatic interpretation of hand-drawn technical diagrams is difficult because the target is a directed graph rather than a collection of visible regions. A connector may appear nearly correct at the pixel level and still produce a wrong source--target relation. We study this gap between visual evidence and structural recovery. <bold>Methods:</bold> A multi-head graph-evidence network predicts node regions, arrow shafts, skeletons, endpoints, local direction, path progress, endpoint offsets, and endpoint tangents. Long-arrow-aware supervision reduces fragmentation in extended connectors, while a separate 34-layer residual network refines symbol regions. A deterministic two-pass assembler first constructs a provisional physical graph and then uses graph context to complete selected attachments, suppress unsupported fallback traces and physical duplicates, recover strongly supported structural deficits, and refine matched node regions. <bold>Results:</bold> On 450 held-out diagrams containing finite automata and flowcharts, the complete framework obtains a node F1 score of 98.57%, maskIntersection over Union (IoU) of 82.71%, soft centerline Dice(Soft-clDice) of 96.72%, a 3.77-pixel endpoint error, a directed Link F1score of 92.49%, and a normalized Graph Edit Distance (nGED) of 0.090.Loop-back and decision-branch recall reach 95.24% and 94.04%,respectively. <bold>Conclusion:</bold> Learning-aligned fields make connector evidence recoverable, but reliable graph extraction also requires delaying ambiguous structural decisions until a provisional graph exists. The two-pass formulation provides that context without introducing a separate semantic reasoner. </p>

Show More

Keywords

graph regions structural node endpoint

Related Articles

PORE

About

Connect