Abstract
<title>Abstract</title> <p>The central dogma of materials informatics is the "Scaling Hypothesis": the assumption that chemical space is a continuous manifold, and that, with sufficient data, deep learning models can approximate any structure-property relationship. While this strategy has achieved near-DFT accuracy for thermodynamic properties like formation energy, prediction accuracy for bandgaps suffers from a fundamental, non-converging error floor (~0.4 eV). By analyzing a comprehensive dataset of binary semiconductors and state-of-the-art Graph Neural Networks (including coGN and foundational embeddings), this work isolates the root cause: unlike the smooth, transferable landscape of total energy, the electronic bandgap manifold is topologically disjoint. Consequently, increasing dataset size fails to improve predictions, instead resulting in negative learning, where increasing the structural diversity of the training set (e.g., adding rocksalt data to zincblende models) degrades prediction accuracy for unseen chemistries. This can be traced back to the discrete quantization of atomic orbitals and the resulting band crossings, which create non-differentiable cusps in the property surface. This work establishes a boundary condition for the "Big Data" paradigm: simple regression on atomic coordinates cannot resolve the intrinsic discontinuities of eigenvalue phenomena. Instead, a more detailed integration of quantum mechanics is required, such as direct entrainment of Hamiltonians as descriptors.</p>