Abstract
<title>Abstract</title> <p>This is one of three companion preprints: this paper (P1) introduces the conservation–divergence spectrum, P2 maps when it can be trusted, and P3 cross-examines what the spectrum and a black-box protein language model each know about the same positions. Proteins in a shared-fold family use highly conserved positions to encode functions common to the whole family, and a smaller set of positions to encode subtype-specific function. How that second, specificity-encoding layer is written into sequence — and whether it obeys any general organizing principle — remains only partly understood. Here we place every structurally equivalent position of a protein family on a plane whose axes are within-class conservation and between-class divergence, and show that the resulting cloud is not an arbitrary scatter but a trend band constrained by an information-theoretic feasibility boundary. The two axes are negatively correlated (Spearman ρ = −0.64), but much of this negative correlation is generated by a matched permutation-null floor (companion preprint P3); the biologically specific signal is the position-wise excess above that null, not the raw trend. The plane's four corners correspond to four selection regimes, and its sparse conserved-yet-divergent corner isolates candidate specificity-determining positions (SDPs). Applied to G-protein-coupling specificity across 219 class-A G-protein-coupled receptors (GPCRs), the framework recovers a coupling code that is distributed over ~ 12 mostly buried positions rather than localized to a binding interface, is consistent with a steric rather than electrostatic encoding through side-chain volume, and is consistent with a phenomenological Boltzmann-competition interpretation among G proteins — a probabilistic rather than a deterministic code — though we do not measure binding free energies directly. The same code shows a spectrum of evolutionary ages, from anchors frozen for ~ 450 My to positions still turning over across vertebrate lineages. Because the method's pattern is consistent with the sequence shadow of a general physical law — non-covalent recognition set by free-energy differences that Boltzmann statistics amplify, a phenomenological interpretation rather than a direct energetic measurement — it transfers across superfamilies: from sequence alone and with no family-specific input, it recovers the textbook S1 specificity determinant of serine proteases (position 189) and the catalytic-loop Tyr-versus-Ser/Thr determinants of protein kinases. We position the framework as a sequence-level ranking lens — a fast, general prospecting instrument that reorders a family's positions by their probability of carrying specificity — not as a per-position decision procedure. The verdict on any single candidate remains with structure, energetics or experiment.</p>