Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Causal discovery faces a dimensionality ladder that no single architecture spans. Cluster-aware methods provide robust inference at d &lt; 200 but collapse with their underlying NOTEARS engine at higher dimensions. The regime d = 200– 500—where most real transcriptomic datasets reside—has no gradient-based solution. We propose Causal Transformer (CT), a self-attention architecture that treats variables as tokens and learns causal edges from multi-head attention, retaining NOTEARS’ acyclicity constraint but constructing W through attention rather than independent optimization. An 80-conffguration sweep reveals a sharp viability boundary: CT discovers near-zero edges at d ≤ 100 but activates at d = 200, producing 1,028 edges versus NOTEARS’ 0. On TCGA-BRCA at d = 250–350 (5 seeds), CT remains operational (62,500–122,500 edges) where NOTEARS collapses. Multi-omics validation exposes a modality-speciffc limit: zero edges on methylation, motivating consensus approaches. An MLP baseline matches raw output but lacks CT’s emergent head specialization and pairwise gradient signals. On 10 TCGA cancer types at d = 200, CT recovers 198 edges on average versus NOTEARS’ 0, with 75% of top-hub genes validated as ClinGen Tier-1 cancer drivers. CT extends DAG-based causal discovery to the frontier of d ≈ 500, operating reliably at d = 250–350 and revealing its O(d 2) memory ceiling near d = 500, all on consumer hardware.</p>

Show More

Keywords

edges notears causal discovery architecture

Related Articles

PORE

About

Connect