Back to Search View Original Cite This Article

Abstract

<jats:p>Confidence in geological CO2 storage forecasts rests on simulators whose discrepancies are, for the most part, characterised against each other rather than against a truth. We wrote a two-phase, two-component, fully implicit CO2 storage simulator from scratch, ran it on the SPE11B benchmark to the specification's own reporting resolution (840 × 120 cells, 1,000 yr), and used the reimplementation as an instrument for asking where the discrepancy against the 32 archived submissions comes from. A clause-by-clause reading of the specification against the code found 7 conformance defects, none of which the verification suite—465 assertions including analytical solutions checked to machine precision—could have detected, because every test encoded the same reading as the code. Repairing 6 of them at fixed grids moved the run 9 and 10 of 54 sampled quantity–time pairs further inside the peer envelope, a measure that cannot be improved by becoming more typical; grid refinement at matched conformance moves that count by less. Because both conformance stages were run at all four grids, the stage change and the resolution change are separated rather than confounded: the largest per-quantity stage effect decays monotonically from 0.88 to 0.03 MAD as the grid refines, and its aggregate, which changes sign twice across the ladder, ends at -0.01 MAD at the reporting grid. Two measured floors calibrate that—a change of time-step policy alone moves the aggregate by 0.46 MAD on identical physics, and a tenfold-finer injection step by 0.06 MAD, more than the finest refinement buys. Separately, a silent clamp on the dissolved-CO2 mass fraction inside the Newton loop created a false convergence wall: it capped the time step at 0.095 yr and consumed 52% of wall-clock time in discarded iterations while every gate stayed green. Of 18 performance interventions measured by isolated paired A/B, 8 survived and 4 first reported the wrong sign. On the Sleipner benchmark model, at a single fixed resolution, five reduced models scored against the observed top-sand plume outline exceed a matched-area baseline by at most 0.035 in overlap—so which mechanism a model represents separates them less than an explicit baseline reveals. On the benchmark, where both axes were varied, the larger discrepancy shifts came from what the model was asked to represent rather than from how finely it was resolved. Benchmark reporting should require conformance to be demonstrated rather than assumed.</jats:p>

Show More

Keywords

than against from rather benchmark

Related Articles

PORE

About

Connect