Completed long-run evidence

Oil-water saturation case study

Static figures from a completed 500-epoch run. The panels show the model input context, including initial-state fields, plus the true water saturation, the prediction, and the signed error at selected timesteps for a held-out case.

Held-out sample A

Oil-water saturation case study input fields for Held-out sample A
Held-out input fields. Continuous rock properties (e.g. permeability, porosity) are shown in the normalised space used by the model (z-score or min-max, depending on the family) — values are not in physical units. Binary fields (well masks, channel indicators) keep their native 0/1 range. See each panel title for which field is shown and at what depth.
Oil-water saturation case study true predicted and error panels for Held-out sample A
True, predicted, and signed-error water saturation panels through prediction time for the same held-out case. Values are denormalised to physical units (see the unit in the panel title).

Held-out sample B

Oil-water saturation case study input fields for Held-out sample B
Held-out input fields. Continuous rock properties (e.g. permeability, porosity) are shown in the normalised space used by the model (z-score or min-max, depending on the family) — values are not in physical units. Binary fields (well masks, channel indicators) keep their native 0/1 range. See each panel title for which field is shown and at what depth.
Oil-water saturation case study true predicted and error panels for Held-out sample B
True, predicted, and signed-error water saturation panels through prediction time for the same held-out case. Values are denormalised to physical units (see the unit in the panel title).

500-Epoch Long-Run Evidence

UNet2D trained to convergence on the held-out test split, rendered alongside the Chen et al. 2025 APT published baseline (italicised reference rows) for direct comparison on the same metrics. Lower is better.

Target Model Size Loss Overall rel-L² Pressure rel-L² Saturation rel-L² APT delta_p APT delta_sw
combined UNet2D 16.38M Rel-Lp + Deriv. 0.0192 0.0172 0.0211 0.801% 0.161%
pressure UNet2D 16.36M Rel-Lp + Deriv. 0.0223 0.0223 - 1.051% -
saturation UNet2D 16.36M Rel-Lp + Deriv. 0.0319 - 0.0319 - 0.190%
reference FNO(APT, Chen 2025) 31.00M Chen et al. - - - 1.850% 1.280%
reference U-FNO(APT, Chen 2025) 33.00M Chen et al. - - - 0.570% 0.660%
reference APT(APT, Chen 2025) 12.00M Chen et al. - - - 0.600% 0.320%

200-Epoch Architecture Comparison

Broader sweep across architectures at a matched 200-epoch training budget. Useful for like-for-like architecture comparison; the 500-epoch table above shows what UNet2D achieves with extended training.

Target Model Size Loss Overall rel-L² Pressure rel-L² Saturation rel-L² APT delta_p APT delta_sw
combined UNet2D 16.38M Rel-Lp + Deriv. 0.0205 0.0176 0.0234 0.823% 0.207%
combined UNet2D 7.30M Rel-Lp 0.0256 0.0243 0.0268 1.138% 0.151%
combined UNet2D 16.38M Rel-Lp 0.0264 0.0284 0.0243 1.498% 0.153%
combined SegResNet2D 14.21M Rel-Lp 0.0380 0.0384 0.0375 1.879% 0.289%
combined SegResNet2D 6.33M Rel-Lp 0.0442 0.0483 0.0401 2.477% 0.389%
combined FNO2D 14.66M Rel-Lp 0.0469 0.0434 0.0504 2.156% 0.242%
combined SwinUNETR2D 15.21M Rel-Lp 0.0535 0.0543 0.0527 2.659% 0.570%
combined FNO2D 2.43M Rel-Lp 0.0547 0.0555 0.0539 2.878% 0.285%
combined SwinUNETR2D 6.79M Rel-Lp 0.0556 0.0566 0.0547 2.725% 0.574%
pressure UNet2D 16.36M Rel-Lp + Deriv. 0.0251 0.0251 - 1.191% -
pressure UNet2D 7.29M Rel-Lp 0.0317 0.0317 - 1.584% -
pressure UNet2D 16.36M Rel-Lp 0.0351 0.0351 - 1.788% -
pressure SwinUNETR2D 15.21M Rel-Lp 0.0420 0.0420 - 2.085% -
pressure SegResNet2D 6.32M Rel-Lp 0.0454 0.0454 - 2.382% -
pressure SegResNet2D 14.20M Rel-Lp 0.0500 0.0500 - 2.593% -
pressure SwinUNETR2D 6.79M Rel-Lp 0.0554 0.0554 - 2.679% -
pressure FNO2D 14.65M Rel-Lp 0.0581 0.0581 - 2.946% -
pressure FNO2D 2.43M Rel-Lp 0.0596 0.0596 - 3.058% -
saturation UNet2D 16.36M Rel-Lp 0.0299 - 0.0299 - 0.168%
saturation UNet2D 7.29M Rel-Lp 0.0319 - 0.0319 - 0.183%
saturation UNet2D 16.36M Rel-Lp + Deriv. 0.0345 - 0.0345 - 0.224%
saturation SegResNet2D 14.20M Rel-Lp 0.0407 - 0.0407 - 0.310%
saturation SegResNet2D 6.32M Rel-Lp 0.0436 - 0.0436 - 0.329%
saturation SwinUNETR2D 15.21M Rel-Lp 0.0479 - 0.0479 - 0.511%
saturation SwinUNETR2D 6.79M Rel-Lp 0.0539 - 0.0539 - 0.501%
saturation FNO2D 14.65M Rel-Lp 0.0565 - 0.0565 - 0.278%
saturation FNO2D 2.43M Rel-Lp 0.0617 - 0.0617 - 0.314%

Leaderboard

Loss
37 results
    Pres (bar)Sat (frac) 
#TargetModelLossEpochsrel-L²MREMAErel-L²MREMAEParams
1combinedUNet3DBadawiCombined20000.01030.00861.4302300.01820.01400.00476520.8M
2combinedUNet3DBadawiCombined2000.01420.01242.1764280.02540.02070.00715620.8M
3combinedUNet2DBadawiCombined40000.01430.01212.0387410.02020.01510.00513229.1M
4combinedUNet2DBadawiCombined40000.01500.01282.1364730.02000.01470.00503216.4M
5combinedUNet2DBadawiCombined40000.01650.01422.3641490.02090.01540.00524029.1M
6combinedUNet2DBadawiCombined5000.01790.01492.4694670.02390.01800.00617516.4M
7pressureUNet2DBadawiSingleField5000.02150.01923.150069---16.4M
8combinedUNet2DRelLp2000.02430.02043.4150110.02680.02010.0067987.3M
9pressureUNet3DBadawiSingleField2000.02840.02684.242195---20.8M
10combinedUNet2DRelLp2000.02840.02594.4944010.02430.01870.00645516.4M
11pressureUNet2DRelLp2000.03170.02784.753270---7.3M
12pressureFNO3DBadawiSingleField2000.03440.03005.154885---5.6M
13pressureUNet2DRelLp2000.03510.03165.364382---16.4M
14combinedSegResNet2DRelLp2000.03840.03365.6363610.03750.03070.01087814.2M
15combinedFNO3DBadawiCombined2000.04150.03576.1024670.05320.04230.0140225.6M
16pressureSwinUNETR2DRelLp2000.04200.03606.253993---15.2M
17combinedFNO2DRelLp2000.04340.03756.4694210.05040.03790.01249014.7M
18pressureSegResNet2DRelLp2000.04540.04067.145418---6.3M
19combinedSegResNet2DRelLp2000.04830.04357.4309820.04010.03330.0113066.3M
20pressureSegResNet2DRelLp2000.05000.04657.779330---14.2M
21combinedSwinUNETR2DRelLp2000.05430.04647.9775820.05270.04340.01664615.2M
22pressureSwinUNETR2DRelLp2000.05540.04538.037558---6.8M
23combinedFNO2DRelLp2000.05550.04828.6347620.05390.04140.0135962.4M
24combinedSwinUNETR2DRelLp2000.05660.04788.1752370.05470.04590.0169386.8M
25pressureFNO2DRelLp2000.05810.05218.838425---14.6M
26pressureFNO2DRelLp2000.05960.05289.173145---2.4M
27saturationFNO2DRelLp200---0.06170.04650.0153502.4M
28saturationFNO2DRelLp200---0.05650.04260.01402714.6M
29saturationFNO3DBadawiSingleField200---0.05550.04310.0144905.6M
30saturationSegResNet2DRelLp200---0.04360.03470.0118776.3M
31saturationSegResNet2DRelLp200---0.04070.03310.01149614.2M
32saturationSwinUNETR2DRelLp200---0.05390.04310.0158616.8M
33saturationSwinUNETR2DRelLp200---0.04790.04030.01500315.2M
34saturationUNet2DRelLp200---0.03190.02430.0081137.3M
35saturationUNet2DRelLp200---0.02990.02270.00762816.4M
36saturationUNet2DBadawiSingleField500---0.03020.02340.00788016.4M
37saturationUNet3DBadawiSingleField200---0.02400.01810.00626420.8M

Metrics Over Time

Combined target variant (multiple fields predicted jointly).

Pressure

SATURATION

Paper Metrics Comparison

Metrics from the shared paper-metric module (utils/paper_metrics.py) evaluated on the held-out test split. Loss column distinguishes the training recipe; lower is better.

Targetpressure

#ModelLossEpochspressure_mre_t1pressure_mre_t2pressure_mre_pinit
1UNet2DBadawiSingleField5002.33181.50351.0500
2UNet2DRelLp2003.61782.24481.5844
3UNet3DBadawiSingleField2003.90582.35241.4141
4UNet2DRelLp2004.05202.45001.7881
5FNO3DBadawiSingleField2004.09742.45011.7183
6SwinUNETR2DRelLp2005.06112.85802.0847
7SegResNet2DRelLp2005.41723.10932.3818
8SegResNet2DRelLp2005.88113.64372.5931
9FNO2DRelLp2006.13464.05602.9461
10FNO2DRelLp2006.15154.12313.0577
11SwinUNETR2DRelLp2006.45323.60392.6792

Targetsaturation

#ModelLossEpochssaturation_mape
1UNet3DBadawiSingleField2000.4538
2UNet2DRelLp2000.5251
3UNet2DBadawiSingleField5000.5556
4UNet2DRelLp2000.5643
5SegResNet2DRelLp2000.9187
6SegResNet2DRelLp2000.9312
7FNO2DRelLp2000.9649
8FNO3DBadawiSingleField2001.0690
9FNO2DRelLp2001.0787
10SwinUNETR2DRelLp2001.3254
11SwinUNETR2DRelLp2001.3373

Targetcombined

#ModelLossEpochspressure_mre_t1pressure_mre_t2pressure_mre_pinitsaturation_mape
1UNet3DBadawiCombined20001.22830.74300.47670.3379
2UNet2DBadawiCombined40001.47800.96050.67960.3534
3UNet2DBadawiCombined40001.55751.00990.71220.3601
4UNet2DBadawiCombined40001.74511.13840.78800.3681
5UNet3DBadawiCombined2001.74780.98760.72550.5377
6UNet2DBadawiCombined5001.92491.22790.82320.4441
7UNet2DRelLp2002.97591.70881.13830.4817
8UNet2DRelLp2003.17771.90511.49810.4735
9SegResNet2DRelLp2004.72852.61871.87880.8947
10FNO2DRelLp2004.74442.91382.15650.8582
11FNO3DBadawiCombined2005.42092.76222.03421.0553
12FNO2DRelLp2005.59173.69362.87830.9702
13SegResNet2DRelLp2006.45003.37942.47700.9206
14SwinUNETR2DRelLp2006.53263.69782.65921.5224
15SwinUNETR2DRelLp2007.33363.87772.72511.4949

About This Benchmark

This benchmark is based on the dataset introduced by Badawi & Gildin (2024), featuring two-phase (oil-water) immiscible flow simulations on a 40x40 Cartesian grid. The simulations are generated using CMG IMEX, a commercial black-oil simulator.

Each simulation spans 10 years with a reporting interval of approximately 0.5 days, producing 366 raw snapshots. These are windowed into 40-timestep sequences for training. The geological models feature heterogeneous permeability fields with varying well counts (3-11 wells) and placements.

The training set comprises 3500 simulations augmented 4x via geometric flips (horizontal, vertical, and combined), yielding 14000 effective training samples. Validation uses 500 simulations and testing uses 200 simulations. Importantly, the test set contains out-of-distribution well configurations not seen during training, testing the model's ability to generalize to novel operational scenarios.

Input features include: log-normalized permeability, producer bottom-hole pressure (BHP), injector rate, binary well location masks, initial pressure, and initial water saturation — 6 channels total.

We also compare against the APT (Approximate Physics Transformer) method from arXiv 2602.11208, which reports delta_p (pressure relative error) and delta_sw (saturation relative error) metrics.

Dataset details

train
3500
train augmented
14000
val
500
test
200
augmentation
4x geometric flips (horizontal, vertical, combined)
test note
Out-of-distribution well configurations (3-11 wells)

Input features

PermeabilityProducer BHPInjector RateWell LocationsP_initSw_init

Paper metrics explained

MMRE
max_t mean(top 5% relative errors)
Maximum Mean Relative Error (Badawi's metric). For each timestep, compute the mean of the top 5% largest relative cell errors, then take the maximum over all timesteps. Captures worst-case spatial accuracy at the most challenging time.
delta_p
mean(||p_pred - p_true||_2 / ||p_true||_2)
APT pressure metric — mean relative L2 error over the test set.
delta_sw
mean(||sw_pred - sw_true||_2 / ||sw_true||_2)
APT saturation metric — mean relative L2 error over the test set.

Training Configuration

Configuration for the best-performing model (UNet3D).

Training
Epochs2000
Learning Rate0.001
Weight Decay0.0001
Max Grad Norm1
Batch Size16
Loss FunctionBadawiCombinedLoss (beta=1, sat_threshold=0.01)
SchedulerCosineAnnealingLR (step_size=20, gamma=0.95, T_max=2000, eta_min=0.0001)
Model Architecture
in channels6
out channels2
base channels32
kernel size3
depth3
num time steps1
squashed outputfalse
flatten outputtrue