Methods
Surrogate model architectures evaluated across Flow Arena and academic reference benchmarks. Each architecture has 2D and/or 3D variants, trained with standardized hyperparameters and loss functions for fair comparison.
U-Net
Encoder-decoder architecture with skip connections, the workhorse of dense prediction tasks.
Fourier Neural Operator
Spectral convolution architecture that learns mappings between function spaces in Fourier domain.
SegResNet
Residual encoder-decoder from MONAI designed for volumetric medical image segmentation.
Swin UNETR
Shifted window vision transformer encoder paired with a U-Net convolutional decoder.
Architecture comparison
U-Net
- 2D Params
- 7.3M
- 3D Params
- 20.8M
Best-performing architecture on arena benchmarks and consistently strong across all dataset families
Purely local receptive field — may miss long-range pressure communication across the reservoir
Fourier Neural Operator
- 2D Params
- 0.8M
- 3D Params
- 8.9M
Global receptive field from the first layer — captures long-range pressure communication naturally
Spectral truncation can blur sharp discontinuities (saturation fronts, fault boundaries)
SegResNet
- 2D Params
- 6.3M
- 3D Params
- 18.8M
Deep residual paths enable learning complex feature transformations without degradation
Originally designed for segmentation (discrete labels), not regression (continuous fields) — may need tuning
Swin UNETR
- 2D Params
- ~25M
- 3D Params
- 64.1M
Self-attention captures long-range spatial dependencies that CNNs may miss
Largest parameter count (64.1M 3D) — slowest to train and most memory-intensive
| Family | Category | 2D Params | 3D Params | Key Strength | Key Limitation |
|---|---|---|---|---|---|
| U-Net | Encoder-Decoder | 7.3M | 20.8M | Best-performing architecture on arena benchmarks and consistently strong across all dataset families | Purely local receptive field — may miss long-range pressure communication across the reservoir |
| Fourier Neural Operator | Neural Operator | 0.8M | 8.9M | Global receptive field from the first layer — captures long-range pressure communication naturally | Spectral truncation can blur sharp discontinuities (saturation fronts, fault boundaries) |
| SegResNet | Encoder-Decoder | 6.3M | 18.8M | Deep residual paths enable learning complex feature transformations without degradation | Originally designed for segmentation (discrete labels), not regression (continuous fields) — may need tuning |
| Swin UNETR | Transformer | ~25M | 64.1M | Self-attention captures long-range spatial dependencies that CNNs may miss | Largest parameter count (64.1M 3D) — slowest to train and most memory-intensive |