Flow Arena
Standardized benchmark for surrogate reservoir simulation models with 12 datasets spanning two property types, three fault configurations, and two well control strategies.
Private data, public evidence.
Flow Arena is the proprietary benchmark layer of this portal. The raw simulation tensors remain private, while the benchmark design, evaluation protocol, aggregate metrics, and model standings are published for scrutiny.
The suite was built to test where surrogate reservoir models break: high-contrast channels, fault transmissibility barriers, variable-throw geometry, and operational control changes. Reference academic benchmarks validate the pipeline; Flow Arena measures industrial robustness.
Dataset axes
Flow Arena comprises 12 datasets formed by the Cartesian product of three axes: property type (Channels, Geostat), fault configuration (No Fault, Zero Throw, Variable Throw), and well control (BHP, Rate).
Throw is the vertical offset between rock layers across a fault plane. No Fault has none; Zero Throw faults perturb only the transmissibility between adjacent cells (no grid distortion); Variable Throw faults add irregular vertical displacement and the resulting grid geometry changes.
Simulations are two-phase oil-water flow run in OPM Flow over 20 reporting steps. Dataset size scales with geological complexity: 1,000 No Fault cases, 2,000 Zero Throw, 4,000 Variable Throw — each faulted structural model also includes 2 random fault-transmissibility multiplier realizations on the same grid.
Stable design, evolving implementation.
The benchmark axes are stable, but simulation schedules, preprocessing details, training recipes, and published results can change as the dataset and models are refined.
The website reads the current simulator protocol from Arena metadata synced from the data pipeline. When the pipeline changes, update the protocol metadata and regenerate the website data instead of editing page copy in multiple places.
Start at the leaderboard, finish at the matrix.
Compare within a dataset
Each dataset fixes property family, fault mode, and well control. Rank models there first before averaging across the full matrix.
Separate pressure from SWAT
Pressure and saturation errors are reported separately because a model can preserve pressure support while smearing the water front.
Look for robustness
A single win is less informative than stable rank across channels, geostatistical fields, zero-throw faults, variable-throw faults, BHP, and rate control.
2 property types x 3 fault regimes x 2 controls
Channels
High-contrast facies fields that create sharper permeability transitions and saturation fronts.
Baseline 8-layer geometry without fault barriers.
Fault transmissibility heterogeneity without vertical displacement.
Variable-throw faulted geometry with 32 ML-grid layers and irregular vertical structure.
Geostat
Continuous geostatistical property fields that test smooth heterogeneity and broad pressure response.
Baseline 8-layer geometry without fault barriers.
Fault transmissibility heterogeneity without vertical displacement.
Variable-throw faulted geometry with 32 ML-grid layers and irregular vertical structure.
Representative Flow Arena geology.
Static 3D renders show the same realization under no-fault and variable-throw geometries. Zero-throw cases use the same fault locations as VT, but their visible grid geometry is close to NF, so the portal shows NF and VT as the clearest geometry contrast.
Geostat
Channels
Generated from Flow Arena grids with representative geostatistical porosity and channel facies arrays. Vertical scale is exaggerated 10x for visual inspection.
Pressure and saturation through time.
A Channels VT case shows raw OPM Flow training-data targets on the simulator CPG grid for two fault-transmissibility multiplier realizations. No model prediction is shown. The panels use the same 3D camera and vertical scale as the geology views, with pressure on the first row and SWAT on the second row.
Grouped error reveals what each axis costs.
These summaries average overall rel-L² across all current model runs in each slice. Lower is better; use them as a directional difficulty map, not a replacement for per-model leaderboards.
Property type
avg rel-L²Fault regime
avg rel-L²Well control
avg rel-L²Where each architecture wins and slips.
Each cell ranks a model within one Flow Arena dataset by overall rel-L². This makes robustness visible without over-weighting runtime.
Best result per cell across all loss variants.
Each cell shows the minimum test rel-L2 achieved by a model on that dataset over every evaluated loss configuration, with rank in each column. Lower is better.
Pressure
SWAT
Isolating one loss configuration at a time exposes the cost of switching loss directly.
abslp_p2 mse What the reported errors mean.
Wins matter, but worst-case behavior matters more.
Flow Arena ranks architectures across all 12 private datasets, exposing whether a model is consistently strong or only wins narrow cases.
What is public, what is restricted, and how access works.
Published
- Benchmark axes, grid dimensions, fields, and sample counts
- Held-out test metrics, per-field errors, and per-timestep trends
- Model parameters and training configuration
- Aggregate rankings and production-ready comparison tables
Restricted
- Raw simulation tensors and full geological realizations
- Complete private train, validation, and test case files
- High-resolution internal exports used for model development
- Access to private data or hosted evaluation by agreement
Evaluation workflow
- Submit a trained model or run an agreed hosted evaluation
- Receive per-dataset, per-field, and per-timestep reports
- Decide which aggregate results can be published
- Use the same protocol for internal model selection