FhSim  3.1.0
Marine systems simulation
Loading...
Searching...
No Matches
Net hydrodynamics: benchmark of load laws and wake evaluation

This page reports what the panel load laws and the net's wake cost, and how accurate the gridded wake is at each grid size and cutoff. It is the benchmark of the net-hydrodynamics plan (WP-B6, MARE-0102). The default WakeGrid and WakeCutoff are the owner's decision (owner question Q12, MENV-0008, lead ruling R21); § 8. Proposed defaults proposes values from the numbers and changes nothing. The owner answered Q12 on 2026-09-27 after § nbench_phase_f: WakeCutoff stays 1e-3 and WakeGrid 96 × 48 × 48, both user-settable. The wake manuals (marenv doc/wake.md § 5.4 "Cost", fhsim_environment doc/user/wakes.md "Cost") can link here; the section anchors are listed in § 10. Summary for the wake manuals. Sections 2 to 10 predate phase F; § 11. Phase F re-benchmark re-measures the wake after it (WP-F6, MARE-0125), when other objects see one lumped source per net, and its decision table replaces § 8's proposal for Q12.

1. Set-up

CPU AMD RYZEN AI MAX+ PRO 395 w/ Radeon 8060S (32 logical CPUs), WSL2, Linux 6.6
Cores One: each benchmark pinned with taskset -c <n>
Build Release, Conan default profile (gcc, C++20); marenv -O3 -mavx2
Code marenv wp/B6-benchmark at 0ed1cd1 (integration head de003ab plus the bench); fhsim_marine_elements wp/B6-benchmark at 0803ca2 (integration head 27b07ae plus the bench), against the editable fhsim_environment 4.0.0 and marenv 2.0.0
Repeats Each figure is the median of at least 5 repeats (7 for the RHS, 15 for the load laws)
Timer CPU time: the thread's for the wake fields and the load laws, the process's (fhsim::test::TestResult::cpuSeconds) for the FhSim runs. Rebuild times are the ones NetWakeCaster logs (wall clock of one rebuild)
Date 2026-09-27

Noise. The machine was shared with other jobs, and its speed drifted between runs of the same binary by up to about 35 % (the same load-law loop ran at 44.6 ns and at 37.9 ns in two back-to-back runs). Within one run the cases are interleaved, so the ratios (Dual against double, Grid against CastWake off, cutoff against cutoff) are stable to about ±10 %; absolute times are good to about ±35 %. Where two passes differ, the table gives the later, quieter one.

Reproducing. Both benchmarks are built only on request and are never run by ctest:

# marenv: a separate build folder on the Release toolchain of `conan build`
cd <marenv worktree>
cmake -S . -B build/bench -DCMAKE_TOOLCHAIN_FILE=build/Release/generators/conan_toolchain.cmake \
-DCMAKE_BUILD_TYPE=Release -DMARENV_WITH_DOC=OFF -DBUILD_TESTING=OFF -DMARENV_BENCH=ON
cmake --build build/bench --target marenv_wake_bench
taskset -c 4 ./build/bench/bench/marenv_wake_bench # about 4 min
# fhsim_marine_elements: in the conan build folder (the libfhsim_environment.so symlink of BUILD.md in place)
cd <fhsim_marine_elements worktree>/build/no_vis/Release
cmake -DFHSIM_ME_BENCH=ON . && cmake --build . --target fhsim_me_bench
cd playpen/bin && taskset -c 4 ./fhsim_me_bench all <work directory> # about 40 min; sections: laws, rhs, rebuild, share

marenv/bench/WakeBench.cpp measures the wake fields on their own, with the cage of test W5. bench/NetBench.cpp measures the load laws and runs a 968-panel Net/NetStructure through fhsim::test::RunTest; it writes its own net file and inputs into the work directory, and its report into <work directory>/benchmark.md.

The two test structures.

  • W5 cage (marenv, tests/WakeGrid_Test.cpp): a cylinder of 40 × 25 = 1000 panels, diameter 6 m, from 1 m to 6 m depth, current along +x, C_D = MF2022 Eq. 10 at Sₙ = 0.2 times |cos θ|. Grid box 26.5 m × 9 m × 7 m (3.5 m upstream of the centre, 23 m downstream). The reference is SourceSetWake (the direct sum, no cutoff) at the 46 230 points of the W5 lattice.
  • 968-panel net (fhsim_marine_elements): a 20 m × 20 m screen of 22 × 22 squares, two triangles each (529 nodes, panels of 0.41 m²), SolidityModel Fixed, Sₙ = 0.2, default law LocalThroughFlow, standing across a 0.5 m/s current. MassFactor 1e9 holds it in place (the set-up of D1's timing net), so every run does the same work and Freeze converges after one rebuild. Explicit Euler with 1 ms steps: one step is one right-hand-side evaluation. WakeLength 0 gives a grid 5 × 20 m long downstream.

2. Gridded wake: accuracy against cost

W5 cage, 1000 sources. Build is one WakeBuilder pass (all sources added, Finish(), since marenv MENV-0058 Finish(velocity)). max |Δd| (W5) is the plan's W5 measure, GriddedWake against SourceSetWake at the lattice points more than 2 cells (2 max_a h_a) from any source; W5's tolerance is 0.01. Because that exclusion shrinks with the cell size, the next column repeats the measure beyond the same 0.83 m (2 cells of the coarsest grid) for every grid, which compares like with like. The last column is the far wake alone, x ≥ 6 m (3 m behind the cage's rear).

Grid Cells Cutoff Build, ms max |Δd| (W5) where, x [m] max |Δd| beyond 0.83 m max |Δd|, x ≥ 6 m
64 × 64 × 64 262 144 1e-3 362 0.160 23.0 0.160 0.160
1e-4 1713 0.020 3.8 0.020 0.0125
3e-5 2409 0.021 3.8 0.021 0.0033
1e-5 2909 0.021 3.8 0.021 0.0011
96 × 48 × 48 (default) 221 184 1e-3 (default) 256 0.160 23.0 0.160 0.160
1e-4 884 0.038 −2.2 0.032 0.0126
3e-5 1249 0.038 −2.2 0.033 0.0033
1e-5 1509 0.038 −2.2 0.033 0.0017
128 × 64 × 64 524 288 1e-3 361 0.160 23.0 0.160 0.160
1e-4 1962 0.026 −2.2 0.020 0.0125
3e-5 2760 0.026 −2.2 0.020 0.0033
1e-5 3378 0.026 −2.2 0.020 0.0012
96 × 96 × 96 884 736 1e-3 652 0.161 23.0 0.161 0.161
1e-4 3226 0.015 −2.2 0.013 0.0126
3e-5 4800 0.014 −2.2 0.013 0.0033
1e-5 5776 0.014 −2.2 0.013 0.0010
64 × 96 × 96 (added) 589 824 1e-3 475 0.161 23.0 0.161 0.161
1e-4 2528 0.016 3.8 0.016 0.0126
3e-5 3443 0.015 3.8 0.015 0.0033
1e-5 3855 0.015 3.8 0.015 0.0010

The largest direct deficit on the lattice is 0.256. 64 × 96 × 96 is not one of the plan's grids; it was added to tell the resolution across the flow from the resolution along it.

What the table shows:

  1. Two separate errors. At cutoff 1e-3 the error is the truncation of the far wake and does not depend on the grid: 0.160 at x = 23 m, where the direct sum gives d ≈ 0.19 (B4's MENV-0004 finding, reproduced on every grid). Below 1e-4 the error is the resolution near and inside the cage (x = −2.2 m is inside it, x = 3.8 m just behind it), and does not depend on the cutoff.
  2. Truncation scales with N × cutoff. The far-wake error is 0.0125 at 1e-4 and 0.0033 at 3e-5, about 0.12 · N · cutoff for these N = 1000 sources, and levels off near 0.001 at 1e-5. This is the tails that each source drops adding up (MENV-0008). The 0.12 is for this geometry and one N; for other N it is an extrapolation.
  3. Resolution is set by the cells across the flow. 64³ and 128 × 64² have the same lateral cells (0.14 m × 0.11 m) and the same error, 0.020, although 128 × 64² has twice the cells along the flow; 96³ and 64 × 96² share 0.094 m × 0.073 m and give 0.013 and 0.015. The default 96 × 48² (0.19 m × 0.15 m) gives 0.033.
  4. No swept setting meets W5 (max |Δd| < 0.01). The best is 96³ at 1e-5, 0.013–0.014. B4 reported 0.009 at 96 × 192² for 23.5 s per build. W5 stays DISABLED_ (R21).
  5. Build cost rises about 5× from cutoff 1e-3 to 1e-4 and another 1.5–2× to 1e-5, and about in proportion to the lateral cell count.

3. Wake rebuild of a net

One rebuild of the 968-panel net in NetWakeCaster (WakeUpdate Off: the build at t = 0, as logged by the caster; this includes the marching build with one load-law evaluation per panel). The cage wake of § 2 crosses the whole grid box; the flat screen's footprints are narrower, so the net's rebuild is cheaper than the cage's build at the same grid and cutoff.

Grid Cutoff 1e-3 1e-4 3e-5 1e-5
64 × 64 × 64 37 ms 400 ms 720 ms 1100 ms
96 × 48 × 48 (default) 25 ms 284 ms 596 ms 823 ms
128 × 64 × 64 46 ms 625 ms 1197 ms 1835 ms
96 × 96 × 96 92 ms 1214 ms 2126 ms 3070 ms
64 × 96 × 96 122 ms 1277 ms 1559 ms 2544 ms
Direct (SourceSetWake, no grid) 1.7 ms

An earlier pass under heavier machine load gave the same picture with times up to 1.8× longer (for example 96³ at 1e-5: 4456 ms).

4. Query cost

One Sample() of each wake field at random points in the W5 grid box (10⁶ queries per repeat for the gridded and disc fields, 2·10⁷/N for the direct sum).

Field Sources µs per Sample()
GriddedWake 96 × 48 × 48 1000 0.043
GriddedWake 64³ / 128 × 64² / 64 × 96² / 96³ 1000 0.043 / 0.047 / 0.050 / 0.053
SourceSetWake (Direct) 100 1.41
SourceSetWake (Direct) 1000 15.1
SourceSetWake (Direct) 10 000 160
PorousDiscWake (since marenv MENV-0060 a one-source SourceSetWake) 1 0.020–0.028

The gridded query does not depend on N and hardly on the grid size (the larger grids fall out of the cache). The direct sum costs 15–16 ns per source, so at N = 1000 it is 350× a grid query.

5. NetStructure right-hand side

968-panel net, CPU time per step from the difference of a 1 s and a 5 s run (which removes the setup and the one build at t = 0), the three cases interleaved within each of 7 repeats. WakeUpdate Off, so no rebuild falls inside the measured interval; the ratio is taken within each repeat.

Wake ms per RHS min to max over the repeats against CastWake off
CastWake off 0.227 0.173 to 0.242 1
Grid, 96 × 48 × 48, cutoff 1e-3 0.256 0.195 to 0.265 1.13
Direct 2.92 2.54 to 3.44 14.7

Two earlier passes gave 1.13–1.14 for Grid and 13.7–15.8 for Direct. The gridded wake adds about 0.03 ms per RHS, 30 ns per panel query, whatever the grid and cutoff (§ 4). The direct wake adds N queries of N sources each: 2.7 ms per RHS here, 2.9 ns per panel-source pair. That is less than the 15 ns of § 4 because every panel of a flat screen lies in the plane of the other sources, upstream of their x_start, where the kernel returns early; in a structure whose panels lie in each other's wakes (a cage, a trawl) the 15 ns per pair of § 4 applies. The direct cost grows as N², so it equals the net's own panel work (about 0.23 µs per panel) at N ≈ 15–80 (15 ns to 3 ns per pair), and a Grid is cheaper per RHS for any larger net.

6. Share of a simulation spent in rebuilds

968-panel net, 60 s of simulation (60 000 steps), Grid 96 × 48 × 48. Measured, median of 5. The net is held in place, so Freeze converges at its first iteration: 2 builds (t = 0 and t = WakeIterInterval = 5 s). Periodic 1 s runs with WakeBlendTime 0.5 s (R22 requires it not to exceed the period).

Cutoff Schedule Rebuilds Rebuild total, s CPU total, s Share in rebuilds
1e-3 Freeze (defaults) 2 0.05 14.4 0.3 %
1e-3 Periodic 10 s 6 0.16 15.1 1.1 %
1e-3 Periodic 1 s 60 1.97 20.9 9.4 %
1e-5 Freeze (defaults) 2 1.36 14.4 9.5 %
1e-5 Periodic 10 s 6 3.86 16.1 24 %
1e-5 Periodic 1 s 60 40.8 50.7 80 %

A second pass gave the same shares within one percentage point. For other settings the share follows from § 3: share = n t_b / (n t_b + 14.35 s), with n rebuilds of t_b each and 14.35 s the rest of this 60 s run. Derived this way (not measured):

Grid Cutoff Freeze (2) Periodic 10 s (6) Periodic 1 s (60)
96 × 48 × 48 1e-4 3.8 % 11 % 54 %
96 × 48 × 48 3e-5 7.7 % 20 % 71 %
64 × 64 × 64 1e-5 13 % 32 % 82 %
96 × 96 × 96 1e-5 30 % 56 % 93 %
64 × 96 × 96 1e-5 26 % 52 % 91 %

The share falls with the length of the simulation under Freeze (a fixed number of rebuilds); under Periodic it does not. A Freeze that needs its full WakeIterations rebuilds after the first build costs 1 + WakeIterations builds: 3 with the 2 of this measurement, 5 with the default 4 since owner ruling R130.

7. Load-law evaluation

One hydrodynamics::Evaluate() through the PanelLoadLaw variant, as NetStructure calls it, on 1024 distinct panels (inflow angle 0–85°, 0.2–1.5 m/s, Sₙ 0.15–0.35, bar directions for TwineCrossFlow). Over Dual<18> every input carries a dense 18-slot gradient, as a panel quantity that depends on the 9 node positions and 9 node velocities does. Median of 15 repeats of 102 400 evaluations, the laws and scalar types interleaved; the quieter of two runs (the other was uniformly about 30 % slower, with the same ratios).

Law double, ns Dual<18>, ns Dual / double
LocalThroughFlow (default) 38 2100 56
ScreenKF2012 32 1650 51
ScreenKF2012, HydroInduction 34 1740 51
ScreenMF2022 5.8 740 127
TwineCrossFlow 99 2950 30

For the 968-panel net, the default law over double is about 37 µs of the 227 µs RHS (§ 5); over Dual<18> (the panel Jacobian) it is about 2 ms per Jacobian evaluation.

8. Proposed defaults

Proposal only; the defaults are unchanged (KernelParams::cutoff = 1e-3, WakeGrid = 96 × 48 × 48); the owner kept them (Q12, 2026-09-27).

WakeCutoff: 1e-5 (from 1e-3).

  • At 1e-3 the gridded wake of a 10³-panel structure loses its far wake: 0.160 against a direct d of 0.19 at 23 m, on every grid. That error is ten times the resolution error of the default grid and is the one systematic error in the table.
  • 1e-4 still leaves 0.0125 in the far wake, above W5's 0.01. 3e-5 leaves 0.0033 and 1e-5 0.0010–0.0017, both below the resolution error of every grid swept.
  • Because the truncation grows as N × cutoff (§ 2, item 2), 1e-5 keeps it at or below about 0.01 up to roughly 10⁴ panels (extrapolated), where 3e-5 would give about 0.03; 1e-5 is the value that holds for the largest nets in the plan (trawls) without a second decision.
  • Cost: 823 ms per rebuild of the 968-panel net on the default grid instead of 25 ms, 1.5 s per build of the 10³-panel cage instead of 0.26 s. Under the default Freeze schedule that is 2–3 builds per simulation, 9.5 % of a 60 s run and proportionally less of a longer one.
  • Consequence to document: with Periodic and a short WakePeriod the rebuilds dominate (80 % at 1 s). A Periodic user who needs short periods should set WakeCutoff 1e-4 (54 % at 1 s, truncation about 0.0125 at 10³ panels), or use Direct for a small net.
  • Alternative for the owner (Q12's "cutoff scaled with N"): a cutoff of about 1e-2 / N would hold the truncation near 0.001 at any N and be cheaper for small nets; it needs code in the caster, which this benchmark does not add.

WakeGrid: keep 96 × 48 × 48.

  • Its remaining error, 0.033, is resolution and sits inside and just behind the structure (x = −2.2 m in the cage); the far wake is right once the cutoff is lowered.
  • Halving that error needs twice the cells across the flow: 96³ gives 0.013 at 3.7× the rebuild (3.1 s against 0.82 s at 1e-5) and 30 % of a 60 s Freeze run. 64 × 96² gives 0.015 at 3.1× and saves little against 96³, so fewer cells along the flow buy little.
  • 64³ gives 0.020 at 1.3× the rebuild; it is the cheapest improvement, but it has a third fewer cells along the flow, which worsens the x_start resolution warning (self-shadowing, W6) on long grids such as the 100 m of the 968-panel net, so it is not proposed as the default.
  • Cells along the flow are the cheap axis to cut and the lateral axes the ones to raise; a user who needs the wake near the structure should raise the second and third WakeGrid counts. The manuals should say so.

WakeEvaluation: keep Grid. Direct is exact and has no build cost, but its RHS cost grows as N² (14.7× the net's RHS at 968 panels, § 5); it suits nets of up to a few tens of panels and reference runs.

9. Against the plan's estimates

Plan §B3 estimate Measured
Build 10–50 ms for 10³ sources 25 ms (968-panel net) and 256 ms (10³-panel cage) at the default grid and cutoff; 0.8 s and 1.5 s at cutoff 1e-5
Grid query about 50 ns 43–53 ns
Direct about 500× a grid lookup 350× at N = 1000 (15.1 µs against 0.043 µs)
Build "paid off after a few tens of right-hand-side evaluations" At the default settings one rebuild (25 ms) is about 110 RHS of the 968-panel net; at 1e-5 (823 ms) about 3600

10. Summary for the wake manuals

The figures a wake manual's cost section needs, with the anchor of their full table (page fhsim_marine_elements_net_benchmark, file fhsim_marine_elements/doc/user/validation/benchmark.md):

11. Phase F re-benchmark

Sections 2 to 10 measure the wake as it was before phase F, when a casting net registered its per-panel field and every other object sampled it. Since F4 (MARE-0129, ADR 0006) the net keeps that field private, for its own panels only, on a grid that covers the structure and one panel diameter beyond its rear (WakeLength 0), and its sources carry their solidity (the MF2022 Eq. 11 near wake, F1). Other objects sample one registered lumped source per net (F2, F4). This section repeats the measurements of § 2 to § 6 for that design (WP-F6, MARE-0125, owner question Q12, lead ruling R34 (d)). No default was changed: WakeCutoff is 1e-3 and WakeGrid 96 × 48 × 48; § 11.5 Decision table for Q12 gives the table the owner chooses from.

Set-up. Machine, build and timers as § 1. Code: marenv wp/F6-rebenchmark at d21c497 (integration head 1ccf0cf plus the bench); fhsim_marine_elements wp/F6-rebenchmark at e224f25 (integration head bab40b4 plus the bench) against the editable fhsim_environment (973c5c6) and marenv. The pre-F4 figures come from fhsim_marine_elements 122c590, the parent of the F4 merge 673da1c on the integration branch, built with the same NetBench.cpp against the same editables: it has the load laws of phase G, like the F4 code, and the pre-F4 NetWakeCaster, so a difference between the two columns is the wake alone. (5d291f0, before phase G, would mix in the new in-plane twine friction.) The marenv and fhsim_environment changes of F1 to F3 are additions; the kernel with Sn = 0 is the pinned one (F1 keeps its tests bit-identical), so the pre-F4 caster behaves in this build as it did before phase F.

Noise and repeats. Each benchmark ran pinned to one core, pre-F4 and F4 at the same time on two cores, in two passes; other agents' builds ran on the machine throughout. Times are medians of 5 repeats (3 for the 96 × 192 × 192 builds and the Direct farm runs); where both passes are shown they are "pass 1 / pass 2". Pairs of identical work differed by up to 35 % between passes and up to 80 % within one (single outliers); the accuracy figures and drags are deterministic.

Reproducing. As § 1, with the sections added for F6:

taskset -c 4 ./build/bench/bench/marenv_wake_bench f6 # marenv, about 5 min; f6-external: § 11.2 alone, f6-step: § 11.4
cd <fhsim_marine_elements build>/playpen/bin
taskset -c 4 ./fhsim_me_bench f6 <work directory> # rebuild-f6, farm and farm-sweep, about 15 min (pre-F4: 25 min)

The farm needs Test/NodeTow from libfhsim_marine_elements_test_objects.so, so the build tree must have been configured with the tests (FH_WITH_TESTS).

The two test structures.

  • W5 cage (marenv, as § 1): 1000 rectangular panel sources, now with Sn = 0.2 on each (what the F4 caster hands the private field). The panel centroids are the source positions. The pre-F4 box is the one the pre-F4 caster built for it (WakeLength 0 = 5 × the structure's size: x from −3.3 m to 34.6 m, y ± 9.2 m, z ± 8.6 m); the F4 box is the structure-only extent of fhsim_environment WakeCaster.cpp PrivateGrid (x ± 3.35 m, y ± 4.5 m, z ± 3.9 m). Both are computed in the bench with the caster's formulas.
  • Farm (fhsim_marine_elements): two Net/NetStructure cages of the W5 size (radius 3 m, 1 m to 6 m depth, walls only, 40 × 12 quads of two triangles: 960 panels, 520 nodes), SolidityModel Fixed, Sₙ = 0.2, default law, in a 0.5 m/s current. Cage 2 stands 20 m behind cage 1, so its rear is at Q12's point, x = 23 m behind cage 1's axis. Test/NodeTow holds every node on a spring of 10⁴ N/m to a fixed anchor (the rigid set-up of the S2 Zhan cylinder); a cage's drag is minus its tow's mean x force over the last 0.1 s of a 0.3 s RK45_i run (steady to six digits after 0.2 s), still water subtracted (0.0000 N). WakeUpdate Off with no blend and no filter: each cage builds once at t = 0, cage 1 first, so cage 2's marching build samples cage 1's field (its logged mean relative flow is 0.44 m/s instead of 0.5 m/s). CastWake off gives 1664.29 N on either cage. The Direct run of the F4 code logs a lumped M = 11.7534 m², which is 1505.90 N at ½ρU² = 128.1 Pa, the drag the tow measures to 0.01 N.

11.1 The private field at the panel centroids

W5 cage, 1000 sources. max |Δd| is GriddedWake against SourceSetWake (the direct sum, no cutoff) at the 1000 panel centroids, the points where the panels read their inflow; mean, rear is the mean |Δd| over the 500 rear-half centroids, where the direct d is 0.099 on average (0.113 at most). Build is one WakeBuilder pass. The last two columns are the farm (§ 11.3 Two cages in line): cage 1's drag, which is its self-shading through this field, against the Direct run of the same code (1505.90 N, identical before and after F4), and the logged rebuild of cage 1.

Box Grid Cutoff Build, ms (pass 1 / 2) max |Δd| mean |Δd|, rear Cage 1 drag against Direct Cage 1 rebuild, ms
pre-F4 (38 m) 96 × 48 × 48 1e-3 35 / 25 0.036 0.012 −0.97 % 27
1e-4 306 / 234 0.032 0.014 −1.78 % 236
1e-5 789 / 777 0.031 0.014 −1.85 % 672
96 × 96 × 96 1e-3 128 / 126 0.036 0.010 −1.09 % 103
1e-4 1221 / 1194 0.036 0.011 −1.69 % 989
1e-5 3244 / 2418 0.036 0.011 −1.75 % 2859
96 × 192 × 192 1e-3 508 / 500 0.054 0.012 −1.82 % 379
1e-4 3847 / 3953 0.054 0.013 −2.42 % 3640
1e-5 11025 / 11164 0.054 0.013 −2.47 % 13306
F4 (6.7 m) 96 × 48 × 48 1e-3 (default) 42 / 44 0.012 0.005 +0.36 % 47
1e-4 73 / 81 0.0088 0.0027 −0.24 % 101
1e-5 137 / 116 0.0087 0.0028 −0.30 % 111
96 × 96 × 96 1e-3 182 / 153 0.0090 0.0060 +0.64 % 159
1e-4 256 / 283 0.0028 0.0007 +0.02 % 308
1e-5 372 / 402 0.0027 0.0006 −0.04 % 412
96 × 192 × 192 1e-3 535 / 573 0.0091 0.0062 +0.68 % 610
1e-4 1033 / 1163 0.0008 0.0005 +0.04 % 1164
1e-5 1565 / 1617 0.0008 0.0002 −0.01 % 1689

What the table shows:

  1. The structure-only box is what makes the private field accurate. The pre-F4 box stretched the same cells over 38 m, 0.40 m along the flow against an x_start of 0.17 m: every pre-F4 row carries the x_start resolution warning, and the error at the centroids, 0.031–0.054, does not fall with the cutoff or the lateral cells (it grows at 96 × 192²). On the F4 box (0.07 m along the flow, no warning) the same cell counts give 0.009–0.012 at cutoff 1e-3 and down to 0.0008.
  2. Two errors again, now inside the structure. At 1e-3 the error is truncation (0.009 on every grid, the grid underestimates d: mean rear 0.093–0.094 against 0.099); below it the resolution error of the grid remains: 0.0088 at 96 × 48², 0.0028 at 96³, 0.0008 at 96 × 192². 1e-4 and 1e-5 differ by less than 0.0002 on every grid.
  3. In loads the errors are small. Cage 1's drag is within +0.36 % of Direct at the defaults and within ±0.7 % in every F4 row; before F4 it was 1–2.5 % low. The load laws themselves differ from each other and from the experiments by 7–25 % (L6, L8, S2).
  4. Cost. On the F4 box a build at 1e-3 costs a little more than on the pre-F4 box (42 against 25–35 ms: the same cells in a box a sixth as long, so each footprint covers more of them), and lowering the cutoff costs much less: 1e-5 is 2–3.3× the 1e-3 build instead of 19–31×.

11.2 What other objects see

W5 cage, 1000 sources with Sn = 0 (the pre-F4 sources). The pre-F4 registered grid is those sources on the pre-F4 box at 96 × 48 × 48; the lumped source is what F4 registers for them: M = Σ M_panel = 13.795 m² (F·ê_U / ½ρU²), frontal rectangle 6 m × 5 m = 30 m² (C_T 0.46, not widened), at the rear-face centre (3, 0, 3.5) m, wake from the rear face (x_start = 0), in a BlendedSourceSetWake (since marenv MENV-0060 SourceSetWake). The reference is the direct sum of the 1000 sources (with Sn = 0.2 it is identical to four digits at every x on the axis below). The lumped M carries no self-shading here, so all three hold the same momentum; in the farm the net's own marching build sets it.

Deficit on the cage axis (y = 0, z = 3.5 m):

x, m Direct sum Pre-F4 grid, 1e-3 Pre-F4 grid, 1e-5 Lumped Lumped − direct Pre-F4 grid 1e-3 − direct
4 0.159 0.190 0.202 0.265 +0.106 +0.031
6 0.205 0.187 0.205 0.265 +0.060 −0.018
10 0.206 0.167 0.205 0.265 +0.060 −0.038
15 0.204 0.125 0.204 0.265 +0.061 −0.079
20 0.200 0.070 0.199 0.263 +0.063 −0.130
23 (Q12) 0.195 0.036 0.194 0.225 +0.031 −0.159
30 0.176 0 0.175 0.165 −0.011 −0.176
40 0.141 0 (outside) 0 (outside) 0.114 −0.027 −0.141
60 0.087 0 (outside) 0 (outside) 0.065 −0.022 −0.087

Over whole cross-sections (y ± 9 m, z −4 to 11 m, 0.25 m spacing), the largest |Δd| against the direct sum, and ∫d dA over ± 12 m:

x, m max |Δd| lumped max |Δd| pre-F4 grid 1e-3 pre-F4 grid 1e-5 ∫d dA, m²: direct / pre-F4 grid 1e-3 / lumped
6 0.155 0.030 0.011
10 0.109 0.052 0.004 6.34 / 4.60 / 4.77
15 0.061 0.093 0.002
23 0.031 0.162 0.002 6.48 / 0.58 / 7.77
30 0.020 0.176 0.001 6.55 / 0 / 7.52

Query cost (random points in the pre-F4 box, µs per Sample(), median of 25 runs of 10⁶ queries, 2·10⁴ for the direct sum; pass 1 / pass 2):

Field other objects sample µs per Sample()
Pre-F4: GriddedWake 96 × 48 × 48, 1000 panel sources 0.032 / 0.033
F4: BlendedSourceSetWake (now SourceSetWake), one lumped source 0.035 / 0.038
Reference: SourceSetWake, 1000 panel sources 16.3 / 17.2

What the tables show:

  1. Q12's far-wake truncation is gone for everybody else. The lumped field is a direct evaluation of one source, so neither WakeCutoff nor WakeGrid enters it, and it has no box: it reaches 40 m and 60 m, where the pre-F4 grid returned nothing. At x = 23 m it gives 0.225 against the direct 0.195 (+0.031); the pre-F4 default gave 0.036 (−0.159).
  2. What remains is the lumping, not a numerical error. One source of C_T 0.46 over the frontal rectangle has the momentum-theory near-wake deficit 1 − √(1 − 0.46) = 0.265 (QF4) from the rear face to about 20 m, where the 1000 overlapping panel wakes compose to 0.20–0.21; past 30 m it is 0.01–0.03 below the direct sum. The cross-section integral follows the composition: the product Π(1 − dᵢ) of 1000 overlapping wakes holds less ∫d dA (6.3–6.6 m²) than one wake of the same M (7.5–7.8 m² at 23–30 m); ∫d(1 − d) dA is 5.7–5.9 m² for the direct sum and exactly M/2 = 6.90 m² for the lumped source at 23 m and 30 m. Which of the two is closer to a real cage's wake is a physics question the benchmark cannot answer; § 11.3 Two cages in line shows what it does to a second cage's drag.
  3. A lumped query costs what a grid query cost (0.035–0.038 against 0.032–0.033 µs) and needs no build and no memory beyond one source. A direct sum over the panels, the only exact alternative, is about 450× dearer per query.
  4. At equal momentum a second cage sees nearly the same inflow. Over the 480 panel centroids of the farm's cage 2 (20 m behind), the mean d is 0.136 (direct) and 0.135 (lumped), and the drag-weighted inflow factor Σ|cos φ|(1 − d)² / Σ|cos φ| is 0.723 and 0.719. With the lumped M reduced by the farm's ratio of lumped M to the sum of the panel M (0.905, § 11.3 Two cages in line item 2) it is 0.747, 3.2 % above the direct sum.

11.3 Two cages in line

The farm of the set-up, end to end (setup, the build at t = 0 of both cages, 0.3 s of RK45_i, 302 steps). Wall time median, min and max of 5 runs (3 for Direct), both passes; drags are deterministic.

Code Wake Cage 1 drag, N Cage 2 drag, N Cage 2 / cage 1 Wall, s (pass 1; pass 2) Rebuild cage 1 / cage 2, ms
either CastWake off 1664.29 1664.29 1 0.9–1.1 —
pre-F4 Grid 96 × 48², 1e-3 (default) 1491.22 1359.49 0.912 1.3 (1.2–1.3); 1.2 (1.2–1.3) 27 / 26
pre-F4 Grid 96 × 48², 1e-5 1478.03 1085.62 0.735 2.5 (2.5–2.8); 2.5 (2.5–2.9) 735 / 623; 654 / 668
pre-F4 Direct (reference) 1505.90 1103.06 0.733 83 (70–89); 72 (70–75) 9 / 25
F4 Grid 96 × 48², 1e-3 (default) 1511.32 1146.74 0.759 1.3 (1.3–1.5); 1.3 (1.3–1.5) 45 / 45; 46 / 48
F4 Grid 96 × 48², 1e-5 1501.41 1141.88 0.761 1.5 (1.5–1.5); 1.5 (1.5–1.8) 120 / 118; 189 / 126
F4 Direct (reference) 1505.90 1144.38 0.760 43 (43–43); 43 (42–43) 10 / 10; 11 / 14

Cage 2 against the Direct run of its own code, at every F6 grid and cutoff (farm-sweep):

Grid Cutoff Pre-F4 cage 2 F4 cage 2
96 × 48 × 48 1e-3 / 1e-4 / 1e-5 +23.3 % / +0.2 % / −1.6 % +0.21 % / −0.18 % / −0.22 %
96 × 96 × 96 1e-3 / 1e-4 / 1e-5 +23.0 % / +0.2 % / −1.6 % +0.40 % / +0.01 % / −0.03 %
96 × 192 × 192 1e-3 / 1e-4 / 1e-5 +22.1 % / −0.6 % / −2.3 % +0.43 % / +0.02 % / −0.01 %

What the tables show:

  1. The default no longer loses the upstream cage. Before F4, cage 2 at the default cutoff felt 23 % more drag than the direct sum gives, because cage 1's registered grid had truncated its wake at cage 2 (§ 2, § 11.2); only 1e-4 or lower fixed it, at 9–25× the rebuild on the default grid. Since F4, cage 2's drag is within 0.5 % of its Direct value at every grid and cutoff: the cutoff and the grid now only act inside each cage.
  2. The lumping changes the shading by 3.7 %, through the momentum it carries. Cage 2 behind cage 1 has 0.760 of cage 1's drag with the lumped wake and 0.733 with the direct sum of cage 1's 960 panels (pre-F4 Direct). The lumped M is cage 1's total force along the flow over the free-stream ½ρ|U_ref|² (design § 1.1): 11.75 m², the drag cage 1 feels with its self-shading. Each panel source's M is its drag over its own inflow's ½ρ|U|² (KF2012Common.h:369, PanelLoadTypes.h:88), so a shaded rear panel keeps its unshaded M, and the panel sources sum to about the unshaded 1664.29 N / 128.1 Pa = 12.99 m²; the lumped source carries 0.905 of that. At equal M the two fields shade a second cage alike (§ 11.2 What other objects see item 4: 0.719 against 0.723); at 0.905 of the M the lumped field gives 3.2 % more inflow factor, close to the 3.7 % in the farm. This follows from R34 (a) and the design's definition of M, not from the numerics; whether the far wake should carry the felt drag (lumped) or the panel sum is a physics question for the owner (§ 11.5 Decision table for Q12).
  3. Cost. At the default settings a farm run costs the same before and after F4 (1.2–1.3 s): the builds are milliseconds either way (27 against 45–48 ms per cage). At 1e-5 the F4 build is 5× cheaper (120–190 against 620–740 ms). Direct is half as dear (43 against 70–89 s), because each cage's panels sum over its own 960 sources and one lumped source instead of 1920.

11.4 Rebuild cost of the lumped step

The F4 rebuild adds, after the marching build: the sum of the panel forces, the frontal rectangle (every panel's three nodes projected on the three frame axes), MakeLumpedSource and one Publish into the registered BlendedSourceSetWake (now SourceSetWake). A replica of those lines on the W5 cage as triangles (2000 panels, 1040 nodes, node positions through a virtual call as in WakePanelSource), median of 5 runs of 2000 calls, three runs:

Step µs per rebuild
Lumped step (panel-force sum, frontal rectangle, MakeLumpedSource, Publish) 10.3 / 10.5 / 10.6 (min 10.2, max 11.1)

That is 0.02 % of the default rebuild of a 960-panel cage (47 ms), and for the 968-panel screen (about half the nodes) well under 1 % of its cheapest rebuild, Direct at 1.7 ms: no measurable cost, as R32 (d) asked.

Whole rebuilds, pre-F4 against F4 (rebuild-f6, the 968-panel screen of § 3, and the farm's cage 1): the Grid rebuilds change with the box (§ 11.1), not with the lumped step. Direct, where the box plays no part: the screen 1.6 / 1.6 ms before and 1.7 / 2.7 ms after (min 1.5 and 1.7 ms); cage 1 of the farm 9.0 / 9.1 ms before and 9.9 / 10.6 ms after. Part of that is the F1 near wake, which the F4 caster now switches on with each source's Sn: one ElementaryDeficit() costs 10.5–10.9 ns with Sn = 0 and 11.2–11.5 ns with Sn = 0.2 (three runs, 10⁶ pairs), +6–9 %, about 0.3 ms over the 4.6·10⁵ source–panel pairs of one Direct marching build of 960 panels. The rest is within the spread of the passes.

The flat 968-panel screen's Grid rebuild (§ 3: 25 ms at the default) is now 12–17 ms at the default and 14–32 ms at 1e-5 (pre-F4 in the same passes: 31–35 and 646–684 ms): its structure-only box is 1.4 m long.

11.5 Decision table for Q12

For the private field of a 10³-panel cage (the only place WakeCutoff and WakeGrid act since F4). Error at the panel centroids from § 11.1, drag of the self-shaded cage against Direct and rebuild time of one 960-panel cage from the farm; the error seen by other objects is the same in every row (§ 11.2, § 11.3: cage 2 within 0.5 % of Direct, lumping 3.7 % from the panel sum).

Grid Cutoff max |Δd| at centroids Cage drag against Direct Rebuild, ms Rebuild against the default
96 × 48 × 48 1e-3 (default) 0.012 +0.36 % 47 1
96 × 48 × 48 1e-4 0.0088 −0.24 % 101 2.2
96 × 48 × 48 1e-5 0.0087 −0.30 % 111 2.4
96 × 96 × 96 1e-3 0.0090 +0.64 % 159 3.4
96 × 96 × 96 1e-4 0.0028 +0.02 % 308 6.6
96 × 96 × 96 1e-5 0.0027 −0.04 % 412 8.8
96 × 192 × 192 1e-3 0.0091 +0.68 % 610 13
96 × 192 × 192 1e-4 0.0008 +0.04 % 1164 25
96 × 192 × 192 1e-5 0.0008 −0.01 % 1689 36
Direct (no grid) — 0 0 10 0.2; RHS 15× (§ 5)

For the share of a run, § 6's formula applies with these rebuild times: under Freeze (2–3 builds) every row except the 96 × 192² ones costs under 1.3 s per simulation.

Recommendation (a recommendation only; the defaults stay WakeCutoff 1e-3 and WakeGrid 96 × 48 × 48; the owner kept them, Q12, 2026-09-27):

  • **WakeCutoff 1e-4.** § 8 proposed 1e-5 because of the far-wake truncation other objects saw; that truncation no longer reaches anybody (§ 11.2, § 11.3), so 1e-5 buys nothing over 1e-4 (≤ 0.0002 on every grid). Inside a 10³-panel cage 1e-3 is already good (drag +0.36 %), but its error is truncation, which grows as N × cutoff (§ 2, item 2): for a net of 10⁴ panels (a trawl) it would be about ten times the 0.009 measured here (extrapolated, not measured), where 1e-4 keeps it near 0.009. The price is 2.2× the rebuild on the default grid, about 0.1 s for a 960-panel cage. Keeping 1e-3 is a sound choice for nets of up to about 10³ panels.
  • **WakeGrid 96 × 48 × 48, unchanged.** At 1e-4 its remaining error is resolution, 0.0088 at the centroids and −0.24 % in the drag, below W5's 0.01. 96 × 96 × 96 cuts it to 0.0028 and +0.02 % for 3× the rebuild of 96 × 48² at the same cutoff (6.6× the present default); the manuals should name it for users who need the self-shading to better than 1 %. 96 × 192² is not worth its cost.
  • **WakeEvaluation Grid, unchanged**, for the RHS cost of Direct on large nets (§ 5, § 8).
  • To decide besides (not a setting in this table): the lumped wake shades a second cage 3.7 % less than the panel sum did (§ 11.3 Two cages in line item 2), because its M is the drag the net feels (design § 1.1), about 10 % below the sum of the panel M, each taken against its own shaded inflow. Recommended: keep the design's definition (the far wake then carries exactly the momentum the net removes, ∫d(1 − d) dA = M/2), and record the difference in the methods manual.