|
FhSim
3.1.0
Marine systems simulation
|
| ID | 0010 |
| Class | ARCHITECTURE |
| Severity | 1 |
| Status | ready |
Models: — (CoRiBoDynamics::ConstraintSolver, CoreBoundThreadPool; consumer fhsim_marine_elements Net/NetStructureWithConstraints, MARE-0119)
Found: 2026-09-28, net-hydrodynamics deep review, worktrees/nethydro/review/deep/net-force-performance.md finding A7, at fhsim_coribo 9980b62
Decision needed: —
Related: Same defect as CORIBO-0009, reported downstream from fhsim_marine_elements MARE-0045. 0009 is the library-API bug (the one-argument constructor); this record adds the deep-review measurements and the inline-dispatch option.
src/ConstraintSolver/ConstraintSolver.cpp:7-10: the one-argument constructor ConstraintSolver(double TimeConstant) builds new CoreBoundThreadPool(std::thread::hardware_concurrency(), 0, false), one worker bound to each hardware thread, whatever OMP_NUM_THREADS, FhSim's NumCores or the SimObject's settings say. The constructors that take a thread count (:13) or a pool (:19) exist; fhsim_marine_elements NetStructureWithConstraints.cpp:167 uses the one-argument form, while its trawl objects go through CreateConstraintSolver (src/trawl_mooring_interaction/ConstraintSolverSetup.cpp:55-65), which reads NumCpuCore and runs a pool of size 0 inline.ConstraintSolver.cpp:90, :95, :150 (WriteDynamic), :243; SparseMatrixBuilder.cpp:182 (sub-multiplications), :211 and :262 (tridiagonal sub-solves), :218. CoreBoundThreadPool::ExecuteTaskSet (src/Utilities/AsynchronousTask.cpp:46-81) takes the pool mutex, queues each task and wakes the worker; WorkUnit::AddToQueue (:173-178) signals m_internal_barrier once per task and WaitUntilIdle (:191-196) blocks on m_external_barrier. Only a pool of size 0 runs the tasks inline (:58-62, :76-80).NetStructureWithConstraints_FreePanel_in.xml, 0.2 s, 200 steps, 1 200 OdeFcn): 7.1 s wall, 0.76 s user, 8.4 s sys with OMP_NUM_THREADS=1; 18.5 s wall and 502 s user with 32. callgrind with system-time collection: 58 % of system time under ConstraintSolver::ComputeDynamic → CoreBoundThreadPool::ExecuteTaskSet → std::condition_variable::notify_one / wait, 270 000 signals and 448 000 waits, about 600 futex round trips per right-hand side.Aside, same file: the affinity masks 0x1 << (i + start_ix) and 0x3 << 2*(i + start_ix) (AsynchronousTask.cpp:10, :15) are shifts of a 64-bit size_t, undefined from 64 (or 32 with the hyperthreading merge) threads; hardware_concurrency() reaches that on large servers.
Small constraint models (net cages with rigid rings, ropes: tens of elements) spend most of their time in thread hand-off rather than in the solve; the default grows worse with the core count of the machine (32 threads: 2.6× the wall time of one). This is the system time reported in MARE-0119.
CoreBoundThreadPool::ExecuteTaskSet, run a task set inline on the calling thread when it is small (a threshold on num or on the work per task), orNetStructureWithConstraints build its solver through a thread count it reads, as CreateConstraintSolver does (fhsim_marine_elements side, MARE-0119).Either keeps every task's arithmetic the same; the tasks of one set write disjoint data, so the results do not depend on which thread runs them (to be confirmed by the test below).
A benchmark of the FreePanel case: wall and system time with the default constructor at 1 and 32 hardware threads before and after, and bit-identical output of the case (the fhsim_marine_elements baseline) with the inline and the threaded pool.
Low: scheduling only. A threshold set too high loses parallel speed-up on large constraint models (long trawl warps with many elements); measure on one.