On topologies where sched_groups inside a NUMA domain have different weights (e.g. a 4-node Arm system with 3 NUMA domain levels), the NUMA load balancer tears a pair of communicating tasks (lmbench bw_pipe -P 1) apart even on an otherwise idle system: ~760 MB/s when split vs ~1500 MB/s when the pair stays on one node. The unequal weights are a property of the distance matrix: with a non-uniform "diameter 3" NUMA topology, e.g. node 0 1 2 3 0: 10 12 35 37 1: 12 10 37 40 2: 35 37 10 12 3: 37 40 12 10 the kernel builds several NUMA sched_domain levels, and inside a level the groups are built by build_overlap_sched_groups() from the spans of the lower-level domains. Because each distance level covers a different set of nodes, some groups end up spanning two nodes while others span only one, so group_weight differs within a single domain (this is the diameter-3 case already documented in the comment above build_overlap_sched_groups()). Two defects are involved: 1. The imbalance is measured in idle-CPU differences. Between groups of different weight this counts capacity, not load, so an idle system computes a large phantom imbalance that actively splits the pair. Patches 1 and 2 fix the periodic and wake paths by comparing busy CPUs instead (no-op when weights are equal). 2. The floating imbalance tolerance is a fixed 2, which is below the busy-CPU difference the scheduler actually observes for a bare pair (2) plus kernel-thread noise (1-2): 3-4 in any balance snapshot. The split is executed by active balance, since the pair is always running and cannot be passively detached. Patch 3 caps the allowance at 4 (pair + noise) instead of scaling it with imb_numa_nr, whose uncapped use regressed large machines in 0-day testing (unixbench fstime -6.8% / fsdisk-w -26.7%) by letting independent throughput tasks accumulate. Measured on a 4-node Arm server (`lmbench bw_pipe -P 1`). The bandwidth is ~1500 MB/s when the reader/writer pair stays on one NUMA node, but drops to ~700 MB/s if split across nodes. * Upstream: Split in 8/10 runs (~760 MB/s). * Patches 1+2: Split in 4/10 runs (~1170 MB/s). However, running completely alone still resulted in 23/23 splits due to the kernel-thread noise mentioned above. * Patches 1+2+3: Zero splits (0/10 with noise, 0/12 alone). Bandwidth stabilized at ~1484 MB/s. All combinations of numa_balancing on/off and background system load present/removed were covered. On an equal-weight-group topology (2-socket x86 server) patches 1-2 are bit-identical no-ops, and patch 3 only raises the tolerance from 2 to at most 4 (versus 6-12 for the uncapped version rejected by 0-day). This is a reworked version of [1]: - the measurement fix is retained (patch 1, hunk 2 of [1]), and extended to the wake path (patch 2, new); - the threshold change (hunk 1 of [1]) is replaced by a capped allowance in patch 3: the uncapped version let large groups of independent throughput tasks accumulate, which 0-day measured as unixbench fstime -6.8% and fsdisk-w -26.7% [2]; - the dst_running change (hunk 3 of [1]) was dropped; no scenario was observed that requires it. [1] https://lore.kernel.org/all/20240524035438.2701479-1-zhangqiao22@huawei.com/ [2] https://lore.kernel.org/all/202406031516.a1956bdc-oliver.sang@intel.com/ Zhang Qiao (3): sched/numa: Use busy CPUs for imbalance with unequal group weights sched/numa: Use busy CPUs for wake-path imbalance with unequal group weights sched/numa: Cap the floating imbalance allowance at pair size kernel/sched/fair.c | 18 ++++++++++++------ 1 file changed, 12 insertions(+), 6 deletions(-) -- 2.18.0