[PATCH 0/3] sched/numa: Keep communicating task pairs local on multi-level NUMA
On topologies where sched_groups inside a NUMA domain have different weights (e.g. a 4-node Arm system with 3 NUMA domain levels), the NUMA load balancer tears a pair of communicating tasks (lmbench bw_pipe -P 1) apart even on an otherwise idle system: ~760 MB/s when split vs ~1500 MB/s when the pair stays on one node. The unequal weights are a property of the distance matrix: with a non-uniform "diameter 3" NUMA topology, e.g. node 0 1 2 3 0: 10 12 35 37 1: 12 10 37 40 2: 35 37 10 12 3: 37 40 12 10 the kernel builds several NUMA sched_domain levels, and inside a level the groups are built by build_overlap_sched_groups() from the spans of the lower-level domains. Because each distance level covers a different set of nodes, some groups end up spanning two nodes while others span only one, so group_weight differs within a single domain (this is the diameter-3 case already documented in the comment above build_overlap_sched_groups()). Two defects are involved: 1. The imbalance is measured in idle-CPU differences. Between groups of different weight this counts capacity, not load, so an idle system computes a large phantom imbalance that actively splits the pair. Patches 1 and 2 fix the periodic and wake paths by comparing busy CPUs instead (no-op when weights are equal). 2. The floating imbalance tolerance is a fixed 2, which is below the busy-CPU difference the scheduler actually observes for a bare pair (2) plus kernel-thread noise (1-2): 3-4 in any balance snapshot. The split is executed by active balance, since the pair is always running and cannot be passively detached. Patch 3 caps the allowance at 4 (pair + noise) instead of scaling it with imb_numa_nr, whose uncapped use regressed large machines in 0-day testing (unixbench fstime -6.8% / fsdisk-w -26.7%) by letting independent throughput tasks accumulate. Measured on a 4-node Arm server (`lmbench bw_pipe -P 1`). The bandwidth is ~1500 MB/s when the reader/writer pair stays on one NUMA node, but drops to ~700 MB/s if split across nodes. * Upstream: Split in 8/10 runs (~760 MB/s). * Patches 1+2: Split in 4/10 runs (~1170 MB/s). However, running completely alone still resulted in 23/23 splits due to the kernel-thread noise mentioned above. * Patches 1+2+3: Zero splits (0/10 with noise, 0/12 alone). Bandwidth stabilized at ~1484 MB/s. All combinations of numa_balancing on/off and background system load present/removed were covered. On an equal-weight-group topology (2-socket x86 server) patches 1-2 are bit-identical no-ops, and patch 3 only raises the tolerance from 2 to at most 4 (versus 6-12 for the uncapped version rejected by 0-day). This is a reworked version of [1]: - the measurement fix is retained (patch 1, hunk 2 of [1]), and extended to the wake path (patch 2, new); - the threshold change (hunk 1 of [1]) is replaced by a capped allowance in patch 3: the uncapped version let large groups of independent throughput tasks accumulate, which 0-day measured as unixbench fstime -6.8% and fsdisk-w -26.7% [2]; - the dst_running change (hunk 3 of [1]) was dropped; no scenario was observed that requires it. [1] https://lore.kernel.org/all/20240524035438.2701479-1-zhangqiao22@huawei.com/ [2] https://lore.kernel.org/all/202406031516.a1956bdc-oliver.sang@intel.com/ Zhang Qiao (3): sched/numa: Use busy CPUs for imbalance with unequal group weights sched/numa: Use busy CPUs for wake-path imbalance with unequal group weights sched/numa: Cap the floating imbalance allowance at pair size kernel/sched/fair.c | 18 ++++++++++++------ 1 file changed, 12 insertions(+), 6 deletions(-) -- 2.18.0
When calculating the imbalance in the lightly-loaded, non-overloaded case, calculate_imbalance() evens out the number of idle CPUs: env->imbalance = local->idle_cpus - busiest->idle_cpus; This is only correct when the local and busiest sched groups have the same weight. On multi-level NUMA topologies (e.g. a 4-node machine where the kernel builds several NUMA sched_domain levels) some sched_group weights differ. In that case, even when both groups are almost completely idle, the idle difference can be large simply because the local group has more CPUs, so the load balancer computes a large imbalance and pulls a pair of communicating tasks apart. Fix this by evening out the number of *busy* CPUs instead. When both groups have the same weight the busy-CPU difference reduces to the idle-CPU difference, so this is a no-op there; the change only affects groups with unequal weights, where busy CPUs is the correct normalized quantity to compare. Measured on a 4-node Arm server (3 NUMA levels, unequal group weights), lmbench bw_pipe -P 1: the communicating pair is split across NUMA nodes in 8/10 runs upstream vs 6/10 with this patch, mean bandwidth 760 -> 1075 MB/s. On an equal-weight-group topology (2-socket x86 server) this patch is a bit-identical no-op. Signed-off-by: Zhang Qiao <zhangqiao22@huawei.com> --- kernel/sched/fair.c | 9 ++++++--- 1 file changed, 6 insertions(+), 3 deletions(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 7455a83a6a990..43edd1f7213a3 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -12868,12 +12868,15 @@ static inline void calculate_imbalance(struct lb_env *env, struct sd_lb_stats *s } else { /* - * If there is no overload, we just want to even the number of - * idle CPUs. + * If there is no overload, we just want to even the + * number of busy CPUs. Busy CPUs is preferred over + * idle CPUs because local and busiest groups can have + * different weights (e.g. multi-level NUMA domains). */ env->migration_type = migrate_task; env->imbalance = max_t(long, 0, - (local->idle_cpus - busiest->idle_cpus)); + (busiest->group_weight - busiest->idle_cpus) - + (local->group_weight - local->idle_cpus)); } #ifdef CONFIG_NUMA -- 2.18.0
The wake path (sched_balance_find_dst_group()) compares idle-CPU counts between the local group and the idlest group, imbalance = abs(local_sgs.idle_cpus - idlest_sgs.idle_cpus), which suffers from the same unequal-group-weight problem as the periodic load balancer: on multi-level NUMA topologies the compared groups can have different weights, so the absolute idle difference is inflated by the capacity difference even when both groups are nearly idle, and a task can be placed on a remote group on wakeup. Use the difference of busy CPUs instead, which is the normalized quantity, and a no-op when group weights are equal. Measured on the same 4-node Arm server together with the periodic-path counterpart: pair splitting 6/10 -> 4/10 runs, and ftrace shows the remaining cross-node splitting is no longer initiated by the wake path (it is executed by active balance and addressed separately). Signed-off-by: Zhang Qiao <zhangqiao22@huawei.com> --- kernel/sched/fair.c | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 43edd1f7213a3..f845c35698b48 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -12593,7 +12593,8 @@ sched_balance_find_dst_group(struct sched_domain *sd, struct task_struct *p, int imb_numa_nr = min(w, sd->imb_numa_nr); } - imbalance = abs(local_sgs.idle_cpus - idlest_sgs.idle_cpus); + imbalance = abs((idlest_sgs.group_weight - idlest_sgs.idle_cpus) - + (local_sgs.group_weight - local_sgs.idle_cpus)); if (!adjust_numa_imbalance(imbalance, local_sgs.sum_nr_running + 1, imb_numa_nr)) { -- 2.18.0
With the imbalance measurement fixed, a bare communicating pair is still always split: tracing shows the split is executed by active balance (the pair is always running and cannot be passively detached), triggered when the real busy-CPU difference of the pair plus kernel-thread noise (3-4 busy CPUs in any balance snapshot) exceeds the fixed floating imbalance tolerance of 2. Raising the tolerance to the domain's imb_numa_nr keeps the pair fully local. However imb_numa_nr scales with topology (7-28 on large machines), and allowing an imbalance up to that size also lets large groups of independent throughput tasks accumulate on fewer nodes, which 0-day measured as unixbench fstime -6.8% and fsdisk-w -26.7% for the uncapped version of this change. The tolerance only needs to cover a communicating pair (2 busy CPUs) plus the transient kernel-thread activity visible in any balance snapshot (~1-2 more busy CPUs), so cap it at 4. On topologies where imb_numa_nr <= 4 the behavior is identical to using imb_numa_nr; on larger topologies the accumulation of independent tasks is bounded. Measured on a 4-node Arm server (lmbench bw_pipe -P 1). The bandwidth is ~1500 MB/s when the reader/writer pair stays on one NUMA node, but drops to ~700 MB/s if split across nodes. Upstream: Split in 8/10 runs (~760 MB/s). Patches 1+2: Split in 4/10 runs (~1170 MB/s). However, running completely alone still resulted in 23/23 splits. With this patch (1+2+3): Zero splits (0/10 with noise, 0/12 alone). Bandwidth stabilized at ~1484 MB/s. Tested across all combinations of numa_balancing and background loads. Signed-off-by: Zhang Qiao <zhangqiao22@huawei.com> --- kernel/sched/fair.c | 6 ++++-- 1 file changed, 4 insertions(+), 2 deletions(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index f845c35698b48..3a6bbd145e7fa 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -2226,7 +2226,6 @@ static inline bool is_core_idle(int cpu) } #ifdef CONFIG_NUMA -#define NUMA_IMBALANCE_MIN 2 static inline long adjust_numa_imbalance(int imbalance, int dst_running, int imb_numa_nr) @@ -2244,8 +2243,11 @@ adjust_numa_imbalance(int imbalance, int dst_running, int imb_numa_nr) /* * Allow a small imbalance based on a simple pair of communicating * tasks that remain local when the destination is lightly loaded. + * The allowance is capped at 4 to cover the pair (2 busy CPUs) + * plus transient kernel-thread noise (1-2 busy CPUs) without + * allowing independent throughput tasks to accumulate. */ - if (imbalance <= NUMA_IMBALANCE_MIN) + if (imbalance <= min(imb_numa_nr, 4)) return 0; return imbalance; -- 2.18.0
participants (1)
-
Zhang Qiao