When calculating the imbalance in the lightly-loaded, non-overloaded case, calculate_imbalance() evens out the number of idle CPUs: env->imbalance = local->idle_cpus - busiest->idle_cpus; This is only correct when the local and busiest sched groups have the same weight. On multi-level NUMA topologies (e.g. a 4-node machine where the kernel builds several NUMA sched_domain levels) some sched_group weights differ. In that case, even when both groups are almost completely idle, the idle difference can be large simply because the local group has more CPUs, so the load balancer computes a large imbalance and pulls a pair of communicating tasks apart. Fix this by evening out the number of *busy* CPUs instead. When both groups have the same weight the busy-CPU difference reduces to the idle-CPU difference, so this is a no-op there; the change only affects groups with unequal weights, where busy CPUs is the correct normalized quantity to compare. Measured on a 4-node Arm server (3 NUMA levels, unequal group weights), lmbench bw_pipe -P 1: the communicating pair is split across NUMA nodes in 8/10 runs upstream vs 6/10 with this patch, mean bandwidth 760 -> 1075 MB/s. On an equal-weight-group topology (2-socket x86 server) this patch is a bit-identical no-op. Signed-off-by: Zhang Qiao <zhangqiao22@huawei.com> --- kernel/sched/fair.c | 9 ++++++--- 1 file changed, 6 insertions(+), 3 deletions(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 7455a83a6a990..43edd1f7213a3 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -12868,12 +12868,15 @@ static inline void calculate_imbalance(struct lb_env *env, struct sd_lb_stats *s } else { /* - * If there is no overload, we just want to even the number of - * idle CPUs. + * If there is no overload, we just want to even the + * number of busy CPUs. Busy CPUs is preferred over + * idle CPUs because local and busiest groups can have + * different weights (e.g. multi-level NUMA domains). */ env->migration_type = migrate_task; env->imbalance = max_t(long, 0, - (local->idle_cpus - busiest->idle_cpus)); + (busiest->group_weight - busiest->idle_cpus) - + (local->group_weight - local->idle_cpus)); } #ifdef CONFIG_NUMA -- 2.18.0