The wake path (sched_balance_find_dst_group()) compares idle-CPU counts between the local group and the idlest group, imbalance = abs(local_sgs.idle_cpus - idlest_sgs.idle_cpus), which suffers from the same unequal-group-weight problem as the periodic load balancer: on multi-level NUMA topologies the compared groups can have different weights, so the absolute idle difference is inflated by the capacity difference even when both groups are nearly idle, and a task can be placed on a remote group on wakeup. Use the difference of busy CPUs instead, which is the normalized quantity, and a no-op when group weights are equal. Measured on the same 4-node Arm server together with the periodic-path counterpart: pair splitting 6/10 -> 4/10 runs, and ftrace shows the remaining cross-node splitting is no longer initiated by the wake path (it is executed by active balance and addressed separately). Signed-off-by: Zhang Qiao <zhangqiao22@huawei.com> --- kernel/sched/fair.c | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index 43edd1f7213a3..f845c35698b48 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -12593,7 +12593,8 @@ sched_balance_find_dst_group(struct sched_domain *sd, struct task_struct *p, int imb_numa_nr = min(w, sd->imb_numa_nr); } - imbalance = abs(local_sgs.idle_cpus - idlest_sgs.idle_cpus); + imbalance = abs((idlest_sgs.group_weight - idlest_sgs.idle_cpus) - + (local_sgs.group_weight - local_sgs.idle_cpus)); if (!adjust_numa_imbalance(imbalance, local_sgs.sum_nr_running + 1, imb_numa_nr)) { -- 2.18.0