With the imbalance measurement fixed, a bare communicating pair is still always split: tracing shows the split is executed by active balance (the pair is always running and cannot be passively detached), triggered when the real busy-CPU difference of the pair plus kernel-thread noise (3-4 busy CPUs in any balance snapshot) exceeds the fixed floating imbalance tolerance of 2. Raising the tolerance to the domain's imb_numa_nr keeps the pair fully local. However imb_numa_nr scales with topology (7-28 on large machines), and allowing an imbalance up to that size also lets large groups of independent throughput tasks accumulate on fewer nodes, which 0-day measured as unixbench fstime -6.8% and fsdisk-w -26.7% for the uncapped version of this change. The tolerance only needs to cover a communicating pair (2 busy CPUs) plus the transient kernel-thread activity visible in any balance snapshot (~1-2 more busy CPUs), so cap it at 4. On topologies where imb_numa_nr <= 4 the behavior is identical to using imb_numa_nr; on larger topologies the accumulation of independent tasks is bounded. Measured on a 4-node Arm server (lmbench bw_pipe -P 1). The bandwidth is ~1500 MB/s when the reader/writer pair stays on one NUMA node, but drops to ~700 MB/s if split across nodes. Upstream: Split in 8/10 runs (~760 MB/s). Patches 1+2: Split in 4/10 runs (~1170 MB/s). However, running completely alone still resulted in 23/23 splits. With this patch (1+2+3): Zero splits (0/10 with noise, 0/12 alone). Bandwidth stabilized at ~1484 MB/s. Tested across all combinations of numa_balancing and background loads. Signed-off-by: Zhang Qiao <zhangqiao22@huawei.com> --- kernel/sched/fair.c | 6 ++++-- 1 file changed, 4 insertions(+), 2 deletions(-) diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index f845c35698b48..3a6bbd145e7fa 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -2226,7 +2226,6 @@ static inline bool is_core_idle(int cpu) } #ifdef CONFIG_NUMA -#define NUMA_IMBALANCE_MIN 2 static inline long adjust_numa_imbalance(int imbalance, int dst_running, int imb_numa_nr) @@ -2244,8 +2243,11 @@ adjust_numa_imbalance(int imbalance, int dst_running, int imb_numa_nr) /* * Allow a small imbalance based on a simple pair of communicating * tasks that remain local when the destination is lightly loaded. + * The allowance is capped at 4 to cover the pair (2 busy CPUs) + * plus transient kernel-thread noise (1-2 busy CPUs) without + * allowing independent throughput tasks to accumulate. */ - if (imbalance <= NUMA_IMBALANCE_MIN) + if (imbalance <= min(imb_numa_nr, 4)) return 0; return imbalance; -- 2.18.0