[PATCH OLK-6.6 00/29] Add support for FEAT_TLBID
This patche set is used to generate PR and AI comments. Fuad Tabba (1): arm64: sysreg: Correct sign definitions for EIESB and DoubleLock Jinjiang Tu (1): arm64: tlbflush: Optimize flush_tlb_mm() by using TLBID Jinqian Yang (14): KVM: arm64: Add internal helpers to support TLBID virtualization KVM: arm64: Update VTLBID(n) when vCPU load KVM: arm64: Add ioctls to support TLBID virtualzation KVM: arm64: Ensure DVMBM is disabled when TLBID is enabled KVM: arm64: Expose tlbididr_el1 to guest KVM: arm64: Expose ID_AA64MMFR4_EL1_TLBID to guest KVM: arm64: Enable read TLBIDIDR_EL1 trap in guest KVM: arm64: Use pCPU-to-vCPU reverse index to avoid vCPU traversal KVM: arm64: Skip TLB invalidation and vCPU IPI when domain map is unchanged KVM: arm64: Replace trace_printk with tracepoint in kvm_tlbidomain_vcpu_load KVM: arm64: Drop KABI_EXTEND_ENUM for TLBIDIDR_EL1 KVM: arm64: Move tlbidomain out of kvm_arch KVM: arm64: Guard vTLBID code with CONFIG_KVM_ARM_VTLBID KVM: arm64: Enable CONFIG_KVM_ARM_VTLBID in openeuler_defconfig Liao Chang (1): acpi: arm64: Add basic skeleton for TLBI Domains Marc Zyngier (4): arm64: cpufeature: Add ID_AA64MMFR4_EL1 handling arm64: sysreg: Add layout for ID_AA64MMFR4_EL1 arm64: cpufeatures: Add missing ID_AA64MMFR4_EL1 to __read_sysreg_by_encoding() arm64: sysreg: Update ID_AA64MMFR4_EL1 description Tian Zheng (1): KVM: arm64: Skip VM-level ID_AA64MMFR4 write on vCPU hotplug Zeng Heng (6): arm64: cpufeature: Add SYSINSTR128 (128-bit System instructions) detection arm64: cpufeature: Add TLBID (Domain-based TLB Invalidation) detection arm64: mm: Track CPUs that a task has run on for TLBID optimization arm64: tlbflush: Optimize flush_tlb_range() by using TLBIP arm64: tlbflush: Optimize flush_tlb_page() by using TLBIP arm64/defconfig: enable CONFIG_ARM64_TLBID Zhou Wang (1): arm64/sysreg: Add TLBID sysreg Documentation/virt/kvm/api.rst | 62 +++ arch/arm64/Kconfig | 29 ++ arch/arm64/configs/openeuler_defconfig | 2 + arch/arm64/include/asm/cpu.h | 2 + arch/arm64/include/asm/cpufeature.h | 19 + arch/arm64/include/asm/kvm_host.h | 55 +++ arch/arm64/include/asm/tlbflush.h | 315 +++++++++++++-- arch/arm64/include/asm/tlbidomain.h | 34 ++ arch/arm64/include/uapi/asm/kvm.h | 2 + arch/arm64/kernel/cpufeature.c | 28 ++ arch/arm64/kernel/cpuinfo.c | 4 + arch/arm64/kernel/smp.c | 3 + arch/arm64/kvm/Kconfig | 12 + arch/arm64/kvm/arm.c | 540 ++++++++++++++++++++++++- arch/arm64/kvm/hisilicon/hisi_virt.c | 9 + arch/arm64/kvm/sys_regs.c | 68 +++- arch/arm64/kvm/sys_regs.h | 6 + arch/arm64/kvm/trace_arm.h | 47 +++ arch/arm64/mm/Kconfig | 19 + arch/arm64/mm/Makefile | 1 + arch/arm64/mm/context.c | 9 +- arch/arm64/mm/test_tlbidomain.c | 176 ++++++++ arch/arm64/mm/tlbidomain.c | 371 +++++++++++++++++ arch/arm64/tools/cpucaps | 4 +- arch/arm64/tools/sysreg | 89 +++- drivers/acpi/arm64/Makefile | 1 + drivers/acpi/arm64/init.c | 2 + drivers/acpi/arm64/init.h | 1 + drivers/acpi/arm64/tlbid.c | 105 +++++ drivers/acpi/tables.c | 9 + include/acpi/actbl2.h | 29 ++ include/linux/acpi.h | 1 + include/uapi/linux/kvm.h | 14 + 33 files changed, 2023 insertions(+), 45 deletions(-) create mode 100644 arch/arm64/include/asm/tlbidomain.h create mode 100644 arch/arm64/mm/test_tlbidomain.c create mode 100644 arch/arm64/mm/tlbidomain.c create mode 100644 drivers/acpi/arm64/tlbid.c -- 2.33.0
From: Marc Zyngier <maz@kernel.org> mainline inclusion from mainline-v6.9-rc1 commit 805bb61f827997c59b92669ae8cc3740ebcf1087 category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 Reference: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?i... ---------------------------------------- Add ID_AA64MMFR4_EL1 to the list of idregs the kernel knows about, and describe the E2H0 field. Reviewed-by: Oliver Upton <oliver.upton@linux.dev> Reviewed-by: Suzuki K Poulose <suzuki.poulose@arm.com> Signed-off-by: Marc Zyngier <maz@kernel.org> Reviewed-by: Catalin Marinas <catalin.marinas@arm.com> Link: https://lore.kernel.org/r/20240122181344.258974-6-maz@kernel.org Signed-off-by: Oliver Upton <oliver.upton@linux.dev> Conflicts: arch/arm64/kernel/cpufeature.c arch/arm64/include/asm/cpu.h arch/arm64/kernel/cpuinfo.c [Fix context conflicts.] Signed-off-by: Zeng Heng <zengheng4@huawei.com> --- arch/arm64/include/asm/cpu.h | 1 + arch/arm64/kernel/cpufeature.c | 7 +++++++ arch/arm64/kernel/cpuinfo.c | 1 + 3 files changed, 9 insertions(+) diff --git a/arch/arm64/include/asm/cpu.h b/arch/arm64/include/asm/cpu.h index 40113c406d7d..56eb090ae48d 100644 --- a/arch/arm64/include/asm/cpu.h +++ b/arch/arm64/include/asm/cpu.h @@ -58,6 +58,7 @@ struct cpuinfo_arm64 { u64 reg_id_aa64mmfr1; u64 reg_id_aa64mmfr2; u64 reg_id_aa64mmfr3; + u64 reg_id_aa64mmfr4; u64 reg_id_aa64pfr0; u64 reg_id_aa64pfr1; u64 reg_id_aa64pfr2; diff --git a/arch/arm64/kernel/cpufeature.c b/arch/arm64/kernel/cpufeature.c index 40abd5ae97b0..547d0ef61035 100644 --- a/arch/arm64/kernel/cpufeature.c +++ b/arch/arm64/kernel/cpufeature.c @@ -468,6 +468,11 @@ static const struct arm64_ftr_bits ftr_id_aa64mmfr3[] = { ARM64_FTR_END, }; +static const struct arm64_ftr_bits ftr_id_aa64mmfr4[] = { + S_ARM64_FTR_BITS(FTR_HIDDEN, FTR_STRICT, FTR_LOWER_SAFE, ID_AA64MMFR4_EL1_E2H0_SHIFT, 4, 0), + ARM64_FTR_END, +}; + static const struct arm64_ftr_bits ftr_ctr[] = { ARM64_FTR_BITS(FTR_VISIBLE, FTR_STRICT, FTR_EXACT, 31, 1, 1), /* RES1 */ ARM64_FTR_BITS(FTR_VISIBLE, FTR_STRICT, FTR_LOWER_SAFE, CTR_EL0_DIC_SHIFT, 1, 1), @@ -816,6 +821,7 @@ static const struct __ftr_reg_entry { ARM64_FTR_REG_OVERRIDE(SYS_ID_AA64MMFR2_EL1, ftr_id_aa64mmfr2, &id_aa64mmfr2_override), ARM64_FTR_REG(SYS_ID_AA64MMFR3_EL1, ftr_id_aa64mmfr3), + ARM64_FTR_REG(SYS_ID_AA64MMFR4_EL1, ftr_id_aa64mmfr4), /* Op1 = 0, CRn = 1, CRm = 2 */ ARM64_FTR_REG(SYS_ZCR_EL1, ftr_zcr), @@ -1116,6 +1122,7 @@ void __init init_cpu_features(struct cpuinfo_arm64 *info) init_cpu_ftr_reg(SYS_ID_AA64MMFR1_EL1, info->reg_id_aa64mmfr1); init_cpu_ftr_reg(SYS_ID_AA64MMFR2_EL1, info->reg_id_aa64mmfr2); init_cpu_ftr_reg(SYS_ID_AA64MMFR3_EL1, info->reg_id_aa64mmfr3); + init_cpu_ftr_reg(SYS_ID_AA64MMFR4_EL1, info->reg_id_aa64mmfr4); init_cpu_ftr_reg(SYS_ID_AA64PFR0_EL1, info->reg_id_aa64pfr0); init_cpu_ftr_reg(SYS_ID_AA64PFR1_EL1, info->reg_id_aa64pfr1); init_cpu_ftr_reg(SYS_ID_AA64PFR2_EL1, info->reg_id_aa64pfr2); diff --git a/arch/arm64/kernel/cpuinfo.c b/arch/arm64/kernel/cpuinfo.c index 492cbb4ff874..76d2eaeb33e4 100644 --- a/arch/arm64/kernel/cpuinfo.c +++ b/arch/arm64/kernel/cpuinfo.c @@ -333,6 +333,7 @@ static void __cpuinfo_store_cpu(struct cpuinfo_arm64 *info) info->reg_id_aa64mmfr1 = read_cpuid(ID_AA64MMFR1_EL1); info->reg_id_aa64mmfr2 = read_cpuid(ID_AA64MMFR2_EL1); info->reg_id_aa64mmfr3 = read_cpuid(ID_AA64MMFR3_EL1); + info->reg_id_aa64mmfr4 = read_cpuid(ID_AA64MMFR4_EL1); info->reg_id_aa64pfr0 = read_cpuid(ID_AA64PFR0_EL1); info->reg_id_aa64pfr1 = read_cpuid(ID_AA64PFR1_EL1); info->reg_id_aa64pfr2 = read_cpuid(ID_AA64PFR2_EL1); -- 2.33.0
From: Marc Zyngier <maz@kernel.org> mainline inclusion from mainline-v6.9-rc1 commit cfc680bb04c54e61faa51a34d8383a0aa25b583f category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 Reference: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?i... ---------------------------------------- ARMv9.5 has infroduced ID_AA64MMFR4_EL1 with a bunch of new features. Add the corresponding layout. This is extracted from the public ARM SysReg_xml_A_profile-2023-09 delivery, timestamped d55f5af8e09052abe92a02adf820deea2eaed717. Reviewed-by: Suzuki K Poulose <suzuki.poulose@arm.com> Signed-off-by: Marc Zyngier <maz@kernel.org> Reviewed-by: Catalin Marinas <catalin.marinas@arm.com> Reviewed-by: Miguel Luis <miguel.luis@oracle.com> Link: https://lore.kernel.org/r/20240122181344.258974-5-maz@kernel.org Signed-off-by: Oliver Upton <oliver.upton@linux.dev> Signed-off-by: Zeng Heng <zengheng4@huawei.com> --- arch/arm64/tools/sysreg | 37 +++++++++++++++++++++++++++++++++++++ 1 file changed, 37 insertions(+) diff --git a/arch/arm64/tools/sysreg b/arch/arm64/tools/sysreg index 68f756a2e39d..74d36b81abd3 100644 --- a/arch/arm64/tools/sysreg +++ b/arch/arm64/tools/sysreg @@ -2289,6 +2289,43 @@ UnsignedEnum 3:0 TCRX EndEnum EndSysreg +Sysreg ID_AA64MMFR4_EL1 3 0 0 7 4 +Res0 63:40 +UnsignedEnum 39:36 E3DSE + 0b0000 NI + 0b0001 IMP +EndEnum +Res0 35:28 +SignedEnum 27:24 E2H0 + 0b0000 IMP + 0b1110 NI_NV1 + 0b1111 NI +EndEnum +UnsignedEnum 23:20 NV_frac + 0b0000 NV_NV2 + 0b0001 NV2_ONLY +EndEnum +UnsignedEnum 19:16 FGWTE3 + 0b0000 NI + 0b0001 IMP +EndEnum +UnsignedEnum 15:12 HACDBS + 0b0000 NI + 0b0001 IMP +EndEnum +UnsignedEnum 11:8 ASID2 + 0b0000 NI + 0b0001 IMP +EndEnum +SignedEnum 7:4 EIESB + 0b0000 NI + 0b0001 ToEL3 + 0b0010 ToELx + 0b1111 ANY +EndEnum +Res0 3:0 +EndSysreg + Sysreg SCTLR_EL1 3 0 1 0 0 Field 63 TIDCP Field 62 SPINTMASK -- 2.33.0
From: Marc Zyngier <maz@kernel.org> mainline inclusion from mainline-v6.9-rc1 commit 87b8cf2387c5ee79576988b2e72b84eeb92c57ec category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 Reference: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?i... ---------------------------------------- When triggering a CPU hotplug scenario, we reparse the CPU feature with SCOPE_LOCAL_CPU, for which we use __read_sysreg_by_encoding() to get the HW value for this CPU. As it turns out, we're missing the handling for ID_AA64MMFR4_EL1, and trigger a BUG(). Funnily enough, Marek isn't completely happy about that. Add the damn register to the list. Fixes: 805bb61f8279 ("arm64: cpufeature: Add ID_AA64MMFR4_EL1 handling") Reported-by: Marek Szyprowski <m.szyprowski@samsung.com> Tested-by: Marek Szyprowski <m.szyprowski@samsung.com> Signed-off-by: Marc Zyngier <maz@kernel.org> Link: https://lore.kernel.org/r/20240212144736.1933112-2-maz@kernel.org Signed-off-by: Oliver Upton <oliver.upton@linux.dev> Conflicts: arch/arm64/kernel/cpufeature.c [Fix context conflicts.] Signed-off-by: Zeng Heng <zengheng4@huawei.com> --- arch/arm64/kernel/cpufeature.c | 1 + 1 file changed, 1 insertion(+) diff --git a/arch/arm64/kernel/cpufeature.c b/arch/arm64/kernel/cpufeature.c index 547d0ef61035..ca1c45d50f8b 100644 --- a/arch/arm64/kernel/cpufeature.c +++ b/arch/arm64/kernel/cpufeature.c @@ -1523,6 +1523,7 @@ u64 __read_sysreg_by_encoding(u32 sys_id) read_sysreg_case(SYS_ID_AA64MMFR1_EL1); read_sysreg_case(SYS_ID_AA64MMFR2_EL1); read_sysreg_case(SYS_ID_AA64MMFR3_EL1); + read_sysreg_case(SYS_ID_AA64MMFR4_EL1); read_sysreg_case(SYS_ID_AA64ISAR0_EL1); read_sysreg_case(SYS_ID_AA64ISAR1_EL1); read_sysreg_case(SYS_ID_AA64ISAR2_EL1); -- 2.33.0
From: Marc Zyngier <maz@kernel.org> mainline inclusion from mainline-v6.16-rc1 commit eef33835bf6f297faa222f48bf941d57d2f8bda0 category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 Reference: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?i... ---------------------------------------- Resync the ID_AA64MMFR4_EL1 with the architectue description. This results in: - the new PoPS field - the new NV2P1 value for the NV_frac field - the new RMEGDI field - the new SRMASK field These fields have been generated from the reference JSON file. Reviewed-by: Joey Gouly <joey.gouly@arm.com> Signed-off-by: Marc Zyngier <maz@kernel.org> Signed-off-by: Zeng Heng <zengheng4@huawei.com> --- arch/arm64/tools/sysreg | 19 ++++++++++++++++--- 1 file changed, 16 insertions(+), 3 deletions(-) diff --git a/arch/arm64/tools/sysreg b/arch/arm64/tools/sysreg index 74d36b81abd3..f6362fe8e4a3 100644 --- a/arch/arm64/tools/sysreg +++ b/arch/arm64/tools/sysreg @@ -2290,12 +2290,21 @@ EndEnum EndSysreg Sysreg ID_AA64MMFR4_EL1 3 0 0 7 4 -Res0 63:40 +Res0 63:48 +UnsignedEnum 47:44 SRMASK + 0b0000 NI + 0b0001 IMP +EndEnum +Res0 43:40 UnsignedEnum 39:36 E3DSE 0b0000 NI 0b0001 IMP EndEnum -Res0 35:28 +Res0 35:32 +UnsignedEnum 31:28 RMEGDI + 0b0000 NI + 0b0001 IMP +EndEnum SignedEnum 27:24 E2H0 0b0000 IMP 0b1110 NI_NV1 @@ -2304,6 +2313,7 @@ EndEnum UnsignedEnum 23:20 NV_frac 0b0000 NV_NV2 0b0001 NV2_ONLY + 0b0010 NV2P1 EndEnum UnsignedEnum 19:16 FGWTE3 0b0000 NI @@ -2323,7 +2333,10 @@ SignedEnum 7:4 EIESB 0b0010 ToELx 0b1111 ANY EndEnum -Res0 3:0 +UnsignedEnum 3:0 PoPS + 0b0000 NI + 0b0001 IMP +EndEnum EndSysreg Sysreg SCTLR_EL1 3 0 1 0 0 -- 2.33.0
From: Fuad Tabba <tabba@google.com> mainline inclusion from mainline-v6.18-rc1 commit f4d4ebc84995178273740f3e601e97fdefc561d2 category: bugfix bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 Reference: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?i... ---------------------------------------- The `ID_AA64MMFR4_EL1.EIESB` field, is an unsigned enumeration, but was incorrectly defined as a `SignedEnum` when introduced in commit cfc680bb04c5 ("arm64: sysreg: Add layout for ID_AA64MMFR4_EL1"). This is corrected to `UnsignedEnum`. Conversely, the `ID_AA64DFR0_EL1.DoubleLock` field, is a signed enumeration, but was incorrectly defined as an `UnsignedEnum`. This is corrected to `SignedEnum`, which wasn't correctly set when annotated as such in commit ad16d4cf0b4f ("arm64/sysreg: Initial unsigned annotations for ID registers"). Signed-off-by: Fuad Tabba <tabba@google.com> Acked-by: Mark Rutland <mark.rutland@arm.com> Signed-off-by: Will Deacon <will@kernel.org> Signed-off-by: Zeng Heng <zengheng4@huawei.com> --- arch/arm64/tools/sysreg | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/arch/arm64/tools/sysreg b/arch/arm64/tools/sysreg index f6362fe8e4a3..b77ebdc4e776 100644 --- a/arch/arm64/tools/sysreg +++ b/arch/arm64/tools/sysreg @@ -1679,7 +1679,7 @@ UnsignedEnum 43:40 TraceFilt 0b0000 NI 0b0001 IMP EndEnum -UnsignedEnum 39:36 DoubleLock +SignedEnum 39:36 DoubleLock 0b0000 IMP 0b1111 NI EndEnum @@ -2327,7 +2327,7 @@ UnsignedEnum 11:8 ASID2 0b0000 NI 0b0001 IMP EndEnum -SignedEnum 7:4 EIESB +UnsignedEnum 7:4 EIESB 0b0000 NI 0b0001 ToEL3 0b0010 ToELx -- 2.33.0
From: Zeng Heng <zengheng4@huawei.com> hulk inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- Add detection for FEAT_SYSINSTR128, the 128-bit System instructions extension, from ID_AA64ISAR2_EL1.SYSINSTR_128. The extension provides the 128-bit form of the System instruction encoding space, to which the SYSP instructions and their TLBIP aliases belong. The feature is detected at runtime, hidden from userspace, and remains disabled when the system does not implement it. Signed-off-by: Zeng Heng <zengheng4@huawei.com> --- arch/arm64/include/asm/cpufeature.h | 5 +++++ arch/arm64/kernel/cpufeature.c | 8 ++++++++ arch/arm64/tools/cpucaps | 2 +- 3 files changed, 14 insertions(+), 1 deletion(-) diff --git a/arch/arm64/include/asm/cpufeature.h b/arch/arm64/include/asm/cpufeature.h index 234533debc05..714f3db48f9f 100644 --- a/arch/arm64/include/asm/cpufeature.h +++ b/arch/arm64/include/asm/cpufeature.h @@ -922,6 +922,11 @@ static inline bool system_supports_tlb_range(void) cpus_have_const_cap(ARM64_HAS_TLB_RANGE); } +static inline bool system_supports_sysinstr128(void) +{ + return cpus_have_const_cap(ARM64_HAS_SYSINSTR128); +} + static inline bool cpus_support_mpam(void) { return IS_ENABLED(CONFIG_ARM64_MPAM) && diff --git a/arch/arm64/kernel/cpufeature.c b/arch/arm64/kernel/cpufeature.c index ca1c45d50f8b..13df5120dfe1 100644 --- a/arch/arm64/kernel/cpufeature.c +++ b/arch/arm64/kernel/cpufeature.c @@ -231,6 +231,7 @@ static const struct arm64_ftr_bits ftr_id_aa64isar2[] = { ARM64_FTR_BITS(FTR_VISIBLE, FTR_NONSTRICT, FTR_LOWER_SAFE, ID_AA64ISAR2_EL1_LUT_SHIFT, 4, 0), ARM64_FTR_BITS(FTR_VISIBLE, FTR_NONSTRICT, FTR_LOWER_SAFE, ID_AA64ISAR2_EL1_CSSC_SHIFT, 4, 0), ARM64_FTR_BITS(FTR_VISIBLE, FTR_NONSTRICT, FTR_LOWER_SAFE, ID_AA64ISAR2_EL1_RPRFM_SHIFT, 4, 0), + ARM64_FTR_BITS(FTR_HIDDEN, FTR_STRICT, FTR_LOWER_SAFE, ID_AA64ISAR2_EL1_SYSINSTR_128_SHIFT, 4, 0), ARM64_FTR_BITS(FTR_HIDDEN, FTR_STRICT, FTR_LOWER_SAFE, ID_AA64ISAR2_EL1_CLRBHB_SHIFT, 4, 0), ARM64_FTR_BITS(FTR_VISIBLE, FTR_STRICT, FTR_LOWER_SAFE, ID_AA64ISAR2_EL1_BC_SHIFT, 4, 0), ARM64_FTR_BITS(FTR_VISIBLE, FTR_STRICT, FTR_LOWER_SAFE, ID_AA64ISAR2_EL1_MOPS_SHIFT, 4, 0), @@ -3334,6 +3335,13 @@ static const struct arm64_cpu_capabilities arm64_features[] = { .matches = has_vip_smt_support, }, #endif + { + .desc = "128-bit System Instructions", + .capability = ARM64_HAS_SYSINSTR128, + .type = ARM64_CPUCAP_SYSTEM_FEATURE, + .matches = has_cpuid_feature, + ARM64_CPUID_FIELDS(ID_AA64ISAR2_EL1, SYSINSTR_128, IMP) + }, {}, }; diff --git a/arch/arm64/tools/cpucaps b/arch/arm64/tools/cpucaps index 8b5b9944884c..7e1c20c32243 100644 --- a/arch/arm64/tools/cpucaps +++ b/arch/arm64/tools/cpucaps @@ -119,7 +119,7 @@ HAS_COPY_OPT HAS_LSUI HAS_FPMR HAS_BBML3 -KABI_RESERVE_12 +HAS_SYSINSTR128 KABI_RESERVE_13 KABI_RESERVE_14 KABI_RESERVE_15 -- 2.33.0
From: Zeng Heng <zengheng4@huawei.com> hulk inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- Add detection for the ARM64 Domain-based TLB Invalidation (TLBID) feature as defined in ARMv9.3-A architecture. TLBID allows TLB invalidation operations to be scoped to a specific domain, avoiding the need for global TLB invalidation broadcasts across the system. This reduces cache coherency traffic and improves performance in virtualization scenarios. Signed-off-by: Zeng Heng <zengheng4@huawei.com> --- arch/arm64/Kconfig | 20 ++++++++++++++++++++ arch/arm64/include/asm/cpufeature.h | 7 +++++++ arch/arm64/kernel/cpufeature.c | 10 ++++++++++ arch/arm64/tools/cpucaps | 2 +- arch/arm64/tools/sysreg | 5 ++++- 5 files changed, 42 insertions(+), 2 deletions(-) diff --git a/arch/arm64/Kconfig b/arch/arm64/Kconfig index 736908afc98a..fe3fa113e5cd 100644 --- a/arch/arm64/Kconfig +++ b/arch/arm64/Kconfig @@ -2303,6 +2303,26 @@ config ARM64_TLB_RANGE The feature introduces new assembly instructions, and they were support when binutils >= 2.30. +config ARM64_TLBID + bool "Enable support for TLBID (TLBI Domains)" + default y + help + TLBI broadcasting all PEs introduces performance noise. By + combining hardware and software, TLBID (TLBI Domain) limit TLBI + to an appropriate scope, avoiding the performance overhead caused + by broadcasting. + + Arm Architecture defines the FEAT_TLBID extension since Armv9.3, + when present, allows the OSPM to target TLB invalidations to + a specific domain. A TLBI domain can be contained in multiple + parent domains. + + The TLBI table lists all the PE and SMMU components and the TLBI + Domains that each component belongs to. + + The feature is detected at runtime, and will remain disabled + if the system does not implement the feature. + config ARM64_MPAM bool "Enable support for MPAM" select ACPI_MPAM if ACPI diff --git a/arch/arm64/include/asm/cpufeature.h b/arch/arm64/include/asm/cpufeature.h index 714f3db48f9f..9e617d40c50b 100644 --- a/arch/arm64/include/asm/cpufeature.h +++ b/arch/arm64/include/asm/cpufeature.h @@ -927,6 +927,13 @@ static inline bool system_supports_sysinstr128(void) return cpus_have_const_cap(ARM64_HAS_SYSINSTR128); } +static inline bool system_supports_tlbid(void) +{ + return IS_ENABLED(CONFIG_ARM64_TLBID) && + system_supports_sysinstr128() && + cpus_have_const_cap(ARM64_HAS_TLBID); +} + static inline bool cpus_support_mpam(void) { return IS_ENABLED(CONFIG_ARM64_MPAM) && diff --git a/arch/arm64/kernel/cpufeature.c b/arch/arm64/kernel/cpufeature.c index 13df5120dfe1..82468372fdd4 100644 --- a/arch/arm64/kernel/cpufeature.c +++ b/arch/arm64/kernel/cpufeature.c @@ -470,6 +470,7 @@ static const struct arm64_ftr_bits ftr_id_aa64mmfr3[] = { }; static const struct arm64_ftr_bits ftr_id_aa64mmfr4[] = { + ARM64_FTR_BITS(FTR_HIDDEN, FTR_STRICT, FTR_LOWER_SAFE, ID_AA64MMFR4_EL1_TLBID_SHIFT, 4, 0), S_ARM64_FTR_BITS(FTR_HIDDEN, FTR_STRICT, FTR_LOWER_SAFE, ID_AA64MMFR4_EL1_E2H0_SHIFT, 4, 0), ARM64_FTR_END, }; @@ -3342,6 +3343,15 @@ static const struct arm64_cpu_capabilities arm64_features[] = { .matches = has_cpuid_feature, ARM64_CPUID_FIELDS(ID_AA64ISAR2_EL1, SYSINSTR_128, IMP) }, +#ifdef CONFIG_ARM64_TLBID + { + .desc = "TLBI Domains", + .capability = ARM64_HAS_TLBID, + .type = ARM64_CPUCAP_SYSTEM_FEATURE, + .matches = has_cpuid_feature, + ARM64_CPUID_FIELDS(ID_AA64MMFR4_EL1, TLBID, IMP) + }, +#endif {}, }; diff --git a/arch/arm64/tools/cpucaps b/arch/arm64/tools/cpucaps index 7e1c20c32243..c0409ba0dcca 100644 --- a/arch/arm64/tools/cpucaps +++ b/arch/arm64/tools/cpucaps @@ -120,7 +120,7 @@ HAS_LSUI HAS_FPMR HAS_BBML3 HAS_SYSINSTR128 -KABI_RESERVE_13 +HAS_TLBID KABI_RESERVE_14 KABI_RESERVE_15 KABI_RESERVE_16 diff --git a/arch/arm64/tools/sysreg b/arch/arm64/tools/sysreg index b77ebdc4e776..1e8dee638509 100644 --- a/arch/arm64/tools/sysreg +++ b/arch/arm64/tools/sysreg @@ -2295,7 +2295,10 @@ UnsignedEnum 47:44 SRMASK 0b0000 NI 0b0001 IMP EndEnum -Res0 43:40 +UnsignedEnum 43:40 TLBID + 0b0000 NI + 0b0001 IMP +EndEnum UnsignedEnum 39:36 E3DSE 0b0000 NI 0b0001 IMP -- 2.33.0
From: Liao Chang <liaochang1@huawei.com> hulk inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- Add the initial infrastructure to support the Armv9.3 FEAT_TLBID extension, which allows OSPM to target TLB invalidations to a specific domain. A TLBI domain can be contained in multiple parent domains. This patch provides: - Kconfig option CONFIG_ARM64_TLBID to enable the feature - TLBIDIDR_EL1 register definition and CPU feature detection - ACPI TLBI table parser to discover TLBI domains for processors and SMMU [1]. - Data structures to manage TLBI domain to CPU mapping. - Help functions to choose the best TLBI domain for a specific set of CPUs. The TLBI table lists all the PE and SMMU components and the TLBI domains that each component belongs to. This information is used to optimize TLB invalidation by targeting specific domains instead of broadcasting to all CPUs. [1] https://github.com/user-attachments/files/25049127/ACPI_TLBID_v3.docx Signed-off-by: Liao Chang <liaochang1@huawei.com> Signed-off-by: Xinyu Zheng <zhengxinyu6@huawei.com> CC: Xia Qinxin <xiaqinxin@huawei.com> CC: Zheng Tian <zhengtian10@huawei.com> CC: Yang Jinqian <yangjinqian1@huawei.com> CC: Zeng Heng <zengheng4@huawei.com> CC: Wang Kefeng <wangkefeng.wang@huawei.com> CC: Yu Jiacheng <yujiacheng3@huawei.com> --- arch/arm64/include/asm/cpu.h | 1 + arch/arm64/include/asm/cpufeature.h | 7 + arch/arm64/include/asm/tlbidomain.h | 34 +++ arch/arm64/kernel/cpufeature.c | 2 + arch/arm64/kernel/cpuinfo.c | 3 + arch/arm64/mm/Kconfig | 19 ++ arch/arm64/mm/Makefile | 1 + arch/arm64/mm/test_tlbidomain.c | 176 +++++++++++++ arch/arm64/mm/tlbidomain.c | 371 ++++++++++++++++++++++++++++ arch/arm64/tools/sysreg | 11 + drivers/acpi/arm64/Makefile | 1 + drivers/acpi/arm64/init.c | 2 + drivers/acpi/arm64/init.h | 1 + drivers/acpi/arm64/tlbid.c | 105 ++++++++ drivers/acpi/tables.c | 9 + include/acpi/actbl2.h | 29 +++ include/linux/acpi.h | 1 + 17 files changed, 773 insertions(+) create mode 100644 arch/arm64/include/asm/tlbidomain.h create mode 100644 arch/arm64/mm/test_tlbidomain.c create mode 100644 arch/arm64/mm/tlbidomain.c create mode 100644 drivers/acpi/arm64/tlbid.c diff --git a/arch/arm64/include/asm/cpu.h b/arch/arm64/include/asm/cpu.h index 56eb090ae48d..ede40c40d410 100644 --- a/arch/arm64/include/asm/cpu.h +++ b/arch/arm64/include/asm/cpu.h @@ -47,6 +47,7 @@ struct cpuinfo_arm64 { u64 reg_gmid; u64 reg_smidr; u64 reg_mpamidr; + u64 reg_tlbididr; u64 reg_id_aa64dfr0; u64 reg_id_aa64dfr1; diff --git a/arch/arm64/include/asm/cpufeature.h b/arch/arm64/include/asm/cpufeature.h index 9e617d40c50b..7a0c8981120b 100644 --- a/arch/arm64/include/asm/cpufeature.h +++ b/arch/arm64/include/asm/cpufeature.h @@ -666,6 +666,13 @@ static inline bool id_aa64pfr1_mte(u64 pfr1) return val >= ID_AA64PFR1_EL1_MTE_MTE2; } +static inline bool id_aa64mmfr4_tlbid(u64 mmfr4) +{ + u32 val = cpuid_feature_extract_unsigned_field(mmfr4, ID_AA64MMFR4_EL1_TLBID_SHIFT); + + return val > 0; +} + void __init setup_cpu_features(void); void check_local_cpu_capabilities(void); diff --git a/arch/arm64/include/asm/tlbidomain.h b/arch/arm64/include/asm/tlbidomain.h new file mode 100644 index 000000000000..1dac3771ff57 --- /dev/null +++ b/arch/arm64/include/asm/tlbidomain.h @@ -0,0 +1,34 @@ +/* SPDX-License-Identifier: GPL-2.0-only */ +/* + * Copyright (C) 2026 Huawei Ltd. + */ +#ifndef __ASM_TLBIDOMAIN_H +#define __ASM_TLBIDOMAIN_H + +#include <linux/types.h> + +struct tlbi_target { + u32 identifier; + s32 *domains; + s32 nr_domain; + struct list_head list; +}; + +#define TLBI_RT_DOMAIN_WIDTH (16) +#define TLBI_INV_DOMAIN (-1) + +#ifdef CONFIG_ARM64_TLBID +void add_target(int type, struct tlbi_target *target); +int pick_best_domain(const cpumask_t *active_cpus); +int __init parse_tlbid_topology(void); +int __init tlbid_data_init(void); +void __init tlbid_data_reset(void); +void __init tlbid_early_init(void); + +#else +static inline int pick_best_domain(const cpumask_t *active_cpus) { return 0; } +static inline void __init tlbid_early_init(void) { } + +#endif /* ! CONFIG_ARM64_TLBID */ +#endif + diff --git a/arch/arm64/kernel/cpufeature.c b/arch/arm64/kernel/cpufeature.c index 82468372fdd4..1491c9273007 100644 --- a/arch/arm64/kernel/cpufeature.c +++ b/arch/arm64/kernel/cpufeature.c @@ -97,6 +97,7 @@ #include <asm/vectors.h> #include <asm/virt.h> #include <asm/vip_smt.h> +#include <asm/tlbidomain.h> /* Kernel representation of AT_HWCAP and AT_HWCAP2 */ static DECLARE_BITMAP(elf_hwcap, MAX_CPU_FEATURES) __read_mostly; @@ -4107,6 +4108,7 @@ void __init setup_cpu_features(void) sve_setup(); sme_setup(); + tlbid_early_init(); minsigstksz_setup(); mpam_extra_caps(); diff --git a/arch/arm64/kernel/cpuinfo.c b/arch/arm64/kernel/cpuinfo.c index 76d2eaeb33e4..39ae0bf1360a 100644 --- a/arch/arm64/kernel/cpuinfo.c +++ b/arch/arm64/kernel/cpuinfo.c @@ -347,6 +347,9 @@ static void __cpuinfo_store_cpu(struct cpuinfo_arm64 *info) if (id_aa64pfr0_32bit_el0(info->reg_id_aa64pfr0)) __cpuinfo_store_cpu_32bit(&info->aarch32); + if (id_aa64mmfr4_tlbid(info->reg_id_aa64mmfr4)) + info->reg_tlbididr = read_cpuid(TLBIDIDR_EL1); + if (IS_ENABLED(CONFIG_ARM64_MPAM) && mpam_detect_is_enabled() && (id_aa64pfr0_mpam(info->reg_id_aa64pfr0) || diff --git a/arch/arm64/mm/Kconfig b/arch/arm64/mm/Kconfig index 315b3322feb5..eeb79a30b2e9 100644 --- a/arch/arm64/mm/Kconfig +++ b/arch/arm64/mm/Kconfig @@ -14,3 +14,22 @@ config PFN_RANGE_ALLOC the linear mapping granule of this range is never larger than PMD. If unsure, say N. + +config TLBID_KUNIT_TEST + bool "KUnit tests for TLBI" if !KUNIT_ALL_TESTS + depends on KUNIT=y + depends on ARM64_TLBID && ACPI + default KUNIT_ALL_TESTS + help + This option enables KUnit unit tests for the arm64 Translation + Lookaside Buffer Domain topology parsing and domain picking + logic. + + It verifies the correct parsing of TLBID ACPI tables and ensures + proper selection of TLB invalidation domains during kernel boot. + + Say Y here if you want to run these unit tests as part of the + kernel test suite. Say N if you are building a production kernel + or do not need boot-time KUnit testing. + + If unsure, say N. diff --git a/arch/arm64/mm/Makefile b/arch/arm64/mm/Makefile index c02aeb729717..4d92d42826bf 100644 --- a/arch/arm64/mm/Makefile +++ b/arch/arm64/mm/Makefile @@ -11,6 +11,7 @@ obj-$(CONFIG_TRANS_TABLE) += trans_pgd.o obj-$(CONFIG_TRANS_TABLE) += trans_pgd-asm.o obj-$(CONFIG_DEBUG_VIRTUAL) += physaddr.o obj-$(CONFIG_ARM64_MTE) += mteswap.o +obj-$(CONFIG_ARM64_TLBID) += tlbidomain.o KASAN_SANITIZE_physaddr.o += n obj-$(CONFIG_KASAN) += kasan_init.o diff --git a/arch/arm64/mm/test_tlbidomain.c b/arch/arm64/mm/test_tlbidomain.c new file mode 100644 index 000000000000..652ecff2079c --- /dev/null +++ b/arch/arm64/mm/test_tlbidomain.c @@ -0,0 +1,176 @@ +// SPDX-License-Identifier: GPL-2.0-only +/* This file is intended to be included into tlbidomain.c + * Copyright (C) 2026 Huawei Ltd. + */ + +#include <kunit/test.h> + +static struct tlbid_data fake_tlbid_data __read_mostly = { + .nr_domain = 1, + .nis = TLBI_RT_DOMAIN_WIDTH, + .nvis = TLBI_RT_DOMAIN_WIDTH, + .nos = TLBI_RT_DOMAIN_WIDTH, + .nvos = TLBI_RT_DOMAIN_WIDTH, +}; + +static struct acpi_madt_generic_interrupt fake_madt_gicc[NR_CPUS]; + +static void __init reset_real_tlbid(void) +{ + tlbid_data_reset(); + memcpy(&tlbid_data, &fake_tlbid_data, sizeof(tlbid_data)); +} + +static int __init setup_fake_tlbid(struct kunit *test) +{ + int cpu, max_domains = 1 << TLBI_RT_DOMAIN_WIDTH; + struct acpi_madt_generic_interrupt *gicc; + struct tlbi_target *target; + + if (num_possible_cpus() == 1) + return 1; + + memcpy(&fake_tlbid_data, &tlbid_data, sizeof(tlbid_data)); + tlbid_data.nis = 8; + tlbid_data.nvis = 0; + tlbid_data.nos = 8; + tlbid_data.nvos = 0; + for_each_possible_cpu(cpu) + memcpy(&fake_madt_gicc[cpu], acpi_cpu_get_madt_gicc(cpu), + sizeof(struct acpi_madt_generic_interrupt)); + + tlbid_data.found_domains = kcalloc(BITS_TO_LONGS(max_domains), + sizeof(long), GFP_KERNEL); + KUNIT_ASSERT_NOT_ERR_OR_NULL(test, tlbid_data.found_domains); + + INIT_LIST_HEAD(&tlbid_data.target_list[ACPI_TLBI_TYPE_PROCESSOR]); + INIT_LIST_HEAD(&tlbid_data.target_list[ACPI_TLBI_TYPE_SMMU]); + + /* + * Fake TLBI domain topology: + * domain0: [0, num_possible_cpus()) + * domain1: [0, (num_possible_cpus() / 2) + * domain2: [num_possible_cpus() / 2, num_possible_cpus()) + */ + for_each_possible_cpu(cpu) { + target = kzalloc(sizeof(*target), GFP_KERNEL); + KUNIT_ASSERT_NOT_ERR_OR_NULL(test, target); + + /* The ACPI Processor UID of the corresponding GICC MADT entry */ + gicc = acpi_cpu_get_madt_gicc(cpu); + gicc->uid = cpu; + target->nr_domain = 2; + target->identifier = cpu; + target->domains = kcalloc(target->nr_domain, + sizeof(*target->domains), GFP_KERNEL); + KUNIT_ASSERT_NOT_ERR_OR_NULL(test, target->domains); + target->domains[0] = 0; + target->domains[1] = (cpu < (num_possible_cpus() / 2)) ? 1 : 2; + add_target(ACPI_TLBI_TYPE_PROCESSOR, target); + } + + parse_tlbid_topology(); + + for_each_possible_cpu(cpu) + memcpy(acpi_cpu_get_madt_gicc(cpu), &fake_madt_gicc[cpu], + sizeof(struct acpi_madt_generic_interrupt)); + + return 0; +} + +static void __init test_parse_tlbid_topology(struct kunit *test) +{ + struct tlbi_domain *actbl[3]; + cpumask_t tmp; + int cpu; + + if (setup_fake_tlbid(test)) + return; + + actbl[0] = &tlbid_data.cpu_domain[num_possible_cpus() - 1][0]; + actbl[1] = &tlbid_data.cpu_domain[(num_possible_cpus() / 2) - 1][0]; + actbl[2] = &tlbid_data.cpu_domain[(num_possible_cpus() / 2) - 1][1]; + + KUNIT_EXPECT_EQ(test, actbl[0]->id, 0); + KUNIT_EXPECT_TRUE(test, cpumask_equal(cpu_possible_mask, + &actbl[0]->cpumask)); + + KUNIT_EXPECT_EQ(test, actbl[1]->id, 1); + cpumask_clear(&tmp); + for_each_possible_cpu(cpu) { + if (cpu >= (num_possible_cpus() / 2)) + continue; + cpumask_set_cpu(cpu, &tmp); + } + KUNIT_EXPECT_TRUE(test, cpumask_equal(&tmp, &actbl[1]->cpumask)); + + KUNIT_EXPECT_EQ(test, actbl[2]->id, 2); + cpumask_clear(&tmp); + for_each_possible_cpu(cpu) { + if (cpu < (num_possible_cpus() / 2)) + continue; + cpumask_set_cpu(cpu, &tmp); + } + KUNIT_EXPECT_TRUE(test, cpumask_equal(&tmp, &actbl[2]->cpumask)); + + for_each_possible_cpu(cpu) { + struct tlbi_domain *domain; + int i; + + for (i = 0; i < tlbid_data.nr_domain; i++) { + domain = &tlbid_data.cpu_domain[cpu][i]; + if (domain != actbl[0] && + domain != actbl[1] && + domain != actbl[2]) { + KUNIT_EXPECT_EQ(test, domain->id, TLBI_INV_DOMAIN); + } + } + } + + reset_real_tlbid(); +} + +static void __init test_pick_best_domain(struct kunit *test) +{ + int domain, cpu; + cpumask_t tmp; + + if (setup_fake_tlbid(test)) + return; + + domain = pick_best_domain(cpu_possible_mask); + KUNIT_EXPECT_EQ(test, domain, 0); + + cpumask_clear(&tmp); + for_each_possible_cpu(cpu) { + if (cpu >= (num_possible_cpus() / 2)) + continue; + cpumask_set_cpu(cpu, &tmp); + } + domain = pick_best_domain(&tmp); + KUNIT_EXPECT_EQ(test, domain, 1); + + cpumask_clear(&tmp); + for_each_possible_cpu(cpu) { + if (cpu < (num_possible_cpus() / 2)) + continue; + cpumask_set_cpu(cpu, &tmp); + } + domain = pick_best_domain(&tmp); + KUNIT_EXPECT_EQ(test, domain, 2); + + reset_real_tlbid(); +} + +static struct kunit_case __refdata tlbidomain_test_cases[] = { + KUNIT_CASE(test_parse_tlbid_topology), + KUNIT_CASE(test_pick_best_domain), + {} +}; + +static struct kunit_suite tlbidomain_test_suite = { + .name = "tlbidomain_test_suite", + .test_cases = tlbidomain_test_cases, +}; + +kunit_test_init_section_suites(&tlbidomain_test_suite); diff --git a/arch/arm64/mm/tlbidomain.c b/arch/arm64/mm/tlbidomain.c new file mode 100644 index 000000000000..a8abbcb94952 --- /dev/null +++ b/arch/arm64/mm/tlbidomain.c @@ -0,0 +1,371 @@ +// SPDX-License-Identifier: GPL-2.0-only +/* + * This file implements the kernel functions rely on FEAT_TLBID + * + * Copyright (C) 2026 Huawei Ltd. + */ + +#define pr_fmt(fmt) "tlbidomain: " fmt + +#include <linux/acpi.h> +#include <linux/cpumask.h> + +#include <asm/cpu.h> +#include <asm/sysreg.h> +#include <asm/tlbidomain.h> +#include <asm/cpufeature.h> + +struct tlbi_domain { + cpumask_t cpumask; + int id; +}; + +struct tlbid_data { + struct list_head target_list[ACPI_TLBI_TYPE_RESERVED]; + unsigned long *found_domains; + u32 nr_domain; + struct tlbi_domain *cpu_domain[NR_CPUS]; + u8 nis; /* TLBIDIDR_EL1.NIS */ + u8 nvis; /* TLBIDIDR_EL1.NVIS */ + u8 nos; /* TLBIDIDR_EL1.NOS */ + u8 nvos; /* TLBIDIDR_EL1.NVOS */ +}; + +/* + * Global TLBI domain data. + * Actual supported domain bits will be reduced based on hardware capability. + */ +static struct tlbid_data tlbid_data __read_mostly = { + .nr_domain = 1, + .nis = TLBI_RT_DOMAIN_WIDTH, + .nvis = TLBI_RT_DOMAIN_WIDTH, + .nos = TLBI_RT_DOMAIN_WIDTH, + .nvos = TLBI_RT_DOMAIN_WIDTH, +}; + +/* + * validate_ids - Validate TLBID domain bit configurations + * + * Validates that NIS/NVIS and NOS/NVOS configurations follow Arm architecture + * constraints: if outer domain bits are zero, inner domain bits must also be + * zero. Outputs error and zeroes out all domain counts if validation fails. + */ +static void validate_ids(u8 nis, u8 nvis, u8 nos, u8 nvos) +{ + bool inv = false; + + switch (nos) { + case 0: + inv = tlbid_data.nvos > 0; + break; + case 1 ... 8: + inv = tlbid_data.nvos > 6; + break; + case 9 ... 16: + inv = tlbid_data.nvos > 5; + break; + } + + if (inv) { + pr_err("Illegal value of NVOS(%u) for NOS(%u)\n", + tlbid_data.nvos, tlbid_data.nos); + goto on_error; + } + + switch (nis) { + case 0: + inv = tlbid_data.nvis > 0; + break; + case 1 ... 8: + inv = tlbid_data.nvis > 6; + break; + case 9 ... 16: + inv = tlbid_data.nvis > 5; + break; + } + + if (!inv) + return; + + pr_err("Illegal value of NVIS(%u) for NIS(%u)\n", + tlbid_data.nvis, tlbid_data.nis); +on_error: + tlbid_data.nis = 0; + tlbid_data.nvis = 0; + tlbid_data.nos = 0; + tlbid_data.nvos = 0; +} + +/* + * find_target - Find TLBI target by type and identifier + * + * Looks up the tlbi_target structure associated with the given type and cpu. + * For PROCESSOR type, converts CPU logical ID to ACPI Processor UID. + */ +static struct tlbi_target *find_target(int type, int cpu) +{ + struct tlbi_target *target = NULL; + struct list_head *head; + int uid; + + switch (type) { + case ACPI_TLBI_TYPE_PROCESSOR: + uid = get_acpi_id_for_cpu(cpu); + break; + default: + pr_err("Only support Processor as TLBI broadcast target\n"); + return NULL; + } + + head = &tlbid_data.target_list[type]; + list_for_each_entry(target, head, list) { + if (target->identifier == uid) + break; + } + + return target; +} + +void add_target(int type, struct tlbi_target *target) +{ + s32 i, id; + + for (i = 0; i < target->nr_domain; i++) { + id = target->domains[i]; + if ((id > 0) && !test_and_set_bit(id, tlbid_data.found_domains)) + tlbid_data.nr_domain += 1; + } + + INIT_LIST_HEAD(&target->list); + list_add_tail(&target->list, &tlbid_data.target_list[type]); +} + +/* + * pick_best_domain - Select optimal TLBI domain for given CPUs + * + * Iterates through cpu_domain table to find a domain whose cpumask exactly + * matches the given active CPUs. This allows targeted TLBI instead of + * broadcasting to all CPUs, improving TLB invalidation efficiency. + * Returns domain ID on success, 0 on failure. + */ +int pick_best_domain(const cpumask_t *active_cpus) +{ + struct tlbi_domain *domain; + cpumask_t tmp; + int nr_cpu; + + /* Domain 0 is the only supported for Innern Shareable TLBI */ + if (unlikely(tlbid_data.nr_domain == 1)) + goto domain0; + + nr_cpu = cpumask_weight(active_cpus); + if (unlikely(nr_cpu == 0)) + goto domain0; + + do { + domain = &tlbid_data.cpu_domain[nr_cpu - 1][0]; + while (domain->id != TLBI_INV_DOMAIN) { + if (cpumask_and(&tmp, active_cpus, &domain->cpumask) && + cpumask_equal(&tmp, active_cpus)) + return domain->id; + domain++; + } + } while (++nr_cpu <= num_possible_cpus()); + +domain0: + return 0; +} + +int __init tlbid_data_init(void) +{ + s32 max_domains = 1 << TLBI_RT_DOMAIN_WIDTH; + int type; + + if (!tlbid_data.nis || !tlbid_data.nos) + return -EINVAL; + + tlbid_data.found_domains = kcalloc(BITS_TO_LONGS(max_domains), + sizeof(long), GFP_KERNEL); + if (!tlbid_data.found_domains) + return -ENOMEM; + + type = ACPI_TLBI_TYPE_PROCESSOR; + while (type != ACPI_TLBI_TYPE_RESERVED) + INIT_LIST_HEAD(&tlbid_data.target_list[type++]); + + return 0; +} + +void __init tlbid_data_reset(void) +{ + struct tlbi_target *target, *next; + int cpu, type; + + for (type = ACPI_TLBI_TYPE_PROCESSOR; + type < ACPI_TLBI_TYPE_RESERVED; type++) { + list_for_each_entry_safe(target, next, + &tlbid_data.target_list[type], list) { + list_del(&target->list); + kfree(target->domains); + kfree(target); + } + } + + for_each_possible_cpu(cpu) { + kfree(tlbid_data.cpu_domain[cpu]); + tlbid_data.cpu_domain[cpu] = NULL; + } + + kfree(tlbid_data.found_domains); + tlbid_data.found_domains = NULL; + tlbid_data.nr_domain = 1; + tlbid_data.nis = 0; + tlbid_data.nos = 0; + tlbid_data.nvis = 0; + tlbid_data.nvos = 0; +} + +/* + * parse_tlbid_topology - Initialize TLBI domain mapping tables + * + * Allocates and populates the cpu_domain table: + * 1. Build CPU-to-domain mapping from ACPI-parsed target data + * 2. Reorganize data for efficient lookup (domain x CPU matrix) + * + * Returns 0 on success, -ENOMEM on allocation failure, -EVINAL on topology + * parse failure + */ +int __init parse_tlbid_topology(void) +{ + struct tlbi_domain *src, *dst; + struct tlbi_target *target; + int cpu, err; + s32 i; + + if (tlbid_data.nr_domain >= (1 << max(tlbid_data.nis, tlbid_data.nos))) { + pr_err("%d TLBI Domains ACPI reports are more than CPU support.\n", + tlbid_data.nr_domain); + err = -EINVAL; + goto on_error; + } + pr_info("found %d TLBI Domain(s).\n", tlbid_data.nr_domain); + + /* + * Allocate a sparse table to map the CPU count to domain list sharing + * the same CPU count. + * + * NOTICE: Each domain is assumed to contain all system-wide SMMUs, + * Any broadcast in a specific domain must wait replies from all SMMUs + * before it complete. + */ + for_each_possible_cpu(cpu) { + tlbid_data.cpu_domain[cpu] = + kcalloc(tlbid_data.nr_domain, + sizeof(struct tlbi_domain), GFP_KERNEL); + if (!tlbid_data.cpu_domain[cpu]) { + err = -ENOMEM; + goto on_error; + } + for (i = 0; i < tlbid_data.nr_domain; i++) { + cpumask_clear(&tlbid_data.cpu_domain[cpu][i].cpumask); + tlbid_data.cpu_domain[cpu][i].id = cpu ? + TLBI_INV_DOMAIN : i; + } + } + + /* Initialize the cpumask associated with each tlbi domain */ + for_each_possible_cpu(cpu) { + target = find_target(ACPI_TLBI_TYPE_PROCESSOR, cpu); + if (!target) { + err = -EINVAL; + goto on_error; + } + for (i = 0; i < target->nr_domain; i++) { + /* Record all CPUs belong to the same domain */ + src = &tlbid_data.cpu_domain[0][target->domains[i]]; + cpumask_set_cpu(cpu, &(src->cpumask)); + } + } + + /* Move each domain data onto the right position in sparse table */ + for (i = 0; i < tlbid_data.nr_domain; i++) { + src = &tlbid_data.cpu_domain[0][i]; + cpu = cpumask_weight(&(src->cpumask)); + + if (cpu == 1) + continue; + + if (cpu == 0) { + pr_err("TLBI Domain %d has no CPU\n", i); + goto on_error; + } + + dst = &tlbid_data.cpu_domain[cpu - 1][0]; + while (dst->id != TLBI_INV_DOMAIN) + dst++; + memcpy(dst, src, sizeof(*dst)); + cpumask_clear(&src->cpumask); + src->id = TLBI_INV_DOMAIN; + } + + return 0; + +on_error: + for_each_possible_cpu(cpu) + kfree(tlbid_data.cpu_domain[cpu]); + return err; +} + +/* + * tlbid_early_init - Early initialization of TLBID hardware capability after + * all CPUs activate. + * + * Called early during boot from setup_system_features() before ACPI is parsed. + * Reads TLBIDIDR_EL1 register on each CPU to discover hardware capabilities + * (NIS, NVIS, NOS, NVOS bits), and takes the minimum across all CPUs. + * Validates the combined configuration via validate_ids(). + */ +void __init tlbid_early_init(void) +{ + u16 nis, nos, nvis, nvos; + u64 tlbididr; + int cpu; + + if (!system_supports_tlbid()) { + tlbid_data.nis = 0; + tlbid_data.nos = 0; + tlbid_data.nvis = 0; + tlbid_data.nvos = 0; + return; + } + + for_each_online_cpu(cpu) { + tlbididr = per_cpu_ptr(&cpu_data, cpu)->reg_tlbididr; + nis = FIELD_GET(TLBIDIDR_EL1_NIS_MASK, tlbididr); + nos = FIELD_GET(TLBIDIDR_EL1_NOS_MASK, tlbididr); + nvis = FIELD_GET(TLBIDIDR_EL1_NVIS_MASK, tlbididr); + nvos = FIELD_GET(TLBIDIDR_EL1_NVOS_MASK, tlbididr); + + if (tlbid_data.nis > nis) + tlbid_data.nis = nis; + + if (tlbid_data.nos > nos) + tlbid_data.nos = nos; + + if (tlbid_data.nvis > nvis) + tlbid_data.nvis = nvis; + + if (tlbid_data.nvos > nvos) + tlbid_data.nvos = nvos; + } + + validate_ids(tlbid_data.nis, tlbid_data.nvis, + tlbid_data.nos, tlbid_data.nvos); + + pr_info("TLBI Domain bits: nis %d nos %d nvis %d nvos %d\n", + tlbid_data.nis, tlbid_data.nos, tlbid_data.nvis, tlbid_data.nvos); +} + +#ifdef CONFIG_TLBID_KUNIT_TEST +#include "test_tlbidomain.c" +#endif diff --git a/arch/arm64/tools/sysreg b/arch/arm64/tools/sysreg index 1e8dee638509..2155315581b2 100644 --- a/arch/arm64/tools/sysreg +++ b/arch/arm64/tools/sysreg @@ -3575,3 +3575,14 @@ Field 5 F Field 4 P Field 3:0 Align EndSysreg + +Sysreg TLBIDIDR_EL1 3 0 10 4 6 +Res0 63:45 +Field 44:40 NVOS +Res0 39:37 +Field 36:32 NOS +Res0 31:13 +Field 12:8 NVIS +Res0 7:5 +Field 4:0 NIS +EndSysreg diff --git a/drivers/acpi/arm64/Makefile b/drivers/acpi/arm64/Makefile index 9497a7777729..2c068491f16e 100644 --- a/drivers/acpi/arm64/Makefile +++ b/drivers/acpi/arm64/Makefile @@ -5,5 +5,6 @@ obj-$(CONFIG_ACPI_GTDT) += gtdt.o obj-$(CONFIG_ACPI_APMT) += apmt.o obj-$(CONFIG_ARM_AMBA) += amba.o obj-$(CONFIG_ACPI_MPAM) += mpam.o +obj-$(CONFIG_ARM64_TLBID) += tlbid.o obj-y += dma.o init.o diff --git a/drivers/acpi/arm64/init.c b/drivers/acpi/arm64/init.c index d0c8aed90fd1..309f29b63f5d 100644 --- a/drivers/acpi/arm64/init.c +++ b/drivers/acpi/arm64/init.c @@ -12,4 +12,6 @@ void __init acpi_arm_init(void) acpi_iort_init(); if (IS_ENABLED(CONFIG_ARM_AMBA)) acpi_amba_init(); + if (IS_ENABLED(CONFIG_ARM64_TLBID)) + acpi_tlbi_init(); } diff --git a/drivers/acpi/arm64/init.h b/drivers/acpi/arm64/init.h index dcc277977194..c036e2622a2f 100644 --- a/drivers/acpi/arm64/init.h +++ b/drivers/acpi/arm64/init.h @@ -5,3 +5,4 @@ void __init acpi_agdi_init(void); void __init acpi_apmt_init(void); void __init acpi_iort_init(void); void __init acpi_amba_init(void); +void __init acpi_tlbi_init(void); diff --git a/drivers/acpi/arm64/tlbid.c b/drivers/acpi/arm64/tlbid.c new file mode 100644 index 000000000000..d6b1ee007e5d --- /dev/null +++ b/drivers/acpi/arm64/tlbid.c @@ -0,0 +1,105 @@ +// SPDX-License-Identifier: GPL-2.0-only +/* + * This file implements handling of + * Arm Translation Lookaside Buffer Invalidate Domain table (TLBI) + * + * Copyright (C) 2026, Huawei Ltd. + */ +#define pr_fmt(fmt) "ACPI: TLBID: " fmt + +#include <linux/acpi.h> +#include <linux/kernel.h> +#include <linux/list.h> +#include <linux/slab.h> + +#include "init.h" + +#include <asm/tlbidomain.h> + +static struct acpi_table_header *acpi_tlbid_table; + +/* + * tlbid_parse_subtable - Parse ACPI TLBI subtable entry + * + * Called for each entry in the ACPI TLBI table. Allocates a tlbi_target + * structure and copies domain IDs from the ACPI table. Builds the + * found_domains bitmap for later domain count calculation. + */ +static int __init tlbid_parse_subtable(union acpi_subtable_headers *header, + unsigned long end) +{ + struct acpi_tlbi_subtable *hdr = (struct acpi_tlbi_subtable *)header; + struct tlbi_target *target; + + if (!hdr) + return -EINVAL; + + target = kzalloc(sizeof(*target), GFP_KERNEL); + if (!target) + return -ENOMEM; + + /* The ACPI Processor UID of the corresponding GICC MADT entry */ + target->nr_domain = hdr->nr_domain; + target->identifier = hdr->identifier; + target->domains = kcalloc(hdr->nr_domain, sizeof(*target->domains), + GFP_KERNEL); + if (!target->domains) { + kfree(target); + return -ENOMEM; + } + + memcpy(target->domains, (void *)acpi_tlbid_table + hdr->offset, + hdr->nr_domain * sizeof(*target->domains)); + + add_target(header->tlbi.type, target); + + return 0; +} + +/* + * acpi_tlbi_init - Initialize TLBID from ACPI table + * + * Main entry point for TLBI table parsing. Gets the ACPI TLBI table, + * parses Processor and SMMU subtable entries, calculates total domain + * count, validates against hardware capability, and calls + * tlbid_chip_data_init() to build domain mapping tables. + * Outputs appropriate error messages on failure. + */ +void __init acpi_tlbi_init(void) +{ + enum acpi_tlbi_target_type type; + acpi_status status; + int ret; + + if (tlbid_data_init() < 0) + goto on_error; + + status = acpi_get_table(ACPI_SIG_TLBI, 0, &acpi_tlbid_table); + if (ACPI_FAILURE(status)) + goto on_error; + + type = ACPI_TLBI_TYPE_PROCESSOR; + while (type != ACPI_TLBI_TYPE_RESERVED) { + ret = acpi_table_parse_entries(ACPI_SIG_TLBI, + sizeof(struct acpi_table_tlbi), + type, tlbid_parse_subtable, 0); + if (ret <= 0) { + pr_err("Failed to get TLBID subtable for %s\n", + type == ACPI_TLBI_TYPE_PROCESSOR ? "Processor":"SMMU"); + goto on_error; + } + type++; + } + + if (parse_tlbid_topology()) { + pr_warn("Failed to initialize TLBI Domain data\n"); + goto on_error; + } + + acpi_put_table(acpi_tlbid_table); + return; + +on_error: + tlbid_data_reset(); + acpi_put_table(acpi_tlbid_table); +} diff --git a/drivers/acpi/tables.c b/drivers/acpi/tables.c index 4fca04c8dd2d..62d786110ebf 100644 --- a/drivers/acpi/tables.c +++ b/drivers/acpi/tables.c @@ -42,6 +42,7 @@ enum acpi_subtable_type { ACPI_SUBTABLE_HMAT, ACPI_SUBTABLE_PRMT, ACPI_SUBTABLE_CEDT, + ACPI_SUBTABLE_TLBI, }; struct acpi_subtable_entry { @@ -287,6 +288,8 @@ acpi_get_entry_type(struct acpi_subtable_entry *entry) return 0; case ACPI_SUBTABLE_CEDT: return entry->hdr->cedt.type; + case ACPI_SUBTABLE_TLBI: + return entry->hdr->tlbi.type; } return 0; } @@ -303,6 +306,8 @@ acpi_get_entry_length(struct acpi_subtable_entry *entry) return entry->hdr->prmt.length; case ACPI_SUBTABLE_CEDT: return entry->hdr->cedt.length; + case ACPI_SUBTABLE_TLBI: + return entry->hdr->tlbi.length; } return 0; } @@ -319,6 +324,8 @@ acpi_get_subtable_header_length(struct acpi_subtable_entry *entry) return sizeof(entry->hdr->prmt); case ACPI_SUBTABLE_CEDT: return sizeof(entry->hdr->cedt); + case ACPI_SUBTABLE_TLBI: + return sizeof(entry->hdr->tlbi); } return 0; } @@ -332,6 +339,8 @@ acpi_get_subtable_type(char *id) return ACPI_SUBTABLE_PRMT; if (strncmp(id, ACPI_SIG_CEDT, 4) == 0) return ACPI_SUBTABLE_CEDT; + if (strncmp(id, ACPI_SIG_TLBI, 4) == 0) + return ACPI_SUBTABLE_TLBI; return ACPI_SUBTABLE_COMMON; } diff --git a/include/acpi/actbl2.h b/include/acpi/actbl2.h index 97b05bd334a5..29dda0318c0f 100644 --- a/include/acpi/actbl2.h +++ b/include/acpi/actbl2.h @@ -26,6 +26,7 @@ */ #define ACPI_SIG_AGDI "AGDI" /* Arm Generic Diagnostic Dump and Reset Device Interface */ #define ACPI_SIG_APMT "APMT" /* Arm Performance Monitoring Unit table */ +#define ACPI_SIG_TLBI "TLBI" /* Translation Lookaside Buffer Invalidate Domain Table */ #define ACPI_SIG_BDAT "BDAT" /* BIOS Data ACPI Table */ #define ACPI_SIG_CCEL "CCEL" /* CC Event Log Table */ #define ACPI_SIG_CDAT "CDAT" /* Coherent Device Attribute Table */ @@ -341,6 +342,34 @@ enum acpi_apmt_node_type { #define ACPI_APMT_OVFLW_IRQ_FLAGS_TYPE_WIRED (0<<1) +/******************************************************************************* + * TLBI - Arm Translation Lookaside Buffer Invalidation Domain Table (TLBI) + * + ******************************************************************************/ +struct acpi_table_tlbi { + struct acpi_table_header header; /* Common ACPI table header */ + u32 nr_subtable; +}; + +struct acpi_tlbi_header { + u16 type; + u16 reserved; + u32 length; +}; + +struct acpi_tlbi_subtable { + struct acpi_tlbi_header header; + u32 identifier; + u32 nr_domain; + u32 offset; +}; + +enum acpi_tlbi_target_type { + ACPI_TLBI_TYPE_PROCESSOR = 1, + ACPI_TLBI_TYPE_SMMU, + ACPI_TLBI_TYPE_RESERVED, +}; + /******************************************************************************* * * BDAT - BIOS Data ACPI Table diff --git a/include/linux/acpi.h b/include/linux/acpi.h index e59a74273e6e..acd1a9241829 100644 --- a/include/linux/acpi.h +++ b/include/linux/acpi.h @@ -127,6 +127,7 @@ union acpi_subtable_headers { struct acpi_hmat_structure hmat; struct acpi_prmt_module_header prmt; struct acpi_cedt_header cedt; + struct acpi_tlbi_header tlbi; }; typedef int (*acpi_tbl_table_handler)(struct acpi_table_header *table); -- 2.33.0
From: Zeng Heng <zengheng4@huawei.com> hulk inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- To utilize TLBID feature effectively, we need to know which CPUs a given mm (address space) has been active on. This patch implements cumulative CPU tracking for each mm: - On new ASID allocation (new_context()): Clear the cpumask to start fresh. This handles both new processes and ASID generation wrap-around cases. - On context switch (check_and_switch_context()): Set the current CPU in mm_cpumask(mm) if TLBID is supported. The tracking is cumulative (CPUs are never cleared except on ASID re-allocation). While this may include CPUs where the task no longer runs, TLB invalidation to a superset remains functionally correct. This infrastructure enables subsequent flush_tlb_mm() can use domain-based invalidation instead of full broadcast when the mm's CPU footprint is limited. Signed-off-by: Zeng Heng <zengheng4@huawei.com> --- arch/arm64/kernel/smp.c | 3 +++ arch/arm64/mm/context.c | 9 +++++++-- 2 files changed, 10 insertions(+), 2 deletions(-) diff --git a/arch/arm64/kernel/smp.c b/arch/arm64/kernel/smp.c index 1131f96fa948..4290231e02e0 100644 --- a/arch/arm64/kernel/smp.c +++ b/arch/arm64/kernel/smp.c @@ -351,6 +351,9 @@ void arch_cpuhp_cleanup_dead_cpu(unsigned int cpu) pr_debug("CPU%u: shutdown\n", cpu); + if (system_supports_tlbid()) + clear_tasks_mm_cpumask(cpu); + /* * Now that the dying CPU is beyond the point of no return w.r.t. * in-kernel synchronisation, try to get the firwmare to help us to diff --git a/arch/arm64/mm/context.c b/arch/arm64/mm/context.c index b2ac06246327..004c08dec441 100644 --- a/arch/arm64/mm/context.c +++ b/arch/arm64/mm/context.c @@ -207,6 +207,9 @@ static u64 new_context(struct mm_struct *mm) asid = find_next_zero_bit(asid_map, NUM_USER_ASIDS, 1); set_asid: + if (system_supports_tlbid()) + cpumask_clear(mm_cpumask(mm)); + __set_bit(asid, asid_map); cur_idx = asid; return asid2ctxid(asid, generation); @@ -215,8 +218,8 @@ static u64 new_context(struct mm_struct *mm) void check_and_switch_context(struct mm_struct *mm) { unsigned long flags; - unsigned int cpu; u64 asid, old_active_asid; + unsigned int cpu = smp_processor_id(); if (system_supports_cnp()) cpu_set_reserved_ttbr0(); @@ -251,7 +254,6 @@ void check_and_switch_context(struct mm_struct *mm) atomic64_set(&mm->context.id, asid); } - cpu = smp_processor_id(); if (cpumask_test_and_clear_cpu(cpu, &tlb_flush_pending)) local_flush_tlb_all(); @@ -262,6 +264,9 @@ void check_and_switch_context(struct mm_struct *mm) arm64_apply_bp_hardening(); + if (system_supports_tlbid()) + cpumask_set_cpu(cpu, mm_cpumask(mm)); + /* * Defer TTBR0_EL1 setting for user threads to uaccess_enable() when * emulating PAN. -- 2.33.0
From: Jinjiang Tu <tujinjiang@huawei.com> hulk inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- This patch optimizes flush_tlb_mm() by using TLBID when available: - Local invalidation: When the mm is only active on the current CPU, use non-shareable TLB invalidation (aside1) for better performance. - Domain-based invalidation: When the mm is active on multiple CPUs but not all, use TLBID to target only the relevant CPUs, reducing cache coherency traffic compared to full broadcast. - Broadcast fallback: When TLBID is not supported or when targeting all CPUs, maintain the existing behavior (aside1is). The scope detection logic (flush_tlb_user_scope()) determines the optimal invalidation strategy based on mm_cpumask() and TLBID capabilities. This optimization is particularly beneficial for: - Workloads with many short-lived processes - Systems with high CPU counts where broadcast TLB shootdowns are costly - Scenarios where process memory is mostly local to a subset of CPUs Signed-off-by: Jinjiang Tu <tujinjiang@huawei.com> Signed-off-by: Zeng Heng <zengheng4@huawei.com> --- arch/arm64/include/asm/tlbflush.h | 53 +++++++++++++++++++++++++++++-- 1 file changed, 50 insertions(+), 3 deletions(-) diff --git a/arch/arm64/include/asm/tlbflush.h b/arch/arm64/include/asm/tlbflush.h index c082cace5683..03ebfc94287f 100644 --- a/arch/arm64/include/asm/tlbflush.h +++ b/arch/arm64/include/asm/tlbflush.h @@ -16,6 +16,7 @@ #include <linux/mmu_notifier.h> #include <asm/cputype.h> #include <asm/mmu.h> +#include <asm/tlbidomain.h> /* * Raw TLBI operations. @@ -165,6 +166,38 @@ static inline unsigned long get_trans_granule(void) (__pages >> (5 * (scale) + 1)) - 1; \ }) +enum tlb_flush_scope { + TLB_FLUSH_SCOPE_LOCAL, + TLB_FLUSH_SCOPE_DOMAIN, + TLB_FLUSH_SCOPE_BROADCAST, +}; + +/* This macro creates a properly formatted VA operand for the TLBID */ +#define __TLBI_DOMAIN(asid, domain) \ + ({ \ + unsigned long __ta = 0; \ + __ta |= FIELD_PREP(GENMASK_ULL(15, 0), domain); \ + __ta |= FIELD_PREP(GENMASK_ULL(63, 48), asid); \ + __ta; \ + }) + +/* + * Determines whether the user tlbi invalidation can be performed only on the + * local CPU or whether it needs to be multicast or broadcast. + */ +static inline enum tlb_flush_scope flush_tlb_user_scope(struct mm_struct *mm) +{ + if (!system_supports_tlbid()) + return TLB_FLUSH_SCOPE_BROADCAST; + + /* check if the tlbflush needs to be sent to other CPUs */ + if (cpumask_any_but(mm_cpumask(mm), smp_processor_id()) >= + nr_cpu_ids) + return TLB_FLUSH_SCOPE_LOCAL; + + return TLB_FLUSH_SCOPE_DOMAIN; +} + /* * TLB Invalidation * ================ @@ -252,12 +285,26 @@ static inline void flush_tlb_all(void) static inline void flush_tlb_mm(struct mm_struct *mm) { + enum tlb_flush_scope scope; unsigned long asid; + int domain; dsb(ishst); - asid = __TLBI_VADDR(0, ASID(mm)); - __tlbi(aside1is, asid); - __tlbi_user(aside1is, asid); + scope = flush_tlb_user_scope(mm); + if (scope == TLB_FLUSH_SCOPE_LOCAL) { + asid = __TLBI_VADDR(0, ASID(mm)); + __tlbi(aside1, asid); + __tlbi_user(aside1, asid); + } else if (scope == TLB_FLUSH_SCOPE_BROADCAST) { + asid = __TLBI_VADDR(0, ASID(mm)); + __tlbi(aside1is, asid); + __tlbi_user(aside1is, asid); + } else { + domain = pick_best_domain(mm_cpumask(mm)); + asid = __TLBI_DOMAIN(ASID(mm), domain); + __tlbi(aside1is, asid); + __tlbi_user(aside1is, asid); + } dsb(ish); mmu_notifier_arch_invalidate_secondary_tlbs(mm, 0, -1UL); } -- 2.33.0
From: Zeng Heng <zengheng4@huawei.com> hulk inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- TLBIP is the 128-bit form of TLBI (an alias of the SYSP system instruction) and is available when FEAT_TLBID is implemented. It carries the TLBI Domain in bits [15:0] of the first register of an even-aligned register pair (Xt, Xt+1), restricting the invalidation to the PEs within that domain. This allows a range invalidation to be multicast to the CPUs of a single domain instead of being broadcast to every CPU in the system. Extend __flush_tlb_range_nosync() to issue the invalidation with TLBIP RVAE1IS/RVALE1IS (or VAE1IS/VALE1IS for a single page) carrying the domain selected by pick_best_domain() when the CPUs the mm has run on fit into a single TLBI Domain, mirroring what flush_tlb_mm() already does for the whole address space. The *E1IS variants are used, like the existing aside1is in flush_tlb_mm(), so the invalidation completes with the usual DSB ISH. Fall back to the regular TLBI broadcast when no suitable domain is found. The TLBIP mnemonic is used when the assembler supports it (CONFIG_AS_HAS_TLBIP), with the instructions emitted as raw encodings otherwise. Since the encoding only holds one register and implies its pair as Xt+1, the operand pair is fixed to x10/x11. TTL64 is set in the operand so that the TTL hint applies to the VMSAv8-64 TLB entries created by the kernel. Signed-off-by: Zeng Heng <zengheng4@huawei.com> --- arch/arm64/Kconfig | 9 ++ arch/arm64/include/asm/tlbflush.h | 191 ++++++++++++++++++++++++++++-- 2 files changed, 193 insertions(+), 7 deletions(-) diff --git a/arch/arm64/Kconfig b/arch/arm64/Kconfig index fe3fa113e5cd..50146c8f8617 100644 --- a/arch/arm64/Kconfig +++ b/arch/arm64/Kconfig @@ -2306,6 +2306,7 @@ config ARM64_TLB_RANGE config ARM64_TLBID bool "Enable support for TLBID (TLBI Domains)" default y + depends on AS_HAS_TLBIP help TLBI broadcasting all PEs introduces performance noise. By combining hardware and software, TLBID (TLBI Domain) limit TLBI @@ -2323,6 +2324,14 @@ config ARM64_TLBID The feature is detected at runtime, and will remain disabled if the system does not implement the feature. +config AS_HAS_TLBIP + def_bool $(as-instr,.arch_extension d128) + help + The TLBIP instruction mnemonic requires the FEAT_D128 extension + of the assembler. When the assembler does not support it, the + kernel falls back to emitting TLBIP instructions as raw + encodings. + config ARM64_MPAM bool "Enable support for MPAM" select ACPI_MPAM if ACPI diff --git a/arch/arm64/include/asm/tlbflush.h b/arch/arm64/include/asm/tlbflush.h index 03ebfc94287f..615e35c1968f 100644 --- a/arch/arm64/include/asm/tlbflush.h +++ b/arch/arm64/include/asm/tlbflush.h @@ -445,6 +445,150 @@ do { \ #define __flush_s2_tlb_range_op(op, start, pages, stride, tlb_level) \ __flush_tlb_range_op(op, start, pages, stride, 0, tlb_level, false) +/* + * TLBIP is the 128-bit form of TLBI (an alias of the SYSP system + * instruction, FEAT_SYSINSTR128) and takes a pair of consecutive, + * even-aligned registers (Xt, Xt+1), of which only Xt is encoded in the + * instruction. With FEAT_TLBID the invalidation is restricted to the PEs + * that are within both the broadcast shareability domain and the TLBI + * Domain held in bits [15:0] of Xt. This allows a range invalidation to + * be multicast to the CPUs of a single domain instead of being broadcast + * to every CPU in the system. + * + * The operand pair of the *E1IS range instructions is laid out as: + * + * Xt : [63:48] ASID, [47:46] TG, [45:44] SCALE, [43:39] NUM, + * [38:37] TTL, [32] TTL64, [15:0] TLBID + * Xt+1: [43:0] BaseADDR[55:12] + * + * and for the non-range VAE1IS/VALE1IS instructions as: + * + * Xt : [63:48] ASID, [47:44] TTL, [32] TTL64, [15:0] TLBID + * Xt+1: [43:0] VA[55:12] + * + * TTL64 is always set so that the TTL hint applies to the VMSAv8-64 TLB + * entries, which are the only ones the kernel creates. When the assembler + * does not support the TLBIP mnemonic (CONFIG_AS_HAS_TLBIP), the + * instructions are emitted as raw encodings instead. + */ +#define __TLBI_PAIR_LO "x10" +#define __TLBI_PAIR_HI "x11" + +#define __TLBIP_PREAMBLE ".arch_extension d128\n" +#define __tlbip(op, lo, hi) \ +do { \ + register unsigned long __tlbip_lo asm (__TLBI_PAIR_LO) = (lo); \ + register unsigned long __tlbip_hi asm (__TLBI_PAIR_HI) = (hi); \ + \ + asm volatile(__TLBIP_PREAMBLE \ + "tlbip " #op ", " __TLBI_PAIR_LO ", " __TLBI_PAIR_HI "\n\t" \ + : : "r" (__tlbip_lo), "r" (__tlbip_hi)); \ +} while (0) + +#define __tlbip_user(insn, lo, hi) \ +do { \ + if (arm64_kernel_unmapped_at_el0()) \ + __tlbip(insn, (lo) | USER_ASID_FLAG, hi); \ +} while (0) + +/* Build the lower register of a TLBIP operand pair carrying @asid. */ +static __always_inline unsigned long +__tlbip_range_lo(unsigned long asid, unsigned long domain, + int scale, int num, int ttl) +{ + return FIELD_PREP(GENMASK_ULL(63, 48), asid) | + FIELD_PREP(GENMASK_ULL(47, 46), get_trans_granule()) | + FIELD_PREP(GENMASK_ULL(45, 44), scale) | + FIELD_PREP(GENMASK_ULL(43, 39), num) | + FIELD_PREP(GENMASK_ULL(38, 37), ttl) | + FIELD_PREP(BIT_ULL(32), 1) | /* TTL64 */ + FIELD_PREP(GENMASK_ULL(15, 0), domain); +} + +/* Build the upper register of a TLBIP operand pair. */ +static __always_inline unsigned long __tlbip_addr_hi(unsigned long addr) +{ + return FIELD_PREP(GENMASK_ULL(43, 0), addr >> PAGE_SHIFT); +} + +/* + * Build the lower register of the non-range TLBIP VAE1IS/VALE1IS operand + * pair. The TTL field uses the same granule+level encoding as + * __tlbi_level(). + */ +static __always_inline unsigned long +__tlbip_va_lo(unsigned long asid, unsigned long domain, int tlb_level) +{ + unsigned long lo = FIELD_PREP(GENMASK_ULL(63, 48), asid) | + FIELD_PREP(BIT_ULL(32), 1) | /* TTL64 */ + FIELD_PREP(GENMASK_ULL(15, 0), domain); + + if (cpus_have_const_cap(ARM64_HAS_ARMv8_4_TTL) && tlb_level) { + unsigned long ttl = (tlb_level & 3) | + (get_trans_granule() << 2); + + lo |= FIELD_PREP(GENMASK_ULL(47, 44), ttl); + } + + return lo; +} + +/* + * __flush_tlb_range_domain_op - Perform a TLBIP range operation restricted + * to a single TLBI Domain + * + * Same algorithm as __flush_tlb_range_op(), except that the invalidation is + * issued with the TLBIP *E1IS instructions carrying @domain, so that only + * the PEs of that domain have to process it. Like the regular TLBI + * instructions they broadcast to the Inner Shareable domain and are + * completed with the usual DSB ISH. + * + * @op: TLBIP operation name for the non-range instruction + * @rop: TLBIP operation name for the range instruction + * @start: The start address of the range + * @pages: Range as the number of pages from 'start' + * @stride: Flush granularity + * @asid: The ASID of the task + * @domain: The TLBI Domain the invalidation targets + * @tlb_level: Translation Table level hint, if known + */ +#define __flush_tlb_range_domain_op(op, rop, start, pages, stride, \ + asid, domain, tlb_level) \ +do { \ + typeof(start) __flush_start = start; \ + typeof(pages) __flush_pages = pages; \ + int num = 0; \ + int scale = 3; \ + unsigned long __lo; \ + \ + while (__flush_pages > 0) { \ + if (!system_supports_tlb_range() || \ + __flush_pages == 1) { \ + __lo = __tlbip_va_lo(asid, domain, tlb_level); \ + __tlbip(op, __lo, \ + __tlbip_addr_hi(__flush_start)); \ + __tlbip_user(op, __lo, \ + __tlbip_addr_hi(__flush_start)); \ + __flush_start += stride; \ + __flush_pages -= stride >> PAGE_SHIFT; \ + continue; \ + } \ + \ + num = __TLBI_RANGE_NUM(__flush_pages, scale); \ + if (num >= 0) { \ + __lo = __tlbip_range_lo(asid, domain, \ + scale, num, tlb_level); \ + __tlbip(rop, __lo, \ + __tlbip_addr_hi(__flush_start)); \ + __tlbip_user(rop, __lo, \ + __tlbip_addr_hi(__flush_start)); \ + __flush_start += __TLBI_RANGE_PAGES(num, scale) << PAGE_SHIFT; \ + __flush_pages -= __TLBI_RANGE_PAGES(num, scale);\ + } \ + scale--; \ + } \ +} while (0) + static inline bool __flush_tlb_range_limit_excess(unsigned long start, unsigned long end, unsigned long pages, unsigned long stride) { @@ -462,12 +606,22 @@ static inline bool __flush_tlb_range_limit_excess(unsigned long start, return false; } +/* + * Issue the TLBI operations for the range [start, end) without a trailing + * DSB, which the caller is responsible for issuing. + * + * When the CPUs the mm has run on fit into a single TLBI Domain, the + * invalidation is multicast to that domain with the TLBIP *E1IS + * instructions, sparing the CPUs outside of it from having to process the + * broadcast. + */ static inline void __flush_tlb_range_nosync(struct mm_struct *mm, - unsigned long start, unsigned long end, - unsigned long stride, bool last_level, - int tlb_level) + unsigned long start, unsigned long end, + unsigned long stride, bool last_level, + int tlb_level) { unsigned long asid, pages; + int domain, scope; start = round_down(start, stride); end = round_up(end, stride); @@ -481,10 +635,33 @@ static inline void __flush_tlb_range_nosync(struct mm_struct *mm, dsb(ishst); asid = ASID(mm); - if (last_level) - __flush_tlb_range_op(vale1is, start, pages, stride, asid, tlb_level, true); - else - __flush_tlb_range_op(vae1is, start, pages, stride, asid, tlb_level, true); + scope = flush_tlb_user_scope(mm); + if (scope == TLB_FLUSH_SCOPE_DOMAIN) { + domain = pick_best_domain(mm_cpumask(mm)); + + if (last_level) + __flush_tlb_range_domain_op(vale1is, rvale1is, + start, pages, stride, asid, domain, + tlb_level); + else + __flush_tlb_range_domain_op(vae1is, rvae1is, + start, pages, stride, asid, domain, + tlb_level); + } else if (scope == TLB_FLUSH_SCOPE_BROADCAST) { + if (last_level) + __flush_tlb_range_op(vale1is, start, pages, stride, + asid, tlb_level, true); + else + __flush_tlb_range_op(vae1is, start, pages, stride, + asid, tlb_level, true); + } else { + if (last_level) + __flush_tlb_range_op(vale1, start, pages, stride, + asid, tlb_level, true); + else + __flush_tlb_range_op(vae1, start, pages, stride, + asid, tlb_level, true); + } mmu_notifier_arch_invalidate_secondary_tlbs(mm, start, end); } -- 2.33.0
From: Zeng Heng <zengheng4@huawei.com> hulk inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- flush_tlb_page() currently broadcasts the invalidation with TLBI VALE1IS regardless of where the mm has run. As with the range flush path, restrict it to the TLBI Domain covering the CPUs the mm has run on when such a domain exists, issuing TLBIP VALE1IS instead. The non-range TLBIP instructions take the VA in the upper register of the pair and the ASID, TTL, TTL64 and TLBID fields in the lower one. No TTL hint is passed, keeping the behaviour identical to the existing __tlbi(vale1is). The page flush helpers are moved after the TLBIP machinery in the header so that they can use it. Signed-off-by: Zeng Heng <zengheng4@huawei.com> --- arch/arm64/include/asm/tlbflush.h | 71 ++++++++++++++++++++----------- 1 file changed, 45 insertions(+), 26 deletions(-) diff --git a/arch/arm64/include/asm/tlbflush.h b/arch/arm64/include/asm/tlbflush.h index 615e35c1968f..b43c12086b4b 100644 --- a/arch/arm64/include/asm/tlbflush.h +++ b/arch/arm64/include/asm/tlbflush.h @@ -309,32 +309,6 @@ static inline void flush_tlb_mm(struct mm_struct *mm) mmu_notifier_arch_invalidate_secondary_tlbs(mm, 0, -1UL); } -static inline void __flush_tlb_page_nosync(struct mm_struct *mm, - unsigned long uaddr) -{ - unsigned long addr; - - dsb(ishst); - addr = __TLBI_VADDR(uaddr, ASID(mm)); - __tlbi(vale1is, addr); - __tlbi_user(vale1is, addr); - mmu_notifier_arch_invalidate_secondary_tlbs(mm, uaddr & PAGE_MASK, - (uaddr & PAGE_MASK) + PAGE_SIZE); -} - -static inline void flush_tlb_page_nosync(struct vm_area_struct *vma, - unsigned long uaddr) -{ - return __flush_tlb_page_nosync(vma->vm_mm, uaddr); -} - -static inline void flush_tlb_page(struct vm_area_struct *vma, - unsigned long uaddr) -{ - flush_tlb_page_nosync(vma, uaddr); - dsb(ish); -} - static inline bool arch_tlbbatch_should_defer(struct mm_struct *mm) { #ifdef CONFIG_ARM64_WORKAROUND_REPEAT_TLBI @@ -687,6 +661,51 @@ static inline void flush_tlb_range(struct vm_area_struct *vma, __flush_tlb_range(vma, start, end, PAGE_SIZE, false, 0); } +static inline void __flush_tlb_page_nosync(struct mm_struct *mm, + unsigned long uaddr) +{ + unsigned long asid = ASID(mm); + int domain, scope; + unsigned long lo; + + dsb(ishst); + + scope = flush_tlb_user_scope(mm); + if (scope == TLB_FLUSH_SCOPE_DOMAIN) { + domain = pick_best_domain(mm_cpumask(mm)); + + lo = __tlbip_va_lo(asid, domain, 0); + __tlbip(vale1is, lo, __tlbip_addr_hi(uaddr)); + __tlbip_user(vale1is, lo, __tlbip_addr_hi(uaddr)); + } else { + unsigned long addr = __TLBI_VADDR(uaddr, asid); + + if (scope == TLB_FLUSH_SCOPE_BROADCAST) { + __tlbi(vale1is, addr); + __tlbi_user(vale1is, addr); + } else { + __tlbi(vale1, addr); + __tlbi_user(vale1, addr); + } + } + + mmu_notifier_arch_invalidate_secondary_tlbs(mm, uaddr & PAGE_MASK, + (uaddr & PAGE_MASK) + PAGE_SIZE); +} + +static inline void flush_tlb_page_nosync(struct vm_area_struct *vma, + unsigned long uaddr) +{ + return __flush_tlb_page_nosync(vma->vm_mm, uaddr); +} + +static inline void flush_tlb_page(struct vm_area_struct *vma, + unsigned long uaddr) +{ + flush_tlb_page_nosync(vma, uaddr); + dsb(ish); +} + static inline void flush_tlb_kernel_range(unsigned long start, unsigned long end) { const unsigned long stride = PAGE_SIZE; -- 2.33.0
From: Zeng Heng <zengheng4@huawei.com> hulk inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- Enable CONFIG_ARM64_TLBID in openeuler_defconfig so that the domain-based TLB invalidation feature is built by default. With TLBID (FEAT_TLBID), flush_tlb_mm() can direct TLB shootdowns to only the CPUs where the address space is active, instead of a full broadcast invalidation. This reduces cache coherency traffic and benefits systems with high CPU counts. Signed-off-by: Zeng Heng <zengheng4@huawei.com> --- arch/arm64/configs/openeuler_defconfig | 1 + 1 file changed, 1 insertion(+) diff --git a/arch/arm64/configs/openeuler_defconfig b/arch/arm64/configs/openeuler_defconfig index aeaaabdb8f68..edf2ed90efb8 100644 --- a/arch/arm64/configs/openeuler_defconfig +++ b/arch/arm64/configs/openeuler_defconfig @@ -574,6 +574,7 @@ CONFIG_ARM64_PTR_AUTH_KERNEL=y # CONFIG_ARM64_AMU_EXTN=y CONFIG_ARM64_TLB_RANGE=y +CONFIG_ARM64_TLBID=y CONFIG_ARM64_MPAM=y # end of ARMv8.4 architectural features -- 2.33.0
From: Zhou Wang <wangzhou1@hisilicon.com> virt inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- Add the information about TLBID-related registers. Signed-off-by: Zhou Wang <wangzhou1@hisilicon.com> Signed-off-by: Jinqian Yang <yangjinqian1@huawei.com> --- arch/arm64/tools/sysreg | 19 ++++++++++++++++++- 1 file changed, 18 insertions(+), 1 deletion(-) diff --git a/arch/arm64/tools/sysreg b/arch/arm64/tools/sysreg index 2155315581b2..c10f6985f5db 100644 --- a/arch/arm64/tools/sysreg +++ b/arch/arm64/tools/sysreg @@ -3178,7 +3178,8 @@ Fields ZCR_ELx EndSysreg Sysreg HCRX_EL2 3 4 1 2 2 -Res0 63:25 +Res0 63:26 +Field 25 VTLBIDEn Field 24 PACMEn Field 23 EnFPM Field 22 GCSEn @@ -3586,3 +3587,19 @@ Field 12:8 NVIS Res0 7:5 Field 4:0 NIS EndSysreg + +Sysreg VTLBID0_EL2 3 4 2 8 0 +Field 63:0 TD +EndSysreg + +Sysreg VTLBID1_EL2 3 4 2 8 1 +Field 63:0 TD +EndSysreg + +Sysreg VTLBID2_EL2 3 4 2 8 2 +Field 63:0 TD +EndSysreg + +Sysreg VTLBID3_EL2 3 4 2 8 3 +Field 63:0 TD +EndSysreg -- 2.33.0
virt inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- Add support functions for TLBID virtualization initialization and ioctls. Signed-off-by: Jinqian Yang <yangjinqian1@huawei.com> Signed-off-by: Zhou Wang <wangzhou1@hisilicon.com> --- Documentation/virt/kvm/api.rst | 62 +++++++++++++ arch/arm64/include/asm/kvm_host.h | 27 ++++++ arch/arm64/kvm/arm.c | 149 ++++++++++++++++++++++++++++++ include/uapi/linux/kvm.h | 14 +++ 4 files changed, 252 insertions(+) diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst index 77c65f16e2a7..911ebf02ea16 100644 --- a/Documentation/virt/kvm/api.rst +++ b/Documentation/virt/kvm/api.rst @@ -6424,6 +6424,52 @@ the capability to be present. `flags` must currently be zero. +4.144 KVM_ARM_GET_VDOMAIN_NUM +------------------------------ + +:Capability: KVM_CAP_ARM_TLBIDOMAIN +:Architectures: arm64 +:Type: vm ioctl +:Parameters: struct kvm_arm_get_vdomain (out) +:Returns: 0 on success, < 0 on error + +Parameters are specified via the following structure:: + +:: + + struct kvm_arm_get_vdomain { + __u8 max_vdomains; + }; + +The ``max_vdomains`` field is the max number of vDomains supported by the +hardware. This value is determined by TLBIDIDR_EL1[NVOS/NVIS]. NVOS/NVIS +is the bit width of the vdomain. + +4.145 KVM_ARM_VCPU_SET_VDOMAIN +------------------------------ + +:Capability: KVM_CAP_ARM_TLBIDOMAIN +:Architectures: arm64 +:Type: vcpu ioctl +:Parameters: struct kvm_arm_set_vdomain (in) +:Returns: 0 on success, < 0 on error + +This ioctl allows KVM to obtain TLBI Domain of Guest. + +Parameters are specified via the following structure:: + +:: + + struct kvm_arm_set_vdomain { + __u32 vdomain_bitmap; + __u8 num_vdomains; + }; + +The ``vdomain_bitmap`` field is a bitmap, where '1' indicates that vDomain +includes this vCPU. For example, ``vCPU0: vdomain_bitmap=0b1011`` indicates +that vCPU0 included in vDomain(0, 1, 3). + +The ``num_vdomains`` field is the number of vDomains. 5. The kvm_run structure ======================== @@ -8218,6 +8264,22 @@ Used to configure and set up the memory for a Realm. The available actions are: enter the realm until it has been activated. ================================= ============================================= +7.39 KVM_CAP_ARM_TLBIDOMAIN +------------------------------------- + +:Architectures: arm64 +:Target: VM +:Parameters: None +:Returns: 0 on success, negative value on error + +This capability enables TLBID virtualization, which Optimizes tlbi broadcast +range to avoid unnecessary broadcasts. + +When this capability is enabled, KVM can obtain the vDomain bitmaps of VM. When +a vCPU is loaded, KVM writes the mapping between the vDomain and the pDomain +into the VTLBID(n). If guest enabled TLBID, hardware will convert vDomain id to +pDomain id for TLBI broadcast. + 8. Other capabilities. ====================== diff --git a/arch/arm64/include/asm/kvm_host.h b/arch/arm64/include/asm/kvm_host.h index 07e8144d55f7..d0a6c5b02807 100644 --- a/arch/arm64/include/asm/kvm_host.h +++ b/arch/arm64/include/asm/kvm_host.h @@ -208,6 +208,32 @@ struct kvm_protected_vm { struct kvm_hyp_memcache teardown_mc; }; +struct vcpu_tlbid_data { + /* Store vDomain bitmap, indexed by vCPU ID */ + u32 vdomain_bitmap; + /* Record the last pCPU where vCPU was running, indexed by vCPU ID */ + int last_pcpu; +}; + +struct tlbidomain { + /* Using xarray to dynamically allocate vCPU-related TLBID info */ + struct xarray vcpu_data_array; + /* The number of vDomains */ + int num_domains; + /* Mapping of pDomain to vDomain, indexed by vDomain ID */ + int *domain_map; + /* Record the bitmap of pCPUs mapped to vDomain, indexed by vDomain ID */ + cpumask_var_t *vdomain_cpumasks; + /* Whether virt-TLBID is enabled by KVM */ + bool kvm_tlbid_enabled; + /* Whether virt-TLBID is enabled by Guest */ + bool guest_tlbid_enabled; + /* TLBIDIDR.NVIS */ + u8 nvis; + /* TLBIDIDR.NIS */ + u8 nis; +}; + struct kvm_arch { struct kvm_s2_mmu mmu; @@ -332,6 +358,7 @@ struct kvm_arch { KABI_EXTEND(u64 midr_el1) KABI_EXTEND(u64 revidr_el1) KABI_EXTEND(u64 aidr_el1) + KABI_EXTEND(struct tlbidomain vdomain) }; struct kvm_vcpu_fault_info { diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c index 324553fea612..1abcb2b2e7e2 100644 --- a/arch/arm64/kvm/arm.c +++ b/arch/arm64/kvm/arm.c @@ -319,6 +319,82 @@ static int kvm_arm_default_max_vcpus(void) return vgic_present ? kvm_vgic_get_max_vcpus() : KVM_MAX_VCPUS; } +static int kvm_arm_init_tlbidomain(struct kvm *kvm) +{ + struct tlbidomain *vdomain; + int i, max_vdomains; + u64 tlbididr; + + if (!system_supports_tlbid()) + return 0; + + tlbididr = read_sysreg_s(SYS_TLBIDIDR_EL1); + vdomain = &kvm->arch.vdomain; + vdomain->nis = FIELD_GET(TLBIDIDR_EL1_NIS_MASK, tlbididr); + vdomain->nvis = FIELD_GET(TLBIDIDR_EL1_NVIS_MASK, tlbididr); + vdomain->num_domains = -1; + vdomain->kvm_tlbid_enabled = true; + max_vdomains = 1 << vdomain->nvis; + + kvm_info("TLBIDIDR: 0x%llx, NIS: %d, NVIS: %d\n", tlbididr, vdomain->nis, vdomain->nvis); + + xa_init(&vdomain->vcpu_data_array); + + vdomain->domain_map = kzalloc(sizeof(*vdomain->domain_map) * max_vdomains, + GFP_KERNEL); + if (!vdomain->domain_map) + goto destroy_vcpu_data_array; + + vdomain->vdomain_cpumasks = kcalloc(max_vdomains, sizeof(cpumask_var_t), + GFP_KERNEL); + if (!vdomain->vdomain_cpumasks) + goto free_domain_map; + + for (i = 0; i < max_vdomains; i++) { + if (!alloc_cpumask_var(&vdomain->vdomain_cpumasks[i], GFP_KERNEL)) + goto free_vdomain_cpumasks; + cpumask_clear(vdomain->vdomain_cpumasks[i]); + } + + for (i = 0; i < max_vdomains; i++) + kvm->arch.vdomain.domain_map[i] = -1; + + return 0; + +free_vdomain_cpumasks: + while (--i >= 0) + free_cpumask_var(vdomain->vdomain_cpumasks[i]); + kfree(vdomain->vdomain_cpumasks); +free_domain_map: + kfree(vdomain->domain_map); +destroy_vcpu_data_array: + xa_destroy(&vdomain->vcpu_data_array); + + return -ENOMEM; +} + +static void free_tlbid_vdomain(struct kvm *kvm) +{ + struct vcpu_tlbid_data *data; + int i, max_vdomains; + unsigned long index; + + max_vdomains = 1 << kvm->arch.vdomain.nvis; + + if (system_supports_tlbid()) { + kfree(kvm->arch.vdomain.domain_map); + + for (i = 0; i < max_vdomains; i++) + free_cpumask_var(kvm->arch.vdomain.vdomain_cpumasks[i]); + kfree(kvm->arch.vdomain.vdomain_cpumasks); + + xa_for_each(&kvm->arch.vdomain.vcpu_data_array, index, data) { + kfree(data); + } + xa_destroy(&kvm->arch.vdomain.vcpu_data_array); + } +} + /** * kvm_arch_init_vm - initializes a VM data structure * @kvm: pointer to the KVM struct @@ -1765,6 +1841,40 @@ static int kvm_vcpu_set_target(struct kvm_vcpu *vcpu, return kvm_reset_vcpu(vcpu); } +static int kvm_arm_tlbidomain_vcpu_init(struct kvm_vcpu *vcpu) +{ + struct vcpu_tlbid_data *data; + int ret; + + if (!system_supports_tlbid()) + return 0; + + data = kzalloc(sizeof(*data), GFP_KERNEL); + if (!data) + return -ENOMEM; + + /* + * When the host has enabled TLBID but the guest has not, + * KVM can map vDomain0 to pDomainN to optimize the TLBI + * instructions for the guest. Therefore, the value is set + * to 0x1, indicating that the vCPU is in vDomain0. If QEMU + * calls KVM_ARM_VCPU_SET_VDOMAIN to configure vdomains, this + * value will be overwritten to support the guest enabling + * TLBID. + */ + data->vdomain_bitmap = 0x1; + data->last_pcpu = -1; + + ret = xa_insert(&vcpu->kvm->arch.vdomain.vcpu_data_array, + vcpu->vcpu_idx, data, GFP_KERNEL); + if (ret) { + kfree(data); + return ret == -EBUSY ? 0 : ret; + } + + return 0; +} + static int kvm_arch_vcpu_ioctl_vcpu_init(struct kvm_vcpu *vcpu, struct kvm_vcpu_init *init) { @@ -1912,6 +2022,37 @@ static int kvm_arm_vcpu_rmm_psci_complete(struct kvm_vcpu *vcpu, return realm_psci_complete(vcpu, target, arg->psci_status); } +static int kvm_arm_vcpu_set_vdomain(struct kvm_vcpu *vcpu, + struct kvm_arm_set_vdomain *vdomain) +{ + u32 vdomain_bitmap = vdomain->vdomain_bitmap; + u8 num_vdomains = vdomain->num_vdomains; + u8 nvis = vcpu->kvm->arch.vdomain.nvis; + u8 max_vdomains = 1 << nvis; + struct vcpu_tlbid_data *data; + + if (num_vdomains > max_vdomains) + return -EINVAL; + + if (vcpu->kvm->arch.vdomain.num_domains != -1 && + vcpu->kvm->arch.vdomain.num_domains != num_vdomains) + return -EINVAL; + + data = xa_load(&vcpu->kvm->arch.vdomain.vcpu_data_array, vcpu->vcpu_idx); + if (!data) { + kvm_err("vCPU%d: Failed to load vcpu_tlbid_data\n", vcpu->vcpu_idx); + return -EINVAL; + } + + data->vdomain_bitmap = vdomain_bitmap; + vcpu->kvm->arch.vdomain.num_domains = num_vdomains; + vcpu->kvm->arch.vdomain.guest_tlbid_enabled = true; + + kvm_info("vCPU%d: vDomain bitmap: 0x%x\n", vcpu->vcpu_idx, vdomain_bitmap); + + return 0; +} + long kvm_arch_vcpu_ioctl(struct file *filp, unsigned int ioctl, unsigned long arg) { @@ -2110,6 +2251,14 @@ static int kvm_vm_set_attr(struct kvm *kvm, struct kvm_device_attr *attr) } } +static int kvm_arm_get_vdomain_num(struct kvm *kvm, + struct kvm_arm_get_vdomain *vdomain) +{ + vdomain->max_vdomains = 1 << kvm->arch.vdomain.nvis; + + return 0; +} + int kvm_arch_vm_ioctl(struct file *filp, unsigned int ioctl, unsigned long arg) { struct kvm *kvm = filp->private_data; diff --git a/include/uapi/linux/kvm.h b/include/uapi/linux/kvm.h index 18ed07a68b2d..6e93bc1e0697 100644 --- a/include/uapi/linux/kvm.h +++ b/include/uapi/linux/kvm.h @@ -1315,6 +1315,7 @@ struct kvm_ppc_resize_hpt { #define KVM_CAP_HYGON_COCO_EXT_CSV3_SP_MGR (1 << 4) #define KVM_CAP_ARM_HW_DIRTY_STATE_TRACK 502 +#define KVM_CAP_ARM_TLBIDOMAIN 503 #define KVM_CAP_ARM_HISI_IPIV 798 #define KVM_CAP_ARM_VIRT_MSI_BYPASS 799 @@ -2358,4 +2359,17 @@ struct kvm_pre_fault_memory { __u64 padding[5]; }; +struct kvm_arm_get_vdomain { + __u8 max_vdomains; +}; + +#define KVM_ARM_GET_VDOMAIN_NUM _IOW(KVMIO, 0xd7, struct kvm_arm_get_vdomain) + +struct kvm_arm_set_vdomain { + __u32 vdomain_bitmap; + __u8 num_vdomains; +}; + +#define KVM_ARM_VCPU_SET_VDOMAIN _IOW(KVMIO, 0xd8, struct kvm_arm_set_vdomain) + #endif /* __LINUX_KVM_H */ -- 2.33.0
virt inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- If the vCPU is first online or is migrated, the domain map needs to be checked whether need to update, and new map needs to be written into the VTLBID(n). Signed-off-by: Jinqian Yang <yangjinqian1@huawei.com> Signed-off-by: Zhou Wang <wangzhou1@hisilicon.com> --- arch/arm64/include/asm/kvm_host.h | 12 ++ arch/arm64/kvm/arm.c | 217 ++++++++++++++++++++++++++++++ 2 files changed, 229 insertions(+) diff --git a/arch/arm64/include/asm/kvm_host.h b/arch/arm64/include/asm/kvm_host.h index d0a6c5b02807..18c95e605d8a 100644 --- a/arch/arm64/include/asm/kvm_host.h +++ b/arch/arm64/include/asm/kvm_host.h @@ -45,6 +45,15 @@ #define KVM_VCPU_MAX_FEATURES 9 #define KVM_VCPU_VALID_FEATURES (BIT(KVM_VCPU_MAX_FEATURES) - 1) +#define NIS_RANGE_MIN_1_8 1 +#define NIS_RANGE_MAX_1_8 8 +#define NIS_RANGE_MIN_9_16 9 +#define NIS_RANGE_MAX_9_16 16 +#define VTLBID_TD_WIDTH_NIS_1_8 8 +#define VTLBID_TD_WIDTH_NIS_9_16 16 +#define VTLBID_EL2_WIDTH 64 +#define VTLBID_EL2_MAX_REGS 4 + #define KVM_REQ_SLEEP \ KVM_ARCH_REQ_FLAGS(0, KVM_REQUEST_WAIT | KVM_REQUEST_NO_WAKEUP) #define KVM_REQ_IRQ_PENDING KVM_ARCH_REQ(1) @@ -57,6 +66,7 @@ #define KVM_REQ_RELOAD_TLBI_DVMBM KVM_ARCH_REQ(8) #define KVM_REQ_RELOAD_WFI_TRAPS KVM_ARCH_REQ(9) #define KVM_REQ_RELOAD_TIMER_EARLY_INJECT KVM_ARCH_REQ(10) +#define KVM_REQ_RELOAD_VTLBID KVM_ARCH_REQ(11) #define KVM_DIRTY_LOG_MANUAL_CAPS (KVM_DIRTY_LOG_MANUAL_PROTECT_ENABLE | \ KVM_DIRTY_LOG_INITIALLY_SET) @@ -232,6 +242,8 @@ struct tlbidomain { u8 nvis; /* TLBIDIDR.NIS */ u8 nis; + /* Used to prevent concurrent modifications to the domain mapping. */ + spinlock_t tlbid_lock; }; struct kvm_arch { diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c index 1abcb2b2e7e2..7985f85ab3cf 100644 --- a/arch/arm64/kvm/arm.c +++ b/arch/arm64/kvm/arm.c @@ -43,6 +43,7 @@ #include <asm/kvm_rme.h> #include <asm/sections.h> #include <asm/kvm_tmi.h> +#include <asm/tlbidomain.h> #include <kvm/arm_hypercalls.h> #include <kvm/arm_pmu.h> #include <kvm/arm_psci.h> @@ -359,6 +360,8 @@ static int kvm_arm_init_tlbidomain(struct kvm *kvm) for (i = 0; i < max_vdomains; i++) kvm->arch.vdomain.domain_map[i] = -1; + spin_lock_init(&kvm->arch.vdomain.tlbid_lock); + return 0; free_vdomain_cpumasks: @@ -822,6 +825,213 @@ static void update_steal_time(struct kvm_vcpu *vcpu) } #endif +static bool is_vcpu_in_vdomain(u32 vdomain_bitmap, int vdomain_idx) +{ + return vdomain_bitmap & BIT(vdomain_idx); +} + +static int update_domain_map(struct kvm_vcpu *vcpu, int vdomain_idx, + int last_pcpu, bool clear_flag) +{ + struct tlbidomain *vdomain = &vcpu->kvm->arch.vdomain; + int pdomain = vdomain->domain_map[vdomain_idx]; + cpumask_var_t new_cpus; + + cpumask_copy(new_cpus, vdomain->vdomain_cpumasks[vdomain_idx]); + + if (unlikely(pdomain == -1)) { + cpumask_set_cpu(vcpu->cpu, new_cpus); + goto update_mask; + } + + if (last_pcpu != -1 && clear_flag) { + /* remove old_cpu, not need to do so when first load */ + if (cpumask_test_cpu(last_pcpu, new_cpus)) { + cpumask_clear_cpu(last_pcpu, new_cpus); + } + } + + cpumask_set_cpu(vcpu->cpu, new_cpus); + +update_mask: + cpumask_copy(vdomain->vdomain_cpumasks[vdomain_idx], new_cpus); + return pick_best_domain(vdomain->vdomain_cpumasks[vdomain_idx]); +} + +static void vtlbidn_clear_set_s(int idx, u64 clear, u64 set) +{ + switch (idx) { + case 0: + sysreg_clear_set_s(SYS_VTLBID0_EL2, clear, set); + break; + case 1: + sysreg_clear_set_s(SYS_VTLBID1_EL2, clear, set); + break; + case 2: + sysreg_clear_set_s(SYS_VTLBID2_EL2, clear, set); + break; + case 3: + sysreg_clear_set_s(SYS_VTLBID3_EL2, clear, set); + break; + default: + BUG_ON(1); + } +} + +static void kvm_arm_update_tlbid_map(struct tlbidomain *vdomain) +{ + int i, pdomain, vtlbid_idx, offset; + int pdomain_bits, nis = vdomain->nis; + int max_vdomains = 1 << vdomain->nvis; + u64 clear, set; + + if (nis >= NIS_RANGE_MIN_1_8 && nis <= NIS_RANGE_MAX_1_8) + pdomain_bits = VTLBID_TD_WIDTH_NIS_1_8; + else if (nis >= NIS_RANGE_MIN_9_16 && nis <= NIS_RANGE_MAX_9_16) + pdomain_bits = VTLBID_TD_WIDTH_NIS_9_16; + + for (i = 0; i < max_vdomains; i++) { + pdomain = vdomain->domain_map[i]; + if (pdomain == -1) { + if (vdomain->domain_map[0] != -1) + pdomain = vdomain->domain_map[0]; + else + pdomain = 0; + } + + vtlbid_idx = i * pdomain_bits / VTLBID_EL2_WIDTH; + offset = i * pdomain_bits % VTLBID_EL2_WIDTH; + + clear = GENMASK(offset + pdomain_bits - 1, offset); + set = pdomain << offset; + + vtlbidn_clear_set_s(vtlbid_idx, clear, set); + } + + sysreg_clear_set_s(SYS_HCRX_EL2, 0, HCRX_EL2_VTLBIDEn); +} + +static void kvm_vcpu_reload_tlbid(struct kvm *kvm) +{ + if (WARN_ON_ONCE(!kvm->arch.vdomain.kvm_tlbid_enabled)) + return; + + preempt_disable(); + kvm_arm_update_tlbid_map(&kvm->arch.vdomain); + preempt_enable(); +} + +static void kvm_tlbidomain_vcpu_load(struct kvm_vcpu *vcpu) +{ + struct kvm *kvm = vcpu->kvm; + struct tlbidomain *vdomain = &kvm->arch.vdomain; + int max_vdomains = 1 << vdomain->nvis; + DECLARE_BITMAP(clear_bits, max_vdomains); + struct vcpu_tlbid_data *current_data, *other_data; + int i; + + /* + * If we support tlbid, but user does not enable it, we can do + * vdomain0 -> pdomainN as well by default. + */ + if (!vdomain->kvm_tlbid_enabled) + return; + + current_data = xa_load(&vdomain->vcpu_data_array, vcpu->vcpu_idx); + if (!current_data) { + BUG(); + return; + } + + /* check if vCPU thread will move to another pCPU */ + if (likely(vcpu->cpu == current_data->last_pcpu)) + return; + + spin_lock(&vcpu->kvm->arch.vdomain.tlbid_lock); + + /* + * check if another vcpu running on the same pCPU as the one + * the vCPU last ran. If so, do not clear last pCPU from the + * vdomain_cpumask. + */ + bitmap_fill(clear_bits, max_vdomains); + for (i = 0; i < kvm->created_vcpus; i++) { + other_data = xa_load(&vdomain->vcpu_data_array, i); + if (WARN_ON_ONCE(!other_data)) + continue; + + if (i == vcpu->vcpu_idx || other_data->last_pcpu == -1) + continue; + + if (other_data->last_pcpu != current_data->last_pcpu) + continue; + + bitmap_andnot(clear_bits, clear_bits, + (unsigned long *)&other_data->vdomain_bitmap, max_vdomains); + } + + if (!vdomain->guest_tlbid_enabled) { + /* + * If QEMU is not configured with vdomains, only vDomain0 needs + * to be mapped. + */ + vdomain->domain_map[0] = update_domain_map(vcpu, + 0, + current_data->last_pcpu, + test_bit(0, clear_bits)); + } else { + /* + * If QEMU is configured with vdomains, all vdomains need to be + * mapped. For all vdomains which includes this vCPU, check if + * vdomain map should be changed as well. + */ + for (i = 0; i < vdomain->num_domains; i++) { + if (!is_vcpu_in_vdomain(current_data->vdomain_bitmap, i)) + continue; + + vdomain->domain_map[i] = update_domain_map(vcpu, + i, + current_data->last_pcpu, + test_bit(i, clear_bits)); + } + } + + kvm_flush_remote_tlbs(kvm); + + /* + * Before this vcpu load, kick other vCPUs out, so maps for other vCPUs + * will be updated during vCPU load. + * + * Add KVM_REQUEST_WAIT to make sure vCPU is out. + */ + kvm_make_all_cpus_request(kvm, KVM_REQ_RELOAD_VTLBID | KVM_REQUEST_WAIT); + + /* update tlbid hardware map register */ + kvm_arm_update_tlbid_map(vdomain); + + current_data->last_pcpu = vcpu->cpu; + + spin_unlock(&vcpu->kvm->arch.vdomain.tlbid_lock); +} + +static void kvm_tlbidomain_vcpu_put(struct kvm_vcpu *vcpu) +{ + struct kvm *kvm = vcpu->kvm; + int i; + + if (!kvm->arch.vdomain.kvm_tlbid_enabled) + return; + + sysreg_clear_set_s(SYS_HCRX_EL2, HCRX_EL2_VTLBIDEn, 0); + + /* + * Clear all VTLBID registers so that no stale vdomain-to-pdomain + * mapping leaks into the next vCPU (or host) that runs on this pCPU. + */ + for (i = 0; i < VTLBID_EL2_MAX_REGS; i++) + vtlbidn_clear_set_s(i, VTLBID0_EL2_TD, 0); +} + void kvm_arch_vcpu_load(struct kvm_vcpu *vcpu, int cpu) { struct kvm_s2_mmu *mmu; @@ -877,6 +1087,8 @@ void kvm_arch_vcpu_load(struct kvm_vcpu *vcpu, int cpu) kvm_tlbi_dvmbm_vcpu_load(vcpu); + kvm_tlbidomain_vcpu_load(vcpu); + /* * When pv_preempted is changed from enabled to disabled, preempted * state will not be updated in kvm_arch_vcpu_put/load. So we must @@ -915,6 +1127,8 @@ void kvm_arch_vcpu_put(struct kvm_vcpu *vcpu) kvm_tlbi_dvmbm_vcpu_put(vcpu); + kvm_tlbidomain_vcpu_put(vcpu); + if (kvm_arm_is_pvsched_valid(&vcpu->arch) && pv_preempted_enable) kvm_update_pvsched_preempted(vcpu, 1); } @@ -1280,6 +1494,9 @@ static int check_vcpu_requests(struct kvm_vcpu *vcpu) if (kvm_check_request(KVM_REQ_RELOAD_TLBI_DVMBM, vcpu)) kvm_hisi_reload_lsudvmbm(vcpu->kvm); + if (kvm_check_request(KVM_REQ_RELOAD_VTLBID, vcpu)) + kvm_vcpu_reload_tlbid(vcpu->kvm); + if (kvm_check_request(KVM_REQ_RELOAD_WFI_TRAPS, vcpu)) { if (single_task_running()) vcpu_clear_wfx_traps(vcpu); -- 2.33.0
virt inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- Expose TLBID virtualization interface to userspace. Signed-off-by: Jinqian Yang <yangjinqian1@huawei.com> --- arch/arm64/kvm/arm.c | 33 ++++++++++++++++++++++++++++++++- 1 file changed, 32 insertions(+), 1 deletion(-) diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c index 7985f85ab3cf..061734d1ad95 100644 --- a/arch/arm64/kvm/arm.c +++ b/arch/arm64/kvm/arm.c @@ -467,6 +467,8 @@ int kvm_arch_init_vm(struct kvm *kvm, unsigned long type) bitmap_zero(kvm->arch.vcpu_features, KVM_VCPU_MAX_FEATURES); + kvm_arm_init_tlbidomain(kvm); + /* Initialise the realm bits after the generic bits are enabled */ if (kvm_is_realm(kvm)) { ret = kvm_init_realm_vm(kvm); @@ -511,6 +513,7 @@ void kvm_arch_destroy_vm(struct kvm *kvm) kvm_arm_teardown_hypercalls(kvm); kvm_destroy_realm(kvm); + free_tlbid_vdomain(kvm); } #ifdef CONFIG_ARM64_HISI_IPIV @@ -663,6 +666,12 @@ int kvm_vm_ioctl_check_extension(struct kvm *kvm, long ext) r = 0; break; #endif + case KVM_CAP_ARM_TLBIDOMAIN: + if (system_supports_tlbid()) + r = 1; + else + r = 0; + break; case KVM_CAP_ARM_RME: r = static_key_enabled(&kvm_rme_is_available); break; @@ -2152,7 +2161,9 @@ static int kvm_arch_vcpu_ioctl_vcpu_init(struct kvm_vcpu *vcpu, kvm_arm_pvtimer_status_set_active(vcpu, false); #endif - return 0; + ret = kvm_arm_tlbidomain_vcpu_init(vcpu); + + return ret; } static int kvm_arm_vcpu_set_attr(struct kvm_vcpu *vcpu, @@ -2403,6 +2414,17 @@ long kvm_arch_vcpu_ioctl(struct file *filp, return -EFAULT; return kvm_arm_vcpu_rmm_psci_complete(vcpu, &req); } + case KVM_ARM_VCPU_SET_VDOMAIN: { + struct kvm_arm_set_vdomain vdomain; + + if (!kvm_vcpu_initialized(vcpu)) + return -ENOEXEC; + + if (copy_from_user(&vdomain, argp, sizeof(vdomain))) + return -EFAULT; + + return kvm_arm_vcpu_set_vdomain(vcpu, &vdomain); + } default: r = -EINVAL; } @@ -2583,6 +2605,15 @@ int kvm_arch_vm_ioctl(struct file *filp, unsigned int ioctl, unsigned long arg) return -EFAULT; return kvm_vm_ioctl_get_reg_writable_masks(kvm, &range); } + case KVM_ARM_GET_VDOMAIN_NUM: { + struct kvm_arm_get_vdomain vdomain; + + kvm_arm_get_vdomain_num(kvm, &vdomain); + if (copy_to_user(argp, &vdomain, sizeof(vdomain))) + return -EFAULT; + + return 0; + } default: return -EINVAL; } -- 2.33.0
virt inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- DVMBM and TLBID are two mutually exclusive features. When BIOS enables DVMBM, the TLBID field in MMFR4 is masked, thus the OS will not enable TLBID. However, when BIOS enables TLBID, the DVMBM field in AIDR remains set to true, causing DVMBM to be enabled in the OS. This issue can be intercepted and handled at the software level. Signed-off-by: Jinqian Yang <yangjinqian1@huawei.com> --- arch/arm64/kvm/hisilicon/hisi_virt.c | 9 +++++++++ 1 file changed, 9 insertions(+) diff --git a/arch/arm64/kvm/hisilicon/hisi_virt.c b/arch/arm64/kvm/hisilicon/hisi_virt.c index 9a04e4052abd..c5dda4cd2f83 100644 --- a/arch/arm64/kvm/hisilicon/hisi_virt.c +++ b/arch/arm64/kvm/hisilicon/hisi_virt.c @@ -243,6 +243,15 @@ bool hisi_dvmbm_supported(void) return false; } + /* + * When TLBID is enabled, DVMBM cannot be enabled. + * After BIOS enables TLBID, the DVMBM field of AIDR + * remains set, therefore interception needs to be + * performed in the code. + */ + if (system_supports_tlbid()) + return false; + /* Determine whether DVMBM is supported by the hardware */ if (!(read_sysreg(aidr_el1) & AIDR_EL1_DVMBM_MASK)) return false; -- 2.33.0
virt inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- Set the NVIS/NVOS fields to populate NIS/NOS and expose TLBIDIDR_EL1 to the guest. Signed-off-by: Jinqian Yang <yangjinqian1@huawei.com> Signed-off-by: Tian Zheng <zhengtian10@huawei.com> --- arch/arm64/include/asm/kvm_host.h | 3 +++ arch/arm64/kvm/sys_regs.c | 28 ++++++++++++++++++++++++++++ 2 files changed, 31 insertions(+) diff --git a/arch/arm64/include/asm/kvm_host.h b/arch/arm64/include/asm/kvm_host.h index 18c95e605d8a..8cea52bd34d2 100644 --- a/arch/arm64/include/asm/kvm_host.h +++ b/arch/arm64/include/asm/kvm_host.h @@ -508,6 +508,9 @@ enum vcpu_sysreg { CNTHV_CTL_EL2, CNTHV_CVAL_EL2, + /* TLBI Domains Identification Register (EL1) */ + KABI_EXTEND_ENUM(TLBIDIDR_EL1) + NR_SYS_REGS /* Nothing after this line! */ }; diff --git a/arch/arm64/kvm/sys_regs.c b/arch/arm64/kvm/sys_regs.c index fc2a1bc3c314..8e478ddd8341 100644 --- a/arch/arm64/kvm/sys_regs.c +++ b/arch/arm64/kvm/sys_regs.c @@ -276,6 +276,17 @@ static bool access_vm_reg(struct kvm_vcpu *vcpu, return true; } +static bool access_tlbididr(struct kvm_vcpu *vcpu, + struct sys_reg_params *p, + const struct sys_reg_desc *r) +{ + if (p->is_write) + return ignore_write(vcpu, p); + + p->regval = vcpu_read_sys_reg(vcpu, r->reg); + return true; +} + static bool access_actlr(struct kvm_vcpu *vcpu, struct sys_reg_params *p, const struct sys_reg_desc *r) @@ -715,6 +726,22 @@ static u64 reset_amair_el1(struct kvm_vcpu *vcpu, const struct sys_reg_desc *r) return amair; } +static u64 reset_tlbididr_el1(struct kvm_vcpu *vcpu, const struct sys_reg_desc *r) +{ + if (!system_supports_tlbid()) + return 0; + + u64 tlbididr = read_sysreg_s(SYS_TLBIDIDR_EL1); + u64 nvis = FIELD_GET(TLBIDIDR_EL1_NVIS_MASK, tlbididr); + u64 nvos = FIELD_GET(TLBIDIDR_EL1_NVOS_MASK, tlbididr); + + tlbididr = FIELD_PREP(TLBIDIDR_EL1_NIS_MASK, nvis); + tlbididr |= FIELD_PREP(TLBIDIDR_EL1_NOS_MASK, nvos); + + vcpu_write_sys_reg(vcpu, tlbididr, TLBIDIDR_EL1); + return tlbididr; +} + static u64 reset_actlr(struct kvm_vcpu *vcpu, const struct sys_reg_desc *r) { u64 actlr = read_sysreg(actlr_el1); @@ -2633,6 +2660,7 @@ static const struct sys_reg_desc sys_reg_descs[] = { { SYS_DESC(SYS_LORN_EL1), trap_loregion }, { SYS_DESC(SYS_LORC_EL1), trap_loregion }, { SYS_DESC(SYS_MPAMIDR_EL1), workaround_bad_mpam_abi }, + { SYS_DESC(SYS_TLBIDIDR_EL1), access_tlbididr, reset_tlbididr_el1, TLBIDIDR_EL1 }, { SYS_DESC(SYS_LORID_EL1), trap_loregion }, { SYS_DESC(SYS_MPAM1_EL1), workaround_bad_mpam_abi }, -- 2.33.0
virt inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- Add the ID_AA64MMFR4_EL1 register and set it to be writable in userspace. When configuring vTLBID in QEMU, expose the TLBID field to userspace. Signed-off-by: Jinqian Yang <yangjinqian1@huawei.com> Signed-off-by: Tian Zheng <zhengtian10@huawei.com> --- arch/arm64/include/uapi/asm/kvm.h | 2 ++ arch/arm64/kvm/arm.c | 4 ++++ arch/arm64/kvm/sys_regs.c | 28 +++++++++++++++++++++++++++- arch/arm64/kvm/sys_regs.h | 2 ++ 4 files changed, 35 insertions(+), 1 deletion(-) diff --git a/arch/arm64/include/uapi/asm/kvm.h b/arch/arm64/include/uapi/asm/kvm.h index 93de3f019e5b..99c0e43c43dd 100644 --- a/arch/arm64/include/uapi/asm/kvm.h +++ b/arch/arm64/include/uapi/asm/kvm.h @@ -253,6 +253,8 @@ struct kvm_arm_counter_offset { #define ARM64_SYS_REG(...) (__ARM64_SYS_REG(__VA_ARGS__) | KVM_REG_SIZE_U64) +#define KVM_REG_ARM_ID_AA64MMFR4_EL1 ARM64_SYS_REG(3, 0, 0, 7, 4) + /* Physical Timer EL0 Registers */ #define KVM_REG_ARM_PTIMER_CTL ARM64_SYS_REG(3, 3, 14, 2, 1) #define KVM_REG_ARM_PTIMER_CVAL ARM64_SYS_REG(3, 3, 14, 2, 2) diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c index 061734d1ad95..be9f64883eac 100644 --- a/arch/arm64/kvm/arm.c +++ b/arch/arm64/kvm/arm.c @@ -51,6 +51,7 @@ static enum kvm_mode kvm_mode = KVM_MODE_DEFAULT; #include "hisilicon/hisi_virt.h" +#include "sys_regs.h" DEFINE_STATIC_KEY_FALSE(kvm_rme_is_available); @@ -2258,6 +2259,7 @@ static int kvm_arm_vcpu_set_vdomain(struct kvm_vcpu *vcpu, u8 nvis = vcpu->kvm->arch.vdomain.nvis; u8 max_vdomains = 1 << nvis; struct vcpu_tlbid_data *data; + int ret; if (num_vdomains > max_vdomains) return -EINVAL; @@ -2278,6 +2280,8 @@ static int kvm_arm_vcpu_set_vdomain(struct kvm_vcpu *vcpu, kvm_info("vCPU%d: vDomain bitmap: 0x%x\n", vcpu->vcpu_idx, vdomain_bitmap); + kvm_update_aa64mmfr4_tlbid(vcpu->kvm); + return 0; } diff --git a/arch/arm64/kvm/sys_regs.c b/arch/arm64/kvm/sys_regs.c index 8e478ddd8341..4924e4ed78ca 100644 --- a/arch/arm64/kvm/sys_regs.c +++ b/arch/arm64/kvm/sys_regs.c @@ -1472,6 +1472,10 @@ static u64 __kvm_read_sanitised_id_reg(const struct kvm_vcpu *vcpu, val &= ID_AA64MMFR3_EL1_TCRX | ID_AA64MMFR3_EL1_S1POE | ID_AA64MMFR3_EL1_S1PIE; break; + case SYS_ID_AA64MMFR4_EL1: + if (!vcpu->kvm->arch.vdomain.guest_tlbid_enabled) + val &= ~ID_AA64MMFR4_EL1_TLBID; + break; case SYS_ID_MMFR4_EL1: val &= ~ARM64_FEATURE_MASK(ID_MMFR4_EL1_CCIDX); break; @@ -2579,7 +2583,7 @@ static const struct sys_reg_desc sys_reg_descs[] = { ID_WRITABLE(ID_AA64MMFR3_EL1, (ID_AA64MMFR3_EL1_TCRX | ID_AA64MMFR3_EL1_S1PIE | ID_AA64MMFR3_EL1_S1POE)), - ID_UNALLOCATED(7,4), + ID_WRITABLE(ID_AA64MMFR4_EL1, ID_AA64MMFR4_EL1_TLBID_MASK), ID_UNALLOCATED(7,5), ID_UNALLOCATED(7,6), ID_UNALLOCATED(7,7), @@ -3799,6 +3803,28 @@ const struct sys_reg_desc *get_reg_by_id(u64 id, return find_reg(¶ms, table, num); } +/* + * Update the VM's stored value for an ID register. + */ +void kvm_update_aa64mmfr4_tlbid(struct kvm *kvm) +{ + u64 val, tlbid_val; + + mutex_lock(&kvm->arch.config_lock); + + val = kvm_read_vm_id_reg(kvm, SYS_ID_AA64MMFR4_EL1); + tlbid_val = read_sanitised_ftr_reg(SYS_ID_AA64MMFR4_EL1); + tlbid_val &= ID_AA64MMFR4_EL1_TLBID_MASK; + + val &= ~ID_AA64MMFR4_EL1_TLBID_MASK; + if (kvm->arch.vdomain.guest_tlbid_enabled) + val |= tlbid_val; + + kvm_set_vm_id_reg(kvm, SYS_ID_AA64MMFR4_EL1, val); + + mutex_unlock(&kvm->arch.config_lock); +} + /* Decode an index value, and find the sys_reg_desc entry. */ static const struct sys_reg_desc * id_to_sys_reg_desc(struct kvm_vcpu *vcpu, u64 id, diff --git a/arch/arm64/kvm/sys_regs.h b/arch/arm64/kvm/sys_regs.h index 3080693719d2..4c5ae75bc9eb 100644 --- a/arch/arm64/kvm/sys_regs.h +++ b/arch/arm64/kvm/sys_regs.h @@ -238,6 +238,8 @@ int kvm_sys_reg_get_user(struct kvm_vcpu *vcpu, const struct kvm_one_reg *reg, int kvm_sys_reg_set_user(struct kvm_vcpu *vcpu, const struct kvm_one_reg *reg, const struct sys_reg_desc table[], unsigned int num); +void kvm_update_aa64mmfr4_tlbid(struct kvm *kvm); + #define AA32(_x) .aarch32_map = AA32_##_x #define Op0(_x) .Op0 = _x #define Op1(_x) .Op1 = _x -- 2.33.0
virt inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- Enable the TLBIDIDR_EL1 trap when the vCPU loading. Disable it when the vCPU putting. Signed-off-by: Jinqian Yang <yangjinqian1@huawei.com> Signed-off-by: Tian Zheng <zhengtian10@huawei.com> --- arch/arm64/kvm/arm.c | 21 +++++++++++++++++++++ arch/arm64/tools/sysreg | 4 +++- 2 files changed, 24 insertions(+), 1 deletion(-) diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c index be9f64883eac..bdeeefe0eea1 100644 --- a/arch/arm64/kvm/arm.c +++ b/arch/arm64/kvm/arm.c @@ -931,6 +931,15 @@ static void kvm_vcpu_reload_tlbid(struct kvm *kvm) preempt_enable(); } +static inline void set_tlbididr_trap(void) +{ + u64 val; + val = read_sysreg_s(SYS_HFGRTR2_EL2); + val &= ~HFGRTR2_EL2_nTLBIDIDR_EL1_MASK; + write_sysreg_s(val, SYS_HFGRTR2_EL2); + isb(); +} + static void kvm_tlbidomain_vcpu_load(struct kvm_vcpu *vcpu) { struct kvm *kvm = vcpu->kvm; @@ -947,6 +956,8 @@ static void kvm_tlbidomain_vcpu_load(struct kvm_vcpu *vcpu) if (!vdomain->kvm_tlbid_enabled) return; + set_tlbididr_trap(); + current_data = xa_load(&vdomain->vcpu_data_array, vcpu->vcpu_idx); if (!current_data) { BUG(); @@ -1024,6 +1035,15 @@ static void kvm_tlbidomain_vcpu_load(struct kvm_vcpu *vcpu) spin_unlock(&vcpu->kvm->arch.vdomain.tlbid_lock); } +static inline void clear_tlbididr_trap(void) +{ + u64 val; + val = read_sysreg_s(SYS_HFGRTR2_EL2); + val |= HFGRTR2_EL2_nTLBIDIDR_EL1_MASK; + write_sysreg_s(val, SYS_HFGRTR2_EL2); + isb(); +} + static void kvm_tlbidomain_vcpu_put(struct kvm_vcpu *vcpu) { struct kvm *kvm = vcpu->kvm; @@ -1032,6 +1052,7 @@ static void kvm_tlbidomain_vcpu_put(struct kvm_vcpu *vcpu) if (!kvm->arch.vdomain.kvm_tlbid_enabled) return; + clear_tlbididr_trap(); sysreg_clear_set_s(SYS_HCRX_EL2, HCRX_EL2_VTLBIDEn, 0); /* diff --git a/arch/arm64/tools/sysreg b/arch/arm64/tools/sysreg index c10f6985f5db..8896f2310d23 100644 --- a/arch/arm64/tools/sysreg +++ b/arch/arm64/tools/sysreg @@ -3001,7 +3001,9 @@ Field 0 nPMECR_EL1 EndSysreg Sysreg HFGRTR2_EL2 3 4 3 1 2 -Res0 63:15 +Res0 63:31 +Field 30 nTLBIDIDR_EL1 +Res0 29:15 Field 14 nACTLRALIAS_EL1 Field 13 nACTLRMASK_EL1 Field 12 nTCR2ALIAS_EL1 -- 2.33.0
virt inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- In kvm_tlbidomain_vcpu_load(), when a vCPU migrates to a different pCPU, the code iterates over all created vCPUs to find those whose last_pcpu matches the current vCPU last_pcpu. This O(logN) loop with xa_load() becomes a significant bottleneck when the number of vCPUs is large. Introduce a per-pCPU linked list (pcpu_vcpu_list) in struct tlbidomain to maintain reverse mapping from pCPU to vCPUs. Each vcpu_tlbid_data gets a pcpu_node to link into the list of its last_pcpu. This allows the clear_bits calculation to only iterate over vCPUs that actually share the same last_pcpu, reducing the complexity from O(logN) to O(1) in the typical case. Signed-off-by: Jinqian Yang <yangjinqian1@huawei.com> --- arch/arm64/include/asm/kvm_host.h | 4 +++ arch/arm64/kvm/arm.c | 46 ++++++++++++++++++++----------- 2 files changed, 34 insertions(+), 16 deletions(-) diff --git a/arch/arm64/include/asm/kvm_host.h b/arch/arm64/include/asm/kvm_host.h index 8cea52bd34d2..d96c41719732 100644 --- a/arch/arm64/include/asm/kvm_host.h +++ b/arch/arm64/include/asm/kvm_host.h @@ -219,6 +219,8 @@ struct kvm_protected_vm { }; struct vcpu_tlbid_data { + /* Node in per-pCPU vCPU list of tlbidomain */ + struct list_head pcpu_node; /* Store vDomain bitmap, indexed by vCPU ID */ u32 vdomain_bitmap; /* Record the last pCPU where vCPU was running, indexed by vCPU ID */ @@ -244,6 +246,8 @@ struct tlbidomain { u8 nis; /* Used to prevent concurrent modifications to the domain mapping. */ spinlock_t tlbid_lock; + /* Per-pCPU list of vCPUs whose last_pcpu equals this pCPU */ + struct list_head *pcpu_vcpu_list; }; struct kvm_arch { diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c index bdeeefe0eea1..3feb8c028268 100644 --- a/arch/arm64/kvm/arm.c +++ b/arch/arm64/kvm/arm.c @@ -363,6 +363,14 @@ static int kvm_arm_init_tlbidomain(struct kvm *kvm) spin_lock_init(&kvm->arch.vdomain.tlbid_lock); + vdomain->pcpu_vcpu_list = kcalloc(nr_cpu_ids, + sizeof(*vdomain->pcpu_vcpu_list), GFP_KERNEL); + if (!vdomain->pcpu_vcpu_list) + goto free_vdomain_cpumasks; + + for (i = 0; i < nr_cpu_ids; i++) + INIT_LIST_HEAD(&vdomain->pcpu_vcpu_list[i]); + return 0; free_vdomain_cpumasks: @@ -392,6 +400,8 @@ static void free_tlbid_vdomain(struct kvm *kvm) free_cpumask_var(kvm->arch.vdomain.vdomain_cpumasks[i]); kfree(kvm->arch.vdomain.vdomain_cpumasks); + kfree(kvm->arch.vdomain.pcpu_vcpu_list); + xa_for_each(&kvm->arch.vdomain.vcpu_data_array, index, data) { kfree(data); } @@ -971,24 +981,25 @@ static void kvm_tlbidomain_vcpu_load(struct kvm_vcpu *vcpu) spin_lock(&vcpu->kvm->arch.vdomain.tlbid_lock); /* - * check if another vcpu running on the same pCPU as the one - * the vCPU last ran. If so, do not clear last pCPU from the - * vdomain_cpumask. + * Remove current vCPU from its last_pcpu list and compute clear_bits + * only when there was a previous pCPU. On first load (last_pcpu == -1), + * no previous pCPU needs to be cleared from vdomain_cpumask. */ - bitmap_fill(clear_bits, max_vdomains); - for (i = 0; i < kvm->created_vcpus; i++) { - other_data = xa_load(&vdomain->vcpu_data_array, i); - if (WARN_ON_ONCE(!other_data)) - continue; - - if (i == vcpu->vcpu_idx || other_data->last_pcpu == -1) - continue; + if (current_data->last_pcpu != -1) { + list_del_init(¤t_data->pcpu_node); - if (other_data->last_pcpu != current_data->last_pcpu) - continue; - - bitmap_andnot(clear_bits, clear_bits, - (unsigned long *)&other_data->vdomain_bitmap, max_vdomains); + /* + * Iterate only vCPUs that share the same last_pcpu to compute + * clear_bits. This is O(1) in the typical case since most pCPUs + * host only one vCPU at a time. + */ + bitmap_fill(clear_bits, max_vdomains); + list_for_each_entry(other_data, + &vdomain->pcpu_vcpu_list[current_data->last_pcpu], + pcpu_node) { + bitmap_andnot(clear_bits, clear_bits, + (unsigned long *)&other_data->vdomain_bitmap, max_vdomains); + } } if (!vdomain->guest_tlbid_enabled) { @@ -1031,6 +1042,8 @@ static void kvm_tlbidomain_vcpu_load(struct kvm_vcpu *vcpu) kvm_arm_update_tlbid_map(vdomain); current_data->last_pcpu = vcpu->cpu; + list_add_tail(¤t_data->pcpu_node, + &vdomain->pcpu_vcpu_list[vcpu->cpu]); spin_unlock(&vcpu->kvm->arch.vdomain.tlbid_lock); } @@ -2112,6 +2125,7 @@ static int kvm_arm_tlbidomain_vcpu_init(struct kvm_vcpu *vcpu) */ data->vdomain_bitmap = 0x1; data->last_pcpu = -1; + INIT_LIST_HEAD(&data->pcpu_node); ret = xa_insert(&vcpu->kvm->arch.vdomain.vcpu_data_array, vcpu->vcpu_idx, data, GFP_KERNEL); -- 2.33.0
virt inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- In kvm_tlbidomain_vcpu_load(), the domain map is recomputed on every vCPU load, followed by an expensive kvm_flush_remote_tlbs() and kvm_make_all_cpus_request() to invalidate TLBs and IPI all vCPUs. However, when a vCPU is scheduled back onto the same pCPU or the pDomain assignment does not change, the domain map remains the same and the invalidation + IPI is unnecessary. Record the old domain map value before recomputing, and compare it with the new one. If no vDomain's mapping has changed, skip the TLB flush and vCPU IPI, jumping directly to update the tlbid hardware map register. This avoids redundant cross-CPU IPIs and TLB shootdowns in the common case where a vCPU migrates back to its previous pCPU, significantly reducing the overhead of the vCPU load path. Signed-off-by: Jinqian Yang <yangjinqian1@huawei.com> --- arch/arm64/kvm/arm.c | 15 ++++++++++++++- 1 file changed, 14 insertions(+), 1 deletion(-) diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c index 3feb8c028268..39d91378c182 100644 --- a/arch/arm64/kvm/arm.c +++ b/arch/arm64/kvm/arm.c @@ -957,7 +957,8 @@ static void kvm_tlbidomain_vcpu_load(struct kvm_vcpu *vcpu) int max_vdomains = 1 << vdomain->nvis; DECLARE_BITMAP(clear_bits, max_vdomains); struct vcpu_tlbid_data *current_data, *other_data; - int i; + bool domain_map_changed = false; + int i, old_map; /* * If we support tlbid, but user does not enable it, we can do @@ -1003,6 +1004,8 @@ static void kvm_tlbidomain_vcpu_load(struct kvm_vcpu *vcpu) } if (!vdomain->guest_tlbid_enabled) { + old_map = vdomain->domain_map[0]; + /* * If QEMU is not configured with vdomains, only vDomain0 needs * to be mapped. @@ -1011,6 +1014,8 @@ static void kvm_tlbidomain_vcpu_load(struct kvm_vcpu *vcpu) 0, current_data->last_pcpu, test_bit(0, clear_bits)); + if (vdomain->domain_map[0] == old_map) + goto unlock; } else { /* * If QEMU is configured with vdomains, all vdomains need to be @@ -1021,11 +1026,18 @@ static void kvm_tlbidomain_vcpu_load(struct kvm_vcpu *vcpu) if (!is_vcpu_in_vdomain(current_data->vdomain_bitmap, i)) continue; + old_map = vdomain->domain_map[i]; vdomain->domain_map[i] = update_domain_map(vcpu, i, current_data->last_pcpu, test_bit(i, clear_bits)); + + if (vdomain->domain_map[i] != old_map) + domain_map_changed = true; } + + if (!domain_map_changed) + goto unlock; } kvm_flush_remote_tlbs(kvm); @@ -1038,6 +1050,7 @@ static void kvm_tlbidomain_vcpu_load(struct kvm_vcpu *vcpu) */ kvm_make_all_cpus_request(kvm, KVM_REQ_RELOAD_VTLBID | KVM_REQUEST_WAIT); +unlock: /* update tlbid hardware map register */ kvm_arm_update_tlbid_map(vdomain); -- 2.33.0
virt inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- In kvm_tlbidomain_vcpu_load(), three trace_printk() calls are used to log vCPU load info and vDomain-to-pDomain mapping. trace_printk() is a debug-only facility that is not meant for production use: it always prints to the ftrace ring buffer when enabled, cannot be dynamically controlled per-event, and may impact performance. Signed-off-by: Jinqian Yang <yangjinqian1@huawei.com> --- arch/arm64/kvm/arm.c | 10 ++++++++ arch/arm64/kvm/trace_arm.h | 47 ++++++++++++++++++++++++++++++++++++++ 2 files changed, 57 insertions(+) diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c index 39d91378c182..16202f89ae6d 100644 --- a/arch/arm64/kvm/arm.c +++ b/arch/arm64/kvm/arm.c @@ -1002,6 +1002,9 @@ static void kvm_tlbidomain_vcpu_load(struct kvm_vcpu *vcpu) (unsigned long *)&other_data->vdomain_bitmap, max_vdomains); } } + trace_kvm_tlbid_vcpu_info(vcpu->vcpu_idx, vcpu->cpu, + current_data->last_pcpu, + current_data->vdomain_bitmap); if (!vdomain->guest_tlbid_enabled) { old_map = vdomain->domain_map[0]; @@ -1014,6 +1017,10 @@ static void kvm_tlbidomain_vcpu_load(struct kvm_vcpu *vcpu) 0, current_data->last_pcpu, test_bit(0, clear_bits)); + trace_kvm_tlbid_domain_mapping(0, + vdomain->vdomain_cpumasks[0], + vdomain->domain_map[0]); + if (vdomain->domain_map[0] == old_map) goto unlock; } else { @@ -1031,6 +1038,9 @@ static void kvm_tlbidomain_vcpu_load(struct kvm_vcpu *vcpu) i, current_data->last_pcpu, test_bit(i, clear_bits)); + trace_kvm_tlbid_domain_mapping(i, + vdomain->vdomain_cpumasks[i], + vdomain->domain_map[i]); if (vdomain->domain_map[i] != old_map) domain_map_changed = true; diff --git a/arch/arm64/kvm/trace_arm.h b/arch/arm64/kvm/trace_arm.h index 5b91456560b5..5f7fd7a487fc 100644 --- a/arch/arm64/kvm/trace_arm.h +++ b/arch/arm64/kvm/trace_arm.h @@ -408,6 +408,53 @@ TRACE_EVENT(kvm_pvspin_kick_vcpu, __entry->vcpu_id, __entry->target_vcpu_id) ); +TRACE_EVENT(kvm_tlbid_vcpu_info, + TP_PROTO(int vcpu_idx, int current_pcpu, int last_pcpu, + u32 vdomain_bitmap), + TP_ARGS(vcpu_idx, current_pcpu, last_pcpu, vdomain_bitmap), + + TP_STRUCT__entry( + __field(int, vcpu_idx ) + __field(int, current_pcpu ) + __field(int, last_pcpu ) + __field(u32, vdomain_bitmap ) + ), + + TP_fast_assign( + __entry->vcpu_idx = vcpu_idx; + __entry->current_pcpu = current_pcpu; + __entry->last_pcpu = last_pcpu; + __entry->vdomain_bitmap = vdomain_bitmap; + ), + + TP_printk("vcpu=%d current_pcpu=%d last_pcpu=%d vdomain_bitmap=0x%x", + __entry->vcpu_idx, __entry->current_pcpu, + __entry->last_pcpu, __entry->vdomain_bitmap) +); + +TRACE_EVENT(kvm_tlbid_domain_mapping, + TP_PROTO(int vdomain_id, const struct cpumask *cpumask, + int pdomain_id), + TP_ARGS(vdomain_id, cpumask, pdomain_id), + + TP_STRUCT__entry( + __field(int, vdomain_id ) + __cpumask(cpumask) + __field(int, pdomain_id ) + ), + + TP_fast_assign( + __entry->vdomain_id = vdomain_id; + __assign_cpumask(cpumask, cpumask_bits(cpumask)); + __entry->pdomain_id = pdomain_id; + ), + + TP_printk("Domain mapping: vDomain=%d, cpumask=%s, pDomain=%d", + __entry->vdomain_id, + __get_cpumask(cpumask), + __entry->pdomain_id) +); + #endif /* _TRACE_ARM_ARM64_KVM_H */ #undef TRACE_INCLUDE_PATH -- 2.33.0
From: Tian Zheng <zhengtian10@huawei.com> virt inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- kvm_update_aa64mmfr4_tlbid() calls kvm_set_vm_id_reg() which triggers KVM_BUG_ON when kvm_vm_has_ran_once() is true, causing a kernel panic on vCPU hotplug. The VM-level ID_AA64MMFR4.TLBID is already set during the first vCPU init and doesn't need to be rewritten. Skip kvm_set_vm_id_reg() when the VM has already started. Fixes: 83eec466269a ("KVM: arm64: Expose ID_AA64MMFR4_EL1_TLBID to guest") Signed-off-by: Tian Zheng <zhengtian10@huawei.com> --- arch/arm64/kvm/sys_regs.c | 11 +++++++++++ 1 file changed, 11 insertions(+) diff --git a/arch/arm64/kvm/sys_regs.c b/arch/arm64/kvm/sys_regs.c index 4924e4ed78ca..82b5b33b1760 100644 --- a/arch/arm64/kvm/sys_regs.c +++ b/arch/arm64/kvm/sys_regs.c @@ -3812,6 +3812,17 @@ void kvm_update_aa64mmfr4_tlbid(struct kvm *kvm) mutex_lock(&kvm->arch.config_lock); + if (kvm_vm_has_ran_once(kvm)) { + /* + * ID_AA64MMFR4 is a VM-level register already set during + * the first vCPU init. Skip the write to avoid triggering + * KVM_BUG_ON in kvm_set_vm_id_reg(). The per-vCPU + * vdomain_bitmap is still updated by the caller. + */ + mutex_unlock(&kvm->arch.config_lock); + return; + } + val = kvm_read_vm_id_reg(kvm, SYS_ID_AA64MMFR4_EL1); tlbid_val = read_sanitised_ftr_reg(SYS_ID_AA64MMFR4_EL1); tlbid_val &= ID_AA64MMFR4_EL1_TLBID_MASK; -- 2.33.0
virt inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- The enum sys_reg member TLBIDIDR_EL1 exists only to back the sys_reg_desc table entry. The guest-visible value is a VM-wide constant synthesized from the hardware TLBIDIDR (the NIS/NOS fields carry the virtual domain counts NVIS/NVOS), so cache it in struct tlbidomain at VM init and let access_tlbididr() read it directly. Removing the enum member leaves NR_SYS_REGS, and therefore the layout of struct kvm_vcpu_arch, unchanged, eliminating the need for KABI_EXTEND_ENUM. As a side effect TLBIDIDR_EL1 is no longer listed by KVM_GET_REG_LIST nor directly readable via KVM_GET_ONE_REG; userspace observes the synthesized value through the guest trap path, which is the intended interface. Signed-off-by: Jinqian Yang <yangjinqian1@huawei.com> --- arch/arm64/include/asm/kvm_host.h | 5 ++--- arch/arm64/kvm/arm.c | 7 +++++++ arch/arm64/kvm/sys_regs.c | 20 ++------------------ 3 files changed, 11 insertions(+), 21 deletions(-) diff --git a/arch/arm64/include/asm/kvm_host.h b/arch/arm64/include/asm/kvm_host.h index d96c41719732..29a47406714c 100644 --- a/arch/arm64/include/asm/kvm_host.h +++ b/arch/arm64/include/asm/kvm_host.h @@ -244,6 +244,8 @@ struct tlbidomain { u8 nvis; /* TLBIDIDR.NIS */ u8 nis; + /* Guest-visible TLBIDIDR value (NIS/NOS carry virtual domain counts) */ + u64 tlbididr_val; /* Used to prevent concurrent modifications to the domain mapping. */ spinlock_t tlbid_lock; /* Per-pCPU list of vCPUs whose last_pcpu equals this pCPU */ @@ -512,9 +514,6 @@ enum vcpu_sysreg { CNTHV_CTL_EL2, CNTHV_CVAL_EL2, - /* TLBI Domains Identification Register (EL1) */ - KABI_EXTEND_ENUM(TLBIDIDR_EL1) - NR_SYS_REGS /* Nothing after this line! */ }; diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c index 16202f89ae6d..b40c4548c7f1 100644 --- a/arch/arm64/kvm/arm.c +++ b/arch/arm64/kvm/arm.c @@ -336,6 +336,13 @@ static int kvm_arm_init_tlbidomain(struct kvm *kvm) vdomain->nvis = FIELD_GET(TLBIDIDR_EL1_NVIS_MASK, tlbididr); vdomain->num_domains = -1; vdomain->kvm_tlbid_enabled = true; + /* + * Guest-visible TLBIDIDR: the NIS/NOS fields carry the virtual + * domain counts (NVIS/NVOS), since the guest uses vDomains. + */ + vdomain->tlbididr_val = FIELD_PREP(TLBIDIDR_EL1_NIS_MASK, vdomain->nvis) | + FIELD_PREP(TLBIDIDR_EL1_NOS_MASK, + FIELD_GET(TLBIDIDR_EL1_NVOS_MASK, tlbididr)); max_vdomains = 1 << vdomain->nvis; kvm_info("TLBIDIDR: 0x%llx, NIS: %d, NVIS: %d\n", tlbididr, vdomain->nis, vdomain->nvis); diff --git a/arch/arm64/kvm/sys_regs.c b/arch/arm64/kvm/sys_regs.c index 82b5b33b1760..1ef578caa7c0 100644 --- a/arch/arm64/kvm/sys_regs.c +++ b/arch/arm64/kvm/sys_regs.c @@ -283,7 +283,7 @@ static bool access_tlbididr(struct kvm_vcpu *vcpu, if (p->is_write) return ignore_write(vcpu, p); - p->regval = vcpu_read_sys_reg(vcpu, r->reg); + p->regval = vcpu->kvm->arch.vdomain.tlbididr_val; return true; } @@ -726,22 +726,6 @@ static u64 reset_amair_el1(struct kvm_vcpu *vcpu, const struct sys_reg_desc *r) return amair; } -static u64 reset_tlbididr_el1(struct kvm_vcpu *vcpu, const struct sys_reg_desc *r) -{ - if (!system_supports_tlbid()) - return 0; - - u64 tlbididr = read_sysreg_s(SYS_TLBIDIDR_EL1); - u64 nvis = FIELD_GET(TLBIDIDR_EL1_NVIS_MASK, tlbididr); - u64 nvos = FIELD_GET(TLBIDIDR_EL1_NVOS_MASK, tlbididr); - - tlbididr = FIELD_PREP(TLBIDIDR_EL1_NIS_MASK, nvis); - tlbididr |= FIELD_PREP(TLBIDIDR_EL1_NOS_MASK, nvos); - - vcpu_write_sys_reg(vcpu, tlbididr, TLBIDIDR_EL1); - return tlbididr; -} - static u64 reset_actlr(struct kvm_vcpu *vcpu, const struct sys_reg_desc *r) { u64 actlr = read_sysreg(actlr_el1); @@ -2664,7 +2648,7 @@ static const struct sys_reg_desc sys_reg_descs[] = { { SYS_DESC(SYS_LORN_EL1), trap_loregion }, { SYS_DESC(SYS_LORC_EL1), trap_loregion }, { SYS_DESC(SYS_MPAMIDR_EL1), workaround_bad_mpam_abi }, - { SYS_DESC(SYS_TLBIDIDR_EL1), access_tlbididr, reset_tlbididr_el1, TLBIDIDR_EL1 }, + { SYS_DESC(SYS_TLBIDIDR_EL1), access_tlbididr }, { SYS_DESC(SYS_LORID_EL1), trap_loregion }, { SYS_DESC(SYS_MPAM1_EL1), workaround_bad_mpam_abi }, -- 2.33.0
virt inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- Keep struct tlbidomain in a global xarray keyed by the kvm pointer instead of embedding it in struct kvm_arch. This leaves the layout of struct kvm_arch untouched and drops the KABI_EXTEND wrapper. Callers look the state up via kvm_tlbidomain(), which costs one RCU-protected xa_load() on the vCPU load path, and must tolerate NULL when the VM has no TLBID support (no hardware support, or failed allocation at VM init). The xarray entry is published at the end of kvm_arm_init_tlbidomain() and erased in free_tlbid_vdomain(), so no entry exists for partially-initialized or unsupported VMs. Signed-off-by: Jinqian Yang <yangjinqian1@huawei.com> --- arch/arm64/include/asm/kvm_host.h | 8 ++- arch/arm64/kvm/arm.c | 115 ++++++++++++++++++++---------- arch/arm64/kvm/sys_regs.c | 16 +++-- 3 files changed, 96 insertions(+), 43 deletions(-) diff --git a/arch/arm64/include/asm/kvm_host.h b/arch/arm64/include/asm/kvm_host.h index 29a47406714c..ed2083b7b206 100644 --- a/arch/arm64/include/asm/kvm_host.h +++ b/arch/arm64/include/asm/kvm_host.h @@ -252,6 +252,13 @@ struct tlbidomain { struct list_head *pcpu_vcpu_list; }; +/* + * Externally-allocated per-VM TLBID domain state, registered in a global + * xarray keyed by the kvm pointer so that struct kvm_arch stays unchanged + * (no KABI impact). Returns NULL when the VM has no TLBID support. + */ +struct tlbidomain *kvm_tlbidomain(struct kvm *kvm); + struct kvm_arch { struct kvm_s2_mmu mmu; @@ -376,7 +383,6 @@ struct kvm_arch { KABI_EXTEND(u64 midr_el1) KABI_EXTEND(u64 revidr_el1) KABI_EXTEND(u64 aidr_el1) - KABI_EXTEND(struct tlbidomain vdomain) }; struct kvm_vcpu_fault_info { diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c index b40c4548c7f1..0025b0499211 100644 --- a/arch/arm64/kvm/arm.c +++ b/arch/arm64/kvm/arm.c @@ -321,17 +321,28 @@ static int kvm_arm_default_max_vcpus(void) return vgic_present ? kvm_vgic_get_max_vcpus() : KVM_MAX_VCPUS; } +static DEFINE_XARRAY(tlbid_domains); + +struct tlbidomain *kvm_tlbidomain(struct kvm *kvm) +{ + return xa_load(&tlbid_domains, (unsigned long)kvm); +} + static int kvm_arm_init_tlbidomain(struct kvm *kvm) { struct tlbidomain *vdomain; - int i, max_vdomains; + int i, max_vdomains, ret; u64 tlbididr; if (!system_supports_tlbid()) return 0; tlbididr = read_sysreg_s(SYS_TLBIDIDR_EL1); - vdomain = &kvm->arch.vdomain; + + vdomain = kzalloc(sizeof(*vdomain), GFP_KERNEL); + if (!vdomain) + return -ENOMEM; + vdomain->nis = FIELD_GET(TLBIDIDR_EL1_NIS_MASK, tlbididr); vdomain->nvis = FIELD_GET(TLBIDIDR_EL1_NVIS_MASK, tlbididr); vdomain->num_domains = -1; @@ -366,9 +377,9 @@ static int kvm_arm_init_tlbidomain(struct kvm *kvm) } for (i = 0; i < max_vdomains; i++) - kvm->arch.vdomain.domain_map[i] = -1; + vdomain->domain_map[i] = -1; - spin_lock_init(&kvm->arch.vdomain.tlbid_lock); + spin_lock_init(&vdomain->tlbid_lock); vdomain->pcpu_vcpu_list = kcalloc(nr_cpu_ids, sizeof(*vdomain->pcpu_vcpu_list), GFP_KERNEL); @@ -378,7 +389,11 @@ static int kvm_arm_init_tlbidomain(struct kvm *kvm) for (i = 0; i < nr_cpu_ids; i++) INIT_LIST_HEAD(&vdomain->pcpu_vcpu_list[i]); - return 0; + ret = xa_err(xa_store(&tlbid_domains, (unsigned long)kvm, vdomain, + GFP_KERNEL)); + if (!ret) + return 0; + i = max_vdomains; free_vdomain_cpumasks: while (--i >= 0) @@ -388,32 +403,37 @@ static int kvm_arm_init_tlbidomain(struct kvm *kvm) kfree(vdomain->domain_map); destroy_vcpu_data_array: xa_destroy(&vdomain->vcpu_data_array); + kfree(vdomain); return -ENOMEM; } static void free_tlbid_vdomain(struct kvm *kvm) { + struct tlbidomain *vdomain; struct vcpu_tlbid_data *data; int i, max_vdomains; unsigned long index; - max_vdomains = 1 << kvm->arch.vdomain.nvis; + vdomain = xa_erase(&tlbid_domains, (unsigned long)kvm); + if (!vdomain) + return; - if (system_supports_tlbid()) { - kfree(kvm->arch.vdomain.domain_map); + max_vdomains = 1 << vdomain->nvis; + + kfree(vdomain->domain_map); - for (i = 0; i < max_vdomains; i++) - free_cpumask_var(kvm->arch.vdomain.vdomain_cpumasks[i]); - kfree(kvm->arch.vdomain.vdomain_cpumasks); + for (i = 0; i < max_vdomains; i++) + free_cpumask_var(vdomain->vdomain_cpumasks[i]); + kfree(vdomain->vdomain_cpumasks); - kfree(kvm->arch.vdomain.pcpu_vcpu_list); + kfree(vdomain->pcpu_vcpu_list); - xa_for_each(&kvm->arch.vdomain.vcpu_data_array, index, data) { - kfree(data); - } - xa_destroy(&kvm->arch.vdomain.vcpu_data_array); - } + xa_for_each(&vdomain->vcpu_data_array, index, data) + kfree(data); + xa_destroy(&vdomain->vcpu_data_array); + + kfree(vdomain); } /** @@ -860,7 +880,7 @@ static bool is_vcpu_in_vdomain(u32 vdomain_bitmap, int vdomain_idx) static int update_domain_map(struct kvm_vcpu *vcpu, int vdomain_idx, int last_pcpu, bool clear_flag) { - struct tlbidomain *vdomain = &vcpu->kvm->arch.vdomain; + struct tlbidomain *vdomain = kvm_tlbidomain(vcpu->kvm); int pdomain = vdomain->domain_map[vdomain_idx]; cpumask_var_t new_cpus; @@ -940,11 +960,13 @@ static void kvm_arm_update_tlbid_map(struct tlbidomain *vdomain) static void kvm_vcpu_reload_tlbid(struct kvm *kvm) { - if (WARN_ON_ONCE(!kvm->arch.vdomain.kvm_tlbid_enabled)) + struct tlbidomain *vdomain = kvm_tlbidomain(kvm); + + if (WARN_ON_ONCE(!vdomain || !vdomain->kvm_tlbid_enabled)) return; preempt_disable(); - kvm_arm_update_tlbid_map(&kvm->arch.vdomain); + kvm_arm_update_tlbid_map(vdomain); preempt_enable(); } @@ -960,9 +982,9 @@ static inline void set_tlbididr_trap(void) static void kvm_tlbidomain_vcpu_load(struct kvm_vcpu *vcpu) { struct kvm *kvm = vcpu->kvm; - struct tlbidomain *vdomain = &kvm->arch.vdomain; - int max_vdomains = 1 << vdomain->nvis; - DECLARE_BITMAP(clear_bits, max_vdomains); + struct tlbidomain *vdomain = kvm_tlbidomain(kvm); + int max_vdomains; + DECLARE_BITMAP(clear_bits, 32); /* 1 << NVIS, NVIS <= 5 */ struct vcpu_tlbid_data *current_data, *other_data; bool domain_map_changed = false; int i, old_map; @@ -971,9 +993,11 @@ static void kvm_tlbidomain_vcpu_load(struct kvm_vcpu *vcpu) * If we support tlbid, but user does not enable it, we can do * vdomain0 -> pdomainN as well by default. */ - if (!vdomain->kvm_tlbid_enabled) + if (!vdomain || !vdomain->kvm_tlbid_enabled) return; + max_vdomains = 1 << vdomain->nvis; + set_tlbididr_trap(); current_data = xa_load(&vdomain->vcpu_data_array, vcpu->vcpu_idx); @@ -986,7 +1010,7 @@ static void kvm_tlbidomain_vcpu_load(struct kvm_vcpu *vcpu) if (likely(vcpu->cpu == current_data->last_pcpu)) return; - spin_lock(&vcpu->kvm->arch.vdomain.tlbid_lock); + spin_lock(&vdomain->tlbid_lock); /* * Remove current vCPU from its last_pcpu list and compute clear_bits @@ -1075,7 +1099,7 @@ static void kvm_tlbidomain_vcpu_load(struct kvm_vcpu *vcpu) list_add_tail(¤t_data->pcpu_node, &vdomain->pcpu_vcpu_list[vcpu->cpu]); - spin_unlock(&vcpu->kvm->arch.vdomain.tlbid_lock); + spin_unlock(&vdomain->tlbid_lock); } static inline void clear_tlbididr_trap(void) @@ -1090,9 +1114,10 @@ static inline void clear_tlbididr_trap(void) static void kvm_tlbidomain_vcpu_put(struct kvm_vcpu *vcpu) { struct kvm *kvm = vcpu->kvm; + struct tlbidomain *vdomain = kvm_tlbidomain(kvm); int i; - if (!kvm->arch.vdomain.kvm_tlbid_enabled) + if (!vdomain || !vdomain->kvm_tlbid_enabled) return; clear_tlbididr_trap(); @@ -2134,6 +2159,7 @@ static int kvm_vcpu_set_target(struct kvm_vcpu *vcpu, static int kvm_arm_tlbidomain_vcpu_init(struct kvm_vcpu *vcpu) { + struct tlbidomain *vdomain; struct vcpu_tlbid_data *data; int ret; @@ -2157,7 +2183,13 @@ static int kvm_arm_tlbidomain_vcpu_init(struct kvm_vcpu *vcpu) data->last_pcpu = -1; INIT_LIST_HEAD(&data->pcpu_node); - ret = xa_insert(&vcpu->kvm->arch.vdomain.vcpu_data_array, + vdomain = kvm_tlbidomain(vcpu->kvm); + if (!vdomain) { + kfree(data); + return 0; + } + + ret = xa_insert(&vdomain->vcpu_data_array, vcpu->vcpu_idx, data, GFP_KERNEL); if (ret) { kfree(data); @@ -2319,29 +2351,33 @@ static int kvm_arm_vcpu_rmm_psci_complete(struct kvm_vcpu *vcpu, static int kvm_arm_vcpu_set_vdomain(struct kvm_vcpu *vcpu, struct kvm_arm_set_vdomain *vdomain) { + struct tlbidomain *td = kvm_tlbidomain(vcpu->kvm); u32 vdomain_bitmap = vdomain->vdomain_bitmap; u8 num_vdomains = vdomain->num_vdomains; - u8 nvis = vcpu->kvm->arch.vdomain.nvis; - u8 max_vdomains = 1 << nvis; + u8 nvis, max_vdomains; struct vcpu_tlbid_data *data; - int ret; + + if (!td) + return -EINVAL; + + nvis = td->nvis; + max_vdomains = 1 << nvis; if (num_vdomains > max_vdomains) return -EINVAL; - if (vcpu->kvm->arch.vdomain.num_domains != -1 && - vcpu->kvm->arch.vdomain.num_domains != num_vdomains) + if (td->num_domains != -1 && td->num_domains != num_vdomains) return -EINVAL; - data = xa_load(&vcpu->kvm->arch.vdomain.vcpu_data_array, vcpu->vcpu_idx); + data = xa_load(&td->vcpu_data_array, vcpu->vcpu_idx); if (!data) { kvm_err("vCPU%d: Failed to load vcpu_tlbid_data\n", vcpu->vcpu_idx); return -EINVAL; } data->vdomain_bitmap = vdomain_bitmap; - vcpu->kvm->arch.vdomain.num_domains = num_vdomains; - vcpu->kvm->arch.vdomain.guest_tlbid_enabled = true; + td->num_domains = num_vdomains; + td->guest_tlbid_enabled = true; kvm_info("vCPU%d: vDomain bitmap: 0x%x\n", vcpu->vcpu_idx, vdomain_bitmap); @@ -2562,7 +2598,12 @@ static int kvm_vm_set_attr(struct kvm *kvm, struct kvm_device_attr *attr) static int kvm_arm_get_vdomain_num(struct kvm *kvm, struct kvm_arm_get_vdomain *vdomain) { - vdomain->max_vdomains = 1 << kvm->arch.vdomain.nvis; + struct tlbidomain *td = kvm_tlbidomain(kvm); + + if (!td) + return -EINVAL; + + vdomain->max_vdomains = 1 << td->nvis; return 0; } diff --git a/arch/arm64/kvm/sys_regs.c b/arch/arm64/kvm/sys_regs.c index 1ef578caa7c0..ec36a6a3262d 100644 --- a/arch/arm64/kvm/sys_regs.c +++ b/arch/arm64/kvm/sys_regs.c @@ -280,10 +280,12 @@ static bool access_tlbididr(struct kvm_vcpu *vcpu, struct sys_reg_params *p, const struct sys_reg_desc *r) { - if (p->is_write) + struct tlbidomain *vdomain = kvm_tlbidomain(vcpu->kvm); + + if (p->is_write || !vdomain) return ignore_write(vcpu, p); - p->regval = vcpu->kvm->arch.vdomain.tlbididr_val; + p->regval = vdomain->tlbididr_val; return true; } @@ -1456,10 +1458,13 @@ static u64 __kvm_read_sanitised_id_reg(const struct kvm_vcpu *vcpu, val &= ID_AA64MMFR3_EL1_TCRX | ID_AA64MMFR3_EL1_S1POE | ID_AA64MMFR3_EL1_S1PIE; break; - case SYS_ID_AA64MMFR4_EL1: - if (!vcpu->kvm->arch.vdomain.guest_tlbid_enabled) + case SYS_ID_AA64MMFR4_EL1: { + struct tlbidomain *vdomain = kvm_tlbidomain(vcpu->kvm); + + if (!vdomain || !vdomain->guest_tlbid_enabled) val &= ~ID_AA64MMFR4_EL1_TLBID; break; + } case SYS_ID_MMFR4_EL1: val &= ~ARM64_FEATURE_MASK(ID_MMFR4_EL1_CCIDX); break; @@ -3792,6 +3797,7 @@ const struct sys_reg_desc *get_reg_by_id(u64 id, */ void kvm_update_aa64mmfr4_tlbid(struct kvm *kvm) { + struct tlbidomain *vdomain = kvm_tlbidomain(kvm); u64 val, tlbid_val; mutex_lock(&kvm->arch.config_lock); @@ -3812,7 +3818,7 @@ void kvm_update_aa64mmfr4_tlbid(struct kvm *kvm) tlbid_val &= ID_AA64MMFR4_EL1_TLBID_MASK; val &= ~ID_AA64MMFR4_EL1_TLBID_MASK; - if (kvm->arch.vdomain.guest_tlbid_enabled) + if (vdomain && vdomain->guest_tlbid_enabled) val |= tlbid_val; kvm_set_vm_id_reg(kvm, SYS_ID_AA64MMFR4_EL1, val); -- 2.33.0
virt inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- Wrap all virtual TLBID (vTLBID) code in KVM behind a new CONFIG_KVM_ARM_VTLBID Kconfig option (default n, depends on KVM && ARM64_TLBID). When disabled, stub functions provide no-op behaviour so the rest of KVM compiles and runs without vTLBID support. Guards are placed at five sites in arm.c, each wrapping only vTLBID-specific functions so non-vTLBID code is unaffected: - Group A: kvm_tlbidomain(), kvm_arm_init_tlbidomain(), free_tlbid_vdomain() and the tlbid_domains xarray - Group B: is_vcpu_in_vdomain(), update_domain_map(), vtlbidn_clear_set_s(), kvm_arm_update_tlbid_map(), kvm_vcpu_reload_tlbid(), set_tlbididr_trap(), kvm_tlbidomain_vcpu_load(), clear_tlbididr_trap(), kvm_tlbidomain_vcpu_put() - Group C: kvm_arm_tlbidomain_vcpu_init() - Group D: kvm_arm_vcpu_set_vdomain() - Group E: kvm_arm_get_vdomain_num() Cross-file functions (kvm_tlbidomain, kvm_update_aa64mmfr4_tlbid) get inline stubs in their headers; access_tlbididr() in sys_regs.c gets a #else stub returning false. KVM_CAP_ARM_TLBIDOMAIN is additionally gated by IS_ENABLED(CONFIG_KVM_ARM_VTLBID). Signed-off-by: Jinqian Yang <yangjinqian1@huawei.com> --- arch/arm64/include/asm/kvm_host.h | 4 ++++ arch/arm64/kvm/Kconfig | 12 +++++++++++ arch/arm64/kvm/arm.c | 33 ++++++++++++++++++++++++++++++- arch/arm64/kvm/sys_regs.c | 11 +++++++++++ arch/arm64/kvm/sys_regs.h | 4 ++++ 5 files changed, 63 insertions(+), 1 deletion(-) diff --git a/arch/arm64/include/asm/kvm_host.h b/arch/arm64/include/asm/kvm_host.h index ed2083b7b206..6ad17d3e9b78 100644 --- a/arch/arm64/include/asm/kvm_host.h +++ b/arch/arm64/include/asm/kvm_host.h @@ -257,7 +257,11 @@ struct tlbidomain { * xarray keyed by the kvm pointer so that struct kvm_arch stays unchanged * (no KABI impact). Returns NULL when the VM has no TLBID support. */ +#ifdef CONFIG_KVM_ARM_VTLBID struct tlbidomain *kvm_tlbidomain(struct kvm *kvm); +#else +static inline struct tlbidomain *kvm_tlbidomain(struct kvm *kvm) { return NULL; } +#endif struct kvm_arch { struct kvm_s2_mmu mmu; diff --git a/arch/arm64/kvm/Kconfig b/arch/arm64/kvm/Kconfig index 689d32b140b7..2a71109ce082 100644 --- a/arch/arm64/kvm/Kconfig +++ b/arch/arm64/kvm/Kconfig @@ -145,4 +145,16 @@ config VIRT_TIMER_EARLY_INJECT The hypervisor can control the early injection latency via the module parameter: timer_early_inject_ns +config KVM_ARM_VTLBID + bool "Virtual TLBID support for KVM guests" + depends on KVM && ARM64_TLBID + default n + help + Enable virtualization of FEAT_TLBID so that each KVM guest can + use TLBI Domains independently of the host. The host TLBID + hardware is partitioned per-VM through VTLBID_EL2 registers + managed on the vCPU load/put path. + + If unsure, say N. + endif # VIRTUALIZATION diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c index 0025b0499211..d641b0dc1d6c 100644 --- a/arch/arm64/kvm/arm.c +++ b/arch/arm64/kvm/arm.c @@ -321,6 +321,7 @@ static int kvm_arm_default_max_vcpus(void) return vgic_present ? kvm_vgic_get_max_vcpus() : KVM_MAX_VCPUS; } +#ifdef CONFIG_KVM_ARM_VTLBID static DEFINE_XARRAY(tlbid_domains); struct tlbidomain *kvm_tlbidomain(struct kvm *kvm) @@ -435,6 +436,10 @@ static void free_tlbid_vdomain(struct kvm *kvm) kfree(vdomain); } +#else +static inline int kvm_arm_init_tlbidomain(struct kvm *kvm) { return 0; } +static inline void free_tlbid_vdomain(struct kvm *kvm) { } +#endif /** * kvm_arch_init_vm - initializes a VM data structure @@ -705,7 +710,7 @@ int kvm_vm_ioctl_check_extension(struct kvm *kvm, long ext) break; #endif case KVM_CAP_ARM_TLBIDOMAIN: - if (system_supports_tlbid()) + if (IS_ENABLED(CONFIG_KVM_ARM_VTLBID) && system_supports_tlbid()) r = 1; else r = 0; @@ -872,6 +877,7 @@ static void update_steal_time(struct kvm_vcpu *vcpu) } #endif +#ifdef CONFIG_KVM_ARM_VTLBID static bool is_vcpu_in_vdomain(u32 vdomain_bitmap, int vdomain_idx) { return vdomain_bitmap & BIT(vdomain_idx); @@ -1130,6 +1136,11 @@ static void kvm_tlbidomain_vcpu_put(struct kvm_vcpu *vcpu) for (i = 0; i < VTLBID_EL2_MAX_REGS; i++) vtlbidn_clear_set_s(i, VTLBID0_EL2_TD, 0); } +#else +static inline void kvm_vcpu_reload_tlbid(struct kvm *kvm) { } +static inline void kvm_tlbidomain_vcpu_load(struct kvm_vcpu *vcpu) { } +static inline void kvm_tlbidomain_vcpu_put(struct kvm_vcpu *vcpu) { } +#endif void kvm_arch_vcpu_load(struct kvm_vcpu *vcpu, int cpu) { @@ -2157,6 +2168,7 @@ static int kvm_vcpu_set_target(struct kvm_vcpu *vcpu, return kvm_reset_vcpu(vcpu); } +#ifdef CONFIG_KVM_ARM_VTLBID static int kvm_arm_tlbidomain_vcpu_init(struct kvm_vcpu *vcpu) { struct tlbidomain *vdomain; @@ -2198,6 +2210,9 @@ static int kvm_arm_tlbidomain_vcpu_init(struct kvm_vcpu *vcpu) return 0; } +#else +static inline int kvm_arm_tlbidomain_vcpu_init(struct kvm_vcpu *vcpu) { return 0; } +#endif static int kvm_arch_vcpu_ioctl_vcpu_init(struct kvm_vcpu *vcpu, struct kvm_vcpu_init *init) @@ -2348,6 +2363,7 @@ static int kvm_arm_vcpu_rmm_psci_complete(struct kvm_vcpu *vcpu, return realm_psci_complete(vcpu, target, arg->psci_status); } +#ifdef CONFIG_KVM_ARM_VTLBID static int kvm_arm_vcpu_set_vdomain(struct kvm_vcpu *vcpu, struct kvm_arm_set_vdomain *vdomain) { @@ -2385,6 +2401,13 @@ static int kvm_arm_vcpu_set_vdomain(struct kvm_vcpu *vcpu, return 0; } +#else +static inline int kvm_arm_vcpu_set_vdomain(struct kvm_vcpu *vcpu, + struct kvm_arm_set_vdomain *vdomain) +{ + return -EINVAL; +} +#endif long kvm_arch_vcpu_ioctl(struct file *filp, unsigned int ioctl, unsigned long arg) @@ -2595,6 +2618,7 @@ static int kvm_vm_set_attr(struct kvm *kvm, struct kvm_device_attr *attr) } } +#ifdef CONFIG_KVM_ARM_VTLBID static int kvm_arm_get_vdomain_num(struct kvm *kvm, struct kvm_arm_get_vdomain *vdomain) { @@ -2607,6 +2631,13 @@ static int kvm_arm_get_vdomain_num(struct kvm *kvm, return 0; } +#else +static inline int kvm_arm_get_vdomain_num(struct kvm *kvm, + struct kvm_arm_get_vdomain *vdomain) +{ + return -EINVAL; +} +#endif int kvm_arch_vm_ioctl(struct file *filp, unsigned int ioctl, unsigned long arg) { diff --git a/arch/arm64/kvm/sys_regs.c b/arch/arm64/kvm/sys_regs.c index ec36a6a3262d..5d2121ac8a88 100644 --- a/arch/arm64/kvm/sys_regs.c +++ b/arch/arm64/kvm/sys_regs.c @@ -276,6 +276,7 @@ static bool access_vm_reg(struct kvm_vcpu *vcpu, return true; } +#ifdef CONFIG_KVM_ARM_VTLBID static bool access_tlbididr(struct kvm_vcpu *vcpu, struct sys_reg_params *p, const struct sys_reg_desc *r) @@ -288,6 +289,14 @@ static bool access_tlbididr(struct kvm_vcpu *vcpu, p->regval = vdomain->tlbididr_val; return true; } +#else +static bool access_tlbididr(struct kvm_vcpu *vcpu, + struct sys_reg_params *p, + const struct sys_reg_desc *r) +{ + return false; +} +#endif static bool access_actlr(struct kvm_vcpu *vcpu, struct sys_reg_params *p, @@ -3795,6 +3804,7 @@ const struct sys_reg_desc *get_reg_by_id(u64 id, /* * Update the VM's stored value for an ID register. */ +#ifdef CONFIG_KVM_ARM_VTLBID void kvm_update_aa64mmfr4_tlbid(struct kvm *kvm) { struct tlbidomain *vdomain = kvm_tlbidomain(kvm); @@ -3825,6 +3835,7 @@ void kvm_update_aa64mmfr4_tlbid(struct kvm *kvm) mutex_unlock(&kvm->arch.config_lock); } +#endif /* Decode an index value, and find the sys_reg_desc entry. */ static const struct sys_reg_desc * diff --git a/arch/arm64/kvm/sys_regs.h b/arch/arm64/kvm/sys_regs.h index 4c5ae75bc9eb..8a2415921ba0 100644 --- a/arch/arm64/kvm/sys_regs.h +++ b/arch/arm64/kvm/sys_regs.h @@ -238,7 +238,11 @@ int kvm_sys_reg_get_user(struct kvm_vcpu *vcpu, const struct kvm_one_reg *reg, int kvm_sys_reg_set_user(struct kvm_vcpu *vcpu, const struct kvm_one_reg *reg, const struct sys_reg_desc table[], unsigned int num); +#ifdef CONFIG_KVM_ARM_VTLBID void kvm_update_aa64mmfr4_tlbid(struct kvm *kvm); +#else +static inline void kvm_update_aa64mmfr4_tlbid(struct kvm *kvm) { } +#endif #define AA32(_x) .aarch32_map = AA32_##_x #define Op0(_x) .Op0 = _x -- 2.33.0
virt inclusion category: feature bugzilla: https://gitcode.com/openeuler/kernel/issues/10066 ---------------------------------------- Enable virtual TLBID support for KVM guests in the openeuler defconfig. This allows guests to use TLBI Domains independently of the host when FEAT_TLBID hardware and a d128-capable toolchain are available. Signed-off-by: Jinqian Yang <yangjinqian1@huawei.com> --- arch/arm64/configs/openeuler_defconfig | 1 + 1 file changed, 1 insertion(+) diff --git a/arch/arm64/configs/openeuler_defconfig b/arch/arm64/configs/openeuler_defconfig index edf2ed90efb8..2baee29c9c7b 100644 --- a/arch/arm64/configs/openeuler_defconfig +++ b/arch/arm64/configs/openeuler_defconfig @@ -821,6 +821,7 @@ CONFIG_VIRTUALIZATION=y CONFIG_KVM=y # CONFIG_NVHE_EL2_DEBUG is not set CONFIG_KVM_ARM_MULTI_LPI_TRANSLATE_CACHE=y +CONFIG_KVM_ARM_VTLBID=y # CONFIG_ENABLE_KVM_FPMR is not set CONFIG_ARCH_VCPU_STAT=y CONFIG_VIRT_VTIMER_IRQ_BYPASS=y -- 2.33.0
反馈: 您发送到kernel@openeuler.org的补丁/补丁集,已成功转换为PR! PR链接地址: https://gitcode.com/openeuler/kernel/merge_requests/29120 邮件列表地址:https://mailweb.openeuler.org/archives/list/kernel@openeuler.org/message/6R2... FeedBack: The patch(es) which you have sent to kernel@openeuler.org mailing list has been converted to a pull request successfully! Pull request link: https://gitcode.com/openeuler/kernel/merge_requests/29120 Mailing list address: https://mailweb.openeuler.org/archives/list/kernel@openeuler.org/message/6R2...
participants (2)
-
Jinqian Yang -
patchwork bot