mailweb.openeuler.org
Manage this list

Keyboard Shortcuts

Thread View

  • j: Next unread message
  • k: Previous unread message
  • j a: Jump to all threads
  • j l: Jump to MailingList overview

Kernel

Threads by month
  • ----- 2026 -----
  • October
  • September
  • August
  • July
  • June
  • May
  • April
  • March
  • February
  • January
  • ----- 2025 -----
  • December
  • November
  • October
  • September
  • August
  • July
  • June
  • May
  • April
  • March
  • February
  • January
  • ----- 2024 -----
  • December
  • November
  • October
  • September
  • August
  • July
  • June
  • May
  • April
  • March
  • February
  • January
  • ----- 2023 -----
  • December
  • November
  • October
  • September
  • August
  • July
  • June
  • May
  • April
  • March
  • February
  • January
  • ----- 2022 -----
  • December
  • November
  • October
  • September
  • August
  • July
  • June
  • May
  • April
  • March
  • February
  • January
  • ----- 2021 -----
  • December
  • November
  • October
  • September
  • August
  • July
  • June
  • May
  • April
  • March
  • February
  • January
  • ----- 2020 -----
  • December
  • November
  • October
  • September
  • August
  • July
  • June
  • May
  • April
  • March
  • February
  • January
  • ----- 2019 -----
  • December
kernel@openeuler.org

  • 25065 discussions
[PATCH OLK-6.6] bpf: Reject bpf_obj_drop() from tracing progs
by Chen Yuxi 30 Sep '26

30 Sep '26
From: Justin Suess <utilityemal77(a)gmail.com> mainline inclusion from mainline-v7.2-rc1 commit 94c8d1c21be40a845357854f98ec07e21bb14bc9 category: bugfix bugzilla: https://atomgit.com/src-openeuler/kernel/issues/19869 Reference: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?… -------------------------------- bpf_obj_drop() runs bpf_obj_free_fields() synchronously for program-allocated objects. When such an object contains NMI unsafe fields, tracing programs that can run from arbitrary instrumented context can reach that destruction from unsafe contexts, including NMI. NMI is likely one instance of this problem, and other instances would include possible unsafe reentrancy. Deferring bpf_obj_drop() is not appealing either: it would add delayed-free machinery to a release operation that otherwise has straightforward synchronous ownership semantics. Reject bpf_obj_drop() and bpf_percpu_obj_drop() from tracing programs that may run from unsafe contexts unless every field in the object's BTF record is explicitly NMI safe. Do not reject sleepable BPF_PROG_TYPE_TRACING programs, since they are not the arbitrary/NMI contexts that motivate the restriction. Note that while bpf_rb_root and bpf_list_head would be NMI safe on their own to free, the objects recursively held by them may not be; be conservative and just mark them as not NMI safe for now. Use a whitelist for the NMI-safe field set instead of listing only known NMI unsafe fields. Locks, async fields, unreferenced kptrs, and refcounts are known to be NMI safe because their destruction is either a no-op, simple state reset, or async cancellation. Referenced kptrs, percpu referenced kptrs, uptrs, graph roots, graph nodes, and any future field type are rejected until audited for arbitrary tracing and NMI contexts. This is less susceptible to future changes in fields that were previously safe by exclusion, and to new fields being added without updating this check. Convert the existing recursive local-object drop success case to a syscall program in the same commit, since this verifier change makes the old tracing program form invalid. The test still exercises bpf_obj_drop() releasing a referenced task kptr from a safe program type. Fixes: ac9f06050a35 ("bpf: Introduce bpf_obj_drop") Signed-off-by: Justin Suess <utilityemal77(a)gmail.com> Co-developed-by: Kumar Kartikeya Dwivedi <memxor(a)gmail.com> Signed-off-by: Kumar Kartikeya Dwivedi <memxor(a)gmail.com> Link: https://lore.kernel.org/r/20260609202548.3571690-2-memxor@gmail.com Signed-off-by: Alexei Starovoitov <ast(a)kernel.org> Conflicts: include/linux/bpf.h, kernel/bpf/verifier.c, tools/testing/selftests/bpf/progs/task_kfunc_success.c, tools/testing/selftests/bpf/prog_tests/task_kfunc.c [Whitelist only the field types present in this tree, add is_bpf_obj_drop_kfunc() matching KF_bpf_obj_drop_impl only and drop the percpu branch (no percpu obj_drop kfuncs here), and convert the selftest with its function body kept. Meanwhile, skip the selftest conversion for it does not call bpf_obj_drop.] Signed-off-by: Chen Yuxi <chenyuxi19(a)huawei.com> --- include/linux/bpf.h | 26 ++++++++++++++++++++++++++ kernel/bpf/verifier.c | 21 +++++++++++++++++++++ 2 files changed, 47 insertions(+) diff --git a/include/linux/bpf.h b/include/linux/bpf.h index 4e9c38b14e04..de601dff733e 100644 --- a/include/linux/bpf.h +++ b/include/linux/bpf.h @@ -452,6 +452,32 @@ static inline bool btf_record_has_field(const struct btf_record *rec, enum btf_f return rec->field_mask & type; } +static inline bool btf_field_is_nmi_safe(enum btf_field_type type) +{ + switch (type) { + case BPF_SPIN_LOCK: + case BPF_TIMER: + case BPF_KPTR_UNREF: + case BPF_REFCOUNT: + return true; + default: + return false; + } +} + +static inline bool btf_record_has_nmi_unsafe_fields(const struct btf_record *rec) +{ + int i; + + if (IS_ERR_OR_NULL(rec)) + return false; + for (i = 0; i < rec->cnt; i++) { + if (!btf_field_is_nmi_safe(rec->fields[i].type)) + return true; + } + return false; +} + static inline void bpf_obj_init(const struct btf_record *rec, void *obj) { int i; diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c index f0a67a9f8d1a..f9531bbc8767 100644 --- a/kernel/bpf/verifier.c +++ b/kernel/bpf/verifier.c @@ -200,6 +200,7 @@ static int release_reference(struct bpf_verifier_env *env, int ref_obj_id); static void invalidate_non_owning_refs(struct bpf_verifier_env *env); static void invalidate_rcu_protected_refs(struct bpf_verifier_env *env); static bool in_rbtree_lock_required_cb(struct bpf_verifier_env *env); +static bool is_tracing_prog_type(enum bpf_prog_type type); static int ref_set_non_owning(struct bpf_verifier_env *env, struct bpf_reg_state *reg); static void specialize_kfunc(struct bpf_verifier_env *env, @@ -11372,6 +11373,11 @@ static int check_reg_allocation_locked(struct bpf_verifier_env *env, struct bpf_ return 0; } +static bool is_bpf_obj_drop_kfunc(u32 func_id) +{ + return func_id == special_kfunc_list[KF_bpf_obj_drop_impl]; +} + static bool is_bpf_list_api_kfunc(u32 btf_id) { return btf_id == special_kfunc_list[KF_bpf_list_push_front_impl] || @@ -12120,6 +12126,7 @@ static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn, struct bpf_reg_state *regs = cur_regs(env); const char *func_name, *ptr_type_name; bool sleepable, rcu_lock, rcu_unlock; + enum bpf_prog_type prog_type = resolve_prog_type(env->prog); struct bpf_kfunc_call_arg_meta meta; struct bpf_insn_aux_data *insn_aux; int err, insn_idx = *insn_idx_p; @@ -12162,6 +12169,20 @@ static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn, if (err < 0) return err; + if (is_bpf_obj_drop_kfunc(meta.func_id) && (is_tracing_prog_type(prog_type) || + /* is_tracing_prog_type() for now doesn't cover non-iterator tracing progs. */ + (prog_type == BPF_PROG_TYPE_TRACING && env->prog->expected_attach_type != BPF_TRACE_ITER + && !env->prog->sleepable))) { + struct btf_struct_meta *struct_meta; + + struct_meta = btf_find_struct_meta(meta.arg_btf, meta.arg_btf_id); + if (struct_meta && btf_record_has_nmi_unsafe_fields(struct_meta->record)) { + verbose(env, "%s cannot be used in tracing programs on types with NMI unsafe fields\n", + func_name); + return -EINVAL; + } + } + if (meta.func_id == special_kfunc_list[KF_bpf_rbtree_add_impl]) { err = push_callback_call(env, insn, insn_idx, meta.subprogno, set_rbtree_add_callback_state); -- 2.34.1
2 1
0 0
[PATCH OLK-5.10] mm/rmap: fix missing barrier between anon_vma init and vma->anon_vma publish
by Jinjiang Tu 30 Sep '26

30 Sep '26
mainline inclusion from mainline-v7.3-rc5 commit b6ac0b3f6013c168f22cad97e79967accacb08e1 category: bugfix bugzilla: https://atomgit.com/openeuler/kernel/issues/10048 Reference: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?… -------------------------------- On arm64 server, we find that a task trying to grab the anon_vma lock triggers hungtask. INFO: task main:2354726 blocked for more than 120 seconds. Tainted: G E 5.10.0-0021.aarch64 #1 "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message. task:main state:D stack: 0 pid:2354726 ppid:2350673 flags:0x00000a01 Call trace: __switch_to+0x7c/0xbc __schedule+0x3b4/0x8a0 schedule+0x50/0xe0 rwsem_down_write_slowpath+0x3cc/0x6cc down_write+0x60/0x260 __anon_vma_prepare+0x6c/0x210 do_anonymous_page+0x258/0x660 handle_pte_fault+0x188/0x214 __handle_mm_fault+0x1b0/0x380 handle_mm_fault+0xf4/0x284 do_page_fault+0x19c/0x494 do_translation_fault+0xcc/0xf8 do_mem_abort+0x48/0xac el0_da+0x44/0x80 el0_sync_handler+0x88/0xb4 el0_sync+0x160/0x180 After analyzing the vmcore, we found the anon_vma->root->rwsem.count is -1. There is another anon_vma whose anon_vma->root->rwsem.count is 1, the anon_vma->root->rwsem.owner shows the lock is held, but the stack of the task shows the task doesn't hold the anon_vma lock. After adding more debugging info, we found __anon_vma_prepare() reuses anon_vma and triggers the UAF of anon_vma->root due to missing memory barrier, leading to locking and unlocking two different anon_vma->root, thus leading to an anon_vma will never be unlocked, and another anon_vma couldn't be locked anymore. This race requires two adjacent VMAs that are not merged but are anon_vma-compatible (e.g., they differ in VMA_ACCESS_FLAGS that can be changed by mprotect()). Two threads fault on each VMA concurrently, both calling __anon_vma_prepare() with only mmap_lock held for reading. THREAD A THREAD B __anon_vma_prepare __anon_vma_prepare find_mergeable_anon_vma() -> NULL anon_vma = anon_vma_alloc(); anon_vma->root = anon_vma; // the two stores may be reordered vma->anon_vma = anon_vma; // finds A's anon_vma anon_vma = find_mergeable_anon_vma(vma); anon_vma_lock_write(anon_vma); // may still see the old root down_write(&anon_vma->root->rwsem); anon_vma_unlock_write(anon_vma); // see the new root, never unlock old up_write(&anon_vma->root->rwsem); thread A triggers page fault and calls __anon_vma_prepare() to prepare anon_vma for the faulting vma. __anon_vma_prepare() allocates and initializes a new anon_vma, and then publishes it to the vma with a plain store. anon_vma_prepare() only requires the mmap_lock to be held for reading, so two threads can fault on adjacent VMAs at the same time. While thread A publishes a new anon_vma, thread B could find the anon_vma via find_mergeable_anon_vma() and then locks anon_vma->root->rwsem. The store to anon_vma->root in anon_vma_alloc() and the store to vma->anon_vma can be reordered. The anon_vma_lock_write() and spin_lock() only provide acquire semantics, which do not prevent prior stores from being reordered after them. The release semantics of the corresponding spin_unlock() and anon_vma_unlock_write() come too late, the store to vma->anon_vma is already published before they take effect. As a result, thread B can observe the following order: vma->anon_vma = anon_vma; anon_vma->root = anon_vma; The anon_vma slab is SLAB_TYPESAFE_BY_RCU, so a newly allocated anon_vma may reuse memory from a previously freed one. The constructor (anon_vma_ctor) does not reset anon_vma->root, and __put_anon_vma() doesn't clear it either, so the old root value persists until anon_vma_alloc() overwrites it. If that store isn't visible, thread B reads a root that points to the old anon_vma and locks it. As a result, thread B can call anon_vma_lock_write() with the old root, and call anon_vma_unlock_write() with the new root, leading to an anon_vma will never be unlocked, and another anon_vma couldn't be locked anymore (its count is dropped from 0 to -1 due to wrong unlock). To fix it, change the plain store `vma->anon_vma = anon_vma` to store release, so that the fields of anon_vma are visible before anon_vma is published to vma->anon_vma. At read side, the load of anon_vma and anon_vma->root have address dependency. According to Documentation/memory-barriers.txt and some investigations, only Alpha needs address-dependency barriers and it has been handled by READ_ONCE() in reusable_anon_vma(). We reproduced this issue in v5.10 with KSM enabled. The kernel doesn't merge commit cf7e7a3503df ("mm: prevent KSM from breaking VMA merging for new VMAs"), so there are many adjacent VMAs that aren't merged but are compatible for anon_vma. Without this fix, our production environment could reproduce this issue about 2-5 times each month. After adding a smp_mb() before anon_vma_lock_write(anon_vma) in __anon_vma_prepare(), which is different to this patch, this issue hasn't been reproduced for one month. Link: https://lore.kernel.org/20260908122924.554373-1-tujinjiang@huawei.com Fixes: 5c341ee1dfc8 ("mm: track the root (oldest) anon_vma") Signed-off-by: Jinjiang Tu <tujinjiang(a)huawei.com> Signed-off-by: Andrew Morton <akpm(a)linux-foundation.org> Reviewed-by: Lance Yang <lance.yang(a)linux.dev> Reviewed-by: Lorenzo Stoakes (ARM) <ljs(a)kernel.org> Acked-by: David Hildenbrand (Arm) <david(a)kernel.org> Acked-by: Vlastimil Babka (SUSE) <vbabka(a)kernel.org> Cc: Minchan Kim <minchan(a)kernel.org> Cc: Harry Yoo <harry(a)kernel.org> Cc: Hiroyouki Kamezawa <kamezawa.hiroyu(a)jp.fujitsu.com> Cc: Jann Horn <jannh(a)google.com> Cc: Jinjiang Tu <tujinjiang(a)huawei.com> Cc: Kefeng Wang <wangkefeng.wang(a)huawei.com> Cc: Larry Woodman <lwoodman(a)redhat.com> Cc: Liam R. Howlett <liam(a)infradead.org> Cc: Nanyong Sun <sunnanyong(a)huawei.com> Cc: Rik van Riel <riel(a)surriel.com> Cc: <stable(a)vger.kernel.org> Conflicts: mm/rmap.c mm/mmap.c mm/vma.c [Context conflicts.] Signed-off-by: Jinjiang Tu <tujinjiang(a)huawei.com> --- mm/mmap.c | 7 +++++++ mm/rmap.c | 6 +++++- 2 files changed, 12 insertions(+), 1 deletion(-) diff --git a/mm/mmap.c b/mm/mmap.c index c0262befcaae..604c0c877c5c 100644 --- a/mm/mmap.c +++ b/mm/mmap.c @@ -1300,6 +1300,13 @@ static int anon_vma_compatible(struct vm_area_struct *a, struct vm_area_struct * * acceptable for merging, so we can do all of this optimistically. But * we do that READ_ONCE() to make sure that we never re-load the pointer. * + * The READ_ONCE() establishes an address dependency between anon_vma and + * any access to its fields, which pairs with the assignment to + * vma->anon_vma performed with release semantics in __anon_vma_prepare(). + * + * This is especially important as anon_vma's are SLAB_TYPESAFE_BY_RCU so + * accessing an uninitialised anon_vma's fields may result in a UAF. + * * IOW: that the "list_is_singular()" test on the anon_vma_chain only * matters for the 'stable anon_vma' case (ie the thing we want to avoid * is to return an anon_vma that is "complex" due to having gone through diff --git a/mm/rmap.c b/mm/rmap.c index 150803a7ffb5..fb07b2b32a91 100644 --- a/mm/rmap.c +++ b/mm/rmap.c @@ -206,7 +206,11 @@ int __anon_vma_prepare(struct vm_area_struct *vma) /* page_table_lock to protect against threads */ spin_lock(&mm->page_table_lock); if (likely(!vma->anon_vma)) { - vma->anon_vma = anon_vma; + /* + * Make anon_vma fields visible before anon_vma is published. + * Paired with an address dependency in reusable_anon_vma(). + */ + smp_store_release(&vma->anon_vma, anon_vma); anon_vma_chain_link(vma, avc, anon_vma); anon_vma->num_active_vmas++; allocated = NULL; -- 2.43.0
2 1
0 0
[PATCH openEuler-1.0-LTS] NFSD: Prevent lock owner use-after-free during client teardown
by Gaosheng Cui 30 Sep '26

30 Sep '26
From: Chuck Lever <cel(a)kernel.org> stable inclusion from stable-6.6.157 commit 1ce74d1b7770e69735e8f8e509807af4ff9c8ee7 category: bugfix bugzilla: https://atomgit.com/src-openeuler/kernel/issues/18928 CVE: CVE-2026-89659 Reference: https://git.kernel.org/pub/scm/linux/kernel/git/stable/linux.git/commit/?id… ------------------------------ commit 5e2fa29d223a9a1e6a948e40b109d09081d1decd upstream. __destroy_client() releases a client's open owners, but a lock owner whose only reference is a blocked lock (nbl) stays on cl_ownerstr_hashtbl. client_has_state() does not count a bare owner, so DESTROY_CLIENTID can reach __destroy_client() with such owners present. __destroy_client() then walks the table, calling remove_blocked_locks() on each owner without a reference. Freeing a blocked lock drops the owner reference held via flc_owner. The per-net laundromat reaps blocked locks from nn->blocked_locks_lru independently of client state. The two paths share blocked_locks_lock only for the list splice, not the owner's lifetime. The laundromat therefore frees the owner as __destroy_client() dereferences it, a NULL dereference in remove_blocked_locks(). nfsd4_release_lockowner() holds a reference across the same call; __destroy_client() does not. Hold cl_lock across the walk, taking a reference and unhashing each owner, then drop it before remove_blocked_locks() and nfs4_put_stateowner(), which take blocked_locks_lock and cl_lock. Reported-by: Wolfgang Walter <linux(a)stwm.de> Closes: https://lore.kernel.org/linux-nfs/6eccafaaaa60651ef091257c3439c46b@stwm.de/ Fixes: 68ef3bc31664 ("nfsd: remove blocked locks on client teardown") Cc: stable(a)vger.kernel.org Reviewed-by: NeilBrown <neil(a)brown.name> Reviewed-by: Jeff Layton <jlayton(a)kernel.org> Link: https://patch.msgid.link/20260709-cel-v4-1-1d519d9be0cb@kernel.org Signed-off-by: Chuck Lever <cel(a)kernel.org> Signed-off-by: Greg Kroah-Hartman <gregkh(a)linuxfoundation.org> Conflicts: fs/nfsd/nfs4state.c [Context differences, commit e0639dc5805a ("NFSD introduce async copy feature") has not being merged] Signed-off-by: Gaosheng Cui <cuigaosheng1(a)huawei.com> --- fs/nfsd/nfs4state.c | 16 +++++++++++++--- 1 file changed, 13 insertions(+), 3 deletions(-) diff --git a/fs/nfsd/nfs4state.c b/fs/nfsd/nfs4state.c index ba6ae4626a6c..d07e620fe72d 100644 --- a/fs/nfsd/nfs4state.c +++ b/fs/nfsd/nfs4state.c @@ -1972,14 +1972,24 @@ __destroy_client(struct nfs4_client *clp) release_openowner(oo); } for (i = 0; i < OWNER_HASH_SIZE; i++) { - struct nfs4_stateowner *so, *tmp; + struct nfs4_stateowner *so; - list_for_each_entry_safe(so, tmp, &clp->cl_ownerstr_hashtbl[i], - so_strhash) { + spin_lock(&clp->cl_lock); + while (!list_empty(&clp->cl_ownerstr_hashtbl[i])) { + so = list_first_entry(&clp->cl_ownerstr_hashtbl[i], + struct nfs4_stateowner, so_strhash); /* Should be no openowners at this point */ WARN_ON_ONCE(so->so_is_open_owner); + nfs4_get_stateowner(so); + unhash_lockowner_locked(lockowner(so)); + spin_unlock(&clp->cl_lock); + remove_blocked_locks(lockowner(so)); + nfs4_put_stateowner(so); + + spin_lock(&clp->cl_lock); } + spin_unlock(&clp->cl_lock); } nfsd4_return_all_client_layouts(clp); nfsd4_shutdown_callback(clp); -- 2.43.0
2 1
0 0
[PATCH openEuler-1.0-LTS] NFSD: Prevent client use-after-free during delegation revoke
by Gaosheng Cui 30 Sep '26

30 Sep '26
From: Chuck Lever <cel(a)kernel.org> mainline inclusion from mainline-v7.3-rc1 commit 4683ca76b3b7e5808338491c6eb3c20e6b4894d5 category: bugfix bugzilla: https://atomgit.com/src-openeuler/kernel/issues/18928 CVE: CVE-2026-89659 Reference: https://git.kernel.org/pub/scm/linux/kernel/git/stable/linux.git/commit/?id… -------------------------------- A delegation stateid holds only a bare pointer to its owning nfs4_client and does not keep it alive. The client survives its stateids only because __destroy_client() drains cl_delegations and cl_revoked before free_client() runs. nfs4_laundromat() breaks that invariant: it unhashes an expired delegation from cl_delegations, drops deleg_lock, then revoke_delegation() relinks it onto cl_revoked under cl_lock. In that window the delegation is on neither list, so client_has_state() can report no remaining state. Every teardown path first requires cl_rpc_users to be zero, but the laundromat holds no such reference. A client whose recalled delegation has just timed out can therefore reach free_client() while revoke_delegation() is still about to dereference cl_lock, a use-after-free. Pin the client with cl_rpc_users across the revoke so teardown blocks until it completes, then reap the delegation from cl_revoked. A client already expiring reaps its own, so skip it and leave the delegation on del_recall_lru. Fixes: 3bd64a5ba171 ("nfsd4: implement SEQ4_STATUS_RECALLABLE_STATE_REVOKED") Cc: stable(a)vger.kernel.org Reviewed-by: NeilBrown <neil(a)brown.name> Reviewed-by: Jeff Layton <jlayton(a)kernel.org> Link: https://patch.msgid.link/20260709-cel-v4-2-1d519d9be0cb@kernel.org Signed-off-by: Chuck Lever <cel(a)kernel.org> Conflicts: fs/nfsd/nfs4state.c [Commit 4683ca76b3b7 pins the client with cl_rpc_users (guarded by a dedicated deleg_lock, waking expiry_wq on release). This tree has neither cl_rpc_users nor deleg_lock nor expiry_wq; it guards client expiry with cl_refcount (checked by mark_client_expired_locked) and protects del_recall_lru with state_lock. So the backport pins the client with atomic_inc(&clp->cl_refcount) under client_lock (taken before state_lock, matching the existing client_lock->state_lock ordering in nfsd_find_all_delegations) and unpins with a plain atomic_dec after revoke_delegation(), skipping the wake_up_all(&expiry_wq) that has no target here. The netns.h comment-only change advertising deleg_lock was dropped since that lock does not exist in this tree.] Co-authored-by: BackportAgent@deepseek-v4-flash Signed-off-by: Gaosheng Cui <cuigaosheng1(a)huawei.com> --- fs/nfsd/nfs4state.c | 13 +++++++++++++ 1 file changed, 13 insertions(+) diff --git a/fs/nfsd/nfs4state.c b/fs/nfsd/nfs4state.c index d07e620fe72d..972a9127fc26 100644 --- a/fs/nfsd/nfs4state.c +++ b/fs/nfsd/nfs4state.c @@ -4825,6 +4825,7 @@ nfs4_laundromat(struct nfsd_net *nn) list_del_init(&clp->cl_lru); expire_client(clp); } + spin_lock(&nn->client_lock); spin_lock(&state_lock); list_for_each_safe(pos, next, &nn->del_recall_lru) { dp = list_entry (pos, struct nfs4_delegation, dl_recall_lru); @@ -4833,15 +4834,27 @@ nfs4_laundromat(struct nfsd_net *nn) new_timeo = min(new_timeo, t); break; } + clp = dp->dl_stid.sc_client; + if (is_client_expired(clp)) + continue; + /* + * Pin the client so it cannot be torn down and freed + * while revoke_delegation() reaps this delegation below. + */ + atomic_inc(&clp->cl_refcount); WARN_ON(!unhash_delegation_locked(dp)); list_add(&dp->dl_recall_lru, &reaplist); } spin_unlock(&state_lock); + spin_unlock(&nn->client_lock); while (!list_empty(&reaplist)) { dp = list_first_entry(&reaplist, struct nfs4_delegation, dl_recall_lru); + clp = dp->dl_stid.sc_client; list_del_init(&dp->dl_recall_lru); revoke_delegation(dp); + /* Unpin without renewing the reaped client's lease. */ + atomic_dec(&clp->cl_refcount); } spin_lock(&nn->client_lock); -- 2.43.0
2 1
0 0
[PATCH openEuler-1.0-LTS] nfsd: gate nfs3 setacl by argp->mask
by Gaosheng Cui 30 Sep '26

30 Sep '26
From: Chris Mason <clm(a)meta.com> stable inclusion from stable-6.6.157 commit 3be1d8611dae4829e4608007db29f4a8748c37b2 category: bugfix bugzilla: https://atomgit.com/src-openeuler/kernel/issues/18934 CVE: CVE-2026-89671 Reference: https://git.kernel.org/pub/scm/linux/kernel/git/stable/linux.git/commit/?id… ------------------------------ commit 453d7198a0ab07a12d46e0575861ac7b932da17e upstream. nfsd3_proc_setacl() calls set_posix_acl() unconditionally for both ACL_TYPE_ACCESS and ACL_TYPE_DEFAULT, passing argp->acl_access and argp->acl_default verbatim. The NFSv3 ACL decoder only populates those pointers when the corresponding mask bit is set: nfs3svc_decode_setaclargs() if (args->mask & NFS_ACL) decode into acl_access if (args->mask & NFS_DFACL) decode into acl_default /* otherwise the pointer stays NULL (pc_argzero) */ nfsd3_proc_setacl() set_posix_acl(.., ACL_TYPE_ACCESS, argp->acl_access) set_posix_acl(.., ACL_TYPE_DEFAULT, argp->acl_default) set_posix_acl(idmap, dentry, type, NULL) is the VFS "remove this ACL type" operation. A NULL pointer that means "the client did not send this arm" is therefore indistinguishable from "the client asked to remove this ACL". A SETACL with mask=NFS_ACL silently drops the directory's default ACL; mask=0 drops both. The sibling nfsd3_proc_getacl() already consults argp->mask before touching each arm; mirror that in setacl. Fix by wrapping each set_posix_acl() call in the matching mask bit check and initializing error to 0 before inode_lock so that a request with neither bit set leaves the on-disk ACLs untouched and returns nfs_ok. The out_drop_lock path and the unconditional posix_acl_release() at out: are preserved; both NULL-tolerate the skipped arms. Fixes: a257cdd0e217 ("[PATCH] NFSD: Add server support for NFSv3 ACLs.") Cc: stable(a)vger.kernel.org Assisted-by: kres:claude-opus-4-7 Reported-by: Chris Mason <clm(a)meta.com> Signed-off-by: Chris Mason <clm(a)meta.com> Link: https://patch.msgid.link/20260530-nfsd-fixes-v2-5-f27e8eb4d974@kernel.org Signed-off-by: Chuck Lever <chuck.lever(a)oracle.com> Signed-off-by: Greg Kroah-Hartman <gregkh(a)linuxfoundation.org> Conflicts: fs/nfsd/nfs3acl.c [Context differences, commit bb4d53d66e4b ("NFSD: use (un)lock_inode instead of fh_(un)lock for file operations") has not been merged, so fh_lock/fh_unlock are kept. And set_posix_acl() in this tree takes (inode, type, acl) without the mnt_idmap/dentry arguments, so the calls are rewritten to the local signature instead of the upstream (&nop_mnt_idmap, fh->fh_dentry, ...) form.] Signed-off-by: Gaosheng Cui <cuigaosheng1(a)huawei.com> --- fs/nfsd/nfs3acl.c | 15 +++++++++++---- 1 file changed, 11 insertions(+), 4 deletions(-) diff --git a/fs/nfsd/nfs3acl.c b/fs/nfsd/nfs3acl.c index 04cc86a520ca..027a0fb0f142 100644 --- a/fs/nfsd/nfs3acl.c +++ b/fs/nfsd/nfs3acl.c @@ -106,10 +106,17 @@ static __be32 nfsd3_proc_setacl(struct svc_rqst *rqstp) fh_lock(fh); - error = set_posix_acl(inode, ACL_TYPE_ACCESS, argp->acl_access); - if (error) - goto out_drop_lock; - error = set_posix_acl(inode, ACL_TYPE_DEFAULT, argp->acl_default); + error = 0; + if (argp->mask & NFS_ACL) { + error = set_posix_acl(inode, ACL_TYPE_ACCESS, + argp->acl_access); + if (error) + goto out_drop_lock; + } + if (argp->mask & NFS_DFACL) { + error = set_posix_acl(inode, ACL_TYPE_DEFAULT, + argp->acl_default); + } out_drop_lock: fh_unlock(fh); -- 2.43.0
2 1
0 0
[PATCH OLK-6.6 0/3] backport devfreq uncore DVFS governor patches
by Jie Zhan 30 Sep '26

30 Sep '26
driver inclusion category: feature bugzilla: https://atomgit.com/openeuler/kernel/issues/9163 ---------------------------------------------------------------------- backport devfreq uncore DVFS governor patches Jie Zhan (2): perf/core: Add perf_pmu_type_by_parent_dev() for in-kernel PMU lookup devfreq: hisi: Add devfreq-event support via L3C PMU Jonathan Cameron (1): perf/hisi-uncore: Assign parents for event_source devices drivers/devfreq/Kconfig | 3 +- drivers/devfreq/hisi_uncore_freq.c | 444 ++++++++++++++++++++++- drivers/perf/hisilicon/hisi_uncore_pmu.c | 1 + include/linux/perf_event.h | 2 + kernel/events/core.c | 28 ++ 5 files changed, 474 insertions(+), 4 deletions(-) -- 2.43.0
2 4
0 0
[PATCH OLK-6.6] bpf: Reject bpf_obj_drop() from tracing progs
by Chen Yuxi 30 Sep '26

30 Sep '26
From: Justin Suess <utilityemal77(a)gmail.com> mainline inclusion from mainline-v7.2-rc1 commit 94c8d1c21be40a845357854f98ec07e21bb14bc9 category: bugfix bugzilla: https://atomgit.com/src-openeuler/kernel/issues/19869 Reference: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?… -------------------------------- bpf_obj_drop() runs bpf_obj_free_fields() synchronously for program-allocated objects. When such an object contains NMI unsafe fields, tracing programs that can run from arbitrary instrumented context can reach that destruction from unsafe contexts, including NMI. NMI is likely one instance of this problem, and other instances would include possible unsafe reentrancy. Deferring bpf_obj_drop() is not appealing either: it would add delayed-free machinery to a release operation that otherwise has straightforward synchronous ownership semantics. Reject bpf_obj_drop() and bpf_percpu_obj_drop() from tracing programs that may run from unsafe contexts unless every field in the object's BTF record is explicitly NMI safe. Do not reject sleepable BPF_PROG_TYPE_TRACING programs, since they are not the arbitrary/NMI contexts that motivate the restriction. Note that while bpf_rb_root and bpf_list_head would be NMI safe on their own to free, the objects recursively held by them may not be; be conservative and just mark them as not NMI safe for now. Use a whitelist for the NMI-safe field set instead of listing only known NMI unsafe fields. Locks, async fields, unreferenced kptrs, and refcounts are known to be NMI safe because their destruction is either a no-op, simple state reset, or async cancellation. Referenced kptrs, percpu referenced kptrs, uptrs, graph roots, graph nodes, and any future field type are rejected until audited for arbitrary tracing and NMI contexts. This is less susceptible to future changes in fields that were previously safe by exclusion, and to new fields being added without updating this check. Convert the existing recursive local-object drop success case to a syscall program in the same commit, since this verifier change makes the old tracing program form invalid. The test still exercises bpf_obj_drop() releasing a referenced task kptr from a safe program type. Fixes: ac9f06050a35 ("bpf: Introduce bpf_obj_drop") Signed-off-by: Justin Suess <utilityemal77(a)gmail.com> Co-developed-by: Kumar Kartikeya Dwivedi <memxor(a)gmail.com> Signed-off-by: Kumar Kartikeya Dwivedi <memxor(a)gmail.com> Link: https://lore.kernel.org/r/20260609202548.3571690-2-memxor@gmail.com Signed-off-by: Alexei Starovoitov <ast(a)kernel.org> Conflicts: include/linux/bpf.h, kernel/bpf/verifier.c [Whitelist only the field types present in this tree, add is_bpf_obj_drop_kfunc() matching KF_bpf_obj_drop_impl only and drop the percpu branch (no percpu obj_drop kfuncs here), and convert the selftest with its function body kept. Meanwhile, skip the selftest conversion for it does not call bpf_obj_drop.] Signed-off-by: Chen Yuxi <chenyuxi19(a)huawei.com> --- include/linux/bpf.h | 26 ++++++++++++++++++++++++++ kernel/bpf/verifier.c | 21 +++++++++++++++++++++ 2 files changed, 47 insertions(+) diff --git a/include/linux/bpf.h b/include/linux/bpf.h index 4e9c38b14e04..de601dff733e 100644 --- a/include/linux/bpf.h +++ b/include/linux/bpf.h @@ -452,6 +452,32 @@ static inline bool btf_record_has_field(const struct btf_record *rec, enum btf_f return rec->field_mask & type; } +static inline bool btf_field_is_nmi_safe(enum btf_field_type type) +{ + switch (type) { + case BPF_SPIN_LOCK: + case BPF_TIMER: + case BPF_KPTR_UNREF: + case BPF_REFCOUNT: + return true; + default: + return false; + } +} + +static inline bool btf_record_has_nmi_unsafe_fields(const struct btf_record *rec) +{ + int i; + + if (IS_ERR_OR_NULL(rec)) + return false; + for (i = 0; i < rec->cnt; i++) { + if (!btf_field_is_nmi_safe(rec->fields[i].type)) + return true; + } + return false; +} + static inline void bpf_obj_init(const struct btf_record *rec, void *obj) { int i; diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c index f0a67a9f8d1a..f9531bbc8767 100644 --- a/kernel/bpf/verifier.c +++ b/kernel/bpf/verifier.c @@ -200,6 +200,7 @@ static int release_reference(struct bpf_verifier_env *env, int ref_obj_id); static void invalidate_non_owning_refs(struct bpf_verifier_env *env); static void invalidate_rcu_protected_refs(struct bpf_verifier_env *env); static bool in_rbtree_lock_required_cb(struct bpf_verifier_env *env); +static bool is_tracing_prog_type(enum bpf_prog_type type); static int ref_set_non_owning(struct bpf_verifier_env *env, struct bpf_reg_state *reg); static void specialize_kfunc(struct bpf_verifier_env *env, @@ -11372,6 +11373,11 @@ static int check_reg_allocation_locked(struct bpf_verifier_env *env, struct bpf_ return 0; } +static bool is_bpf_obj_drop_kfunc(u32 func_id) +{ + return func_id == special_kfunc_list[KF_bpf_obj_drop_impl]; +} + static bool is_bpf_list_api_kfunc(u32 btf_id) { return btf_id == special_kfunc_list[KF_bpf_list_push_front_impl] || @@ -12120,6 +12126,7 @@ static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn, struct bpf_reg_state *regs = cur_regs(env); const char *func_name, *ptr_type_name; bool sleepable, rcu_lock, rcu_unlock; + enum bpf_prog_type prog_type = resolve_prog_type(env->prog); struct bpf_kfunc_call_arg_meta meta; struct bpf_insn_aux_data *insn_aux; int err, insn_idx = *insn_idx_p; @@ -12162,6 +12169,20 @@ static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn, if (err < 0) return err; + if (is_bpf_obj_drop_kfunc(meta.func_id) && (is_tracing_prog_type(prog_type) || + /* is_tracing_prog_type() for now doesn't cover non-iterator tracing progs. */ + (prog_type == BPF_PROG_TYPE_TRACING && env->prog->expected_attach_type != BPF_TRACE_ITER + && !env->prog->sleepable))) { + struct btf_struct_meta *struct_meta; + + struct_meta = btf_find_struct_meta(meta.arg_btf, meta.arg_btf_id); + if (struct_meta && btf_record_has_nmi_unsafe_fields(struct_meta->record)) { + verbose(env, "%s cannot be used in tracing programs on types with NMI unsafe fields\n", + func_name); + return -EINVAL; + } + } + if (meta.func_id == special_kfunc_list[KF_bpf_rbtree_add_impl]) { err = push_callback_call(env, insn, insn_idx, meta.subprogno, set_rbtree_add_callback_state); -- 2.34.1
2 1
0 0
[PATCH openEuler-1.0-LTS] nfsd: RCU-protect cl_cb_session to fix use-after-free on session teardown
by Jinjiang Tu 30 Sep '26

30 Sep '26
From: Jeff Layton <jlayton(a)kernel.org> mainline inclusion from mainline-v7.3-rc1 commit 01c5d5f58a5db9b0ee5afba2e49d3157788687b2 category: bugfix bugzilla: https://atomgit.com/src-openeuler/kernel/issues/18799 CVE: CVE-2026-89708 Reference: https://git.kernel.org/pub/scm/linux/kernel/git/stable/linux.git/commit/?id… -------------------------------- After a DESTROY_SESSION the per-session teardown path can free a session while rpciod still holds an inflight callback rpc_task that dereferences clp->cl_cb_session. nfsd4_probe_callback_sync() flushes cl_callback_wq, but once nfsd4_run_cb_work() has called rpc_call_async() the rpc_task lives on rpciod; flushing the workqueue does not wait for it. rpc_shutdown_client() does drain rpciod tasks, but uses a 1-second wait_event_timeout — tasks stuck in rpc_delay() (e.g. 2-second NFS4ERR_DELAY retries) can outlive the drain. destroy path rpciod ------------ ------ unhash_session(ses) nfsd4_probe_callback_sync(clp) flush_workqueue(cl_callback_wq) /* returns; rpc_task still live */ nfsd4_put_session_locked(ses) free_session(ses) -> kfree(ses) nfsd4_cb_sequence_done() reads cb_clp->cl_cb_session /* freed slab */ A second window exists in nfsd4_process_cb_update(). When __nfsd4_find_backchannel() returns NULL because unhash_session() has already removed the destroyed session from cl_sessions, setup_callback_client() takes the v4.1 early return so clp->cl_cb_session = ses never fires and the field retains a pointer to the about-to-be-freed session. Fix both by converting cl_cb_session to an RCU-protected pointer: - Move the cl_cb_session = ses assignment in setup_callback_client() to after rpc_create() succeeds, so it is only published when a working backchannel exists. Clear cl_cb_session on the error return in nfsd4_process_cb_update(). Both stores use rcu_assign_pointer(). - Annotate cl_cb_session with __rcu. All rpciod-side readers use rcu_read_lock()/rcu_dereference() and check for NULL, bailing to the appropriate error or requeue path: encode_cb_sequence4args(), decode_cb_sequence4resok(), nfsd41_cb_get_slot(), nfsd41_cb_release_slot(), nfsd4_cb_prepare(), and nfsd4_cb_sequence_done(). - Switch __free_session() from kfree() to kfree_rcu() so the session slab is not reclaimed until after an RCU grace period, guaranteeing that rpciod readers inside rcu_read_lock() never dereference freed memory. - Pass the session pointer to the nfsd_cb_seq_status and nfsd_cb_free_slot tracepoints instead of having them re-read cl_cb_session. - nfsd4_cb_prepare() calls rpc_exit() when the session is NULL, routing through the done/release path to requeue the callback. Fixes: dcbeaa68dbbd ("nfsd4: allow backchannel recovery") Cc: stable(a)vger.kernel.org Reported-by: Chris Mason <clm(a)meta.com> Signed-off-by: Chris Mason <clm(a)meta.com> Signed-off-by: Jeff Layton <jlayton(a)kernel.org> Link: https://patch.msgid.link/20260530-nfsd-fixes-v2-2-f27e8eb4d974@kernel.org Signed-off-by: Chuck Lever <chuck.lever(a)oracle.com> Conflicts: fs/nfsd/nfs4callback.c fs/nfsd/nfs4state.c fs/nfsd/state.h [The target 4.18 tree lacks the nfsd_cb_seq_status/nfsd_cb_free_slot tracepoints and the modern cb slot-table (grab_slot/cb_held_slot, se_cb_seq_nr[] array, se_slots xarray, se_slot_gen/se_target_maxslots) that this upstream fix is built on; it uses the legacy single cl_cb_slot_busy slot and a scalar se_cb_seq_nr with a flexible-array se_slots[]. The RCU conversion was therefore adapted to those structures: only the cl_cb_session readers/writers and the kfree_rcu teardown are backported, the tracepoint-hunk and slot-table rework are skipped as not present in this tree.] Co-authored-by: BackportAgent@deepseek-v4-flash Signed-off-by: Hulk Robot <hulkrobot(a)huawei.com> Signed-off-by: Jinjiang Tu <tujinjiang(a)huawei.com> --- fs/nfsd/nfs4callback.c | 59 ++++++++++++++++++++++++++++++++++++------ fs/nfsd/nfs4state.c | 4 +-- fs/nfsd/state.h | 3 ++- 3 files changed, 55 insertions(+), 11 deletions(-) diff --git a/fs/nfsd/nfs4callback.c b/fs/nfsd/nfs4callback.c index 5ed13b765517..5a439124c6ef 100644 --- a/fs/nfsd/nfs4callback.c +++ b/fs/nfsd/nfs4callback.c @@ -353,12 +353,19 @@ static void encode_cb_sequence4args(struct xdr_stream *xdr, const struct nfsd4_callback *cb, struct nfs4_cb_compound_hdr *hdr) { - struct nfsd4_session *session = cb->cb_clp->cl_cb_session; + struct nfsd4_session *session; __be32 *p; if (hdr->minorversion == 0) return; + rcu_read_lock(); + session = rcu_dereference(cb->cb_clp->cl_cb_session); + if (!session) { + rcu_read_unlock(); + return; + } + encode_nfs_cb_opnum4(xdr, OP_CB_SEQUENCE); encode_sessionid4(xdr, session); @@ -370,6 +377,7 @@ static void encode_cb_sequence4args(struct xdr_stream *xdr, xdr_encode_empty_array(p); /* csa_referring_call_lists */ hdr->nops++; + rcu_read_unlock(); } /* @@ -396,21 +404,32 @@ static void encode_cb_sequence4args(struct xdr_stream *xdr, static int decode_cb_sequence4resok(struct xdr_stream *xdr, struct nfsd4_callback *cb) { - struct nfsd4_session *session = cb->cb_clp->cl_cb_session; + struct nfsd4_session *session; int status = -ESERVERFAULT; __be32 *p; u32 dummy; + rcu_read_lock(); + session = rcu_dereference(cb->cb_clp->cl_cb_session); + if (!session) { + rcu_read_unlock(); + cb->cb_seq_status = -NFS4ERR_BADSESSION; + return -NFS4ERR_BADSESSION; + } + /* * If the server returns different values for sessionID, slotID or * sequence number, the server is looney tunes. */ p = xdr_inline_decode(xdr, NFS4_MAX_SESSIONID_LEN + 4 + 4 + 4 + 4); - if (unlikely(p == NULL)) + if (unlikely(p == NULL)) { + rcu_read_unlock(); goto out_overflow; + } if (memcmp(p, session->se_sessionid.data, NFS4_MAX_SESSIONID_LEN)) { dprintk("NFS: %s Invalid session id\n", __func__); + rcu_read_unlock(); goto out; } p += XDR_QUADLEN(NFS4_MAX_SESSIONID_LEN); @@ -418,12 +437,14 @@ static int decode_cb_sequence4resok(struct xdr_stream *xdr, dummy = be32_to_cpup(p++); if (dummy != session->se_cb_seq_nr) { dprintk("NFS: %s Invalid sequence number\n", __func__); + rcu_read_unlock(); goto out; } dummy = be32_to_cpup(p++); if (dummy != 0) { dprintk("NFS: %s Invalid slotid\n", __func__); + rcu_read_unlock(); goto out; } @@ -431,6 +452,7 @@ static int decode_cb_sequence4resok(struct xdr_stream *xdr, * FIXME: process highest slotid and target highest slotid */ status = 0; + rcu_read_unlock(); out: cb->cb_seq_status = status; return status; @@ -801,9 +823,8 @@ static int setup_callback_client(struct nfs4_client *clp, struct nfs4_cb_conn *c if (!conn->cb_xprt || !ses) return -EINVAL; clp->cl_cb_conn.cb_xprt = conn->cb_xprt; - clp->cl_cb_session = ses; args.bc_xprt = conn->cb_xprt; - args.prognumber = clp->cl_cb_session->se_cb_prog; + args.prognumber = ses->se_cb_prog; args.protocol = conn->cb_xprt->xpt_class->xcl_ident | XPRT_TRANSPORT_BC; args.authflavor = ses->se_cb_sec.flavor; @@ -820,6 +841,8 @@ static int setup_callback_client(struct nfs4_client *clp, struct nfs4_cb_conn *c rpc_shutdown_client(client); return PTR_ERR(cred); } + if (clp->cl_minorversion != 0) + rcu_assign_pointer(clp->cl_cb_session, ses); clp->cl_cb_client = client; clp->cl_cb_cred = cred; return 0; @@ -926,6 +949,10 @@ static void nfsd4_cb_prepare(struct rpc_task *task, void *calldata) cb->cb_seq_status = 1; cb->cb_status = 0; if (minorversion) { + if (!rcu_access_pointer(clp->cl_cb_session)) { + rpc_exit(task, -EIO); + return; + } if (!cb->cb_holds_slot && !nfsd41_cb_get_slot(clp, task)) return; cb->cb_holds_slot = true; @@ -936,7 +963,7 @@ static void nfsd4_cb_prepare(struct rpc_task *task, void *calldata) static bool nfsd4_cb_sequence_done(struct rpc_task *task, struct nfsd4_callback *cb) { struct nfs4_client *clp = cb->cb_clp; - struct nfsd4_session *session = clp->cl_cb_session; + struct nfsd4_session *session; bool ret = true; if (!clp->cl_minorversion) { @@ -958,6 +985,15 @@ static bool nfsd4_cb_sequence_done(struct rpc_task *task, struct nfsd4_callback if (!cb->cb_holds_slot) goto need_restart; + rcu_read_lock(); + session = rcu_dereference(clp->cl_cb_session); + if (!session) { + rcu_read_unlock(); + clear_bit(0, &clp->cl_cb_slot_busy); + rpc_wake_up_next(&clp->cl_cb_waitq); + goto need_restart; + } + switch (cb->cb_seq_status) { case 0: /* @@ -978,16 +1014,21 @@ static bool nfsd4_cb_sequence_done(struct rpc_task *task, struct nfsd4_callback ret = false; break; case -NFS4ERR_DELAY: - if (!rpc_restart_call(task)) + if (!rpc_restart_call(task)) { + rcu_read_unlock(); goto out; + } rpc_delay(task, 2 * HZ); + rcu_read_unlock(); return false; case -NFS4ERR_BADSLOT: + rcu_read_unlock(); goto retry_nowait; case -NFS4ERR_SEQ_MISORDERED: if (session->se_cb_seq_nr != 1) { session->se_cb_seq_nr = 1; + rcu_read_unlock(); goto retry_nowait; } break; @@ -1000,7 +1041,8 @@ static bool nfsd4_cb_sequence_done(struct rpc_task *task, struct nfsd4_callback clear_bit(0, &clp->cl_cb_slot_busy); rpc_wake_up_next(&clp->cl_cb_waitq); dprintk("%s: freed slot, new seqid=%d\n", __func__, - clp->cl_cb_session->se_cb_seq_nr); + session->se_cb_seq_nr); + rcu_read_unlock(); if (task->tk_flags & RPC_TASK_KILLED) goto need_restart; @@ -1149,6 +1191,7 @@ static void nfsd4_process_cb_update(struct nfsd4_callback *cb) err = setup_callback_client(clp, &conn, ses); if (err) { nfsd4_mark_cb_down(clp, err); + rcu_assign_pointer(clp->cl_cb_session, ses); return; } } diff --git a/fs/nfsd/nfs4state.c b/fs/nfsd/nfs4state.c index ba6ae4626a6c..fdd657d38d6d 100644 --- a/fs/nfsd/nfs4state.c +++ b/fs/nfsd/nfs4state.c @@ -1713,7 +1713,7 @@ static void nfsd4_del_conns(struct nfsd4_session *s) static void __free_session(struct nfsd4_session *ses) { free_session_slots(ses); - kfree(ses); + kfree_rcu(ses, rcu_head); } static void free_session(struct nfsd4_session *ses) @@ -2207,7 +2207,7 @@ static struct nfs4_client *create_client(struct xdr_netobj name, clear_bit(0, &clp->cl_cb_slot_busy); copy_verf(clp, verf); rpc_copy_addr((struct sockaddr *) &clp->cl_addr, sa); - clp->cl_cb_session = NULL; + RCU_INIT_POINTER(clp->cl_cb_session, NULL); clp->net = net; return clp; } diff --git a/fs/nfsd/state.h b/fs/nfsd/state.h index c6e4ea41d93a..3c39a0b5ddf6 100644 --- a/fs/nfsd/state.h +++ b/fs/nfsd/state.h @@ -256,6 +256,7 @@ struct nfsd4_session { struct list_head se_conns; u32 se_cb_prog; u32 se_cb_seq_nr; + struct rcu_head rcu_head; struct nfsd4_slot *se_slots[]; /* forward channel slots */ }; @@ -337,7 +338,7 @@ struct nfs4_client { #define NFSD4_CB_FAULT 3 int cl_cb_state; struct nfsd4_callback cl_cb_null; - struct nfsd4_session *cl_cb_session; + struct nfsd4_session __rcu *cl_cb_session; /* for all client information that callback code might need: */ spinlock_t cl_lock; -- 2.43.0
2 1
0 0
[PATCH openEuler-1.0-LTS] SUNRPC: harden gss_krb5_unwrap_v2 against short tokens
by superdcc97@163.com 30 Sep '26

30 Sep '26
From: Chris Mason <clm(a)meta.com> mainline inclusion from mainline-v7.3-rc1 commit 6959297aaa9572783d620a226d73c3fb94494888 category: bugfix bugzilla: https://atomgit.com/src-openeuler/kernel/issues/18878 CVE: CVE-2026-89542 Reference: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?… -------------------------------- gss_krb5_unwrap_v2() reads the EC and RRC header fields at ptr+4 and ptr+6 before validating that the token is at least GSS_KRB5_TOK_HDR_LEN (16) bytes long, and its rotate_left() helper passes buf->len - base to xdr_buf_subsegment() without verifying that base <= buf->len. When a caller hands in a sub-16-byte token, or a token whose declared len leaves base past the end of the buffer, three distinct failures follow: gss_krb5_unwrap_v2(offset, len, buf) ptr = buf->head[0].iov_base + offset ec = *(ptr + 4) /* OOB read on short head */ rrc = *(ptr + 6) /* OOB read on short head */ rotate_left(offset + 16, buf, rrc) xdr_buf_subsegment(buf, &subbuf, base, buf->len - base) /* u32 wrap when base > len */ _rotate_left(&subbuf, shift) shift %= buf->len /* divide-by-zero when base == len */ After decryption, the cleanup arithmetic has the same shape: movelen = min_t(unsigned int, buf->head[0].iov_len, len); movelen -= offset + GSS_KRB5_TOK_HDR_LEN + headskip; BUG_ON(offset + GSS_KRB5_TOK_HDR_LEN + headskip + movelen > buf->head[0].iov_len); The BUG_ON re-adds the value just subtracted, so it reduces to min(A, B) > A and is permanently false; it cannot catch the unsigned underflow of movelen, which then drives a ~UINT_MAX-byte memmove(). Add four defense-in-depth guards inside the unwrap core so it is safe regardless of what its callers validate: - reject tokens with len - offset < GSS_KRB5_TOK_HDR_LEN before touching ptr+4/ptr+6; - bail from rotate_left() when buf->len <= base, covering both the underflow and zero-length cases; - return early from _rotate_left() when buf->len is zero, so the shift %= buf->len modulo cannot fault; - replace the dead BUG_ON with a live check that returns GSS_S_DEFECTIVE_TOKEN before the movelen subtraction. Fixes: de9c17eb4a91 ("gss_krb5: add support for new token formats in rfc4121") Cc: stable(a)vger.kernel.org Assisted-by: kres (claude-opus-4-7) Signed-off-by: Chris Mason <clm(a)meta.com> Reviewed-by: Jeff Layton <jlayton(a)kernel.org> Link: https://patch.msgid.link/20260524010213.557424-5-cel@kernel.org Signed-off-by: Chuck Lever <chuck.lever(a)oracle.com> Conflicts: net/sunrpc/auth_gss/gss_krb5_wrap.c [commit e01b2c79f4af is not backport] Signed-off-by: Dong Chenchen <dongchenchen2(a)huawei.com> --- net/sunrpc/auth_gss/gss_krb5_wrap.c | 11 +++++++++-- 1 file changed, 9 insertions(+), 2 deletions(-) diff --git a/net/sunrpc/auth_gss/gss_krb5_wrap.c b/net/sunrpc/auth_gss/gss_krb5_wrap.c index b85899bf9d79..33debe19cfa4 100644 --- a/net/sunrpc/auth_gss/gss_krb5_wrap.c +++ b/net/sunrpc/auth_gss/gss_krb5_wrap.c @@ -420,6 +420,8 @@ static void _rotate_left(struct xdr_buf *buf, unsigned int shift) int shifted = 0; int this_shift; + if (!buf->len) + return; shift %= buf->len; while (shifted < shift) { this_shift = min(shift - shifted, LOCAL_BUF_LEN); @@ -432,6 +434,8 @@ static void rotate_left(u32 base, struct xdr_buf *buf, unsigned int shift) { struct xdr_buf subbuf; + if (buf->len <= base) + return; xdr_buf_subsegment(buf, &subbuf, base, buf->len - base); _rotate_left(&subbuf, shift); } @@ -507,6 +511,9 @@ gss_unwrap_kerberos_v2(struct krb5_ctx *kctx, int offset, struct xdr_buf *buf) if (kctx->gk5e->decrypt_v2 == NULL) return GSS_S_FAILURE; + if ((s64)buf->len - (s64)offset <= GSS_KRB5_TOK_HDR_LEN) + return GSS_S_DEFECTIVE_TOKEN; + ptr = buf->head[0].iov_base + offset; if (be16_to_cpu(*((__be16 *)ptr)) != KG2_TOK_WRAP) @@ -573,9 +580,9 @@ gss_unwrap_kerberos_v2(struct krb5_ctx *kctx, int offset, struct xdr_buf *buf) * head buffer space rather than that actually occupied. */ movelen = min_t(unsigned int, buf->head[0].iov_len, buf->len); + if (movelen < offset + GSS_KRB5_TOK_HDR_LEN + headskip) + return GSS_S_DEFECTIVE_TOKEN; movelen -= offset + GSS_KRB5_TOK_HDR_LEN + headskip; - BUG_ON(offset + GSS_KRB5_TOK_HDR_LEN + headskip + movelen > - buf->head[0].iov_len); memmove(ptr, ptr + GSS_KRB5_TOK_HDR_LEN + headskip, movelen); buf->head[0].iov_len -= GSS_KRB5_TOK_HDR_LEN + headskip; buf->len -= GSS_KRB5_TOK_HDR_LEN + headskip; -- 2.43.0
2 1
0 0
[PATCH OLK-6.6] bpf: Reject bpf_obj_drop() from tracing progs
by Chen Yuxi 30 Sep '26

30 Sep '26
From: Justin Suess <utilityemal77(a)gmail.com> mainline inclusion from mainline-v7.2-rc1 commit 94c8d1c21be40a845357854f98ec07e21bb14bc9 category: bugfix bugzilla: https://atomgit.com/src-openeuler/kernel/issues/19869 Reference: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?… -------------------------------- bpf_obj_drop() runs bpf_obj_free_fields() synchronously for program-allocated objects. When such an object contains NMI unsafe fields, tracing programs that can run from arbitrary instrumented context can reach that destruction from unsafe contexts, including NMI. NMI is likely one instance of this problem, and other instances would include possible unsafe reentrancy. Deferring bpf_obj_drop() is not appealing either: it would add delayed-free machinery to a release operation that otherwise has straightforward synchronous ownership semantics. Reject bpf_obj_drop() and bpf_percpu_obj_drop() from tracing programs that may run from unsafe contexts unless every field in the object's BTF record is explicitly NMI safe. Do not reject sleepable BPF_PROG_TYPE_TRACING programs, since they are not the arbitrary/NMI contexts that motivate the restriction. Note that while bpf_rb_root and bpf_list_head would be NMI safe on their own to free, the objects recursively held by them may not be; be conservative and just mark them as not NMI safe for now. Use a whitelist for the NMI-safe field set instead of listing only known NMI unsafe fields. Locks, async fields, unreferenced kptrs, and refcounts are known to be NMI safe because their destruction is either a no-op, simple state reset, or async cancellation. Referenced kptrs, percpu referenced kptrs, uptrs, graph roots, graph nodes, and any future field type are rejected until audited for arbitrary tracing and NMI contexts. This is less susceptible to future changes in fields that were previously safe by exclusion, and to new fields being added without updating this check. Convert the existing recursive local-object drop success case to a syscall program in the same commit, since this verifier change makes the old tracing program form invalid. The test still exercises bpf_obj_drop() releasing a referenced task kptr from a safe program type. Fixes: ac9f06050a35 ("bpf: Introduce bpf_obj_drop") Signed-off-by: Justin Suess <utilityemal77(a)gmail.com> Co-developed-by: Kumar Kartikeya Dwivedi <memxor(a)gmail.com> Signed-off-by: Kumar Kartikeya Dwivedi <memxor(a)gmail.com> Link: https://lore.kernel.org/r/20260609202548.3571690-2-memxor@gmail.com Signed-off-by: Alexei Starovoitov <ast(a)kernel.org> Conflicts: include/linux/bpf.h, kernel/bpf/verifier.c [Whitelist only the field types present in this tree, add is_bpf_obj_drop_kfunc() matching KF_bpf_obj_drop_impl only and drop the percpu branch (no percpu obj_drop kfuncs here), and convert the selftest with its function body kept. Meanwhile, skip the selftest conversion for it does not call bpf_obj_drop.] Signed-off-by: Chen Yuxi <chenyuxi19(a)huawei.com> --- include/linux/bpf.h | 26 ++++++++++++++++++++++++++ kernel/bpf/verifier.c | 21 +++++++++++++++++++++ 2 files changed, 47 insertions(+) diff --git a/include/linux/bpf.h b/include/linux/bpf.h index 4e9c38b14e04..de601dff733e 100644 --- a/include/linux/bpf.h +++ b/include/linux/bpf.h @@ -452,6 +452,32 @@ static inline bool btf_record_has_field(const struct btf_record *rec, enum btf_f return rec->field_mask & type; } +static inline bool btf_field_is_nmi_safe(enum btf_field_type type) +{ + switch (type) { + case BPF_SPIN_LOCK: + case BPF_TIMER: + case BPF_KPTR_UNREF: + case BPF_REFCOUNT: + return true; + default: + return false; + } +} + +static inline bool btf_record_has_nmi_unsafe_fields(const struct btf_record *rec) +{ + int i; + + if (IS_ERR_OR_NULL(rec)) + return false; + for (i = 0; i < rec->cnt; i++) { + if (!btf_field_is_nmi_safe(rec->fields[i].type)) + return true; + } + return false; +} + static inline void bpf_obj_init(const struct btf_record *rec, void *obj) { int i; diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c index f0a67a9f8d1a..f9531bbc8767 100644 --- a/kernel/bpf/verifier.c +++ b/kernel/bpf/verifier.c @@ -200,6 +200,7 @@ static int release_reference(struct bpf_verifier_env *env, int ref_obj_id); static void invalidate_non_owning_refs(struct bpf_verifier_env *env); static void invalidate_rcu_protected_refs(struct bpf_verifier_env *env); static bool in_rbtree_lock_required_cb(struct bpf_verifier_env *env); +static bool is_tracing_prog_type(enum bpf_prog_type type); static int ref_set_non_owning(struct bpf_verifier_env *env, struct bpf_reg_state *reg); static void specialize_kfunc(struct bpf_verifier_env *env, @@ -11372,6 +11373,11 @@ static int check_reg_allocation_locked(struct bpf_verifier_env *env, struct bpf_ return 0; } +static bool is_bpf_obj_drop_kfunc(u32 func_id) +{ + return func_id == special_kfunc_list[KF_bpf_obj_drop_impl]; +} + static bool is_bpf_list_api_kfunc(u32 btf_id) { return btf_id == special_kfunc_list[KF_bpf_list_push_front_impl] || @@ -12120,6 +12126,7 @@ static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn, struct bpf_reg_state *regs = cur_regs(env); const char *func_name, *ptr_type_name; bool sleepable, rcu_lock, rcu_unlock; + enum bpf_prog_type prog_type = resolve_prog_type(env->prog); struct bpf_kfunc_call_arg_meta meta; struct bpf_insn_aux_data *insn_aux; int err, insn_idx = *insn_idx_p; @@ -12162,6 +12169,20 @@ static int check_kfunc_call(struct bpf_verifier_env *env, struct bpf_insn *insn, if (err < 0) return err; + if (is_bpf_obj_drop_kfunc(meta.func_id) && (is_tracing_prog_type(prog_type) || + /* is_tracing_prog_type() for now doesn't cover non-iterator tracing progs. */ + (prog_type == BPF_PROG_TYPE_TRACING && env->prog->expected_attach_type != BPF_TRACE_ITER + && !env->prog->sleepable))) { + struct btf_struct_meta *struct_meta; + + struct_meta = btf_find_struct_meta(meta.arg_btf, meta.arg_btf_id); + if (struct_meta && btf_record_has_nmi_unsafe_fields(struct_meta->record)) { + verbose(env, "%s cannot be used in tracing programs on types with NMI unsafe fields\n", + func_name); + return -EINVAL; + } + } + if (meta.func_id == special_kfunc_list[KF_bpf_rbtree_add_impl]) { err = push_callback_call(env, insn, insn_idx, meta.subprogno, set_rbtree_add_callback_state); -- 2.34.1
2 1
0 0
  • ← Newer
  • 1
  • 2
  • 3
  • 4
  • 5
  • ...
  • 2507
  • Older →

HyperKitty Powered by HyperKitty