commit 725bd2f3c81d54edccb66fead01d4c0c222e2231 Author: Greg Kroah-Hartman Date: Sat Oct 3 12:39:03 2026 +0200 Linux 6.18.55 Link: https://lore.kernel.org/r/20260930152340.591469096@linuxfoundation.org Tested-by: Florian Fainelli Tested-by: Peter Schneider Tested-by: Pavel Machek (CIP) Tested-by: Ron Economos Tested-by: Miguel Ojeda Tested-by: Brett A C Sheffield Tested-by: Wentao Guan Tested-by: Barry K. Nathan Signed-off-by: Greg Kroah-Hartman commit fa047f2a025003f5654ebd76539f16a6ead0401a Author: Jiayuan Chen Date: Tue Sep 1 18:47:35 2026 +0800 bpf: Reject key-less BTF for hash maps commit 0895a0c0734703be5532f3883c42db95615fd98b upstream. map_check_btf() allows a key-less BTF (btf_key_type_id == 0) only for maps that have a ->map_check_btf callback, and leaves the actual decision to that callback. Hash maps used to have no ->map_check_btf, so a key-less BTF was rejected outright. That changed when htab and rhtab gained a ->map_check_btf to register a dtor - htab in commit 1df97a7453ee ("bpf: Register dtor for freeing special fields") and rhtab in commit 6905f8601298 ("bpf: Allow special fields in resizable hashtab"). Neither looks at the key, so a key-less hash map now passes map_check_btf() and gets created. Reading it back through bpffs feeds the key type_id 0 into btf_type_seq_show(); btf_type_by_id() returns the void type, kind_ops[BTF_KIND_UNKN] is NULL, and btf_type_show() dereferences it: RIP: 0010:btf_type_show+0x223/0x2e0 kernel/bpf/btf.c:8232 RSP: 0018:ffffc9000399f868 EFLAGS: 00010206 RAX: dffffc0000000000 RBX: 0000000000000000 RCX: 0000000000000000 RDX: 0000000000000005 RSI: 0000000000000000 RDI: 0000000000000028 RBP: 0000000000000000 R08: 0000000000000001 R09: 0000000000000000 R10: ffffc9000399f970 R11: 0000000000000001 R12: ffffffff9b96b140 R13: ffffc9000399f8e0 R14: ffff88803d393c00 R15: 0000000000000003 CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 CR2: 0000200000000000 CR3: 000000003d213000 CR4: 0000000000352ef0 DR0: 0000000039ae8f55 DR1: 0000000000000000 DR2: 0000000000000000 DR3: 0000000000000000 DR6: 00000000ffff0ff0 DR7: 0000000000000400 Call Trace: btf_type_seq_show_flags+0xca/0x120 kernel/bpf/btf.c:8250 htab_map_seq_show_elem+0x12e/0x350 kernel/bpf/hashtab.c:1669 map_seq_show+0x13d/0x1e0 kernel/bpf/inode.c:293 traverse.part.0.constprop.0+0x107/0x650 fs/seq_file.c:112 traverse fs/seq_file.c:99 [inline] seq_read_iter+0x93f/0x1270 fs/seq_file.c:196 seq_read+0x344/0x4d0 fs/seq_file.c:163 vfs_read+0x1e4/0xb40 fs/read_write.c:572 ksys_pread64 fs/read_write.c:764 [inline] __do_sys_pread64 fs/read_write.c:772 [inline] __se_sys_pread64 fs/read_write.c:769 [inline] __x64_sys_pread64+0x1eb/0x250 fs/read_write.c:769 do_syscall_x64 arch/x86/entry/syscall_64.c:61 [inline] do_syscall_64+0x123/0x790 arch/x86/entry/syscall_64.c:84 entry_SYSCALL_64_after_hwframe+0x77/0x7f Reject a key-less BTF in htab_map_check_btf() and rhtab_map_check_btf(), restoring the previous behavior. Fixes: 1df97a7453ee ("bpf: Register dtor for freeing special fields") Fixes: 6905f8601298 ("bpf: Allow special fields in resizable hashtab") Reported-by: syzbot+37b56485bbbf90ad8489@syzkaller.appspotmail.com Closes: https://lore.kernel.org/all/6a8f4e88.27659fcc.2ceef7.0008.GAE@google.com/T/ Signed-off-by: Jiayuan Chen Acked-by: Ihor Solodrai Link: https://lore.kernel.org/r/20260901104924.346187-2-jiayuan.chen@linux.dev Signed-off-by: Alexei Starovoitov Signed-off-by: Greg Kroah-Hartman commit 98fd50aaaeb2a5854c9af1cf1e406c295147c428 Author: Bernard Pidoux Date: Sun Jun 14 19:15:27 2026 +0200 rose: don't warn on stray CALL_ACCEPTED/CLEAR_CONFIRMATION in state 3 rose_state3_machine() logs "ROSE: unknown %02X in state 3" at KERN_WARNING for any frame type it does not handle during data transfer. Two of them reach that default arm routinely on real AX.25 links and alarm sysops watching the console, even though they are harmless: - CALL_ACCEPTED (0x0F): a late or duplicated Call Accepted that arrives after the socket has already moved to STATE_3, typically a retransmission on a slow link. - CLEAR_CONFIRMATION (0x17): crossed clearing, or a clear confirmation left over from a previous incarnation of a reused logical channel. In both cases the frame is simply dropped: no state change, no teardown, nothing freed. Only the noise is a problem. Drop these two frame types silently. Keep reporting any other, genuinely unexpected frame type, but through net_warn_ratelimited() so a misbehaving peer cannot flood the kernel log. Signed-off-by: Bernard Pidoux Signed-off-by: Greg Kroah-Hartman commit 5454ea6e22e7d38ef86facf2f305af5c33b83500 Author: Bernard Pidoux Date: Thu Sep 10 15:35:58 2026 +0200 rose: use timer_shutdown_sync() for t0timer teardown rose_t0timer_expiry() re-arms itself via rose_start_t0timer() at its own tail. rose_neigh_put() and rose_remove_neigh() stop it with timer_delete_sync() before freeing (or unlinking) the neighbour, but that only guarantees the callback is not running *at the moment the call returns* -- it does nothing to stop the very invocation that was just waited out from re-arming the timer on its way out. That re-arm races the kfree() in rose_neigh_put(): the timer can fire again on freed memory, and rose_t0timer_expiry() -> rose_transmit_restart_request() -> rose_send_frame() -> ax25_send_frame() dereferences the freed neigh->digipeat, use-after-free. This matches a syzbot report (KASAN slab-use-after-free read in ax25_find_cb()) that stayed open with both cause and fix bisection failing -- consistent with a hole that timer_delete_sync() alone cannot close for a self-rearming timer, a case documented in its own kerneldoc ("there is no way to get this correct with timer_delete_sync()"). Use timer_shutdown_sync() instead in both call sites. Unlike timer_delete_sync(), it also marks the timer so that any further add_timer()/mod_timer() on it is silently ignored, which closes the window regardless of how the self-rearm and the free are interleaved. ftimer's handler is a no-op and never re-arms, but it is switched the same way for consistency: both timers are being torn down for good in both of these call sites. Reported-by: syzbot+caa052a0958a9146870d@syzkaller.appspotmail.com Closes: https://syzkaller.appspot.com/bug?extid=caa052a0958a9146870d Signed-off-by: Bernard Pidoux Signed-off-by: Greg Kroah-Hartman commit bcf4c43f797ced865ec26fde1e8b1ca1dbcc9836 Author: Bernard Pidoux Date: Mon Sep 7 14:39:32 2026 +0200 rose: guard rose_transmit_link() against a NULL neighbour rose_kick() and rose_write_internal() call rose_transmit_link(skb, rose->neighbour) without checking that rose->neighbour is still set. Since commit e8eb0c6faa88 ("rose: clear neighbour pointer after rose_neigh_put() in state machines"), the state machines routinely set rose->neighbour to NULL once the underlying AX.25 link is gone, while leaving the socket in ROSE_STATE_3. rose_kick() only checks the state, not the neighbour, so a write() on such a socket reaches rose_transmit_link() with neigh == NULL, which immediately dereferences neigh->loopback and crashes: Unable to handle kernel NULL pointer dereference at virtual address 0000000000000036 ... pc : rose_transmit_link+0x14/0x1c8 [rose] lr : rose_kick+0xec/0x198 [rose] Call trace: rose_transmit_link+0x14/0x1c8 [rose] (P) rose_kick+0xec/0x198 [rose] rose_sendmsg+0x260/0x3c8 [rose] Reproduced on f6bvp-8 (a real ROSE/FPAC node) when the BBS daemon wrote to a socket whose neighbour had just been cleared. Confirmed by recompiling rose_link.o with a static_assert on offsetof(struct rose_neigh, loopback), which matches the faulting address (0x36) exactly. Drop the skb and return early when neigh is NULL, matching what rose_transmit_link() already does for other frames it cannot send. Fixes: e8eb0c6faa88 ("rose: clear neighbour pointer after rose_neigh_put() in state machines") Signed-off-by: Bernard Pidoux Signed-off-by: Greg Kroah-Hartman commit 7d10439345e2ba96c5b7e7891b22dc5097b9d98c Author: Bernard Pidoux Date: Fri Sep 4 16:29:44 2026 +0200 rose: keep the route needed for routing loop detection rose_route_frame() walks the route list before parsing the Call Request facilities. When a CALL REQUEST arrives on a LCI that already belongs to a route, the "Remove an existing unused route" branch deletes that route with rose_remove_route() and breaks out of the walk, so that a new call is able to reuse a stale LCI. The routing loop detection that follows looks for a route recording the same random number and the same pair of callsigns as the incoming call. The route that has just been deleted is exactly the one it would have matched, so a call coming back to us around a routing loop is not detected: it is routed again, a new route is allocated, and the next copy repeats the cycle. Two nodes whose routes point at each other then relay the same call to each other as fast as the CPU allows. This was observed on an AX.25 network where two nodes each believed the route to a third one went through the other. The same CALL REQUEST, carrying the same random number, was re-sent every 400 us, interleaved with CLEAR REQUEST "Not obtainable", diagnostic 120, until the AX.25 link collapsed. Over nine days one node emitted 204 million frames on a link that normally carries a few frames per minute. The outgoing LCI stayed at 001 the whole time, because deleting the route freed it for immediate reuse. Parse the facilities before walking the route list, and in both branches remove the route only when the incoming call differs from the one it records. When it is the same call, keep the route and clear the call with diagnostic 120, which is what the loop detection would have done. The predicate is factored out as rose_same_call() and reused by the loop detection itself. Tested with two ROSE nodes routing a call back to each other. Before this change the kernel relayed 377 CALL REQUEST in ten seconds, never emitted a CLEAR, and the AX.25 link dropped. After it, the first call is relayed once, the next copy is cleared with diagnostic 120, and the link stays up. Signed-off-by: Bernard Pidoux Signed-off-by: Greg Kroah-Hartman commit 6c401ae426e20e53a3398a8e79afd0964c0f5110 Author: Tejun Heo Date: Wed Sep 30 08:45:28 2026 -0400 sched_ext: Derive SCX_RQ_IN_WAKEUP from the core enqueue flags [ Upstream commit df5cdc2c832ca4e8a6d774596b9005558761a403 ] schedule_deferred_locked() skips scheduling a deferred action while SCX_RQ_IN_WAKEUP is set and relies on the task_woken_scx() call that follows a wakeup enqueue to run it. enqueue_task_scx() sets the flag from the merged enqueue flags, which include the flags stashed for a remote activation. move_remote_task_to_local_dsq() thus sets SCX_RQ_IN_WAKEUP on the destination rq when the moved task was woken up, although no task_woken_scx() follows that activation. An IMMED insert into a busy destination requests a local reenqueue during that enqueue. The request gets linked but not scheduled and stays pending until an unrelated wakeup or preemption on that CPU runs the deferred actions. The IMMED task sits behind the running task in the meantime. If nothing runs them before the scheduler is disabled, the request outlives the scheduler and points into its freed per-cpu area, which the next scheduler dereferences from run_deferred(). Test the core enqueue flags for the wakeup bit. Only the core's wakeup path is followed by task_woken_scx(). Fixes: 57ccf5ccdc56 ("sched_ext: Fix enqueue_task_scx() truncation of upper enqueue flags") Cc: stable@vger.kernel.org # v7.1+ Reported-by: Andrea Righi Link: https://lore.kernel.org/all/20260916145807.3250167-1-arighi@nvidia.com/ Signed-off-by: Tejun Heo Reviewed-by: Andrea Righi [ Retargeted the patch to kernel/sched/ext.c with its older extra_enq_flags context. ] Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit d3c8fb7094c538195fc532e66a932675b3addd0d Author: Shardul Bankar Date: Wed Feb 4 22:34:40 2026 +0530 hfsplus: avoid double unload_nls() on mount failure commit ebebb04baefdace1e0dc17f7779e5549063ca592 upstream. The recent commit "hfsplus: ensure sb->s_fs_info is always cleaned up" [1] introduced a custom ->kill_sb() handler (hfsplus_kill_super) that cleans up the s_fs_info structure (including the NLS table) on superblock destruction. However, the error handling path in hfsplus_fill_super() still calls unload_nls() before returning an error. Since the VFS layer calls ->kill_sb() when fill_super fails, this results in unload_nls() being called twice for the same sbi->nls pointer: once in hfsplus_fill_super() and again in hfsplus_kill_super() (via delayed_free). Remove the explicit unload_nls() call from the error path in hfsplus_fill_super() to rely solely on the cleanup in ->kill_sb(). [1] https://lore.kernel.org/r/20251201222843.82310-3-mehdi.benhadjkhelifa@gmail.com/ Reported-by: Al Viro Link: https://lore.kernel.org/r/20260203043806.GF3183987@ZenIV/ Signed-off-by: Shardul Bankar Link: https://lore.kernel.org/r/20260204170440.1337261-1-shardul.b@mpiricsoftware.com Signed-off-by: Viacheslav Dubeyko Signed-off-by: Greg Kroah-Hartman commit 500c9259f5a668d079e3a9a8d8a3be056701a35e Author: Liew Rui Yan Date: Tue Sep 8 06:47:38 2026 -0700 mm/damon/core: fix unconditionally skip last region commit b3723b596b548c837a766aae3553c14a7b15af2b upstream. Once quota set, the charge_{target,addr}_from unconditionally skips and resets at the last region of the tracked target, so the last region can be skipped even when it has not been processed. Example: 1. Target has 2 regions: R1 (0-100 bytes) and R2 (100-200 bytes). 2. Quota is configured to process only 100 bytes per window. 3. Window 1: Processes R1 (0-100). Quota is full. charge_{target, addr}_from is saved at (Target, 100). 4. Window 2: The loop reaches R2. Because R2 is damon_last_region(t), the old code unconditionally returns true, skipping R2 entirely and resetting the charge_{target,addr}_from. Result: R2 is permanently skipped even though it has never been processed. However, it is important to note that this is a very minor issue. This is because it is triggered only when the previous window saved/kept charge_{target,addr}_from, and in the next window, all regions except the last region were skipped by damos_skip_charged_region(). Fix this by only resetting the charge_{target,addr}_from when last region is reached, only skipping when it is applied or cannot split. Link: https://lore.kernel.org/20260908134739.96919-1-sj@kernel.org Fixes: 50585192bc2e ("mm/damon/schemes: skip already charged targets and regions") Signed-off-by: Liew Rui Yan Reviewed-by: SJ Park Signed-off-by: SJ Park Signed-off-by: Andrew Morton Cc: # v5.16.x Signed-off-by: SJ Park Signed-off-by: Greg Kroah-Hartman commit f6660e9a3e1e6cb084a6288f9ff64290a70c5025 Author: Chunfeng Song Date: Thu Sep 10 05:51:10 2026 +0000 rust: net: phy: fix off-by-one bit positions in device status accessors commit 6fb0a9d9071f1ff0cc5cfc0782302d9c90d642cb upstream. The hand-written bitfield offsets in is_link_up(), is_autoneg_enabled() and is_autoneg_completed() were correct when the abstraction was merged: at that time autoneg, link, and autoneg_complete were at bits 13, 14, and 15 of struct phy_device's first bitfield unit. Commit 2796ff1e3dca ("net: phy: add flag is_genphy_driven to struct phy_device") later inserted is_genphy_driven just before autoneg, shifting the three fields up by one, so the accessors now read: is_link_up() reads bit 14 = autoneg is_autoneg_enabled() reads bit 13 = is_genphy_driven is_autoneg_completed() reads bit 15 = link The official ax88796b Rust driver uses all three accessors in its read_status() implementation, so it inherits the bug. phy_attach_direct() sets is_genphy_driven only when it falls back to the generic driver, and ax88796b has a real driver, so is_genphy_driven stays 0. The broken is_autoneg_enabled() therefore reads bit 13 as 0, compares it against AUTONEG_ENABLE (1), and always returns false, so read_status() never reaches the resolve_aneg_linkmode() call. The ordinary bindgen accessors take &self. Calling them through (*phydev).link() would create a shared reference to the complete bindings::phy_device, which is not appropriate for an object wrapped in Opaque. Use the bindgen-generated raw accessors (link_raw(), autoneg_raw(), and autoneg_complete_raw()) instead. They retain the bit positions and endianness handling generated from the C layout without creating a Rust reference to the complete phy_device. Drop the hand-written numbers together with the TODO comment that marked them as a stopgap. The raw accessors are only emitted by bindgen 0.71 and later, and were added at the Rust-for-Linux project's request, so this fix can only be backported to stable branches whose minimum bindgen version is at least that, hence the scope on the Cc: stable line below. Found by a static equivalence audit (C2RustDrv, a C-to-Rust driver migration tool) that compares hand-written bitfield offsets against the bindgen layout of struct phy_device. Verified by building the bindings and checking the generated accessors; no runtime testing was possible without PHY hardware. Fixes: 2796ff1e3dca ("net: phy: add flag is_genphy_driven to struct phy_device") Cc: stable@vger.kernel.org # Only 7.1.y and later (requires bindgen's raw pointer accessors). Link: https://github.com/rust-lang/rust-bindgen/issues/2674 Signed-off-by: Chunfeng Song Reviewed-by: FUJITA Tomonori Link: https://patch.msgid.link/20260910055110.167110-1-springbreeze@stu.pku.edu.cn Signed-off-by: Jakub Kicinski [ For 6.18.y, re-do the commit manually adjusting the bit numbers to make it work with the `bindgen` version available there as Chunfeng mentions in [1]. - Miguel ] Link: https://lore.kernel.org/stable/SEWP216MB9770116738884C46F156A686FAF98D2@SEWP216MB977011.KORP216.PROD.OUTLOOK.COM/ [1] Signed-off-by: Miguel Ojeda Signed-off-by: Greg Kroah-Hartman commit d50418db3b862afaf981e2581ca3da9d84fc1f5d Author: Greg Kroah-Hartman Date: Wed Sep 30 14:18:48 2026 +0200 Revert "rust: net: phy: fix off-by-one bit positions in device status accessors" This reverts commit f22d349940cb34e56f6e8cd86d69c36d3998b33a which is commit 6fb0a9d9071f1ff0cc5cfc0782302d9c90d642cb upstream. It was not meant to be backported before 7.1.y -- its tag said: Cc: stable@vger.kernel.org # Only 7.1.y and later (requires bindgen's raw pointer accessors). Cc: FUJITA Tomonori Cc: Jakub Kicinski Link: https://lore.kernel.org/stable/SEWP216MB9770116738884C46F156A686FAF98D2@SEWP216MB977011.KORP216.PROD.OUTLOOK.COM/ Acked-by: Chunfeng Song Signed-off-by: Miguel Ojeda Signed-off-by: Greg Kroah-Hartman commit c525772c304c971b0d3a1f2898dedbae653efe90 Author: Xie Bo Date: Wed Sep 30 07:44:12 2026 -0400 RISC-V: KVM: Serialize IMSIC attributes with vCPU migration [ Upstream commit 8ae12ccaec6ec74945d8c1ef39f2c1b8df779abc ] KVM device ioctls are not serialized against KVM_RUN. As a result, kvm_riscv_aia_imsic_rw_attr() can snapshot the physical CPU and HGEI of an IMSIC VS-file before a concurrent vCPU migration releases it. The HGEI can then be allocated to another vCPU before imsic_vsfile_rw() uses the stale tuple. A GET or SET attribute may consequently access the new owner's interrupt file. Serialize the entire IMSIC attribute operation with the target vCPU mutex. This prevents the VS-file from being migrated and recycled until the attribute access completes. Acquire the mutex killably so that the device ioctl remains interruptible while waiting for KVM_RUN to finish. Fixes: db8b7e97d613 ("RISC-V: KVM: Add in-kernel virtualization of AIA IMSIC") Cc: stable@vger.kernel.org Signed-off-by: Xie Bo Reviewed-by: Anup Patel Link: https://lore.kernel.org/r/20260810-imsic-attr-race-v2-1-00ed95ad321e@ultrarisc.com Signed-off-by: Anup Patel Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 5fdc644a4fea48410dff51805e8ba69ea517e635 Author: Jiakai Xu Date: Wed Sep 30 07:44:11 2026 -0400 RISC-V: KVM: Fix null pointer dereference in kvm_riscv_aia_imsic_rw_attr() [ Upstream commit aeb1d17d1af5924f7357d7204a293bd8fc06ea13 ] Add a null pointer check for imsic_state before dereferencing it in kvm_riscv_aia_imsic_rw_attr(). While the function checks that the vcpu exists, it doesn't verify that the vcpu's imsic_state has been initialized, leading to a null pointer dereference when accessed. The crash manifests as: Unable to handle kernel paging request at virtual address dfffffff00000006 ... kvm_riscv_aia_imsic_rw_attr+0x2d8/0x854 arch/riscv/kvm/aia_imsic.c:958 aia_set_attr+0x2ee/0x1726 arch/riscv/kvm/aia_device.c:354 kvm_device_ioctl_attr virt/kvm/kvm_main.c:4744 [inline] kvm_device_ioctl+0x296/0x374 virt/kvm/kvm_main.c:4761 vfs_ioctl fs/ioctl.c:51 [inline] ... The fix adds a check to return -ENODEV if imsic_state is NULL and moves isel assignment after imsic_state NULL check. Fixes: 5463091a51cfaa ("RISC-V: KVM: Expose IMSIC registers as attributes of AIA irqchip") Signed-off-by: Jiakai Xu Signed-off-by: Jiakai Xu Reviewed-by: Anup Patel Link: https://lore.kernel.org/r/20260127072219.3366607-1-xujiakai2025@iscas.ac.cn Signed-off-by: Anup Patel Stable-dep-of: 8ae12ccaec6e ("RISC-V: KVM: Serialize IMSIC attributes with vCPU migration") Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 4a04ed91575819fb1859311113e95700b709bec7 Author: Ilya Maximets Date: Tue Sep 29 20:20:25 2026 -0400 net/sched: act_ct: avoid modifying shared unconfirmed ct entry [ Upstream commit f85009dfcd65e5969526b0db7a49b5413746e630 ] In a case where skb with an unconfirmed ct entry gets cloned, we may end up processing both again but with different sets of extensions. The series of events: 1. The first clone wants to commit and runs the helpers wiring up the extension pointer into the expectation list. 2. Then it looses the confirmation keeping the entry unconfirmed. 3. Second clone now wants to commit labels or run NAT and adds the new extension for that breaking the pointer in the expectation list causing UAF on the destruction path later. While this is possible to trigger, there should be no practical network pipeline where we need to process both clones without modifications in the same zone. So, let's just reset the entry in case for some reason we got an skb with a shared one. This doesn't affect any known use cases, but avoids any potential problems with sharing and modification of the unconfirmed ct entry. Unlike openvswitch module, act_ct allows for NAT without commit. Changing that would be a uAPI break. So, act_ct needs to reset on NAT regardless of the commit flag to avoid reallocation of the extension space. This, however, doesn't really change the picture for sensible networking cases as there should be no need to run the same packet twice (before and after the clone) through conntrack without packet header or zone changes and without commit. The fixes tag points to the introduction of helpers, since that's the main UAF trigger for the sharing. Fixes: a21b06e73191 ("net: sched: add helper support in act_ct") Cc: stable@vger.kernel.org Reported-by: Axel Mierczuk Signed-off-by: Ilya Maximets Reviewed-by: Aaron Conole Reviewed-by: Xin Long Reviewed-by: Jamal Hadi Salim Link: https://patch.msgid.link/20260921145655.3167436-5-i.maximets@ovn.org Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 812c16a255e7773a544e88b39ae6daed25544d30 Author: Ilya Maximets Date: Tue Sep 29 20:20:22 2026 -0400 net/sched: act_ct: fix helper UAF due to extensions realloc [ Upstream commit dad19b59da050cb60d3f7023dac2a042a84bf0bd ] While calling the helpers, a raw pointer to the extensions area is wired into expectations list: -> nf_ct_helper() -> helper->help() -> nf_ct_expect_related_report() -> nf_ct_expect_insert() -> hlist_add_head_rcu(&exp->lnode, &master_help->expectations) In case the connection is not confirmed yet, more extensions can be added afterwards with *_ext_add() calls reallocating the extension space and leaving the now invalid pointer in the expectations list that is later accessed while removing the expectation. Make sure that helpers are called at the end after all the other extensions are already added. Note that the helper rejection now leaves the mark and labels set, but that's not different from how the NAT was handled before or how the mark and the labels were handled on confirmation failure. And there are no atomicity guarantees provided by the API anyway. Fixes: a21b06e73191 ("net: sched: add helper support in act_ct") Cc: stable@vger.kernel.org Reported-by: Axel Mierczuk Signed-off-by: Ilya Maximets Reviewed-by: Xin Long Reviewed-by: Jamal Hadi Salim Reviewed-by: Aaron Conole Link: https://patch.msgid.link/20260921145655.3167436-7-i.maximets@ovn.org Signed-off-by: Jakub Kicinski [ Preserved the existing add_helper condition that upstream had removed. ] Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 9ba16ce8f3faaba0f95b39c546d84a3c28a337a4 Author: Melody Wang Date: Tue Sep 29 09:31:08 2026 -0400 x86/sev: Make vTPM SVSM calls preemption-safe [ Upstream commit 6c43c72748fffd29dec15cd1f31e9a32949bc437 ] Two functions in the SVSM vTPM guest implementation do not disable preemption when fetching the SVSM Calling Area Address (CAA). The SVSM CAA is a per-CPU structure. When a thread is preempted and migrated to a different CPU after fetching the per-CPU CAA, the SVSM call will execute on the new CPU with the original CPU's CAA. Which is wrong. Move the CAA fetching operation inside svsm_perform_call_protocol() which disables interrupts around the SVSM call and thus runs preemption-safe. Fixes: 770de678bc28 ("x86/sev: Add SVSM vTPM probe/send_command functions") Signed-off-by: Melody Wang Signed-off-by: Borislav Petkov (AMD) Reviewed-by: Stefano Garzarella Cc: stable@vger.kernel.org Link: https://patch.msgid.link/a5bc0d4a2c462a0089109e145c21626b244b2ff0.1789345277.git.huibo.wang@amd.com Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit e43d42bd711e7773b3a2c45b980b0be0c0a97466 Author: Nikunj A Dadhania Date: Tue Sep 29 09:31:07 2026 -0400 x86/sev: Remove redundant ghcbs_initialized checks around __sev_{get,put}_ghcb() [ Upstream commit 9d8460a1c7a6b0f2dc6302e5d0f31d4e8c2a7913 ] After 3645eb7e3915 ("x86/fred: Fix early boot failures on SEV-ES/SNP guests"), __sev_{get,put}_ghcb() handle the early-boot GHCB fallback internally, making the ghcbs_initialized guards in __set_pages_state() and svsm_perform_call_protocol() redundant. Remove them. Also initialize state->ghcb to NULL in the early-boot path of __sev_get_ghcb() so that the ghcb_state is well-defined for all callers, even though __sev_put_ghcb() currently returns early before reading it. No functional change intended. Suggested-by: Tom Lendacky Signed-off-by: Nikunj A Dadhania Signed-off-by: Borislav Petkov (AMD) Reviewed-by: Tom Lendacky Link: https://patch.msgid.link/20260518102230.3394603-1-nikunj@amd.com Stable-dep-of: 6c43c72748ff ("x86/sev: Make vTPM SVSM calls preemption-safe") Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 524becfebcd0f74327581765ee90e8144a3a8f4a Author: Borislav Petkov (AMD) Date: Tue Sep 29 09:31:06 2026 -0400 x86/sev: Carve out the SVSM code into a separate compilation unit [ Upstream commit e21279b73ef6c7d27237912c914f2db0c5a74786 ] Move the SVSM-related machinery into a separate compilation unit in order to keep sev/core.c slim and "on-topic". No functional changes. Reviewed-by: Tom Lendacky Signed-off-by: Borislav Petkov (AMD) Link: https://patch.msgid.link/20251204124809.31783-4-bp@kernel.org Stable-dep-of: 6c43c72748ff ("x86/sev: Make vTPM SVSM calls preemption-safe") Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit b2a9d6ff593b51e201a777d0113993ad037f03e4 Author: Borislav Petkov (AMD) Date: Tue Sep 29 09:31:05 2026 -0400 x86/sev: Add internal header guards [ Upstream commit f01c6489ad6ceb8d06eae0c5123fc6cf39276ff1 ] All headers need guards ifdeffery. Reviewed-by: Tom Lendacky Signed-off-by: Borislav Petkov (AMD) Link: https://patch.msgid.link/20251204124809.31783-3-bp@kernel.org Stable-dep-of: 6c43c72748ff ("x86/sev: Make vTPM SVSM calls preemption-safe") Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 1cad4f235e1502244996da3fa476cb0551e69bf4 Author: Borislav Petkov (AMD) Date: Tue Sep 29 09:31:04 2026 -0400 x86/sev: Move the internal header [ Upstream commit c1e8980fabf5d0106992a430284fac28bba053a6 ] Move the internal header out of the usual include/asm/ include path because having an "internal" header there doesn't really make it internal - quite the opposite - that's the normal arch include path. So move where it belongs and make it really internal. No functional changes. Reviewed-by: Tom Lendacky Signed-off-by: Borislav Petkov (AMD) Link: https://lore.kernel.org/r/20251204145716.GDaTGhTEHNOtSdTkEe@fat_crate.local Stable-dep-of: 6c43c72748ff ("x86/sev: Make vTPM SVSM calls preemption-safe") Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 99607d955be3bff72a885334193948378bc29e28 Author: Dapeng Mi Date: Wed Sep 30 07:14:25 2026 -0400 perf/x86/intel: Remove incorrect LionCove PEBS data-source constraints [ Upstream commit 7c944595cc43a664019edc22505b6d6f6039be5e ] On Lion Cove, PEBS data source is valid only for these events: - MEM_TRANS_RETIRED.LOAD_LATENCY (0x1cd) - MEM_TRANS_RETIRED.STORE_SAMPLE (0x2cd) The perfmon database (https://github.com/intel/perfmon) previously tagged additional memory events such as MEM_INST_RETIRED.STLB_MISS_LOADS with L1_Hit_Indication, implying PEBS data-source support, which is incorrect. The database has since been fixed, but intel_lnc_pebs_event_constraints[] still follows the old definition and marks those events as data-source capable. As a result, get_data_src() may decode data-source information for events that do not provide valid PEBS data-source data and mislead users. Remove those non-data-source memory events from the Lion Cove PEBS constraint table so matching falls back to the regular non-PEBS constraints, which already provide the same counter constraints. Also update lnc_latency_data() to decode LOAD/STORE flags explicitly when setting memory operation direction, for consistency with other *_latency_data() helpers. Fixes: a932aa0e868f ("perf/x86: Add Lunar Lake and Arrow Lake support") Signed-off-by: Dapeng Mi Signed-off-by: Peter Zijlstra (Intel) Signed-off-by: Ingo Molnar Cc: Link: https://patch.msgid.link/20260917015234.981153-6-dapeng1.mi@linux.intel.com Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 7ba97aa16778e9d9795a5aa8d1a99ffc5561efc3 Author: Dapeng Mi Date: Wed Sep 30 07:14:24 2026 -0400 perf/x86/intel: Update event constraints and cache_extra_regsfor LNL [ Upstream commit 331c3e4fa39a87560c09bdd878652090ae040b69 ] Update perf hard-coded event constraints and cache_extra_regs[] for Lunarlake according to the latest LNL perfmon events (V1.22). LNL introduces new extra register values for the OCR L3 cache events, so introduce lnc_hw_cache_extra_regs[] and skt_hw_cache_extra_regs[] to reflect the changes. LNL perfmon events: https://github.com/intel/perfmon/blob/main/LNL/events/lunarlake_lioncove_core.json https://github.com/intel/perfmon/blob/main/LNL/events/lunarlake_skymont_core.json Signed-off-by: Dapeng Mi Signed-off-by: Peter Zijlstra (Intel) Link: https://patch.msgid.link/20260515061143.338553-7-dapeng1.mi@linux.intel.com [Backport to 6.18: retain the Lion Cove PEBS constraint additions needed by 7c944595cc43 ("perf/x86/intel: Remove incorrect LionCove PEBS data-source constraints"). Keep the matching regular counter masks for TOPDOWN.MEMORY_BOUND_SLOTS (0x10a4, 0x8) and MEM_INST_RETIRED.ANY (0x87d0, 0x3ff), since that fix removes their PEBS-specific entries and falls back to the regular constraint table. Omit the unrelated cache extra-register tables and PMU initialization changes, which depend on cache-event refactoring and intel_pmu_init_cmt() absent from 6.18. Also omit the unrelated fixed-counter, non-PEBS OCR, 0x00e0, and Skymont constraint updates. No functions are added.] Stable-dep-of: 7c944595cc43 ("perf/x86/intel: Remove incorrect LionCove PEBS data-source constraints") Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit edc0309282aab70e1cda91214c6a465fd39e71b8 Author: Dapeng Mi Date: Wed Sep 30 07:14:22 2026 -0400 perf/x86/intel: Fix GRT PEBS load/store direction for latency events, to fix sample classification [ Upstream commit 89dc568e8c0be60e05e5fcd0b528c79077d7b84b ] On Gracemont, intel_grt_pebs_event_constraints[] applies LAT_CONSTRAINT constraints to MEM_UOPS_RETIRED.{LOAD,STORE}_LATENCY, but does not set explicit LOAD/STORE flags for those events. The PEBS latency path (pebs_latency_data(), via __grt_latency_data()) uses the event flags to determine memory operation direction. Without an explicit STORE flag, samples from MEM_UOPS_RETIRED.STORE_LATENCY can be misclassified as LOADs. Set explicit LOAD/STORE flags in intel_grt_pebs_event_constraints[] for: - MEM_UOPS_RETIRED.LOAD_LATENCY - MEM_UOPS_RETIRED.STORE_LATENCY Also update __grt_latency_data() to explicitly interpret these flags when assigning the sampled memory operation direction. This fixes incorrect STORE sample classification. Fixes: 39a41278f041 ("perf/x86/intel: Fix PEBS memory access info encoding for ADL") Signed-off-by: Dapeng Mi Signed-off-by: Peter Zijlstra (Intel) Signed-off-by: Ingo Molnar Cc: # v7.2+ Link: https://patch.msgid.link/20260917015234.981153-2-dapeng1.mi@linux.intel.com Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 51199c793bc057456cc1bad986593beab56aeef5 Author: Dapeng Mi Date: Wed Sep 30 07:14:21 2026 -0400 perf/x86/intel: Update event constraints and cache_extra_regsfor ADL [ Upstream commit 4ef863352bcde482d65722ed721c3fd3967a67d6 ] Update perf hard-coded event constraints and cache_extra_regs[] for Alderlake according to the latest ADL perfmon events (V1.39). One important note is that ADL has differences on the L3/node related OCR events although it shares same uarch with SPR server, e.g., ADL has different extra MSR values and no node events. So some variants of structures and functions are introduced to reflect these differences, like adl_glc_hw_cache_event_ids[], adl_glc_hw_cache_extra_regs[] and intel_pmu_init_glc_hybrid(), etc. Please note these changes would temporarily impact other platforms like MTL/ARL-U which shares hard-coded event structures, but it would be fixed soon in subsequent patches. ADL perfmon events: https://github.com/intel/perfmon/blob/main/ADL/events/alderlake_goldencove_core.json https://github.com/intel/perfmon/blob/main/ADL/events/alderlake_gracemont_core.json Signed-off-by: Dapeng Mi Signed-off-by: Peter Zijlstra (Intel) Link: https://patch.msgid.link/20260515061143.338553-5-dapeng1.mi@linux.intel.com [Stable dependency adaptation] Retain only the Gracemont STORE_LATENCY (event 0x6d0) PEBS counter-mask update from 0xf to 0x3f in intel_grt_pebs_event_constraints[]. This is the preimage required by target 89dc568e8c0be60e05e5fcd0b528c79077d7b84b. The existing latency decoder and LOAD/STORE constraint macros are already available in this stable tree, so no new functions are needed. Drop all core.c changes: the cache-event tables, offcore MSR masks, fixed-event aliases and hybrid initialization changes are unrelated to the target. Their upstream context includes cache tables absent from 6.18, and retaining them would bring in an unnecessary helper function and broader platform changes. [ sashal: Reduced backport -- upstream 4ef863352bcde touches 2 file(s), this backport carries 1. Not backported here: arch/x86/events/intel/core.c This note is generated from the file lists only; see the resolution record for the reasoning. ] Stable-dep-of: 89dc568e8c0b ("perf/x86/intel: Fix GRT PEBS load/store direction for latency events, to fix sample classification") Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 24e79508b237dd2a69bc6fb6ac5cf7f126178307 Author: Yuqi Xu Date: Tue Sep 29 21:21:31 2026 -0400 net: ipconfig: bound DHCP option construction [ Upstream commit e47a1958e12abc3a17b5231a4f21c8f1bf662e08 ] ic_dhcp_init_options() appends the hostname (option 12), vendor-class (option 60) and client-ID (option 61) options into the fixed 312-byte bootp_pkt.exten[] buffer. Only the client-ID branch checked the remaining space; the hostname and vendor-class writes were unbounded. A 64-byte hostname together with the maximum 252-byte dhcpclass= identifier needs 18 + (2 + 64) + (2 + 252) = 338 of the 312 available bytes even before the terminating END marker, so the vendor-class memcpy runs past the end of exten[]. With CONFIG_FORTIFY_SOURCE this is reported as a field-spanning write and, when the kernel is booted with panic_on_warn=1, aborts boot with a panic. Route the optional options through a common helper that makes sure the option, its 2-byte header and the END marker all fit and drops an option that would not. Configurations with short options keep sending exactly the same bytes as before. Fixes: 130c0f47fdf9 ("ipconfig: send host-name in DHCP requests") Cc: stable@vger.kernel.org Reported-by: Vega Assisted-by: LLM Signed-off-by: Yuqi Xu Reviewed-by: Ren Wei Reviewed-by: Simon Horman Link: https://patch.msgid.link/7808dfbfa2162dfd0b19f59aff5742d6e0db2abb.1789798023.git.xuyuqiabc@gmail.com Signed-off-by: Paolo Abeni Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 8fbfa861dfd4fe6a188942a3cc495c75b545e196 Author: Thorsten Blum Date: Tue Sep 29 21:21:30 2026 -0400 net: ipconfig: Remove outdated comment and indent code block [ Upstream commit e405b3c9d4aaa10972525fe2ea5cf94224020561 ] The comment has been around ever since commit 1da177e4c3f4 ("Linux-2.6.12-rc2") and can be removed. Remove it and indent the code block accordingly. Signed-off-by: Thorsten Blum Link: https://patch.msgid.link/20260109121128.170020-2-thorsten.blum@linux.dev Signed-off-by: Jakub Kicinski Stable-dep-of: e47a1958e12a ("net: ipconfig: bound DHCP option construction") Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit bda487567d0b7b15dfa876511085c0666b9efad5 Author: Théo Lebrun Date: Tue Sep 29 21:13:00 2026 -0400 net: macb: fix dma_alloc_coherent() leak on macb_alloc() error paths [ Upstream commit 23d42b9a3bcd55b17d3b371544fedc708c2397e9 ] Fix 3 leaks in macb_alloc() error paths: - Tx buffer allocated but crossing a 4G boundary: Tx leaked. - Rx buffer allocation fails: Tx leaked. - Rx buffer allocated but crossing a 4G boundary: Tx & Rx leaked. This is because our error handling calls macb_free(bp) which in turn frees the buffers stored in bp->queues[0], but nothing has been stored in there. Fix by storing allocated buffers into bp->queues[0] ASAP. Fixes: 78d901897b3c ("net: macb: single dma_alloc_coherent() for DMA descriptors") Cc: stable@vger.kernel.org Signed-off-by: Théo Lebrun Reviewed-by: Nicolai Buchwitz Link: https://patch.msgid.link/20260918-macb-alloc-leak-v1-1-aba9a3d4f6e3@bootlin.com Signed-off-by: Paolo Abeni Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit aa4845a710ae9fee6a3831f5c338904cea77c3aa Author: Chen Changcheng Date: Tue Sep 29 20:04:27 2026 -0400 HID: alps: unregister DualPoint Stick input device on remove [ Upstream commit aa9dde93e05a837645fdfd577eea71f0733f0694 ] alps_input_configured() allocates a second input device ("DualPoint Stick") with input_allocate_device() and registers it, but the alps_driver struct has no .remove handler and input2 is not tracked in hdev->inputs. The default remove path (hid_hw_stop -> hidinput_disconnect) only iterates hdev->inputs, so input2 is never unregistered and leaks on every device removal. Add a .remove handler that stops the device first (preventing URB callbacks from touching input2 during teardown) and then unregisters input2. Fixes: 2562756dde55 ("HID: add Alps I2C HID Touchpad-Stick support") Cc: stable@vger.kernel.org Signed-off-by: Chen Changcheng Signed-off-by: Jiri Kosina Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 48c5c02bf6fde304fa7ed10e23879df87d010851 Author: Bastien Nocera Date: Tue Sep 29 20:04:26 2026 -0400 HID: hid-alps: Use pm_ptr instead of #ifdef CONFIG_PM [ Upstream commit 5e130f58629a2819c9d053507ad368160d25eb65 ] This increases build coverage and allows to drop an #ifdef. Signed-off-by: Bastien Nocera Signed-off-by: Jiri Kosina Stable-dep-of: aa9dde93e05a ("HID: alps: unregister DualPoint Stick input device on remove") Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit d0f6244f4af4479a05491067b470ba46dd1f693e Author: Liz Fong-Jones Date: Tue Sep 29 15:47:37 2026 -0400 PCI: Fix BAR resize for devices on a root bus [ Upstream commit d58384c22739848efe14b34e9586e4f1242f33c0 ] pci_do_resource_release_and_resize() releases device BARs that share a bridge window with the BAR being resized, but when the device sits directly on a root bus (pdev->bus->self == NULL) it then skips resource assignment entirely and returns success, leaving the BARs it just released unassigned (IORESOURCE_UNSET). Skipping pbus_reassign_bridge_resources() is correct in that case -- there is no bridge window to adjust -- but the device BARs still have to be reassigned. Before the BAR release was consolidated into the PCI core, this case worked for amdgpu because the driver released the BARs itself and then called pci_assign_unassigned_bus_resources() unconditionally after the resize, which assigns unassigned device BARs also on a root bus. Commit db92e3fef53e ("drm/amdgpu: Remove driver side BAR release before resize") removed that call, so nothing assigns the released BARs anymore. This breaks amdgpu completely on the SolidRun HoneyComb LX2K (NXP LX2160A, arm64, ACPI), where ACPI doesn't expose the Root Port so the GPU endpoint appears directly on a "root bus" of its segment: amdgpu 0004:01:00.0: BAR 0 [mem 0xa400000000-0xa40fffffff 64bit pref]: releasing amdgpu 0004:01:00.0: BAR 2 [mem 0xa410000000-0xa4101fffff 64bit pref]: releasing amdgpu 0004:01:00.0: sw_init of IP block failed -19 amdgpu 0004:01:00.0: amdgpu_device_ip_init failed amdgpu 0004:01:00.0: Fatal error during GPU init No error is logged because the resize path reports success; amdgpu then finds BAR 0 IORESOURCE_UNSET and bails out with -ENODEV. When there is no upstream bridge, call pci_bus_assign_resources() on the root bus to place the BARs released above, using the same alignment-sorted algorithm as normal enumeration instead of a manual per-BAR loop. This also walks the rest of the hierarchy under the root bus, as pci_assign_unassigned_bus_resources() used to for amdgpu before commit db92e3fef53e ("drm/amdgpu: Remove driver side BAR release before resize") removed that call -- the core-side fix that commit asked for ("such a problem should be fixed inside pci_resize_resource() instead"). pci_bus_assign_resources() returns void, so failure is detected by checking whether the released BARs are still assigned afterward; if not, roll back as in the bridged case. This is stricter than the bridged path -- it fails on any unplaced resource, not just required ones -- since a root bus typically has one shared window, and failing loudly seemed better than leaving something silently unassigned. The root bus path also had a locking bug that any fix here necessarily touches: the old "goto out" jumped to up_read(&pci_bus_sem) without a matching down_read() (as does the "goto restore" taken when pci_dev_res_add_to_list() fails in the release loop). Take pci_bus_sem before the BAR release loop so every path through the function holds it exactly once. Fixes: 337b1b566db0 ("PCI: Fix restoring BARs on BAR resize rollback path") Link: https://bugs.launchpad.net/ubuntu/+source/linux-hwe-7.0/+bug/2159596 Suggested-by: Ilpo Järvinen Assisted-by: Claude:claude-fable-5 checkpatch Assisted-by: Claude:claude-sonnet-5 Signed-off-by: Liz Fong-Jones [bhelgaas: commit log] Signed-off-by: Bjorn Helgaas Reviewed-by: Ilpo Järvinen Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260918035633.566823-1-lizf@honeycomb.io [ Replaced unavailable resource_assigned(dev_res->res) with the equivalent dev_res->res->parent check. ] Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 85b1112945f21d43b166c0e27f03647f51e4afc0 Author: Christian Brauner Date: Tue Sep 29 14:47:16 2026 -0400 super: make iterate_supers_type() deletion-safe [ Upstream commit 2d2a2d7aa98741b58f54cacc99b52024e4d865f9 ] iterate_supers_type() drops sb_lock while invoking the callback and keeps only a passive reference to the current superblock. That reference keeps the object allocated, but does not keep its s_instances node linked. After the iterator releases s_umount, final teardown can unlink the current s_instances node. The iterator then advances through a reinitialized node. With the current hlist it stops without visiting the remaining superblocks. The unlink moved from generic_shutdown_super() to kill_super_notify(), but the cursor lifetime has been unsafe since the helper was introduced. The CIFS DFS lookup can consequently miss a matching superblock and return -EINVAL. Move removal from fs_supers to put_super(), alongside removal from super_blocks, so a passive reference keeps both list nodes linked. Keep the filesystem module reference until then, since unlinking s_instances may touch type->fs_supers. Make sget_fc() skip SB_DEAD superblocks before invoking test(), and set SB_DEAD under sb_lock to serialize with those callbacks. This allows kernfs to free its private information after kill_anon_super() returns. Keep matching SB_DYING superblocks until SB_DEAD is set so concurrent mounts still wait for teardown before retrying. Fixes: 43e15cdbefea ("new helper: iterate_supers_type()") Reported-by: Karl Mehltretter Closes: https://lore.kernel.org/r/20260903013336.92081-1-kmehltretter@gmail.com Suggested-by: Jan Kara Cc: stable@vger.kernel.org Tested-by: Karl Mehltretter [kmehltretter: supplied the commit message] Signed-off-by: Karl Mehltretter Link: https://patch.msgid.link/20260909193034.7467-1-kmehltretter@gmail.com Reviewed-by: Jan Kara Signed-off-by: Christian Brauner (Amutable) [ Adapted s_passive reference counting to s_count using the existing __put_super() helper. ] Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit ccf96a4b8626a001c13c5836d8357f31c57bbd8e Author: Longlong Xia Date: Tue Sep 29 14:46:57 2026 -0400 mm/hugetlb: do not dissolve gigantic pages without runtime support [ Upstream commit a363c62a653cc8b3e21da9545fa4e028ef50f9c3 ] dissolve_free_hugetlb_folio() doesn't check hstate_is_gigantic_no_runtime(h) though remove_hugetlb_folio()/ update_and_free_hugetlb_folio() silently bail for such folios, so it frees a still-listed folio and, on vmemmap restore failure, the add_hugetlb_folio() rollback corrupts the free list. Link: https://lore.kernel.org/20260823044118.1097121-2-xialonglong2025@163.com Fixes: 6eb4e88a6d27 ("hugetlb: create remove_hugetlb_page() to separate functionality") Signed-off-by: Longlong Xia Signed-off-by: Andrew Morton Assisted-by: Codex:gpt-5.6-sol Acked-by: Muchun Song Cc: David Hildenbrand Cc: Miaohe Lin Cc: Michal Hocko Cc: Oscar Salvador Cc: Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 2ab4b6815d10ec5cf281adc238563276dfb63541 Author: Usama Arif Date: Tue Sep 29 14:46:56 2026 -0400 mm/hugetlb: create hstate_is_gigantic_no_runtime helper [ Upstream commit a743e0af503a633e4ca68a100d9b2a1a071fe8ae ] This is a common condition used to skip operations that cannot be performed on gigantic pages when runtime support is disabled. This helper is introduced as the condition will exist even more when allowing "overcommit" of gigantic hugepages. No functional change intended with this patch. Link: https://lkml.kernel.org/r/20251009172433.4158118-1-usamaarif642@gmail.com Signed-off-by: Usama Arif Suggested-by: Andrew Morton Reviewed-by: Shakeel Butt Reviewed-by: Kefeng Wang Acked-by: David Hildenbrand Acked-by: Oscar Salvador Cc: Johannes Weiner Cc: Muchun Song Cc: Rik van Riel Cc: SeongJae Park Signed-off-by: Andrew Morton Stable-dep-of: a363c62a653c ("mm/hugetlb: do not dissolve gigantic pages without runtime support") Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 4cef512a955e4008ae4436a4fc97726976df1fdf Author: Jinjiang Tu Date: Tue Sep 29 13:15:15 2026 -0400 mm/rmap: fix missing barrier between anon_vma init and vma->anon_vma publish [ Upstream commit b6ac0b3f6013c168f22cad97e79967accacb08e1 ] On arm64 server, we find that a task trying to grab the anon_vma lock triggers hungtask. INFO: task main:2354726 blocked for more than 120 seconds. Tainted: G E 5.10.0-0021.aarch64 #1 "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message. task:main state:D stack: 0 pid:2354726 ppid:2350673 flags:0x00000a01 Call trace: __switch_to+0x7c/0xbc __schedule+0x3b4/0x8a0 schedule+0x50/0xe0 rwsem_down_write_slowpath+0x3cc/0x6cc down_write+0x60/0x260 __anon_vma_prepare+0x6c/0x210 do_anonymous_page+0x258/0x660 handle_pte_fault+0x188/0x214 __handle_mm_fault+0x1b0/0x380 handle_mm_fault+0xf4/0x284 do_page_fault+0x19c/0x494 do_translation_fault+0xcc/0xf8 do_mem_abort+0x48/0xac el0_da+0x44/0x80 el0_sync_handler+0x88/0xb4 el0_sync+0x160/0x180 After analyzing the vmcore, we found the anon_vma->root->rwsem.count is -1. There is another anon_vma whose anon_vma->root->rwsem.count is 1, the anon_vma->root->rwsem.owner shows the lock is held, but the stack of the task shows the task doesn't hold the anon_vma lock. After adding more debugging info, we found __anon_vma_prepare() reuses anon_vma and triggers the UAF of anon_vma->root due to missing memory barrier, leading to locking and unlocking two different anon_vma->root, thus leading to an anon_vma will never be unlocked, and another anon_vma couldn't be locked anymore. This race requires two adjacent VMAs that are not merged but are anon_vma-compatible (e.g., they differ in VMA_ACCESS_FLAGS that can be changed by mprotect()). Two threads fault on each VMA concurrently, both calling __anon_vma_prepare() with only mmap_lock held for reading. THREAD A THREAD B __anon_vma_prepare __anon_vma_prepare find_mergeable_anon_vma() -> NULL anon_vma = anon_vma_alloc(); anon_vma->root = anon_vma; // the two stores may be reordered vma->anon_vma = anon_vma; // finds A's anon_vma anon_vma = find_mergeable_anon_vma(vma); anon_vma_lock_write(anon_vma); // may still see the old root down_write(&anon_vma->root->rwsem); anon_vma_unlock_write(anon_vma); // see the new root, never unlock old up_write(&anon_vma->root->rwsem); thread A triggers page fault and calls __anon_vma_prepare() to prepare anon_vma for the faulting vma. __anon_vma_prepare() allocates and initializes a new anon_vma, and then publishes it to the vma with a plain store. anon_vma_prepare() only requires the mmap_lock to be held for reading, so two threads can fault on adjacent VMAs at the same time. While thread A publishes a new anon_vma, thread B could find the anon_vma via find_mergeable_anon_vma() and then locks anon_vma->root->rwsem. The store to anon_vma->root in anon_vma_alloc() and the store to vma->anon_vma can be reordered. The anon_vma_lock_write() and spin_lock() only provide acquire semantics, which do not prevent prior stores from being reordered after them. The release semantics of the corresponding spin_unlock() and anon_vma_unlock_write() come too late, the store to vma->anon_vma is already published before they take effect. As a result, thread B can observe the following order: vma->anon_vma = anon_vma; anon_vma->root = anon_vma; The anon_vma slab is SLAB_TYPESAFE_BY_RCU, so a newly allocated anon_vma may reuse memory from a previously freed one. The constructor (anon_vma_ctor) does not reset anon_vma->root, and __put_anon_vma() doesn't clear it either, so the old root value persists until anon_vma_alloc() overwrites it. If that store isn't visible, thread B reads a root that points to the old anon_vma and locks it. As a result, thread B can call anon_vma_lock_write() with the old root, and call anon_vma_unlock_write() with the new root, leading to an anon_vma will never be unlocked, and another anon_vma couldn't be locked anymore (its count is dropped from 0 to -1 due to wrong unlock). To fix it, change the plain store `vma->anon_vma = anon_vma` to store release, so that the fields of anon_vma are visible before anon_vma is published to vma->anon_vma. At read side, the load of anon_vma and anon_vma->root have address dependency. According to Documentation/memory-barriers.txt and some investigations, only Alpha needs address-dependency barriers and it has been handled by READ_ONCE() in reusable_anon_vma(). We reproduced this issue in v5.10 with KSM enabled. The kernel doesn't merge commit cf7e7a3503df ("mm: prevent KSM from breaking VMA merging for new VMAs"), so there are many adjacent VMAs that aren't merged but are compatible for anon_vma. Without this fix, our production environment could reproduce this issue about 2-5 times each month. After adding a smp_mb() before anon_vma_lock_write(anon_vma) in __anon_vma_prepare(), which is different to this patch, this issue hasn't been reproduced for one month. Link: https://lore.kernel.org/20260908122924.554373-1-tujinjiang@huawei.com Fixes: 5c341ee1dfc8 ("mm: track the root (oldest) anon_vma") Signed-off-by: Jinjiang Tu Signed-off-by: Andrew Morton Reviewed-by: Lance Yang Reviewed-by: Lorenzo Stoakes (ARM) Acked-by: David Hildenbrand (Arm) Acked-by: Vlastimil Babka (SUSE) Cc: Minchan Kim Cc: Harry Yoo Cc: Hiroyouki Kamezawa Cc: Jann Horn Cc: Jinjiang Tu Cc: Kefeng Wang Cc: Larry Woodman Cc: Liam R. Howlett Cc: Nanyong Sun Cc: Rik van Riel Cc: Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit a56e5bcc43310efa45032158059c5362eeb19c01 Author: Lorenzo Stoakes Date: Tue Sep 29 13:15:14 2026 -0400 mm/rmap: allocate anon_vma_chain objects unlocked when possible [ Upstream commit bfc2b13b05a1343bb60a85d840fd8956731866c5 ] There is no reason to allocate the anon_vma_chain under the anon_vma write lock when cloning - we can in fact assign these to the destination VMA safely as we hold the exclusive mmap lock and therefore preclude anybody else accessing these fields. We only need take the anon_vma write lock when we link rbtree edges from the anon_vma to the newly established AVCs. This also allows us to eliminate the weird GFP_NOWAIT, GFP_KERNEL dance introduced in commit dd34739c03f2 ("mm: avoid anon_vma_chain allocation under anon_vma lock"), further simplifying this logic. This should reduce lock anon_vma contention, and clarifies exactly where the anon_vma lock is required. We cannot adjust __anon_vma_prepare() in the same way as this is only protected by VMA read lock, so we have to perform the allocation here under the anon_vma write lock and page_table_lock (to protect against racing threads), and we wish to retain the lock ordering. With this change we can simplify cleanup_partial_anon_vmas() even further - since we allocate AVC's without any lock taken and do not insert anything into the interval tree until after the allocations are tried, we can remove all logic pertaining to this and just free up AVC's only. Link: https://lkml.kernel.org/r/624bf1ac0bde4871fcfca2c8c8e294b6d8f7ae7b.1768746221.git.lorenzo.stoakes@oracle.com Signed-off-by: Lorenzo Stoakes Reviewed-by: Suren Baghdasaryan Reviewed-by: Liam R. Howlett Cc: Barry Song Cc: Chris Li Cc: David Hildenbrand Cc: Harry Yoo Cc: Jann Horn Cc: Michal Hocko Cc: Mike Rapoport Cc: Pedro Falcato Cc: Rik van Riel Cc: Shakeel Butt Cc: Vlastimil Babka Signed-off-by: Andrew Morton [Stable dependency adaptation for 6.18: Keep the split of anon_vma_chain_link() into anon_vma_chain_assign() and explicit interval-tree insertion at all three callers. Retain the fork change that assigns the chain before taking the anon_vma write lock; the interval-tree insertion remains protected by that lock. Drop the two-pass clone allocation and cleanup_partial_anon_vmas() changes. This tree lacks the prerequisite clone assertions, early unfaulted-VMA handling, simplified root locking, and partial-clone cleanup. Preserve the existing GFP_NOWAIT/GFP_KERNEL fallback, root locking, list traversal, reference accounting, and allocation-failure cleanup instead. No new functions are introduced. The helper split supplies the context needed for b6ac0b3f6013 ("mm/rmap: fix missing barrier between anon_vma init and vma->anon_vma publish") to merge cleanly while retaining this tree's interval-tree API. The memory barrier fix itself is left to that target commit.] Stable-dep-of: b6ac0b3f6013 ("mm/rmap: fix missing barrier between anon_vma init and vma->anon_vma publish") Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit fcbedba553a9c430079ee225df2dc7ab556a2f83 Author: shechenglong Date: Tue Sep 29 11:42:36 2026 -0400 drm/client: fix restore of partially initialized client [ Upstream commit 1fca688e9443003e33cf30453e7a7560367656c9 ] I got a null-ptr-deref report when closing a DRM file descriptor: WARNING: drivers/gpu/drm/drm_atomic.c:2031 at __drm_atomic_helper_set_config+0x18e/0x1b0 [drm] Call Trace: drm_client_modeset_commit_atomic+0x16b/0x220 [drm] drm_client_modeset_commit_locked+0x56/0x160 [drm] drm_client_modeset_commit+0x21/0x40 [drm] __drm_fb_helper_restore_fbdev_mode_unlocked.part.0+0x7b/0x80 drm_fbdev_client_restore+0xe/0x20 [drm_client_lib] drm_client_dev_restore+0x9f/0xc0 [drm] drm_release+0xc5/0xe0 [drm] The warning is followed by a NULL pointer dereference: BUG: kernel NULL pointer dereference, address: 0000000000000008 RIP: __drm_fb_helper_restore_fbdev_mode_unlocked.part.0+0x41/0x80 [drm_kms_helper] Call Trace: drm_fbdev_client_restore+0xe/0x20 [drm_client_lib] drm_client_dev_restore+0x9f/0xc0 [drm] drm_release+0xc5/0xe0 [drm] __fput+0xdc/0x2b0 __x64_sys_close+0x39/0x80 do_syscall_64+0x8d/0x460 entry_SYSCALL_64_after_hwframe+0x76/0x7e drm_client_register() adds the DRM client to the device client list before invoking the initial hotplug callback. If the hotplug callback fails, the client remains registered. For the fbdev client, a failure during drm_fb_helper_initial_config() causes the partially initialized fbdev helper to be cleaned up. drm_fb_helper_fini() releases fb_helper->info and leaves it NULL. The fbdev client therefore remains registered even though there is no fully initialized framebuffer device. Later, when userspace closes the DRM file descriptor, drm_release() can invoke the restore callbacks of registered DRM clients: drm_release() drm_client_dev_restore() drm_fbdev_client_restore() drm_fb_helper_restore_fbdev_mode_unlocked() drm_fbdev_client_restore() currently restores the fbdev state unconditionally. For a partially initialized fbdev client this can submit an incomplete modeset state and subsequently access fbdev state which has not been initialized, resulting in the warning and NULL pointer dereference above. drm_fbdev_client_unregister() already uses fb_helper->info to distinguish a fully probed framebuffer device from a partially initialized client. Use the same condition in drm_fbdev_client_restore() and skip restore if no framebuffer device has been successfully initialized. Signed-off-by: shechenglong Reviewed-by: Thomas Zimmermann Fixes: 5d08c44e47b9 ("drm/fbdev: Add memory-agnostic fbdev client") Signed-off-by: Thomas Zimmermann Cc: # v6.13+ Link: https://patch.msgid.link/20260907035147.1339-1-shechenglong@xfusion.com Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 933145e2bbb4cab6007b26d17f7f1722f14c5440 Author: Thomas Zimmermann Date: Tue Sep 29 11:42:35 2026 -0400 drm/client: Pass force parameter to client restore [ Upstream commit 943240d342f148896733eb6c7b223a08aa1f520a ] Add force parameter to client restore and pass value through the layers. The only currently used value is false. If force is true, the client should restore its display even if it does not hold the DRM master lock. This is be required for emergency output, such as sysrq. While at it, inline drm_fb_helper_lastclose(), which is a trivial wrapper around drm_fb_helper_restore_fbdev_mode_unlocked(). Signed-off-by: Thomas Zimmermann Reviewed-by: Jocelyn Falempe Link: https://patch.msgid.link/20251110154616.539328-2-tzimmermann@suse.de Stable-dep-of: 1fca688e9443 ("drm/client: fix restore of partially initialized client") Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 0589706ace3bdd47a821d66036d36402918cc2a7 Author: Thomas Zimmermann Date: Tue Sep 29 11:42:34 2026 -0400 drm/client: Remove holds_console_lock parameter from suspend/resume [ Upstream commit 7910d69376cde30e5871970d97d1a2e360568474 ] No caller of the client resume/suspend helpers holds the console lock. The last such cases were removed from radeon in the patch series at [1]. Now remove the related parameter and the TODO items. v2: - update placeholders for CONFIG_DRM_CLIENT=n Signed-off-by: Thomas Zimmermann Link: https://patchwork.freedesktop.org/series/151624/ # [1] Reviewed-by: Rodrigo Vivi Acked-by: Rodrigo Vivi Reviewed-by: Petr Vorel Reviewed-by: Andi Shyti Acked-by: Danilo Krummrich Reviewed-by: Lyude Paul Reviewed-by: Jocelyn Falempe Link: https://lore.kernel.org/r/20251001143709.419736-1-tzimmermann@suse.de Stable-dep-of: 1fca688e9443 ("drm/client: fix restore of partially initialized client") Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 817dd706ee6165812225581bebda038153120ba3 Author: Wentao Liang Date: Tue Sep 29 11:15:31 2026 -0400 drm/amdgpu: Fix acpi device leak in amdgpu_acpi_enumerate_xcc() [ Upstream commit a997baa61179b450bd55c4810c7ccfed3b753a54 ] amdgpu_acpi_enumerate_xcc() looks up each XCC ACPI device with acpi_dev_get_first_match_dev(), which takes a reference to the device. The reference is dropped with acpi_dev_put() after the XCC info is initialized, but if the kzalloc_obj() allocation of the XCC info fails the function returns -ENOMEM without releasing the reference, leaking the last reference to the ACPI device. Drop the ACPI device reference on the allocation failure path before returning. Fixes: 4d5275ab0b18 ("drm/amdgpu: Add parsing of acpi xcc objects") Reviewed-by: Lijo Lazar Signed-off-by: Wentao Liang Signed-off-by: Alex Deucher (cherry picked from commit 9211ef48b31ec66999cf55e04d0cbc60cd855fd5) Cc: stable@vger.kernel.org [ adapted the hunk to the existing kzalloc() failure block with braces and DRM_ERROR() logging. ] Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 13f299763f6c065210206ea0f17a149ec6c500bd Author: Ankit Nautiyal Date: Tue Sep 29 10:12:45 2026 -0400 drm/i915/quirks: Limit eDP rate to HBR2 on HP Pavilion Plus 14-ew1 [ Upstream commit 5271d81f99dd01d983d439930eb056952485e15e ] The eDP panel on the HP Pavilion Plus Laptop 14-ew1xxx advertises HBR3 while leaving the TPS4 support bit clear. The output however flickers, once link is trained with HBR3. Until commit 8c9006283e4b ("Revert "drm/i915/dp: Reject HBR3 when sink doesn't support TPS4"") such sinks were capped at HBR2 by the TPS4 check which incidentally kept this panel stable. That check was reverted because other panels legitimately need HBR3 without advertising TPS4, and the per-machine QUIRK_EDP_LIMIT_RATE_HBR2 was introduced to handle the affected machines instead. Add the machine to the list of devices that need the QUIRK_EDP_LIMIT_RATE_HBR2. Fixes: 8c9006283e4b ("Revert "drm/i915/dp: Reject HBR3 when sink doesn't support TPS4"") Reported-by: Annoy Cc Closes: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16743 Cc: # v6.18+ Tested-by: Annoy Cc Signed-off-by: Ankit Nautiyal Reviewed-by: Nemesa Garg Link: https://patch.msgid.link/20260907034555.2753846-1-ankit.k.nautiyal@intel.com (cherry picked from commit 550b703fdbb2a2022faa75b4b11ab135241afbd9) Signed-off-by: Jani Nikula Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 97e40aa29c10c72cb4fc02c39a63a64e3f17705d Author: Jouni Högander Date: Tue Sep 29 10:12:44 2026 -0400 drm/i915/psr: Disable PSR2 on Xiaomi Book Pro 14 2026 as a quirk [ Upstream commit 5e79af5db00b2a5f4667aaf16cfe4ecb7759383b ] Add new quirk (QUIRK_DISABLE_PSR2) for disabling PSR2 as a quirk for problematic setups. Apply this newly added quirk on Xiaomi Book Pro 14 2026. v2: logging adjusted Closes: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/7677 Signed-off-by: Jouni Högander Acked-by: Jani Nikula Reviewed-by: Mika Kahola Link: https://patch.msgid.link/20260417102350.28328-1-jouni.hogander@intel.com Stable-dep-of: 5271d81f99dd ("drm/i915/quirks: Limit eDP rate to HBR2 on HP Pavilion Plus 14-ew1") Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 66f814e8af4627c769c4d148bb8efb48cef79571 Author: Jouni Högander Date: Tue Sep 29 10:12:43 2026 -0400 drm/i915/psr: Fixes for Dell XPS DA14260 quirk [ Upstream commit 1de647abdfda9dc307503d0a85152161850ba52c ] Dell seems to be changing device ID even within same device model. Due to this we need to ignore device ID when applying quirk for Dell XPS 14 DA14260. Do this by adding DEVICE_ID_ANY and assign it to Dell XPS 14 DA14260 quirk. Also apply the quirk only for eDP Panel Replay. Fixes: 45c77d4bf8d4 ("drm/i915/psr: Disable Panel Replay on Dell XPS 14 DA14260 as a quirk") Cc: Mika Kahola Signed-off-by: Jouni Högander Reviewed-by: Mika Kahola Link: https://patch.msgid.link/20260320080403.1396926-1-jouni.hogander@intel.com Stable-dep-of: 5271d81f99dd ("drm/i915/quirks: Limit eDP rate to HBR2 on HP Pavilion Plus 14-ew1") Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 782b5210e89e76ddaa228ec64a1620244dd2ccfc Author: Jouni Högander Date: Tue Sep 29 10:12:42 2026 -0400 drm/i915/psr: Disable Panel Replay on Dell XPS 14 DA14260 as a quirk [ Upstream commit 45c77d4bf8d4d15453d709b9b828e498898e0751 ] Add new quirk (QUIRK_DISABLE_PANEL_REPLAY) for disabling Panel Replay as quirk for problematic setups. Apply this newly added quirk on Dell XPS 14 DA14260 if specific panel model is installed. We are observing problems with Dell XPS 14 DA14260. This device has certain LGD panel model which seems to be problematic. We have seen other LGD panel model with same OUI is working fine. Due to this we can't apply the quirk only based on panel OUI. There are also cases where same device model has differing panel model. We don't want to disable Panel Replay on such devices. Best we can do is to apply the quirk based on both device model and panel model. Closes: https://gitlab.freedesktop.org/drm/xe/kernel/-/issues/7521 Signed-off-by: Jouni Högander Reviewed-by: Mika Kahola Link: https://patch.msgid.link/20260317062402.1888624-1-jouni.hogander@intel.com Stable-dep-of: 5271d81f99dd ("drm/i915/quirks: Limit eDP rate to HBR2 on HP Pavilion Plus 14-ew1") Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 5815329988acf1d6c40f25ce4591014f43a44e9e Author: Viken Dadhaniya Date: Tue Sep 29 10:02:59 2026 -0400 i2c: qcom-geni: Fix hardcoded clock index in SE_GENI_CLK_SEL [ Upstream commit cb97bf3d4f91453b881acaf8e9f0cc47bb40b604 ] qcom_geni_i2c_conf() writes a hardcoded 0 to SE_GENI_CLK_SEL, which selects an index from the hardware clock performance table. This always picks the first table entry regardless of the actual source clock configuration. On platforms where the matching entry is not at index 0, the wrong source clock divider is active and the I2C bus runs at an incorrect frequency. Use geni_se_clk_freq_match() in geni_i2c_clk_map_idx() to find the performance table index for the source clock (32 MHz or 19.2 MHz). Store the resolved index in a new clk_idx field in geni_i2c_dev and write it to SE_GENI_CLK_SEL instead of the hardcoded 0. Fixes: 37692de5d523 ("i2c: i2c-qcom-geni: Add bus driver for the Qualcomm GENI I2C controller") Signed-off-by: Viken Dadhaniya Cc: # v4.19+ Reviewed-by: Mukesh Kumar Savaliya Signed-off-by: Andi Shyti Link: https://patch.msgid.link/20260921-i2c-fix-se-clk-conf-v2-1-8b5537ceff2d@oss.qualcomm.com [ Applied initialization error handling in geni_i2c_probe() because geni_i2c_resources_init() is absent. ] Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit f42562dd027dc4ed103fae17b2206e73ea1963c4 Author: Xiang Mei Date: Tue Sep 29 08:24:55 2026 -0400 vlan: require the MAC header to be present in __vlan_insert_inner_tag() [ Upstream commit ab888242fce4f16f6c4d4c6ec53939ad36aa3b3a ] __vlan_insert_inner_tag() only guarantees head room via skb_cow_head(), never that mac_len bytes of MAC header are present. Its ETH_HLEN wrappers - __vlan_insert_tag() under skb_vlan_push(), and vlan_insert_tag() under validate_xmit_vlan() on the generic transmit path - therefore rewrite the first 16 bytes at skb->data: a 12-byte memmove plus two 2-byte stores at +12 and +14. No caller supplies the bound, while the pop helpers use skb_ensure_writable()/pskb_may_pull(). An IFF_TUN device has hard_header_len == 0, so packet_snd() accepts a one-byte AF_PACKET/SOCK_RAW frame. The first vlan push only sets a hwaccel tag; the next - clsact "action vlan push" or bpf_skb_vlan_push() - enters the helper with skb->len still 1. The head comes from skbuff_small_head without __GFP_ZERO, so each push drags bytes from beyond skb->tail into the frame. After three the one-byte send leaves as 13 bytes carrying 11 bytes of uninitialised slab: 0000: 5a b3 62 12 80 88 ff ff 00 b3 62 12 81 `------------------------------' only 0x5a was sent; the rest is slab, here the top 56 bits of a linear-map address Require the MAC header the helper rewrites to be present, so such a frame is dropped rather than transmitted. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Cc: stable@vger.kernel.org Reported-by: co+0ea1ac045375cf05@bugs.sh Signed-off-by: Xiang Mei Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260915083152.705309-1-xmei5@asu.edu Signed-off-by: Jakub Kicinski [ adjusted context around skb_cow_head(skb, VLAN_HLEN) because the branch lacks upstream meta_len handling. ] Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 241e67b60308cdce25bba66616a109759edd1d40 Author: Hui Peng Date: Tue Sep 29 08:10:20 2026 -0400 cgroup/cpuset: Return PERR_NOCPUS in remote_partition_enable() on subpartitions_cpus conflict [ Upstream commit 31c88350b7dd1522792f726f79607f31bb55c50f ] When a remote partition is created underneath an existing local partition via a non-partition (PRS_MEMBER) intermediate cgroup, update_prstate() sees parent->partition_root_state == PRS_MEMBER and calls remote_partition_enable(). Commit 86888c7bd117 ("cgroup/cpuset: Add warnings to catch inconsistency in exclusive CPUs") replaced the cpumask_intersects(tmp->new_cpus, subpartitions_cpus) error check in remote_partition_enable() with WARN_ON_ONCE(). As a result, remote_partition_enable() emits a warning and proceeds to enable the remote partition on CPUs that are already owned by the ancestor local partition in subpartitions_cpus. This can be reproduced on Linux 7.3.0-rc3 with: mkdir -p /tmp/cg1 mount -t cgroup2 none /tmp/cg1 echo "+cpuset" > /tmp/cg1/cgroup.subtree_control mkdir /tmp/cg1/A echo 1 > /tmp/cg1/A/cpuset.cpus echo 1 > /tmp/cg1/A/cpuset.cpus.exclusive echo root > /tmp/cg1/A/cpuset.cpus.partition echo "+cpuset" > /tmp/cg1/A/cgroup.subtree_control mkdir /tmp/cg1/A/B echo 1 > /tmp/cg1/A/B/cpuset.cpus echo 1 > /tmp/cg1/A/B/cpuset.cpus.exclusive echo "+cpuset" > /tmp/cg1/A/B/cgroup.subtree_control mkdir /tmp/cg1/A/B/D echo 1 > /tmp/cg1/A/B/D/cpuset.cpus echo 1 > /tmp/cg1/A/B/D/cpuset.cpus.exclusive echo root > /tmp/cg1/A/B/D/cpuset.cpus.partition which triggers: WARNING: kernel/cgroup/cpuset.c:1594 at remote_partition_enable+0x1c1/0x300 and leaves both /tmp/cg1/A and /tmp/cg1/A/B/D as active root partitions claiming exclusive CPU 1. Fix this by returning PERR_NOCPUS when tmp->new_cpus intersects subpartitions_cpus in remote_partition_enable(), matching the error code used by remote_cpus_update() for the same subpartitions_cpus conflict, and add a regression test case to tools/testing/selftests/cgroup/test_cpuset_prs.sh. Tested in QEMU on Linux 7.3.0-rc3 using the reproducer above and tools/testing/selftests/cgroup/test_cpuset_prs.sh. Fixes: 86888c7bd117 ("cgroup/cpuset: Add warnings to catch inconsistency in exclusive CPUs") Suggested-by: Guopeng Zhang Cc: stable@vger.kernel.org Assisted-by: LLM Signed-off-by: Hui Peng Reviewed-by: Waiman Long Signed-off-by: Tejun Heo Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit d186e0171c8a4f8df5fa1a954f09ddc63d18b0e3 Author: Waiman Long Date: Tue Sep 29 08:10:19 2026 -0400 cgroup/cpuset: Ensure domain isolated CPUs stay in root or isolated partition [ Upstream commit b1034a690129acd8995137bf4462470b4a2aa690 ] Commit 4a74e418881f ("cgroup/cpuset: Check partition conflict with housekeeping setup") is supposed to ensure that domain isolated CPUs designated by the "isolcpus" boot command line option stay either in root partition or in isolated partitions. However, the required check wasn't implemented when a remote partition was created or when an existing partition changed type from "root" to "isolated". Even though this is a relatively minor issue, we still need to add the required prstate_housekeeping_conflict() call in the right places to ensure that the rule is strictly followed. The following steps can be used to reproduce the problem before this fix. # fmt -1 /proc/cmdline | grep isolcpus isolcpus=9 # cd /sys/fs/cgroup/ # echo +cpuset > cgroup.subtree_control # mkdir test # echo 9 > test/cpuset.cpus # echo isolated > test/cpuset.cpus.partition # cat test/cpuset.cpus.partition isolated # cat test/cpuset.cpus.effective 9 # echo root > test/cpuset.cpus.partition # cat test/cpuset.cpus.effective 9 # cat test/cpuset.cpus.partition root With this fix, the last few steps will become: # echo root > test/cpuset.cpus.partition # cat test/cpuset.cpus.effective 0-8,10-95 # cat test/cpuset.cpus.partition root invalid (partition config conflicts with housekeeping setup) Reported-by: Chen Ridong Signed-off-by: Waiman Long Reviewed-by: Chen Ridong Signed-off-by: Tejun Heo Stable-dep-of: 31c88350b7dd ("cgroup/cpuset: Return PERR_NOCPUS in remote_partition_enable() on subpartitions_cpus conflict") Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 8a64ac9f03fc9d41ca8b0e70d6a5d2d01e6b67a8 Author: Waiman Long Date: Tue Sep 29 08:10:18 2026 -0400 cgroup/cpuset: Move up prstate_housekeeping_conflict() helper [ Upstream commit 6cfeddbf4ade9202849d75c27c4d0c82b42c73d1 ] Move up the prstate_housekeeping_conflict() helper so that it can be used in remote partition code. Signed-off-by: Waiman Long Reviewed-by: Chen Ridong Signed-off-by: Tejun Heo Stable-dep-of: 31c88350b7dd ("cgroup/cpuset: Return PERR_NOCPUS in remote_partition_enable() on subpartitions_cpus conflict") Signed-off-by: Sasha Levin Signed-off-by: Greg Kroah-Hartman commit 438bf73ca9775aadc6df8b5c49c998ee17f3cdb4 Author: Jakub Kicinski Date: Wed Apr 29 15:29:38 2026 -0700 net: tls: fix silent data drop under pipe back-pressure [ Upstream commit 7e7be31bfdb066c1c780dcd6b1224078fc54063f ] tls_sw_splice_read() uses len when advancing rxm->offset / rxm->full_len after skb_splice_bits(), rather than copied (the actual number of bytes successfully spliced into the pipe). When the destination pipe cannot accept all the requested bytes, splice_to_pipe() returns fewer bytes than len, and 'len - copied' of data is effectively skipped over. Fixes: e062fe99cccd ("tls: splice_read: fix accessing pre-processed records") Link: https://patch.msgid.link/20260429222944.2139041-2-kuba@kernel.org Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 66a2179d02668f8df44b1a249899006f3ba38ea0 Author: Darrick J. Wong Date: Mon Sep 14 22:38:38 2026 -0700 xfs: fix cursor and pointer handling when recovering iunlink buckets commit 65f39d09d73718611cee40399179323b5d4ead00 upstream. LOLLM pointed out a bug in xlog_recover_iunlink_bucket: 1. We don't null out prev_ip after releasing it, which can lead to UAF problems if the inodegc flush call in the loop fails. at which point I noticed even more bugs: 2. If the inodegc flush inside the loop fails, we also leak @ip. 3. We set prev_agino to agino having already advanced agino, which results in inodes with i_prev_unlinked set to itself. 4. If we exit the bottom of the loop with prev_ip set, then prev_ip aliases ip and we also set its i_prev_unlinked to itself. Bugs 3 and 4 introduce loops into the unlinked list, though these loops don't surface because we immediately flush each unlinked inode after loading it. Fix all of these issues. Cc: stable@vger.kernel.org # v6.0 Fixes: 04755d2e5821b3 ("xfs: refactor xlog_recover_process_iunlinks()") Signed-off-by: Darrick J. Wong Assisted-by: LOLLM # finding obvious bugs Reviewed-by: Christoph Hellwig Signed-off-by: Carlos Maiolino Signed-off-by: Greg Kroah-Hartman commit 86b17699bd0c2248f33d27c00cf472f3b3ae361a Author: Darrick J. Wong Date: Mon Sep 14 22:38:23 2026 -0700 xfs: fix blockgc group quota scanning when usrquota isn't enforced commit f8f6382ff13109e19d0fc1d0224ff7d29641e56a upstream. LOLLM noticed the copy-paste error here -- if user quotas aren't enforced but we're near the group quota limit, we fail to set FLAG_GID and hence we might not actually free any preallocations, causing unnecessary EDQUOT. Fix that. Cc: stable@vger.kernel.org # v5.12 Fixes: c237dd7c709432 ("xfs: flush eof/cowblocks if we can't reserve quota for inode creation") Signed-off-by: Darrick J. Wong Assisted-by: LOLLM # finding obvious bugs Reviewed-by: Christoph Hellwig Signed-off-by: Carlos Maiolino Signed-off-by: Greg Kroah-Hartman commit 930a9c1c321efdaa3e7f8caecb0479cee3baf342 Author: Darrick J. Wong Date: Mon Sep 14 22:39:10 2026 -0700 xfs: don't let hidden_space go negative in xfs_metafile_resv_init commit 476582d754cdc5110f806001417fea6c77824c13 upstream. LOLLM points out that if the amount of fdblocks that we can reserve for a metadata btree file goes below the space already used by that file, then the hidden_space subtraction can underflow, causing xfs_dec_fdblocks to subtract a huge amount of space. We never want the target to be less than the used sapce, so fix the logic that adjusts dblocks_avail downwards. Also fix an error in the adjacent comment. Cc: stable@vger.kernel.org # v6.15 Fixes: 1df8d75030b787 ("xfs: make metabtree reservations global") Signed-off-by: Darrick J. Wong Assisted-by: LOLLM # finding obvious bugs Reviewed-by: Christoph Hellwig Signed-off-by: Carlos Maiolino Signed-off-by: Greg Kroah-Hartman commit 09ebf777ad4e21231f1ebf561a03beec8d1bec8d Author: Darrick J. Wong Date: Mon Sep 14 22:37:05 2026 -0700 xfs: call xfs_dquot_set_prealloc_limits if we installed default rtb limits commit 065f3ce5936e68da75f3dc18201d290073d78f3c upstream. Now that we have quotas for the realtime volume, we also have precomputed watermark limits for the realtime block counts. These precomputations should be done any time we change the rtb limits, which means that xfs_qm_adjust_dqlimits needs to ensure that if we installed a default rtb limit. Cc: stable@vger.kernel.org # v6.13 Fixes: 5dd70852b03901 ("xfs: create quota preallocation watermarks for realtime quota") Signed-off-by: Darrick J. Wong Reviewed-by: Christoph Hellwig Signed-off-by: Carlos Maiolino Signed-off-by: Greg Kroah-Hartman commit d46492fb8018d12d00651b06b9d240e747110b38 Author: Darrick J. Wong Date: Mon Sep 14 22:39:25 2026 -0700 xfs: fix wild memcpy access when formatting ondisk rtrefcount btree roots commit fe2f9135df43db849e74f03956d322ac20b59af7 upstream. LOLLM noticed that the inode btree root formatting methods copy too many bytes -- there's only one set of keys in node blocks, not two. This causes memory corruption of whatever's beyond the buffers. Cc: stable@vger.kernel.org # v6.14 Fixes: f0415af60f482a ("xfs: wire up a new metafile type for the realtime refcount") Signed-off-by: Darrick J. Wong Assisted-by: LOLLM # finding obvious bugs Reviewed-by: Christoph Hellwig Signed-off-by: Carlos Maiolino Signed-off-by: Greg Kroah-Hartman commit 543bc5fd7d8ec662686208090fc41d766a94b4e2 Author: Darrick J. Wong Date: Mon Sep 14 22:37:20 2026 -0700 xfs: fix rtgroup repair estimations commit 41c4c41cf6c44f98db2916e1781f537d9ba6461a upstream. When I added online fsck for realtime reflink, I forgot to update xrep_calc_rtgroup_resblks to factor in the size of the refcount btree when it guesses how much space we need to start a repair. This hasn't been a huge problem in practice because there are few filesystems with (a) realtime, (b) rtgroups, (c) reflink, and (d) no rmap. But let's fix this before someone stumbles upon it, especially since LOLLM flagged this for me. Cc: stable@vger.kernel.org # v6.14 Fixes: 83ccffc489975d ("xfs: online repair of the realtime refcount btree") Signed-off-by: Darrick J. Wong Assisted-by: LOLLM # finding obvious bugs Reviewed-by: Christoph Hellwig Signed-off-by: Carlos Maiolino Signed-off-by: Greg Kroah-Hartman commit 6aabf4f351cd8ee56b85982285cc50c8cbaaeb30 Author: Darrick J. Wong Date: Thu Sep 10 22:54:58 2026 -0700 xfs: check di_forkoff correctly in scrub commit e9193f2f1ce32d02b9230094ffdbfab715ab6137 upstream. The di_forkoff check in xchk_dinode is incorrect, according to LOLLM. XFS_DFORK_BOFF returns a byte count relative to the start of the literal area, not the start of the inode. Therefore, this check won't flag di_forkoff values that are larger than the literal area but not the inode size itself. Fix this check; sadly the old APTR code was correct. Cc: stable@vger.kernel.org # v6.8 Fixes: 6b5d917780219d ("xfs: dont cast to char * for XFS_DFORK_*PTR macros") Signed-off-by: Darrick J. Wong Assisted-by: LOLLM # finding obvious bugs Reviewed-by: Christoph Hellwig Signed-off-by: Carlos Maiolino Signed-off-by: Greg Kroah-Hartman commit f99c672471797bcfd5c162b0c548d6f4564defff Author: Darrick J. Wong Date: Thu Sep 10 22:54:11 2026 -0700 xfs: use the correct reservations for rtrmap/refcount recovery commit 471e0b6e2ddac9e16b8dc2153d6e5575fefb8a2f upstream. LOLLM noticed that we might reserve the wrong number of blocks for recovering rtrmap and rtrefcount updates after a crash. Fix that. Cc: stable@vger.kernel.org # v6.14 Fixes: 5e0679d1c62f25 ("xfs: support recovering rmap intent items targetting realtime extents") Signed-off-by: Darrick J. Wong Assisted-by: LOLLM # finding obvious bugs Reviewed-by: Christoph Hellwig Signed-off-by: Carlos Maiolino Signed-off-by: Greg Kroah-Hartman commit adb80d620dabebf2c60b53cac95aaf9234264d69 Author: Darrick J. Wong Date: Thu Sep 10 22:53:55 2026 -0700 xfs: don't call xfs_exchange_range_finish for a dry run commit 8fc18580ec17f90beac4c933fbe4c74dcd3b7f36 upstream. LOLLM noticed that we strip file privileges and whatnot even for a dry run. We also shouldn't flush dirty data to disk or trim COW staging events for a dry run. Neither of those behaviors are allowed by the manpage, so fix that by exiting early on DRY_RUN in various functions. Cc: stable@vger.kernel.org # v6.10 Fixes: 42672471f938cd ("xfs: bind together the front and back ends of the file range exchange code") Signed-off-by: Darrick J. Wong Assisted-by: LOLLM # finding obvious bugs Reviewed-by: Christoph Hellwig Signed-off-by: Carlos Maiolino Signed-off-by: Greg Kroah-Hartman commit 770d21b52c91b5d8af9e9e6958b34f39bc7bfc11 Author: Darrick J. Wong Date: Thu Sep 10 22:53:39 2026 -0700 xfs: check padding field in xfs_ioc_commit_range commit 3083ba8dde765a9ab2337f3db68d00724a6b1202 upstream. LOLLM points out that we don't check the ioctl padding field here, so let's do that. I don't think there are many users yet since exchrange requires a new feature flag, so it's a good time to try to plug this hole. Cc: stable@vger.kernel.org # v6.12 Fixes: 398597c3ef7fb1 ("xfs: introduce new file range commit ioctls") Signed-off-by: Darrick J. Wong Assisted-by: LOLLM # finding obvious bugs Reviewed-by: Christoph Hellwig Signed-off-by: Carlos Maiolino Signed-off-by: Greg Kroah-Hartman commit 09b7a63e11c7cec725346f3f93abe437ff0bec61 Author: Darrick J. Wong Date: Wed Sep 9 23:01:18 2026 -0700 xfs: use correct jiffies comparison function in xchk_maybe_relax commit 984aab2d905a8557fafb27cd9e8713d6d12b3437 upstream. LOLLM points out that we're supposed to use time_after_eq, not a raw >= operation here, or else jiffies wraps can go unnoticed. Fix this. Cc: stable@vger.kernel.org # v6.10 Fixes: 271557de7cbfde ("xfs: reduce the rate of cond_resched calls inside scrub") Signed-off-by: Darrick J. Wong Assisted-by: LOLLM # finding obvious bugs Reviewed-by: Christoph Hellwig Signed-off-by: Carlos Maiolino Signed-off-by: Greg Kroah-Hartman commit 6a65a68097a1903988b256d5b367da7b0863a1d5 Author: Darrick J. Wong Date: Wed Sep 9 23:00:47 2026 -0700 xfs: release orphanage dir inode if chown fails commit 1c32cdc986467eaffeedb6c5334852809555b82d upstream. LOLLM points out that we leak the igrab'd reference to the orphanage directory inode if chowning it fails. Fix that. Cc: stable@vger.kernel.org # v6.10 Fixes: 1e58a8ccf2597c ("xfs: move orphan files to the orphanage") Signed-off-by: Darrick J. Wong Assisted-by: LOLLM # finding obvious bugs Reviewed-by: Christoph Hellwig Signed-off-by: Carlos Maiolino Signed-off-by: Greg Kroah-Hartman commit ac76f8df55980e3e9cddfe8143ec89e17859b848 Author: Darrick J. Wong Date: Wed Sep 9 23:00:31 2026 -0700 xfs: fix attr fork block count checks in xrep_inode_blockcounts commit bb991b7f79dd34cc5f24db0f736bf75630c970e7 upstream. LOLLM points out that a file has an attr fork, it will call xchk_inode_count_blocks to set @ablocks to the number of fsblocks mapped by the attr fork; but then it'll compare @blocks (aka the count of fsblocks mapped by the data fork). We already checked that and we never do anything with @acount, so I think this is clearly a bug. Fix the comparison. Cc: stable@vger.kernel.org # v6.8 Fixes: 2d295fe65776d1 ("xfs: repair inode records") Signed-off-by: Darrick J. Wong Assisted-by: LOLLM # finding obvious bugs Reviewed-by: Christoph Hellwig Signed-off-by: Carlos Maiolino Signed-off-by: Greg Kroah-Hartman commit ef7a56c6630ff8b61168bcdf7fe2e49b1a04afd5 Author: Darrick J. Wong Date: Wed Sep 9 23:00:16 2026 -0700 xfs: don't assert when XFS_SCRUB_TYPE_HEALTHY scans return corruption commit afbccf99f7f82117cba9ad4b0b006692030f49e8 upstream. XFS_SCRUB_TYPE_HEALTHY is a synthentic scrub type so that xfs_scrub can tell the kernel "Hey, I finished a scan and saw no problems" and have the kernel forget that it saw indirect evidence of corruption. Unfortunately, as LOLLM points out, it's possible for the health system to record a new corruption just before xfs_scrub gets to XFS_SCRUB_TYPE_HEALTHY. In this case, the existing logic doesn't return early and instead wanders into unknown regions of type_to_health_flag and trips the assert because HEALTHY doesn't have a group assignment. Fix the logic so that we always return early for a HEALTHY scrub type, even if we decide not to call xchk_mark_all_healthy. Cc: stable@vger.kernel.org # v6.9 Fixes: a1f3e0cca41036 ("xfs: update health status if we get a clean bill of health") Signed-off-by: Darrick J. Wong Assisted-by: LOLLM # finding obvious bugs Reviewed-by: Christoph Hellwig Signed-off-by: Carlos Maiolino Signed-off-by: Greg Kroah-Hartman commit c2ac475ddd737875563553dec67e744fedc11de7 Author: Zihan Xi Date: Wed Sep 16 15:29:25 2026 +0000 smb: client: validate POSIX create context length commit fa2e9900dd2a3f5a1e7ef5a8c5e8d435feedbfcc upstream. parse_posix_ctxt() reads the fixed nlink, reparse_tag, and mode fields before checking that the POSIX create context contains them. A short context can pass the generic checks and still make these fixed-width reads run past its declared data. The current in-tree smb2_open_file() path passes a NULL posix pointer, so this handler is not reached on the ordinary open path. Still require the POSIX data to cover all three fields before reading them because the helper performs those unguarded reads. Keep the existing soft-failure behavior so malformed optional metadata does not fail the open. Fixes: 69dda3059e7a ("cifs: add SMB2_open() arg to return POSIX data") Cc: stable@vger.kernel.org Reported-by: Vega Assisted-by: LLM Co-developed-by: Luxing Yin Signed-off-by: Luxing Yin Signed-off-by: Zihan Xi Tested-by: Frank Sorenson Signed-off-by: Paulo Alcantara Signed-off-by: Greg Kroah-Hartman commit 31107f79ad33bdff8c233c88f9dc0cd3cc51f960 Author: Zihan Xi Date: Wed Sep 16 15:29:29 2026 +0000 smb: client: preserve create-context parsing errors commit 2e828035d5d736d904c238ae7ec3f77c0d6270bf upstream. smb2_compound_op() saves the result from compound_send_recv() in tmp_rc. For SMB2_OP_OPEN_QUERY it then parses the CREATE contexts, but the final assignment of rc from tmp_rc discards a parsing error. A malformed create-context response can therefore be reported as successful to smb2_query_path_info(). Keep a create-context parsing error in tmp_rc so it survives per-command response processing and is returned to the caller. Fixes: b07687edee99 ("cifs: Improve SMB2+ stat() to work also without FILE_READ_ATTRIBUTES") Cc: stable@vger.kernel.org Reported-by: Vega Assisted-by: LLM Co-developed-by: Luxing Yin Signed-off-by: Luxing Yin Signed-off-by: Zihan Xi Tested-by: Frank Sorenson Signed-off-by: Paulo Alcantara Signed-off-by: Greg Kroah-Hartman commit abab661eed00fa102d3c35ddf1cce4f687e718c1 Author: Zihan Xi Date: Wed Sep 16 15:29:27 2026 +0000 smb: client: clean up failed cached directory opens commit d2ff5fb93ea83034025850266b5eed391f96b825 upstream. open_cached_dir() sends CREATE and QUERY_INFO as a compound request. If the CREATE succeeds but a later command returns an error, the function must retain the CREATE FID so common cleanup can issue SMB2_close(). It also must not treat a response error as a valid CREATE. Validate the CREATE response before using its fields, record the FIDs, and mark the handle open before handling errors from later compound commands. Move the -EREMCHG reconnect handling before response validation so a missing response does not hide the reconnect request. Count the handle when it is marked open; confirmed close responses decrement the counter, while existing close retry behavior remains best effort on transport failures. Fixes: b0f6df737a1c ("cifs: cache FILE_ALL_INFO for the shared root handle") Cc: stable@vger.kernel.org Reported-by: Vega Assisted-by: LLM Co-developed-by: Luxing Yin Signed-off-by: Luxing Yin Signed-off-by: Zihan Xi Tested-by: Frank Sorenson Signed-off-by: Paulo Alcantara Signed-off-by: Greg Kroah-Hartman commit 05a7fd80024ff32b9b40ab1478bb172eeab1c386 Author: Zihan Xi Date: Wed Sep 16 15:29:24 2026 +0000 smb: client: fix create context out-of-bounds reads commit 67f4c1c6a1b51e203d986779299824d1c2c590a6 upstream. smb2_parse_contexts() validates the complete create-context area but does not limit each record to its Next field before dispatching it. A malformed chain can therefore expose bytes beyond the current context to a handler. The QFid handler also used a full response-structure cast although it only reads DiskFileId. The SMB2/SMB3 lease parsers made the same layout assumption: they read LeaseState and LeaseFlags at canonical offsets rather than at DataOffset. A valid non-canonical DataOffset could therefore yield unrelated in-bounds data, while a short DataLength was still accepted. Limit each context to its Next value, reject offsets before the context header, and reject malformed chains. Bound the name range by the current context and do not dispatch a known handler when DataLength is zero. Read the QFid DiskFileId only when the context data covers that field. Parse the lease context from DataOffset and require DataLength to match the v1 or v2 lease_context size used by ksmbd. A size mismatch skips lease parsing without failing the open. Fixes: b8c32dbb0deb ("CIFS: Request SMB2.1 leases") Fixes: f047390a097e ("CIFS: Add create lease v2 context for SMB3") Fixes: 89a5bfa350fa ("smb3: optimize open to not send query file internal info") Cc: stable@vger.kernel.org Reported-by: Vega Assisted-by: LLM Co-developed-by: Luxing Yin Signed-off-by: Luxing Yin Signed-off-by: Zihan Xi Tested-by: Frank Sorenson Signed-off-by: Paulo Alcantara Signed-off-by: Greg Kroah-Hartman commit 8f5b30cc12c409ff62891e47ed51e6bf0945a3ea Author: Hui Peng Date: Sat Sep 19 11:25:18 2026 +0000 Bluetooth: RFCOMM: fix NULL dereference of dlc->session in RFCOMM_CONNINFO commit 46f8ffd0a1f1eb6cbc94946a92c11ef601e228a1 upstream. The RFCOMM_CONNINFO getsockopt handler accepts a socket that is not connected as long as deferred setup is enabled: if (sk->sk_state != BT_CONNECTED && !rfcomm_pi(sk)->dlc->defer_setup) { err = -ENOTCONN; break; } l2cap_sk = rfcomm_pi(sk)->dlc->session->sock->sk; dlc->defer_setup is set in rfcomm_sock_init() when rfcomm_connect_ind() creates a child socket for an incoming connection on a listening socket that has BT_DEFER_SETUP enabled. It is never cleared afterwards. The session, however, can go away underneath it. rfcomm_recv_disc() forces the dlc state before tearing it down: d->state = BT_CLOSED; __rfcomm_dlc_close(d, err); The RFCOMM_DEFER_SETUP early return in __rfcomm_dlc_close() only covers BT_CONNECT, BT_CONFIG, BT_OPEN and BT_CONNECT2, so with the state already BT_CLOSED that switch does not match and the function falls through to rfcomm_dlc_unlink(), which sets d->session = NULL, while d->defer_setup stays 1. A getsockopt(SOL_RFCOMM, RFCOMM_CONNINFO) on the accepted socket after that point therefore skips the -ENOTCONN path -- sk->sk_state is BT_CLOSED, but dlc->defer_setup is still set -- and dereferences the NULL session. No race is needed: once the DISC has been processed, the dereference is unconditional. Reproduced on a KASAN kernel under QEMU with a BR/EDR peer emulated over /dev/vhci: the peer brings up an ACL link, opens L2CAP on the RFCOMM PSM, starts a session and sends SABM for a channel bound with BT_DEFER_SETUP, and sends DISC for that dlci after the socket has been accepted. getsockopt(SOL_RFCOMM, RFCOMM_CONNINFO) on the accepted socket then hits: Oops: general protection fault, probably for non-canonical address 0xdffffc0000000002: 0000 [#1] SMP KASAN PTI KASAN: null-ptr-deref in range [0x0000000000000010-0x0000000000000017] CPU: 1 UID: 0 PID: 150 Comm: init Tainted: G B 7.3.0-rc3-g5dd1818b15d9 Hardware name: QEMU Standard PC (i440FX + PIIX, 1996) RIP: 0010:rfcomm_sock_getsockopt+0x529/0x780 Call Trace: do_sock_getsockopt+0x3ad/0x7d0 __sys_getsockopt+0x10e/0x1b0 __x64_sys_getsockopt+0xc2/0x160 do_syscall_64+0xda/0x4b0 entry_SYSCALL_64_after_hwframe+0x77/0x7f 0x10 is the offset of sock in struct rfcomm_session; rfcomm_sock_getsockopt_old() is inlined into rfcomm_sock_getsockopt(). Commit 43a556b2fd43 ("Bluetooth: RFCOMM: take rfcomm_mutex for the deferred setup accept") fixed the same "a remote DISC clears the session while deferred setup is still flagged" problem in rfcomm_dlc_accept(); this is the remaining instance of it, in the getsockopt path. Deferred setup only leaves a socket usable here once it has reached BT_CONNECT2, so restrict the exception to that state and check that a session is actually present before following it. Fixes: bb23c0ab8246 ("Bluetooth: Add support for deferring RFCOMM connection setup") Cc: stable@vger.kernel.org Assisted-by: LLM Signed-off-by: Hui Peng Signed-off-by: Luiz Augusto von Dentz Signed-off-by: Greg Kroah-Hartman commit 416546cd6a54e9b68f2b91bc680b1ee17dcd77a1 Author: Aldo Ariel Panzardo Date: Tue Sep 15 12:59:52 2026 -0300 Bluetooth: mgmt: fix race in read_unconf_index_list() commit b5dbb41b212c50c095a4dbee3017a84fe94f033b upstream. read_unconf_index_list() counts unconfigured controllers before allocating its response, then checks the device flags again while filling it. hci_dev_list_lock stabilizes list membership, but it does not serialize the per-device flags. During asynchronous controller setup, the worker can set HCI_UNCONFIGURED and clear HCI_SETUP between the two passes. A controller omitted from the allocation count can then become eligible for the fill pass, causing an out-of-bounds write to rp->index[]. Allocate space for every device on hci_dev_list. Since list membership cannot change while hci_dev_list_lock is held, the response remains large enough regardless of flag transitions. The reported count and response length still include only eligible unconfigured controllers. Fixes: 73d1df2a7a10 ("Bluetooth: Add support for Read Unconfigured Index List command") Cc: stable@vger.kernel.org Signed-off-by: Aldo Ariel Panzardo Signed-off-by: Luiz Augusto von Dentz Signed-off-by: Greg Kroah-Hartman commit c6a288e6001bcce16553c3853e0f5a45eaaf85c7 Author: Aldo Ariel Panzardo Date: Tue Sep 15 13:02:39 2026 -0300 Bluetooth: L2CAP: validate frame length before control and FCS access commit 6c78a213d9070b610c7f418af2c25b66180b7e37 upstream. l2cap_data_rcv() unpacks either a two-byte or four-byte control field without first ensuring that it is present. A short ERTM or streaming-mode frame can therefore cause an out-of-bounds read. There is a second short-frame case when CRC16 is enabled. After the control field is pulled, l2cap_check_fcs() subtracts two from skb->len without checking it. If fewer than two bytes remain, the subtraction wraps; skb_trim() leaves the buffer unchanged and the subsequent FCS load reads past the logical end of the frame. Validate that the frame contains both its control field and, when enabled, its FCS before either field is accessed. Fixes: 1c2acffb76d4 ("Bluetooth: Add initial support for ERTM packets transfers") Fixes: fcc203c30d72 ("Bluetooth: Add support for FCS option to L2CAP") Cc: stable@vger.kernel.org Signed-off-by: Aldo Ariel Panzardo Signed-off-by: Luiz Augusto von Dentz Signed-off-by: Greg Kroah-Hartman commit ed4c026cc4094dfb7a329b9d8df0ddd0be1b2458 Author: Aldo Ariel Panzardo Date: Tue Sep 15 13:04:30 2026 -0300 Bluetooth: ISO: release unused CIS holds after channel attach commit 0fcd4dad555c96e0bd3a1b8c569f989be85c7341 upstream. hci_bind_cis() and hci_connect_cis() return one hci_conn hold for the ISO layer. A new channel association consumes that hold, which is eventually released by iso_conn_free(). There are two cases where iso_chan_add() does not create an association: it returns success when the socket is already attached to the same iso_conn, and it returns -EBUSY when another socket is attached. The hold returned for the current call is unused in both cases. This occurs when deferred setup calls iso_connect_cis() again for its existing socket, or when another socket attempts to reuse the CIS. Detect the idempotent case while the connection is locked and release the unused hold after iso_chan_add(). Also release it on -EBUSY. Do not drop it for other errors: a newly allocated iso_conn releases the transferred hold when its last temporary reference is put. Fixes: 69997d50ec57 ("Bluetooth: ISO: handle bound CIS cleanup via hci_conn") Cc: stable@vger.kernel.org Signed-off-by: Aldo Ariel Panzardo Signed-off-by: Luiz Augusto von Dentz Signed-off-by: Greg Kroah-Hartman commit 4e49206d9574a4c109480198e73aea34ad0ed7f6 Author: Aldo Ariel Panzardo Date: Tue Sep 15 13:03:32 2026 -0300 Bluetooth: ISO: balance the parent hold in hci_bind_bis() commit 4c94557dd02569efa6c1072a0439addaef9a5224 upstream. hci_conn_link() takes a lifetime reference to its parent with hci_conn_get(), but only takes an operational hold on the child. hci_conn_unlink() later balances both a hold and a reference on the parent. The SCO and CIS paths pass a parent acquired from a connect helper, so it already has a hold. For an additional BIS, hci_bind_bis() obtains the parent from hci_conn_hash_lookup_big(), which returns a bare pointer. Unlinking the child then drops the parent's existing hold and can schedule it for disconnection while its socket is still using it. Take a hold on the parent before linking it and drop that hold if linking fails. A successful link transfers the hold to hci_conn_unlink(). Fixes: fa224d0c094a ("Bluetooth: ISO: Reassociate a socket with an active BIS") Cc: stable@vger.kernel.org Signed-off-by: Aldo Ariel Panzardo Signed-off-by: Luiz Augusto von Dentz Signed-off-by: Greg Kroah-Hartman commit 173457fa3e99ef4188dd0fe63e4dee98d0e441a9 Author: Aldo Ariel Panzardo Date: Tue Sep 15 13:03:58 2026 -0300 Bluetooth: hci_sock: reject out-of-range OCF values commit e93fad891c72deb84cae49430163b384ebcc92b1 upstream. The raw HCI socket security filter has 128 OCF bits per supported OGF, but masks the 10-bit OCF with 127 before looking up the command. An unprivileged socket can therefore submit a reserved OCF that aliases an allowlisted command modulo 128. A conforming controller should reject reserved opcodes. Nevertheless, the security decision must apply to the opcode that will actually be sent, especially since controller-specific behavior is outside the host stack's control. Reject OCF values that cannot be represented by the security filter instead of aliasing them onto an unrelated command. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Cc: stable@vger.kernel.org Signed-off-by: Aldo Ariel Panzardo Signed-off-by: Luiz Augusto von Dentz Signed-off-by: Greg Kroah-Hartman commit f894f6292d5122cd62d30600736b4a0ae79f7770 Author: Aldo Ariel Panzardo Date: Tue Sep 15 13:03:07 2026 -0300 Bluetooth: hci_sock: validate event length before filtering commit b0a6cf99afd57a39598b1beca0e86ef5004980de upstream. is_filtered_packet() reads the event code from skb->data[0] without first checking that the skb is nonempty. When an opcode filter is configured, it also reads the command opcode at offsets 3 or 4 without checking that a Command Complete or Command Status event is long enough. hci_send_to_sock() invokes the filter before hci_event_packet() validates the event header. A malformed event supplied by a controller or a vhci device can therefore cause an out-of-bounds read. Keep the unmasked event code for the opcode checks. The masked value is needed for the 64-bit event bitmap, but using it to identify command events aliases event codes above 0x3f. In particular, Synchronous Train Complete (0x4f) was treated as Command Status (0x0f) even though its payload has no opcode. Reject actual command events that are too short for the field being inspected. A truncated command event cannot match a configured opcode. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Cc: stable@vger.kernel.org Signed-off-by: Aldo Ariel Panzardo Signed-off-by: Luiz Augusto von Dentz Signed-off-by: Greg Kroah-Hartman commit 3ff5a8d949e3a2a435d0a5ee05074cd371b0158a Author: Aldo Ariel Panzardo Date: Tue Sep 15 13:04:29 2026 -0300 Bluetooth: hci_conn: fix CIS hold ownership on reuse commit e06d549fcd4a0ba381ed67ddf1ab3c7a6ca4314c upstream. Commit 69997d50ec57 ("Bluetooth: ISO: handle bound CIS cleanup via hci_conn") made hci_bind_cis() and hci_connect_cis() return a connection with one hold for the ISO layer. hci_bind_cis() currently takes that hold only after configuring a CIS, so its BT_CONNECTED and matching BT_BOUND paths return a bare lookup result. Its configuration failure path can likewise call hci_conn_drop() before taking a hold. Take the hold before any state-dependent return or configuration error so every successful return follows the documented ownership contract and every error drop is balanced. hci_connect_cis() also assumes hci_conn_link() always takes a new CIS hold before dropping the one returned by hci_bind_cis(). However, the helper returns an existing link without taking another hold. In that case, preserve the CIS hold for the caller and drop the redundant LE hold because the existing link already owns its parent hold. Returning early also avoids changing an existing CIS back to BT_CONNECT. Fixes: 69997d50ec57 ("Bluetooth: ISO: handle bound CIS cleanup via hci_conn") Cc: stable@vger.kernel.org Signed-off-by: Aldo Ariel Panzardo Signed-off-by: Luiz Augusto von Dentz Signed-off-by: Greg Kroah-Hartman commit 57ccaa818307b73220bcdc7626e67aaa92bb483e Author: Tan Chi Date: Mon Sep 14 11:11:46 2026 +0800 RISC-V: KVM: Fix HSM hart status error propagation commit 41e81f7e3ef96594fb840445343c0ee7723aa550 upstream. kvm_sbi_hsm_vcpu_get_status() returns SBI_ERR_INVALID_PARAM when the requested hart does not exist. However, the HART_STATUS case returns from the SBI handler without storing this error in retdata->err_val. As a result, a guest querying the status of a non-existent hart observes SBI_SUCCESS instead of SBI_ERR_INVALID_PARAM. Use the common SBI error handling path for HART_STATUS after saving a valid hart state in retdata->out_val. This preserves the returned error when kvm_sbi_hsm_vcpu_get_status() fails. Fixes: bae0dfd74e01 ("RISC-V: KVM: Modify SBI extension handler to return SBI error code") Cc: stable@vger.kernel.org Assisted-by: LLM Signed-off-by: Tan Chi Reviewed-by: Anup Patel Link: https://lore.kernel.org/r/20260914031146.446157-1-tanchi25@mails.ucas.ac.cn Signed-off-by: Anup Patel Signed-off-by: Greg Kroah-Hartman commit 5567ce11304dbff27ebbd3a2a9a82c43cab0bd02 Author: Xie Bo Date: Mon Aug 10 13:15:44 2026 +0800 RISC-V: KVM: Propagate interrupted G-stage faults commit f41fb17143df890855d3de980f8eb92dfd595817 upstream. __kvm_faultin_pfn() reports an interrupted host page fault with KVM_PFN_ERR_SIGPENDING. RISC-V currently handles it as a generic error PFN and returns -EFAULT. Return -EINTR for the signal-pending sentinel so callers can distinguish an interrupted fault from an invalid userspace mapping. Do not log the expected interruption as a vCPU exit error. Fixes: 9d05c1fee837 ("RISC-V: KVM: Implement stage2 page table programming") Cc: stable@vger.kernel.org Signed-off-by: Xie Bo Reviewed-by: Anup Patel Link: https://lore.kernel.org/r/20260810051544.3953925-3-xb@ultrarisc.com Signed-off-by: Anup Patel Signed-off-by: Greg Kroah-Hartman commit ea91f05a55b807ead3166546a2c75f4185ca645f Author: Myeonghun Pak Date: Sat Aug 1 01:35:50 2026 +0900 RISC-V: KVM: Synchronize hrtimer callback during teardown commit aaad136d56d91252517272b68cd533e5714698d5 upstream. The non-Sstc hrtimer callback clears next_set before its final uses of the enclosing vCPU. If teardown observes next_set as false while the callback is still running, kvm_riscv_vcpu_timer_cancel() skips hrtimer_cancel() and kvm_destroy_vcpus() can free the vCPU before the callback enters kvm_riscv_vcpu_set_interrupt(). A guest can arm the timer with SBI TIME and request shutdown with SBI legacy shutdown or SRST. A VMM that honors KVM_EXIT_SYSTEM_EVENT and destroys the VM supplies the teardown side of the race; no post-launch host ioctl is needed to arm or request teardown. On upstream master 62cc90241548, generic KASAN reported: BUG: KASAN: slab-use-after-free in do_raw_spin_lock Write of size 4 at addr ff60000005e58898 kvm_riscv_vcpu_set_interrupt kvm_riscv_vcpu_hrtimer_expired __hrtimer_run_queues hrtimer_interrupt The object was allocated by KVM_CREATE_VCPU and freed concurrently by: kvm_destroy_vcpus kvm_arch_destroy_vm kvm_destroy_vm __fput For deterministic validation, I added mdelay(1000) immediately after the existing next_set = false assignment. This only widens the existing post-clear callback window. A no-delay trace build naturally reached the callback-after-teardown-start/before-deinit ordering in 12 of 200 runs, but 1,500 stock-kernel stress iterations did not produce a KASAN report, so natural reproduction is timing-sensitive. Always invoke hrtimer_cancel() for an initialized timer. Preserve the existing -EINVAL result when the timer is no longer set, but only after synchronizing with a running callback. With this patch, hrtimer_cancel() blocked for the full widened callback window before vCPU destruction. KASAN reported no error in 100 fixed-and-widened runs or 200 fix-only timing-sweep runs. Fixes: 3a9f66cb25e1 ("RISC-V: KVM: Add timer functionality") Cc: stable@vger.kernel.org Assisted-by: OpenAI:GPT-5.6 Signed-off-by: Myeonghun Pak Reviewed-by: Anup Patel Link: https://lore.kernel.org/r/20260731163550.46991-1-mhun512@gmail.com Signed-off-by: Anup Patel Signed-off-by: Greg Kroah-Hartman commit e5dd2a91a7ae3a276e6d79806127e923936e241b Author: Fuad Tabba Date: Tue Sep 8 12:07:10 2026 +0100 KVM: arm64: Transfer the hyp stack pages out of the host stage-2 commit 3a8c562892b96f35bba1e00d5e455a15963bbb92 upstream. fix_host_ownership() walks only the linear-map alias of each memblock region, and the per-CPU hyp stack, mapped in the private VA range for its guard page, has none. Walk each stack's VA range with the same walker. Fixes: 1a919b17ef012 ("KVM: arm64: Add guard pages for pKVM (protected nVHE) hypervisor stack") Reported-by: Hiroyuki Katsura Cc: stable@vger.kernel.org Signed-off-by: Fuad Tabba Reviewed-by: Vincent Donnefort Tested-by: Vincent Donnefort Reviewed-by: Marc Zyngier Link: https://patch.msgid.link/20260908110713.1540304-2-fuad.tabba@linux.dev Signed-off-by: Oliver Upton Signed-off-by: Greg Kroah-Hartman commit 970de111c087488e53cadf5480b94e0f695db985 Author: Lorenzo Stoakes (ARM) Date: Tue Sep 1 18:29:00 2026 +0100 KVM: arm64: nv: Fix null ptr deref on nested wp/unmap, teardown race commit 4c74e233cdedd11592775fae2a6243e67ca3f891 upstream. Commit 7270cc9157f4 ("KVM: arm64: nv: Handle VNCR_EL2 invalidation from MMU notifiers") introduced VNCR_EL2 invalidation in both kvm_nested_s2_unmap() and kvm_nested_s2_wp(). However at the point of this being performed concurrent stage 2 teardown of a nested guest can cause kvm->arch.mmu.pgt to be set to NULL. This happens in kvm_flush_shadow_all() -> kvm_arch_flush_shadow_all() -> kvm_free_stage2_pgd() and is performed under the kvm->mmu_lock. Commit ec14c272408a ("KVM: arm64: nv: Unmap/flush shadow stage 2 page tables") introduced the teardown of the entire nested MMU range, which then invokes stage2_apply_range() with resched=true: mmu_notifier_invalidate_range_start() -> ... -> kvm_mmu_notifier_invalidate_range_start() -> kvm_mmu_unmap_gfn_range() -> kvm_unmap_gfn_range() -> kvm_nested_s2_unmap() -> kvm_stage2_unmap_range() -> __unmap_stage2_range() -> stage2_apply_range() This means that stage2_apply_range() can drop the kvm->mmu_lock and thus concurrent progress can be made in lockstep with kvm_arch_flush_shadow_all(). If kvm_arch_flush_shadow_all() advances ahead of stage2_apply_range() and completes its operation it guarantees a NULL pointer deref. Since kvm_free_stage2_pgd() is performed under the kvm->mmu_lock this will either be observed NULL or not and serialised against kvm_free_stage2_pgd(). Resolve the issue by abstracting the invalidation to a new function, kvm_invalidate_vncr_ipa_all(), and check that the pgt is non-NULL before dereferencing it. Fixes: 7270cc9157f4 ("KVM: arm64: nv: Handle VNCR_EL2 invalidation from MMU notifiers") Cc: stable@vger.kernel.org Reviewed-by: Marc Zyngier Signed-off-by: Lorenzo Stoakes (ARM) Tested-by: Jonathan Davies Link: https://patch.msgid.link/20260901-kvm-arm-nested-virt-fix-v3-2-b154676f7e4c@kernel.org Signed-off-by: Oliver Upton Signed-off-by: Greg Kroah-Hartman commit c23ed110d27ba99cf3a881986514f79aacedf476 Author: Lorenzo Stoakes (ARM) Date: Tue Sep 1 18:28:59 2026 +0100 KVM: arm64: Fix spurious warning for benign stage 2 teardown race commit 38b70fc453c3112f1a62583b89903ae41116cc27 upstream. kvmtool was used to establish an L1 guest with 8 CPUs and 8 GiB of RAM, an L2 guest with 4 CPUs and 4 GiB of RAM and an L3 guest with 2 CPUs and 2 GiB of RAM, all of which was then exited. Under memory pressure in the L0 host warnings were observed due to migration triggered by compaction: WARNING: arch/arm64/kvm/mmu.c:336 at __unmap_stage2_range+0x64/0x80, CPU#5: kcompactd0/66 Which was, in turn, triggered by an MMU notifier for the host invalidation: mmu_notifier_invalidate_range_start() -> ... -> kvm_mmu_notifier_invalidate_range_start() -> kvm_mmu_unmap_gfn_range() -> kvm_unmap_gfn_range() -> kvm_nested_s2_unmap() -> kvm_stage2_unmap_range() -> __unmap_stage2_range() -> stage2_apply_range() <- -EINVAL, triggering a WARN_ON() Racing with L0's teardown of stage 2 page tables: exit_mm() -> mmput() -> __mmput() -> exit_mmap() -> mmu_notifier_release() -> ... -> kvm_mmu_notifier_release() -> kvm_flush_shadow_all() -> kvm_arch_flush_shadow_all() -> kvm_free_stage2_pgd() -> [ acquire kvm->mmu_lock for write ] -> mmu->pgt = NULL [ among other tasks ] -> [ release kvm->mmu_lock for write ] It turns out there is a benign race resulting in a spurious warning: Thread A - notify: migration | Thread B - notify: release -------------------------------|--------------------------------- < kvm->mmu_lock held > | stage2_apply_range() | get mmu->pgt, check !NULL | ... | kvm_arch_flush_shadow_all() cond_resched_rwlock_write(); | < contend, sleep kvm->mmu_lock > < drop kvm->mmu_lock > | < acquire kvm->mmu_lock> | ... | kvm_free_stage2_pgd() | mmu->pgt = NULL | < invalidate MMU > | ... | < release kvm->mmu_lock > [ scheduled ] | stage2_apply_range() | < loop to next > | get, mmu->pgt, check !NULL | is NULL, return -EINVAL | __unmap_stage2_range() | WARN_ON(-EINVAL) <--- entirely spurious - the race was handled correctly. Fix the spurious warning by updating stage2_apply_range() to no longer treat concurrent PGT teardown on lock release as an error - whether the walker is tearing down page tables or doing something else this is a legitimate reason to abort the operation without error. This keeps the warning in place for all other circumstances. In practice only __unmap_stage2_range() actually does anything with the error so this only impacts that. Fixes: ec14c272408a ("KVM: arm64: nv: Unmap/flush shadow stage 2 page tables") Cc: stable@vger.kernel.org Reviewed-by: Yuan Yao Reviewed-by: Marc Zyngier Signed-off-by: Lorenzo Stoakes (ARM) Link: https://patch.msgid.link/20260901-kvm-arm-nested-virt-fix-v3-1-b154676f7e4c@kernel.org Signed-off-by: Oliver Upton Signed-off-by: Greg Kroah-Hartman commit 8530a05b304d312fe7dfbc5df7d3e87cc31b256a Author: Karl Mehltretter Date: Mon Aug 10 02:56:16 2026 +0200 KVM: arm64: Fix AArch32 DBGBXVR handling commit 6b1bca1b1ab77f60a62087337bfe6e2f0efb9e6d upstream. The consolidation of the breakpoint and watchpoint register accessors switched DBGBXVR from trap_bvr() to trap_dbg_wb_reg(). The latter selects backing storage with demux_wb_reg(), which only handles Op2 values 4 through 7. Since DBGBXVR uses Op2 1, an AArch32 guest access hits KVM_BUG_ON() and marks the VM dead. DBGBXVR aliases DBGBVR_EL1[63:32], and its AA32(HI) descriptor already selects the upper half. Map Op2 1 to dbg_bvr[] alongside Op2 4, restoring the pre-regression behavior. Fixes: 3ce9f3357e9e ("KVM: arm64: Fold DBGxVR/DBGxCR accessors into common set") Cc: stable@vger.kernel.org Assisted-by: Claude:claude-opus-5 Signed-off-by: Karl Mehltretter Reviewed-by: Marc Zyngier Link: https://patch.msgid.link/20260810005616.13227-1-kmehltretter@gmail.com Signed-off-by: Oliver Upton Signed-off-by: Greg Kroah-Hartman commit 37dfb676fafd1eef36fe0bb649607afbb0d5521b Author: Sean Christopherson Date: Wed Sep 23 09:37:21 2026 -0700 KVM: SEV: Do cache maintenance on the source VM during intra-host migration commit 93de2a6a4b91b72607136dd656edf03fb399d27f upstream. Manually perform cache maintenance on the source VM during intra-host migration to ensure no stale data is left in CPU caches after the VM is destroyed. Because the source VM is "converted" to a non-SEV VM, KVM's memory reclaim flows won't trigger cache maintenance, e.g. when all guest memory is reclaimed in response to detaching from the mmu_notifier. Note, relying on the destination VM to do cache maintenance isn't an option as KVM doesn't require identical guest memory configurations, i.e. the source VM may have access to memory that the destination VM does not. Enforcing equivalent memory configurations is infeasible, as it would require a *deep* comparison of memslots, e.g. to verify that not only are the memslot identical, but what the memslots point at is also identical. Fixes: b56639318bb2 ("KVM: SEV: Add support for SEV intra host migration") Cc: stable@vger.kernel.org Reported-by: Stefan Teodorescu Signed-off-by: Sean Christopherson Message-ID: <20260923163721.1584779-3-seanjc@google.com> Signed-off-by: Paolo Bonzini Signed-off-by: Greg Kroah-Hartman commit 6aef7cf800098749ff5d660222efaf6e01e98c09 Author: Sean Christopherson Date: Wed Sep 23 09:37:20 2026 -0700 KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV commit 12c1f6e03f944e399bd2c88441dca5dc702b95a5 upstream. Unconditionally free SEV's "have run CPUs" cpumask in the VM destroy path, i.e. even for what appear to be non-SEV VMs, as an SEV VM becomes a non-SEV VM if its state is intra-host migrated. Alternatively, the mask could be freed in sev_migrate_from() when "converting" the source VM, but that gets annoying because ideally KVM would nullify the mask to guard against UAF, and nullifying the mask would need be conditioned on CPUMASK_OFFSTACK=y. Freeing the mask during sev_migrate_from() is also not robust against other KVM bugs, though that's kind of a moot point since any such bugs would show up even if sev->active is never set. I.e. KVM must get that side of things correct. But, that's not a great reason to add more code just to make things marginally less robust. Fixes: 6f38f8c57464 ("KVM: SVM: Flush cache only on CPUs running SEV guest") Cc: stable@vger.kernel.org Reported-by: Stefan Teodorescu Signed-off-by: Sean Christopherson Message-ID: <20260923163721.1584779-2-seanjc@google.com> Signed-off-by: Paolo Bonzini Signed-off-by: Greg Kroah-Hartman commit 066a53a0bc19ee7f022ed02c9b2c73ef49a0d3b4 Author: Zeng Chi Date: Mon Sep 21 18:24:42 2026 +0800 KVM: Don't treat reserved xarray entries as having memory attributes commit 277d3623d99a4fc2623bfb7d191b649ca380605d upstream. kvm_vm_set_mem_attributes() reserves an xarray entry for every gfn in the range before storing the new attributes, so that the store loop can't fail partway through. If one of the reservations fails, e.g. with -ENOMEM, the entries that were already reserved are left in the array. That is harmless as far as xa_reserve() is concerned, as the reserved entries read back as NULL via xa_load(), but it confuses the "does this range have no attributes at all" check: if (!attrs) return !xas_find(&xas, end - 1); A reserved entry is XA_ZERO_ENTRY, not NULL, and xas_find() returns it as present. So a leftover reservation makes KVM report that a fully shared range has attributes even though kvm_get_memory_attributes() returns none for every gfn in the range. On x86, the next time mixed-attribute tracking is recomputed for the range (memslot creation, or a later attribute change that straddles the 2MiB page), hugepage_has_attrs() treats a fully shared 2MiB range as mixed and refuses to map it with a hugepage, until userspace happens to set attributes on the range again. Drop the shortcut and handle the !attrs case in the per-index loop, using xas_next_entry() to find the next non-NULL entry. xas_next_entry() is essentially an optimized xas_find(), so the effective change is that the !attrs lookup now goes through xas_retry() like the attrs != 0 case, i.e. reserved entries are skipped and retry entries restart the walk. Don't check the index when no entry is found, as the xarray leaves the xas index in a bogus state in that case; no entry simply means the rest of the range has no attributes. KVM never stores a non-NULL entry with a value of zero (clearing stores NULL), but such an entry would be returned by xas_next_entry() and trip the index check, so WARN if one is ever seen. Fixes: 5a475554db1e ("KVM: Introduce per-page memory attributes") Cc: stable@vger.kernel.org Suggested-by: Sean Christopherson Cc: David Ballesteros Signed-off-by: Zeng Chi Link: https://patch.msgid.link/20260921102442.1232375-1-zeng_chi911@163.com [sean: expand comment to elaborate on xarray APIs, split optimization out] Signed-off-by: Sean Christopherson Signed-off-by: Greg Kroah-Hartman commit 7dcaf0107523fc2ca3918d65c8ade06600f44ee7 Author: David Ballesteros Date: Tue Sep 15 17:53:57 2026 +0000 KVM: Ensure memory attributes xarray nodes are accounted to the caller's memcg commit 382e5d514b6f35bdda2ab9044b4eed23d2ec4254 upstream. Explicitly instantiate the memory attributes xarray with XA_FLAGS_ACCOUNT to ensure that all allocations are accounted to the memcg. Frustratingly, memory allocations done in the "fastpath" do not honor the passed in gfp, even for an explicit xa_reserve(). Only the rare, slow path __xas_nomem() honors the original gfp. E.g. xa_reserve(..., GFP_KERNEL_ACCOUNT) | -> ... | -> __xa_cmpxchg_raw() | -> xas_store() <== does not take @gfp | -> xas_create() | -> xas_alloc() The bug was confirmed by observing that a process in a cgroup limited to 256 MiB grew radix_tree_node slab by ~512 MiB while its memory.current stayed near 0. Fixes: 5a475554db1e ("KVM: Introduce per-page memory attributes") Cc: stable@vger.kernel.org Assisted-by: Claude-Code:claude-opus-5 Signed-off-by: David Ballesteros Link: https://patch.msgid.link/20260915175335.138547-4-davimaba.v@proton.me [sean: rewrite changelog, tag for stable] Signed-off-by: Sean Christopherson Signed-off-by: Greg Kroah-Hartman commit cf441a0db681dfad5e89b015d6be11365159d7a9 Author: Anthony Krowiak Date: Tue Aug 18 15:33:49 2026 -0400 s390/vfio-ap: fix KVM GISC and page leak when queue removed from host config commit 65e05ec252a9b79e75930d3c4dd42d8877db04c5 upstream. Three related problems exist in the handling of KVM interrupt and page resources when a queue is removed from the host's AP configuration while assigned to a mediated device (mdev). Problem 1: ~~~~~~~~~ AP_RESPONSE_Q_NOT_AVAIL not handled in vfio_ap_mdev_reset_queue() When the AP bus removes a queue device whose adapter or domain has been removed from the host's AP configuration, vfio_ap_mdev_remove_queue() is called. If the queue is still in the host's AP configuration at that point, it calls vfio_ap_mdev_reset_queue(), which issues a PQAP(ZAPQ). Since the adapter is already gone from the host configuration, ap_zapq() returns AP_RESPONSE_Q_NOT_AVAIL (0x01). This response code is not handled in vfio_ap_mdev_reset_queue()'s switch statement and falls through to the default case, which issues a WARN but does not call vfio_ap_free_aqic_resources(). As a result, if IRQ handling was enabled for the queue by the guest, the KVM GISC registration and the pinned guest page holding the notification indicator byte (NIB) are both leaked. This is fixed by adding AP_RESPONSE_Q_NOT_AVAIL to the same case as AP_RESPONSE_DECONFIGURED and AP_RESPONSE_CHECKSTOPPED in vfio_ap_mdev_reset_queue(). Like those response codes, Q_NOT_AVAIL indicates the queue is not operational and no further reset attempts are possible; the correct action is to free the IRQ resources immediately. Problem 2: ~~~~~~~~~ AP_RESPONSE_Q_NOT_AVAIL not handled in apq_status_check() In vfio_ap_mdev_reset_queue(), there are four cases that indicate a queue reset has not yet completed, in which case apq_reset_check() is queued to a work queue to verify completion of the reset operation. This function uses the PQAP(TAPQ) function to get the queue's status and calls apq_status_check() to verify whether the reset has completed, failed or needs to be executed again. As described in Problem #1 above, apq_reset_check() does not specifically check for AP_RESPONSE_Q_NOT_AVAIL, thereby potentially leaking KVM GISC registration and the pinned guest page holding the NIB. This is fixed by adding a case statement for AP_RESPONSE_Q_NOT_AVAIL to apq_status_check() and returning -ENODEV for that case. The caller, apq_reset_check() will then check for this return code and call vfio_ap_free_aqic_resources() to prevent the leak. Problem 3: ~~~~~~~~~ vfio_ap_free_aqic_resources() leaks saved_isc when kvm is NULL vfio_ap_free_aqic_resources() guards the call to kvm_s390_gisc_unregister() with: if (q->saved_isc != VFIO_AP_ISC_INVALID && !WARN_ON(!(q->matrix_mdev && q->matrix_mdev->kvm))) If matrix_mdev->kvm is NULL -- which can happen when vfio_ap_mdev_unset_kvm() has already run and cleared kvm before a subsequent cleanup path reaches this function -- the WARN_ON fires and the entire block is skipped. This leaves q->saved_isc set to a non-invalid value, creating a potential double-free on any subsequent call to this function. When kvm is NULL the KVM guest is already torn down, so kvm_s390_gisc_unregister() need not and cannot be called; however, q->saved_isc must always be cleared. Fix this by separating the kvm_s390_gisc_unregister() call from the q->saved_isc reset. The WARN_ON now guards only the genuinely impossible case of matrix_mdev being NULL. A NULL kvm is handled gracefully by skipping only the unregister call, and q->saved_isc = VFIO_AP_ISC_INVALID is set unconditionally whenever saved_isc was not already invalid. Additionally, add an else clause to the host-config check in vfio_ap_mdev_remove_queue() to call vfio_ap_free_aqic_resources() directly when the queue is not in the host's AP configuration. This serves as a backstop: when the AP bus fires the driver .remove callback after an adapter is removed from the host config, the queue is by definition no longer addressable, so vfio_ap_mdev_reset_queue() would always return Q_NOT_AVAIL. The else clause handles this case directly without the unnecessary ap_zapq() call, and ensures cleanup occurs even if kvm has already been set to NULL by a prior call to vfio_ap_mdev_unset_kvm(). Fixes: b9bd10c43456d ("s390/vfio-ap: do not reset queue removed from host config") Cc: stable@vger.kernel.org Signed-off-by: Anthony Krowiak Reviewed-by: Matthew Rosato Acked-by: Halil Pasic Signed-off-by: Claudio Imbrenda Message-ID: <20260818193349.1877940-2-akrowiak@linux.ibm.com> Signed-off-by: Greg Kroah-Hartman commit af295f0eab68370cb5527a1c3710de06a7204cf0 Author: Niklas Schnelle Date: Wed Sep 16 17:14:13 2026 +0200 s390/pci: Don't report recovery success on skipped recovery commit 3a43be7a1fd06a35cf9e621b88283b3b6e7d281c upstream. When a PCI device is already in the permanent failure state, recovery is skipped, but the SCLP recovery report still shows success. Fix this by changing the status string to explicitly state that recovery was skipped due to permanent failure. Cc: stable@vger.kernel.org Fixes: 4ec6054e7321 ("s390/pci: Report PCI error recovery results via SCLP") Signed-off-by: Niklas Schnelle Reviewed-by: Benjamin Block Reviewed-by: Farhan Ali Signed-off-by: Heiko Carstens Signed-off-by: Greg Kroah-Hartman commit eef5c8e1cca6fc218ceab4a75c5844af4c1cb577 Author: Niklas Schnelle Date: Wed Sep 16 17:14:12 2026 +0200 s390/pci: Report SCLP status on error events when no pdev is associated commit a1120bea9bc8ea9d9ab2f9904a63b9a228bf2ffc upstream. With commit 4ec6054e7321 ("s390/pci: Report PCI error recovery results via SCLP") SCLP reports are generated when recovery is performed in response to an error event. If such an error event arrives but no pdev is currently associated with the zdev, e.g. because it was removed or not yet probed, no report is generated. Fix this by generating a report specific to an error event for a zdev without an associated pdev. Cc: stable@vger.kernel.org Fixes: 4ec6054e7321 ("s390/pci: Report PCI error recovery results via SCLP") Signed-off-by: Niklas Schnelle Reviewed-by: Benjamin Block Reviewed-by: Farhan Ali Signed-off-by: Heiko Carstens Signed-off-by: Greg Kroah-Hartman commit 743a8458978d13d84a1cf74bcd29b9959a75d3a0 Author: Niklas Schnelle Date: Wed Sep 16 17:14:11 2026 +0200 s390/pci: Fix missing device lock in zpci_report_status() commit 0261aef4b15efcee2860ab857e5cb05e9bfa47b0 upstream. When pdev is non-NULL, zpci_report_status() accesses the device's driver. To get a consistent state matching the recovery, the device lock needs to be held. Do so by expanding the existing device lock critical section. The lock only needs to be held when the pdev is non-NULL, so extract the pdev-specific reporting into a helper function which also adds a lockdep assertion to detect calls without the device lock held. Cc: stable@vger.kernel.org Fixes: 4ec6054e7321 ("s390/pci: Report PCI error recovery results via SCLP") Signed-off-by: Niklas Schnelle Reviewed-by: Benjamin Block Reviewed-by: Farhan Ali Signed-off-by: Heiko Carstens Signed-off-by: Greg Kroah-Hartman commit dd5ed64dadb07691aa7a012b5c9be0c0dda2abe9 Author: Niklas Schnelle Date: Wed Sep 16 17:14:10 2026 +0200 s390/pci: Fix leak of struct pci_dev reference in zpci_report_status() commit 09b7040a1b79f4f61cdad6d9af972e045bb7c498 upstream. In zpci_report_status(), a reference to the pdev associated with the zdev being reported about is acquired using pci_get_slot(). This reference needs to be dropped with pci_dev_put(), but this call is missing, thus leaking the reference. On subsequent hot unplug, this will cause the struct pci_dev to not be released, leaking memory and preventing reattach. At the same time, the only existing caller already holds a pdev reference. So instead of reacquiring and then dropping another reference, simply pass the existing pdev pointer to zpci_report_status(). This gets rid of the need for pci_get_slot() as well as the zdev->zbus check. Cc: stable@vger.kernel.org Fixes: 4ec6054e7321 ("s390/pci: Report PCI error recovery results via SCLP") Signed-off-by: Niklas Schnelle Reviewed-by: Benjamin Block Reviewed-by: Farhan Ali Signed-off-by: Heiko Carstens Signed-off-by: Greg Kroah-Hartman commit 748ec2f7f5df9b4fc742546a8bfc8aa5d847353e Author: Peter Oberparleiter Date: Mon Sep 21 15:16:48 2026 +0200 s390/cmf: Fix virtual vs physical address confusion commit f4d04425e66af2ecee9d1a49ae0484436f3c2fd1 upstream. The measurement block address is an absolute address. Define the associated schib_config and schib fields as dma64_t to enable automatic detection of incorrect assignments. Also add the missing virt_to_dma64() translation. Without this fix, a wrong address will be used by firmware when storing extended format channel measurement data on kernels built with CONFIG_RANDOMIZE_IDENTITY_BASE=y. Fixes: 14edd0d73bfe ("s390/cmf: fix virtual vs physical address confusion") Cc: stable@vger.kernel.org Signed-off-by: Peter Oberparleiter Reviewed-by: Heiko Carstens Signed-off-by: Heiko Carstens Signed-off-by: Greg Kroah-Hartman commit c3bbcb32deecc06ebb1762398399b566d5ff2fc9 Author: Vineeth Vijayan Date: Tue Sep 22 22:48:39 2026 +0200 s390/cio: Fix NULL pointer dereference in ccw_device_get_util_str() commit 5b76268dac968612f7283d59b539036de955b7d9 upstream. The channel path registry entry associated with a CHPID may be removed while the subchannel's PMCW still references that CHPID. In this case, chpid_to_chp() can return NULL, leading to a NULL pointer dereference. Add the missing NULL check before dereferencing the returned pointer. Fixes: 199652309a4d ("s390/cio: add helper to query utility strings per given ccw device") Cc: stable@vger.kernel.org Signed-off-by: Vineeth Vijayan Reviewed-by: Peter Oberparleiter Signed-off-by: Heiko Carstens Signed-off-by: Greg Kroah-Hartman commit 02664a8d8d5929e941db2c2c09cec316321652cc Author: Ilya Titov Date: Thu Sep 3 12:17:18 2026 +0300 pinctrl: sunxi: keep a shadow copy of the data register output latches commit a13f7f5d14af9baf34eb12c25962f8d3542b281d upstream. On Allwinner SoCs, reading a bank's data register returns the pin level, not the output latch, for pins that are muxed as inputs. Writing a GPIO therefore corrupts the output latches of all input-muxed pins in the same bank: the read-modify-write in sunxi_pinctrl_gpio_set() reads back their pin levels and writes those into their latches. This breaks emulated open-drain lines (e.g. a bit-banged I2C bus from i2c-gpio). Such a line is released high by muxing it as input and letting the pull-up raise it, so any concurrent GPIO write in the same bank stores 1 into its latch. Driving the line low afterwards is a non-atomic data-then-mux sequence in sunxi_pinctrl_gpio_direction_output(); if the poisoning write lands between the two steps, the pin actively drives high (push-pull) instead of low. Observed in practice as sporadic glitches on a T507 board bit-banging I2C on port E while other PE GPIOs are toggled. On a scope the failure is unmistakable: on a clock pulse where SCL should fall to GND, the line instead steps *above* its idle high level for the whole low phase — the pad drives a strong push-pull 3.3 V high, higher than the level the pull-up sustains on the loaded bus — before the next transition recovers it. The same can hit SDA, corrupting data instead of clocks. Steps to reproduce on any sunxi board with a bit-banged (i2c-gpio) bus: # background: toggle any other GPIO of the same bank, e.g. line 21 gpioset -c --toggle 100us 21=0 & # foreground: keep the bit-banged bus busy while :; do i2cdetect -y 0x50 0x57; done # watch SCL/SDA with a scope or logic analyzer: sporadic clock-low # phases driven high (above the pull-up level) instead of low The bank spinlock cannot help: the racing write is a perfectly valid whole-register RMW that faithfully writes back what the hardware returned. There are no set/clear registers on this IP to write a single bit atomically. Fix it the same way gpio-mmio handles hardware whose data register read does not return the output latch: keep a shadow copy of each bank's latches, base the read-modify-write on the shadow, and only write the register. The shadow is seeded from the hardware at probe time so pins left in output mode by the bootloader keep their state. Pins that reach output mode through the gpiolib paths write their value (and thereby their shadow bit) before the mux switch in sunxi_pinctrl_gpio_direction_output(); pins muxed to gpio_out directly through a pinmux node bypass that path, so sunxi_pmx_set() refreshes their shadow bit from the latch (readable once the pin is in output mode) to keep them driving their pre-existing level. Seeding the shadow reads the PIO registers at probe time, which requires the bus clock to be enabled. The clock was only requested at the very end of probe, after devm_pinctrl_register() had already claimed the pin hogs described in the device tree - which mux pins, and thus access registers, with the clock still gated. Move the request ahead of both. Boards whose bootloader leaves the PIO clock running are unaffected, which is why the pre-existing hog problem has gone unnoticed since commit 950707c0eb5c ("pinctrl: sunxi: add clock support"). Fixes: df7b34f4c3d2 ("pinctrl: sunxi: Fix gpio_set behaviour") Cc: stable@vger.kernel.org Signed-off-by: Ilya Titov Signed-off-by: Linus Walleij Signed-off-by: Greg Kroah-Hartman commit 1deeeca89ea1207457266fa6071394738aa5c50a Author: Myeonghun Pak Date: Sun Sep 13 00:03:12 2026 -0400 pinctrl: single: free the IRQ on domain creation failure commit 1d9bb9c870632adea249cfdec49ac37e6f069164 upstream. pcs_irq_init_chained_handler() requests a shared IRQ on affected SoCs, but its domain creation failure path only removes a chained handler. That does not release the action installed by request_irq(). The probe can continue without interrupt support while leaving the shared IRQ action registered. Use pcs_irq_free() to undo the appropriate type of handler registration. At this point pcs->domain is NULL, so the helper only releases the parent IRQ handler. Then mark the IRQ invalid, as the other initialization error paths already do, to prevent another release from a later probe unwind or remove. This issue was identified during our ongoing static-analysis research while reviewing kernel code. Fixes: 3e6cee1786a1 ("pinctrl: single: Add support for wake-up interrupts") Cc: stable@vger.kernel.org Assisted-by: LLM Co-developed-by: Ijae Kim Signed-off-by: Ijae Kim Signed-off-by: Myeonghun Pak Signed-off-by: Linus Walleij Signed-off-by: Greg Kroah-Hartman commit a23f9f9b56b795cb3dbb55357f171c502108cba2 Author: Puranjay Mohan Date: Mon Aug 10 06:35:35 2026 -0700 perf/core: Run sched_task() for PMUs with only CPU-wide events commit 3d8d74100954a3b17e5c5e37adfe14e16b1db103 upstream. perf_pmu_sched_task() returns early when cpuctx->task_ctx is set and leaves the work to perf_ctx_sched_task_cb(), which only walks ctx->pmu_ctx_list. A PMU whose events are all CPU-wide is not on that list, so nothing calls its sched_task(). With perf record -b -e cycles -a -- ls armv8pmu_sched_task() is skipped on every switch to a task that has a perf context but no event on that PMU, and BRBE records leak across the task boundary. intel_pmu_lbr_add() calls perf_sched_cb_inc() unconditionally too, so LBR records leak the same way on x86. Drop the early return and skip only the CPCs that perf_ctx_sched_task_cb() handles. That one needs a gate of its own to make the split exact: it tests cpc->sched_cb_usage, which perf_sched_cb_inc() sets per CPU for every branch stack user, so a task with an event for that PMU pinned to another CPU would be handled twice. On x86 the second __intel_pmu_lbr_restore() finds lbr_stack_state == LBR_NONE and calls intel_pmu_lbr_reset(), throwing away the callstack the first one restored. cpc->task_epc is set only while a task context is scheduled in, and there is one epc per PMU on ctx->pmu_ctx_list, so the two gates are inverses. For the CPCs perf_pmu_sched_task() picks up, the callback now runs outside the perf_ctx_disable() and perf_ctx_enable() pair in perf_event_context_sched_in(). __perf_pmu_sched_task() disables the PMU around the call itself. Fixes: bd2756811766 ("perf: Rewrite core context handling") Signed-off-by: Puranjay Mohan Signed-off-by: Peter Zijlstra (Intel) Tested-by: Yifan Wu Link: https://patch.msgid.link/20260810133540.1947118-3-puranjay@kernel.org Cc: stable@vger.kernel.org Signed-off-by: Greg Kroah-Hartman commit 249a2bd87a0e57dc28f36e992807edc9cae309b5 Author: Puranjay Mohan Date: Mon Aug 10 06:35:34 2026 -0700 perf/core: Fix NULL pmu_ctx passed to pmu->sched_task() commit 36bb85cf36cab15fb611cb44b78a5df06e4e69a2 upstream. perf_pmu_sched_task() returns early when cpuctx->task_ctx is set, and cpc->task_epc is only non-NULL while a task context is scheduled in on this CPU. __perf_pmu_sched_task() therefore always passes NULL: Unable to handle kernel NULL pointer dereference at virtual address 00 pc : armv8pmu_sched_task+0x14/0x50 Call trace: armv8pmu_sched_task+0x14/0x50 (P) perf_pmu_sched_task+0xac/0x108 __perf_event_task_sched_out+0x6c/0xe0 Pass &cpc->epc instead, the CPU-wide context for this PMU, which the function already dereferences a few lines up to find pmu. armv8pmu_sched_task() is the only in-tree implementation that dereferences the argument, and it only reads ->pmu, so the oops needs BRBE, added in v6.17. Fixes: bd2756811766 ("perf: Rewrite core context handling") Signed-off-by: Puranjay Mohan Signed-off-by: Peter Zijlstra (Intel) Tested-by: Yifan Wu Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260810133540.1947118-2-puranjay@kernel.org Signed-off-by: Greg Kroah-Hartman commit 5023c70853ad9829c53d8e9def5ee969b4e1b931 Author: Helge Deller Date: Sat Sep 19 22:09:20 2026 +0200 parisc: Increase kernel stack size to 32kb commit 94b7e3a7e871ae27d4935c76959dfc61829f27f9 upstream. For 64-bit Linux kernels, increase the default kernel stack size (THREAD_SIZE_ORDER) to 32 kB, in order to avoid kernel crashes which have been triggered recently when building the debian vtk9 package with gcc 17: stackcheck: kworker/u128:0 will most likely overflow kernel stack (sp:179a83af0, stk bottom-top:179a80000-179a84000) Kernel panic - not syncing: low stack detected by irq handler - check messages CPU: 2 UID: 0 PID: 30760 Comm: kworker/u128:0 Tainted: G W 6.18.46-dirty #1 NONE Tainted: [W]=WARN Hardware name: 9000/800/rp3440 Workqueue: writeback wb_workfn (flush-259:0) Backtrace: [<000000004022f050>] show_stack+0x70/0x90 [<000000004022378c>] dump_stack_lvl+0x124/0x190 [<000000004022382c>] dump_stack+0x34/0x48 [<000000004020212c>] vpanic+0x204/0x648 [<00000000402025c4>] panic+0x54/0x58 [<0000000040232230>] do_cpu_irq_mask+0x3f8/0x440 [<0000000040227070>] intr_return+0x0/0xc Signed-off-by: Helge Deller Reported-by: John David Anglin Cc: stable@vger.kernel.org # v6.18+ Signed-off-by: Greg Kroah-Hartman commit ee56b5e18f566c0b5a183379dcd63d402cebf2f8 Author: Yehyeong Lee Date: Sat Aug 1 22:36:35 2026 +0900 scsi: libiscsi_tcp: Check the data direction of a Data-In PDU commit bce07e2f37b5e4a427d36fd6b1c14067b27591db upstream. The Data-In branch of iscsi_tcp_hdr_dissect() resolves the ITT to a task and copies the PDU's data segment into that command's scatterlist without asking whether the command was reading. iscsi_tcp_r2t_rsp() in the same file does ask, and rejects an R2T for a command that is not DMA_TO_DEVICE. A target that answers a WRITE command's ITT with a Data-In therefore has the initiator write target-supplied bytes into the pages that write was about to send. Those are the caller's own pinned pages for an O_DIRECT write, and page cache pages for a buffered one. Observed against a test target that emits one 512-byte Data-In naming a 128 KB write's ITT, after the R2T for that write. With O_DIRECT the caller's buffer ends up holding 512 bytes of the target's data while pwrite() returns 131072. Buffered is quieter: pwrite() and fsync() both succeed, nothing is logged, and reading those blocks back returns the target's bytes out of the page cache without a command going on the wire. Check the direction before using the scatterlist, the way the R2T path already does. Cc: stable@vger.kernel.org Signed-off-by: Yehyeong Lee Reviewed-by: Mike Christie Link: https://patch.msgid.link/20260801133635.1986706-1-yhlee@isslab.korea.ac.kr Fixes: a081c13e39b5 ("[SCSI] iscsi_tcp: split module into lib and lld") Signed-off-by: Martin K. Petersen (Oracle) Signed-off-by: Greg Kroah-Hartman commit cc444e4ad43461374dd128573eccb8979d2adb4c Author: Myeonghun Pak Date: Sun Sep 13 00:26:25 2026 -0400 nfc: trf7970a: power down on startup RX gain failure commit d2acbde7e67df44efa8f0963462d1192e7694ffc upstream. trf7970a_startup() powers up the device before applying the optional RX gain reduction. If the register read or write fails, it returns without undoing that power-up. Probe's unwind only drops the separate regulator references acquired by probe, leaving the additional VIN enable from startup unbalanced. The system resume caller also has no power-down on this error. Call trf7970a_power_down() before returning the RX gain error to deassert the enable GPIOs, release the startup VIN reference and restore the powered-off state. Runtime PM has not been enabled yet, so the full shutdown helper is not appropriate here. Preserve the original SPI error. This issue was identified during our ongoing static-analysis research while reviewing kernel code. Fixes: 5d69351820ea ("NFC: trf7970a: Create device-tree parameter for RX gain reduction") Cc: stable@vger.kernel.org Assisted-by: OpenAI:GPT-5.6 Co-developed-by: Ijae Kim Signed-off-by: Ijae Kim Signed-off-by: Myeonghun Pak Reviewed-by: Paul Geurts Link: https://patch.msgid.link/20260913042625.31296-1-mhun512@gmail.com Signed-off-by: David Heidelberg Signed-off-by: Greg Kroah-Hartman commit 14b1467cfc6630583650f33538174c5a1ec685ab Author: Doruk Tan Ozturk Date: Sat Jul 11 14:36:51 2026 +0200 nfc: port100: reject frames whose declared length exceeds the received data commit 092c6a605cbd6414ef499834c2e0da69c2c3388e upstream. port100_recv_response() passes the URB transfer buffer to port100_rx_frame_is_valid(), which checksums le16_to_cpu(frame->datalen) bytes of frame->data. datalen is a 16-bit field supplied by the device and is never checked against the number of bytes actually received (urb->actual_length), so a device reporting a datalen larger than the received frame makes port100_data_checksum() read out of bounds past the transfer buffer. Reject a response whose declared frame size does not fit the received length before validating it. Found by 0sec (https://0sec.ai) using automated source analysis; the missing bound is evident from source. Compile-tested. Fixes: 562d4d59b8a1 ("NFC: Sony Port-100 Series driver") Cc: stable@vger.kernel.org Assisted-by: 0sec:claude-opus-4-8 Signed-off-by: Doruk Tan Ozturk Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260711123651.32595-1-doruk@0sec.ai Signed-off-by: David Heidelberg Signed-off-by: Greg Kroah-Hartman commit 810f204ef8f61f0c930cf6b3280050b3ee6a8ae7 Author: Aamir Ahmed Date: Tue Sep 15 19:54:27 2026 +0100 nfc: llcp: drop truncated I/RR/RNR PDUs in nfc_llcp_recv_hdlc() commit 273f9d667cde649f8de9d72b1303cc2f4b658c50 upstream. nfc_llcp_recv_hdlc() reads the sequence byte skb->data[2], via nfc_llcp_ns()/nfc_llcp_nr(), before any length check. The receive path only guarantees the two-byte LLCP header -- __nfc_llcp_recv() checks it with pskb_may_pull() and nfc_llcp_recv_agf() admits two-byte inner PDUs -- so a two-byte I, RR or RNR PDU reads one byte of uninitialised skb tailroom. The byte becomes N(R)/N(S); a peer can already set those with a well-formed PDU, so this is acting on uninitialised memory, not new peer control. Guard the read with pskb_may_pull(), as commit 95674f506c63 ("nfc: llcp: reject PDUs shorter than the LLCP header") did for the two-byte header, so the sequence byte is present and linear before it is read. RR and RNR PDUs are LLCP_HEADER_SIZE + LLCP_SEQUENCE_SIZE bytes and an I PDU is longer, so no valid frame is rejected; a truncated PDU is malformed, so return without a DM reply. Fixes: d646960f7986 ("NFC: Initial LLCP support") Cc: stable@vger.kernel.org Assisted-by: LLM Signed-off-by: Aamir Ahmed Reviewed-by: Simon Horman Link: https://patch.msgid.link/AS8P251MB0001789BBF04B72745C7D96BC8BA2@AS8P251MB0001.EURP251.PROD.OUTLOOK.COM Signed-off-by: David Heidelberg Signed-off-by: Greg Kroah-Hartman commit 841af260ff477f6191872bc3b2e15488615c1827 Author: Luxiao Xu Date: Wed Sep 9 13:19:24 2026 +0800 nfc: fix use-after-free in nfc_get_local_general_bytes commit dcab71a7011918f6fdba7adcec02d217dcb84b8d upstream. Commit 6709d4b7bc2e ("net: nfc: Fix use-after-free caused by nfc_llcp_find_local") attempted to fix a use-after-free (UAF) issue by invoking nfc_llcp_local_put(local) after accessing local->gb. However, if the reference count drops to zero, local is freed immediately, leading to a use-after-free when callers access the returned pointer. Alternative approaches using dynamic allocation (e.g. kmemdup) introduced memory leaks because callers consistently treat the returned pointer as borrowed memory. Fix this properly by refactoring nfc_llcp_general_bytes() and nfc_get_local_general_bytes() to accept a caller-provided output buffer (out_gb) and its maximum length (gb_max_len). The general bytes are safely copied into out_gb before calling nfc_llcp_local_put(local), ensuring safe lifetime management without ownership transfer complications. Update all callers across drivers (microread, pn533, pn544, st21nfca, digital_dep, and nci) to provide their own destination buffers and pass them to nfc_get_local_general_bytes(). Fixes: 6709d4b7bc2e ("net: nfc: Fix use-after-free caused by nfc_llcp_find_local") Cc: stable@vger.kernel.org Reported-by: Vega Assisted-by: LLM Signed-off-by: Luxiao Xu Signed-off-by: Ren Wei Reviewed-by: Simon Horman Link: https://patch.msgid.link/3cbaac3bee23f8ff3a3284ed32d347696eb1d208.1788841683.git.rakukuip@gmail.com Signed-off-by: David Heidelberg Signed-off-by: Greg Kroah-Hartman commit 12cf803edfffef96f6415f0b5dbc7ec88842ea0f Author: Aohan Mei Date: Mon Sep 14 19:51:47 2026 +0800 netfilter: nf_tables: skip expired catchall elements on insert and delete commit 70194dc37670bd08e44b471389861cc01bd3a3c9 upstream. nft_setelem_catchall_insert() looks up duplicates with nft_set_elem_active() only, while nft_set_catchall_lookup() and the dump path additionally skip expired elements. Once a catchall element with a timeout expires, this predicate drift makes it invisible to userspace dumps, yet it still blocks re-insertion: with NLM_F_EXCL the request fails with -EEXIST, and without it the request reports success but silently inserts nothing. The stale entry only goes away when the (user-tunable) gc interval elapses, so the catchall rule may silently stop matching for an arbitrarily long time after its first expiration. The delete path shows the same drift: nft_setelem_catchall_deactivate() picks the first active-next entry in the catchall list, so with an expired entry still pending GC it retires the stale entry instead of the fresh one, and it deactivates an element that userspace no longer sees instead of failing with -ENOENT. Align both walks with the lookup and dump predicates: only an element that is active and not expired counts as a duplicate or delete candidate, using the per-netns timestamp taken at transaction start, in line with the set backend .insert/.deactivate and catchall GC sync paths. Reported-by: TencentOS Corvus AI Cc: stable@vger.kernel.org Fixes: aaa31047a6d2 ("netfilter: nftables: add catch-all set element support") Assisted-by: CodeBuddy:Kimi-K3 Signed-off-by: Aohan Mei Signed-off-by: Pablo Neira Ayuso Signed-off-by: Greg Kroah-Hartman commit 191239b0939692e92459d8d32cad14247fa93f89 Author: Luxiao Xu Date: Sun Sep 6 21:29:55 2026 +0800 netfilter: ip6t_rt: fix zero-address non-strict match out-of-bounds read commit 82313c169eddc02b1bf5ba6b427803e272d3ec42 upstream. rt_mt6_check() permits rules to be configured with rtinfo->addrnr == 0 even when address matching (IP6T_RT_FST_MASK) is requested. In the IP6T_RT_FST_NSTRICT path, rt_mt6() evaluates packet routing addresses against rtinfo->addrs[i] and terminates backwards at the bottom of the loop: if (ipv6_addr_equal(ap, &rtinfo->addrs[i])) { i++; } if (i == rtinfo->addrnr) break; When addrnr is 0, if the first packet address matches rtinfo->addrs[0], i is incremented to 1. Because i is now strictly greater than addrnr (0), the loop termination condition (i == rtinfo->addrnr) is bypassed and will never be satisfied. If a crafted IPv6 packet contains matching routing addresses, i will advance past IP6T_RT_HOPS (16). The subsequent call to ipv6_addr_equal() reads beyond struct ip6t_rt, triggering UBSAN/KASAN out-of-bounds warnings or kernel panics. Fix this by: 1. Rejecting rules in rt_mt6_check() where IP6T_RT_FST_MASK is set but rtinfo->addrnr is zero. 2. In rt_mt6(), moving the termination condition (i < rtinfo->addrnr) into the for-loop header condition and removing the backwards break at the end of the loop body. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Cc: stable@vger.kernel.org Reported-by: Vega Suggested-by: Florian Westphal Assisted-by: LLM Signed-off-by: Luxiao Xu Signed-off-by: Ren Wei Signed-off-by: Pablo Neira Ayuso Signed-off-by: Greg Kroah-Hartman commit eafe081ea77d25b5c8e601680671b4ecb4520e7d Author: Weiming Shi Date: Sun Sep 6 16:44:10 2026 +0800 netfilter: ip6t_rpfilter: reject routes without inet6_dev commit 1b9b5323725e458906c7620a3bc10398b51ad954 upstream. ip6_route_lookup() can return an error-free route whose rt6i_idev is NULL. Lowering an external nexthop device's MTU below IPV6_MIN_MTU tears down its inet6_dev while fib6_ifdown() leaves routes using nexthop objects in the FIB. An unprivileged user can construct this state with rtnetlink in a private user and network namespace, then trigger a NULL dereference through an IPv6 rpfilter lookup: Oops: general protection fault, probably for non-canonical address 0xdffffc0000000000 KASAN: null-ptr-deref in range [0x0000000000000000-0x0000000000000007] RIP: rpfilter_mt (net/ipv6/netfilter/ip6t_rpfilter.c:75) Call Trace: ip6t_do_table (net/ipv6/netfilter/ip6_tables.c:316) nf_hook_slow (net/netfilter/core.c:619) ipv6_rcv (net/ipv6/ip6_input.c:351) __netif_receive_skb_one_core (net/core/dev.c:6216) process_backlog (net/core/dev.c:6680) __napi_poll (net/core/dev.c:7739) net_rx_action (net/core/dev.c:7959) handle_softirqs (kernel/softirq.c:622) do_softirq.part.0 (kernel/softirq.c:523) __local_bh_enable_ip (kernel/softirq.c:450) __dev_queue_xmit (net/core/dev.c:4913) packet_sendmsg (net/packet/af_packet.c:3139) __sys_sendto (net/socket.c:2252) __x64_sys_sendto (net/socket.c:2259) do_syscall_64 (arch/x86/entry/syscall_64.c:94) entry_SYSCALL_64_after_hwframe (arch/x86/entry/entry_64.S:121) Kernel panic - not syncing: Fatal exception in interrupt Reject routes without an inet6_dev immediately after lookup. Such routes are not eligible for reverse-path filtering, and the check protects all later rt6i_idev dereferences. Fixes: e26f9a480fb6 ("netfilter: add ipv6 reverse path filter match") Reported-by: co+459f67f4d8af8ce6@bugs.sh Closes: https://lore.kernel.org/all/VtWUkE8QzJt5CroTj2V2v3ZQ0gwbXZ7nq7I3@bugs.sh/ Suggested-by: Florian Westphal Assisted-by: Claude:gpt-5 Cc: stable@vger.kernel.org Signed-off-by: Weiming Shi Signed-off-by: Pablo Neira Ayuso Signed-off-by: Greg Kroah-Hartman commit 8a611f33216f43c86287860d60be4a3d4517ca87 Author: Wentao Liang Date: Thu Sep 17 11:58:11 2026 +0000 net: usb: lan78xx: Fix URB reference leak in lan78xx_submit_deferred_urbs() commit 17741334d00bf5ebd37f8c1c36bc9c146a351deb upstream. usb_get_from_anchor() hands over a reference to the URB, which the caller must release. lan78xx_submit_deferred_urbs() never does, so every deferred Tx URB keeps an extra reference: the counter grows on each suspend/resume cycle and the URBs are never freed when the buffers are released. Drop the reference after submitting, and on the path that drops the packet instead of submitting it. Fixes: 5f4cc6e25148 ("lan78xx: Fix race conditions in suspend/resume handling") Cc: stable@vger.kernel.org Signed-off-by: Wentao Liang Link: https://patch.msgid.link/20260917115811.2150119-1-vulab@iscas.ac.cn Signed-off-by: Paolo Abeni Signed-off-by: Greg Kroah-Hartman commit 7d2377868feb7471683cb16586415501d2f4c338 Author: Ming Wang Date: Sun Sep 20 15:44:59 2026 +0800 net: usb: cdc_mbim: add MeiG Smart SRM821 to ZLP whitelist commit f75f21ef36285e5f56ee0c428bd2909ee81165b9 upstream. The MeiG Smart SRM821 5G module (0x2dee:0x4d53) crashes and drops off the USB bus when it receives a Zero Length Packet (ZLP) after sending or receiving an NTB of exactly 16384 bytes (tx_max). According to the MBIM specification, devices do not require a ZLP if the NTB size is exactly dwNtbOutMaxSize. However, the cdc_mbim driver defaults to sending ZLPs for devices not explicitly whitelisted to accommodate non-conformant hardware. This default behavior breaks the strictly conformant MeiG SRM821 module. Add this device to the ZLP conformance whitelist (cdc_mbim_info) so the driver will pad the NTB to avoid sending ZLPs, preventing the device firmware from crashing. Cc: stable@vger.kernel.org Signed-off-by: Ming Wang Link: https://patch.msgid.link/20260920074500.826121-1-wangming01@loongson.cn Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit 43ff11501ea6a07fd4d76acbbbe9d39ff1e7d8c1 Author: Zhang Yunfei Date: Fri Sep 11 17:11:23 2026 +0800 net: txgbe: fix FDIR filter restore for VF rules commit 651010592bdce7005c1179498327e51bfc4fe1a5 upstream. txgbe_fdir_filter_restore() reprograms every filter from txgbe->fdir_filter_list after a reset. It extracts the ring part of filter->action with ethtool_get_flow_spec_ring() and maps it onto a PF rx ring, silently dropping the VF part of the cookie that txgbe_add_ethtool_fdir_entry() stores there (input->action = fsp->ring_cookie). For a rule directed at a VF, restore therefore reprograms the filter to the PF queue with the same ring index: after any down/up or txgbe_reinit_locked(), traffic matching the rule is steered to the PF instead of the VF. Handle VF rules the same way txgbe_add_ethtool_fdir_entry() does: validate vf against wx->num_vfs and ring against wx->num_rx_queues_per_pool, and map the ring onto the absolute queue index ((vf - 1) * wx->num_rx_queues_per_pool) + ring. Fixes: 7a91722e0dd4 ("net: txgbe: Support the FDIR rules assigned to VFs") Cc: stable@vger.kernel.org Signed-off-by: Zhang Yunfei Link: https://patch.msgid.link/20260911091123.798931-1-zhangyunfei1@kylinos.cn Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit 9447ddd8b9242c4e8e46b2407db8008935fd2f73 Author: Abhishek Ojha Date: Wed Sep 16 19:19:28 2026 -0400 net: phy: micrel: Advance register data pointer in write loop commit 95c4d54ed02283e9a09e8cd7360e384daa67a741 upstream. lanphy_write_reg_data() does not advance the data pointer while iterating over the register table. As a result, it writes the first entry num times and leaves the remaining errata registers unconfigured. Single-entry tables are unaffected, but tables with multiple entries leave every entry after the first unapplied. Advance the data pointer after each successful write so every table entry is applied in order. Fixes: c8732e933925 ("net: phy: micrel: lan8842 errata") Cc: stable@vger.kernel.org Signed-off-by: Abhishek Ojha Reviewed-by: Andrew Lunn Link: https://patch.msgid.link/20260916231928.1336305-1-abhishek.ojha@savoirfairelinux.com Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit 30ff01b4c7e419bb70f40f01dcd7f6af680808b3 Author: Guangshuo Li Date: Mon Sep 21 23:42:02 2026 +0800 net: ena: fix MMIO read buffer leak on probe failure commit 9476b4468862927297c94c440863cd8ed1e7cc83 upstream. ena_device_init() initializes the MMIO read mechanism with ena_com_mmio_reg_read_request_init(), which allocates a coherent DMA buffer for MMIO read responses. The normal removal path releases this buffer through ena_com_mmio_reg_read_request_destroy(). However, if ena_probe() fails after ena_device_init() succeeds, the error path destroys the admin resources and eventually frees ena_dev without destroying the MMIO read request, leaving the coherent DMA buffer allocated. Call ena_com_mmio_reg_read_request_destroy() in the probe error path before releasing the remaining device resources. This issue was found by manual code inspection. Fixes: 1738cd3ed342 ("net: ena: Add a driver for Amazon Elastic Network Adapters (ENA)") Cc: stable@vger.kernel.org Signed-off-by: Guangshuo Li Link: https://patch.msgid.link/20260921154202.471662-3-lgs201920130244@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit e8a2f48bbff1f85fe0466c7f446e83585efd59b8 Author: Guangshuo Li Date: Mon Sep 21 23:42:01 2026 +0800 net: ena: fix PHC cleanup on probe failure commit 0958ea4355e2e9220ad4e13da3b7d94f365ed34f upstream. ena_probe() initializes the PHC as part of ena_device_init(), but the probe failure path does not destroy it before freeing the PHC private data. The normal removal path calls ena_phc_destroy() through ena_destroy_device() before ena_phc_free(). However, if probe fails after ena_device_init() succeeds, the error path reaches ena_phc_free() without unregistering the PTP clock or destroying the device PHC resources. Call ena_phc_destroy() in the probe error path before freeing the PHC private data. This issue was found by manual code inspection. Cc: stable@vger.kernel.org tags and describe this as a consistency cleanup Fixes: e0ea34158ee8 ("net: ena: Add PHC support in the ENA driver") Cc: stable@vger.kernel.org Signed-off-by: Guangshuo Li Cc: stable Link: https://patch.msgid.link/20260921154202.471662-2-lgs201920130244@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit ab3c4a4c85f2e1e8d04b5c1770dea8412d21c7fe Author: Fourie Zhang Date: Sun Sep 20 19:08:43 2026 +0800 net: bridge: mdb: restart port group walk after deletion commit ab1404ac81154a89fb61ac50ae9a04cd8d4834dc upstream. br_mdb_flush_pgs() keeps a pointer-to-pointer cursor while walking mp->ports. br_multicast_del_pg() can re-enter the same MDB entry through br_multicast_sg_del_exclude_ports() and unlink other port groups. If the cursor points into one of those groups, the next iteration dereferences a stale cursor and can leave mp->ports pointing at freed memory. A following RTM_GETMDB exposes the dangling pointer: BUG: KASAN: slab-use-after-free in br_mdb_dump Read of size 8 br_mdb_dump rtnl_mdb_dump rtnl_dumpit netlink_dump Reset the cursor to mp->ports after every deletion. The deletion removes at least the selected group, so the restarted walk always makes progress. Fixes: a6acb535afb2 ("bridge: mdb: Add MDB bulk deletion support") Cc: stable@vger.kernel.org Signed-off-by: Fourie Zhang Acked-by: Nikolay Aleksandrov Link: https://patch.msgid.link/20260920110852.60293-1-fouriezhang@tencent.com Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit 4b7161804d8f31227551901f93947bb4eea4690b Author: Gajdos Tamás Date: Mon Sep 21 11:13:33 2026 +0200 net: atl1e: fix soft lockup on out-of-range hw_next_to_clean read commit 374bf9e4b90f979e052332c4faca2d745c491a12 upstream. Same issue as atl1c (see the first commit in this series, "net: atl1c: fix soft lockup on out-of-range tpd_cons read"): the hardware can report an out-of-range hw_next_to_clean (seen as 0xffff) while the PCIe link/MAC is resetting. An out-of-range value can never be reached and the loop below would spin forever. Treat it as "nothing new to clean" instead. Fixes: a6a5325239c202 ("atl1e: Atheros L1E Gigabit Ethernet driver") Cc: stable@vger.kernel.org Signed-off-by: Gajdos Tamás Link: https://patch.msgid.link/20260921091334.3571525-3-tamas@rimpianto.com Signed-off-by: Paolo Abeni Signed-off-by: Greg Kroah-Hartman commit 586847400b1168a88917287ac02c2c91a620b841 Author: Gajdos Tamás Date: Mon Sep 21 11:13:32 2026 +0200 net: atl1c: fix soft lockup on out-of-range tpd_cons read commit 36c2009d90f2210ef92e6f4f2850e8b57b09e754 upstream. The hardware can report an out-of-range tpd_cons (seen as 0xffff) while the PCIe link/MAC is resetting. An out-of-range value can never be reached and the loop below would spin forever. To avoid a soft lockup treat it as "nothing new to clean" instead. Reproduced on two machines, same NIC (Qualcomm Atheros AR8151 v2.0, 4-port), triggered by rebooting a Mikrotik CCR2004 PCIe card that the ports are directly linked to: - Ubuntu 26.04.1 LTS, kernel 7.0.0-31-generic. The link-flap precursor, before the lockup was captured with a full trace elsewhere: atl1c 0000:05:00.0 enp5s0f0: NETDEV WATCHDOG: CPU: 4: transmit queue 2 timed out 489984 ms atl1c 0000:05:00.0: MAC state machine can't be idle since disabled for 10ms second atl1c 0000:05:00.0: atl1c: enp5s0f0 NIC Link is Up<65535 Mbps Full Duplex> 65535 (0xffff) here is the same value tpd_cons reads back once the loop below gets stuck. - Proxmox VE, kernel 7.0.14-11-pve. Same NIC/trigger, this time caught by the soft lockup watchdog with a full stack trace: watchdog: BUG: soft lockup - CPU#12 stuck for 354s! [napi/eth%d-0:329] CPU: 12 UID: 0 PID: 329 Comm: napi/eth%d-0 Tainted: P O L 7.0.14-11-pve #1 PREEMPT(lazy) RIP: 0010:atl1c_clean_tx+0x142/0x2d0 [atl1c] Call Trace: __napi_poll+0x32/0x1e0 napi_threaded_poll_loop+0x286/0x2e0 napi_threaded_poll+0xfd/0x140 kthread+0xf7/0x130 ret_from_fork+0x2da/0x3a0 ret_from_fork_asm+0x1a/0x30 Fixes: 43250ddd75a35d ("atl1c: Atheros L1C Gigabit Ethernet driver") Cc: stable@vger.kernel.org Signed-off-by: Gajdos Tamás Link: https://patch.msgid.link/20260921091334.3571525-2-tamas@rimpianto.com Signed-off-by: Paolo Abeni Signed-off-by: Greg Kroah-Hartman commit 5f23080a625dd7855ab91485c8d691a54a1979f0 Author: Ilya Maximets Date: Mon Sep 21 16:55:45 2026 +0200 net: openvswitch: conntrack: fix helper UAF due to extensions realloc commit 1a4151e6be57b098b7a5ebfbde58585e83200cdc upstream. While calling the helpers, a raw pointer to the extensions area is wired into expectations list: -> nf_ct_helper() -> helper->help() -> nf_ct_expect_related_report() -> nf_ct_expect_insert() -> hlist_add_head_rcu(&exp->lnode, &master_help->expectations) In case the connection is not confirmed yet, more extensions can be added afterwards with *_ext_add() calls reallocating the extension space and leaving the now invalid pointer in the expectations list that is later accessed while removing the expectation. Make sure that helpers are called at the end after all the other extensions are already added. Note that the helper rejection now leaves the mark and labels set, but that's not different from how the NAT was handled before or how the mark and the labels were handled on confirmation failure. And there are no atomicity guarantees provided by the API anyway. Fixes: cae3a2627520 ("openvswitch: Allow attaching helpers to ct action") Cc: stable@vger.kernel.org Reported-by: Axel Mierczuk Signed-off-by: Ilya Maximets Reviewed-by: Aaron Conole Link: https://patch.msgid.link/20260921145655.3167436-4-i.maximets@ovn.org Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit caefcd433ecb29aaf2aa0be5416c042bcca777e0 Author: Ilya Maximets Date: Mon Sep 21 16:55:44 2026 +0200 net: openvswitch: conntrack: remove 'add_helper' dead code commit 5e6c14dd42a1c1fe938e573dc6c9098145b2b0c4 upstream. This variable can only become 'true' when the connection is not confirmed, but it is only checked when it is confirmed. So, it can be treated as being always false and just removed. Fixes: 3c1860543fcc ("openvswitch: add nf_ct_is_confirmed check before assigning the helper") Cc: stable@vger.kernel.org Signed-off-by: Ilya Maximets Reviewed-by: Aaron Conole Link: https://patch.msgid.link/20260921145655.3167436-3-i.maximets@ovn.org Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit e2829ae5459ddb5fce526b2e84697edf6c0c3cef Author: Ilya Maximets Date: Mon Sep 21 16:55:43 2026 +0200 net: openvswitch: conntrack: avoid modifying shared unconfirmed ct entry commit 26b2bd70d22457556e2fa01cbf1192cb1a94d619 upstream. In a case where skb with an unconfirmed ct entry gets cloned, we may end up committing both but with different sets of extensions. The series of events: 1. The first clone wants to commit and runs the helpers wiring up the extension pointer into the expectation list. 2. Then it looses the confirmation keeping the entry unconfirmed. 3. Second clone now wants to commit labels and adds the new extension for that breaking the pointer in the expectation list causing UAF on the destruction path later. While this is possible to trigger, there should be no practical network pipeline where committing both clones without modifications into the same zone is needed. So, let's just reset the entry in case for some reason we got an skb with a shared one during commit. This doesn't affect any known use cases, but avoids any potential problems with sharing and modification of the unconfirmed ct entry. The fixes tag points to the introduction of helpers, since that's the main UAF trigger for the sharing. Fixes: cae3a2627520 ("openvswitch: Allow attaching helpers to ct action") Cc: stable@vger.kernel.org Reported-by: Axel Mierczuk Signed-off-by: Ilya Maximets Reviewed-by: Aaron Conole Link: https://patch.msgid.link/20260921145655.3167436-2-i.maximets@ovn.org Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit 729737267e7c1d6fd4525c12477d3c4d6d87bf2f Author: Wentao Liang Date: Thu Sep 17 11:08:28 2026 +0000 net: hisilicon: hns_dsaf_mac: fix mdio device leak in hns_mac_register_phy() commit 999e8295bc41d6ce45b8e54f88150efa96f3f01e upstream. hns_dsaf_find_platform_device() returns the mdio platform device with its reference count incremented. hns_mac_register_phy() never drops that reference, so the mdio device can not be released. Release the reference on both the deferred probe and the normal path. Fixes: 1d1afa2ebf82 ("net: hns: register phy device in each mac initial sequence") Cc: stable@vger.kernel.org Signed-off-by: Wentao Liang Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260917110828.2148390-1-vulab@iscas.ac.cn Signed-off-by: Paolo Abeni Signed-off-by: Greg Kroah-Hartman commit 0f573b0576252d44fb27ab8470ea873bc6408ef3 Author: Gajdos Tamás Date: Mon Sep 21 11:13:34 2026 +0200 net: atl1: fix soft lockup on out-of-range cmb_tpd_next_to_clean read commit 43e746821f5f5afbbf68e388bf9fbe221e03bfca upstream. Same issue as atl1c (see the first commit in this series, "net: atl1c: fix soft lockup on out-of-range tpd_cons read"): the hardware can report an out-of-range cmb_tpd_next_to_clean (seen as 0xffff) while the PCIe link/MAC is resetting. An out-of-range value can never be reached and the loop below would spin forever. Treat it as "nothing new to clean" instead. Fixes: f3cc28c797604f ("Add Attansic L1 ethernet driver.") Cc: stable@vger.kernel.org Signed-off-by: Gajdos Tamás Link: https://patch.msgid.link/20260921091334.3571525-4-tamas@rimpianto.com Signed-off-by: Paolo Abeni Signed-off-by: Greg Kroah-Hartman commit ba664d83638640cf454b65167339cbf20dfd6f10 Author: Zijie Huang Date: Mon Sep 21 01:36:19 2026 +0800 net: arp: terminate device name before lookup commit d8b6529e80bcb4fb8177121404cbb3377acaebd2 upstream. The ARP ioctl copies a user-provided struct arpreq into a stack object. Its arp_dev field may contain IFNAMSIZ bytes without a NUL terminator. Such input is passed to dev_get_by_name_rcu() or __dev_get_by_name(), where strcmp() can read past the end of the stack object when a matching alternative interface name exists. Terminate the field before the lookup to prevent the out-of-bounds read. Fixes: 36fbf1e52bd3 ("net: rtnetlink: add linkprop commands to add and delete alternative ifnames") Cc: stable@vger.kernel.org Reported-by: Vega Signed-off-by: Zijie Huang Signed-off-by: Ren Wei Reviewed-by: Ido Schimmel Link: https://patch.msgid.link/fabf02a70787d17299e4b3153eadffaf20d154b3.1789910973.git.milkory@outlook.com Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit aca91d8fa0ecbc990ebf9a688c684dda78a43c00 Author: Myeonghun Pak Date: Mon Sep 21 20:09:14 2026 -0400 net: airoha: npu: cancel wdt_work after releasing the WDT IRQ commit 4bdee8060d1e4581624e68fbd369b1afb14df4bc upstream. airoha_npu_remove() calls cancel_work_sync() on each core's wdt_work, but the watchdog IRQ that queues it is requested with devm_request_irq() and is freed only after .remove() returns. airoha_npu_wdt_handler() can therefore schedule_work() again once the cancel has returned. struct airoha_npu, which contains the work, is devm_kzalloc()'d and is freed in that same unwind, so the late work dereferences freed memory. Register the work with devm_work_autocancel() before devm_request_irq() and drop .remove(). Devres runs in reverse order, so the IRQ is freed before cancel_work_sync(), including when probe fails. A cancel left in .remove() cannot get that order. Initializing the work first also stops a pending watchdog interrupt from queuing an uninitialized work item. Probe currently calls INIT_WORK() only after devm_request_irq(). This issue was identified during our ongoing static-analysis research while reviewing kernel code. Fixes: 23290c7bc190 ("net: airoha: Introduce Airoha NPU support") Cc: stable@vger.kernel.org # 6.15+ Co-developed-by: Ijae Kim Signed-off-by: Ijae Kim Signed-off-by: Myeonghun Pak Acked-by: Lorenzo Bianconi Link: https://patch.msgid.link/20260922000914.542068-1-mhun512@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit 968550a439f64bd1a6c0d88efb92e4f082f84a26 Author: Weiming Shi Date: Mon Sep 14 14:51:23 2026 +0800 net/sched: reject IDR error pointers when deleting actions commit c82b797abe668d0b668601a93ba2c0b071a63574 upstream. tcf_action_delete() drops the reference held by its lookup before calling tcf_idr_delete_index() with the saved action index. An unlocked classifier can remove that action and reserve the same IDR slot with ERR_PTR(-EBUSY) in between. tcf_idr_delete_index() only checks the lookup result for NULL. It therefore treats the reservation as a tc_action and dereferences tcfa_bindcnt. A hardware execution breakpoint was used to schedule the interleaving without changing the kernel source. KASAN reported this decoded trace: BUG: KASAN: null-ptr-deref in tca_action_gd+0x5b9/0x1010 Read of size 4 at addr 0000000000000010 by task poc/150 Oops: general protection fault, probably for non-canonical address 0xdffffc0000000002 RIP: tca_action_gd+0x5c0/0x1010: arch_atomic_read at arch/x86/include/asm/atomic.h:23 raw_atomic_read at include/linux/atomic/atomic-arch-fallback.h:457 atomic_read at include/linux/atomic/atomic-instrumented.h:33 tcf_idr_delete_index at net/sched/act_api.c:766 tcf_action_delete at net/sched/act_api.c:1859 tcf_del_notify at net/sched/act_api.c:2014 tca_action_gd at net/sched/act_api.c:2064 R13: 0000000000000010 R15: fffffffffffffff0 Kernel panic - not syncing: Fatal exception R15 contains ERR_PTR(-EBUSY), and adding the tcfa_bindcnt offset produces the address in R13. With the guard applied, the same reproducer returned -ENOENT without a KASAN report or panic. Treat error pointers as absent and return -ENOENT. Fixes: 0190c1d452a9 ("net: sched: atomically check-allocate action") Cc: stable@vger.kernel.org Reported-by: Xiang Mei Signed-off-by: Weiming Shi Link: https://patch.msgid.link/20260914065123.4109709-2-bestswngs@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit 3803dc8477f08e0dac4050cafee0de43f32727be Author: Andrea Parri Date: Thu Sep 17 13:55:42 2026 +0200 net/mlx5e: fix swapped IPv6 IPsec policy masks commit 10de7ed8ef4840da9ca21de4c29578657ac367db upstream. IPv6 XFRM policies may use different source and destination prefix lengths. mlx5e_ipsec_policy_mask() builds the corresponding masks independently, but setup_fte_addr6() installs each mask in the opposite address field. When the prefix lengths differ, this makes the source match use the destination prefix and the destination match use the source prefix. The resulting hardware rule can both miss traffic covered by the policy and match traffic outside it. Install each mask in its corresponding match field. Fixes: ca7992f52c2c ("net/mlx5e: Properly match IPsec subnet addresses") Cc: stable@vger.kernel.org Signed-off-by: Andrea Parri Reviewed-by: Tariq Toukan Link: https://patch.msgid.link/20260917115542.177675-1-parri.andrea@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit 0c336ecabaf9d39d307a72d52740b4bb5f0473df Author: Ralf Lici Date: Thu Sep 17 14:27:23 2026 +0200 net/mlx5e: advertise MACsec offload only when supported commit 4581c3d2adc3c73a019bc38db64ca11f28bbd7fd upstream. Commit 339ccec8d43d ("net/mlx5: Enable MACsec offload feature for VLAN interface") added NETIF_F_HW_MACSEC unconditionally to vlan_features so that VLAN devices could inherit MACsec offload support. mlx5e_build_nic_netdev subsequently copies vlan_features into hw_features and features. As a result, all mlx5e NIC netdevices advertise MACsec hardware offload, even when the firmware does not support it and the driver does not install macsec_ops. Set the MACsec feature bits in mlx5e_macsec_build_netdev, after device capabilities have been validated. This preserves MACsec-over-VLAN support and the ethtool feature control on capable devices, without advertising either on unsupported hardware. Fixes: 339ccec8d43d ("net/mlx5: Enable MACsec offload feature for VLAN interface") Cc: stable@vger.kernel.org Reviewed-by: Tariq Toukan Signed-off-by: Ralf Lici Link: https://patch.msgid.link/20260917122724.654639-1-ralf@mandelbit.com Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit 99d8dc51663747078c031acac32edc6d90c8f1f7 Author: Wentao Liang Date: Thu Sep 17 11:31:31 2026 +0000 net/mlx5: Fix rev_entry reference leak in mlx5_tc_ct_shared_counter_get() commit 0bf6bb567f0edaa771e7dd208ef98da50e6a4485 upstream. When the reverse entry is found but its counter is already being released, refcount_inc_not_zero() fails and the reference taken by mlx5_tc_ct_entry_get() is never dropped before falling through to create_counter. Drop it so the reverse entry is not kept alive forever by a shared counter lookup that did not use it. Fixes: 1edae2335adf ("net/mlx5e: CT: Use the same counter for both directions") Cc: stable@vger.kernel.org Signed-off-by: Wentao Liang Reviewed-by: Tariq Toukan Link: https://patch.msgid.link/20260917113131.2149024-1-vulab@iscas.ac.cn Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit c3129e6f26e18c0886ede7265f187a7ec289e1d1 Author: Wei Jie LAW <98lawweijie@gmail.com> Date: Wed Sep 9 09:15:13 2026 +0800 HID: wacom: fix OOB read in wacom_wac_pen_serial_enforce() commit 9aa237cf66495b2426ddde8532e9b08a0ed83aaa upstream. The 'wacom_wac_pen_serial_enforce()' function may calculate and pass an invalid offset to hid_field_extract(), resulting in memory reads at incorrect addresses -- possibly beyond the end of the report. If a field in the HID descriptor lists more usages than its Report Count actually reserves space for, the function's inner 'j' will walk past the end of the field: for (i = 0; i < report->maxfield; i++) { for (j = 0; j < report->field[i]->maxusage; j++) { ... value = hid_field_extract(hdev, raw_data + 1, offset + j * size, size); A descriptor listing 12288 usages against Report Count 1 has the loop extract the usage at index 12287 from bit offset 98296 -- about 12 KB past a 2-byte received report. The value is stored in wacom_wac->serial[0] and can reach userspace as an MSC_SERIAL event, making this an information disclosure. Clamp the loop to field->report_count, the number of value slots the report holds. Value slots past the last declared usage are still scanned; they reuse that usage (HID 1.11, 6.2.2.8). Verified on v6.12.105 with a UHID reproducer: a 2-byte report from such a descriptor trips KASAN before the patch and not after it. Fixes: 83417206427b ("HID: wacom: Queue events with missing type/serial data for later processing") Suggested-by: Jason Gerecke Cc: stable@vger.kernel.org Assisted-by: Claude:claude-opus-5 Assisted-by: GLM:glm-5.3 Signed-off-by: Wei Jie Law <98lawweijie@gmail.com> Reviewed-by: Jason Gerecke Signed-off-by: Jiri Kosina Signed-off-by: Greg Kroah-Hartman commit 5695ec5339d4dd9a23e34abc4a18b8a3e4b85951 Author: Junjie Cao Date: Mon Aug 24 11:14:19 2026 +0800 HID: quirks: add ALWAYS_POLL quirk for SDINNOVATION gaming keyboard commit cdb669a3b8f844aca71fc3224990157d61562165 upstream. The SDINNOVATION gaming keyboard (USB ID 36ae:feab) stops reporting input events after its RGB lighting mode is switched about twice. Disabling USB autosuspend and unbinding the other HID interfaces make no difference; the issue does not occur on Windows. HID_QUIRK_ALWAYS_POLL alone resolves it, verified on 7.1.8 via usbhid.quirks=0x36ae:0xfeab:0x400. Reported-by: Marco Carvalho Link: https://bugzilla.redhat.com/show_bug.cgi?id=2514627 Cc: stable@vger.kernel.org Signed-off-by: Junjie Cao Signed-off-by: Benjamin Tissoires Signed-off-by: Greg Kroah-Hartman commit 925da1d2a24787f24872de13a4680437c54de179 Author: Chen Changcheng Date: Fri Aug 14 15:06:21 2026 +0800 HID: alps: fix use-after-free on input2 registration failure commit d3aba3442798ce4a4c8ce3104d7b286d61e605f9 upstream. alps_input_configured() stores data->input2 before calling input_register_device(). If registration fails, input_free_device() frees the input device but data->input2 still points to the freed memory. alps_input_configured() calls hid_hw_open() before allocating input2, so URBs are already active and raw_event can fire during the failure window. A U1_SP_ABSOLUTE_REPORT_ID report arriving then causes u1_raw_event() to dereference the freed data->input2 -> use-after-free. Fix by only storing input2 into drvdata after successful registration and adding a NULL guard in the raw_event path. Fixes: 2562756dde55 ("HID: add Alps I2C HID Touchpad-Stick support") Cc: stable@vger.kernel.org Signed-off-by: Chen Changcheng Signed-off-by: Jiri Kosina Signed-off-by: Greg Kroah-Hartman commit eda9bcb721f15a93746a99eabda33d80b86404fd Author: Willem de Bruijn Date: Fri Sep 18 20:47:29 2026 -0400 packet: use ubuf_info completion for TX_RING packets commit 9518405613863d0bf0700927a21367f5942cf058 upstream. tpacket_snd sends skbs with frags pointing into its ring slots. Slots are released when skb->destructor is called. A call to skb_orphan calls skb->destructor before the skb is freed. This can cause the slot to be reused while still linked into the skb. Switch to standard zerocopy completion (ubuf_info) so the slot is only released once all references to the payload are freed or copied. Restore skb->destructor to standard sock_wfree. The ubuf_info completion callback can be called with a NULL skb, but only from net_zcopy_put and related API, used by zerocopy implementations that hold their own reference on the uarg, such as MSG_ZEROCOPY. This uarg is only ever completed from skb_zcopy_clear, so skb is always set. To prevent userspace from aliasing in-flight state on shared ring slots, allocate tpacket_uarg per packet, rather than per slot. This adds a small allocation to the transmit path. Use standard kmalloc to allow backporting to stable kernels. The uarg holds an sk_wmem_alloc reference, rather than an sk_refcnt reference. packet_free_tx_ring waits on sk_wmem_alloc before freeing the ring pages. Always allocate vec->deferred for tx_ring so page-backed rings also wait on sk_wmem_alloc when skb_copy_ubufs drops page refs before calling tpacket_ubuf_complete. Drop the tx_ring.pg_vec test that tpacket_destruct_skb performed before accessing the slot. The sk_wmem_alloc reference now guarantees that the slot is valid. The test is also not sufficient by itself, as it reads pg_vec without pg_vec_lock, so it can race with packet_set_ring. As a result a slot is released when its payload is copied, which can be before transmission (e.g., in skb_orphan_frags_rx). If copied before skb_tx_timestamp() is called, no slot timestamp is recorded, similar to when skb_orphan() was called early in the datapath before this patch. Revert the now unused previous skb_zcopy_.._nouarg infra. Depends on commit 992cc9f94ca9 ("net/packet: defer vmalloc TX_RING free until skbs finish"). Reported-by: Katherine Leaver Reported-by: Bjoern Doebel Closes: https://lore.kernel.org/netdev/20260909085542.3370986-1-doebel@amazon.de/ Fixes: 5cd8d46ea156 ("packet: copy user buffers before orphan or clone") Cc: stable@vger.kernel.org Signed-off-by: Willem de Bruijn Link: https://patch.msgid.link/20260919004748.1463985-3-willemdebruijn.kernel@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit 879c911f77b8a7862aa16079b04efc3fdc5d32ec Author: Angel J Date: Fri Sep 18 14:55:40 2026 -0500 PCI: of_property: Omit bus properties without a subordinate bus commit 8805840aad73df7146778be243a196d48b4f6430 upstream. A bridge (a device with a Type 1 header) may not have a secondary bus allocated (pdev->subordinate), e.g., if there are no available bus numbers or the bridge secondary/subordinate bus numbers are not writable. The dynamic OF helpers of_pci_prop_bus_range() and of_pci_prop_intr_map() dereference pdev->subordinate without checking it. When CONFIG_PCI_DYNAMIC_OF_NODES is enabled, this can cause a NULL pointer dereference and early boot hang. Generate 'bus-range' and 'interrupt-map' properties only when a subordinate bus exists. Keep the node and its remaining properties for bridges without one. The problem was latent since 407d1a51921e ("PCI: Create device tree node for bridge"), but wasn't reachable until 1f340724419e ("PCI: of: Create device tree PCI host bridge node"), which appeared in v6.15. Before 1f340724419e, of_pci_make_dev_node() returned early because the parent OF node was missing. Fixes: 407d1a51921e ("PCI: Create device tree node for bridge") Signed-off-by: Angel J [bhelgaas: move pdev->subordinate test to callees, commit log] Signed-off-by: Bjorn Helgaas Cc: stable@vger.kernel.org # v6.6+ Link: https://patch.msgid.link/20260918195540.GA1187209@bhelgaas Signed-off-by: Greg Kroah-Hartman commit e59cb7eda7c06334a0faf5aa44adb6acf2a1287a Author: SJ Park Date: Mon Sep 7 10:03:56 2026 -0700 mm/damon/vaddr: avoid hw-driven pte updates during damon_hugetlb_mkold() commit 39c0ceedd54557bdc1542de08d22b2ed33e534e4 upstream. damon_hugetlb_mkold() reads the page table entry into a local variable, unsets the accessed bit in the variable, and updates the page table entry with the updated variable value. If hardware updates the same page table entry in parallel, the hw updates could be lost. For example, hardware-updated dirty bits might be lost. Avoid the parallel updates by clearing the page table entry when reading it together, using huge_ptep_get_and_clear(). If a parallel write to the memory is made after the clearing, the hw will see the page table entry is cleared, trigger page fault and wait until it is handled. The page fault handling will wait for damon_hugetlb_mkold() due to the page table lock. Because hugetlbfs is an in-memory file system and hugetlb pages cannot be reclaimed, no critical issue is expected to my best knowledge. But definitely this is a nasty bug that should be fixed sooner rather than later. The issue was discovered [1] by Sashiko. Link: https://lore.kernel.org/20260907170358.100168-1-sj@kernel.org Link: https://lore.kernel.org/20260830160545.98969-1-sj@kernel.org [1] Fixes: 49f4203aae06 ("mm/damon: add access checking for hugetlb pages") Signed-off-by: SJ Park Signed-off-by: Andrew Morton Cc: Baolin Wang Cc: # 5.17.x Signed-off-by: Greg Kroah-Hartman commit 850cc4a84b36a0062dcd0391acd2450e20ab0767 Author: Jaewook You Date: Mon Sep 14 22:23:52 2026 +0900 mm/hugetlb: preserve mremap address delta when skipping page tables commit 9bdad082d44bdcf93716973dcba6be77e8a06e7b upstream. move_hugetlb_page_tables() optimizes mremap() by advancing to the last entry in the page table when the source page table does not exist, either initially or after unsharing a PMD table. The common loop increment then steps to the first entry in the next page table. However, the code advances both the source and destination addresses to the last entries in their respective page tables, which is wrong. The destination address must be advanced only by the same amount as the source address. If the source and destination offsets within their page tables differ, the destination address can be advanced too far, causing follow-up issues. Fix this by advancing the destination address by the source advance distance. With a reproducer, we were able to trigger a kernel panic on x86-64. With this fix in place, we can no longer reproduce the issue. Link: https://lore.kernel.org/20260914132352.472-1-jaewook376@gmail.com Fixes: e95a9851787b ("hugetlb: skip to end of PT page mapping when pte not present") Fixes: 4ddb4d91b82f ("hugetlb: do not update address in huge_pmd_unshare") Signed-off-by: Jaewook You Signed-off-by: Andrew Morton Acked-by: David Hildenbrand (Arm) Cc: Johan Hovold Cc: Muchun Song Cc: Oscar Salvador Cc: Assisted-by: LLM Signed-off-by: Greg Kroah-Hartman commit b1db1ec5c894130c23bcec822a5d860e24355024 Author: Karl Mehltretter Date: Sat Sep 19 19:11:00 2026 +0200 gpio: tps65219: Fix TPS65214 GPIO direction programming commit 270437f3fe62516f16482742a7762a075e7a9457 upstream. GPIO_LINE_DIRECTION_OUT and GPIO_LINE_DIRECTION_IN have the values 0 and 1, respectively, while the TPS65214 GPIO_CONFIG field is BIT(1). regmap_update_bits() masks the supplied value, so passing either direction value clears the field and selects input mode. Translate the GPIO direction to the register encoding used by tps65214_gpio_get_direction(), setting GPIO_CONFIG for output and clearing it for input. Fixes: 1b6ab07c0c80 ("gpio: tps65219: Add support for TI TPS65214 PMIC") Cc: stable@vger.kernel.org Assisted-by: LLM Signed-off-by: Karl Mehltretter Link: https://patch.msgid.link/20260919171100.90430-4-kmehltretter@gmail.com Signed-off-by: Bartosz Golaszewski Signed-off-by: Greg Kroah-Hartman commit ac3c7eb6a44ba8ac8e74f81fb5d64cbbc250ed06 Author: Karl Mehltretter Date: Sat Sep 19 19:10:59 2026 +0200 gpio: tps65219: Use the variant-specific direction callback commit 93cf8cedeaaa05714f709b539cfb976e0b80c830 upstream. The TPS65214 template installs its own get_direction callback because its direction bit is in GENERAL_CONFIG. The shared get and direction callbacks nevertheless call tps65219_gpio_get_direction() directly and interpret the unrelated TPS65219 MFP bit. On TPS65214 this can reject reads from an input and skip the change from input to output. Call the callback selected by the gpio_chip template instead. Fixes: 1b6ab07c0c80 ("gpio: tps65219: Add support for TI TPS65214 PMIC") Cc: stable@vger.kernel.org Assisted-by: LLM Signed-off-by: Karl Mehltretter Link: https://patch.msgid.link/20260919171100.90430-3-kmehltretter@gmail.com Signed-off-by: Bartosz Golaszewski Signed-off-by: Greg Kroah-Hartman commit fabe897e61e45ddfe83f6d375525fbd11a3bc593 Author: Karl Mehltretter Date: Sat Sep 19 19:10:58 2026 +0200 gpio: tps65219: Fix GPIO input value reads commit 4cbe530c0233c7413aaaeb029a4f32dd6aadacbb upstream. TPS65219_MFP_GPIO_STATUS_MASK is already BIT(4). Passing it to BIT() again tests bit 16, which cannot be set in the 8-bit MFP_CTRL register, so GPIO0 is always reported low when configured as an input. Test the register value with the mask directly. Fixes: 57e30e00bd5b ("gpio: tps65219: add GPIO support for TPS65219 PMIC") Cc: stable@vger.kernel.org Assisted-by: LLM Signed-off-by: Karl Mehltretter Reviewed-by: Jonathan Cormier Link: https://patch.msgid.link/20260919171100.90430-2-kmehltretter@gmail.com Signed-off-by: Bartosz Golaszewski Signed-off-by: Greg Kroah-Hartman commit f2f1683955b95c9b958069da3050f7e8e129ff88 Author: Peiyang He Date: Wed Sep 9 17:11:14 2026 +0800 drm/virtio: fix NULL pointer dereference on fence allocation failure commit 846b3c64fe3e77d9db20a7e3e62dbbb637c773e1 upstream. virtio_gpu_fence_alloc() can fail due to memory pressure and return NULL, but its caller like virtio_gpu_init_submit() never checks it. Later, virtio_gpu_init_submit() passes the NULL fence to virtio_gpu_fence_event_create(), which unconditionally dereferences it. Found when fuzzing the virtio driver with Syzkaller: Oops: general protection fault, probably for non-canonical address 0xdffffc0000000012: 0000 [#1] SMP KASAN NOPTI KASAN: null-ptr-deref in range [0x0000000000000090-0x0000000000000097] CPU: 1 UID: 0 PID: 9991 Comm: syz.0.121 Not tainted 7.2.0 #4 PREEMPT(full) Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS rel-1.17.0-0-gb52ca86e094d-prebuilt.qemu.org 04/01/2014 RIP: 0010:virtio_gpu_fence_event_create drivers/gpu/drm/virtio/virtgpu_submit.c:295 [inline] RIP: 0010:virtio_gpu_init_submit drivers/gpu/drm/virtio/virtgpu_submit.c:398 [inline] RIP: 0010:virtio_gpu_execbuffer_ioctl+0xc78/0x1aa0 drivers/gpu/drm/virtio/virtgpu_submit.c:505 Code: 85 ed 0f 85 21 09 00 00 e8 05 5a c9 fb 48 8b 44 24 10 48 8d b8 90 00 00 00 48 b8 00 00 00 00 00 fc ff df 48 89 fa 48 c1 ea 03 <80> 3c 02 00 0f 85 9a 0d 00 00 48 8b 44 24 10 4c 89 b0 90 00 00 00 RSP: 0018:ffffc900039dfad0 EFLAGS: 00010216 RAX: dffffc0000000000 RBX: ffffc900039dfdd8 RCX: ffffffff85f6fd3d RDX: 0000000000000012 RSI: ffffffff85f6fd4b RDI: 0000000000000090 RBP: 0000000000000000 R0virtio_gpu_virgl_process_cmd: ctrl 0x102, error 0x1203 R10: 0000000000000000 R11: 0000000000000000 R12: ffff8880132c4000 R13: 0000000000000000 R14: ffff888073b6c700 R15: 000000000000003b FS: 00007fab480b96c0(0000) GS:ffff8880eb6e9000(0000) knlGS:0000000000000000 CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 CR2: 00007effbf5e55a8 CR3: 0000000048d19000 CR4: 0000000000350ef0 Call Trace: drm_ioctl_kernel+0x1f4/0x3e0 drivers/gpu/drm/drm_ioctl.c:817 drm_ioctl+0x5f4/0xc70 drivers/gpu/drm/drm_ioctl.c:914 vfs_ioctl fs/ioctl.c:51 [inline] __do_sys_ioctl fs/ioctl.c:597 [inline] __se_sys_ioctl fs/ioctl.c:583 [inline] __x64_sys_ioctl+0x18e/0x210 fs/ioctl.c:583 do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline] do_syscall_64+0x116/0x800 arch/x86/entry/syscall_64.c:94 entry_SYSCALL_64_after_hwframe+0x77/0x7f RIP: 0033:0x7fab471a82bd Code: ff c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e fa 48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 c7 c1 b0 ff ff ff f7 d8 64 89 01 48 RSP: 002b:00007fab480b9018 EFLAGS: 00000246 ORIG_RAX: 0000000000000010 RAX: ffffffffffffffda RBX: 00007fab47435fa0 RCX: 00007fab471a82bd RDX: 00002000000000c0 RSI: 00000000c0406442 RDI: 0000000000000003 RBP: 00007fab480b9080 R08: 0000000000000000 R09: 0000000000000000 R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000001 R13: 00007fab47436038 R14: 00007fab47435fa0 R15: 00007ffe85ab0740 Modules linked in: ---[ end trace 0000000000000000 ]--- RIP: 0010:virtio_gpu_fence_event_create drivers/gpu/drm/virtio/virtgpu_submit.c:295 [inline] RIP: 0010:virtio_gpu_init_submit drivers/gpu/drm/virtio/virtgpu_submit.c:398 [inline] RIP: 0010:virtio_gpu_execbuffer_ioctl+0xc78/0x1aa0 drivers/gpu/drm/virtio/virtgpu_submit.c:505 Code: 85 ed 0f 85 21 09 00 00 e8 05 5a c9 fb 48 8b 44 24 10 48 8d b8 90 00 00 00 48 b8 00 00 00 00 00 fc ff df 48 89 fa 48 c1 ea 03 <80> 3c 02 00 0f 85 9a 0d 00 00 48 8b 44 24 10 4c 89 b0 90 00 00 00 RSP: 0018:ffffc900039dfad0 EFLAGS: 00010216 RAX: dffffc0000000000 RBX: ffffc900039dfdd8 RCX: ffffffff85f6fd3d RDX: 0000000000000012 RSI: ffffffff85f6fd4b RDI: 0000000000000090 RBP: 0000000000000000 R08: 0000000000000005 R09: 0000000000000000 R10: 0000000000000000 R11: 0000000000000000 R12: ffff8880132c4000 R13: 0000000000000000 R14: ffff888073b6c700 R15: 000000000000003b FS: 00007fab480b96c0(0000) GS:ffff888098ae9000(0000) knlGS:0000000000000000 CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 CR2: 00007f24c3759000 CR3: 0000000048d19000 CR4: 0000000000350ef0 ---------------- Code disassembly (best guess): 0: 85 ed test %ebp,%ebp 2: 0f 85 21 09 00 00 jne 0x929 8: e8 05 5a c9 fb call 0xfbc95a12 d: 48 8b 44 24 10 mov 0x10(%rsp),%rax 12: 48 8d b8 90 00 00 00 lea 0x90(%rax),%rdi 19: 48 b8 00 00 00 00 00 movabs $0xdffffc0000000000,%rax 20: fc ff df 23: 48 89 fa mov %rdi,%rdx 26: 48 c1 ea 03 shr $0x3,%rdx * 2a: 80 3c 02 00 cmpb $0x0,(%rdx,%rax,1) <-- trapping instruction 2e: 0f 85 9a 0d 00 00 jne 0xdce 34: 48 8b 44 24 10 mov 0x10(%rsp),%rax 39: 4c 89 b0 90 00 00 00 mov %r14,0x90(%rax) Fix by checking virtio_gpu_fence_alloc() in virtio_gpu_init_submit() and returning -ENOMEM before any later code can dereference the NULL fence. Fixes: 70d1ace56db6 ("drm/virtio: Conditionally allocate virtio_gpu_fence") Cc: stable@vger.kernel.org Signed-off-by: Peiyang He Assisted-by: Codex:gpt-5.5 Signed-off-by: Dmitry Osipenko Link: https://patch.msgid.link/00EFE4BA92889B14+20260909091114.2622550-1-peiyang_he@smail.nju.edu.cn Signed-off-by: Greg Kroah-Hartman commit 8fa78900fd24247277c03dc7b8a2864cb12d4340 Author: Peiyang He Date: Tue Sep 8 20:13:23 2026 +0800 drm/virtio: fix memory leak of fence event on execbuffer failure commit b74aad23d99b279bb34d135795f39a6d8ecdc075 upstream. virtio_gpu_execbuffer_ioctl() reserves a DRM event with drm_event_reserve_init() when VIRTGPU_EXECBUF_RING_IDX selects a ring that userspace has enabled polling for. virtio_gpu_init_submit() does this before the BO handles, the command buffer, the syncobj arrays and the in-fence are processed, so every later error path runs with the event already pending, including plain argument validation failures such as an invalid bo_handle or an in-syncobj that carries no fence. On those paths, virtio_gpu_cleanup_submit() drops the out-fence without cancelling the event. The fence is freed without ever having been emitted, taking the only driver-side pointer to the event with it. Closing the DRM file does not help. drm_events_release() unlinks pending events but deliberately leaves the freeing to the driver's later drm_send_event(), which never runs for an orphaned event, so the allocation is leaked permanently. Found when fuzzing the virtio driver with Syzkaller: BUG: memory leak unreferenced object 0xffff88802c176e80 (size 96): comm "syz.1.367", pid 10561, jiffies 4294960122 hex dump (first 32 bytes): 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 ................ c8 6e 17 2c 80 88 ff ff 00 00 00 00 00 00 00 00 .n.,............ backtrace (crc e1973c6b): kmemleak_alloc_recursive include/linux/kmemleak.h:44 [inline] slab_post_alloc_hook mm/slub.c:4597 [inline] slab_alloc_node mm/slub.c:4917 [inline] __kmalloc_cache_noprof+0x49d/0x6f0 mm/slub.c:5485 _kmalloc_noprof include/linux/slab.h:988 [inline] _kzalloc_noprof include/linux/slab.h:1309 [inline] virtio_gpu_fence_event_create drivers/gpu/drm/virtio/virtgpu_submit.c:282 [inline] virtio_gpu_init_submit drivers/gpu/drm/virtio/virtgpu_submit.c:398 [inline] virtio_gpu_execbuffer_ioctl+0xbbf/0x1aa0 drivers/gpu/drm/virtio/virtgpu_submit.c:505 drm_ioctl_kernel+0x1f4/0x3e0 drivers/gpu/drm/drm_ioctl.c:817 drm_ioctl+0x5f4/0xc70 drivers/gpu/drm/drm_ioctl.c:914 vfs_ioctl fs/ioctl.c:51 [inline] __do_sys_ioctl fs/ioctl.c:597 [inline] __se_sys_ioctl fs/ioctl.c:583 [inline] __x64_sys_ioctl+0x18e/0x210 fs/ioctl.c:583 do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline] do_syscall_64+0x116/0x800 arch/x86/entry/syscall_64.c:94 entry_SYSCALL_64_after_hwframe+0x77/0x7f Fix by cancelling and freeing the DRM event on the execbuffer error path before dropping the fence. Clear the fence's event pointer after cancellation so it does not retain a dangling pointer. Fixes: cd7f5ca33585 ("drm/virtio: implement context init: add virtio_gpu_fence_event") Cc: stable@vger.kernel.org Signed-off-by: Peiyang He Assisted-by: Codex:gpt-5.6-luna Signed-off-by: Dmitry Osipenko Link: https://patch.msgid.link/D320EAB5680C1411+20260908121323.2405044-1-peiyang_he@smail.nju.edu.cn Signed-off-by: Greg Kroah-Hartman commit c2d4689c3ea929314cbbb05740ddee8a77fb5e4b Author: Matthew Brost Date: Thu Sep 17 13:31:58 2026 -0700 drm/xe: Keep walking on SVM eviction failure commit be1df8badae513e01d9575398438716cfae18655 upstream. The desired behavior for SVM eviction failures, which can occur due to various uncontrollable races, is for TTM to continue walking the LRU list and look for another eviction candidate. This is expressed by returning -ENOSPC from the ->move() callback. Adjust the SVM eviction failure path because of races in ->move() to return -ENOSPC so that TTM continues searching for another buffer to evict. Fixes: 3ca608dc7561 ("drm/xe: Basic SVM BO eviction") Cc: stable@vger.kernel.org Signed-off-by: Matthew Brost Reviewed-by: Himal Prasad Ghimiray Link: https://patch.msgid.link/20260917203158.292823-1-matthew.brost@intel.com Signed-off-by: Rodrigo Vivi (cherry picked from commit 36a86c23588b8f57c9d20feb4cf5a2ab27e3baba) Signed-off-by: Rodrigo Vivi Signed-off-by: Greg Kroah-Hartman commit c306b6419fd00ffac089a3c78bd9a00b1c55b64f Author: Szymon Acedański Date: Wed Sep 16 19:30:30 2026 +0200 drm/xe: Limit sg segment size to PAGE_SIZE on Xen PV commit 141008dec73521ccf64878517460cec8b3297251 upstream. Fix display corruption on Xen PV dom0, where DMA buffers are not guaranteed machine-contiguous, in which case bounce buffering kicks in, breaking xe's memory coherency assumptions. Apply the same workaround i915 carries in i915_sg_segment_size() since commit 78a07fe777c4 ("drm/i915: stop abusing swiotlb_max_segment"). Fixes: dd08ebf6c352 ("drm/xe: Introduce a new DRM driver for Intel GPUs") Reported-by: Marek Marczykowski-Górecki Closes: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8382 Link: https://lore.kernel.org/xen-devel/aYtznP_tT6xNPwf-@mail-itl/ Link: https://lore.kernel.org/all/20221020110308.1582518-1-hch@lst.de/ # i915 counterpart Cc: Christoph Hellwig Cc: Robert Beckett Cc: stable@vger.kernel.org # v6.8+ Signed-off-by: Szymon Acedański Reviewed-by: Thomas Hellström Signed-off-by: Thomas Hellström Link: https://patch.msgid.link/20260916173030.3223833-1-accek@invisiblethingslab.com (cherry picked from commit 77f704158f099b952681f207478a22d5b8218edb) Signed-off-by: Rodrigo Vivi Signed-off-by: Greg Kroah-Hartman commit 2bcbfc5e7bfa9cc76818c54db698da5bdc13d9ff Author: Tangudu Tilak Tirumalesh Date: Wed Sep 16 15:35:44 2026 +0530 drm/xe: harden adjust_idledly() against divide-by-zero and overflow commit 90f467577ebc28c5bc4ba2f0f1b83a4419fb0740 upstream. adjust_idledly() has several corner-case issues flagged during review: 1. If xe_gt_clock_init() failed to recognise the crystal clock, gt->info.timestamp_base is 0, which makes idledly_units_ps also 0. The subsequent DIV_ROUND_CLOSEST(..., idledly_units_ps) is then a divide-by-zero and panics the kernel. 2. The tick-to-ns conversions are done in u32: idledly * idledly_units_ps, (maxcnt - 1) * 1000 Both overflow u32 before DIV_ROUND_CLOSEST() sees them. 3. If IDLE_WAIT_TIME reads back as 0, maxcnt evaluates to 0 and the maxcnt - 1 clamp wraps to 0xFFFFFFFF in u32. 4. The register only stores whole ticks, so the clamped ns value has to be converted to ticks and back. DIV_ROUND_CLOSEST() can round that conversion up past maxcnt: maxcnt = 640 ns, one tick = 666664 ps clamp: maxcnt - 1 = 639 ns ns -> ticks: 639000 / 666664 = 0.958 -> rounds to 1 tick tick -> ns: 1 * 666664 / 1000 = 667 ns 667 ns is programmed into RING_IDLEDLY, but 667 >= maxcnt (640), so xe_gt_WARN_ON() fires again on every subsequent init. Return early if timestamp_base is 0 (the unknown-crystal path). Do the conversions in u64 via the *_ULL() helpers so they cannot wrap. Clamp with a floor (DIV_ROUND_DOWN_ULL) so the programmed delay stays strictly below maxcnt, and guard the maxcnt == 0 case with a zero delay while still writing RING_IDLEDLY so INHIBIT_SWITCH_UNTIL_PREEMPTED is cleared. v2: Drop the redundant warn on the timestamp_base == 0 path; xe_gt_clock_init() already warns on an unrecognised crystal clock. Keep the early return to avoid the divide-by-zero. - Vinay v3: Field-mask the RING_IDLEDLY write with REG_FIELD_PREP(IDLE_DELAY, ...) instead of writing the raw tick count, which could clobber INHIBIT_SWITCH_UNTIL_PREEMPTED and reserved bits. Split the inhibit-switch clear from the maxcnt clamp so a set inhibit bit no longer forces a needless delay overwrite when the delay itself is already valid. Use gt_to_xe(gt) instead of gt_to_xe(hwe->gt). Fixes: d2de4410a88f ("drm/xe: Apply Wa_16023105232") Cc: stable@vger.kernel.org Assisted-by: GitHub_Copilot:claude-opus-4.8 Signed-off-by: Tangudu Tilak Tirumalesh Reviewed-by: Vinay Belgaumkar Link: https://patch.msgid.link/20260916100545.779894-2-tilak.tirumalesh.tangudu@intel.com Signed-off-by: Matt Roper (cherry picked from commit d864065ea25e9d12897c175de9176ce46677e176) Signed-off-by: Rodrigo Vivi Signed-off-by: Greg Kroah-Hartman commit ba678c2b02e0d36b53deb5a6d5e5ea044d36def7 Author: Matthew Auld Date: Fri Sep 18 14:10:35 2026 +0100 drm/xe/vm: nuke PTs only after unlinking contested VMAs commit 24a22fb3c731474b68e986af6804db450fb88617 upstream. In xe_vm_close_and_put(), external-BO VMAs are queued on the contested list for deferred destruction via xe_vma_destroy_unlocked(). However, xe_vm_pt_destroy() was previously invoked before processing contested VMAs, destroying vm->pt_root while those VMAs were still linked to their respective buffer objects (vm_bo->list.gpuva). If a concurrent thread evicts one of those shared buffer objects, xe_bo_trigger_rebind() holding only bo->resv walks the BO's VMAs and, in fault mode, calls xe_vm_invalidate_vma() -> xe_pt_zap_ptes(). Because vm->pt_root[tile->id] is already NULL, dereferencing pt->level causes a NULL ptr deref. Fix this by deferring xe_vm_free_scratch() and xe_vm_pt_destroy() until after all contested VMAs have been unlinked and destroyed. User is reporting hitting a NULL ptr deref in xe_pt_zap_ptes(), which could be explained by this race. Assisted-by: LLM Fixes: b06d47be7c83 ("drm/xe: Port Xe to GPUVA") Link: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/9290 Signed-off-by: Matthew Auld Cc: Thomas Hellström Cc: Matthew Brost Cc: # v6.12+ Reviewed-by: Thomas Hellström Reviewed-by: Matthew Brost Link: https://patch.msgid.link/20260918131034.598078-2-matthew.auld@intel.com (cherry picked from commit c2863648959489767f08892fd6e90577d2ea0b6a) Signed-off-by: Rodrigo Vivi Signed-off-by: Greg Kroah-Hartman commit 19f5a5839b308697824ff46e6213a2b4cb677771 Author: Peiyang He Date: Fri Sep 18 10:53:12 2026 +0800 drm/nouveau: don't bump pin count on failed re-pin in nouveau_bo_pin_locked() commit 6a6870d3077faa501ca97760057ddca22b68418d upstream. nouveau_bo_pin_locked() checks whether an already pinned BO is in a memory domain compatible with a new pin request. When the domains are incompatible, it sets -EBUSY but still calls ttm_bo_pin() before returning. Callers treat a failed nouveau_bo_pin() as not having acquired a new pin, so the extra pin count is never decreased by a matching unpin. This triggers the warning in ttm_bo_release(): WARN_ON_ONCE(bo->pin_count); Found when fuzzing the nouveau driver with a modified Syzkaller: WARNING: drivers/gpu/drm/ttm/ttm_bo.c:256 at ttm_bo_release+0x827/0x9e0 drivers/gpu/drm/ttm/ttm_bo.c:256, CPU#1: syz.3.24/2212 Modules linked in: CPU: 1 UID: 0 PID: 2212 Comm: syz.3.24 Not tainted 7.2.0 #24 PREEMPT(lazy) nouveau 0000:01:00.0: gsp:msg fn:103 len:0x40/0x20 res:0x19 resp:0x19 Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-debian-1.16.3-2 04/01/2014 RIP: 0010:ttm_bo_release+0x827/0x9e0 drivers/gpu/drm/ttm/ttm_bo.c:256 Code: 02 00 0f 85 51 01 00 00 48 8b 7b 08 e8 d2 20 01 00 e9 80 fd ff ff e8 d8 15 c0 fe 90 0f 0b 90 e9 e1 f8 ff ff e8 ca 15 c0 fe 90 <0f> 0b 90 e9 a4 f8 ff ff e8 bc 15 c0 fe be 03 00 00 00 4c 89 e7 e8 msg: 00000000: 05 00 d0 c1 04 00 f0 f1 01 30 00 00 2d 90 00 00 .........0..-... RSP: 0018:ffffc9000f5cf710 EFLAGS: 00010293 RAX: 0000000000000000 RBX: ffff888018e5d2a8 RCX: ffffffff82bb1b36 RDX: ffff888017b68000 RSI: 0000000000000004 RDI: ffff888018e5d2a8 msg: 00000010: 19 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 ................ RBP: ffff88801261c720 R08: 0000000000000001 R09: ffffed10031cba55 R10: ffff888018e5d2ab R11: 00000000000000f3 R12: ffff888018e5d290 R13: ffff888018e5d2d4 R14: ffff88801b219c18 R15: dffffc0000000000 FS: 0000000000000000(0000) GS:ffff8880e0f6f000(0000) knlGS:0000000000000000 CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 CR2: 0000001b31223ffc CR3: 0000000028e00005 CR4: 0000000000770ef0 PKRU: 80000000 Call Trace: kref_put include/linux/kref.h:65 [inline] ttm_bo_put drivers/gpu/drm/ttm/ttm_bo.c:325 [inline] ttm_bo_fini+0x55/0x80 drivers/gpu/drm/ttm/ttm_bo.c:330 nouveau_gem_object_del+0xb2/0x1b0 drivers/gpu/drm/nouveau/nouveau_gem.c:90 drm_gem_object_free+0x5f/0x90 drivers/gpu/drm/drm_gem.c:1165 kref_put include/linux/kref.h:65 [inline] __drm_gem_object_put include/drm/drm_gem.h:562 [inline] drm_gem_object_put include/drm/drm_gem.h:575 [inline] nouveau_abi16_chan_fini.constprop.0+0x44f/0x5a0 drivers/gpu/drm/nouveau/nouveau_abi16.c:195 nouveau 0000:01:00.0: syz.2.23[2209]: Unknown handle 0x00000000 nouveau_abi16_fini+0x1d0/0x340 drivers/gpu/drm/nouveau/nouveau_abi16.c:225 nouveau_drm_postclose+0x18b/0x3e0 drivers/gpu/drm/nouveau/nouveau_drm.c:1284 nouveau 0000:01:00.0: syz.2.23[2209]: validate_init drm_file_free.part.0+0x6d6/0xb60 drivers/gpu/drm/drm_file.c:267 drm_file_free drivers/gpu/drm/drm_file.c:237 [inline] drm_close_helper.isra.0+0x11a/0x160 drivers/gpu/drm/drm_file.c:290 drm_release+0x1ab/0x330 drivers/gpu/drm/drm_file.c:438 __fput+0x39c/0xa60 fs/file_table.c:512 nouveau 0000:01:00.0: syz.2.23[2209]: validate: -2 task_work_run+0x15a/0x230 kernel/task_work.c:233 exit_task_work include/linux/task_work.h:40 [inline] do_exit+0x82b/0x25a0 kernel/exit.c:1009 do_group_exit+0xc2/0x280 kernel/exit.c:1152 get_signal+0x1d6e/0x1f30 kernel/signal.c:3046 arch_do_signal_or_restart+0x7d/0x6e0 arch/x86/kernel/signal.c:337 __exit_to_user_mode_loop kernel/entry/common.c:66 [inline] exit_to_user_mode_loop+0xdf/0x440 kernel/entry/common.c:101 __exit_to_user_mode_prepare include/linux/irq-entry-common.h:207 [inline] syscall_exit_to_user_mode_prepare include/linux/irq-entry-common.h:230 [inline] syscall_exit_to_user_mode include/linux/entry-common.h:318 [inline] do_syscall_64+0x4f8/0x690 arch/x86/entry/syscall_64.c:100 entry_SYSCALL_64_after_hwframe+0x77/0x7f RIP: 0033:0x7f12bac8594d Code: Unable to access opcode bytes at 0x7f12bac85923. RSP: 002b:00007f12b96e70d8 EFLAGS: 00000246 ORIG_RAX: 00000000000000ca RAX: 0000000000000001 RBX: 00007f12baf15fa8 RCX: 00007f12bac8594d RDX: 00000000000f4240 RSI: 0000000000000081 RDI: 00007f12baf15fac RBP: 00007f12baf15fa0 R08: 00007f12baee8000 R09: 0000000000000000 R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000000 R13: 00007f12baf16038 R14: 0000000000000006 R15: 00007ffe2ed394b0 irq event stamp: 47867 hardirqs last enabled at (47883): [] __up_console_sem+0x66/0x70 kernel/printk/printk.c:347 hardirqs last disabled at (47892): [] __up_console_sem+0x4b/0x70 kernel/printk/printk.c:345 softirqs last enabled at (47880): [] __do_softirq kernel/softirq.c:656 [inline] softirqs last enabled at (47880): [] invoke_softirq kernel/softirq.c:496 [inline] softirqs last enabled at (47880): [] __irq_exit_rcu+0x137/0x1c0 kernel/softirq.c:735 softirqs last disabled at (47875): [] __do_softirq kernel/softirq.c:656 [inline] softirqs last disabled at (47875): [] invoke_softirq kernel/softirq.c:496 [inline] softirqs last disabled at (47875): [] __irq_exit_rcu+0x137/0x1c0 kernel/softirq.c:735 Fix by calling ttm_bo_pin() only when the existing placement is compatible with the new pin request. This matches the correct behavior in other DRM drivers such as amdgpu_bo_pin() in amdgpu. Cc: stable@vger.kernel.org Fixes: ad76b3f7c7a0 ("drm/nouveau: teach nouveau_bo_pin() how to force a contig vram allocation") Signed-off-by: Peiyang He Assisted-by: LLM Reviewed-by: Lyude Paul Signed-off-by: Lyude Paul Link: https://patch.msgid.link/EACEF2F4E098413F+20260918025312.2814889-1-peiyang_he@smail.nju.edu.cn Signed-off-by: Greg Kroah-Hartman commit 35151e9a8274174ccbb92e3b3ba7dc7a008dcada Author: Jonghyuk Kim(MalHyuk) Date: Wed Sep 2 10:27:16 2026 +0900 drm/nouveau: RCU-free the scheduler-containing nouveau_sched commit f7eae6d8d768fabd6b59779ca7da79e02c74e113 upstream. struct nouveau_sched embeds a struct drm_gpu_scheduler (base). nouveau_sched_destroy() calls nouveau_sched_fini() (which does drm_sched_fini(&sched->base)) and then frees the object with plain kfree(sched). drm_sched_fence_get_timeline_name() returns fence->sched->name, and the scheduler fence keeps a .release callback so it is not ops-detached on signalling. A finished fence exported to userspace via drm_syncobj / sync_file therefore keeps pointing at &sched->base after nouveau_sched_destroy(), and a later get_timeline_name() -- reachable unprivileged through SYNC_IOC_FILE_INFO -- dereferences freed memory (KASAN slab-use-after-free read). Per the dma-fence lifetime contract the exporter must keep the data backing a signalled fence alive for an RCU grace period. Free the scheduler-containing object with kfree_rcu() instead of kfree(). Fixes: 5f03a507b29e ("drm/nouveau: implement 1:1 scheduler - entity relationship") Cc: stable@vger.kernel.org Signed-off-by: Jonghyuk Kim(MalHyuk) Reviewed-by: Lyude Paul Signed-off-by: Lyude Paul Link: https://patch.msgid.link/20260902012717.880724-1-malhyuk97@gmail.com Signed-off-by: Greg Kroah-Hartman commit 59aa1f8c95941e623891d46c805e8fb1257b5a3d Author: Wentao Liang Date: Wed Sep 16 18:03:42 2026 +0000 drm/nouveau: Fix runtime PM leak in nouveau_connector_detect() commit 1e04611d3735543bd80a67d9d13dc13f503746fb upstream. If nvif_outp_edid_get() fails, nouveau_connector_detect() returns early without dropping the runtime PM reference taken at the start of the function, keeping the device powered on until the next successful detect. Balance the reference on the error path like the other exit paths do. Fixes: 0cd7e0718139 ("drm/nouveau/disp: add output method to fetch edid") Cc: stable@vger.kernel.org Signed-off-by: Wentao Liang Reviewed-by: Lyude Paul Signed-off-by: Lyude Paul Link: https://patch.msgid.link/20260916180342.2090360-1-vulab@iscas.ac.cn Signed-off-by: Greg Kroah-Hartman commit 030e93fc91d55d9487b3b85fc0712bb795d77834 Author: Wentao Liang Date: Wed Sep 16 18:02:02 2026 +0000 drm/nouveau: Fix gem reference leak in validate_init() commit 5ea72f7b7139b123713a7983448f910bc4514d9e upstream. On the ttm_bo_reserve() failure and "vma not found" error paths, the loop breaks without adding the looked-up object to any validate list, so the reference taken by drm_gem_object_lookup() is never released; validate_fini() only walks the spliced lists. Drop the reference before breaking out on both paths. Fixes: 19ca10d82e33bcfe ("drm/nouveau/gem: lookup VMAs for buffers referenced by pushbuf ioctl") Cc: stable@vger.kernel.org Signed-off-by: Wentao Liang Reviewed-by: Lyude Paul Signed-off-by: Lyude Paul Link: https://patch.msgid.link/20260916180202.2090231-1-vulab@iscas.ac.cn Signed-off-by: Greg Kroah-Hartman commit 109a8d2d5b7d2719ebe09df3f9241023900b3846 Author: Peiyang He Date: Wed Sep 16 18:31:38 2026 +0800 drm/nouveau: fix double-free in nvif_vmm_dtor commit 97077ac87afe9e91ec074ef0be64454e7ccbf344 upstream. On failure, nouveau_cli_init() calls nouveau_cli_fini() to tear the client down. Then, nouveau_drm_open() also enters into its cleanup path and calls nouveau_cli_fini() AGAIN. nouveau_cli_fini() calls nouveau_vmm_fini(): void nouveau_vmm_fini(struct nouveau_vmm *vmm) { nouveau_svmm_fini(&vmm->svmm); nvif_vmm_dtor(&vmm->vmm); vmm->cli = NULL; } Inside nvif_vmm_dtor(), vmm->page is freed unconditionally: void nvif_vmm_dtor(struct nvif_vmm *vmm) { kfree(vmm->page); nvif_object_dtor(&vmm->object); } vmm->page is never cleared after being freed, so the second call of nvif_vmm_dtor() will cause a double-free. Found by fuzzing the nouveau driver with a modified Syzkaller: BUG: KASAN: double-free in nvif_vmm_dtor+0x31/0x50 drivers/gpu/drm/nouveau/nvif/vmm.c:194 Free of addr ffff888010fcdc30 by task syz.0.173/2567 CPU: 1 UID: 0 PID: 2567 Comm: syz.0.173 Not tainted 7.2.0 #24 PREEMPT(lazy) Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-debian-1.16.3-2 04/01/2014 Call Trace: __dump_stack lib/dump_stack.c:94 [inline] dump_stack_lvl+0x95/0xe0 lib/dump_stack.c:120 print_address_description mm/kasan/report.c:378 [inline] print_report+0xcb/0x5a0 mm/kasan/report.c:482 kasan_report_invalid_free+0xaa/0xd0 mm/kasan/report.c:557 check_slab_allocation+0xe4/0x110 mm/kasan/common.c:235 kasan_slab_pre_free include/linux/kasan.h:199 [inline] slab_free_hook mm/slub.c:2622 [inline] slab_free mm/slub.c:6377 [inline] kfree+0x192/0x590 mm/slub.c:6692 nvif_vmm_dtor+0x31/0x50 drivers/gpu/drm/nouveau/nvif/vmm.c:194 nouveau_vmm_fini+0x16/0x50 drivers/gpu/drm/nouveau/nouveau_vmm.c:127 nouveau_cli_fini+0x10e/0x210 drivers/gpu/drm/nouveau/nouveau_drm.c:225 nouveau_drm_open+0x24e/0x740 drivers/gpu/drm/nouveau/nouveau_drm.c:1255 drm_file_alloc+0x5f2/0xad0 drivers/gpu/drm/drm_file.c:176 drm_open_helper+0x1d7/0x4a0 drivers/gpu/drm/drm_file.c:335 drm_open+0x190/0x3d0 drivers/gpu/drm/drm_file.c:388 drm_stub_open+0x1f2/0x460 drivers/gpu/drm/drm_drv.c:1211 chrdev_open+0x21c/0x660 fs/char_dev.c:411 do_dentry_open+0x59d/0x12b0 fs/open.c:947 vfs_open+0x82/0x390 fs/open.c:1052 do_open fs/namei.c:4700 [inline] path_openat+0x2345/0x3420 fs/namei.c:4863 do_file_open+0x207/0x460 fs/namei.c:4892 do_sys_openat2+0xd1/0x1d0 fs/open.c:1368 do_sys_open fs/open.c:1374 [inline] __do_sys_openat fs/open.c:1390 [inline] __se_sys_openat fs/open.c:1385 [inline] __x64_sys_openat+0x144/0x200 fs/open.c:1385 do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline] do_syscall_64+0x115/0x690 arch/x86/entry/syscall_64.c:94 entry_SYSCALL_64_after_hwframe+0x77/0x7f RIP: 0033:0x7fc6d687594d Code: ff c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e fa 48 89 f8 48 89 f7 48 89 d6 48 89 ca 4d 89 c2 4d 89 c8 4c 8b 4c 24 08 0f 05 <48> 3d 01 f0 ff ff 73 01 c3 48 c7 c1 b0 ff ff ff f7 d8 64 89 01 48 RSP: 002b:00007fc6d5295008 EFLAGS: 00000246 ORIG_RAX: 0000000000000101 RAX: ffffffffffffffda RBX: 00007fc6d6b06180 RCX: 00007fc6d687594d RDX: 0000000000022501 RSI: 0000200000000000 RDI: ffffffffffffff9c RBP: 00007fc6d691c303 R08: 0000000000000000 R09: 0000000000000000 R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000000 R13: 00007fc6d6b06218 R14: 00007fc6d6b06180 R15: 00007ffd9451d760 Allocated by task 2567 on cpu 1 at 163.593900s: kasan_save_stack+0x24/0x50 mm/kasan/common.c:57 kasan_save_track+0x17/0x60 mm/kasan/common.c:78 poison_kmalloc_redzone mm/kasan/common.c:398 [inline] __kasan_kmalloc+0xaa/0xb0 mm/kasan/common.c:415 kasan_kmalloc include/linux/kasan.h:263 [inline] __do_kmalloc_node mm/slub.c:5334 [inline] __kmalloc_noprof+0x304/0x7c0 mm/slub.c:5359 _kmalloc_noprof include/linux/slab.h:992 [inline] nvif_vmm_ctor+0x3c0/0x7e0 drivers/gpu/drm/nouveau/nvif/vmm.c:237 nouveau_vmm_init+0x40/0x90 drivers/gpu/drm/nouveau/nouveau_vmm.c:134 nouveau_cli_init+0x7b9/0xe10 drivers/gpu/drm/nouveau/nouveau_drm.c:293 nouveau_drm_open+0x236/0x740 drivers/gpu/drm/nouveau/nouveau_drm.c:1243 drm_file_alloc+0x5f2/0xad0 drivers/gpu/drm/drm_file.c:176 drm_open_helper+0x1d7/0x4a0 drivers/gpu/drm/drm_file.c:335 drm_open+0x190/0x3d0 drivers/gpu/drm/drm_file.c:388 drm_stub_open+0x1f2/0x460 drivers/gpu/drm/drm_drv.c:1211 chrdev_open+0x21c/0x660 fs/char_dev.c:411 do_dentry_open+0x59d/0x12b0 fs/open.c:947 vfs_open+0x82/0x390 fs/open.c:1052 do_open fs/namei.c:4700 [inline] path_openat+0x2345/0x3420 fs/namei.c:4863 do_file_open+0x207/0x460 fs/namei.c:4892 do_sys_openat2+0xd1/0x1d0 fs/open.c:1368 do_sys_open fs/open.c:1374 [inline] __do_sys_openat fs/open.c:1390 [inline] __se_sys_openat fs/open.c:1385 [inline] __x64_sys_openat+0x144/0x200 fs/open.c:1385 do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline] do_syscall_64+0x115/0x690 arch/x86/entry/syscall_64.c:94 entry_SYSCALL_64_after_hwframe+0x77/0x7f Freed by task 2567 on cpu 1 at 163.601355s: kasan_save_stack+0x24/0x50 mm/kasan/common.c:57 kasan_save_track+0x17/0x60 mm/kasan/common.c:78 kasan_save_free_info+0x3b/0x60 mm/kasan/generic.c:584 poison_slab_object mm/kasan/common.c:253 [inline] __kasan_slab_free+0x61/0x80 mm/kasan/common.c:285 kasan_slab_free include/linux/kasan.h:235 [inline] slab_free_hook mm/slub.c:2677 [inline] slab_free mm/slub.c:6377 [inline] kfree+0x383/0x590 mm/slub.c:6692 nvif_vmm_dtor+0x31/0x50 drivers/gpu/drm/nouveau/nvif/vmm.c:194 nouveau_vmm_fini+0x16/0x50 drivers/gpu/drm/nouveau/nouveau_vmm.c:127 nouveau_cli_fini+0x10e/0x210 drivers/gpu/drm/nouveau/nouveau_drm.c:225 nouveau_cli_init+0x593/0xe10 drivers/gpu/drm/nouveau/nouveau_drm.c:324 nouveau_drm_open+0x236/0x740 drivers/gpu/drm/nouveau/nouveau_drm.c:1243 drm_file_alloc+0x5f2/0xad0 drivers/gpu/drm/drm_file.c:176 drm_open_helper+0x1d7/0x4a0 drivers/gpu/drm/drm_file.c:335 drm_open+0x190/0x3d0 drivers/gpu/drm/drm_file.c:388 drm_stub_open+0x1f2/0x460 drivers/gpu/drm/drm_drv.c:1211 chrdev_open+0x21c/0x660 fs/char_dev.c:411 do_dentry_open+0x59d/0x12b0 fs/open.c:947 vfs_open+0x82/0x390 fs/open.c:1052 do_open fs/namei.c:4700 [inline] path_openat+0x2345/0x3420 fs/namei.c:4863 do_file_open+0x207/0x460 fs/namei.c:4892 do_sys_openat2+0xd1/0x1d0 fs/open.c:1368 do_sys_open fs/open.c:1374 [inline] __do_sys_openat fs/open.c:1390 [inline] __se_sys_openat fs/open.c:1385 [inline] __x64_sys_openat+0x144/0x200 fs/open.c:1385 do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline] do_syscall_64+0x115/0x690 arch/x86/entry/syscall_64.c:94 entry_SYSCALL_64_after_hwframe+0x77/0x7f The buggy address belongs to the object at ffff888010fcdc30 which belongs to the cache kmalloc-16 of size 16 The buggy address is located 0 bytes inside of 16-byte region [ffff888010fcdc30, ffff888010fcdc40) The buggy address belongs to the physical page: page: refcount:0 mapcount:0 mapping:0000000000000000 index:0x0 pfn:0x10fcd flags: 0x100000000000000(node=0|zone=1) page_type: f5(slab) raw: 0100000000000000 ffff88800d441640 dead000000000100 dead000000000122 raw: 0000000000000000 0000000000550055 00000000f5000000 0000000000000000 page dumped because: kasan: bad access detected Memory state around the buggy address: ffff888010fcdb00: fc fc 00 04 fc fc fc fc fa fb fc fc fc fc fa fb ffff888010fcdb80: fc fc fc fc fa fb fc fc fc fc 00 07 fc fc fc fc >ffff888010fcdc00: fa fb fc fc fc fc fa fb fc fc fc fc fa fb fc fc ^ ffff888010fcdc80: fc fc fa fb fc fc fc fc 00 04 fc fc fc fc fa fb ffff888010fcdd00: fc fc fc fc 00 00 fc fc fc fc fa fb fc fc fc fc Fix by removing the redundant teardown in nouveau_drm_open(), since nouveau_cli_init() already does the cleanup work. Also clear vmm->page after its freeing. Cc: stable@vger.kernel.org Fixes: 20d8a88e557a ("drm/nouveau: tidy up the client init/fini interfaces") Signed-off-by: Peiyang He Assisted-by: LLM Reviewed-by: Lyude Paul Signed-off-by: Lyude Paul Link: https://patch.msgid.link/03BA723D9E5FF725+20260916103138.2651605-1-peiyang_he@smail.nju.edu.cn Signed-off-by: Greg Kroah-Hartman commit 6d497216c10af3dc9747b399295f385fabbac9e5 Author: Wentao Liang Date: Wed Sep 16 18:00:36 2026 +0000 drm/nouveau: Fix bridge reference leak in nv1a_ram_new() commit 67b4411538c8341692548429d43256f25be99f7a upstream. pci_get_domain_bus_and_slot() takes a reference to the PCI device, which is never released once the memory size has been read from its config space. Drop the reference before returning. Fixes: 2fa6d6cdaf283c05 ("drm/nouveau: deprecate pci_get_bus_and_slot()") Cc: stable@vger.kernel.org Signed-off-by: Wentao Liang Reviewed-by: Lyude Paul Signed-off-by: Lyude Paul Link: https://patch.msgid.link/20260916180036.2090118-1-vulab@iscas.ac.cn Signed-off-by: Greg Kroah-Hartman commit 9132215d2c90e216c52a52b4ae59eab35454724f Author: Guangshuo Li Date: Sat Aug 8 21:41:37 2026 +0800 drm/nouveau: fix autosuspend cleanup during teardown commit fefd9480ec361969f1a836df46326a1801062c26 upstream. nouveau_drm_device_init() calls pm_runtime_use_autosuspend(), but nouveau_drm_device_fini() does not call the matching pm_runtime_dont_use_autosuspend(). If the autosuspend delay is set to a negative value while autosuspend is enabled, the runtime PM core increments usage_count to prevent runtime suspend. Without calling pm_runtime_dont_use_autosuspend() during teardown, this reference is not dropped and usage_count remains unbalanced. The documentation for pm_runtime_use_autosuspend() also notes that it is important to undo it with pm_runtime_dont_use_autosuspend() at driver exit time, unless runtime PM was initially enabled with devm_pm_runtime_enable(). Add the missing pm_runtime_dont_use_autosuspend() call to the common device teardown path. This issue was found by manual code inspection. Fixes: 5addcf0a5f0f ("nouveau: add runtime PM support (v0.9)") Cc: stable@vger.kernel.org Signed-off-by: Guangshuo Li Reviewed-by: Lyude Paul Signed-off-by: Lyude Paul Link: https://patch.msgid.link/20260808134137.2864847-1-lgs201920130244@gmail.com Signed-off-by: Greg Kroah-Hartman commit 0d29bf51f8dddcf8556bf76710d3dd3024a1a717 Author: Peiyang He Date: Mon Sep 7 13:22:04 2026 +0800 drm/nouveau/uvmm: fix UAF in nouveau_uvmm_sm when BO is in TTM_PL_SYSTEM commit 3359a372efb6d585c97019ee1b7f1874442bcebe upstream. nouveau_uvmm_sm() calls op_map(), which passes bo->resource through nouveau_mem() to nouveau_uvma_map(). nouveau_uvmm_vmm_map() then reads mem->mem.type. But this is only valid when bo->resource is backed by struct nouveau_mem, as is the case for VRAM and TT resources. If the BO is left in TTM_PL_SYSTEM, bo->resource is only a struct ttm_resource. Treating it as struct nouveau_mem makes the mem->mem.type read past the end of the resource, causing a KASAN: slab-use-after-free Read in nouveau_uvmm_sm report: BUG: KASAN: slab-use-after-free in nouveau_uvmm_vmm_map drivers/gpu/drm/nouveau/nouveau_uvmm.c:152 [inline] BUG: KASAN: slab-use-after-free in nouveau_uvma_map drivers/gpu/drm/nouveau/nouveau_uvmm.c:199 [inline] BUG: KASAN: slab-use-after-free in op_map drivers/gpu/drm/nouveau/nouveau_uvmm.c:849 [inline] BUG: KASAN: slab-use-after-free in nouveau_uvmm_sm.constprop.0+0x6ab/0x900 drivers/gpu/drm/nouveau/nouveau_uvmm.c:903 Read of size 1 at addr ffff888127d3e3a0 by task kworker/0:1/11 CPU: 0 UID: 0 PID: 11 Comm: kworker/0:1 Not tainted 7.2.0 #5 PREEMPT(lazy) Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-debian-1.16.3-2 04/01/2014 Workqueue: nouveau_sched_wq_2224 drm_sched_run_job_work Call Trace: __dump_stack lib/dump_stack.c:94 [inline] dump_stack_lvl+0x95/0xe0 lib/dump_stack.c:120 print_address_description mm/kasan/report.c:378 [inline] print_report+0xcb/0x5a0 mm/kasan/report.c:482 kasan_report+0xca/0x100 mm/kasan/report.c:595 nouveau_uvmm_vmm_map drivers/gpu/drm/nouveau/nouveau_uvmm.c:152 [inline] nouveau_uvma_map drivers/gpu/drm/nouveau/nouveau_uvmm.c:199 [inline] op_map drivers/gpu/drm/nouveau/nouveau_uvmm.c:849 [inline] nouveau_uvmm_sm.constprop.0+0x6ab/0x900 drivers/gpu/drm/nouveau/nouveau_uvmm.c:903 nouveau_uvmm_sm_unmap drivers/gpu/drm/nouveau/nouveau_uvmm.c:932 [inline] nouveau_uvmm_bind_job_run+0xd6/0x250 drivers/gpu/drm/nouveau/nouveau_uvmm.c:1532 nouveau_job_run drivers/gpu/drm/nouveau/nouveau_sched.c:350 [inline] nouveau_sched_run_job+0x62/0xd0 drivers/gpu/drm/nouveau/nouveau_sched.c:364 drm_sched_run_job_work+0x356/0xa10 drivers/gpu/drm/scheduler/sched_main.c:1061 process_one_work+0x8a5/0x1900 kernel/workqueue.c:3322 process_scheduled_works kernel/workqueue.c:3405 [inline] worker_thread+0x5dd/0xd80 kernel/workqueue.c:3486 kthread+0x31d/0x420 kernel/kthread.c:436 ret_from_fork+0x662/0x940 arch/x86/kernel/process.c:158 ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245 Allocated by task 2224 on cpu 0 at 66.550027s: kasan_save_stack+0x24/0x50 mm/kasan/common.c:57 kasan_save_track+0x17/0x60 mm/kasan/common.c:78 poison_kmalloc_redzone mm/kasan/common.c:398 [inline] __kasan_kmalloc+0xaa/0xb0 mm/kasan/common.c:415 kasan_kmalloc include/linux/kasan.h:263 [inline] __do_kmalloc_node mm/slub.c:5334 [inline] __kmalloc_noprof+0x304/0x7c0 mm/slub.c:5359 _kmalloc_noprof include/linux/slab.h:992 [inline] dma_resv_list_alloc+0x27/0x90 drivers/dma-buf/dma-resv.c:106 dma_resv_reserve_fences+0x60e/0xa30 drivers/dma-buf/dma-resv.c:205 ttm_bo_alloc_resource+0x12c/0xbd0 drivers/gpu/drm/ttm/ttm_bo.c:721 ttm_bo_validate+0x1bc/0x4a0 drivers/gpu/drm/ttm/ttm_bo.c:856 ttm_bo_init_reserved+0x2c3/0x570 drivers/gpu/drm/ttm/ttm_bo.c:970 nouveau_bo_init+0x159/0x2c0 drivers/gpu/drm/nouveau/nouveau_bo.c:359 nouveau_gem_new+0x234/0x5f0 drivers/gpu/drm/nouveau/nouveau_gem.c:272 nouveau_gem_ioctl_new+0x1eb/0x420 drivers/gpu/drm/nouveau/nouveau_gem.c:352 drm_ioctl_kernel+0x192/0x350 drivers/gpu/drm/drm_ioctl.c:817 drm_ioctl+0x4f8/0xb40 drivers/gpu/drm/drm_ioctl.c:914 nouveau_drm_ioctl+0xea/0x2c0 drivers/gpu/drm/nouveau/nouveau_drm.c:1338 vfs_ioctl fs/ioctl.c:51 [inline] __do_sys_ioctl fs/ioctl.c:597 [inline] __se_sys_ioctl fs/ioctl.c:583 [inline] __x64_sys_ioctl+0x180/0x1d0 fs/ioctl.c:583 do_syscall_x64 arch/x86/entry/syscall_64.c:63 [inline] do_syscall_64+0x115/0x690 arch/x86/entry/syscall_64.c:94 entry_SYSCALL_64_after_hwframe+0x77/0x7f Freed by task 2223 on cpu 0 at 66.554063s: kasan_save_stack+0x24/0x50 mm/kasan/common.c:57 kasan_save_track+0x17/0x60 mm/kasan/common.c:78 kasan_save_free_info+0x3b/0x60 mm/kasan/generic.c:584 poison_slab_object mm/kasan/common.c:253 [inline] __kasan_slab_free+0x61/0x80 mm/kasan/common.c:285 kasan_slab_free include/linux/kasan.h:235 [inline] slab_free_hook mm/slub.c:2677 [inline] __rcu_free_sheaf_prepare+0xb6/0x2e0 mm/slub.c:2928 rcu_free_sheaf+0x1b/0x120 mm/slub.c:5978 rcu_do_batch kernel/rcu/tree.c:2645 [inline] rcu_core+0x521/0x1490 kernel/rcu/tree.c:2897 handle_softirqs+0x1b1/0x8a0 kernel/softirq.c:622 __do_softirq kernel/softirq.c:656 [inline] invoke_softirq kernel/softirq.c:496 [inline] __irq_exit_rcu+0x137/0x1c0 kernel/softirq.c:735 irq_exit_rcu+0x9/0x20 kernel/softirq.c:752 instr_sysvec_apic_timer_interrupt arch/x86/kernel/apic/apic.c:1062 [inline] sysvec_apic_timer_interrupt+0x70/0x80 arch/x86/kernel/apic/apic.c:1062 asm_sysvec_apic_timer_interrupt+0x1a/0x20 arch/x86/include/asm/idtentry.h:674 The buggy address belongs to the object at ffff888127d3e380 which belongs to the cache kmalloc-96 of size 96 The buggy address is located 32 bytes inside of freed 96-byte region [ffff888127d3e380, ffff888127d3e3e0) The buggy address belongs to the physical page: page: refcount:0 mapcount:0 mapping:0000000000000000 index:0x0 pfn:0x127d3e flags: 0x200000000000000(node=0|zone=2) page_type: f5(slab) raw: 0200000000000000 ffff888100041280 dead000000000122 0000000000000000 raw: 0000000000000000 0000000000200020 00000000f5000000 0000000000000000 page dumped because: kasan: bad access detected Memory state around the buggy address: ffff888127d3e280: 00 00 00 00 00 00 00 00 00 fc fc fc fc fc fc fc ffff888127d3e300: 00 00 00 00 00 00 00 00 00 00 00 fc fc fc fc fc >ffff888127d3e380: fa fb fb fb fb fb fb fb fb fb fb fb fc fc fc fc ^ ffff888127d3e400: fa fb fb fb fb fb fb fb fb fb fb fb fc fc fc fc ffff888127d3e480: 00 00 00 00 00 00 00 00 00 00 fc fc fc fc fc fc Fix by resetting the placement to the BO's valid domains before calling nouveau_bo_validate(), matching the handling in nouveau_uvmm_bo_validate(), so map jobs do not run for SYSTEM resources; Reject BO that cannot reside in VRAM or GART; Also skip op_map() when the GPUVA has been invalidated, matching the handling in the unmap and remap paths. Found when fuzzing the nouveau driver with a modified Syzkaller. Fixes: b88baab82871 ("drm/nouveau: implement new VM_BIND uAPI") Cc: stable@vger.kernel.org Signed-off-by: Peiyang He Assisted-by: Codex:gpt-5.5 Reviewed-by: Lyude Paul Signed-off-by: Lyude Paul Link: https://patch.msgid.link/0D77BEC410CE0129+20260907052204.1431488-1-peiyang_he@smail.nju.edu.cn Signed-off-by: Greg Kroah-Hartman commit c83261854e19a58ff80d6ae46715f0b0ed9e8776 Author: Wentao Liang Date: Wed Sep 16 10:01:39 2026 +0000 drm/amdgpu: Fix vmid_wait fence leak in amdgpu_ring_init() commit aea841bc62a76242396610d22d8ff40c13065f64 upstream. amdgpu_ring_init() initializes ring->vmid_wait with a reference to the stub fence taken via dma_fence_get_stub(). When a later step of the initialization fails, e.g. amdgpu_fence_driver_init_ring(), a writeback slot allocation or the ring buffer allocation, the function returns an error without releasing the stub fence reference and the reference is leaked if the ring is torn down without amdgpu_ring_fini(). Move the stub fence assignment to the end of the initialization, right before the ring is registered with the GPU scheduler, where no further failure is possible. The stub fence is only consumed by command submission handling in amdgpu_ids.c once the ring is up and running, so nothing reads it during the error-prone part of the initialization. Fixes: 48e9fbd1a284 ("drm/amdgpu: initialize the vmid_wait with the stub fence") Signed-off-by: Wentao Liang Signed-off-by: Alex Deucher (cherry picked from commit f2b96986851203e9c50ca0d13aaa3581ca3e8ebd) Cc: stable@vger.kernel.org Signed-off-by: Greg Kroah-Hartman commit c28e74395ab09922bbf5be42bc5a6ed1ba4f3c37 Author: Wentao Liang Date: Wed Sep 16 09:58:05 2026 +0000 drm/amdgpu: Fix runtime PM leak in amdgpu_debugfs_test_ib_show() commit 2b86ab1bd6673c525adda88819d7658ba9e784ec upstream. amdgpu_debugfs_test_ib_show() resumes the device with pm_runtime_get_sync() before taking the reset domain semaphore with down_write_killable(). If the write lock acquisition is interrupted, the function returns without calling pm_runtime_put_autosuspend(), leaking the runtime PM reference acquired for the device and keeping the GPU awake. Drop the runtime PM reference on the interrupted down_write_killable() error path before returning. Fixes: 6049db43d6dd ("drm/amdgpu: change reset lock from mutex to rw_semaphore") Signed-off-by: Wentao Liang Signed-off-by: Alex Deucher (cherry picked from commit ec30a576c2d4c0364549e6c04218f50704ef56c8) Cc: stable@vger.kernel.org Signed-off-by: Greg Kroah-Hartman commit 0bfb0419a1480d71b53bb65f656764318993974c Author: Wentao Liang Date: Wed Sep 16 09:55:36 2026 +0000 drm/amdgpu: Fix last_update fence leak in amdgpu_vm_init() commit b4f7b4459b1b155e4c4977a6482b5df2cf08758c upstream. amdgpu_vm_init() initializes vm->last_update, vm->last_unlocked and vm->last_tlb_flush with references to the stub fence taken via dma_fence_get_stub(). The error label at the end of the function releases the last_unlocked and last_tlb_flush references with dma_fence_put(), but the reference stored in vm->last_update is never dropped, so whenever the page table root creation, the reservation of the root BO or the PASID registration fails, the stub fence reference leaks. Drop the vm->last_update reference together with the other stub fence references on the error path. Fixes: 187916e6ed9d ("drm/amdgpu: install stub fence into potential unused fence pointers") Signed-off-by: Wentao Liang Signed-off-by: Alex Deucher (cherry picked from commit e7979c84fc05a176bdf855ee664871b1648404c9) Cc: stable@vger.kernel.org Signed-off-by: Greg Kroah-Hartman commit d197b15d4beb9dc1785b66afb42091655e6fda64 Author: Ivan Lipski Date: Fri Aug 21 00:04:53 2026 -0400 drm/amd/display: Bump frame warning limit for clang builds of dml commit 18779dd84515db093fedb4ebaf0998c9b165a5fb upstream. [Why&How] When building the DML files with clang without any sanitizer or LTO, the following -Wframe-larger-than errors break the build under CONFIG_WERROR: display_mode_vba_30.c: error: stack frame size (2512) exceeds limit (2048) in 'dml30_ModeSupportAndSystemConfigurationFull' display_mode_vba_31.c: error: stack frame size (2416) exceeds limit (2048) in 'dml31_ModeSupportAndSystemConfigurationFull' display_mode_vba_314.c: error: stack frame size (2392) exceeds limit (2048) in 'dml314_ModeSupportAndSystemConfigurationFull' Clang consistently spills more than gcc, pushing the frame past the 2048 byte limit. Apply an existing approach of increasing the warn stack size to the non-sanitizer path so plain clang builds use a 3072 byte limit. Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5642 Signed-off-by: Ivan Lipski Signed-off-by: Alex Deucher (cherry picked from commit 21711b6e66bb7b41b1aec67b2d99aafe768c8fcb) Cc: stable@vger.kernel.org Signed-off-by: Greg Kroah-Hartman commit 89a11c7efe387294aab342c3ddcdff0a4744cb0a Author: Wentao Liang Date: Wed Sep 16 09:49:56 2026 +0000 drm/amd/display: Fix dc stream excess put in dm_update_crtc_state() commit c5fd4eaad50d620c7e09ac2082b2fb55ee54170e upstream. In dm_update_crtc_state(), when a modeset is required the newly created stream is stored in dm_new_crtc_state->stream and an extra reference is taken with dc_stream_retain(). The reference returned by create_validate_stream_for_sink() is released as an extra reference at the skip_modeset label, leaving the stream owned by the new CRTC state. If amdgpu_dm_check_crtc_color_mgmt() fails afterwards, the code jumps to the fail label which releases new_stream again. Since the extra reference was already released at skip_modeset, this drops the reference owned by dm_new_crtc_state->stream and the stream is released while the atomic state still points to it, leading to a premature free of the dc stream. Set new_stream to NULL after releasing the extra reference at the skip_modeset label so that a later goto fail cannot release the reference owned by the new CRTC state. Fixes: 7cd4b70091a5 ("drm/amd/display: Rework CRTC color management") Signed-off-by: Wentao Liang Signed-off-by: Alex Deucher (cherry picked from commit 102a47065a62dc8f6bbbb47cf082a2934282eb08) Cc: stable@vger.kernel.org Signed-off-by: Greg Kroah-Hartman commit fe933dda0bb1814cbc4e68a9b21792f8b959d8c3 Author: Imre Deak Date: Mon Sep 7 20:44:13 2026 +0300 drm/i915/dp_mst: Fix configuring TUs for a disconnected stream commit a443e0b8d647c1401b110d9f919d8c6cb8607260 upstream. During an atomic commit after all the MST stream CRTC state is computed the driver ensures that the sum of TUs of all the streams on a given MST topology link is within limits (63 for 8b10 and 64 for 128b132b). For a disconnected stream the DRM MST core's BW verification doesn't ensure this, because the topology state it uses for this is destroyed as soon as the stream (i.e. MST connector/port) is disconnected. The driver should keep the link state valid even for such disconnected streams, as userspace may disable them one-by-one only in a deferred way. Ensure the link's sum of TUs stays within limits in this case by simply reusing the maximum link BPP limit from the stream's (i.e. CRTC's) old state. The disconnection can happen either via the whole topology getting disconnected or via only the given stream's port getting disconnected. Check for both of these conditions separately, as a connector gets unregistered after a link disconnect event only in a deferred way. Cc: stable@vger.kernel.org # v6.10+ Link: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16073 Link: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16384 Reviewed-by: Luca Coelho Signed-off-by: Imre Deak Link: https://patch.msgid.link/20260907174413.741851-2-imre.deak@intel.com (cherry picked from commit ee00f8fbb2b202002ab90834e02e9ba372773a36) Signed-off-by: Jani Nikula Signed-off-by: Greg Kroah-Hartman commit 369f4e32643746b35189ad2ceefa2307577ca4e7 Author: Imre Deak Date: Mon Sep 7 20:44:12 2026 +0300 drm/i915/dp_mst: Fix configuring FEC for a disconnected stream commit acbe9a3b60b9a6ace8ef11fe898f586251c84592 upstream. During an atomic commit after all the MST stream CRTC state is computed the driver ensures that the FEC is configured the same way (enabled or disabled) for all the streams on a given MST topology's link. drm_dp_mst_port_downstream_of_parent() used to determine if a stream is downstream of an MST port will return false if the whole topology is disconnected, since in that case it can't verify that the port/ parent_port passed to it is in the given MST topology. This is a problem during the above FEC configuration check, since intel_dp_mst_check_dsc_change()->get_pipes_downstream_of_mst_ports() will not return all the stream CRTCs/pipes for the topology as expected. Since passing parent_port==NULL to get_pipes_downstream_of_mst_port() is meant to return all the streams for the given topology (i.e. mst_mgr) skip checking if an MST port is downstream of a parent port in this case. This fixes a problem where the FEC configuration check explained above failed to ensure that all streams' FEC is configured the same way if the topology was disconnected, leading to a FEC state mismatch error. Cc: stable@vger.kernel.org # v6.10+ Closes: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16073 Closes: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16384 Reviewed-by: Luca Coelho Signed-off-by: Imre Deak Link: https://patch.msgid.link/20260907174413.741851-1-imre.deak@intel.com (cherry picked from commit 270681fbffbba2b6ccf5b7e3c34b8b563b36167f) Signed-off-by: Jani Nikula Signed-off-by: Greg Kroah-Hartman commit cc890d3f1b15bf6481dc7270fba0b4fbba572813 Author: Christian König Date: Thu Sep 3 13:36:21 2026 +0200 drm/i915: fix incorrect RCU teardown order commit d2da6696e0c4e60414706e607029d0bb0330c67e upstream. i915_gem_busy_ioctl uses dma_resv_for_each_fence_unlocked() to iterate over the fences in an GEM object without holding a reference but only the RCU read side lock. What can happen here is that the GEM object is destroyed concurrently while i915_gem_busy_ioctl is still running. This won't free the GEM objects memory, but still drops all the dma_fence references. Now when dma_resv_for_each_fence_unlocked() sees a destroyed dma_fence it assumes that a new fence list was installed and re-starts the loop. But in the case of a destroyed GEM object a new fence list is never installed, only the old one freed and therefore the iteration never finishes resulting in an endless loop. The solution is to drop the fence references only after the RCU grace period. The fixes tag is not necessary the patch introducing the problem, but the one making it so worse that we need to address it. This problem was pointed out by Sashiko-bot. Signed-off-by: Christian König Fixes: 912ff2ebd695 ("drm/i915: use the new iterator in i915_gem_busy_ioctl v2") CC: stable@vger.kernel.org Reviewed-by: Tvrtko Ursulin Signed-off-by: Tvrtko Ursulin Link: https://lore.kernel.org/r/20260903113621.54660-1-christian.koenig@amd.com (cherry picked from commit 5113479556025093bf8133bb2dcaa33be2d50921) Signed-off-by: Jani Nikula Signed-off-by: Greg Kroah-Hartman commit bbef7450e2b423fb8531c68b3ef8febd5153d7e5 Author: Brajesh Gupta Date: Tue Sep 22 09:56:56 2026 +0530 drm/imagination: Fix page count for page table for map() interface commit 0a8224058a5835297dcf4a46bbcd16f77a9fe424 upstream. The GPU virtual start address wasn't included in the calculation for the amount of page tables required for mapping a BO object in map() interface. It resulted in map failure later due to not enough pages at L0/L1 level. Update pvr_mmu_op_context_create() interface to pass device address as well to allow correct calculation for page table memory. If L0 tables cover 2MB (0x200000), the range defined by device address 0x80001ff000 (general heap at 2MB - 4KB) and size 0x2000 (two 4KB pages) requires two L0 pages to be mapped, but without the base address a range of 0x2000 computes to a single L0 page which is not enough. Fixes: ff5f643de0bf ("drm/imagination: Add GEM and VM related code") Reviewed-by: Alexandru Dadu Reviewed-by: Alessio Belle Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260922-mmu_fix-v4-2-12f1a871456a@imgtec.com Signed-off-by: Brajesh Gupta Signed-off-by: Greg Kroah-Hartman commit afe5b4d66d45b278ecd0e3d6889fb208ce2e5339 Author: Brajesh Gupta Date: Tue Sep 22 09:56:55 2026 +0530 drm/imagination: Propagate map failures correctly from pvr_mmu_map_sgl() commit 7b824c293a6b56de8285a97984c507cba56bc4c4 upstream. Map failure from pvr_mmu_map_sgl() interface was not returned correctly to pvr_mmu_map() interface. This resulted in pvr_mmu_map() interface to continue instead of returning an error to caller. Fix it by returning a proper error code from pvr_mmu_map_sgl() interface. Call stack for crash: [ 1179.286237] Unable to handle kernel NULL pointer dereference at virtual address 0000000000000008 [ 1179.295067] Mem abort info: [ 1179.297877] ESR = 0x0000000096000004 [ 1179.301656] EC = 0x25: DABT (current EL), IL = 32 bits [ 1179.306987] SET = 0, FnV = 0 [ 1179.310048] EA = 0, S1PTW = 0 [ 1179.313198] FSC = 0x04: level 0 translation fault [ 1179.318096] Data abort info: [ 1179.320993] ISV = 0, ISS = 0x00000004, ISS2 = 0x00000000 [ 1179.326483] CM = 0, WnR = 0, TnD = 0, TagAccess = 0 [ 1179.331546] GCS = 0, Overlay = 0, DirtyBit = 0, Xs = 0 [ 1179.336895] user pgtable: 4k pages, 48-bit VAs, pgdp=000000009822a000 [ 1179.343402] [0000000000000008] pgd=0000000000000000, p4d=0000000000000000 [ 1179.350243] Internal error: Oops: 0000000096000004 [#2] SMP [ 1179.355908] Modules linked in: powervr gpu_sched drm_shmem_helper drm_gpuvm drm_exec xhci_plat_hcd xhci_hcd dwc3 usbcore usb_common snd_soc_simple_card snd_soc_simple_card_utils dwc3_am62 at24 sa2ul sha512 libsha512 sha256 authenc sch_fq_codel fuse dm_mod ipv6 [ 1179.378992] CPU: 1 UID: 1000 PID: 680 Comm: deqp-vk Tainted: G D 6.17.0 #1 PREEMPT [ 1179.388120] Tainted: [D]=DIE [ 1179.390994] Hardware name: Texas Instruments AM625 SK (DT) [ 1179.396467] pstate: 00000005 (nzcv daif -PAN -UAO -TCO -DIT -SSBS BTYPE=--) [ 1179.403415] pc : pvr_mmu_op_context_unmap_curr_page+0x6c/0x134 [powervr] [ 1179.410140] lr : pvr_mmu_op_context_unmap_curr_page+0x58/0x134 [powervr] [ 1179.416848] sp : ffff8000839ab8c0 [ 1179.420153] x29: ffff8000839ab8c0 x28: 0000000000000001 x27: 000000008f386000 [ 1179.427283] x26: ffff000016d1df98 x25: 0000000000247000 x24: 00000000000001e6 [ 1179.434413] x23: 0000000000000002 x22: 000000000000ffff x21: 0000000000000247 [ 1179.441540] x20: 0000000000000245 x19: ffff000016d1df60 x18: 0000000000000002 [ 1179.448668] x17: 0000000000000000 x16: 0000000000000000 x15: 0000000000000001 [ 1179.455793] x14: 0000000000060810 x13: ffff80007fffffff x12: ffff000004190480 [ 1179.462921] x11: ffff8000853f7000 x10: ffff8000811ae000 x9 : ffff0000041900b8 [ 1179.470051] x8 : 0000000000000000 x7 : 00000000990c4001 x6 : 0000000000000007 [ 1179.477177] x5 : ffff000016d1df60 x4 : 0000000000000000 x3 : ffff00000a7d8000 [ 1179.484306] x2 : 00000000000001ff x1 : 0000000000000000 x0 : 0000000000000000 [ 1179.491433] Call trace: [ 1179.493872] pvr_mmu_op_context_unmap_curr_page+0x6c/0x134 [powervr] (P) [ 1179.500582] pvr_mmu_map+0x31c/0x388 [powervr] [ 1179.505027] pvr_vm_gpuva_map+0x40/0x88 [powervr] [ 1179.509732] __drm_gpuvm_sm_map+0x250/0x44c [drm_gpuvm] [ 1179.514952] drm_gpuvm_sm_map+0x48/0x5c [drm_gpuvm] [ 1179.519822] pvr_vm_bind_op_exec+0x64/0x70 [powervr] [ 1179.524785] pvr_vm_map+0x1f8/0x2a8 [powervr] [ 1179.529142] pvr_ioctl_vm_map+0x12c/0x188 [powervr] [ 1179.534018] drm_ioctl_kernel+0xb8/0x128 [ 1179.537941] drm_ioctl+0x21c/0x4ec [ 1179.541337] __arm64_sys_ioctl+0xac/0x108 [ 1179.545344] invoke_syscall+0x44/0x100 [ 1179.549091] el0_svc_common.constprop.0+0x40/0xe0 [ 1179.553790] do_el0_svc+0x1c/0x28 [ 1179.557106] el0_svc+0x34/0xf0 [ 1179.560159] el0t_64_sync_handler+0xd0/0xe4 [ 1179.564334] el0t_64_sync+0x198/0x19c [ 1179.567996] Code: 54000300 35000360 f9402261 79409a62 (f9400421) [ 1179.574081] ---[ end trace 0000000000000000 ]--- Fixes: ff5f643de0bf ("drm/imagination: Add GEM and VM related code") Reviewed-by: Alexandru Dadu Reviewed-by: Alessio Belle Cc: stable@vger.kernel.org Link: https://patch.msgid.link/20260922-mmu_fix-v4-1-12f1a871456a@imgtec.com Signed-off-by: Brajesh Gupta Signed-off-by: Greg Kroah-Hartman commit e2ee0616260080376123997488b5cf8a3870b8d9 Author: Dongliang Qin Date: Tue Sep 22 11:15:45 2026 +0800 rds: ib: Clear the sg list when mapping an MR fails commit 58eb1b3325edac42dc6df72c80962bd53a3c8ca7 upstream. rds_ib_map_frmr() stores the caller's scatterlist in the MR before DMA mapping and registration can fail. On failure, __rds_rdma_map() unpins the pages and frees the scatterlist, but rds_ib_free_frmr() can still return the MR to the pool with the stale pointer set. This leaves the pool with a dangling scatterlist and can lead to local privilege escalation. KASAN detects the resulting use-after-free when the MR is later torn down: BUG: KASAN: slab-use-after-free in __rds_ib_teardown_mr Read of size 8 Call Trace: __rds_ib_teardown_mr rds_ib_unreg_frmr rds_ib_flush_mr_pool rds_ib_flush_mrs rds_free_mr rds_setsockopt Store the scatterlist in the MR only after DMA mapping succeeds. If DMA mapping fails, return directly while the MR fields remain clear; the caller keeps ownership of the scatterlist and its pinned pages. If a later registration step fails, unmap the scatterlist and clear the MR fields before returning. Fixes: 1659185fb4d0 ("RDS: IB: Support Fastreg MR (FRMR) memory registration mode") Cc: stable@vger.kernel.org Signed-off-by: Dongliang Qin Reviewed-by: Allison Henderson Link: https://patch.msgid.link/20260922031546.3874605-1-cccccccccccc777777@gmail.com Signed-off-by: Paolo Abeni Signed-off-by: Greg Kroah-Hartman commit e35cb500dd14882f8dee61783fccd1823fd3ab0d Author: Hui Peng Date: Mon Sep 21 05:10:01 2026 +0000 mctp: route: iterate socket tag list in mctp_lookup_prealloc_tag() commit 26cc0e69cce062cd3aa6fae33074684669c35a71 upstream. When a socket transmits a packet with MCTP_TAG_PREALLOC set, mctp_lookup_prealloc_tag() iterates over the per-netns &mns->keys list and matches netid, req_tag, peer_addr, and manual_alloc, without checking whether tmp->sk == &msk->sk. This allows any MCTP socket in the same network namespace to use and consume another socket's preallocated tag. Iterate the socket's own tag list (&msk->keys via sklist) instead of the namespace-wide &mns->keys list in mctp_lookup_prealloc_tag(), ensuring that only tags allocated by msk are matched. Tested in QEMU against Linux 7.3.0-rc3 by allocating a manual tag (0x18) on socket A via SIOCMCTPALLOCTAG for peer EID 9 and sending a 4-byte message with MCTP_TAG_PREALLOC from socket B in the same network namespace. On the unfixed kernel, sendto(sock_b) using socket A's preallocated tag succeeds (ret = 4); with this patch applied, sendto(sock_b) fails with -ENOENT (errno = 2) while sendto(sock_a) succeeds (ret = 4). Fixes: 63ed1aab3d40 ("mctp: Add SIOCMCTP{ALLOC,DROP}TAG ioctls for tag control") Suggested-by: Jeremy Kerr Cc: stable@vger.kernel.org Signed-off-by: Hui Peng Link: https://patch.msgid.link/20260921051002.1656692-1-benquike@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit e06f8b6e27cc616d3c587e9784c4a033eb310a16 Author: Zixuan Chai Date: Thu Sep 24 09:26:05 2026 +0800 llc: reserve device headroom for allocated frames commit 72b5b9a28b996e09b8b5b944370c79851bb68f52 upstream. llc_alloc_frame() reserves link-layer headroom using the device type. This is insufficient for stacked Ethernet devices such as VLAN devices, where vlan_dev_hard_header() pushes a VLAN header before the lower device's Ethernet header. An LLC response on such a device can therefore underflow skb headroom in eth_header(). Use LL_RESERVED_SPACE() to account for the device's actual required headroom while preserving the existing LLC device-type check. Fixes: bf9ae5386bca ("llc: use dev_hard_header") Cc: stable@vger.kernel.org Reported-by: VEGA Signed-off-by: Zixuan Chai Signed-off-by: Ren Wei Reviewed-by: Eric Dumazet Link: https://patch.msgid.link/20260924012613.2533934-1-weir@nebusec.ai Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit 58a417a19d80141fff3abe6004e2e47285506e73 Author: Ridham Khurana Date: Tue Sep 22 09:20:58 2026 +0000 gpio: zynq: fix runtime PM leak on request error path commit e9438ab5328a177c9c0e5df87eb92a7162841e98 upstream. pm_runtime_get_sync() leaves the usage counter incremented even when it fails, and zynq_gpio_request() returns the error without dropping it. gpiolib does not call ->free() when ->request() fails, so zynq_gpio_free(), which holds the only matching pm_runtime_put(), never runs. The reference is leaked and the controller can no longer runtime-suspend, so its clock stays enabled. Switch to pm_runtime_resume_and_get(), which only increments the usage counter on success. Fixes: 3242ba117e9b ("gpio: Add driver for Zynq GPIO controller") Cc: stable@vger.kernel.org Signed-off-by: Ridham Khurana Link: https://patch.msgid.link/20260922092102.1053513-1-khurana.ridham222@gmail.com Signed-off-by: Bartosz Golaszewski Signed-off-by: Greg Kroah-Hartman commit cf3d043582e106a0e57002bec8a6de8929ba3e1c Author: Bartosz Golaszewski Date: Tue Sep 22 10:52:28 2026 +0200 gpio: cdev: fix kernel stack leak to user-space in error path commit 1feb5d39b05afd902ed9fc902ec5b15be03a4bdb upstream. If we fail to acquire the GPIO chip guard in gpio_desc_to_lineinfo(), we return immediately before zeroing the info struct we'll end up passing to the user-space later in lineinfo_get_v1(). This may leak the kernel stack contents. Make gpio_desc_to_lineinfo() return int so that the -ENODEV returned on failure to acquire the guard can be propagated to the callers. While not strictly necessary: move the memset() before trying to acquire the SRCU read lock too for good measure. Fixes: d83cee3d2bb1 ("gpio: protect the pointer to gpio_chip in gpio_device with SRCU") Cc: stable@vger.kernel.org Reported-by: Sashiko Closes: https://sashiko.dev/#/patchset/20260912123529.7951-1-tzungbi%40kernel.org?part=3 Reviewed-by: Kent Gibson Link: https://patch.msgid.link/20260922-gpio-cdev-stack-leak-fixes-v3-1-7a0c7a4299d5@oss.qualcomm.com Signed-off-by: Bartosz Golaszewski Signed-off-by: Greg Kroah-Hartman commit a9f738487c11fb512dadef097b56020239c21810 Author: Wentao Liang Date: Wed Sep 16 09:47:01 2026 +0000 gpio: arizona: Fix runtime PM leak in arizona_gpio_direction_out() commit e9d810279f84b30738f7790c0ed15f8dd5b9024a upstream. Switching a persistent GPIO line from input to output acquires a runtime PM reference on the parent device, but if the subsequent regmap_update_bits() fails the reference is never dropped and no later direction_in() can balance it since the direction was never changed. Drop the reference on the update failure path. Fixes: 27a49ed17e22 ("gpio: arizona: Add support for GPIOs that need to be maintained") Cc: stable@vger.kernel.org Signed-off-by: Wentao Liang Reviewed-by: Charles Keepax Link: https://patch.msgid.link/20260916094701.2007509-1-vulab@iscas.ac.cn Signed-off-by: Bartosz Golaszewski Signed-off-by: Greg Kroah-Hartman commit 6fcfc362344edea745a69f3a705de50b731f3418 Author: Hui Peng Date: Mon Sep 21 04:40:25 2026 +0000 ipv6: sr: enforce exact attribute length for SEG6_ATTR_DST commit 2d959c75c27f90e9ec489d18ce5ee6b852ad4741 upstream. In seg6_genl_policy, SEG6_ATTR_DST is defined with .type = NLA_BINARY and .len = sizeof(struct in6_addr). For NLA_BINARY, .len only enforces the maximum payload length and permits shorter payloads (e.g., 0 bytes). When seg6_genl_set_tunsrc() copies sizeof(struct in6_addr) bytes via kmemdup(val, sizeof(*val), GFP_KERNEL), a short SEG6_ATTR_DST attribute triggers a 16-byte out-of-bounds read past skb->tail into uninitialized skb->head memory, which is stored in sdata->tun_src and leaked back to userspace via SEG6_CMD_GET_TUNSRC. Switch SEG6_ATTR_DST in seg6_genl_policy to NLA_POLICY_EXACT_LEN(sizeof(struct in6_addr)) so that generic netlink validation rejects any attribute whose length is not exactly sizeof(struct in6_addr) with -ERANGE. Tested in QEMU against Linux 7.3.0-rc3 by sending a SEG6_CMD_SET_TUNSRC Generic Netlink message with a 0-byte SEG6_ATTR_DST attribute followed by SEG6_CMD_GET_TUNSRC. On the unfixed kernel, SEG6_CMD_SET_TUNSRC succeeds (err = 0) and SEG6_CMD_GET_TUNSRC leaks 16 bytes of uninitialized kernel heap memory (tun_src = 836a61ecc4d25a1042a8d60411cfb378); with this patch applied, SEG6_CMD_SET_TUNSRC is rejected by netlink policy validation with -ERANGE (-34) and tun_src remains zeroed. Fixes: 915d7e5e5930 ("ipv6: sr: add code base for control plane support of SR-IPv6") Cc: stable@vger.kernel.org Signed-off-by: Hui Peng Reviewed-by: Hangbin Liu Reviewed-by: Justin Iurman Reviewed-by: Andrea Mayer Link: https://patch.msgid.link/20260921044025.1535982-1-benquike@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit 1d42828565889632068fbc2983c34d57c66af4e0 Author: Norbert Szetei Date: Wed Sep 16 21:57:53 2026 +0200 ipv6: do not let ipv6_find_hdr() return an offset past the packet end commit ee319bd3a0e976af5087cbe59ebc50a66f31d202 upstream. ipv6_find_hdr() walks the extension header chain, skipping each header by the length that header itself declares. ipv6_optlen() returns up to 2048, and the skip is never checked against skb->len, so the offset stored in *offset can point past the end of the packet. openvswitch installs that offset as the transport header, and update_ipv6_checksum() then reads and writes the transport checksum field out of bounds: BUG: KASAN: slab-use-after-free in inet_proto_csum_replace16+0x445/0x470 Read of size 2 at addr ffff88810b754b06 by task ovs_ipv6_oob/629 CPU: 4 UID: 1000 PID: 629 Comm: ovs_ipv6_oob Tainted: G N 7.3.0-rc3+ #348 Call Trace: inet_proto_csum_replace16+0x445/0x470 set_ipv6_addr+0x3dd/0x460 do_execute_actions+0x6a3d/0x7c40 ovs_execute_actions+0xfd/0x480 ovs_packet_cmd_execute+0xc38/0xf20 genl_rcv_msg+0x59e/0x870 netlink_rcv_skb+0x18b/0x450 genl_rcv+0x2d/0x40 netlink_unicast+0x6bc/0xa20 The buggy address belongs to the object at ffff88810b754980 which belongs to the cache skbuff_small_head of size 704 The buggy address is located 390 bytes inside of freed 704-byte region [ffff88810b754980, ffff88810b754c40) Other callers use that offset too, so bound it here rather than in one caller. Reject a header whose declared length does not fit in the packet. ipv6_find_hdr() already fails with -EBADMSG on a malformed chain, so this adds no new failure mode. Fixes: f8f626754ebe ("ipv6: Move ipv6_find_hdr() out of Netfilter code.") Suggested-by: Ilya Maximets Suggested-by: Eric Dumazet Cc: stable@vger.kernel.org Signed-off-by: Norbert Szetei Reviewed-by: Ido Schimmel Reviewed-by: Ilya Maximets Link: https://patch.msgid.link/8F80BA1A-DDFD-432D-9075-242A3435FEB5@doyensec.com Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit 599d0498a39319285afa908762e0c9fad29f772e Author: Fan Wu Date: Tue Sep 22 20:13:49 2026 -0700 ipe: protect the dm-verity root hash with RCU commit 2776e9c28513a1c855a94792b292cbcc533418c8 upstream. ipe_bdev_setintegrity() frees the old root hash when dm-verity publishes a new one on ->preresume, while policy evaluation can still be dereferencing it. Protect the root hash with RCU. The evaluation path already runs under rcu_read_lock(). Fixes: e155858dd995 ("ipe: add support for dm-verity as a trust provider") Cc: stable@vger.kernel.org Assisted-by: LLM [FW: remove model name according to latest guideline] Signed-off-by: Fan Wu Signed-off-by: Greg Kroah-Hartman commit 48f7d0233c3cc8d986d1177d0e5dd4af435e7265 Author: Fan Wu Date: Tue Sep 22 20:13:48 2026 -0700 ipe: fix use-after-free when auditing a newly loaded policy commit 9814077275eca36ebf8d510d2f076d235ff9f51a upstream. new_policy() audits the policy after ipe_new_policyfs_node() publishes it and drops the new directory's inode lock. A concurrent delete can free the policy while ipe_audit_policy_load() is still using it. Audit the successful load under that lock. Fixes: f44554b5067b ("audit,ipe: add IPE auditing support") Cc: stable@vger.kernel.org Assisted-by: LLM [FW: remove model name according to latest guideline] Signed-off-by: Fan Wu Signed-off-by: Greg Kroah-Hartman commit c2a5e591858629b6c0e284010d2ccdbc22fd0a43 Author: Wentao Liang Date: Thu Sep 17 11:01:35 2026 +0000 fsl/fman: Fix clk reference leak in read_dts_node() commit a644f09b2090ad22a13fbcf9d141084f573108ef upstream. of_clk_get() returns a clock with its reference count incremented, but read_dts_node() only uses it to read the rate and never calls clk_put(). The clock is not stored anywhere, so the reference cannot be released later either. Release the clock once its rate has been read, which also covers the error path taken when the rate is zero. Fixes: 414fd46e7762 ("fsl/fman: Add FMan support") Cc: stable@vger.kernel.org Signed-off-by: Wentao Liang Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260917110135.2148068-1-vulab@iscas.ac.cn Signed-off-by: Paolo Abeni Signed-off-by: Greg Kroah-Hartman commit fd3bf32279df98751364b3609eb100a3e42f253b Author: Josef Bacik Date: Wed Sep 9 18:01:07 2026 +0000 writeback: report a Tasks-RCU quiescent state per cgwb drain pass commit 407a5d205179a4ab186571b0e16ec42725dc77bc upstream. cleanup_offline_cgwbs_workfn() drains a dying cgwb by calling cleanup_offline_cgwb() until it returns false, with a cond_resched() between passes. On a CONFIG_PREEMPTION kernel that cond_resched() does nothing: _cond_resched() is a plain "return 0", and under PREEMPT_DYNAMIC the full and lazy modes disable it. Since commit 7dadeaa6e851 ("sched: Further restrict the preemption modes") those are the only two models on the architectures with PREEMPT_LAZY support, arm64 and x86 among them, so the drain loop never reports a Tasks-RCU quiescent state. A worker draining a cgwb with millions of attached inodes runs for minutes. On a 6.18 arm64 host in lazy mode the cgwb worker drained one dying cgroup's writeback domain for over 11 minutes. A BPF program unlink (bpf_trampoline_unlink_prog -> bpf_trampoline_update -> unregister_ftrace_direct -> ftrace_shutdown -> synchronize_rcu_tasks()) waited on that grace period while holding the trampoline mutex, 42 tasks queued behind it in D state, and the hung task detector fired at 614 s and panicked the host. Any BPF or ftrace detach during a long drain inherits the drain's length. Fix this by calling cond_resched_tasks_rcu_qs() so we do not stall out anybody who calls sycnrhonize_rcu_tasks(). We put this in a do { } while loop because if we have many small cgroups cleanup_offline_cgwb() will return false and we will never call cond_resched_tasks_rcu_qs(), creating the same problem. Link: https://lore.kernel.org/20260909-cgwb-tasks-rcu-qs-v1-1-967a7754771f@toxicpanda.com Fixes: c22d70a162d3 ("writeback, cgroup: release dying cgwbs by switching attached inodes") Signed-off-by: Josef Bacik Signed-off-by: Andrew Morton Link: https://lore.kernel.org/bpf/9d444098-7c03-4163-af12-bd0a79a51443@paulmck-laptop/ Assisted-by: LLM Acked-by: Tejun Heo Reviewed-by: Roman Gushchin Reviewed-by: Jan Kara Acked-by: Lorenzo Stoakes (ARM) Cc: David Hildenbrand Cc: Dennis Zhou Cc: Liam R. Howlett Cc: Matthew Wilcox (Oracle) Cc: Michal Hocko Cc: Mike Rapoport Cc: "Paul E . McKenney" Cc: Suren Baghdasaryan Cc: Vlastimil Babka Cc: Signed-off-by: Greg Kroah-Hartman commit 1dccf5dc2f8fe688270b261ca4216e03c6a1e674 Author: Pavankumar Kondeti Date: Fri Sep 25 15:25:12 2026 +0530 workqueue: Fix NULL current_pwq deref in flush dependency check commit db6365ced4d5855e321f772b240c0e473bcfcdd5 upstream. check_flush_dependency() uses current_wq_worker() to determine whether the caller is a workqueue worker and then dereferences worker->current_pwq to test whether the current workqueue is WQ_MEM_RECLAIM. current_wq_worker() only means that %current has PF_WQ_WORKER set. A kworker can reach check_flush_dependency() while it is not executing a work item. One such path is worker_thread() acting as the pool manager, where create_worker() does GFP_KERNEL allocation and the allocation path invokes the OOM notifier. In that state worker->current_pwq is NULL because current_pwq is set only by process_one_work() and cleared again after the work function returns. [ 416.760634][ T375] Call trace: [ 416.760638][ T375] check_flush_dependency+0x80/0x120 (P) [ 416.760648][ T375] __flush_work+0x98/0x224 [ 416.760657][ T375] flush_work+0x30/0x44 [ 416.760665][ T375] ... [ 416.760710][ T375] blocking_notifier_call_chain+0x58/0xa0 [ 416.760719][ T375] out_of_memory+0xb4/0x458 [ 416.760730][ T375] __alloc_pages_may_oom+0x11c/0x1a8 [ 416.760739][ T375] __alloc_pages_slowpath+0x314/0x46c [ 416.760746][ T375] __alloc_frozen_pages_noprof+0x110/0x1a4 [ 416.760753][ T375] new_slab+0x12c/0x484 [ 416.760759][ T375] ___slab_alloc+0x7a8/0xc7c [ 416.760765][ T375] __slab_alloc+0x74/0xd8 [ 416.760772][ T375] __kmalloc_cache_node_noprof+0x2ac/0x304 [ 416.760779][ T375] alloc_worker+0x28/0x60 [ 416.760785][ T375] create_worker+0x4c/0x20c [ 416.760790][ T375] worker_thread+0xe8/0x2b8 [ 416.760796][ T375] kthread+0x1a8/0x200 [ 416.760805][ T375] ret_from_fork+0x10/0x20 Guard the WQ_MEM_RECLAIM-worker warning with worker->current_pwq. If the kworker is not currently executing a work item, there is no current workqueue to diagnose with that warning. The PF_MEMALLOC warning is left unchanged so explicit reclaim context flushing a !WQ_MEM_RECLAIM target is still reported. Fixes: fca839c00a12 ("workqueue: warn if memory reclaim tries to flush !WQ_MEM_RECLAIM workqueue") Cc: stable@vger.kernel.org Assisted-by: LLM Signed-off-by: Pavankumar Kondeti Signed-off-by: Tejun Heo Signed-off-by: Greg Kroah-Hartman commit dd4b92067238771b604b59fdda8efed6d70ee823 Author: Patrick Lu (Anthropic) Date: Fri Sep 11 18:49:49 2026 +0000 writeback: bound cleanup_offline_cgwb() rescans by rotating scanned inodes commit f6988c90671e83db79df1b7b9d6fdb0e5947fd84 upstream. cleanup_offline_cgwb() prepares at most WB_MAX_INODES_PER_ISW inodes per call and is called again until the dying wb is drained, but every call walks wb->b_attached and then wb->b_dirty_time from the same end. Inodes already prepared (they stay on the list with I_WB_SWITCH set until the switch worker runs) and inodes that cannot be switched (I_FREEING, I_WILL_FREE, !SB_ACTIVE, DAX, already on the target wb) stay where they are, so each pass rescans a growing run of them under wb->list_lock and a full drain is quadratic in the number of inodes on the list. With ~17M inodes attached to one dying cgwb we saw this end in soft lockups, with CPUs reported stuck for 21-48s. Walk both lists from the oldest end and move every scanned inode to the newest end, so the next pass starts where the previous one stopped and the drain becomes linear. b_attached is unordered, so nobody sees the reorder there. b_dirty_time is ordered by dirtied_when, but the oldest unscanned inode stays at the end move_expired_inodes() picks from, sync takes the whole list regardless of order, and prepared inodes leave the list as soon as the switch work runs and get a new dirtied_time_when on the new wb anyway, so the only inodes left out of order are the ones that can never switch (DAX), and only on the dying wb. Fixes: c22d70a162d3 ("writeback, cgroup: release dying cgwbs by switching attached inodes") Cc: stable@vger.kernel.org Acked-by: Tejun Heo Acked-by: Roman Gushchin Signed-off-by: Patrick Lu (Anthropic) Link: https://patch.msgid.link/20260911-wb-cgwb-rotate-v2-1-a9ab253a1295@gmail.com Reviewed-by: Jan Kara Signed-off-by: Christian Brauner (Amutable) Signed-off-by: Greg Kroah-Hartman commit 37ec7353f4ac495b28284b0fcb0e53d21d1b3f3f Author: Christian Brauner Date: Wed Sep 9 11:03:18 2026 +0200 fs/ntfs3: use d_instantiate_new() in ntfs_create_inode() and murder syzbot's "WARNING in do_new_mount" saga commit 1abd643f3783ea8f8e273c18697ff0413aa92dc7 upstream. ntfs_create_inode() creates a new inode via ntfs_new_inode(). It hashes it with insert_inode_locked() and so it's marked as I_NEW until unlock_new_inode(). ntfs 3 calls d_instantiate() in between though... Since the dentry was already hashed by the lookup before the create any path walk finds it without touching the parent's i_rwsem and so can lock the inode. If the inode is a directory unlock_new_inode() calls lockdep_annotate_inode_mutex_key() and marks i_rwsem with the i_mutex_dir_key class. That resets the count and the owner of a lock somebody else may already hold by now... syzbot has been spamming us with the same godforsaken bug "WARNING in do_new_mount" since 2023. I can't take it anymore so I went looking. Afaict, syzbot's executor chdirs into a freshly mounted ntfs3 image, creates a directory and then mounts some pseudofs on it. Everytime the mkdir() takes longer than syzbot waits mount() runs concurrently: mkdir("./sys") mount(NULL, "./sys", "sysfs") ntfs_create_inode() d_instantiate() user_path_at() finds the dentry do_lock_mount() inode_lock(inode) namespace_lock() unlock_new_inode() lockdep_annotate_inode_mutex_key() init_rwsem(&inode->i_rwsem) unlock_mount() inode_unlock(inode) The mount side then releases a lock that according to the rwsem nobody holds: DEBUG_RWSEMS_WARN_ON((rwsem_owner(sem) != current) && ...): count = 0x0, magic = 0xffff888043a854e8, owner = 0x0, curr 0xffff888000244880, list empty WARNING: CPU: 0 PID: 5346 at kernel/locking/rwsem.c:1368 __up_write Call Trace: inode_unlock include/linux/fs.h:877 [inline] unlock_mount fs/namespace.c:2892 [inline] do_new_mount_fc fs/namespace.c:3828 [inline] do_new_mount+0x777/0xa40 fs/namespace.c:3887 On PREEMPT_RT the same thing shows up as DEBUG_LOCKS_WARN_ON(rt_mutex_owner(lock) != current) WARNING: kernel/locking/rtmutex_common.h:193 at rt_mutex_slowunlock The up_write() underflows the reset count. A following inode_lock() on that directory then never returns. A path walk into the new directory racing with the mkdir() corrupts the lock the same way via inode_lock_shared() in lookup_slow(). Switch to d_instantiate_new() and drop the trailing unlock_new_inode(). All error paths bail out before that point with I_NEW still set and keep using discard_new_inode(). May we never see this fscking bug report again. Link: https://patch.msgid.link/20260909-work-ntfs3-d_instantiate_new-v1-1-2db697162ce8@kernel.org Fixes: 82cae269cfa9 ("fs/ntfs3: Add initialization of super block") Reviewed-by: Jan Kara Cc: stable@vger.kernel.org # v5.15+ Reported-by: syzbot+2a13ad6914e6fcec716c@syzkaller.appspotmail.com Closes: https://lore.kernel.org/6a9beced.a5e650b3.26d8a.000b.GAE@google.com Signed-off-by: Christian Brauner (Amutable) Signed-off-by: Greg Kroah-Hartman commit b1229b210ef5a39c58be665e6798203751411842 Author: Hui Peng Date: Mon Sep 21 04:59:20 2026 +0000 fou: reject omitted FOU_ATTR_IPPROTO on FOU_ENCAP_DIRECT commit d22609f3d13fc5baacd92c222731b03c593401db upstream. Commit 7a9bc9e3f423 ("fou: Don't allow 0 for FOU_ATTR_IPPROTO.") added NLA_POLICY_MIN(NLA_U8, 1) to fou_nl_policy[FOU_ATTR_IPPROTO], which rejects an explicitly supplied FOU_ATTR_IPPROTO == 0 attribute with -ERANGE. However, FOU_ATTR_IPPROTO is an optional netlink attribute. When a user sends FOU_CMD_ADD with FOU_ATTR_TYPE set to FOU_ENCAP_DIRECT and omits FOU_ATTR_IPPROTO entirely, nla_policy validation succeeds and parse_nl_config() leaves cfg->protocol as 0 (from memset(cfg, 0, sizeof(*cfg))). fou_create() then creates a FOU_ENCAP_DIRECT socket with fou->protocol == 0. In fou_udp_recv(), returning -fou->protocol to udp_queue_rcv_one_skb() triggers IP protocol resubmission when fou->protocol > 0, whereas returning 0 tells the UDP tunnel layer that the skb was consumed without freeing it. When fou->protocol == 0, every packet received on the socket returns 0 from fou_udp_recv() and leaks the sk_buff. Reject FOU_ENCAP_DIRECT when !cfg->protocol in fou_create() so that creating a direct encapsulation port without FOU_ATTR_IPPROTO fails with -EINVAL while leaving FOU_CMD_DEL and FOU_CMD_GET (which share parse_nl_config()) unaffected. Tested in QEMU against Linux 7.3.0-rc3 by sending a FOU_CMD_ADD Generic Netlink request with FOU_ATTR_PORT = 5555 and FOU_ATTR_TYPE = FOU_ENCAP_DIRECT while omitting FOU_ATTR_IPPROTO. On the unfixed kernel, FOU_CMD_ADD succeeds (err = 0), FOU_CMD_GET reports fou->type = 1 and fou->protocol = 0, and sending 4000 UDP packets to 127.0.0.1:5555 leaks all 4000 sk_buffs (SUnreclaim in /proc/meminfo grows from 41456 kB to 59008 kB, +17552 kB); with this patch applied, FOU_CMD_ADD is rejected with -EINVAL (-22). Fixes: 23461551c006 ("fou: Support for foo-over-udp RX path") Fixes: 7a9bc9e3f423 ("fou: Don't allow 0 for FOU_ATTR_IPPROTO.") Cc: stable@vger.kernel.org Signed-off-by: Hui Peng Reviewed-by: Hangbin Liu Link: https://patch.msgid.link/20260921045920.1613098-1-benquike@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit 2f7042429a8e939d5c8770b1b9c050da219310d7 Author: Guopeng Zhang Date: Thu Sep 24 17:15:36 2026 +0800 cgroup/pids: Restore pids.events notifications in local mode commit 1765a153d985c231357145e26798f9408db10e42 upstream. A fork rejected by the pids controller increments the counter reported by pids.events. When local event accounting is selected, however, pids_event() returns after notifying only events_local_file, leaving pids.events pollers asleep. On legacy hierarchies, pids.events.local does not exist. With pids_localevents, pids.events reports the same local counter. In both cases, pids.events changes without generating a notification. This can be reproduced with a pids_localevents mount: mkdir /tmp/test mount -t cgroup2 -o pids_localevents none /tmp/test mkdir /tmp/test/t echo 1 > /tmp/test/t/pids.max cat /tmp/test/t/pids.events # max 0 timeout 3 inotifywait -e modify /tmp/test/t/pids.events & sh -c 'echo $$ > /tmp/test/t/cgroup.procs; (true &)' 2>/dev/null wait cat /tmp/test/t/pids.events # max 1 Without this patch, inotifywait times out without reporting an event. Notify pids.events before returning from the local event path. Fixes: 3f26a885a068 ("cgroup/pids: Add pids.events.local") Cc: stable@vger.kernel.org # v6.11+ Signed-off-by: Guopeng Zhang Signed-off-by: Tejun Heo Signed-off-by: Greg Kroah-Hartman commit caf8b2d9b30c3bdb9fd7ea96cd5236da1f5580e4 Author: Holger Dengler Date: Fri Aug 14 16:20:10 2026 +0200 crypto: s390/hmac - Generate intermediate CV for API partial block handling commit 10396a2d6d41d594975b6ece712278570c3c970c upstream. The API partial block handling requires a intermediate chaining value (CV). The internal function hash_data() sets the function code correctly, so also call cpacf_kimd() instruction for intermediate CV generation, as cpacf_klmd() always generate the final hash value. Cc: stable@vger.kernel.org # 6.15+ Fixes: 08811169ac01 ("crypto: s390/hmac - Use API partial block handling") Signed-off-by: Holger Dengler Reviewed-by: Harald Freudenberger Signed-off-by: Herbert Xu Signed-off-by: Greg Kroah-Hartman commit 7504b677c4370c9e4b153c553ff8831c093b1ff6 Author: Andrea Parri Date: Tue Sep 22 16:55:30 2026 +0200 bpf: fs/xattr: don't assume the inode is locked in path_unlink/path_rmdir commit 35d442ed1f86465e49df3119fb898f985186db13 upstream. bpf_lsm_has_d_inode_locked() makes the verifier rewrite bpf_[set|remove]_dentry_xattr() to the _locked variants, which assume that the caller already holds the inode's i_rwsem. The path_unlink and path_rmdir hooks are listed, but security_path_unlink() and security_path_rmdir() run before vfs_unlink()/vfs_rmdir() take the victim inode's i_rwsem, so a sleepable BPF LSM program attached to either hook mutates the victim's xattrs without the lock held. Drop the two path hooks from d_inode_locked_hooks so that the verifier keeps the locking bpf_[set|remove]_dentry_xattr() variants, which take the lock themselves. Fixes: 56467292794b8 ("bpf: fs/xattr: Add BPF kfuncs to set and remove xattrs") Cc: stable@vger.kernel.org Signed-off-by: Andrea Parri Link: https://patch.msgid.link/20260922145530.369775-1-parri.andrea@gmail.com Signed-off-by: Christian Brauner (Amutable) Signed-off-by: Greg Kroah-Hartman commit 4d97a141e85a034566c33c7ed854c8a3e8a268f1 Author: Myeonghun Pak Date: Mon Sep 21 21:46:05 2026 -0400 bna: prevent IOC timer rearm during teardown commit 77b1718e39e5c9f6956fb60807326af20baf889d upstream. bna: prevent IOC timer rearm during teardown bnad_pci_remove() and the probe disable_ioceth path call timer_delete_sync() for ioc_timer, sem_timer and hb_timer, but not for iocpf_timer. bnad_iocpf_timeout() then takes bnad->bna_lock after free_netdev() has freed the struct bnad. Deleting iocpf_timer last does not fix this. sem_timer and iocpf_timer rearm each other: bnad_iocpf_sem_timeout() can arm iocpf_timer, and bnad_iocpf_timeout() arms sem_timer from bfa_ioc_hw_sem_get() when the semaphore is busy. timer_delete_sync() only waits out its own callback. bnad_ioceth_disable() can time out and leave that callback live. Shut all four IOC timers down with timer_shutdown_sync() on both paths, so a later mod_timer() is ignored. This issue was identified during our ongoing static-analysis research while reviewing kernel code. Fixes: 1d32f7696286 ("bna: IOC failure auto recovery fix") Cc: stable@vger.kernel.org Assisted-by: LLM Co-developed-by: Ijae Kim Signed-off-by: Ijae Kim Signed-off-by: Myeonghun Pak Link: https://patch.msgid.link/20260922014605.588040-1-mhun512@gmail.com Signed-off-by: Paolo Abeni Signed-off-by: Greg Kroah-Hartman commit 1603abb0c50420e97cf4a1feddbd5e183c5fdf86 Author: Matthias Goergens Date: Thu Sep 24 01:52:03 2026 +0800 ata: libata-scsi: bound the ATA passthru sense descriptor writes commit 80320b278fea07ffcda3f57b67b61658e0a4e1ca upstream. When an ATA PASS-THROUGH command to an ATAPI device fails, the sense buffer holds the device's REQUEST SENSE reply, and ata_scsi_set_passthru_sense_fields() trusts its additional length byte, sb[7], when adding the ATA Status Return descriptor. A faulty or malicious device can use that to make the kernel read and write past the 96-byte buffer in three ways: - scsi_sense_desc_find() is passed sb[7] + 8 as the buffer length, so its clamp against sb[7] does nothing and the walk runs off the end. - A type-9 descriptor found near the end is filled in unchecked. - A new descriptor at sb[8 + len] needs len + 22 bytes, not len + 14, so len 75..82 writes up to 8 bytes past the end. Reproduced with KASAN under qemu, with the emulated ATAPI REQUEST SENSE reply patched: BUG: KASAN: slab-out-of-bounds in scsi_sense_desc_find+0x1a5/0x210 BUG: KASAN: slab-out-of-bounds in ata_scsi_qc_complete+0x1a15/0x1a50 Both are gone with this patch, and a valid descriptor is still filled in. Fixes: 97981926224a ("ata: libata-scsi: Do not overwrite valid sense data when CK_COND=1") Cc: stable@vger.kernel.org Reviewed-by: Damien Le Moal Signed-off-by: Matthias Goergens Link: https://lore.kernel.org/r/20260923175203.1576825-1-matthias.goergens@gmail.com Signed-off-by: Niklas Cassel Signed-off-by: Greg Kroah-Hartman commit cdccb760be1f129ae2774ec924e3a85d6dd61c40 Author: David Carlier Date: Sun Sep 6 13:14:16 2026 +0100 arm64: errata: match the target implementation CPU's own MIDR commit b7403afb7a5f85073243df10238b3483958ad69e upstream. __is_affected_midr_range() is handed the MIDR and REVIDR of one target implementation CPU, but tests the erratum's range with is_midr_in_range(), which re-scans all of target_impl_cpus[] and ignores the @midr argument. The range test is thus constant across the per-CPU loop in is_affected_midr_range() and only answers "is any target CPU in range". Since just the fixed_revs REVIDR check uses the iteration's own registers, an out-of-range target CPU can decide whether a MIDR_FIXED() exemption applies. A VM then enables a workaround whose only in-range CPU is fixed silicon, e.g. erratum 2658417 on a Cortex-A510 r1p1 with REVIDR_EL1[25] set. Factor the range test into __is_midr_in_range(), which takes an explicit MIDR, and use it in __is_affected_midr_range(). Fixes: 86edf6bdcf05 ("smccc/kvm_guest: Enable errata based on implementation CPUs") Cc: stable@vger.kernel.org Assisted-by: Claude:claude-opus-5 Signed-off-by: David Carlier Signed-off-by: Will Deacon Signed-off-by: Greg Kroah-Hartman commit 071f51188a5d8ec10535541255587976d12fc7d8 Author: Fuad Tabba Date: Tue Sep 22 19:14:30 2026 +0100 arm64/boot: Disable trapping of PMZR_EL0 writes to EL2 commit 2bc6b218717b9d08f466f88209251d54bc09b207 upstream. __init_el2_fgt2() writes one mask to both HDFGRTR2_EL2 and HDFGWTR2_EL2. PMZR_EL0 is write-only, so its trap bit, nPMZR_EL0, exists only in HDFGWTR2_EL2 and is therefore never set: a PMZR_EL0 write from the host traps to EL2, where the nVHE hypervisor has no handler and BUG()s. The kernel never writes PMZR_EL0, but kernel.perf_user_access=1 has the PMU driver set PMUSERENR_EL0.UEN for a task with a user-read event, so a write from EL0 reaches the trap and takes the host down without a panic message. Accumulate the HDFGWTR2_EL2 bits separately, as __init_el2_fgt() already does for HDFGWTR_EL2, and set nPMZR_EL0 with the other FEAT_PMUv3p9 bits. Fixes: 858c7bfcb35e1 ("arm64/boot: Enable EL2 requirements for FEAT_PMUv3p9") Cc: stable@vger.kernel.org Signed-off-by: Fuad Tabba Reviewed-by: Anshuman Khandual Reviewed-by: Oliver Upton Signed-off-by: Will Deacon Signed-off-by: Greg Kroah-Hartman commit 6279211e7d2456ecd195d236c9ee5e5c6fa5a17a Author: Dairui Zhang Date: Wed Sep 23 13:01:01 2026 +0800 af_packet: fix integer overflow in prb_calc_retire_blk_tmo() commit 56d82862a0a243ac14ba11b6d7b57ddc2d064b95 upstream. prb_calc_retire_blk_tmo() computes in 32-bit int arithmetic: mbits = (blk_size_in_bytes * 8) / (1024 * 1024); If I'm reading the validation right, tp_block_size is user controlled and packet_set_ring() only rejects values that are <= 0 as int or not page aligned, so a 256MiB block goes right through (and alloc_one_pg_vec_page() even has a vzalloc fallback for it). 0x10000000 * 8 wraps to INT_MIN, and on a NIC reporting 1 Gbps (div == 1) the function ends up returning -2047. The condition is actually (8 * size) mod 2^32 >= 2^31 && div == 1, so the trigger set is [256,512), [768,1024), [1280,1536) and [1792,2048) MiB. Other sizes wrap to non-negative values and faster links divide the unsigned value back below 2^31, which is why this doesn't blow up for everyone. What makes it fatal is what happens next in init_prb_bdqc(): p1->interval_ktime = ms_to_ktime(prb_calc_retire_blk_tmo(...)); hrtimer_start(&p1->retire_blk_timer, p1->interval_ktime, HRTIMER_MODE_REL_SOFT); A negative relative timeout expires immediately. The callback unconditionally returns HRTIMER_RESTART, and hrtimer_forward() turns the negative interval into hrtimer_resolution: if (interval < hrtimer_resolution) interval = hrtimer_resolution; So the SOFT timer re-fires at the maximum rate forever, holding sk_receive_queue.lock each pass. One CPU spins in softirq until the socket is closed. Repeat with more rings and the machine is gone. The overflow itself is ancient - it was introduced together with TPACKET_V3 in f6fb8f100b80 ("af-packet: TPACKET_V3 flexible buffer implementation."). Its effect prior to f7460d2989fa ("net: af_packet: Use hrtimer to do the retire operation", v6.18) was not as clear-cut, though: the return value was stored into an unsigned short retire_blk_tov, so a negative result was truncated, and a 0-jiffy delay loop could be programmed as well. Neither is nearly as detrimental as the immediate maximum-rate spin the hrtimer conversion turned it into. (Unrelated to CVE-2019-20812 - that one was the ethtool failure path returning 0, which now returns DEFAULT_PRB_RETIRE_TOV.) Reproducer, needs CAP_NET_RAW (a --network host container has it by default) and a 1 Gbps NIC (QEMU e1000 works): int fd = socket(AF_PACKET, SOCK_RAW, htons(ETH_P_ALL)); bind(fd, ...); int v = TPACKET_V3; setsockopt(fd, SOL_PACKET, PACKET_VERSION, &v, sizeof(v)); struct tpacket_req3 req = { .tp_block_size = 0x10000000, .tp_block_nr = 1, .tp_frame_size = 2048, .tp_frame_nr = 0x10000000 / 2048, .tp_retire_blk_tov = 0, }; setsockopt(fd, SOL_PACKET, PACKET_RX_RING, &req, sizeof(req)); Compute in 64 bits instead. The operands are already bounded by the existing validation, so nothing else changes. If you'd prefer a different fix, just say so and I'll respin. Fixes: f6fb8f100b80 ("af-packet: TPACKET_V3 flexible buffer implementation.") Cc: stable@vger.kernel.org Signed-off-by: Dairui Zhang Reviewed-by: Willem de Bruijn Link: https://patch.msgid.link/20260923050101.1510064-1-zhangdairui@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit ec23855274a73906b02f0d124fefd9cb3c008b1d Author: Aohan Mei Date: Mon Sep 21 17:37:04 2026 +0800 sctp: discard the rest of the packet on a stale-cookie error commit 4498467a8af06cfa3d71cb04bd7c4170dec8f449 upstream. When an association is in COOKIE-ECHOED state and the peer sends a bundled [ERROR(Stale Cookie)][DATA] packet from one of its non-primary addresses, processing the ERROR chunk takes the non-fatal stale-cookie retry path sctp_sf_do_5_2_6_stale(), which queues SCTP_CMD_DEL_NON_PRIMARY while keeping the association alive. sctp_cmd_del_non_primary() removes every non-primary transport - including the very transport this packet arrived on, which is still referenced by the receive lookup and shared by all chunks of the packet via chunk->transport. sctp_assoc_rm_peer() does redirect asoc->peer.last_data_from away from the removed transport, but right afterwards the bundled DATA chunk makes sctp_assoc_bh_rcv() re-register asoc->peer.last_data_from = chunk->transport unconditionally, undoing the redirection with the just-removed transport. Once the packet is done, the receive reference is dropped and the transport is RCU-freed, while the surviving association keeps the dangling last_data_from. A later FWD-TSN (or the delayed SACK timer) makes sctp_gen_sack() dereference it (->param_flags and friends), and sctp_make_sack()/sctp_outq_select_transport() may write to the freed object and link it into the live transport list. This is a use-after-free triggerable by any malicious SCTP peer (or a local unprivileged user acting as one) with no capabilities required: BUG: KASAN: slab-use-after-free in sctp_do_sm+0x498a/0x5660 Read of size 4 at addr ffff88800e1e356c by task poc/115 Call Trace: sctp_do_sm <- sctp_assoc_bh_rcv <- sctp_inq_push <- sctp_rcv <- ip_protocol_deliver_rcu <- ip_rcv Allocated: sctp_transport_new <- sctp_assoc_add_peer <- sctp_process_init (INIT-ACK processing) Freed: kfree <- sctp_transport_destroy_rcu <- rcu_core (call_rcu queued by sctp_transport_put at end of sctp_rcv) The buggy address is located 364 bytes inside of freed 1024-byte region [ffff88800e1e3400, ffff88800e1e3800), cache kmalloc-1k Note that commit 03a9d10ecf71 ("sctp: drop a chunk if its transport was removed") only covers the window between the receive lookup and the chunk processing (e.g. an ASCONF DEL-IP racing the socket backlog); here the transport is removed *while* the packet is being processed, by an earlier chunk of the same packet, so the drop in sctp_inq_push() does not reach this path. Verified with the bundled [ERROR(Stale Cookie)][DATA] + FWD-TSN reproducer: the KASAN report above still fires with that commit applied, and is gone with this patch on top. Fix it by discarding the rest of the packet on this path, as suggested by Xin. After the stale-cookie ERROR has sent the association back to COOKIE-WAIT and removed the non-primary transports, the remaining chunks of the packet can only run against the restarted handshake while referencing the removed arrival transport through chunk->transport: besides the last_data_from registration above, sctp_cmd_setup_t2() and the sctp_make_*() reply builders would also copy that pointer into association-lifetime state that sctp_assoc_rm_peer() has already sanitized. Let the peer retransmit them, in line with what sctp_inq_push() does for chunks whose transport was removed before processing. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Suggested-by: Xin Long Reported-by: TencentOS Corvus AI Cc: stable@vger.kernel.org Signed-off-by: Aohan Mei Acked-by: Xin Long Link: https://patch.msgid.link/20260921093707.1432184-1-ljp1205831794@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit 7ba87509c54cd519e7ed362e0c85e8431f066a18 Author: Willem de Bruijn Date: Thu Sep 24 11:44:12 2026 -0400 tcp: prevent collapsing skbs across boundary in rtx queue commit fc6d80eb504458d6416b75a94188b268c95c6533 upstream. tcp_write_collapse_fence() sets TCP_SKB_CB(skb)->eor = 1 on tcp_write_queue_tail(sk) to prevent skbs queued after a switch to device encryption from being collapsed into earlier skbs. The fence is a no-op if all earlier data has already been transmitted when the switch happens: sk->sk_write_queue is empty. The not yet acknowledged earlier skbs wait in sk->tcp_rtx_queue with eor 0. On a subsequent retransmit or SACK shift, tcp_retrans_try_collapse() or tcp_shift_skb_data() can then merge an skb queued after the switch into one queued before it. Both users of the fence are affected: - psp: devices only encrypt skbs with skb->decrypted set. The merged skb keeps decrypted = 0 from the earlier skb, so merged data sent after psp_sock_assoc_set_tx() is retransmitted in cleartext. - tls device offload: the merged skb straddles the start marker set in tls_set_device_offload(). The software fallback (fill_sg_in() returns -EINVAL) and the mlx5, nfp and funeth drivers cannot handle such an skb and drop it. Every retransmit rebuilds the same skb, so the connection stalls. Fix this in two places, for defense in depth: 1. Fall back to tcp_rtx_queue_tail(sk) in tcp_write_collapse_fence() when tcp_write_queue_tail(sk) is NULL. 2. Check !skb_cmp_decrypted(to, from) in tcp_skb_can_collapse(), as tcp_skb_can_collapse_rx() does on receive. skb_shift(), which both collapse paths call, already has a DEBUG_NET_WARN_ON_ONCE() for this condition. Fixes: e8f69799810c ("net/tls: Add generic NIC offload infrastructure") Cc: stable@vger.kernel.org Signed-off-by: Willem de Bruijn Reviewed-by: Eric Dumazet Reviewed-by: Daniel Zahka Link: https://patch.msgid.link/20260924154427.953800-1-willemdebruijn.kernel@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit 12cc3a1eac5a2a06b67a7b85f0f48ff241a94d1b Author: Eric Dumazet Date: Sun Sep 13 04:42:33 2026 +0000 tipc: reject invalid and unexpected GRP_ACK_MSG to prevent bc_ackers underflow commit 99cc2a62e07a44a22254d7beca9ef1f8ad886d0d upstream. Commit 48a5fe38772b ("tipc: fix bc_ackers underflow on duplicate GRP_ACK_MSG") rejected duplicate/stale ACKs in tipc_group_proto_rcv() by returning early when less_eq(acked, m->bc_acked). However, that check remains incomplete in two ways: 1. When grp->bc_ackers is zero (e.g. on a quiet group, when replicast ACKs were not requested, or after all expected members have already acknowledged), an unexpected GRP_ACK_MSG with acked > m->bc_acked passes less_eq() and unconditionally decrements grp->bc_ackers. Because bc_ackers is a u16, this wraps to 65535, causing tipc_group_bc_cong() to permanently report congestion and blocking all future group broadcasts on the socket. 2. During an active broadcast round (grp->bc_ackers > 0), the sender transmits packet S and advances grp->bc_snd_nxt to S + 1. Receivers increment their expected counter to S + 1 upon consuming packet S, so the only valid ACK value for the current round is strictly acked == grp->bc_snd_nxt. However, tipc_group_update_bc_members() initializes each member's m->bc_acked to prev = grp->bc_snd_nxt - 1 (S - 1 before increment). This leaves a 2-sequence gap (S - 1 to S + 1) in sequence space. An incoming ACK is therefore neither rejected as duplicate nor prevented from decrementing grp->bc_ackers if an unexpected or stale value (such as S) is received. A member sending acked = S followed by acked = S + 1 could decrement grp->bc_ackers twice in the same round, prematurely clearing bc_ackers or underflowing it. Fix this by: - Dropping GRP_ACK_MSG immediately if grp->bc_ackers is zero. - Requiring acked == grp->bc_snd_nxt and rejecting duplicates where m->bc_acked == acked. Because replicast broadcast rounds are strictly sequential, only grp->bc_snd_nxt can be acknowledged, and each member can acknowledge at most once per round. Note that a related pre-existing issue in tipc_group_delete_member() (where grp->bc_ackers decrementing to zero upon member departure does not restore *grp->open or trigger a socket wakeup) will be addressed in a separate patch. Fixes: 48a5fe38772b ("tipc: fix bc_ackers underflow on duplicate GRP_ACK_MSG") Fixes: 2f487712b893 ("tipc: guarantee that group broadcast doesn't bypass group unicast") Reported-by: James Burton Cc: stable@vger.kernel.org Signed-off-by: Eric Dumazet Link: https://patch.msgid.link/20260913044233.193927-1-edumazet@google.com Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit ca40f4df4ab2f7840aff5992e0b6fd323329f515 Author: Willem de Bruijn Date: Fri Sep 18 20:47:28 2026 -0400 virtio_net: copy zerocopy frags in start_xmit without NAPI commit 07e1a9408b6c2f9d0cfb757b67dabb52da7a32b2 upstream. Virtio-net without NAPI frees completed skbs lazily on the next start_xmit. Senders waiting for in-flight zerocopy buffers can deadlock if they cannot transmit more packets, as then no completed packets will be freed. When !use_napi, virtio-net already calls skb_orphan to avoid waiting up for transmitted skbs to be freed. For zerocopy packets that require deep copying on orphan (i.e. those that do not set SKBFL_DONT_ORPHAN, such as PACKET_TX_RING), call skb_orphan_frags before orphaning to release the buffers. This fixes the tpacket_snd slot reuse bug on skb_orphan for virtio-net, and prevents PACKET_TX_RING from running out of slots. This fix also touches vhost_net zerocopy packets, which also do not set SKBFL_DONT_ORPHAN. This is fine: vhost_net packets only encounter virtio-net in nested virtualization, and only if napi_tx is explicitly disabled (it has been default-enabled since Linux 4.12). In that rare case, copying the frags is desirable anyway to prevent holding guest descriptors pinned across unbounded intervals. This is a prerequisite for the next patch, which converts PACKET_TX_RING to standard zerocopy completion. Without this patch first, a bounded ring sender can stall indefinitely behind a virtio-net virtqueue that cannot reclaim. Fixes: 5cd8d46ea156 ("packet: copy user buffers before orphan or clone") Cc: stable@vger.kernel.org Cc: mst@redhat.com Cc: jasowangio@gmail.com Signed-off-by: Willem de Bruijn Link: https://patch.msgid.link/20260919004748.1463985-2-willemdebruijn.kernel@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Greg Kroah-Hartman commit 33da1ff56f955fba4c705785db02d6bd4f99341c Author: Mario Limonciello Date: Tue Sep 8 14:05:59 2026 -0500 x86/PCI: Disable enhanced atomics on AMD NBIO 7.7 and 7.11 commit 4fde448225123442c5796f54b7a4400e2d3cbaf6 upstream. Multiple users report data corruption during 64-bit DMA transfers on systems with AMD NBIO 7.7 and 7.11 controllers. This occurs when BIOS enables AMD "enhanced atomic operations" on PCIe Root Ports. When enhanced atomics are enabled, any 64-bit DMA access may be corrupted. Disable enhanced atomics using SMN for NBIO 7.7 and 7.11 based models. Reported-by: Mikael Etienne Closes: https://lore.kernel.org/178789300872.392066.15963676631650361573@gmail.com/ Reported-by: Arthur Husband Closes: https://lore.kernel.org/20260406222335.379935-1-artmoty@gmail.com/ Reported-by: Alvin Lim Closes: https://lore.kernel.org/20260621100844.1224301-1-alvinwylim@gmail.com/ Signed-off-by: Mario Limonciello [bhelgaas: commit log, s/IOVA/DMA/ in comment] Signed-off-by: Bjorn Helgaas Cc: stable@vger.kernel.org Cc: David Laight Cc: John Smith Cc: Lennert Buytenhek Cc: Niklas Cassel Cc: Roland Waltersson Link: https://patch.msgid.link/20260908190600.226485-2-mario.limonciello@amd.com Signed-off-by: Greg Kroah-Hartman commit f50d2e8cc1f1dceb437c402fa3a80e0d6da3f521 Author: Masami Hiramatsu (Google) Date: Tue Sep 22 13:24:55 2026 +0900 x86/mce: Fix hardware debug register corruption on task migration commit b8d1d5b63a8ef532038eebd9d97d406860385668 upstream. In exc_machine_check_user(), local_db_save() and local_db_restore() are invoked in the outer entry stubs (DEFINE_IDTENTRY_MCE_USER, DEFINE_FREDENTRY_MCE, and DEFINE_IDTENTRY_RAW), surrounding exc_machine_check_user(). However, exc_machine_check_user() calls irqentry_exit_to_user_mode(), which handles pending thread work and may schedule() if TIF_NEED_RESCHED is set. If the task migrates to another CPU during schedule(), local_db_restore() runs on the new CPU with the dr7 state saved from the old CPU. This corrupts the new CPU's DR7 hardware debug register and leaves the old CPU's DR7 disabled. In short, local_db_save() and local_db_restore() pair must be run on the same CPU. To fix this, move local_db_save() and local_db_restore() inside exc_machine_check_user() and exc_machine_check_kernel(). In exc_machine_check_user(), DR7 is saved and restored strictly around do_machine_check() to avoid schedule() during migration. In exc_machine_check_kernel(), local_db_save() is called at the entry point to prevent early memory accesses from triggering nested #DB exceptions, and restored on all exits. Fixes: cd840e424f27 ("x86/entry, mce: Disallow #DB during #MC") Assisted-by: LLM Signed-off-by: Masami Hiramatsu (Google) Signed-off-by: Borislav Petkov (AMD) Acked-by: Peter Zijlstra (Intel) Cc: Link: https://patch.msgid.link/179005109564.388919.3937970081044095776.stgit@devnote2 Signed-off-by: Greg Kroah-Hartman commit a557816c6e38e9473d99aa9ea8c662f55fc85847 Author: Pablo Neira Ayuso Date: Mon Sep 28 09:58:47 2026 +0200 netfilter: nf_tables: join hook list via splice_list_rcu() in commit phase [ Upstream commit a6134e62dba2ea4f760b29d5226907f447c92400 ] Publish new hooks in the list into the basechain/flowtable using splice_list_rcu() to ensure netlink dump list traversal via rcu is safe while concurrent ruleset update is going on. Fixes: 78d9f48f7f44 ("netfilter: nf_tables: add devices to existing flowtable") Fixes: b9703ed44ffb ("netfilter: nf_tables: support for adding new devices to an existing netdev chain") Signed-off-by: Pablo Neira Ayuso Signed-off-by: Sasha Levin Signed-off-by: Benjamin Robin (Schneider Electric) Signed-off-by: Sasha Levin commit 7b32eba3fad845d9026caa3a959862d0fdb3bb46 Author: Pablo Neira Ayuso Date: Mon Sep 28 09:58:46 2026 +0200 rculist: add list_splice_rcu() for private lists [ Upstream commit f902877b635551513729bdf9a8d1422c4aab7741 ] This patch adds a helper function, list_splice_rcu(), to safely splice a private (non-RCU-protected) list into an RCU-protected list. The function ensures that only the pointer visible to RCU readers (prev->next) is updated using rcu_assign_pointer(), while the rest of the list manipulations are performed with regular assignments, as the source list is private and not visible to concurrent RCU readers. This is useful for moving elements from a private list into a global RCU-protected list, ensuring safe publication for RCU readers. Subsystems with some sort of batching mechanism from userspace can benefit from this new function. The function __list_splice_rcu() has been added for clarity and to follow the same pattern as in the existing list_splice*() interfaces, where there is a check to ensure that the list to splice is not empty. Note that __list_splice_rcu() has no documentation for this reason. Reviewed-by: Paul E. McKenney Signed-off-by: Pablo Neira Ayuso Signed-off-by: Benjamin Robin (Schneider Electric) Signed-off-by: Sasha Levin commit 162bf5f73f1b4a656ce3a57bfd7f19e371579a34 Author: Mark Amirkan Date: Sun Sep 13 10:30:05 2026 +0000 mptcp: return sk_wait_data() errors from recvmsg() [ Upstream commit 60404266ef3e0a1cd8f7a164060e0c83efb72f4b ] Commit 581302298524 ("mptcp: error out earlier on disconnect") made mptcp_recvmsg() stop when sk_wait_data() returns an error. The error is stored in err, but the function then jumps to a path which returns copied. When no data was copied, recvmsg() therefore returns zero and reports a false EOF. Store the result in copied, which is the value returned by the function. This also keeps the usual partial-read result when data was copied before the error. A recvmsg() blocked in one thread reproduces the issue when another thread disconnects the same MPTCP socket with connect(AF_UNSPEC). Before this change recvmsg() returns zero; afterwards it returns -EPIPE. Fixes: 581302298524 ("mptcp: error out earlier on disconnect") Cc: stable@vger.kernel.org Signed-off-by: Mark Amirkan Reviewed-by: Matthieu Baerts (NGI0) Link: https://patch.msgid.link/20260913-b4-send-mptcp-recv-error-v1-1-4eaa3684a8b8@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 9fe317b999aa8b816e3d1f47b5088b579c8ec2d0 Author: Zhan Xusheng Date: Fri Sep 18 21:29:15 2026 +0800 sched/core: Account PSI IRQ time to the execution context, not the scheduling context [ Upstream commit a0bb6fac53fa7cf1cadb487b43d4c9276a6b82e3 ] psi_account_irqtime() has two callers which share rq->psi_irq_time, and they disagree about the context: __schedule() passes the outgoing rq->curr, sched_tick() passes rq->donor. Under proxy execution the donor is blocked on a mutex while rq->curr burns the CPU. The tick charges PSI_IRQ_FULL to the donor's cgroup and advances the timestamp, so the call from __schedule() then finds delta <= 0 and charges nothing. The delta is not counted twice, it lands on the wrong cgroup. Pass rq->curr, which is what the call read before commit af0c8b2bf67b ("sched: Split scheduler and execution contexts") renamed 'curr' to 'donor' across sched_tick(). Without CONFIG_SCHED_PROXY_EXEC the two rq members are a union, so this only changes anything where that option is set, and it depends on EXPERT. Fixes: af0c8b2bf67b ("sched: Split scheduler and execution contexts") Signed-off-by: Zhan Xusheng Signed-off-by: Peter Zijlstra (Intel) Signed-off-by: Ingo Molnar Link: https://patch.msgid.link/20260918132915.1236312-1-zhanxusheng@xiaomi.com Signed-off-by: Sasha Levin commit 23250517cd475367653f1e99396967f54c6f1b7c Author: Namhyung Kim Date: Sun Sep 20 16:16:39 2026 -0700 perf/core: Fix a refcount leak in attach_perf_ctx_data() [ Upstream commit cca4980630b3c7a85f53cb43c6018184ce5d4e37 ] The attach_perf_ctx_data() can race on global and !global cases. The global case is protected by global_ctx_data_rwsem and shares a single reference count using perf_ctx_data.global field. But when it races with !global case, it may miss to set the global field and result in a reference count leak. CPU1 CPU2 ---------------------------------------------------------------- attach_task_ctx_data(.global=1) attach_task_ctx_data(.global=0) cd1 = alloc_perf_ctx_data(); cd2 = alloc_perf_ctx_data(); // { .global = 0, .refcount = 1 }; try_cmpxchg(); // success, // task->perf_ctx_data = cd2 try_cmpxhg(); // fail; old = cd2 refcount_inc_not_zero(&old->refcount); // success // old.refcount = 2 free_perf_ctx_data(cd1); Then later detach_global_ctx_data() will see the data but it's not marked as global, so it won't call detach_task_ctx_data(). Fixes: 506e64e710ff ("perf: attach/detach PMU specific data") Assisted-by: Sashiko.dev:Gemini-3.1-pro Signed-off-by: Namhyung Kim Signed-off-by: Peter Zijlstra (Intel) Link: https://patch.msgid.link/20260920231639.11910-1-namhyung@kernel.org Signed-off-by: Sasha Levin commit 654fc8445ca8d73b8b023e5be4e2f5a92b4d657d Author: Sean Christopherson Date: Mon Sep 21 12:14:12 2026 -0700 perf/x86/intel: Make @data a mandatory param for intel_guest_get_msrs() [ Upstream commit a391618e1d563f099e4c2a704f45d08329ccdf7c ] Drop "support" for passing a NULL @data/@kvm_pmu param when getting guest MSRs. KVM, the only in-tree user, unconditionally passes a non-NULL pointer, and carrying code that suggests @data may be NULL is confusing, e.g. incorrectly implies that there are scenarios where KVM doesn't pass a PMU context. Fixes: 8183a538cd95 ("KVM: x86/pmu: Add IA32_DS_AREA MSR emulation to support guest DS") Signed-off-by: Sean Christopherson Signed-off-by: Peter Zijlstra (Intel) Signed-off-by: Ingo Molnar Reviewed-by: Jim Mattson Reviewed-by: Dapeng Mi Link: https://patch.msgid.link/20260921191418.950933-5-seanjc@google.com Signed-off-by: Sasha Levin commit 0688289753450f1f5ed3ea4ac347de75b2df8496 Author: Sean Christopherson Date: Mon Sep 21 12:14:11 2026 -0700 perf/x86/intel: Don't pointlessly context switch DS_AREA (and PEBS config) if PEBS is unused [ Upstream commit d06260e99eb93d2942b7af4ccd789eb8a6c829d3 ] When filling the list of MSRs to be loaded by KVM on VM-Enter and VM-Exit, load the guest values for DS_AREA and (conditionally) MSR_PEBS_DATA_CFG if and only if PEBS will be active in the guest, i.e. only if a PEBS record may be generated while running the guest. As shown by the !pebs_ept path, it's perfectly safe to run with the host's DS_AREA, so long as PEBS-enabled counters are disabled via PERF_GLOBAL_CTRL. Omitting DS_AREA and MSR_PEBS_DATA_CFG when PEBS is unused saves two MSR writes per MSR on each VMX transition, i.e. eliminates two/four pointless MSR writes on each VMX roundtrip when PEBS isn't being used by the guest. Fixes: c59a1f106f5c ("KVM: x86/pmu: Add IA32_PEBS_ENABLE MSR emulation for extended PEBS") Signed-off-by: Sean Christopherson Signed-off-by: Peter Zijlstra (Intel) Signed-off-by: Ingo Molnar Reviewed-by: Jim Mattson Reviewed-by: Dapeng Mi Link: https://patch.msgid.link/20260921191418.950933-4-seanjc@google.com Signed-off-by: Sasha Levin commit 8cb5a01cb5f128771a7bf9708b71852e636a5636 Author: Sean Christopherson Date: Mon Sep 21 12:14:10 2026 -0700 perf/x86/intel: Don't write PEBS_ENABLED on host<=>guest xfers if CPU has PEBS isolation, to fix stuck PEBS_ENABLED [ Upstream commit 4b64dbdc5861477f148e13d1ed127e7fe7182e4f ] When filling the list of MSRs to be loaded by KVM on VM-Enter and VM-Exit, *never* insert an entry for PEBS_ENABLED if the CPU properly isolates PEBS events, in which case disabling counters via PERF_GLOBAL_CTRL is sufficient to prevent unwanted PEBS events in the guest (or host). Because perf loads PEBS_ENABLE with the unfiltered cpu_hw_events.pebs_enabled, i.e. with both host and guest masks, there is no need to load different values for the guest versus host, perf+KVM can and should simply control which counters are enabled/disabled via PERF_GLOBAL_CTRL. Avoiding touching PEBS_ENABLED "fixes" a bug where PEBS_ENABLED can end up with "stuck" bits if a PEBS event is throttled between generating the list and actually entering the guest (Intel CPUs can't arbtitrarily block NMIs). Fixes in quotes because leaving PEBS_ENABLED as-is doesn't fix the underlying problem of perf (via PMIs) being able to modify state after the perf<=>KVM handoff. But not writing PEBS_ENABLED is desirable no matter what, as stating the obvious, leaving PEBS_ENABLED as-is avoids three MSR writes on every VMX transition: one each on entry/exit, and one more explicit WRMSR to zero PEBS_ENABLED before VM-Entry (KVM assumes the only reason PEBS_ENABLED is in the load list is if the CPU lacks PEBS isolation and thus needs a quiescent period). Opportunistically add comments to (better) explain the rules for generating the set of PEBS counters that will be active while the guest is running, along with a FIXME for the suspected hack-a-fix where perf disables guest PEBS if _any_ PEBS event is configured to count in the host (commit 854250329c02 ("KVM: x86/pmu: Disable guest PEBS temporarily in two rare situations") doesn't explain the motivation, at all). Fixes: c59a1f106f5c ("KVM: x86/pmu: Add IA32_PEBS_ENABLE MSR emulation for extended PEBS") Signed-off-by: Sean Christopherson Signed-off-by: Peter Zijlstra (Intel) Signed-off-by: Ingo Molnar Reviewed-by: Dapeng Mi Link: https://patch.msgid.link/20260921191418.950933-3-seanjc@google.com Signed-off-by: Sasha Levin commit c205e3fc9962c6a0f7718b5d49a0f028682536a1 Author: Sean Christopherson Date: Mon Sep 21 12:14:09 2026 -0700 perf/x86/intel: Ensure KVM guest PEBS path doesn't set unwanted PERF_GLOBAL_CTRL bits [ Upstream commit cec38d5c098a350dcf084d345025136ade7e6d1e ] When reinstating PEBS counters into PERF_GLOBAL_CTRL for a KVM guest, mask the value with perf's desired/original PERF_GLOBAL_CTRL value to ensure KVM doesn't unintentionally set reserved bits in PERF_GLOBAL_CTRL. E.g. if the guest's PEBS_ENABLE value had bit 63, "Enable Precise Store", set, then using the raw guest PEBS value would propagate bit 63 to the guest's PERF_GLOBAL_CTRL value (which thankfully would be a failed VM-Entry, not a VMX Abort). The only reason this bug isn't reachable is because KVM doesn't support "Enable Precise Store" (which is probably a KVM bug?), i.e. bit 63 can't be set in kvm_pmu->pebs_enable and thus not in arr[pebs_enable].guest. In other words, this _should_ be a glorified NOP in the current code base. Fixes: c59a1f106f5c ("KVM: x86/pmu: Add IA32_PEBS_ENABLE MSR emulation for extended PEBS") Signed-off-by: Sean Christopherson Signed-off-by: Peter Zijlstra (Intel) Signed-off-by: Ingo Molnar Reviewed-by: Dapeng Mi Link: https://patch.msgid.link/20260921191418.950933-2-seanjc@google.com Signed-off-by: Sasha Levin commit 4bcb371019097d0e357f097902587401c509bbc9 Author: Hui Peng Date: Sat Sep 19 20:48:08 2026 +0000 autofs: fix sbi->pipe file reference leak in autofs_kill_sb() [ Upstream commit aa5e44b29ffe4eaa08cc2237fd65bc2596bc023e ] When autofs_fill_super() fails before clearing AUTOFS_SBI_CATATONIC (for example, when find_get_pid() fails on an invalid pgrp mount option, or when an fs_context is closed before mounting), deactivate_locked_super() invokes autofs_kill_sb() -> autofs_catatonic_mode(sbi). Because AUTOFS_SBI_CATATONIC is still set in sbi->flags, autofs_catatonic_mode() returns early without calling fput(sbi->pipe), permanently leaking the pipe struct file reference. Explicitly release sbi->pipe in autofs_kill_sb() if it is still non-NULL after autofs_catatonic_mode(). Fixes: ebc921ca9b92 ("autofs: copy autofs4 to autofs") Signed-off-by: Hui Peng Link: https://patch.msgid.link/20260919204808.2812930-1-benquike@gmail.com Signed-off-by: Christian Brauner (Amutable) Signed-off-by: Sasha Levin commit c7ee4bea64f559ea8e03260014775b64c720cb5b Author: Eric Dumazet Date: Thu Sep 24 08:29:51 2026 +0000 vlan: ensure sufficient headroom in vlan_dev_hard_header() [ Upstream commit cd5dd68267c4238795fadaf02b3575ca3f8a6500 ] Callers that only reserve ETH_HLEN or less (such as llc_alloc_frame()), or skbs allocated before dynamic device/headroom changes (e.g. toggling VLAN_FLAG_REORDER_HDR or bonding/team switching slaves), can reach vlan_dev_hard_header() with insufficient headroom and trigger skb_under_panic(). Use skb_cow_head() in vlan_dev_hard_header() when VLAN_FLAG_REORDER_HDR is not set to ensure sufficient headroom for the VLAN header(s) and the underlying device hard header. Use READ_ONCE() to read dev->hard_header_len and dev->needed_headroom as they can be updated concurrently under RTNL (e.g. in vlan_transfer_features()) while vlan_dev_hard_header() runs locklessly on the transmit path. Also avoid LL_RESERVED_SPACE(dev) here so that the extra HH_DATA_MOD alignment padding does not trigger unnecessary pskb_expand_head() reallocations on inner stacked VLAN devices after the outer VLAN header has been pushed. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Reported-by: Zixuan Chai Closes: https://lore.kernel.org/netdev/cover.1789987105.git.petalzu987@gmail.com/ Link: https://lore.kernel.org/netdev/179022851638.2160803.1808206741379444999@kernel.org/ Cc: Hangbin Liu Signed-off-by: Eric Dumazet Link: https://patch.msgid.link/20260924082951.1599377-5-edumazet@google.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 7f8b56c82919c0365b7e9587c1c18e9266dd9730 Author: Eric Dumazet Date: Thu Sep 24 08:29:50 2026 +0000 net/sched: sch_teql: fix shadowed err in __teql_resolve() [ Upstream commit 907b978e82cb4c1c245fc2985bb27c5d5c88c8f6 ] __teql_resolve() declares an inner 'int err;' inside the 'if (neigh_event_send(n, skb_res) == 0)' block, shadowing the outer 'int err = 0;'. As a result, a negative return from dev_hard_header() is written to the inner variable and __teql_resolve() still returns 0. Remove the shadowed variable and set the outer err to -EINVAL when dev_hard_header() returns a negative error. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Closes: https://lore.kernel.org/netdev/179022851638.2160803.1808206741379444999@kernel.org/ Cc: Jamal Hadi Salim Cc: Jiri Pirko Signed-off-by: Eric Dumazet Link: https://patch.msgid.link/20260924082951.1599377-4-edumazet@google.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 15d00994a1bd279ac87b2e3fde2c955fd99a2579 Author: Eric Dumazet Date: Thu Sep 24 08:29:49 2026 +0000 bridge: check llc_mac_hdr_init() return value in br_send_bpdu() [ Upstream commit ac704ff08e511c87643799c385f55ecd69b85e03 ] If llc_mac_hdr_init() fails (for instance if the port device type does not support LLC or dev_hard_header() fails), br_send_bpdu() should drop the skb instead of resetting the mac header to the LLC payload and transmitting a malformed frame. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Closes: https://lore.kernel.org/netdev/179022851638.2160803.1808206741379444999@kernel.org/ Cc: Nikolay Aleksandrov Cc: Ido Schimmel Cc: bridge@lists.linux.dev Signed-off-by: Eric Dumazet Acked-by: Nikolay Aleksandrov Link: https://patch.msgid.link/20260924082951.1599377-3-edumazet@google.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 41dcae82bece9f106b4dff0ee777bba8d157c8a8 Author: Eric Dumazet Date: Thu Sep 24 08:29:48 2026 +0000 llc: fix skb UAF and leaks on llc_mac_hdr_init() failure [ Upstream commit 72f9dd522f8d6c5a00be9695c7bb74631eb5069e ] In llc_conn_ac_resend_i_xxx_x_set_0_or_send_rr(), if llc_mac_hdr_init() fails, kfree_skb(skb) is called instead of kfree_skb(nskb). This leaks the newly allocated nskb, reads from the freed skb via LLC_I_GET_NR(pdu), and double-frees skb when llc_conn_state_process() drops its reference. In llc_sap_action_send_xid_r() and llc_sap_action_send_test_r(), nskb is leaked if llc_mac_hdr_init() returns an error. Free nskb in all three error paths. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Closes: https://lore.kernel.org/netdev/179022851638.2160803.1808206741379444999@kernel.org/ Signed-off-by: Eric Dumazet Link: https://patch.msgid.link/20260924082951.1599377-2-edumazet@google.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 2dc3ae3942e151fef254b7bbb65c832192ad2069 Author: Eric Dumazet Date: Thu Sep 24 00:42:52 2026 +0000 gve: DQO: reject TSO packets with an out of range MSS [ Upstream commit 296c83b5ccc808c080865eb20fd7a477b0355bb7 ] gve_prep_tso() notes that the device requires the MSS to be <= 9728, but does not enforce it, assuming the 9K MTU enforced by the hypervisor and the 64KB limit on TSO sizes are enough. This does not hold for packets that were not generated locally. A guest behind a tap, or any packet socket user, can provide an arbitrary gso_size in virtio_net_hdr. Layer 2 forwarding does not check the MTU for GSO packets (is_skb_forwardable()), and gso_features_check() only bounds skb->len and gso_segs, never gso_size. Such a packet reaches gve_tx_fill_tso_ctx_desc(), which puts gso_size into the mss field of the TSO context descriptor. This field is 14 bits wide, so a gso_size of 16384 is silently turned into an MSS of zero. Drop these packets from gve_prep_tso(), and make sure that gve_features_check_dqo() leaves their GSO bits alone: skb_segment() splits at gso_size regardless of the MTU, so falling back to software segmentation would give the device non TSO packets bigger than the 9728 bytes it supports. Note that the device can still be given oversized non TSO packets when the stack segments in software for other reasons, for instance after TSO has been disabled with ethtool. This is a generic issue, because the MTU check is skipped for GSO packets in the forwarding path, and is addressed separately. Fixes: a57e5de476be ("gve: DQO: Add TX path") Signed-off-by: Eric Dumazet Reviewed-by: Harshitha Ramamurthy Link: https://patch.msgid.link/20260924004252.1196328-3-edumazet@google.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 089e04db44754788f28690a9993c70dcf99cb638 Author: Eddie Phillips Date: Thu Sep 24 00:42:51 2026 +0000 gve: fix TX drop when GSO MSS is too small for hw [ Upstream commit 3b430ea6234087957b0d3cd181e3116722b59819 ] The device has a strict requirement that the minimum MSS (gso_size) for TSO/GSO packets must be at least 88 bytes. If a packet below this threshold is pushed to the hardware, it can cause hardware to silently drop the packet, leading to increased latency and retransmissions. Currently, this is validated too late in the transmit pipeline (gve_prep_tso), leading to silent drops. Fix this by moving the validation into the .ndo_features_check callback (gve_features_check_dqo). If we detect a GSO packet with a gso_size smaller than GVE_TX_MIN_TSO_MSS_DQO, we clear the GSO feature flags for this packet. Fixes: a57e5de476be ("gve: DQO: Add TX path") Signed-off-by: Eddie Phillips Signed-off-by: Eric Dumazet Reviewed-by: Harshitha Ramamurthy Link: https://patch.msgid.link/20260924004252.1196328-2-edumazet@google.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit feab3880a72b43913d332c06b75bb80526fa47ba Author: Eric Dumazet Date: Wed Sep 23 14:59:42 2026 +0000 gve: DQO: fix header length used by gve_can_send_tso() for UDP GSO [ Upstream commit 83769c23fb1879edc916a526ba424285033baf2d ] gve_can_send_tso() computes how many buffers each segment of a GSO packet would span, and for this it needs the length of the headers that the device replicates in front of every segment. It unconditionally uses skb_tcp_all_headers(), which reads the doff field of the TCP header. SKB_GSO_UDP_L4 packets have no TCP header: tcp_hdrlen() then reads one byte of the UDP payload, and header_len can be anything in [0, 60] instead of the transport offset plus the eight bytes of the UDP header that gve_prep_tso() programs into the TSO context descriptor. A wrong header length shifts all the segment boundaries computed in the loop, so the number of buffers per segment can be over or under estimated. In the first case, GSO is needlessly disabled for this packet by gve_features_check_dqo() and the stack has to segment it. In the second case, the driver hands the device a packet whose segments span more than GVE_TX_MAX_DATA_DESCS buffers. Use the UDP header length for SKB_GSO_UDP_L4 packets, matching what gve_prep_tso() does. Fixes: 014c607f86ab ("gve: add support for UDP GSO for DQO format") Closes: https://lore.kernel.org/netdev/CANn89i+MS4L60sFQ49=-f-mibeveUfcrpVkD5X+Qy6SOnEpd6w@mail.gmail.com/ Signed-off-by: Eric Dumazet Cc: Ankit Garg Cc: Harshitha Ramamurthy Cc: Joshua Washington Cc: Willem de Bruijn Reviewed-by: Ankit Garg Reviewed-by: Harshitha Ramamurthy Link: https://patch.msgid.link/20260923145942.731365-1-edumazet@google.com Signed-off-by: Jakub Kicinski Stable-dep-of: 3b430ea62340 ("gve: fix TX drop when GSO MSS is too small for hw") Signed-off-by: Sasha Levin commit 2f855c52a2e1d3c40d186549ac46cbd88b38eb49 Author: Ankit Garg Date: Fri Mar 6 22:48:16 2026 +0000 gve: add support for UDP GSO for DQO format [ Upstream commit 014c607f86abc903d7bf46e13373d89392e371fe ] Enable support for UDP GSO when using DQO format. Advertise the feature flag during device initialization and enable offload by default. Signed-off-by: Ankit Garg Reviewed-by: Willem de Bruijn Signed-off-by: Harshitha Ramamurthy Link: https://patch.msgid.link/20260306224816.3391551-1-hramamurthy@google.com Signed-off-by: Jakub Kicinski Stable-dep-of: 3b430ea62340 ("gve: fix TX drop when GSO MSS is too small for hw") Signed-off-by: Sasha Levin commit 6d4a5ce75b6e316f86fe399b16df98be715050a5 Author: Ankit Garg Date: Tue Mar 3 11:55:49 2026 -0800 gve: Enable hw-gro by default if device supported [ Upstream commit 3c398063ef01b02d7efd31662154fe70fd28ace6 ] Change the driver's default behavior to enable hw-gro whenever supported for device. Performance observations: - We observed ~10% improvement in RX single stream throughput across various MTU sizes. - No change in TCP_RR/TCP_CRR latencies Signed-off-by: Ankit Garg Reviewed-by: Willem de Bruijn Reviewed-by: Harshitha Ramamurthy Signed-off-by: Joshua Washington Link: https://patch.msgid.link/20260303195549.2679070-5-joshwash@google.com Signed-off-by: Paolo Abeni Stable-dep-of: 3b430ea62340 ("gve: fix TX drop when GSO MSS is too small for hw") Signed-off-by: Sasha Levin commit adbfc707ff6b994bc453e61a72d96acadbe45a95 Author: Ankit Garg Date: Tue Mar 3 11:55:46 2026 -0800 gve: Advertise NETIF_F_GRO_HW instead of NETIF_F_LRO [ Upstream commit e637c244b954426b84340cbc551ca0e2a32058ce ] The device behind DQO format has always coalesced packets per stricter hardware GRO spec even though it was being advertised as LRO. Update advertised capability to match device behavior. Signed-off-by: Ankit Garg Reviewed-by: Willem de Bruijn Reviewed-by: Harshitha Ramamurthy Signed-off-by: Joshua Washington Link: https://patch.msgid.link/20260303195549.2679070-2-joshwash@google.com Signed-off-by: Paolo Abeni Stable-dep-of: 3b430ea62340 ("gve: fix TX drop when GSO MSS is too small for hw") Signed-off-by: Sasha Levin commit 6c7937db1d561f6ace8aded6652cd1f83988a187 Author: Ankit Garg Date: Thu Nov 6 11:27:45 2025 -0800 gve: Allow ethtool to configure rx_buf_len [ Upstream commit d235bb213f411ace8317bcca3740a1008628ea9c ] Add support for getting and setting the RX buffer length via the ethtool ring parameters (`ethtool -g`/`-G`). The driver restricts the allowed buffer length to 2048 (SZ_2K) by default and allows 4096 (SZ_4K) based on device options. As XDP is only supported when the `rx_buf_len` is 2048, the driver now enforces this in two places: 1. In `gve_xdp_set`, rejecting XDP programs if the current buffer length is not 2048. 2. In `gve_set_rx_buf_len_config`, rejecting buffer length changes if XDP is loaded and the new length is not 2048. Signed-off-by: Ankit Garg Reviewed-by: Harshitha Ramamurthy Reviewed-by: Jordan Rhee Reviewed-by: Willem de Bruijn Signed-off-by: Joshua Washington Link: https://patch.msgid.link/20251106192746.243525-4-joshwash@google.com Signed-off-by: Jakub Kicinski Stable-dep-of: 3b430ea62340 ("gve: fix TX drop when GSO MSS is too small for hw") Signed-off-by: Sasha Levin commit bcaedb514a5d28465ab104e060c5e80bfd03d456 Author: Ankit Garg Date: Thu Nov 6 11:27:44 2025 -0800 gve: Use extack to log xdp config verification errors [ Upstream commit 091a3b6ff2b98354270cb9278faad6d17d5aa27d ] Plumb extack as it allows us to send more detailed error messages back and append 'gve' suffix to method name per convention. NL_SET_ERR_MSG_FMT_MOD doesn't support format string longer than 80 chars so keeping netdev warning with actual queue count details. Signed-off-by: Ankit Garg Reviewed-by: Harshitha Ramamurthy Reviewed-by: Willem de Bruijn Signed-off-by: Joshua Washington Link: https://patch.msgid.link/20251106192746.243525-3-joshwash@google.com Signed-off-by: Jakub Kicinski Stable-dep-of: 3b430ea62340 ("gve: fix TX drop when GSO MSS is too small for hw") Signed-off-by: Sasha Levin commit a0d644b8ff883c51bca26f133283e1c39da2d023 Author: Ankit Garg Date: Thu Oct 16 18:25:42 2025 -0700 gve: Consolidate and persist ethtool ring changes [ Upstream commit c30fd916c4d7760ccf654e52370b0a31be885789 ] Refactor the ethtool ring parameter configuration logic to address two issues: unnecessary queue resets and lost configuration changes when the interface is down. Previously, `gve_set_ringparam` could trigger multiple queue destructions and recreations for a single command, as different settings (e.g., header split, ring sizes) were applied one by one. Furthermore, if the interface was down, any changes made via ethtool were discarded instead of being saved for the next time the interface was brought up. This patch centralizes the configuration logic. Individual functions like `gve_set_hsplit_config` are modified to only validate and stage changes in a temporary config struct. The main `gve_set_ringparam` function now gathers all staged changes and applies them as a single, combined configuration: 1. If the interface is up, it calls `gve_adjust_config` once. 2. If the interface is down, it saves the settings directly to the driver's private struct, ensuring they persist and are used when the interface is brought back up. Signed-off-by: Ankit Garg Reviewed-by: Harshitha Ramamurthy Reviewed-by: Jordan Rhee Reviewed-by: Willem de Bruijn Signed-off-by: Joshua Washington Link: https://patch.msgid.link/20251017012614.3631351-1-joshwash@google.com Signed-off-by: Jakub Kicinski Stable-dep-of: 3b430ea62340 ("gve: fix TX drop when GSO MSS is too small for hw") Signed-off-by: Sasha Levin commit 5f86408eba31bdf3073fb864d04b711473a07453 Author: Coia Prant Date: Wed Sep 23 20:37:13 2026 +0800 net: ethernet: stmmac: dwmac-rk: fix bulk clock leak when the PHY clock fails [ Upstream commit 8db67bb6a1fffa4df68fbbc22e39943aeeff9178 ] gmac_clk_enable() enables the bulk clocks first and then the optional PHY clock. If clk_prepare_enable() on the PHY clock fails, the function returns without rolling back the bulk clocks, and bsp_priv->clk_enabled stays false, so the later gmac_clk_enable(bsp_priv, false) becomes a no-op and the bulk clock references are leaked. Add the missing clk_bulk_disable_unprepare() on that failure path. Fixes: ea449f7fa0bf ("net: ethernet: stmmac: dwmac-rk: rework optional clock handling") Reviewed-by: Maxime Chevallier Reviewed-by: Heiko Stuebner Acked-by: Lorenzo Bianconi Signed-off-by: Coia Prant Link: https://patch.msgid.link/20260923123713.3137146-1-coiaprant@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit d990b786abfd6b6fc4274aaa6d07a20a9df159ab Author: Ginger Li Date: Tue Sep 22 16:09:09 2026 +0800 tipc: Fix a data race on mon->peer_cnt in mon_timeout() [ Upstream commit 8e1937fed6738460554ec123c64839e2445e7d53 ] mon_timeout() evaluates dom_size(mon->peer_cnt) before it takes mon->lock, while mon->peer_cnt is updated under that lock by tipc_mon_add_peer() and tipc_mon_remove_peer(). The value can therefore be stale, and the decision whether the local domain has to be recomputed can be based on an outdated member count. Read mon->peer_cnt inside the write_lock_bh(&mon->lock) protected region. Fixes: 35c55c9877f8 ("tipc: add neighbor monitoring framework") Signed-off-by: Ginger Li Reviewed-by: Tung Nguyen Link: https://patch.msgid.link/20260922080909.21123-1-ginger.jzllee@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 3d65a2eb4449fc9658fb860dea893025d028d90e Author: Sidraya Jayagond Date: Tue Sep 22 09:31:49 2026 +0200 net/smc: fix UAF on lgr list traversal in smcr_port_err() [ Upstream commit 61cb282fe97b3b0ba32ca09417a693162bf4ae3f ] smcr_port_err() traverses smc_lgr_list.list without holding smc_lgr_list.lock, allowing a concurrent smc_lgr_terminate_sched() to free an lgr while it is still being dereferenced. Hold smc_lgr_list.lock across the traversal. Update smc_ib_gid_check() to call smcr_port_err() after releasing the lock. Fixes: 541afa10c126 ("net/smc: add smcr_port_err() and smcr_link_down() processing") Reviewed-by: Mahanta Jambigi Signed-off-by: Sidraya Jayagond Reviewed-by: Dust Li Link: https://patch.msgid.link/20260922073149.474762-1-sidraya@linux.ibm.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 8468d05aa08dc430b5675beadf6f6fb8943f170b Author: Sang-Hoon Choi Date: Tue Sep 22 03:15:59 2026 +0900 nfp: hold IPsec RX state under the XArray lock [ Upstream commit 1a983a4e14c635c40354be110cd9a1a5c94e01e6 ] nfp_net_ipsec_rx() drops the XArray lock before taking a reference to the xfrm_state it found. The delete path can erase the entry and drop the last state reference in that interval. RX can then try to increment a zero refcount after the state has been queued for destruction. The driver queues firmware invalidation asynchronously; the delete path does not wait for the command to complete or drain pending RX processing. The XFRM garbage collector waits for an RCU grace period before freeing the state. That delays reclamation but does not make acquiring a reference from zero valid. Take the xfrm_state reference before releasing the XArray lock so xa_erase() cannot run between lookup and reference acquisition. Fixes: 57f273adbcd4 ("nfp: add framework to support ipsec offloading") Reported-by: Changyul Lee Signed-off-by: Sang-Hoon Choi Link: https://patch.msgid.link/179001455912.44752.17153022439349797877.idr-bug-92@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 4e7a4af364c2048bafb213bf0ef4a89dcbe1bfd4 Author: Aleksei Sviridkin Date: Mon Sep 21 01:20:44 2026 +0300 net: phylink: record the PHY only once bringup cannot fail [ Upstream commit a940003f44e7e441c228151dd212642152700ec8 ] phylink_bringup_phy() stores the PHY in pl->phydev before its last fallible step: on a MAC whose phylink ops implement LPI, phy_eee_rx_clock_stop() can fail with a real MDIO error. The callers unwind with phy_detach(), which knows nothing about pl->phydev, so a pointer to a PHY that is no longer attached outlives the failed connect. What that costs depends on how the caller got here. phylink_connect_phy() goes through phylink_attach_phy(), which refuses to attach while pl->phydev is set, turning a transient MDIO error into a permanent -EBUSY. The SFP path is worse than that: sfp_sm_probe_phy() answers the failure with phy_device_remove() and phy_device_free(), and it assigns sfp->mod_phy only past that error return, so nothing clears pl->phydev and it is left pointing at a freed phy_device that phylink_resolve() and the ethtool helpers go on reading. phylink_fwnode_phy_connect() has no such check, so a later connect overwrites the stale pointer and hides the problem. A disconnect does not: phylink_disconnect_phy() hands that pointer to phy_disconnect(), and the second phy_detach() on the same PHY drops references the first one already released. Found while making a DSA port survive a PHY whose driver arrives after the switch probes: keeping the port across a failed connect and retrying is what makes this window reachable. Publish the pointer after the last call that can fail instead of unwinding it afterwards. Nothing between the two points reads pl->phydev, and the registration that follows cannot fail: phy_request_interrupt() falls back to polling on its own. The PHY-side state keeps the order it had, so no MDIO operation moves relative to another. Fixes: 03abf2a7c654 ("net: phylink: add EEE management") Signed-off-by: Aleksei Sviridkin Link: https://patch.msgid.link/20260920222044.1752860-1-f@lex.la Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 631aa4cb45099f09e9385dd786bd291c6270bfc4 Author: Yilin Zhang Date: Thu Sep 24 12:49:00 2026 +0800 tcp: fix use-after-free of retransmit_skb_hint in tcp_send_synack() [ Upstream commit fe99bbeee5c5dbd3abc30721a8079ced59649d97 ] When tcp_send_synack() replaces the cloned SYN skb at the head of the retransmit queue with a copy, it frees the original with tcp_rtx_queue_unlink_and_free() and only repairs tp->highest_sack. tp->retransmit_skb_hint keeps pointing at the freed skbuff_fclone_cache object. The dangling hint is read in tcp_verify_retransmit_hint() and used as the root of the rbtree walk in tcp_xmit_retransmit_queue(). An unprivileged TFO client (sendmsg(MSG_FASTOPEN)) can arm the hint with an attacker-supplied ICMP fragmentation-needed message, after which a simultaneous open frees the armed SYN skb: BUG: KASAN: slab-use-after-free in tcp_mark_skb_lost (net/ipv4/tcp_input.c:1316) Read of size 4 at addr ffff88800604d928 by task swapper/1/0 Call Trace: tcp_mark_skb_lost (net/ipv4/tcp_input.c:1316) tcp_simple_retransmit (net/ipv4/tcp_input.c:3158) tcp_v4_err (net/ipv4/tcp_ipv4.c:587) Sync the hint to the copy. Fixes: c31b70c9968f ("tcp: Add logic to check for SYN w/ data in tcp_simple_retransmit") Reported-by: Kimi Security Team Tested-by: Weiming Shi Signed-off-by: Yilin Zhang Reviewed-by: Eric Dumazet Link: https://patch.msgid.link/8a9dff4063a2745653b7e88ceb745d75efa16e68.1790224474.git.yilinzhang@moonshot.ai Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 2357c682f6eaf639dec88c68b3c0c93a3b802635 Author: Aleksei Sviridkin Date: Fri Sep 18 04:50:20 2026 +0300 net: dsa: mt7530: leave the MDIO IRQ mappings to regmap-irq [ Upstream commit 0d80ba0a204c6a16bd7778b50de578dff107c0fe ] mt7530_remove_common() disposes the per-PHY interrupt mappings from .remove, but the regmap-irq chip that owns the domain is devm-registered, so its parent interrupt is only freed once .remove has returned. The switch's own regmap-irq thread can therefore still dispatch on a mapping that is already gone: irq_find_mapping() returns 0, irq_to_desc() returns NULL and handle_nested_irq() locks desc->lock without checking it. The attached PHYs have not given those interrupts back yet either, which the kernel warns about a moment before the fault. regmap_del_irq_chip() disposes the same mappings itself, after freeing the parent interrupt and before removing the domain, so there is nothing left for the driver to do here. Until it runs the descriptors stay alive, and a late dispatch on one of them is harmless: dsa_unregister_switch() has freed the PHY handlers by then, so handle_nested_irq() finds no action and returns. Fixes: 254f6b272e3b ("dsa: mt7530: Utilize REGMAP_IRQ for interrupt handling") Signed-off-by: Aleksei Sviridkin Link: https://patch.msgid.link/20260918015020.2518315-3-f@lex.la Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 62c0e38b2e5034184d4502ce77bd178b29e59b83 Author: Aleksei Sviridkin Date: Fri Sep 18 04:50:19 2026 +0300 net: dsa: mt7530: fix NULL dereference on unbind of MT7531 and MT7621 [ Upstream commit c2cdef41e0b4d8ed23a5b41e6ad4e64594e055e4 ] The core and io supplies are only requested for ID_MT7530: both the devm_regulator_get() in probe and the regulator_enable() in mt7530_setup() are guarded by the switch id, but mt7530_remove() disables them unconditionally. On an MT7621 or an MT7531 both pointers are still NULL from devm_kzalloc(), so rmmod or a sysfs unbind calls regulator_disable() on NULL. Fixes: ddda1ac116c8 ("net: dsa: mt7530: support the 7530 switch on the Mediatek MT7621 SoC") Signed-off-by: Aleksei Sviridkin Link: https://patch.msgid.link/20260918015020.2518315-2-f@lex.la Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 2b42adf3b8ee9a0e1d2629aca3c5a02a1c438066 Author: Haseeb Malik Date: Mon Sep 21 16:40:30 2026 -0400 macsec: initialize SecY before registering the netdevice [ Upstream commit c2de369c5c5b8599ca10fd5ca8d11fcd845c1331 ] Creating a MACsec device with MAC offload over an LRO-capable lower device triggers a warning in rtmsg_ifinfo_build_skb() when IPv4 forwarding is enabled by default. register_netdevice() invokes inetdev_init(), which disables LRO and emits a NETDEV_FEAT_CHANGE notification. This reaches macsec_fill_info() before macsec_add_dev() initializes the SecY. key_len is still zero, so macsec_fill_info() returns -EMSGSIZE and trips the WARN_ON in rtmsg_ifinfo_build_skb(), even though the skb has enough space. Even without the warning, notifications during registration can report uninitialized SecY attributes, including the SCI. This ordering has existed since the driver was introduced. Initialize the SecY and apply the new-link attributes before registration. Move MAC address inheritance into macsec_newlink() so the SCI can also be initialized before registration-time notifications report it. Move the per-CPU statistics and metadata destination allocation into ndo_init(), and release partial allocations on failure. Fixes: c09440f7dcb3 ("macsec: introduce IEEE 802.1AE driver") Reported-by: syzbot+f2f6312ad1b5a0bfe316@syzkaller.appspotmail.com Closes: https://syzkaller.appspot.com/bug?extid=f2f6312ad1b5a0bfe316 Suggested-by: Sabrina Dubroca Link: https://lists.openwall.net/linux-kernel/2026/08/19/552 Signed-off-by: Haseeb Malik Reviewed-by: Sabrina Dubroca Link: https://patch.msgid.link/20260921-fix-macsec-net-v3-1-accf94f93f5e@gmail.com Signed-off-by: Paolo Abeni Signed-off-by: Sasha Levin commit dc106fe25e9a85f26d728014fa1ac4a0c0d86365 Author: David Dai Date: Fri Sep 18 16:11:55 2026 -0500 bonding: crypto offload enabled, non-offload slave failover, rekey failed [ Upstream commit 00efbbd40bd5fd92c67b7cf1aab8904fa59a96f6 ] Create a bonding device (i.e. bond0) in active-backup mode, 2 slaves. Active slave: offload capable interface (i.e. eth1), primary interface. Backup slave: non-offload capable interface(i.e. eth2). Configure strongswan service swantl.conf child SA "hw_offload = crypto" Start strongswan service IPSec Crytpo Offload is enabled on top of bond0. i.e. ip xfrm state |grep offload crypto offload parameters: dev bond0 dir out mode crypto crypto offload parameters: dev bond0 dir in mode crypto Active slave eth1 takes adavantage of IPSec Crypto Offload capability. If active slave eth1 is down for any reason (i.e. eth1 link down): ip link set down dev eth1 non-offload capable interface eth2 failover to becomes active slave. The existing SAs can continue use software IPsec after failover. Traffic still keeps going properly. However if eth1 link had not recovered yet, strongswan service does new child SA rekey, or uses swanctl command to do new child SA rekey, it will fail because active slave eth2 doesn't support crypto offload. In bond_ipsec_add_sa routine, it returns -EINVAL now, which is treated as fatal error by xfrm_dev_state_add routine in kernel xfrm. To make the non-offload active slave survive the child SA rekey, need to make bond_ipsec_add_sa routine returns -EOPNOTSUPP instead when active slave doesn't support IPsec Crypto offload, the xfrm will gracefully fallback to create new SA using Software IPsec. Network traffic can keep going. After offload capable interface eth1 link is up, becomes active slave, next time strongswan child SA rekey will create a new SA which enables crypto offload again. Fixes: 18cb261afd7b ("bonding: support hardware encryption offload to slaves") Signed-off-by: David Dai Reviewed-by: Hangbin Liu Link: https://patch.msgid.link/20260918211155.1664493-1-zdai@linux.ibm.com Signed-off-by: Paolo Abeni Signed-off-by: Sasha Levin commit e0560db0ea6e10c283ebe7c259a7af8d3b6e61a6 Author: Pengpeng Hou Date: Sun Sep 20 11:43:29 2026 +0800 drm/imagination: clamp freelist reconstruction requests [ Upstream commit 45585c3aa285854face65293acc95eff73063d6d ] The firmware reconstruction count controls accesses to the request's fixed freelist ID array and the copy into the fixed response array. Neither access currently bounds the count to those protocol arrays. Clamp the count to the request capacity, which is shared by the response layout, and use that count consistently for reconstruction and response publication. Keep the firmware recovery exchange instead of dropping an oversized request without a response, as discussed with the firmware maintainer. The issue was found by our static-analysis tool. Fixes: 6eedddab733b ("drm/imagination: Implement free list and HWRT create and destroy ioctls") Assisted-by: gpt 5 Signed-off-by: Pengpeng Hou Reviewed-by: Alessio Belle Link: https://patch.msgid.link/20260920034329.16614-1-hppiscas@163.com Signed-off-by: Brajesh Gupta Signed-off-by: Sasha Levin commit 2cc1c3977acbffc417fb5b57699d1df3dcffb5fd Author: Norbert Szetei Date: Mon Sep 21 17:03:57 2026 +0200 net: xps: reject an out of range traffic class [ Upstream commit 4da3b7b8b50f3e2fde54a4c18a82a8e3f6223910 ] Only the entries below dev->num_tc are valid in dev->tc_to_txq[], and dev->prio_tc_map[] may only name classes below it. netdev_set_num_tc() lowers dev->num_tc without touching either array. netdev_txq_to_tc() walks all TC_MAX_QUEUE slots and netdev_get_prio_tc_map() returns the entry as it stands, so a leftover entry is handed out as a traffic class >= dev->num_tc. Taking that class from netdev_txq_to_tc(), __netif_set_xps_queue() rejects only a negative one and indexes an XPS map sized for dev->num_tc classes: tci = j * num_tc + tc; RCU_INIT_POINTER(new_dev_maps->attr_map[tci], map); attr_map[] holds nr_ids * num_tc entries and j runs over the ids named in the mask, so a class that is not below num_tc pushes tci past the end of the map for the last ids and the store overruns it. Any caller that lowers num_tc leaves such entries behind, and mqprio_destroy() tears down with netdev_set_num_tc(dev, 0) rather than netdev_reset_tc(). After mqprio with 8 classes then 1, tc_to_txq[1..7] still describe txq 1..7. The splat is from an XPS write to txq 2 on a veth with 8 rx queues: attr_map[] has 8 * 1 entries, tci = j + 2, and j == 6 stores one past the end of the 88-byte map: BUG: KASAN: slab-out-of-bounds in __netif_set_xps_queue (net/core/dev.c:2954) Write of size 8 at addr ffff88813016bc58 by task xps_oob/634 __netif_set_xps_queue (net/core/dev.c:2954) xps_rxqs_store (net/core/net-sysfs.c:1880) netdev_queue_attr_store (net/core/net-sysfs.c:1390) Allocated by task 634: __kmalloc_noprof (mm/slub.c:5439) __netif_set_xps_queue (net/core/dev.c:2937) The buggy address is located 0 bytes to the right of allocated 88-byte region [ffff88813016bc00, ffff88813016bc58) Reject a class the map has no room for. Fixes: 184c449f91fe ("net: Add support for XPS with QoS via traffic classes") Signed-off-by: Norbert Szetei Link: https://patch.msgid.link/162DD16F-54C6-444A-9E09-0B8CB3D591F2@doyensec.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit a3c5eb6e382dd426a5e7eb08608897ed57167ff7 Author: Sanghyun Park Date: Fri Sep 18 12:26:58 2026 +0900 vxlan: use one headroom snapshot for neighbour replies [ Upstream commit 481506a756dcd828ef42391cb08f38d8d96d38fc ] vxlan_na_create() samples LL_RESERVED_SPACE() to size the reply skb and then samples it again to reserve headroom. A concurrent vxlan_changelink() can update needed_headroom between the two reads, creating a TOCTOU race. The second value can exceed the allocation and make the Ethernet header write out of bounds. The race is reproducible on the unpatched kernel. It occurred when vxlan_na_create() generated a neighbour reply while vxlan_changelink() changed the link headroom. KASAN caught a four-byte write two bytes beyond a 704-byte skbuff_small_head allocation. Snapshot the headroom once and use that value for both allocation and reservation. Fixes: 4b29dba9c085 ("vxlan: fix nonfunctional neigh_reduce()") Signed-off-by: Sanghyun Park Link: https://patch.msgid.link/20260918032842.502409-2-sanghyun.park.cnu@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit bdd41b96e814300c66222cf4e47ecb8aef995656 Author: Xuanqiang Luo Date: Mon Sep 21 11:18:59 2026 +0800 ip_gre: Reject enabling collect metadata through changelink [ Upstream commit a3f315be9d30eeb6938d11fa17fd4b32d52f7c42 ] ipgre_netlink_parms() can enable collect_md on an existing GRE, GRETAP or ERSPAN device. Unlike newlink, changelink does not enforce metadata tunnel uniqueness. Converting a non-metadata device can therefore replace the metadata receive entry for another device of the same type in the same netns. Deleting either device then clears the shared entry, breaking metadata receive lookup for the surviving device. If parameter validation fails after collect_md is set, deleting the modified device can also clear an entry it never owned. Reject enabling metadata mode in both changelink callbacks before any encapsulation or tunnel parameters are modified. Allow requests that repeat the metadata attribute on an existing metadata device. Fixes: 2e15ea390e6f ("ip_gre: Add support to collect tunnel metadata.") Signed-off-by: Xuanqiang Luo Reviewed-by: Ido Schimmel Reviewed-by: Hangbin Liu Link: https://patch.msgid.link/20260921031859.9283-1-xuanqiang.luo@linux.dev Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit c3fab2f1d8a6e555b5783943bd12d675fdd7f9b9 Author: Florian Fainelli Date: Mon Sep 21 15:00:21 2026 -0700 net: bcmgenet: mask DMA_TIMEOUT_MASK when reading DMA_RING0_TIMEOUT [ Upstream commit d64e277b955be4506931802837499b62c8f3968a ] bcmgenet_get_coalesce() reads DMA_RING0_TIMEOUT to calculate rx_coalesce_usecs without masking out bits outside DMA_TIMEOUT_MASK (16 bits). If upper bits are non-zero or contain status/flags, the computed value of rx_coalesce_usecs returned to userspace via ethtool becomes corrupted. Mask the register read with DMA_TIMEOUT_MASK before computing the timeout in microseconds. Fixes: 4a29645bfe6c ("net: bcmgenet: Implement RX coalescing control knobs") Reviewed-by: Nicolai Buchwitz Signed-off-by: Florian Fainelli Link: https://patch.msgid.link/20260921220021.281418-6-florian.fainelli@broadcom.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit d68899b50caf8b54592a42c753e413cd4abd8e11 Author: Florian Fainelli Date: Mon Sep 21 15:00:20 2026 -0700 net: bcmgenet: validate Ethernet address in bcmgenet_set_mac_addr [ Upstream commit 273941c85fc2632cd3e56ddff737b9245de7697d ] bcmgenet_set_mac_addr() did not check whether the provided MAC address is a valid Ethernet address before applying it. Userspace could configure an invalid address (such as all zeroes or a multicast address) while the interface is down. Add a call to is_valid_ether_addr() and return -EADDRNOTAVAIL if the MAC address is not valid. Fixes: 1c1008c793fa ("net: bcmgenet: add main driver file") Reviewed-by: Nicolai Buchwitz Signed-off-by: Florian Fainelli Link: https://patch.msgid.link/20260921220021.281418-5-florian.fainelli@broadcom.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit b3ebe8983d86fb83cc5bbeae63f627462fe1754d Author: Florian Fainelli Date: Mon Sep 21 15:00:19 2026 -0700 net: bcmgenet: do not skip WoL power up on GENET V1 [ Upstream commit cbbc1aee7776c7fa1d89e6cb963a23e58c495dca ] bcmgenet_power_up() had an early check for bcmgenet_has_ext(priv) before dispatching by power mode. GENET V1 does not have the EXT block (unlike GENET V2+), which causes bcmgenet_power_up() to immediately return 0. As a consequence, when waking up from GENET_POWER_WOL_MAGIC on GENET V1, bcmgenet_wol_power_up_cfg() is never invoked to disable the WoL clock, clear wake event masks, and restore normal PHY and MAC operations. Move the bcmgenet_has_ext() checks to the GENET_POWER_PASSIVE and GENET_POWER_CABLE_SENSE cases where the EXT registers are actually accessed, allowing GENET_POWER_WOL_MAGIC cleanup to execute on all hardware versions. Fixes: c3ae64ae0c08 ("net: bcmgenet: handle GENET_POWER_WOL_MAGIC") Reviewed-by: Nicolai Buchwitz Signed-off-by: Florian Fainelli Link: https://patch.msgid.link/20260921220021.281418-4-florian.fainelli@broadcom.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 21a0de48587790833748248248765c160b561006 Author: Florian Fainelli Date: Mon Sep 21 15:00:18 2026 -0700 net: bcmgenet: initialize u64 stats seq counter for all queues [ Upstream commit 3aeaa609fda19c09d5298c9fedaaa3b6229601b5 ] bcmgenet_gstrings_stats statically defines ethtool statistics for queues 0 through GENET_MAX_MQ_CNT (4). However, bcmgenet_probe() only initialized the u64_stats_sync seq counter up to priv->hw_params->rx_queues and priv->hw_params->tx_queues. Since priv->hw_params->rx_queues is 0 across all hardware versions (and priv->hw_params->tx_queues is 0 on GENET V1), rings 1..4 have uninitialized u64_stats_sync structures. When ethtool -S is run on 32-bit kernels, bcmgenet_get_ethtool_stats() reads stats from rx_rings[1..4], causing lockdep warnings due to the uninitialized sequence counters. Initialize the sequence counters for all GENET_MAX_MQ_CNT + 1 queues. Fixes: ffc2c8c4a714 ("net: bcmgenet: Initialize u64 stats seq counter") Reviewed-by: Nicolai Buchwitz Signed-off-by: Florian Fainelli Link: https://patch.msgid.link/20260921220021.281418-3-florian.fainelli@broadcom.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit aefa837c26674b159a64e8224cd738eac647dcfc Author: Florian Fainelli Date: Mon Sep 21 15:00:17 2026 -0700 net: bcmgenet: fix 64-bit RTNL stats reading in ethtool on 32-bit systems [ Upstream commit 0e2bec77ea62895416600c90588f593516572bca ] When bcmgenet was converted to 64-bit statistics, STAT_RTNL members were switched to point into struct rtnl_link_stats64, whose fields are 64-bit (__u64) regardless of architecture. However, bcmgenet_get_ethtool_stats() retained a legacy check: if (sizeof(unsigned long) != sizeof(u32) && s->stat_sizeof == sizeof(unsigned long)) On 32-bit systems, sizeof(unsigned long) == sizeof(u32), causing this condition to evaluate to false. As a result, 64-bit RTNL stats fields were read via *(u32 *)p. On 32-bit Big-Endian systems (such as MIPS BE), this reads the high 32 bits and returns 0 until the counter exceeds 4GB; on 32-bit Little-Endian systems (such as 32-bit ARM), the value is truncated to 32 bits. Fix this by checking if s->stat_sizeof == sizeof(u64) so 64-bit fields are always read as 64-bit values. Fixes: 59aa6e3072aa ("net: bcmgenet: switch to use 64bit statistics") Reviewed-by: Nicolai Buchwitz Signed-off-by: Florian Fainelli Link: https://patch.msgid.link/20260921220021.281418-2-florian.fainelli@broadcom.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 501baaeca7d128fa0a2f6953cf5873c5b0d7cf89 Author: Lorenzo Bianconi Date: Mon Sep 21 16:46:18 2026 +0200 net: stmmac: clear stale buf->page after recycling on skb build failure [ Upstream commit 0a7822e34a0bfde31b194ac3da3253e5032b44cc ] In stmmac_rx(), when napi_build_skb() fails the descriptor page is recycled back to the page pool with page_pool_recycle_direct(), but buf->page is left pointing at the recycled page, unlike every other consumption site in the function which clears the pointer after handing the page away. With the stale pointer stmmac_rx_refill() skips the replacement allocation and programs the already-recycled page back into the RX descriptor. Clear buf->page on the napi_build_skb() failure path to keep the buffer lifecycle consistent with the other consumption sites. Fixes: df542f669307 ("net: stmmac: Switch to zero-copy in non-XDP RX path") Signed-off-by: Lorenzo Bianconi Reviewed-by: Maxime Chevallier Link: https://patch.msgid.link/20260921-stmmac-fix-napi-build-skb-error-v1-1-3d54bf6d9bb6@oss.qualcomm.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 0f7ed27586acf65785a019e3c46bad31a1a3519f Author: Ivan Delalande Date: Fri Sep 18 15:47:15 2026 -0700 tg3: use random MAC address when tg3_get_device_address fails [ Upstream commit 4eb3f195ef08c5acaed87958297e41cc49588dde ] Some of the tg3 NICs we use (BCM57762) reset the SRAM MAC address to the placeholder address on link flaps, tg3_chip_reset, etc. We've typically fixed it from userspace, but since e4c00ba7274b ("tg3: replace placeholder MAC address with device property") was merged, tg3 just fails probe as we don't have a way to get it through the generic device_get_mac_address infrastructure as fallback on our systems. Make the driver assign a random address in this condition instead of being fatal for probe. Set deferred_probe_reason through dev_warn_probe if the address isn't yet available from the provider. Fixes: e4c00ba7274b ("tg3: replace placeholder MAC address with device property") Suggested-by: Jakub Kicinski Link: https://lore.kernel.org/netdev/20260909191751.651aa5c4@kernel.org/ Signed-off-by: Ivan Delalande Link: https://patch.msgid.link/20260918224715.GA654128@visor Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit e40990861074a51fe0149183abb0d58782c82aa2 Author: Johan Almbladh Date: Wed Sep 23 12:51:58 2026 +0200 bpf: Fix BSWAP 32 and 16 on MIPS64 [ Upstream commit 8110ba09777873443db286b3cbb89b0e6311c554 ] The 16/32-bit byteswap implementations for MIPS64r1 and earlier do not have an explicit zero extension afterwards. The input is first sign-extended to 64 bits, and the byteswap sequence can then leave the result sign-extended depending on the value of the low bits. Add the missing zero-extension. Found with test_bpf on MIPS64r1 emulated by QEMU. Fixes: fbc802de6b10 ("mips, bpf: Add new eBPF JIT for 64-bit MIPS") Signed-off-by: Johan Almbladh Signed-off-by: Alexei Starovoitov Link: https://patch.msgid.link/20260923105158.3514342-2-johan.almbladh@anyfinetworks.com Signed-off-by: Sasha Levin commit 876ef04c75c2751acfd355606fca3eced71dec65 Author: Johan Almbladh Date: Wed Sep 23 12:51:57 2026 +0200 bpf: Fix immediate JMP JEQ/JNE on MIPS32 [ Upstream commit db762fd96be225bd06161c9631160c755d891693 ] An addu instruction was emitted instead of addiu, causing the immediate value 1 to be interpreted as register $at. This made the comparison result invalid when the immediate operand was negative. Note that $at is mapped to BPF_REG_AX, which is used for constant blinding. Fix the instruction to use the immediate form. Found with test_bpf on MIPS32r1 emulated by QEMU. Fixes: eb63cfcd2ee8 ("mips, bpf: Add eBPF JIT for 32-bit MIPS") Signed-off-by: Johan Almbladh Signed-off-by: Alexei Starovoitov Link: https://patch.msgid.link/20260923105158.3514342-1-johan.almbladh@anyfinetworks.com Signed-off-by: Sasha Levin commit 6d76ddce4f1bae32d14642d0fe034bfdf8d7c797 Author: Ido Schimmel Date: Tue Sep 22 16:12:39 2026 +0300 vrf: Stop corrupting skb->csum when capturing CHECKSUM_COMPLETE packets [ Upstream commit ab7aa05c06ae340e5c7530bb78fa8d23794e460b ] The VRF device is an Ethernet device but it can have non-Ethernet ports such as IP tunnels. Before the cited commit, capturing packets from such ports on the VRF device resulted in these packets being detected as malformed since they lack an Ethernet header. The cited commit fixed it by pushing a dummy Ethernet header to such packets before the capture and pulling it afterwards. In the case of CHECKSUM_COMPLETE packets it also updated skb->csum with the checksum of the dummy Ethernet header. This is wrong as skb->csum should not include the checksum of the Ethernet header ("checksum of the _whole_ packet as seen by netif_rx()"). This also means that L4 protocols receive a corrupted skb->csum and potentially drop the packet, as is the case with UDP packets whose checksum was completed by software. Fix by removing the unnecessary call to skb_postpush_rcsum(). Fixes: 048939088220 ("vrf: add mac header for tunneled packets when sniffer is attached") Reported-by: Stefano Sasso Closes: https://lore.kernel.org/netdev/CALtE316UtL3x7LL6uxfXzx8rW6AbzYPeDOb478hqJCr_-dj=Wg@mail.gmail.com/ Signed-off-by: Ido Schimmel Reviewed-by: David Ahern Reviewed-by: Eric Dumazet Reviewed-by: Andrea Mayer Link: https://patch.msgid.link/20260922131239.2509494-1-idosch@nvidia.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit cfa643f2ca36c6f913fe1a2ceb12412ebb0fd962 Author: Shihuang Liu Date: Sat Sep 19 21:36:04 2026 +0800 net: skbuff: fix pull-bound underflow in skb_checksum_setup_ipv6() [ Upstream commit 3b4e0b0c008a8c1b474730248cd5b873026c74bd ] skb_maybe_pull_tail() subtracts skb_headlen(skb) from the unsigned max argument and passes the result to __pskb_pull_tail() as a signed int. The function does not ensure that max is at least skb_headlen(skb). This can happen while parsing IPv6 extension headers when an skb already has a linear area larger than MAX_IPV6_HDR_LEN. Once the parser needs data beyond the linear area, max - skb_headlen(skb) wraps and is converted to a negative delta. __pskb_pull_tail() then passes that negative length to skb_copy_bits(), where it can become a very large copy length. Pass the requested length itself as the pull bound at the three extension-header call sites, so the delta can no longer go negative. Fixes: 1431fb31ecba ("xen-netback: fix fragment detection in checksum setup") Suggested-by: Eric Dumazet Signed-off-by: Shihuang Liu Link: https://patch.msgid.link/20260919133604.50948-1-shlomojune6@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 25bb7c36225220d30f404d7e29d2e052bc5f4b99 Author: Jakub Kicinski Date: Mon Sep 21 16:18:56 2026 -0700 veth: manage XDP program pointers during channel resize [ Upstream commit 7104a370714346b667712913dc16abf14bbc97ed ] veth_set_channels() tears down XDP resources for removed RX queues without clearing rq->xdp_prog. If the program is then detached or replaced, those queues keep the old pointer after bpf_prog_put(). A later channel increase can re-enable NAPI and run the freed program. BUG: unable to handle page fault for address: ffffc90000256048 Oops: Oops: 0000 [#1] SMP KASAN NOPTI RIP: veth_xdp_rcv_skb (include/linux/filter.h:779 include/net/xdp.h:696 drivers/net/veth.c:820) Call Trace: veth_xdp_rcv (drivers/net/veth.c:941) veth_poll (drivers/net/veth.c:986) __napi_poll (net/core/dev.c:7787) net_rx_action (net/core/dev.c:7850 net/core/dev.c:8007) handle_softirqs (kernel/softirq.c:645) Kernel panic - not syncing: Fatal exception in interrupt Fixes: 4752eeb3d891 ("veth: implement support for set_channel ethtool op") Signed-off-by: Weiming Shi Acked-by: Stanislav Fomichev Reviewed-by: Jiayuan Chen Reviewed-by: Jason Xing Link: https://patch.msgid.link/20260921231856.1798630-1-kuba@kernel.org Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 63bcc2fee6419963e45e7c2cd698fa862409601d Author: Nicolai Buchwitz Date: Tue Sep 22 15:06:38 2026 +0200 net: bcmgenet: stop Tx NAPI before disabling the queues [ Upstream commit 7e87508b5c4d81210d0a736ed01962e52f5c4c56 ] bcmgenet_netif_stop() and the Wake-on-LAN branch of bcmgenet_suspend() both disable the Tx queues first and stop Tx NAPI several steps later. A completion in flight calls netif_tx_wake_queue() in between, and nothing stops the queue again, so a transmit can reach the rings after they have been freed. Close is safe because dev_deactivate_many() stops the qdisc first. bcmgenet_suspend() does not, so stop Tx NAPI before the queues on both paths. KASAN on a Raspberry Pi CM4, driven from an MTU change because suspend freezes user space before the callback runs: BUG: KASAN: use-after-free in bcmgenet_xmit+0x17f8/0x2258 Write of size 8 at addr ffffff8055844a68 by task ksoftirqd/0/14 bcmgenet_xmit+0x17f8/0x2258 dev_hard_start_xmit+0x13c/0x588 sch_direct_xmit+0x108/0x340 __dev_queue_xmit+0x1190/0x3848 Fixes: 254f3239dd07 ("net: bcmgenet: revise suspend/resume") Signed-off-by: Nicolai Buchwitz Link: https://patch.msgid.link/20260922130639.1660797-1-nb@tipi-net.de Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 284166f34b8fb2c9c0dfd39f611d6740b2591d1a Author: Victor Nogueira Date: Sun Sep 20 14:07:01 2026 -0300 net/sched: act_gate: budget the per-entry list in get_fill_size [ Upstream commit cfa165cbfbed9d0f4bbc22fef4309f595a3ab187 ] tcf_gate_get_fill_size returns only the TCA_GATE_PARMS size, but tcf_gate_dump also emits three 64-bit timestamps, the clock id, flags, priority and the variable-length TCA_GATE_ENTRY_LIST nest. The per-entry nest is unbounded: parse_gate_list places no cap on the number of sched-entries, so a gate with many entries can push the real dump well past the skb that tca_get_fill allocates from this size. RTM_NEWACTION then fails the add-notify with -EINVAL while the action is already committed to the IDR, and a subsequent RTM_GETACTION on the installed gate also returns -EINVAL because its dump no longer fits. Fix this by accounting for the missing fields in tcf_gate_get_fill_size along with all elements in the entries list. Note that sizing the reply from the action lets an oversized gate install cleanly for the first time: with the input unbounded by parse_gate_list, the sized skb can now grow well above NLMSG_GOODSIZE per netlink request (a transient GFP_KERNEL allocation reachable only with namespace-local CAP_NET_ADMIN). Overload from a malicious netns admin is hardening material, not net, per the discussion at https://lore.kernel.org/netdev/20260914191108.55a1a4f1@kernel.org/; a follow-up patch for net-next will cap the sched-entry count. Fixes: 4e76e75d6aba ("net sched actions: calculate add/delete event message size") Reported-by: Sashiko Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260824153903.4143642-1-victor@mojatatu.com Tested-by: hybris Co-developed-by: Jamal Hadi Salim Signed-off-by: Jamal Hadi Salim Signed-off-by: Victor Nogueira Link: https://patch.msgid.link/QDISC-3BLH.v1.20260914203033@mojatatu.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit e72cbedc4596e808d22dd2b9fff666ece76f8dd6 Author: Deepanshu Kartikey Date: Wed Sep 23 09:26:27 2026 +0530 nfc: pn533: fix OOB read in pn533_acr122_is_rx_frame_valid() [ Upstream commit b61732f47316d45f27706db7812950145d3327b5 ] frame->ccid.datalen is read directly from the USB response frame and used, unchecked, as an index into frame->data[]. A malicious or malfunctioning device can set this field to an arbitrary value, causing the driver to read far outside the received buffer. Bound ccid.datalen against the maximum possible ACR122 frame size before using it. This replaces the existing datalen == 0 check, since datalen < 2 already covers that case and additionally rejects datalen == 1, which would still underflow the "datalen - 2" offset used below. Fixes: 9815c7cf22da ("NFC: pn533: Separate physical layer from the core implementation") Reported-by: syzbot+1853daab1a47603d4678@syzkaller.appspotmail.com Closes: https://syzkaller.appspot.com/bug?extid=1853daab1a47603d4678 Tested-by: syzbot+1853daab1a47603d4678@syzkaller.appspotmail.com Assisted-by: LLM Signed-off-by: Deepanshu Kartikey Link: https://patch.msgid.link/20260923035627.6210-1-kartikey406@gmail.com Signed-off-by: David Heidelberg Signed-off-by: Sasha Levin commit 8978dff632519712d1fa432991e1f2e857f90239 Author: Ömer Mete Kaya Date: Tue Sep 8 19:18:01 2026 +0300 nfc: llcp: fix slab-out-of-bounds reads when logging service names [ Upstream commit 7dcf371a35632f035baf77bcf2c129165f772ce4 ] nfc_llcp_wks_sap() and nfc_llcp_build_sdreq_tlv() pass non-null- terminated strings to pr_debug() using the %s format specifier. The buffers are allocated via kmemdup() or come from netlink attributes and are not guaranteed to be null-terminated, causing __dynamic_pr_debug() to read beyond the allocated region: KASAN: slab-out-of-bounds Read in __dynamic_pr_debug Fix both call sites by using %.*s with the explicit length to limit the output to the actual length of the string. Fixes: d9b8d8e19b07 ("NFC: llcp: Service Name Lookup netlink interface") Reported-by: syzbot+1e3df0852e82c21ca418@syzkaller.appspotmail.com Closes: https://syzkaller.appspot.com/bug?extid=1e3df0852e82c21ca418 Signed-off-by: Ömer Mete Kaya Link: https://patch.msgid.link/20260908161952.731468-1-omermetekaya0@gmail.com Signed-off-by: David Heidelberg Signed-off-by: Sasha Levin commit f234197f3bda3aea5d4b1e591d6bcb3660a8b3fc Author: Ömer Mete Kaya Date: Wed Sep 9 15:14:31 2026 +0300 nfc: llcp: fix WKS SAP hijacking via prefix match in nfc_llcp_wks_sap() [ Upstream commit 408cff6bd60636df201274d320edfdfde9ed41db ] nfc_llcp_wks_sap() compares only service_name_len bytes, so a short service_name like "u" matches longer WKS strings like "urn:nfc:sn:snep". Fix by requiring exact length match before strncmp(). Fixes: d646960f7986 ("NFC: Initial LLCP support") Signed-off-by: Ömer Mete Kaya Link: https://patch.msgid.link/20260909121437.33744-1-omermetekaya0@gmail.com Signed-off-by: David Heidelberg Signed-off-by: Sasha Levin commit 2386d107df6d217ca08c8166f07c689f24f98df1 Author: Ömer Mete Kaya Date: Wed Sep 9 15:16:22 2026 +0300 nfc: llcp: fix -ENOMEM on connect with zero-length service name [ Upstream commit c04981e42d94f39c1dba965cc462a046e946a6c5 ] When service_name_len is 0, kmemdup() returns ZERO_SIZE_PTR which passes the NULL check, causing nfc_llcp_send_connect() to attempt building a zero-length service name TLV and fail with -ENOMEM. Fix by setting service_name to NULL directly when service_name_len is 0. Fixes: d646960f7986 ("NFC: Initial LLCP support") Signed-off-by: Ömer Mete Kaya Link: https://patch.msgid.link/20260909122029.34081-1-omermetekaya0@gmail.com Signed-off-by: David Heidelberg Signed-off-by: Sasha Levin commit fdd95e819d828c975aaeb258a3fed5dd7fd136d0 Author: Pengpeng Hou Date: Sun Aug 30 21:29:58 2026 +0800 nfc: st21nfca: validate ISO15693 inventory length [ Upstream commit 7f2ea5ed588c03d481f0301e6c3d4240132383fb ] The ISO15693 inventory helper removes a two-byte prefix without checking that it exists, then accepts a one-byte remainder before reading data[1] as the DSFID. Require the prefix and at least two remaining bytes before copying the UID data and reading the DSFID. Fixes: 7974728094d3 ("NFC: st21nfca: Add ISO15693 Reader/Writer support") Signed-off-by: Pengpeng Hou Link: https://patch.msgid.link/20260830132958.6397-1-pengpeng@iscas.ac.cn Signed-off-by: David Heidelberg Signed-off-by: Sasha Levin commit cbde7f5d326172d6dbb55cca0df42fa7e1ed8115 Author: Cong Nguyen Date: Mon Sep 14 19:11:29 2026 +0700 nfc: llcp: fix sdreq TLV list leak on parse/alloc/send failure [ Upstream commit 66f4300206b82b0b143ef0d9be90cd8d29f23a47 ] nfc_genl_llc_sdreq() builds a list of TLV nodes while walking nested netlink attrs, but 3 error paths (nested-attr parse failure, TLV alloc ENOMEM, nfc_llcp_send_snl_sdreq() failure) all skip freeing what was already queued. Route them through a new free_list label, mirroring the SDRES path in the same file which already does this. Harmless on the success path too -- send_snl_sdreq() drains the list as it moves nodes, so it's already empty by the time free_list runs. Fixes: d9b8d8e19b07 ("NFC: llcp: Service Name Lookup netlink interface") Assisted-by: Claude:claude-opus-4 Signed-off-by: Cong Nguyen Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260914121129.2098606-1-congnt264@gmail.com Signed-off-by: David Heidelberg Signed-off-by: Sasha Levin commit 0b32dcc8a3f9471316e41a3e7bcdf946ed7b1bed Author: Chris Gellermann Date: Fri Sep 4 18:42:52 2026 +0200 nfc: virtual_ncidev: Add missing ioctl compat handler [ Upstream commit 51814683e28fc64eceb415962376956c3cfc75a7 ] The compat handler for ioctls to the virtual nci device is missing. So, nci-specific ioctls of a compat task return with -1 and errno set to ENOTTY. Add a handler. The handling of an ioctl() call of a compat task to get the index of virtual nci device (IOCTL_GET_NCIDEV_IDX) lands in the default case of the ioctl compat handler (see fs/ioctl.c): COMPAT_SYSCALL_DEFINE3(ioctl, ...) { ... default: error = do_vfs_ioctl(fd_file(f), fd, cmd, ...); if (error != -ENOIOCTLCMD) break; if (fd_file(f)->f_op->compat_ioctl) error = fd_file(f)->f_op->compat_ioctl(fd_file(f), cmd, arg); if (error == -ENOIOCTLCMD) error = -ENOTTY; ... } There, do_vfs_ioctl() returns -ENOIOCTLCMD and compat_ioctl is not set for virtual_ncidev_fops, i.e. f_op->compat_ioctl == NULL. So, the ioctl() syscall returns with -1 and errno set to ENOTTY to the compat task. To fix this, use the compat_ptr_ioctl helper for compat handling here. It shall be used for ioctls that "either ignore the argument or pass a pointer to a compatible data type". The driver's sole ioctl takes a user void pointer and copies nfc_dev->idx to it, a 4-byte integer across all ABIs. This issue has been found by running the nci_dev kernel selftest as rv64 binary on top of a CHERI kernel, where the ioctl() ends up in the ioctl compat handler, similar to a 32-bit application on top of a 64-bit kernel. Fixes: e624e6c3e777 ("nfc: Add a virtual nci device driver") Signed-off-by: Chris Gellermann Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260904164252.18351-1-christian.gellermann@codasip.com Signed-off-by: David Heidelberg Signed-off-by: Sasha Levin commit 27cf04ba4e7ef9392dfbbc239df82b86166234ea Author: Chris Gellermann Date: Fri Sep 4 11:59:15 2026 +0200 selftests/nci: Fix out-of-bounds store on thread join [ Upstream commit 6be581aeffc215bfc77939cd59902b0dbc4af23e ] The NCI test collects the exit status of its helper threads by passing the address of an int to pthread_join(): int status; ... pthread_join(thread_t, (void **) &status); pthread_join() stores a void pointer to the memory location. On 64-bit systems, a void pointer is wider than an int, so the store overruns the 4 bytes of space allocated on the stack for the integer and corrupts the adjacent stack. On our CHERI system, this caused a fault due to a capability bounds violation. Fix this by introducing a helper that joins a thread through a void pointer and converts the result back to an integer, which is what the helper threads return. While here, also fix the logic in disconnect_tag() if the helper thread creation failed. Previously, it would have joined a thread that was never created when pthread_create() failed. Fixes: f595cf1242f3 ("selftests: Add nci suite") Signed-off-by: Chris Gellermann Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260904095915.3372241-1-christian.gellermann@codasip.com Signed-off-by: David Heidelberg Signed-off-by: Sasha Levin commit f9cbe5c68d93a0c950c6258abe6bb55406f55969 Author: Chaithanya Lagisetty Date: Tue Sep 1 07:06:18 2026 +0000 selftests: nci: Fix uninitialized family ID on missing attribute [ Upstream commit eda518d2cdb6074a0bcdfa06af291616bcb5c421 ] get_family_id() walks the generic netlink CTRL_CMD_GETFAMILY reply looking for the CTRL_ATTR_FAMILY_ID attribute and returns the parsed value in the local variable "id". If the reply does not carry that attribute, the parsing loop never assigns "id" and the function returns an indeterminate stack value, which the caller stores in self->fid and uses for subsequent netlink requests. Initialize "id" to 0 so a missing attribute yields a deterministic (invalid) family ID instead of a garbage value. Fixes: f595cf1242f3 ("selftests: Add nci suite") Signed-off-by: Chaithanya Lagisetty Reviewed-by: Hangbin Liu Link: https://patch.msgid.link/20260901070618.3299012-1-nagachaithanya9911@gmail.com Signed-off-by: David Heidelberg Signed-off-by: Sasha Levin commit da8a07757383614a8380a0f7dd54749d2687b82e Author: Lee Jones Date: Wed Sep 2 12:30:31 2026 +0000 nfc: llcp: Fix race condition in accept_queue lifecycle [ Upstream commit c3eef2f988a3db9690369d7cef9a3344dd9788d3 ] In nfc_llcp_socket_release(), sockets and listener accept queues are walked under the local sockets rwlock and bh_lock_sock(). However, bh_lock_sock() does not synchronise against process-context lock_sock() held by nfc_llcp_accept_dequeue() during accept(). Because socket_release() does not check sock_owned_by_user(), both paths can concurrently unlink and release the same child socket, resulting in use-after-free or a NULL pointer dereference of child->parent in nfc_llcp_accept_unlink(). Fix this synchronisation race by having nfc_llcp_socket_release() use process-context lock_sock() instead of bh_lock_sock(): 1. Pop sockets from the local sockets list under the write lock using nfc_llcp_sock_list_pop() so lock_sock() can be acquired without holding the rwlock. 2. Because lock_sock() can sleep, defer the final release of the nfc_llcp_local structure to a workqueue (release_work). This avoids a sleeping-in-atomic bug when the last local reference is dropped from softirq context. Additionally, hold a single device reference on local from registration until final destruction. 3. In nfc_llcp_local_get(), use kref_get_unless_zero() to prevent resurrecting a local object whose teardown has been scheduled. 4. In llcp_sock_accept(), verify that the listener socket state is still LLCP_LISTEN after waking from schedule_timeout() to prevent hangs if the listener is closed concurrently. 5. When unlinking unaccepted child sockets during listener release, unlink them from local->sockets, call sock_orphan(), and drop their initial sk_alloc creation reference via sock_put(). 6. Make nfc_llcp_accept_unlink() idempotent by guarding parent access with a NULL check. Fixes: 50b78b2a6500 ("NFC: Fix sleeping in atomic when releasing socket") Signed-off-by: Lee Jones Link: https://patch.msgid.link/20260902123033.1169067-1-lee@kernel.org Signed-off-by: David Heidelberg Signed-off-by: Sasha Levin commit 3316903fdc3a306fa56b3ac3297c2464a858c5b9 Author: Lei Zhu Date: Wed Jul 29 15:24:26 2026 +0800 selftests: nci: Correct pthread_create return value check [ Upstream commit 3d8afc5243ea2ee803d98e69eb4a01748167ac1b ] The pthread_create() functions returns 0 on success and a positive value on failure. Modify the return value check to correctly detect failure cases. Fixes: 72696bd8a09d ("selftests: nci: Extract the start/stop discovery function") Signed-off-by: Lei Zhu Link: https://patch.msgid.link/20260729072426.303484-1-zhulei_szu@163.com Signed-off-by: David Heidelberg Signed-off-by: Sasha Levin commit b94cf09720543cc34256c7aa88f8059cbe7c0419 Author: Aldo Ariel Panzardo Date: Thu Jul 16 20:26:57 2026 -0300 nfc: llcp: Fix list corruption / refcount desync in nfc_llcp_recv_dm() [ Upstream commit bf1460acdf8cf5a07c819f59785d40f20d113099 ] nfc_llcp_recv_dm() handles DM(NOBOUND)/DM(REJ) for a socket that is still linked on local->connecting_sockets: it looks the socket up with nfc_llcp_connecting_sock_get(), sets sk->sk_state = LLCP_CLOSED and returns, without taking the socket lock and without unlinking the socket from the connecting_sockets list. llcp_sock_release() selects the list to unlink from by sk_state: a socket in LLCP_CONNECTING is unlinked from connecting_sockets, otherwise from the sockets list. Because recv_dm left the socket physically on connecting_sockets but in the LLCP_CLOSED state, release() takes the else branch and calls nfc_llcp_sock_unlink(&local->sockets, sk). That runs sk_del_node_init() while holding sockets.lock, i.e. it removes the socket from the connecting_sockets hlist under the wrong lock. A concurrent connect() linking another socket onto connecting_sockets under connecting_sockets.lock then mutates the same hlist unserialized, which corrupts the list and desyncs the sk_add_node()/sk_del_node_init() sock_hold()/__sock_put() pairing. An unprivileged local process holding LLCP sockets, with the DM supplied by the remote peer over an established LLCP link, can drive this to leak kernel sockets without bound (the mis-decrement goes through the non-freeing __sock_put() path, so the object is never released), leading to memory exhaustion / DoS. This is the same class of bug that was fixed in the sibling handler nfc_llcp_recv_cc() by commit b493ea2765cc ("nfc: llcp: Fix use-after-free race in nfc_llcp_recv_cc()"); recv_dm did not receive the equivalent fix. Fix it the same way: take lock_sock(), re-check that the socket is still hashed (release() may have won the race), and for the NOBOUND/REJ case unlink it from connecting_sockets before moving it to LLCP_CLOSED. The unlink drops the connecting_sockets membership reference via sk_del_node_init(), leaving the socket unhashed, so the later nfc_llcp_sock_unlink() in llcp_sock_release() becomes a no-op and no double put occurs. Fixes: a69f32af86e3 ("NFC: Socket linked list") Signed-off-by: Aldo Ariel Panzardo Link: https://patch.msgid.link/20260716232657.203145-1-qwe.aldo@gmail.com Signed-off-by: David Heidelberg Signed-off-by: Sasha Levin commit 57919d0662d8856211b794140172e51ca389f23d Author: Pengpeng Hou Date: Wed Jul 15 16:44:05 2026 +0800 nfc: st21nfca: validate received frame size [ Upstream commit a653c01ce447f10c36b901646888c0330363af4f ] st21nfca_hci_i2c_repack() trims a received frame at its EOF marker before removing byte stuffing. It then assumes the truncated frame contains the LLC header and two CRC bytes, and it unconditionally reads the byte after an escape marker. A malformed frame can place EOF immediately after the start marker or can end its data portion with an escape marker. The former leaves too few bytes for check_crc(), while the latter makes the unstuffing loop read past the current skb length. Require the minimum framing bytes both before and after unstuffing. Use separate input and output cursors while removing byte stuffing, and reject an escape marker without its encoded byte. This keeps malformed frames within the received frame boundary before CRC processing. Fixes: 3096e25a3e40 ("NFC: st21nfca: Fix incorrect byte stuffing revocation") Signed-off-by: Pengpeng Hou Link: https://patch.msgid.link/20260715084405.41546-1-pengpeng@iscas.ac.cn Signed-off-by: David Heidelberg Signed-off-by: Sasha Levin commit ed2e00bb22c73156bf8030b6204deabb25cf4eb5 Author: Pengpeng Hou Date: Wed Jul 15 16:43:25 2026 +0800 nfc: nfcmrvl: validate helper command length before pull [ Upstream commit 686f942332b1667f13f3b8d6a2f50bcfbf42e277 ] The firmware download receive path removes the NCI data header and reads the helper command before validating the remaining packet length. A short frame can therefore reach the data access before the malformed packet is rejected. Validate the complete helper command length before stripping the NCI data header. Fixes: 3194c6870158 ("NFC: nfcmrvl: add firmware download support") Signed-off-by: Pengpeng Hou Link: https://patch.msgid.link/20260715084325.40276-1-pengpeng@iscas.ac.cn Signed-off-by: David Heidelberg Signed-off-by: Sasha Levin commit 940b626854de200e6187777d42114727daca617c Author: Weiming Shi Date: Sun Sep 20 21:23:04 2026 +0800 bpf: Reject dev-bound-only programs on other devices [ Upstream commit 6db1ce73e9853f533eb7f413f14ba00f8ec6f80d ] __bpf_offload_dev_match() falls back to comparing offdev pointers after an exact netdev mismatch. Bound-only programs normally have NULL offdevs, so unrelated netdevs compare equal. A bound-only program on an offload-registered netdev can instead inherit a real offdev and match a sibling port. With CAP_BPF and CAP_NET_ADMIN, a caller can use bpf(BPF_LINK_CREATE) with a different target ifindex to run metadata kfuncs specialized for the bound driver on the target driver's xdp_buff. Running a veth-bound program on tun reads beyond tun's bare stack xdp_buff as a veth_xdp_buff. Oops: general protection fault, probably for non-canonical address KASAN: null-ptr-deref in range [0x0000000000000010-0x0000000000000017] RIP: 0010:veth_xdp_rx_timestamp (drivers/net/veth.c:1673) Call Trace: ... tun_build_skb (drivers/net/tun.c:1739) tun_get_user (drivers/net/tun.c:1856) tun_chr_write_iter (drivers/net/tun.c:2091) vfs_write (fs/read_write.c:595 fs/read_write.c:687) ksys_write (fs/read_write.c:739) do_syscall_64 (arch/x86/entry/syscall_64.c:84) entry_SYSCALL_64_after_hwframe (arch/x86/entry/entry_64.S:121) Kernel panic - not syncing: Fatal exception in interrupt Restrict non-offloaded programs to exact netdev matches and retain the shared-offdev fallback only for genuinely offloaded multi-port programs. Fixes: 2b3486bc2d23 ("bpf: Introduce device-bound XDP programs") Reported-by: Signed-off-by: Weiming Shi Signed-off-by: Alexei Starovoitov Link: https://lore.kernel.org/bpf/20260917161335.1020405-2-bestswngs@gmail.com/ Link: https://patch.msgid.link/20260920132303.4109240-3-bestswngs@gmail.com Signed-off-by: Sasha Levin commit 9923e1c537f91f3530b52d4efce85f3d9c131bf8 Author: Bernard Ladenthin Date: Fri Sep 18 15:39:37 2026 +0200 net/sched: fix potential stack infoleak in em_text_dump() [ Upstream commit 9c572a83037a7dcd653ba3a9cc468c16b857d0c9 ] em_text_dump() allocates struct tcf_em_text on the stack without zeroing it. strscpy() writes the algorithm name and a NUL terminator into conf.algo[], leaving the remaining bytes uninitialised. nla_put_nohdr() then copies the full struct to the netlink response. KMSAN on Linux 7.2-rc6 reports two kernel-infoleak splats from this path, one triggered via "tc filter show" and one via a raw RTM_GETTFILTER dump: BUG: KMSAN: kernel-infoleak in _copy_to_iter+0x1c9/0x2620 nla_put_nohdr+0x83/0x130 em_text_dump+0x291/0x550 Local variable conf created at: em_text_dump+0x5d/0x550 Bytes 168-179 of 199 are uninitialized I am not certain whether this constitutes a real security problem in practice: the test was conducted in a controlled KMSAN environment and the leaked stack bytes may or may not carry sensitive data on actual production kernels. I am reporting it because KMSAN flagged it as a kernel-infoleak and the fix is straightforward. I can provide a userspace reproducer on request. The original code used strncpy() which zero-pads to the destination size. Commit b04202d6065c ("net/sched: replace strncpy with strscpy") replaced it with strscpy(), which does not pad, creating this condition. Zero-initialising the struct closes it. Fixes: b04202d6065c ("net/sched: replace strncpy with strscpy") Link: https://lore.kernel.org/netdev/20250327143733.187438-1-richard120310@gmail.com/ Assisted-by: Claude:claude-sonnet-4-6 [KMSAN] Signed-off-by: Bernard Ladenthin Acked-by: Jamal Hadi Salim Link: https://patch.msgid.link/20260918133953.12494-1-bernard.ladenthin@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 92c9f2ea634d92adff5403d231ebeda683c5685a Author: Björn Töpel Date: Fri Sep 18 13:46:40 2026 +0200 eth: fbnic: Avoid rounding zero ring sizes [ Upstream commit 0160953d8eec75c3c55562158c46442ff1fd410b ] roundup_pow_of_two() is undefined for zero. ethtool permits a zero ring size to reach the driver, where the minimum-size check should reject it. Leave zero unchanged while rounding nonzero ring sizes. The minimum-size check then rejects zero deterministically without changing the established behavior for other values. Fixes: 6cbf18a05c06 ("eth: fbnic: support ring size configuration") Reported-by: Sashiko Link: https://lore.kernel.org/netdev/178971206933.22033.236948278674126701@kernel.org/ Suggested-by: Alexander Duyck Signed-off-by: Björn Töpel Reviewed-by: Joe Damato Link: https://patch.msgid.link/20260918114641.1281172-1-bjorn@kernel.org Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 3ea177bdba25f5e56f82260ee3b6b4123e8d1c97 Author: Maxime Chevallier Date: Thu Sep 17 23:53:38 2026 +0200 net: stmmac: selftests: Account for alignment shift on dwmac1000 for Jumbo test [ Upstream commit c4ac6e94eb9423126bda907a7f2933284f7450ff ] On dwmac1000, we currently only support single-descriptor frames. The Jumbo test started failing when NET_IP_ALIGN was added to align the IP header, as this tests tries to send the biggest possible frame. On dwmac1000 the DMA transfer is aligned on 4-bytes, so adding a 2-byte shift at the start-of-buffer address means it takes a whole extra 4-byte DMA burst to receive the Jumbo packet, causing it to spill over the next descriptor. This doesn't seem to happen on dwmac4 and xgmac that appear to correctly handle unaligned xfers (only tested on dwmac4) Let's account for that in the Jumbo test, reduce the size of our big packet by the align size. Fixes: 23680bf5f8c6 ("net: stmmac: restore NET_IP_ALIGN in the RX DMA offset") Reviewed-by: Nicolai Buchwitz Signed-off-by: Maxime Chevallier Link: https://patch.msgid.link/20260917215339.2022523-8-maxime.chevallier@bootlin.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit b17959fb4d55f6a277fd14759dfaf45a96ff6e34 Author: Maxime Chevallier Date: Thu Sep 17 23:53:36 2026 +0200 net: stmmac: dwmac4: Use the correct bufzise when the len is exactly 8K [ Upstream commit b42e7012773a0e81e97a2dda6ef907f5147a6658 ] DMA bufsize selection isn't made on the MTU but the actual frame length, so including the L2 header. On DWMAC4, if the len is exactly BUF_SIZE_8KiB, the next larger size is incorrectly selected. Lets fix the comparison and while at it, rename the parameter from len to mtu. Fixes: c3efed5ad1b0 ("net: stmmac: Enable dwmac4 jumbo frame more than 8KiB"). Signed-off-by: Maxime Chevallier Reviewed-by: Nicolai Buchwitz Link: https://patch.msgid.link/20260917215339.2022523-6-maxime.chevallier@bootlin.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 4c4217aaa5e4bb733919784f85d9a72f75f81a49 Author: Maxime Chevallier Date: Thu Sep 17 23:53:35 2026 +0200 net: stmmac: selftests: Capture all packets for vlan checks [ Upstream commit 960db6f65788c21249ea04a947d5c01e38d19294 ] While we use vlan_vid_add to trigger the tag filtering machinery in the driver, there's no netdev associated to the VLAN. This causes the skb to arrive with empty skb->vlan_tci fields, as the packet is marked OTHERHOST in __netif_receive_skb_core(), and we fail our validation. Let's use the proxy mechanism introduced for DSA, that registers a ETH_P_ALL packet handler that runs earlier, before the vlan netdev lookup, then filters for the correct ethertype before passing an skb clone to our validation function. As we may receive external frames with the right tag from the outside, let's move the address check in the vlan validation function earlier. Fixes: 091810dbded9 ("net: stmmac: Introduce selftests support") Reviewed-by: Nicolai Buchwitz Signed-off-by: Maxime Chevallier Link: https://patch.msgid.link/20260917215339.2022523-5-maxime.chevallier@bootlin.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit e7a8bbb6149b4f9dbfd0d12c5bdca404fbf45aac Author: Maxime Chevallier Date: Thu Sep 17 23:53:34 2026 +0200 net: stmmac: selftests: Check the dev->features for S-TAG offload testing [ Upstream commit ba804b23d76d278ee475b8427fa7c5623ce5e270 ] The S-TAG offload insertion incorrectly checks the dvlan (double vlan) DMA cap, which is different than S-TAG support. Use NETIF_F_HW_VLAN_STAG_TX to check if the feature is supported instead. Note that this flag isn't set in stmmac yet, but contrary to ARP offload, this is a feature that has a chance to get there eventually so let's leave the selftest here for now. It'll report -EOPNOTSUPP in the meantime. Fixes: 091810dbded9 ("net: stmmac: Introduce selftests support") Reviewed-by: Nicolai Buchwitz Signed-off-by: Maxime Chevallier Link: https://patch.msgid.link/20260917215339.2022523-4-maxime.chevallier@bootlin.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 74da0c1cede8ae7f5dd17021a0307c21dc0e06c9 Author: Maxime Chevallier Date: Thu Sep 17 23:53:32 2026 +0200 net: stmmac: selftests: Support running selftests on DSA conduits [ Upstream commit d68acbf93531abdb5b02b21994cd4c15a3c95b42 ] Most stmmac selftests rely on dev_add_pack() to add custom handlers, that validate the packets sent to ourselves through MAC loopback. However, when the stmmac-driven interface is a DSA CPU conduit, all frames that are received have ETH_P_XDSA as a protocol, even though they don't actually contain any tag as they come from the loopback and not the switch. This will prevent any incoming packet to match our packet handlers. Let's register a ETH_P_ALL packet handler when we detect that we're a DSA conduit, and use a proxy packet handler to filter the h_proto. As this allows external frames to be received through our .func(), the packet handler is added after the dev->addr field is populated in our selftest attributes. Note that we may still receive incoming packets from the switch, but these frames shouldn't interfere with the very specific frames used for selftests, and stmmac selftests in general aren't safe against external traffic interferences. This was validated on a WPQ864 devkit for IPQ8064, that has the SoC connected to a QCA8k switch. The ARP offload's packet handler is left alone, this feature is just not implemented in stmmac and due for removal. Fixes: 091810dbded9 ("net: stmmac: Introduce selftests support") Reviewed-by: Nicolai Buchwitz Signed-off-by: Maxime Chevallier Link: https://patch.msgid.link/20260917215339.2022523-2-maxime.chevallier@bootlin.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit d18cb8fd1b6ba95c85de644203cfe1b071ff0c0f Author: Xin Long Date: Mon Sep 21 14:03:45 2026 -0400 sctp: hold asoc or transport before mod_timer() in timer handlers [ Upstream commit cae23ae3f7887a1cf8a75da38edcebeef695040f ] Take the association or transport reference before rearming a timer in the timer handlers. The existing code calls mod_timer() before taking the reference needed by the rearmed timer without holding the sock lock. This creates a race with timer cleanup: if the timer is deleted after mod_timer() returns but before the reference is taken, the cleanup path can drop the timer's reference and destroy the transport or association. The timer handler then takes a reference on the already freed object and eventually drops it, causing a refcount underflow. Hold the object before mod_timer() and drop the reference if mod_timer() reports that the timer was already pending in timer handlers. Apply the same ordering to the proto-unreachable path, which can rearm a transport timer outside the timer handlers without holding the sock lock. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Reported-by: Tangxin Xie Signed-off-by: Xin Long Link: https://patch.msgid.link/c31b5e3ee2b7274e804f5eba2f21e2412e7eef7a.1790013825.git.lucien.xin@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 650dcd938e5632e3858cf1f3b9efbc7642c4fd0b Author: Bernardo Soares Date: Fri Sep 18 10:59:31 2026 +0100 net/mlx5: Bridge, don't fail unlink of untracked/unsupported peer ports [ Upstream commit 2e51097c982b7b22382fda3202ba29f0ea33e8c0 ] mlx5_esw_bridge_vport_unlink() returns -EINVAL when the port isn't tracked by this instance's br_offloads. This is reachable on a sibling instance that registered its notifier after the port was already enslaved: it never saw the NETDEV_CHANGEUPPER link event, so peer_link() never created a peer port for it, but it does see the later unlink event and fails. Return 0 instead, and give mlx5_esw_bridge_vport_peer_unlink() the same merged_eswitch capability guard peer_link() already has, since without it peer_link() likewise never creates a port to unlink. This also matters beyond the -EINVAL itself: mlx5_esw_bridge_switchdev_port_event() runs on the per-netns netdev_chain, and notifier_from_errno(-EINVAL) sets NOTIFY_STOP_MASK, which call_netdevice_notifiers_info() checks to stop calling further listeners on that chain - so the old -EINVAL silently dropped the event for any listener registered later on the same chain, even though none of it was visible to user space since __netdev_upper_dev_unlink() discards the return value. Fixes: c358ea1741bc ("net/mlx5: Bridge, allow merged eswitch connectivity") Signed-off-by: Bernardo Soares Reviewed-by: Mark Bloch Link: https://patch.msgid.link/20260918095931.29792-3-bsoares.it@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit cdeca5de4e18630ca927c4a90bfeed70bea51266 Author: Bernardo Soares Date: Fri Sep 18 10:59:30 2026 +0100 net/mlx5: Bridge, don't fail switchdev events of sibling eswitch ports [ Upstream commit 35e6f970f553954d92ba20afa885139a8e7dd0d7 ] mlx5 registers the bridge offload switchdev notifiers once per eswitch instance, but the notifier chains are global, so every instance sees every event and must filter out the ones that aren't its own. The existing filter, mlx5_esw_bridge_dev_same_hw(), only checks that the event netdevice sits on the same HCA - intentional for merged eswitch, where one bridge can span representors of several eswitches on one HCA - but same-HCA doesn't mean the instance actually has that port: peer ports are only created reactively from NETDEV_CHANGEUPPER, so an instance brought up after a sibling PF's port was already enslaved has none. The port object and attribute handlers claim the event anyway once same-HW passes, then fail the port lookup and return -EINVAL, which gets reported to user space even though the owning instance already handled it (e.g. "bridge vlan add ... RTNETLINK answers: Invalid argument"). Fix by filtering on the tracked port instead. The same gap exists in the generic recursive lower-device walk used by attribute changes on a bridge with more than one representor enslaved directly: mlx5_esw_bridge_lower_rep_vport_num_vhca_id_get() is entered with the bridge master netdevice, falls through to its generic netdev_for_each_lower_dev() loop, and returns as soon as the recursion into any one lower device yields a non-NULL rep - the underlying base case, mlx5_esw_bridge_rep_vport_num_vhca_id_get(), only checks mlx5_esw_bridge_dev_same_hw(), not ownership by the calling instance's br_offloads. mlx5_esw_bridge_lag_rep_get(), used for the LAG-master case, already filters on mlx5_esw_bridge_dev_same_esw() per candidate and so cannot select a sibling's rep; it is not the source of this bug. On a merged-eswitch HCA with a bridge spanning representors of more than one eswitch instance directly, the walk can return a sibling's rep instead of continuing to the one the calling instance actually owns, so the attribute change fails the same way as above. Fix by checking mlx5_esw_bridge_port_exists() at the point each rep is picked, same as the previous fix did for the notifier filter. Fixes: c358ea1741bc ("net/mlx5: Bridge, allow merged eswitch connectivity") Signed-off-by: Bernardo Soares Cc: Vlad Buslov Cc: Saeed Mahameed Reviewed-by: Mark Bloch Link: https://patch.msgid.link/20260918095931.29792-2-bsoares.it@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit cea597519b9ca6d691c98d92460ead8ece2a07af Author: Zhao Gongyi Date: Thu Sep 17 20:10:16 2026 +0800 bpf, sockmap: Reject max_entries > INT_MAX in sock_map_alloc [ Upstream commit 814a81c842bd88f6bd8a4ce550d560df071a5d03 ] sock_map_alloc() only rejects max_entries == 0 and otherwise allows any u32 value. sock_map_free() then walks the sks[] array with a signed int iterator: int i; for (i = 0; i < stab->map.max_entries; i++) struct sock **psk = &stab->sks[i]; When a SOCKMAP is created with max_entries = 0xffffffff (UINT_MAX), the allocation of 32 GiB can succeed on large-memory hosts. During free the counter reaches 0x80000000, wraps to INT_MIN, is sign-extended by movslq and turned into a ~16 GiB negative offset from stab->sks, pointing far below the allocation. The faulting access is an xchg() write in sock_map_free(). Without KASAN, the same out-of-bounds write can fault on an unmapped vmalloc page or corrupt an unrelated allocation if that vmalloc address is populated. On a KASAN kernel with CONFIG_KASAN_VMALLOC=y, the shadow check for that address hits an unmapped shadow page and oopses first: BUG: unable to handle page fault for address: fffff521b59c5a00 RIP: 0010:kasan_check_range+0x107/0x190 Call Trace: sock_map_free+0x93/0x190 map_create+0x68d/0xb30 __sys_bpf+0x21e/0x2e70 Vmcore confirmed stab->map.max_entries == 0xffffffff, stab->sks == 0xffffc911ace2d000, and the faulting address sks + (s64)INT_MIN * 8 exactly at 0xffffc90dace2d000. The same buggy path is reached on the normal close()/bpf_map_free_deferred() path whenever such a map is destroyed. sock_map_alloc() used to bound its allocation size through bpf_map_charge_init(), but the bound was dropped when rlimit-based memory accounting was removed. Reject max_entries > INT_MAX at creation time so the signed iterator in sock_map_free() never sees a value that would overflow. Triggered by syzkaller and reproduced on both a 6.6-based KASAN kernel and the upstream v7.3-rc2 kernel. Fixes: 0d2c4f964050 ("bpf: Eliminate rlimit-based memory accounting for sockmap and sockhash maps") Signed-off-by: Zhao Gongyi Signed-off-by: Alexei Starovoitov Link: https://patch.msgid.link/20260917121016.48171-1-zhaogongyi@bytedance.com Signed-off-by: Sasha Levin commit d95dd5a6c6cd08d5ae466d3aed27b3cf59155a6c Author: Emil Tsalapatis Date: Tue Sep 22 17:20:20 2026 +0000 bpf: Fix bpf_sock context code generation [ Upstream commit 4a4852376e3a2727ea40e61143d6d7c22bb6dfad ] Currently, the ctx access code reads the rx_queue_mapping field with either a 4-byte or 2-byte load. The rest of the bits in the register are marked known zero by the verifier. However, the emitted ctx access code places in the register on certain the special value (-1) using BPF_MOV_IMM64, which gets sign-extended to turn on all the bits in the register. By shifting this value right, the program ends up with a value at runtime above what the verifier assumes is possible. Fix this by ensuring the read value is as wide as the assumed size. Use MOV32 instructions instead of MOV64 instructions to keep the upper bits zero as assumed by the verifier. Also properly report the size of the destination variable (the bpf_sock field, 4 bytes) instead of the source (the socket field, 2 bytes). Fixes: c3c16f2ea6d2 ("bpf: Add rx_queue_mapping to bpf_sock") Reported-by: Nicholas Carlini Suggested-by: Nicholas Carlini Signed-off-by: Emil Tsalapatis Signed-off-by: Alexei Starovoitov Reviewed-by: Jiayuan Chen Link: https://patch.msgid.link/20260922172028.6269-4-emil@etsalapatis.com Signed-off-by: Sasha Levin commit 1196903a27cac5c929c7517e987a8d159f20c5e5 Author: Emil Tsalapatis Date: Tue Sep 22 17:20:18 2026 +0000 bpf: Fix bounds check for skb-backed dynptrs [ Upstream commit ed6eec97b534979dcf28b40c389cee57bd6561d4 ] The skb_pointer_if_linear() function checks whether a memory region of length len starting at offset off into the skb is in the linear area, and returns a pointer to the region if so. The check currently subtracts between skb_headlen and offset of the check, and since skb_headlen is unsigned the subtraction can underflow. This causes the bounds check to spuriously pass and generate an arbitrary pointer of the form *(skb->data + off). The only user of this helper is currently skb-backed BPF dynptr code. Returning the wrong pointer leads to the dynptr erroneously being backed with invalid memory. Ensure the subtraction cannot underflow, and fail the check if it would. Use u64 arithmetic to also prevent overflow when calculating (skb_headlen(skb) - off) since off is unsigned. Fixes: 6f5a630d7c57 ("bpf, net: Introduce skb_pointer_if_linear().") Reported-by: Nicholas Carlini Signed-off-by: Emil Tsalapatis Signed-off-by: Alexei Starovoitov Reviewed-by: Jiayuan Chen Link: https://patch.msgid.link/20260922172028.6269-2-emil@etsalapatis.com Signed-off-by: Sasha Levin commit fb1e2df8e2e28876d4ed2c671bc24fcd71ed3ac2 Author: Manaf Meethalavalappu Pallikunhi Date: Tue Sep 22 17:49:02 2026 +0530 thermal: gov_step_wise: Fix stale mitigation vote with non-zero lower bounds [ Upstream commit ec0d89150a9381d591344a9f6f5428655c227a7f ] When two or more thermal zones bind to a common cooling device and one zone uses a non-zero instance->lower value, there is a bug where the instance holds a stale mitigation vote even after its trip is cleared. Problem scenario: - thermal-zone1: Trip at 50°C, cooling-map with lower=0 - thermal-zone2: Trip at 55°C, cooling-map with lower=2 - Both zones share the same cooling device (e.g., CPU) Issue flow: 1. Both trips trigger, zone1 requests state 5, zone2 also mitigates 2. Zone2 trip clears (temp < 53°C due to hysteresis) 3. When throttle=false and trend=THERMAL_TREND_DROPPING: - Current code checks: if (cur_state <= instance->lower) return THERMAL_NO_TARGET - Since cur_state (5) > instance->lower (2), it returns instance->lower (2) - This is the BUG where it returns instance->lower even though trip is cleared 4. Zone2's passive polling stops (tz->passive reaches 0) - no more updates for zone2 5. Zone2's stale vote of 2 persists indefinitely 6. Even when zone1 wants to reduce cooling to state, the cooling device cannot go below state 2 due to zone2's stale vote When a trip is cleared (throttle == false), always return THERMAL_NO_TARGET instead of instance->lower. Remove the unnecessary check comparing cur_state with instance->lower. Since passive polling is already deactivated when the trip is cleared, the instance should always be deactivated regardless of its current cooling state. This ensures that instances with non-zero lower bounds do not retain stale mitigation votes after their trips are cleared. Fixes: 042a3d80f118 ("thermal: core: Move passive polling management to the core") Signed-off-by: Manaf Meethalavalappu Pallikunhi Link: https://patch.msgid.link/20260922-step_wise_multi_zone_stale_vote_fix-v1-1-789f68dab229@oss.qualcomm.com Signed-off-by: Rafael J. Wysocki Signed-off-by: Sasha Levin commit 0f5d8953cb63a1c3c8fe1b88b31c57fdb4793dc6 Author: Coia Prant Date: Sun Sep 20 01:20:21 2026 +0800 net: pcs: xpcs: fix clock reference leak on xpcs_init_clks failure [ Upstream commit 9892d71cf0ce3ff3d4fed2d9a3968fd4feb1c918 ] xpcs_init_clks() takes references with clk_bulk_get_optional() and then enables them with clk_bulk_prepare_enable(). If the enable step fails, the function returns without dropping the references. xpcs_create() handles the failure through out_free_data, which calls xpcs_free_data() but never xpcs_clear_clks(), so the clk references are leaked. Add the missing clk_bulk_put() on the enable failure path. The prepare/enable side is already rolled back by clk_bulk_prepare_enable() itself. Fixes: f6bb3e9d98c2 ("net: pcs: xpcs: Add Synopsys DW xPCS platform device driver") Signed-off-by: Coia Prant Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260919172021.2336748-1-coiaprant@gmail.com Signed-off-by: Paolo Abeni Signed-off-by: Sasha Levin commit 54e98316ce5c183b64b7f616e19d0eb2224d1dec Author: Jakub Kicinski Date: Fri Sep 18 15:29:48 2026 -0700 genetlink: report the real command id for dump-only ops in policy dumps [ Upstream commit 261e8a37ecbaf462cdf9c336d2b2f5056088401a ] The op-to-policy map a CTRL_CMD_GETPOLICY dump returns is the only way for userspace to find out which policy index belongs to which command. ctrl_dumppolicy_put_op() tags the nest with doit->cmd, but an op which only has a dumpit has no doit and every path which fills the split ops in zeroes it out, so those entries all claim to be command 0. nlctrl's own CTRL_CMD_GETPOLICY and NETDEV_CMD_QSTATS_GET are both in that group: [{'family-id': 16, 'op-policy': {'do': 0, 'dump': 0, 'op-id': 3}}, {'family-id': 16, 'op-policy': {'dump': 1, 'op-id': 0}}, ctrl_fill_info() gets this right - it uses the iterator's cmd for CTRL_ATTR_OP_ID - so the two introspection interfaces of the same family contradict each other today. Pass the command in rather than reconstructing it from doit->cmd | dumpit->cmd inside the helper, both callers already have it. Fixes: 26588edbef60 ("genetlink: support split policies in ctrl_dumppolicy_put_op()") Signed-off-by: Jakub Kicinski Link: https://patch.msgid.link/20260918222949.4190284-1-kuba@kernel.org Signed-off-by: Paolo Abeni Signed-off-by: Sasha Levin commit 41e444535c60f1dd6f8c625fd66dcfe77ba3aeb8 Author: bui duc phuc Date: Fri Sep 18 11:28:04 2026 +0700 net: ethernet: ti: netcp: fix pm_runtime usage counter leak on error [ Upstream commit ac4334522e4ba4a3b6710dd5d4cc98092824b8ca ] pm_runtime_get_sync() leaves the runtime PM usage counter incremented even when it fails, but the error path in netcp_probe() does not call pm_runtime_put_noidle() to balance it, leaking a reference each time resume fails. Use pm_runtime_resume_and_get() instead, which automatically drops the usage counter on failure, fixing the leak. Fixes: 84640e27f230 ("net: netcp: Add Keystone NetCP core ethernet driver") Signed-off-by: bui duc phuc Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260918042804.13101-1-phucduc.bui@gmail.com Signed-off-by: Paolo Abeni Signed-off-by: Sasha Levin commit ca9986cc824f2feb79b4a83e5f6f4c24c2140448 Author: Muhammad Bilal Date: Sun Sep 20 00:19:37 2026 +0500 net: spacemit: clear TX descriptor on fragment mapping failure [ Upstream commit 2d14720beb58870b15a52b236c3ab0be0e06e915 ] emac_tx_mem_map() writes TX_DESC_0_OWN into the ring descriptor for every slot beyond old_head as soon as that slot's memset()'d local copy is committed with "*tx_desc_addr = tx_desc", i.e. before the buffers for that slot have necessarily all been mapped successfully. If emac_tx_map_frag() then fails on a later fragment, the err_free_skb path calls emac_free_tx_buf() to unmap and drop the skb, but leaves the already-written descriptor memory untouched, and tx_ring->head is never advanced past old_head (the "tx_ring->head = head" store is skipped by the goto). So a slot between old_head and the rolled-back head can be left with TX_DESC_0_OWN set and buffer_addr_{1,2} pointing at DMA mappings that emac_free_tx_buf() just tore down, while software considers that slot free again. The next successful emac_tx_mem_map() call only rebuilds old_head itself; if the DMA engine auto-advances into the following descriptor once it finishes old_head's packet, it will fetch that stale, already-unmapped address. emac_tx_clean_desc() already treats emac_free_tx_buf() and clearing the descriptor as a pair when reclaiming completed descriptors; do the same in the mapping failure path. Fixes: bfec6d7f2001 ("net: spacemit: Add K1 Ethernet MAC") Signed-off-by: Muhammad Bilal Reviewed-by: Vivian Wang Reviewed-by: Troy Mitchell Link: https://patch.msgid.link/20260919191937.271202-1-meatuni001@gmail.com Signed-off-by: Paolo Abeni Signed-off-by: Sasha Levin commit 9c6af88fbefbe77608b77409df23a1617e28398c Author: Mikhail Zaslonko Date: Wed Sep 16 18:06:26 2026 +0200 s390/debug: Fix NULL pointer dereference in debug_info_copy() [ Upstream commit 012bfcd5a51082d5a65f096dfb9ca652b5267965 ] When debug_register_static() fails, it clears areas, active_pages and active_entries but leaves the area bounds unchanged. Copying such an area, either by opening its view file or via debug_dump(), makes debug_info_copy() dereference the NULL pointers. Skip the copy loop when the source has no areas. Closes: https://lore.kernel.org/r/20260903132123.12F271F00A3F@smtp.kernel.org Fixes: d72541f94512 ("s390/debug: add early tracing support") Signed-off-by: Mikhail Zaslonko Reviewed-by: Peter Oberparleiter Acked-by: Heiko Carstens Signed-off-by: Heiko Carstens Signed-off-by: Sasha Levin commit 36831afe8eba058535884c8fe81481e41454e0db Author: Mikhail Zaslonko Date: Fri Sep 18 17:13:51 2026 +0200 s390/debug: Do not register views for failed static debug areas [ Upstream commit 28e29992b034acffc9342df216c06097825ce610 ] __REGISTER_STATIC_DEBUG_INFO() calls debug_register_view() unconditionally, even when debug_register_static() has failed. In that case _debug_register() was never reached and id->debugfs_root_entry is still NULL, so debugfs_create_file() places the view file in the debugfs root directory. For sclp_err this leaves a /sys/kernel/debug/hex_ascii file with nothing to indicate which debug log it belongs to. debug_register_static() is not exported and the macro is its only caller, so let it return an error code and skip the view registration when it fails. No debugfs files are created for such an area then. Reproduce by booting with s390dbf=sclp_err::100000000. The sclp_err registration fails, no s390dbf/sclp_err/ directory is created, and a hex_ascii file appears in the debugfs root instead. Fixes: d72541f94512 ("s390/debug: add early tracing support") Signed-off-by: Mikhail Zaslonko Reviewed-by: Heiko Carstens Signed-off-by: Heiko Carstens Signed-off-by: Sasha Levin commit 8b661a045082ca756dfd09a4d47f7c1466deb472 Author: Nemesa Garg Date: Wed Sep 9 16:33:32 2026 +0530 drm/i915/psr: Clear stale sel fetch enable bits on sel fetch disable [ Upstream commit 2777ec9852277a06ae68fee0c4f1a32783e4a999 ] Selective fetch is dropped while pipe CRC is active, and the planes keep their SEL_FETCH_PLANE_CTL / SEL_FETCH_CUR_CTL enable bit set in hardware over that. A plane disabled while selective fetch is off never gets the bit cleared, as the disable path is guarded by enable_psr2_sel_fetch. Once selective fetch comes back the hardware resumes fetching for a plane that is no longer enabled and keeps its DDB range reserved. Clear the bits as selective fetch is turned off instead. Atomic check has both the old and the new crtc state, so record the transition there and let the plane and cursor arm paths write the registers to 0 for that commit. v2: Drop the old_crtc_state->hw.active check. [Jouni] Fixes: b1f5279b5981 ("drm/i915/psr: Move plane sel fetch configuration into plane source files") Closes: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/8739 Assisted-by: Copilot:Claude-Opus-5 Signed-off-by: Nemesa Garg Reviewed-by: Jouni Högander Signed-off-by: Suraj Kandpal Link: https://patch.msgid.link/20260909110332.3528029-3-nemesa.garg@intel.com (cherry picked from commit a4c0e7f80429eda6990960971aebd4e4b9533cc6) Signed-off-by: Jani Nikula Signed-off-by: Sasha Levin commit 6515c41503eef849abd2c7364f34e4bb25110f90 Author: Myeonghun Pak Date: Thu Sep 17 14:33:36 2026 -0400 tg3: clean up PHYLIB resources on probe failure [ Upstream commit a92e1a412c53dc0d9ad639e7abf8b3fc70a5b6ad ] tg3_get_invariants() can register an MDIO bus and connect a PHY for USE_PHYLIB devices. If tg3_init_one() later fails, its common error path releases the mappings and netdev without undoing those PHYLIB resources. Disconnect the PHY and unregister the MDIO bus before the remaining teardown. Guard PHY cleanup with USE_PHYLIB to match tg3_phy_init(), and call tg3_mdio_fini() unconditionally to match tg3_mdio_init(). The existing IS_CONNECTED and MDIOBUS_INITED flags make both helpers safe when initialization only completed partially. This issue was identified during our ongoing static-analysis research while reviewing kernel code. Fixes: 158d7abdae85 ("tg3: Add mdio bus registration") Assisted-by: OpenAI:GPT-5.6 Co-developed-by: Ijae Kim Signed-off-by: Ijae Kim Signed-off-by: Myeonghun Pak Link: https://patch.msgid.link/20260917183336.36239-1-mhun512@gmail.com Signed-off-by: Paolo Abeni Signed-off-by: Sasha Levin commit c21daf37d5ce9a4e234b7c42238ac1e6041e28e4 Author: Shardul Bankar Date: Thu Sep 17 14:55:33 2026 +0530 udp: remove a disconnected socket from the 4-tuple hash table [ Upstream commit 9e95b1a94c9c49b4ba722251bbca2759b9c51737 ] A UDP socket bound to a specific address and port keeps its entry in the 4-tuple hash table after it is disconnected: sk binds to 127.0.0.1:21001 sk connects to 127.0.0.2:20001 // filed in the 4-tuple table sk disconnects, connect(AF_UNSPEC) // still filed, peer now 0.0.0.0:0 __udp_disconnect() takes a socket out of that table only as a side effect of ->rehash() or ->unhash(), and it skips ->rehash() when SOCK_BINDADDR_LOCK is set and ->unhash() when SOCK_BINDPORT_LOCK is set. commit 6996a2d2d0a6 ("udp: Unhash auto-bound connected sk from 4-tuple hash table when disconnected.") fixed the same end state for a wildcard-bound socket, by a path this one does not take. The entry is counted whether or not anything hits it. hash4_cnt on the hash2 slot stays raised for as long as the socket lives, so udp_has_hash4() keeps sending every packet for that address and port through the 4-tuple lookup first. On IPv6 it can also be hit. __udp_disconnect() does not clear sk_v6_daddr, so udp_v6_rehash() files the entry under the peer the socket was connected to with a zero dport, and inet6_match() compares that same field: a datagram from the former peer with a zero source port matches, and source port zero is accepted on receive. On IPv4 the peer is cleared, so a match would need a zero source address as well, which the routing layer rejects as martian. The stale sk_v6_daddr is a separate defect, not addressed here; removing the entry closes this path either way. The entry can also be relocated. __udp_disconnect() clears sk_bound_dev_if, so a subsequent SO_BINDTODEVICE calls ->rehash(), and because the receive address is still specific udp_lib_rehash() moves the entry instead of removing it, into the bucket that (rcv_saddr, num, 0, 0) hashes to -- a pure function of the address and port, so every socket reaching this state on one address and port collects in one bucket. The bucket cannot be chosen from outside, as udp_ehashfn() is seeded with a per-boot secret. This last one became reachable only with commit 644f9108f3a5 ("udp: Make rehash4 independent in udp_lib_rehash()"), which moved the hash4 handling out of a branch a disconnected socket does not take; the stale entry itself dates from the commit in Fixes. Take the socket out of the table before __udp_disconnect() runs, while it still matches how it was filed. This also reaches the wildcard case ahead of udp_lib_rehash()'s udp_unhash4() branch, leaving that branch unreachable from udp_disconnect(); removing it belongs in net-next. udp_disconnect() and udp_abort() are the only UDP entries into __udp_disconnect(), which is shared with raw, ping and l2tp sockets that are not struct udp_sock: ping_prot.obj_size is sizeof(struct inet_sock), so udp_hashed4() on one would read past the allocation. Fixes: 78c91ae2c6de ("ipv4/udp: Add 4-tuple hash for connected socket") Assisted-by: LLM Signed-off-by: Shardul Bankar Reviewed-by: Kuniyuki Iwashima Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-2-718891af0d7a@mpiricsoftware.com Signed-off-by: Paolo Abeni Signed-off-by: Sasha Levin commit 7a0bbe5654ecba9b289927ca9cdec5ec2afcfa55 Author: Shardul Bankar Date: Thu Sep 17 14:55:32 2026 +0530 udp: relocate a connected socket in the 4-tuple hash table on re-connect [ Upstream commit 5fd0783b99d4af98f65cd58b56ec203d1d426104 ] A connected UDP socket that connects again to a different peer is not re-filed in the 4-tuple hash table: sk binds to 127.0.0.1:21001 sk connects to 127.0.0.2:20001 // filed under hash(sk, peer1) sk connects to 127.0.0.3:20002 // still filed under hash(sk, peer1) packet from 127.0.0.3:20002 // hash(sk, peer2) misses, so the // lookup falls back to scoring the // hash2 chain for this address // and port udp_lib_hash4() returns early when the socket is already hashed, assuming ->rehash() relocates it. ->rehash() runs from __ip{4,6}_datagram_connect() only while the receive address is unset, which a second connect never is: the first connect assigns it, whether the socket was bound to a specific address or to the wildcard. commit 644f9108f3a5 ("udp: Make rehash4 independent in udp_lib_rehash()") added that early return and named connect(AF_UNSPEC) as the way around it. That workaround does not help a socket with both SOCK_BINDADDR_LOCK and SOCK_BINDPORT_LOCK set, because __udp_disconnect() skips ->rehash() for the first and ->unhash() for the second. Delivery is correct either way. Relocate the socket when the hash it is filed under differs from the one requested, which is what commit 78c91ae2c6de ("ipv4/udp: Add 4-tuple hash for connected socket") did before the early return became unconditional. It is done here under hslot->lock, which that version did not take, to match udp_lib_rehash() and udp_lib_unhash(). hslot2 is unchanged, so hash4_cnt needs no adjustment, as in udp_lib_rehash(). A first connect is unaffected, and IPv6 shares the code. With 500 sockets on the port, a re-connected socket measured 522,553 pps without this change and 2,055,078 with it. The UDP side was noted as remaining work in [1]. Link: https://lore.kernel.org/netdev/apnHqmYZQ4yzOP4N@v4bel/ [1] Fixes: 644f9108f3a5 ("udp: Make rehash4 independent in udp_lib_rehash()") Assisted-by: LLM Signed-off-by: Shardul Bankar Reviewed-by: Kuniyuki Iwashima Link: https://patch.msgid.link/20260917-udp_hash4_fix_v1-v1-1-718891af0d7a@mpiricsoftware.com Signed-off-by: Paolo Abeni Signed-off-by: Sasha Levin commit ed902d5546fb9c142b3216d8eec914fa1bc0aa80 Author: Kuniyuki Iwashima Date: Sun Sep 20 19:14:32 2026 +0000 ipv6: Fix dst leak for uncached routes. [ Upstream commit be31fe6333f534155e6b408f1ef6d77974bb41aa ] ip6_route_output_flags(), ip6_rt_put_flags(), and ip6_dst_check() detect an uncached route by list_empty(&rt->dst.rt_uncached), which replaced the static DST_NOCACHE flag check in commit a4c2fd7f7891 ("net: remove DST_NOCACHE flag"). When a device is unregistered, rt6_uncached_list_flush_dev() unlinks uncached routes tied to the device from rt6_uncached_list. Previously, they were moved to another list with list_move() (__list_del_entry() + list_add()), and since commit 98aa546af5e4 ("inet: remove (struct uncached_list)->quarantine"), the routes are just unlinked with list_del_init(). If list_del_init() runs concurrently, list_empty() evaluates to true; ip6_route_output_flags() calls dst_hold_safe() incorrectly and ip6_rt_put_flags() skips ip6_rt_put(), leaking dst, and thus dev tied via rt->from as well. The same race is partially fixed by commit 9a6f0c4d5796 ("dst: fix races in rt6_uncached_list_del() and rt_del_uncached_list()"). Let's check rt6->dst.rt_uncached_list instead. Note that IPv4 does not have the same issue. Fixes: 98aa546af5e4 ("inet: remove (struct uncached_list)->quarantine") Signed-off-by: Kuniyuki Iwashima Reviewed-by: Hangbin Liu Reviewed-by: Xuanqiang Luo Reviewed-by: Ido Schimmel Reviewed-by: Eric Dumazet Link: https://patch.msgid.link/20260920191558.2990636-1-kuniyu@google.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 6ab6c9d120b30c94579192ef83a581a3dda6d79c Author: Pengpeng Hou Date: Sun Sep 20 11:47:45 2026 +0800 net: usb: sr9700: include receive overhead in the length check [ Upstream commit c06bde80ae7a7b595732f7cabcb92cf08db9d56a ] The receive fixup subtracts the Ethernet CRC from the reported packet length, but compares that payload length against the whole remaining receive buffer. The following copy starts after the three-byte header, and the cursor advance consumes both that header and the four-byte CRC. Require the payload to fit after SR_RX_OVERHEAD before copying it or advancing to the next packet. The loop already ensures that the remaining buffer is larger than the overhead, so the subtraction is safe. The issue was found by our static-analysis tool. Fixes: c9b37458e956 ("USB2NET : SR9700 : One chip USB 1.1 USB2NET SR9700Device Driver Support") Reviewed-by: Ethan Nelson-Moore Tested-by: Ethan Nelson-Moore Signed-off-by: Pengpeng Hou Link: https://patch.msgid.link/20260920034745.18468-1-hppiscas@163.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit e3c21e8b5390603ddbdf8bdcd1efc5de3eac65b5 Author: Ivan Vecera Date: Thu Sep 17 16:37:36 2026 +0200 dpll: use exact lookup for reference sync pin id [ Upstream commit 7cce782d8327b7291334c4a304cf3fd909a74d9d ] dpll_pin_ref_sync_state_set() looks up the reference sync pin in the pin->ref_sync_pins xarray, which is keyed by the sync pin's id (see dpll_pin_ref_sync_pair_add() using xa_insert() with ref_sync_pin->id). The pin id to operate on is supplied by userspace via DPLL_A_PIN_ID. The lookup however used xa_find() with a ULONG_MAX limit, which returns the first present entry with an index greater than or equal to the requested id, not the entry stored exactly at that id. If userspace passes an id that is not paired as a reference sync pin, but another pin with a higher id is present in the xarray, xa_find() silently returns that wrong pin and the subsequent ref_sync_set() operates on it. The request only fails when the given id is larger than every present key. Use xa_load() for an exact-key lookup instead, mirroring the deletion path in dpll_pin_ref_sync_pair_del(). Fixes: 58256a26bfb3 ("dpll: add reference sync get/set") Signed-off-by: Ivan Vecera Link: https://patch.msgid.link/20260917143736.526221-1-ivecera@redhat.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 649ffc76ed5483b2caf35ed90d3787538efc4e53 Author: Kuniyuki Iwashima Date: Fri Sep 18 08:22:05 2026 +0000 ipv6: Prevent rt6_insert_exception() for dying fib6_info. [ Upstream commit 0346ec2f080b40d95ed05b853bb9226289e75212 ] Before the cited commit, fib6_nh_flush_exceptions() always set from->exception_bucket_flushed = 1 under rt6_exception_lock to prevent rt6_insert_exception() from inserting a new exception for a dying fib6_info. The flag was replaced with the FIB6_EXCEPTION_BUCKET_FLUSHED bit stored in nh->rt6i_exception_bucket. The problem is that now the bit is only set when the bucket is not NULL and fib6_nh_flush_exceptions() is called from fib6_nh_release() after fib6_ref has already reached zero. If rt6_insert_exception() is called while the target fib6_info is being removed via fib6_purge_rt(), a new exception could be created successfully because rt6_flush_exceptions() no longer sets the bit. This creates a reference cycle between the fib6_info and the exception route, leaking the fib6_info, its nexthop device, and all per-CPU routes in fib6_nh->rt6i_pcpu, which stalls netdev unregistration. [ 34.680602] unregister_netdevice: waiting for gre6 to become free. Usage count = 68 [ 44.920675] unregister_netdevice: waiting for gre6 to become free. Usage count = 68 [ 55.176582] unregister_netdevice: waiting for gre6 to become free. Usage count = 68 Let's call fib6_drop_pcpu_from() before rt6_flush_exceptions(), to set fib6_destroying before rt6_exception_lock, and check f6i->fib6_destroying in rt6_insert_exception(). Note that FIB6_EXCEPTION_BUCKET_FLUSHED logic is dead and we can clean it up in net-next. Fixes: cc5c073a693f ("ipv6: Move exception bucket to fib6_nh") Signed-off-by: Kuniyuki Iwashima Reviewed-by: Ido Schimmel Link: https://patch.msgid.link/20260918082209.2853582-1-kuniyu@google.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit ff8cff3aff673613c39449328731859c2b99af7f Author: Ratheesh Kannoth Date: Wed Sep 16 07:51:11 2026 +0530 octeontx2-af: Fix memory scaling limitation in SR-IOV mode [ Upstream commit d06f2ebf67ff2962fe00d687e4f0d4703eb41a12 ] The original code used DMA_ATTR_FORCE_CONTIGUOUS, which could exhaust the CMA pool when a large number of VFs were requested. Fix this by switching to the DMA streaming API. This is equivalent on Octeon platforms, which provide full I/O coherency via the SMMU. Cc: Leon Romanovsky Fixes: 73d33dbc0723 ("octeontx2-af: Use DMA_ATTR_FORCE_CONTIGUOUS attribute in DMA alloc") Signed-off-by: Ratheesh Kannoth Reviewed-by: Leon Romanovsky Link: https://patch.msgid.link/20260916022111.1083017-1-rkannoth@marvell.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit be00deee60044ffa820f9436415caeb665b2cda0 Author: Hui Peng Date: Sat Sep 19 11:25:14 2026 +0000 Bluetooth: RFCOMM: Reject short EA=0 frames in rfcomm_recv_frame() [ Upstream commit 6d91041bb38b97e2feb625123cc0529d7b83a0e1 ] While rfcomm_recv_frame() verifies that skb->len is at least sizeof(*hdr) + 1 (4 bytes: 3-byte header + 1-byte FCS), an RFCOMM frame with an extended 2-byte length field (!__test_ea(hdr->len)) has a 4-byte header plus a 1-byte FCS (5 bytes minimum, sizeof(*hdr) + 2). When a 4-byte RFCOMM frame with EA == 0 arrives: 1. The initial skb->len < sizeof(*hdr) + 1 check passes (4 < 4 is false). 2. Trimming the FCS byte decrements skb->len to 3. 3. If __check_fcs() succeeds, skb_pull(skb, 4) fails (4 > 3) and returns NULL without advancing skb->data. 4. Because the return value of skb_pull() is ignored, the un-pulled 3-byte struct rfcomm_hdr remains at skb->data and is either queued as application payload via rfcomm_recv_data() or parsed as a multiplexer control command via rfcomm_recv_mcc() on DLCI 0. Fix this by extending the length check in rfcomm_recv_frame() to also require skb->len >= sizeof(*hdr) + 2 when !__test_ea(hdr->len). Fixes: b230e5bf501c ("Bluetooth: RFCOMM: validate skb length in rfcomm_recv_frame") Assisted-by: LLM Signed-off-by: Hui Peng Signed-off-by: Luiz Augusto von Dentz Signed-off-by: Sasha Levin commit 8ec5838a35b03ac0cf3602452a369898051ffaf9 Author: Ravindra Date: Tue Sep 15 10:42:15 2026 +0530 Bluetooth: btintel_pcie: validate device-supplied DMA indices [ Upstream commit 37a11129345337efd6eef8e62b03b6348cd0dd8b ] In btintel_pcie_msix_rx_handle(), the driver processes RX completion descriptors (urbd1) written by the PCIe device into DMA-coherent memory. urbd1->frbd_tag (a 16-bit field fully controlled by the device firmware via DMA) is used directly as an array index into rxq->bufs[] without any bounds check. rxq->bufs[] has only BTINTEL_PCIE_RX_DESCS_COUNT (64) entries, while frbd_tag can be any value 0-65535. A malicious or malfunctioning device can write an out-of-range frbd_tag, causing the driver to dereference an out-of-bounds data_buf pointer. Additionally, cr_hia is read from a DMA-shared index array also writable by the device; if the device sets cr_hia >= rxq->count, the while-loop never terminates because cr_tia is wrapped via modulo rxq->count and can never equal an out-of-range cr_hia. Add bounds validation for cr_hia and frbd_tag in the RX path, and cr_hia in the TX path. Log invalid values with bt_dev_err before returning. Fixes: c2b636b3f788 ("Bluetooth: btintel_pcie: Add support for PCIe transport") Signed-off-by: Ravindra Signed-off-by: Luiz Augusto von Dentz Signed-off-by: Sasha Levin commit 8f31031326994d1abfcc6ebb6927ef96e51fb9ff Author: Hui Peng Date: Sat Sep 19 22:17:38 2026 +0000 Bluetooth: bnep: fix out-of-bounds reads on short RX/TX frames and control fallthrough [ Upstream commit f0ca020cbb9bb7f3f4ea8ba1dfcf30a282aec91e ] Fix multiple out-of-bounds reads in Bluetooth BNEP frame processing: 1. In bnep_rx_frame() and bnep_ctrl_frame() (net/bluetooth/bnep/core.c), use pskb_may_pull() to verify the BNEP header, control type byte, filter count, and extension headers exist before reading them, and return 0 after handling BNEP_CONTROL instead of falling through to Ethernet frame submission when no extension headers follow. 2. In bnep_net_xmit() (net/bluetooth/bnep/netdev.c), verify skb->len >= ETH_HLEN with pskb_may_pull() before reading the 14-byte Ethernet header to prevent an out-of-bounds heap read and infoleak on short AF_PACKET TX frames. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Assisted-by: LLM Signed-off-by: Hui Peng Signed-off-by: Luiz Augusto von Dentz Signed-off-by: Sasha Levin commit a8f3a829f3a53dec4e7480b8eabb2c5abd265de5 Author: Ralf Lici Date: Fri Aug 28 15:00:09 2026 +0200 ovpn: reject invalid peer VPN addresses [ Upstream commit 5940f3407b78062442cb01f541ef6eed709fc380 ] In MP mode, ovpn uses peer VPN addresses as lookup keys for selecting the peer that should receive outgoing tunnel packets. The netlink configuration path currently accepts address values that cannot sensibly identify a VPN peer, such as multicast, broadcast or loopback addresses. Reject invalid peer VPN addresses when creating or updating an MP peer. Keep accepting the unspecified address as the internal unset value, provided that at least one VPN address family remains configured. Fixes: 1d36a36f6d53 ("ovpn: implement peer add/get/dump/delete via netlink") Signed-off-by: Ralf Lici Signed-off-by: Antonio Quartulli Signed-off-by: Sasha Levin commit 7efbc134253c12f5bd641342460c526793d6acf0 Author: Ralf Lici Date: Fri Aug 28 15:00:08 2026 +0200 ovpn: reject multipeer peers without VPN addresses [ Upstream commit 025af3a0a892514f9f27f186338ba3d44365547a ] In MP mode, ovpn uses the peer VPN addresses to select the peer for outgoing tunnel packets. Peer creation currently requires a VPN IPv4 or IPv6 attribute, but it only checks for the presence of the attribute and not for a usable address value. This allows userspace to create an MP peer with only unspecified VPN addresses, or to update an existing peer so that both VPN address families become unspecified. Such a peer cannot be selected through the VPN address hash tables. Reject MP peer creation or update when the resulting peer would not have at least one VPN address configured. This changes such configurations from being accepted to being rejected, but they have never been usable because the peer cannot be selected through the VPN address hash tables. Fixes: 1d36a36f6d53 ("ovpn: implement peer add/get/dump/delete via netlink") Signed-off-by: Ralf Lici Signed-off-by: Antonio Quartulli Signed-off-by: Sasha Levin commit daab972893aec78e034932794573ec8c0581565d Author: Ralf Lici Date: Fri Aug 28 15:00:07 2026 +0200 ovpn: reject duplicate peer VPN addresses [ Upstream commit d25e885b31a0f2808d936f95c9a558a8a792669b ] In MP mode, ovpn uses the peer VPN addresses as lookup keys for selecting the peer that should receive an outgoing tunnel packet. However, the netlink peer configuration path does not currently reject duplicate VPN addresses. If two peers are configured with the same VPN address, both can be inserted in the VPN address hash table and lookups return whichever peer is found first. This makes peer selection ambiguous and dependent on hash insertion order. Reject peer creation or update when the resulting VPN address is already assigned to another peer. Ignore unspecified addresses because those are not inserted in the VPN address hash tables. This changes such configurations from being accepted to being rejected, but they have never worked reliably because peer selection is ambiguous. Fixes: 1d36a36f6d53 ("ovpn: implement peer add/get/dump/delete via netlink") Signed-off-by: Ralf Lici Signed-off-by: Antonio Quartulli Signed-off-by: Sasha Levin commit 12ee72319805adf3e8e620f78f85190d6d3908a3 Author: Ralf Lici Date: Fri Aug 28 15:00:06 2026 +0200 ovpn: always unhash old VPN addresses before rehashing [ Upstream commit b43beccb3713fafada57814b0a652f4a876eb75f ] ovpn_peer_hash_vpn_ip updates the per-peer VPN address hash entries after userspace changes a peer VPN address. The current code removes an old hash entry only when the new address for that family is not the unspecified address. When an address is cleared to 0.0.0.0 or ::, its hash node therefore remains linked in the bucket selected by the old address. The address comparison performed during lookup prevents the old address from matching, but the table retains a stale entry until the peer is removed or another address is configured for that family. Always remove both old VPN address hash entries before conditionally adding the currently configured addresses back. This ensures that a cleared address leaves its hash node unhashed. Fixes: 1d36a36f6d53 ("ovpn: implement peer add/get/dump/delete via netlink") Signed-off-by: Ralf Lici Signed-off-by: Antonio Quartulli Signed-off-by: Sasha Levin commit c44a04722f7dae9a45a101b8e100931ab59f1c18 Author: Ralf Lici Date: Fri Aug 28 16:50:27 2026 +0200 ovpn: replace bind when clearing stale local source [ Upstream commit 7d8104988f423572df1f3347ce578037b1043f34 ] The UDP output fallback clears bind->local in place when the remembered source address is no longer usable. The bind is RCU-published and read locklessly by concurrent TX, so an IPv6 reader can observe a torn address. Retry the route lookup with source address autoselection without modifying the bind. After a successful lookup, revalidate the bind and route key under peer->lock, reset the dst cache, and best-effort publish a replacement bind with a wildcard local address. Do not cache the resolved dst when clearing the local source. Replacing the source invalidates all per-CPU cache entries, while dst_cache_set_ip4 and dst_cache_set_ip6 update only the current CPU slot. The current packet can still use the resolved route; if bind allocation fails, a later cache miss retries the repair. Fixes: 08857b5ec5d9 ("ovpn: implement basic TX path (UDP)") Signed-off-by: Ralf Lici Signed-off-by: Antonio Quartulli Signed-off-by: Sasha Levin commit 21dbc9f4d0bdd3b22f91beabc7840064bf11594f Author: Ralf Lici Date: Fri Aug 28 16:50:26 2026 +0200 ovpn: replace bind when learning local endpoint [ Upstream commit aea934a221ec6a867221e5b765f65f1857befd53 ] struct ovpn_bind is published through peer->bind with RCU, but local endpoint learning updates bind->local in place under peer->lock. UDP TX reads the field without that lock. In particular, a concurrent IPv6 update can therefore result in a torn address read. Use ovpn_peer_reset_sockaddr to publish a replacement bind when learning a new local endpoint, just as a remote endpoint change does. Preserve the current remote address and reset the dst cache only after the new bind has been published successfully. Track remote endpoint changes separately so that float notification and transport-address rehashing remain limited to actual peer floats. Fixes: f0281c1d3732 ("ovpn: add support for updating local or remote UDP endpoint") Signed-off-by: Ralf Lici Signed-off-by: Antonio Quartulli Signed-off-by: Sasha Levin commit 324ebbde06cd535648bf86339d3224300b2177b9 Author: Ralf Lici Date: Fri Aug 28 16:50:25 2026 +0200 ovpn: validate peer state before caching UDP dst [ Upstream commit fa603710bdb9aea33c0d9cc2c05ed24d84f58753 ] UDP route lookup runs without peer->lock while the bind is protected by RCU. The route key is snapshotted separately. Either can change while the lookup is in progress. The TX path currently checks only the route key before publishing the looked-up dst. If the bind changes but the route key does not, a dst resolved from the old endpoint can be installed in the cache after the bind replacement. Compare both the bind pointer and the route key under peer->lock before updating the cache. The RCU read-side critical section keeps the old bind alive throughout the lookup, so pointer identity is sufficient to detect a replacement. Fixes: f0281c1d3732 ("ovpn: add support for updating local or remote UDP endpoint") Signed-off-by: Ralf Lici Signed-off-by: Antonio Quartulli Signed-off-by: Sasha Levin commit db97cfe9a5070b5cbe5520f9213cb0db446d9f03 Author: Ralf Lici Date: Fri Aug 28 16:50:24 2026 +0200 ovpn: track UDP socket route key for peer dst cache [ Upstream commit 7c66b7a4ae80a9309e6dc1d24b7b6b897e6348eb ] ovpn stores the route used to transmit UDP packets in a per-peer dst cache. A cached dst is only valid for the route lookup inputs used when it was resolved. Some of those inputs are mutable while userspace still owns the UDP socket. In particular, changes to the socket mark or UDP source port do not invalidate ovpn's peer dst cache, so ovpn can keep using a route selected with an old socket route key. Replace the cached mark with a route key containing the socket-owned lookup inputs currently used by ovpn, and reset the peer dst cache when the key changes. Before storing a newly looked-up dst, recheck the route key under the peer lock so a dst resolved for stale socket state is not published. Fixes: 08857b5ec5d9 ("ovpn: implement basic TX path (UDP)") Signed-off-by: Ralf Lici Signed-off-by: Antonio Quartulli Signed-off-by: Sasha Levin commit ba3877bdcc9aaa85cb6c03d628feb2fd7812f219 Author: Ralf Lici Date: Fri Aug 28 16:50:23 2026 +0200 ovpn: skip UDP source validation for unspecified addresses [ Upstream commit 77393b4d72dfeb764b2af2b848acc659f6fcfd0a ] ovpn validates the cached local UDP source address before reusing or refreshing a peer dst cache. This is only meaningful when a concrete source address is selected. For IPv6, calling ipv6_chk_addr with :: checks whether the unspecified address itself is configured on the host. A peer may legitimately have bind->local.ipv6 set to :: when no local endpoint was configured or after a stale learned address was cleared. In that case the source should be left unspecified and selected by ip6_dst_lookup_flow(). For IPv4, inet_confirm_addr(..., local = 0, ...) asks for local address autoselection rather than validating a chosen source. Skip the precheck there as well and let ip_route_output_flow select or reject the source. Only validate non-zero/non-any source addresses. Fixes: 08857b5ec5d9 ("ovpn: implement basic TX path (UDP)") Signed-off-by: Ralf Lici Signed-off-by: Antonio Quartulli Signed-off-by: Sasha Levin commit 8c1d7994a18ec6a5ccdb970c0316e6cfc4e4f360 Author: Ralf Lici Date: Fri Aug 28 16:50:22 2026 +0200 ovpn: preserve IPv6 scope id for netlink peer endpoints [ Upstream commit 7a6d08ee0f0e30023d18779bb314db8fd9a3b6d4 ] ovpn accepts OVPN_A_PEER_REMOTE_IPV6_SCOPE_ID and reports bind->remote.in6.sin6_scope_id in peer dumps, but the netlink endpoint parser never copied the attribute into the sockaddr_in6 used to create or update the peer bind. As a result, an IPv6 link-local remote endpoint configured through netlink loses its interface scope, unlike on the peer float path where ipv6_iface_scope_id populates the field. The UDPv6 output path then builds a flow with flowi6_oif set to zero and route lookup can fail or select the wrong interface. Copy the scope id when parsing non-v4-mapped IPv6 remote endpoints. The existing precheck already rejects the scope-id attribute for IPv4 and v4-mapped IPv6 remotes. Fixes: 1d36a36f6d53 ("ovpn: implement peer add/get/dump/delete via netlink") Signed-off-by: Ralf Lici Signed-off-by: Antonio Quartulli Signed-off-by: Sasha Levin commit ebffd70e4877bb7554974cf975a27872f2bb225f Author: Li Youhong Date: Fri Sep 4 09:49:58 2026 +0800 drm/bridge: samsung-dsim: fix TE GPIO lifetime for host attach [ Upstream commit ada667890773e033d2f40dc94176e3beb930b516 ] When the Exynos DSI driver was generalized into samsung-dsim, the TE GPIO acquisition was switched from gpiod_get_optional() to devm_gpiod_get_optional() while keeping the matching gpiod_put() calls. That combination is wrong for a managed descriptor. However, dropping the puts and keeping the managed get is also wrong: samsung_dsim_register_te_irq() runs from the DSI host attach callback, and host detach/reattach can happen without destroying the device that owns the managed action. A second attach would then request the GPIO again without having released it. Switch back to a non-managed gpiod_get_optional() and keep the explicit gpiod_put() on the request_irq() error path and in samsung_dsim_unregister_te_irq(). Fixes: e7447128ca4a ("drm: bridge: Generalize Exynos-DSI driver into a Samsung DSIM bridge") Suggested-by: Luca Ceresoli Signed-off-by: Li Youhong Reviewed-by: Luca Ceresoli Tested-by: Luca Ceresoli Link: https://patch.msgid.link/20260904014958.1572918-1-dayou5941@163.com [Luca: remove unnecessary comment] Signed-off-by: Luca Ceresoli Signed-off-by: Sasha Levin commit d6b59a9fece4d339010acda2f3e21a2ff9701e0e Author: Benjamin Leggett Date: Fri Aug 14 18:20:11 2026 -0400 drm/virtio: sync shmem backing on guest-bound transfers [ Upstream commit 598c1c3e895590f845e04455d5580ea28ffde666 ] virtio_gpu_cmd_transfer_to_host_{2d,3d}() sync the shmem backing for the device before the transfer, but nothing syncs for the CPU when a transfer runs the other way. That breaks two ways. Where the DMA layer bounces, the device writes into the bounce buffer while the guest keeps reading the original pages. Where DMA is not coherent, the device writes memory while the CPU keeps stale cache lines, because nothing reaches arch_sync_dma_for_cpu(). Either way DRM_IOCTL_VIRTGPU_TRANSFER_FROM_HOST hands back stale data. Sashiko originally found this in https://lore.kernel.org/dri-devel/20260806231002.27B4D1F000E9@smtp.kernel.org but the suggestion there to fix this with dma_sync_sgtable_for_cpu() isn't a sufficient fix, for two reasons. - The transfer is asynchronous. virtio_gpu_cmd_transfer_from_host_3d() only queues the command, so a sync there would run before the device had written anything. It belongs on completion, and ahead of any fence signalling. A waiter woken by the fence would otherwise race the sync and read the backing pages regardless. It needs its own pass over the reclaim list rather than a step inside the existing one, because virtio_gpu_fence_event_process() also signals every earlier fence in the same context, so any entry in that loop may signal an earlier entry's fence. - The transfer is also partial, carrying an offset, a level and a box. Where the mapping bounces, a sync for the CPU copies the whole mapping back, so unless the mapping is primed first the regions the device did not write come back holding whatever the bounce buffer contained, discarding data the guest owned. So the fix: Prime the mapping before queueing, tag the vbuffer, and sync for the CPU on completion before the fence is signalled. A second transfer must not snapshot the mapping while an earlier one is still in flight, or the snapshot would predate whatever the CPU wrote once the earlier fence signalled and the later sync would discard it. To mitigate this, wait for outstanding fences under the reservation before priming. Neither sync copies anything unless the mapping genuinely bounces: swiotlb_sync_single_for_cpu() and its Xen counterpart look the address up in the bounce pool and return early when it is absent. On a platform with non-coherent DMA they still perform the necessary cache maintenance. The range cannot be narrowed to the box, since for a non-blob resource virtio_gpu_transfer_from_host_ioctl() rejects a caller-supplied stride and layer_stride, leaving the layout to the host and the guest with no way to work out which bytes the device writes. A host3d guest blob does carry both, so its extent could be bounded, but the sync is left whole there too rather than special-cased: priming makes the untouched regions round-trip unchanged either way. Behaviour changes worth noting: - TRANSFER_FROM_HOST can now block, where before it returned as soon as the command was queued. Repeated readbacks of one resource serialise, and a readback can wait behind an earlier queued command that touched it, since virtio_gpu_array_add_fence() tags uploads, execbufs and plane flushes alike with DMA_RESV_USAGE_WRITE. -ERESTARTSYS was already possible here via dma_resv_lock_interruptible(). - A CPU write racing an in-flight transfer to the same resource is now lost, where before it survived and the transfer was lost instead. Priming captures the pages as of queueing, so a write landing before completion is overwritten by the sync. - TRANSFER_TO_HOST can also block now, but only while a guest-bound transfer on the same resource is outstanding, which happens only for callers that issue both without waiting. - Where a batch of completions contains a guest-bound transfer, the sync pass delays fence signalling for the whole batch. Only bounced pages are copied and the swiotlb pool bounds it. A batch with no such transfer is unaffected. Tested under QEMU on x86 with swiotlb=force and virtio-vga-gl iommu_platform=on, which forces both preconditions required to hit the original bug. Fixes: a3b815f09bb8 ("drm/virtio: add iommu support.") Reported-by: Sashiko AI review Closes: https://lore.kernel.org/dri-devel/20260806231002.27B4D1F000E9@smtp.kernel.org/ Signed-off-by: Benjamin Leggett Signed-off-by: Dmitry Osipenko Link: https://patch.msgid.link/20260814-virtgpu-from-host-sync-v4-1-64dd736b1779@edera.io Signed-off-by: Sasha Levin commit e7df00bb036e037aa8d5a4ea42a6ae7ca46969f4 Author: Dmitry Osipenko Date: Fri Sep 11 17:42:03 2026 +0300 Revert "drm/virtio: Allow importing prime buffers when 3D is enabled" [ Upstream commit 1e3b08de63274d0b009e99ef51cd6a9c0c6bf08c ] Guest userspace may import udmabuf to vrend. Vrend doesn't support guest blobs, and thus, further 3d operations with the imported blob are failing. Typical scenario of the problem shown with mouse cursor RGBA image imported into virtio-gpu, which previously was rejected by virtio-gpu driver. Revert enabling guest blobs importing into vrend to fix the regression. Link: https://gitlab.freedesktop.org/virgl/virglrenderer/-/work_items/674 Fixes: df4dc947c46b ("drm/virtio: Allow importing prime buffers when 3D is enabled") Signed-off-by: Dmitry Osipenko Reviewed-by: Val Packett Link: https://patch.msgid.link/20260911144204.2089401-1-dmitry.osipenko@collabora.com Signed-off-by: Sasha Levin commit 5c2cb3dd1447752b73d4f44d6b7434bc5259bfec Author: Junrui Luo Date: Tue Sep 15 15:39:13 2026 +0800 drm/virtio: release the GEM object on virtio_gpu_vram_create() errors [ Upstream commit 036d28db1818af2f9d80db771f5405da84d7732d ] virtio_gpu_vram_create() frees the object with a bare kfree(vram) on both error paths after drm_gem_private_object_init() has run, and on the second one after drm_gem_create_mmap_offset() has linked obj->vma_node into the device's VMA offset manager. The freed object stays in that interval tree, so a later lookup or insertion walks freed memory, and the dma_resv and gpuva lock are never destroyed. Call drm_gem_object_release() before kfree() on both paths. Fixes: 16845c5d5409 ("drm/virtio: implement blob resources: implement vram object") Assisted-by: Claude:claude-opus-5 Signed-off-by: Junrui Luo Signed-off-by: Dmitry Osipenko Link: https://patch.msgid.link/20260915-fixes-v2-4-a0d799e4db66@outlook.com Signed-off-by: Sasha Levin commit 81c6769e069d703154087377d1749af2f1b74e49 Author: Junrui Luo Date: Tue Sep 15 15:39:12 2026 +0800 drm/virtio: fix object leaks in virtio_gpu_resource_create_blob_ioctl() [ Upstream commit 24b6d5c7641412c9ebef0d4c8b888d49a0e6b880 ] virtio_gpu_resource_create_blob_ioctl() calls drm_gem_object_release() on both the virtio_gpu_resource_assign_uuid() and drm_gem_handle_create() error paths instead of dropping the reference it owns, so obj->funcs->free() never runs and the virtio_gpu_object, the resource id and the host-side resource are leaked. Use drm_gem_object_put() instead. Fixes: 897b4d1acaf5 ("drm/virtio: implement blob resources: resource create blob ioctl") Assisted-by: Claude:claude-opus-5 Signed-off-by: Junrui Luo Signed-off-by: Dmitry Osipenko Link: https://patch.msgid.link/20260915-fixes-v2-3-a0d799e4db66@outlook.com Signed-off-by: Sasha Levin commit b516446a25bf365cfffe3d7884ea6637f7f70862 Author: Junrui Luo Date: Tue Sep 15 15:39:11 2026 +0800 drm/virtio: fix object leak in virtio_gpu_resource_create_ioctl() [ Upstream commit 477bc3068fc3777b9d8ffd79e265b0dfdf2d3a6b ] virtio_gpu_resource_create_ioctl() calls drm_gem_object_release() on the drm_gem_handle_create() error path instead of dropping the reference it owns, so obj->funcs->free() never runs and the virtio_gpu_object, its pages and sg table, the resource id and the host-side resource are leaked. Use drm_gem_object_put() instead. Fixes: 62fb7a5e1096 ("virtio-gpu: add 3d/virgl support") Assisted-by: Claude:claude-opus-5 Signed-off-by: Junrui Luo Signed-off-by: Dmitry Osipenko Link: https://patch.msgid.link/20260915-fixes-v2-2-a0d799e4db66@outlook.com Signed-off-by: Sasha Levin commit a93dce15dfef200c346b050a2acbed516cc62706 Author: Junrui Luo Date: Tue Sep 15 15:39:10 2026 +0800 drm/virtio: fix object leak when drm_gem_handle_create() fails [ Upstream commit 36570ef2244cc4d7563b1f0157bc0f032498638c ] virtio_gpu_gem_create() owns the reference taken by virtio_gpu_object_create(). On the drm_gem_handle_create() error path it calls drm_gem_object_release() instead of dropping that reference. drm_gem_object_release() is the inverse of drm_gem_object_init() and does not touch the reference count or call obj->funcs->free(), so it is only correct as the last step of a destructor, as in virtio_gpu_cleanup_object(). Using it here leaves the bo at refcount 1 with no remaining reference, so virtio_gpu_free_object() never runs and the shmem pages, sg table and virtio_gpu_object are leaked. Since virtio_gpu_object_create() has already set bo->created, VIRTIO_GPU_CMD_RESOURCE_UNREF is not queued either, leaking the host-side resource and the resource id. drm_gem_handle_create_tail() drops the handle reference on all of its internal error paths, so the caller only has to drop its own. Use drm_gem_object_put(), matching the success path below. Fixes: dc5698e80cf7 ("Add virtio gpu driver.") Reported-by: Yuhao Jiang Assisted-by: Claude:claude-opus-5 Signed-off-by: Junrui Luo Signed-off-by: Dmitry Osipenko Link: https://patch.msgid.link/20260915-fixes-v2-1-a0d799e4db66@outlook.com Signed-off-by: Sasha Levin commit 0fa25ba50a382d79c0cf060214c2095756836f5f Author: Jamal Hadi Salim Date: Thu Sep 17 06:57:32 2026 -0400 net/sched: sch_hfsc: bound the classify inner-filter walk with a drift budget [ Upstream commit 8a60ade2277e1f0e0d0578d565354e52292fa46d ] hfsc_classify() applies the "filter may only point downwards" level check only when the filter result carries no bound class. A filter created with a flowid gets res.class set once at bind time, so the check never runs for it during classification. hfsc_adjust_levels() can later raise a class's level without revalidating existing bindings, leaving two binds that were each legal at bind time pointing at each other; the classify walk then bounces between two interior classes forever with the qdisc lock held and BH disabled — a soft lockup from a single packet. The stuck walk trips the watchdog: watchdog: BUG: soft lockup - CPU#3 stuck for 13s! [ping:444] RIP: 0010:u32_classify+0x542/0x17f0 ... tcf_classify+0x66/0xa0 hfsc_enqueue+0x166/0xdf0 Bound the traversal with a budget of non-descending hops, the only way a configured walk can move without descending the class tree once levels drift after bind time. The budget is cumulative over the whole walk and is deliberately not reset on a descending hop: a chain that alternates a descent with a lateral hop would return the budget every lap and never trip. Descending hops never decrement it, so legitimately deep trees are unaffected and a terminating lateral chain still classifies normally. Drop the packet with a rate-limited warning once the budget is exhausted, mirroring the merged HTB fix. This is a follow-up to commit 729c4896ab82 ("net/sched: sch_htb: limit htb_classify inner-class filter hops"), which bounded the same classify loop on the HTB side but left the HFSC walk unbounded. Conditions to recreate the bug: - CONFIG_NET_SCHED, CONFIG_NET_SCH_HFSC, CONFIG_NET_CLS_U32, CONFIG_LOCKUP_DETECTOR. - Build a cycle with two legal-at-bind-time flowid binds and a level drift: class X 1:1 (child of root) with leaf child 1:10; class Y 1:2 (sibling of X) with children 1:20 and 1:200; root u32 filter flowid 1:1; filter on X flowid 1:2 (legal when Y is a leaf); after Y's level rises to 2, filter on Y flowid 1:1 (legal then). Send one packet (ping on the device). Unfixed kernel: classify spins with the qdisc lock held; with softlockup_panic=1 it panics. - Reachable from unprivileged user via unshare -Urn (CAP_NET_ADMIN). Fixes: a2f79227138c ("net_sched: sch_hfsc: fix classification loops") Reported-by: Sashiko (gemini + nipa) Closes: https://lore.kernel.org/netdev/QDISC-CTUU.v2.20260913192614@mojatatu.com/ Link: https://sashiko.dev/#/patchset/QDISC-CTUU.v2.20260913192614@mojatatu.com Link: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/QDISC-CTUU.v2.20260913192614%40mojatatu.com Reviewed-by: Victor Nogueira Tested-by: hybris Signed-off-by: Jamal Hadi Salim Link: https://patch.msgid.link/QDISC-CTUU.v3.20260916184908@mojatatu.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit cd12d3e08620257fbcaa262785d7f4b9473afc7a Author: Yuqi Xu Date: Sat Sep 19 16:45:02 2026 +0800 bpf: Check params size before reading reserved fields [ Upstream commit a11212910cf09b2fe8db9afa41ef60c4f81879c5 ] bpf_crypto_ctx_create() is a kfunc whose second argument is declared with the __sz annotation, so the verifier only guarantees that params__sz bytes of params are valid. The function nevertheless reads params->reserved[0] and params->reserved[1] (offsets 14 and 15) before comparing params__sz against the size of struct bpf_crypto_params, so a BPF program can pass a shorter buffer and have the kernel read past the region that was validated for it. Move the size check in front of the reserved field reads. Fixes: 3e1c6f35409f ("bpf: make common crypto API for TC/XDP programs") Reported-by: Vega Signed-off-by: Yuqi Xu Signed-off-by: Alexei Starovoitov Reviewed-by: Ren Wei Link: https://patch.msgid.link/4f3ab4b03e79017e215521743996555439bf0bb3.1789802413.git.xuyuqiabc@gmail.com Signed-off-by: Sasha Levin commit dd9c6f8f05170223edec2dac476606cd40d9acb7 Author: Aamir Ahmed Date: Tue Sep 15 00:06:58 2026 +0100 net: usb: catc: bound the RX packet length in catc_rx_done() [ Upstream commit 9d565b6b72fe3f41fd43636e143072848105189f ] catc_rx_done() walks a multi-packet URB, reading a two-byte length from each packet header. Its bound, pkt_len > urb->actual_length, ignores the header offset and compares against the whole transfer rather than the bytes left from pkt_start, so a crafted packet header makes skb_copy_to_linear_data() read past the buffer. A length below ETH_HLEN is also accepted, including zero, and eth_type_trans() then reads a MAC header from the uninitialised tailroom of a shorter skb. The is_f5u011 branch takes its length straight from the transfer, so a zero-length URB reaches the same path. Track the bytes remaining from the current packet, and reject a header that does not fit, a length past what is left, and a length below an Ethernet header. A transfer shorter than an Ethernet header, including a zero-length one, previously became a runt skb passed to netif_rx() and counted as received; it is now counted in rx_length_errors and ends the walk. Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2") Signed-off-by: Aamir Ahmed Reviewed-by: Simon Horman Link: https://patch.msgid.link/AS8P251MB00015FD7716F38C345619B56C8BB2@AS8P251MB0001.EURP251.PROD.OUTLOOK.COM Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 9c11e543fec32e92b3aa3f693339ca3ee79ea9dd Author: Kumar Kartikeya Dwivedi Date: Mon Sep 14 15:24:42 2026 +0200 bpf: Bound ownership depth through local kptrs and graph roots [ Upstream commit bfc888f04588f591851e95c974954cfca58e6c19 ] Program-allocated objects can own other local objects through referenced kptrs. bpf_obj_free_fields() follows those pointers through __bpf_obj_drop_impl() synchronously, before the object storage is freed through RCU. A self-referential local kptr type therefore permits arbitrarily deep object chains, and dropping the head can exhaust the kernel stack. Long acyclic type chains have the same problem. btf_check_and_fixup_fields() still assumes referenced kptrs only point to kernel types and checks ownership through list and rbtree roots only. Its existing rule is sufficient for graph-only cycles: the target of each graph edge must contain a node, so every type in a cycle has both a root and a node. The rule rejects such a type owning another root, breaking every cycle. It also limits graph-only chains to three types, or two if the first type contains a node, and conservatively rejects longer acyclic chains. The missing local-kptr edges, rather than a missed graph-only cycle, are the bug introduced by support for bpf_kptr_xchg() into local kptrs. Replace that restriction with one bounded ownership walk covering graph roots and local referenced kptrs. Run it after all BTF records have been fixed up, reject cycles and paths deeper than eight record-bearing types, and cache each type's suffix depth while checking it against the remaining budget. This also permits the longer acyclic graph-only layouts rejected by the old rule; update their existing BTF tests accordingly. Keep the bound independent of MAX_CALL_FRAMES because recursive destruction can run below a BPF call chain. A plain local pointee without special-field metadata adds only a final non-recursing drop. Non-owning kptrs and kernel-BTF kptrs do not recurse through local records and remain outside the walk. Include local percpu-kptr edges too, although allocation of percpu objects with special fields is currently forbidden, so that relaxing that restriction cannot bypass the ownership bound. btf_check_and_fixup_fields() continues to initialize graph_root.value_rec, including for separately allocated map records. The ownership relationships belong to immutable program BTF and only need validation at BTF load time. Fixes: b0966c724584 ("bpf: Support bpf_kptr_xchg into local kptr") Reported-by: Nicholas Carlini Suggested-by: Nicholas Carlini Signed-off-by: Kumar Kartikeya Dwivedi Signed-off-by: Alexei Starovoitov Link: https://patch.msgid.link/20260914132444.2564218-2-memxor@gmail.com Signed-off-by: Sasha Levin commit c9338057df1025f578fe7d0ff08bada87b8b838d Author: Yiqi Sun Date: Tue Sep 15 17:50:17 2026 +0800 sctp: avoid livelock while updating retransmit path [ Upstream commit d2c31b837406395e576afeb25958c98e9938f3f6 ] sctp_assoc_update_retran_path() can loop forever when every remaining transport, including retran_path, is SCTP_UNCONFIRMED: the state check runs before the wraparound test, so the loop cannot observe that it has completed a full pass. Fix this by considering a transport only when it is not UNCONFIRMED, then checking whether the walk has returned to retran_path. This makes the full-pass termination independent of the transport state while preserving the existing fallback selection semantics. Also restore the NULL guard around the retran_path assignment. In the all-UNCONFIRMED case there is no eligible replacement transport, and installing NULL would leave later retransmit-path users and the debug print with a NULL path. Fixes: 4c47af4d5eb2 ("net: sctp: rework multihoming retransmission path selection to rfc4960") Signed-off-by: Yiqi Sun Acked-by: Xin Long Link: https://patch.msgid.link/20260915095017.942213-1-sunyiqixm@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit c7a254c16c6fb295dfa255a4b00674535f70b93e Author: Alexander Duyck Date: Mon Sep 14 14:10:33 2026 -0700 eth: fbnic: Handle FW mailbox completions flagged with an error [ Upstream commit 1b97a269a5bdde20d4e69511f27649c9cb82b7c7 ] The firmware can complete a mailbox descriptor while also setting FW_ERR to indicate it could not process the request, for example on a mailbox DMA error. The completion carries no valid data. The driver did not check FW_ERR. On the Rx mailbox it would sync and parse the stale page as a normal message, and on the Tx mailbox it silently freed the request. If the initial capabilities exchange in fbnic_mbx_poll_tx_ready() hit FW_ERR -- on the Tx request or on the Rx response descriptor -- no response was parsed and the poll spun until it timed out even though the ring was healthy. Check FW_ERR on both mailboxes. Count it per-mailbox in fbnic_fw_mbx.resp_error, which is also shown in debugfs, warn (rate limited, since the bit is firmware controlled), and drop the Rx page instead of parsing it. In fbnic_mbx_poll_tx_ready() re-issue the capabilities request when either the Tx or the Rx resp_error counter advances, so a FW_ERR on the request or on its response triggers a retry rather than a timeout. A valid capabilities response is honored before the retry check, so a response parsed in the same poll as an unrelated FW_ERR is not discarded. The counters are mailbox-wide rather than keyed to the capabilities request; that is sufficient here because the exchange runs during bring-up before any other mailbox traffic, and any spurious retry is bounded by the existing 10s timeout. Fixes: da3cde08209e ("eth: fbnic: Add FW communication mechanism") Signed-off-by: Alexander Duyck Reviewed-by: Simon Horman Link: https://patch.msgid.link/178942023343.7700.9423398932961964439.stgit@ahduyck-xeon-server.home.arpa Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit a85a019b3011a1573edc5131829f13d18c7f799c Author: Mike Marciniszyn (Meta) Date: Tue Jan 27 15:06:43 2026 -0500 eth fbnic: Add debugfs hooks for firmware mailbox [ Upstream commit ac7b803c2dd57938cbf1e9a3ada23471925a5d46 ] This patch adds reporting the Rx and Tx information interfacing with the firmware. The result of reading fbnic/fw_mbx is: Rx Rdy: 1 Head: 11 Tail: 10 Idx Len E Addr F H Raw ---------------------------------- 00 4096 0 000101fea000 0 1 1000000101fea001 01 4096 0 000101feb000 0 1 1000000101feb001 . . . 15 4096 0 000101fe9000 0 1 1000000101fe9001 Tx Rdy: 1 Head: 4 Tail: 4 Idx Len E Addr F H Raw ---------------------------------- 00 0004 1 00010321b000 1 1 000440010321b003 01 0004 1 00010228d000 1 1 000440010228d003 . . . 15 0004 1 00010321b000 1 1 000440010321b003 Signed-off-by: Mike Marciniszyn (Meta) Link: https://patch.msgid.link/20260127200644.11640-2-mike.marciniszyn@gmail.com Signed-off-by: Jakub Kicinski Stable-dep-of: 1b97a269a5bd ("eth: fbnic: Handle FW mailbox completions flagged with an error") Signed-off-by: Sasha Levin commit 6a112552d30557b406dd118a79c750fab45041a6 Author: Mohsin Bashir Date: Wed Jan 14 16:33:51 2026 -0800 eth: fbnic: Reuse RX mailbox pages [ Upstream commit 301ae0d5391a19d9897c342c874240b36f7bc85f ] Currently, the RX mailbox frees and reallocates a page for each received message. Since FW Rx messages are processed synchronously, and nothing hold these pages (unlike skbs which we hand over to the stack), reuse the pages and put them back on the Rx ring. Now that we ensure the ring is always fully populated we don't have to worry about filling it up after partial population during init, either. Update fbnic_mbx_process_rx_msgs() to recycle pages after message processing. Signed-off-by: Mohsin Bashir Link: https://patch.msgid.link/20260115003353.4150771-4-mohsin.bashr@gmail.com Signed-off-by: Jakub Kicinski Stable-dep-of: 1b97a269a5bd ("eth: fbnic: Handle FW mailbox completions flagged with an error") Signed-off-by: Sasha Levin commit 8a052cd4b2caee1f90cf0e3874377066939c90e7 Author: Alexander Duyck Date: Mon Sep 14 14:10:25 2026 -0700 eth: fbnic: Set AW_FLUSH_MODE alongside AW_FLUSH when flushing the mailbox [ Upstream commit 8947f13e436a4ff5eed9f8f019b2865a07af4bb2 ] When tearing down the FW mailbox Rx ring, fbnic_mbx_reset_desc_ring() writes AW_CFG with FLUSH set and everything else, BME included, cleared. Clearing BME halts the device's writes to the host but leaves the staged requests parked in the PUL write pipeline rather than draining them, so on the write path FLUSH alone never terminates the outstanding requests and the flush the firmware waits on never completes. Add the FLUSH_MODE definition and set both bits so the staged writes drain out of the pipeline on their own. BME stays cleared, so nothing lands on the host; it is restored later in fbnic_mbx_init_desc_ring() when the ring is rebuilt, once the outstanding writes are gone. The read path is unaffected. AR_CFG has no equivalent mode bit and AR_FLUSH terminates the outstanding reads by itself, so it is left as is. Both writes remain plain stores rather than read-modify-writes. That is deliberate: the matching write in fbnic_mbx_init_desc_ring() restores BME and the TLP attributes, and clears both flush bits as a side effect. Fixes: 3b12f00ddd08 ("fbnic: Gate AXI read/write enabling on FW mailbox") Signed-off-by: Alexander Duyck Reviewed-by: Simon Horman Link: https://patch.msgid.link/178942022583.7700.11050671998277309744.stgit@ahduyck-xeon-server.home.arpa Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit c5b5aee41b9f5ed0743344be1c7b544555fa3cb9 Author: Alexander Duyck Date: Mon Sep 14 14:10:18 2026 -0700 eth: fbnic: reset num_napi when the napi vectors are freed [ Upstream commit 4bcc4a92c603fe7f062cea22e20da2e0ad6b12c3 ] fbn->num_napi is the count of live napi vectors, each of which owns an IRQ. The PM path had freed them without clearing the count. fbnic_pm_suspend() tears the datapath down via ndo_stop() and frees the IRQs, but leaves netif_running() true so resume knows to re-open. Resume rebuilds the datapath in __fbnic_pm_resume() and fbnic_reset_queues() sets num_napi and __fbnic_open() re-allocates the vectors. When the datapath is torn down but never rebuilt, num_napi is left pointing at freed vectors under 2 different scenarios: - a PCIe error recovery that fails (fbnic_err_slot_reset() -> __fbnic_pm_resume() returns an error -> PCI_ERS_RESULT_DISCONNECT), so .resume never runs; or - an __fbnic_open() that fails partway on resume and unwinds, freeing the vectors after fbnic_reset_queues() has already set num_napi. The netdev is then running with num_napi > 0 but napi[] freed, and the eventual remove/unbind close re-enters fbnic_down() -> fbnic_dbg_down() and dereferences the freed vectors: BUG: kernel NULL pointer dereference, address: 0000000000000210 RIP: fbnic_dbg_down+0x28 Clear num_napi when the vectors are freed: in the suspend teardown (a good resume re-establishes it before __fbnic_open()) and on the resume open failure. A redundant ndo_stop() then walks an empty napi[]. The normal ndo_stop() down/up cycle is untouched and keeps num_napi for the next ndo_open(). Fixes: bc6107771bb4 ("eth: fbnic: Allocate a netdevice and napi vectors with queues") Signed-off-by: Alexander Duyck Reviewed-by: Simon Horman Link: https://patch.msgid.link/178942021809.7700.10804028989308077839.stgit@ahduyck-xeon-server.home.arpa Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 6c903188b67617070017305a325fa326f270a43d Author: Björn Töpel Date: Mon Sep 14 14:10:04 2026 -0700 eth: fbnic: Handle maximum standalone channels [ Upstream commit 1f4c73064a50f53d596c6f1d06d2d700f43c4b32 ] Standalone channels use one NAPI vector for each Tx and Rx queue. fbnic's allocation path excludes FBNIC_MAX_TXQS from that layout. A 64-Tx/64-Rx configuration therefore records 128 vectors but allocates only 64, leaving NULL entries that resource setup dereferences. Include the maximum vector count in standalone allocation. Fixes: bc6107771bb4 ("eth: fbnic: Allocate a netdevice and napi vectors with queues") Signed-off-by: Björn Töpel Reviewed-by: Simon Horman Link: https://patch.msgid.link/178942020457.7700.13129750616387075931.stgit@ahduyck-xeon-server.home.arpa Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 8dc5667770912be6909002b94a834a97bb9f68dd Author: Kuniyuki Iwashima Date: Wed Sep 16 23:09:24 2026 +0000 ip6_gre: Call ip6erspan_tunnel_unlink_md() in ip6erspan_changelink(). [ Upstream commit dd47bcf279f1083f09bf5266890b26263361022b ] The cited commit accidentally added ip6gre_tunnel_unlink_md() in ip6erspan_changelink(). Let's correct it to ip6erspan_tunnel_unlink_md(). Fixes: b80d0b93b991 ("net: ip6_gre: fix tunnel metadata device sharing.") Signed-off-by: Kuniyuki Iwashima Reviewed-by: Xuanqiang Luo Reviewed-by: Ido Schimmel Link: https://patch.msgid.link/20260916230927.378957-1-kuniyu@google.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 7254b702e556158e34b4d7c168b5d8e7da5978af Author: Lorenzo Bianconi Date: Wed Sep 16 15:30:13 2026 +0200 net: ethernet: mtk_eth_soc: unregister net_devices in case of probe failure [ Upstream commit 310d1ac61a4d5a2ca8356a3a48d263acf54503ce ] If register_netdev() fails for one of the MTK_MAX_DEVS devices in mtk_probe(), the error path jumps to err_deinit_ppe, skipping mtk_unreg_dev(). The previously registered net_devices are then freed by mtk_free_dev() while still in NETREG_REGISTERED state, hitting the BUG_ON(dev->reg_state != NETREG_UNREGISTERED). Route the register_netdev() failure to err_unreg_netdev so the net_devices registered so far are properly unregistered before being freed. Fixes: 8a8a9e89f801 ("net: ethernet: mediatek: cleanup error path inside mtk_hw_init") Signed-off-by: Lorenzo Bianconi Link: https://patch.msgid.link/20260916-mtk_eth_soc-netdev-fix-v1-1-5dac50eb65b1@oss.qualcomm.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 40aeafbcbc7bc4018e8f74bbbb4d29ed8846f57f Author: Jérémy Jean Date: Tue Sep 15 12:48:07 2026 +0000 net: gue: reject invalid REMCSUM offsets [ Upstream commit 2566866fc30965d915d0b52b5c3323b362619f0e ] The REMCSUM option carries an absolute checksum start and checksum field offset. gue_remcsum() passes them to skb_remcsum_process(), whose partial path stores offset - start in the u16 skb->csum_offset variable. If offset is less than start, this underflows. A forwarded packet can retain CHECKSUM_PARTIAL and reach a NETIF_F_HW_CSUM driver which trusts the metadata, leading skb_copy_and_csum_dev() to write two bytes about 64 KiB beyond the destination buffer. Reject reversed tuples in validate_gue_flags(), after the existing length validation, so all GUE parsers enforce the ordering in one place. Fixes: fe881ef11cf0 ("gue: Use checksum partial with remote checksum offload") Signed-off-by: Jérémy Jean Link: https://patch.msgid.link/20260915124806.2852293-2-Jeremy.Jean@oss.cyber.gouv.fr Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 00734f3684fcec42c319bedfe68a5d357c720b53 Author: Nguyen Ngoc Thang Date: Tue Sep 15 22:08:16 2026 +0700 net/sched: act_ct: don't WARN on benign flow_offload_alloc() failure [ Upstream commit 47abe7a5c4eb53269aca3506446f851572a059a3 ] flow_offload_alloc() returns NULL when the conntrack entry is dying (e.g. raced with a conntrack flush) or when the GFP_ATOMIC allocation fails; both are expected under load and neither is a kernel bug. This path runs from softirq on every committed packet, so with panic_on_warn=1 an unprivileged user can panic the box just by racing a conntrack flush against a `tc ... action ct commit` classifier. Reproduced with a custom repro under QEMU: a small, fixed set of UDP flows through `tc filter ... action ct commit` on lo, raced against threads flooding bare ctnetlink CT_DELETE (flush) requests. Hits WARNING: net/sched/act_ct.c:437 (tcf_ct_flow_table_add(), inlined into tcf_ct_act() in this build) within ~15s on the unpatched kernel; same setup is clean on the patched kernel. The fix itself is behavior-preserving: both branches already did `goto err_alloc` before and after, only the WARN is removed. Fixes: 64ff70b80fd4 ("net/sched: act_ct: Offload established connections to flow table") Reported-by: syzbot+6cc37aba98dac721c415@syzkaller.appspotmail.com Closes: https://syzkaller.appspot.com/bug?extid=6cc37aba98dac721c415 Signed-off-by: Nguyen Ngoc Thang Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260915150816.36487-1-ngocthang2710.1999@gmail.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit ca29834c8c146c944063b55728b9de8adf0cdcbb Author: Giuseppe Ranieri Date: Thu Sep 17 21:51:14 2026 +0000 drm/nouveau/disp: don't reject HDMI config on cards without SCDC [ Upstream commit 1717fcc5be575d4768279148ae9465a8b13d4339 ] nv50_hdmi_enable() passes the sink's SCDC capability from its EDID straight through to nvif_outp_hdmi(). On pre-Maxwell-2 cards there is no hdmi->scdc callback, so nvkm_uoutp_mthd_hdmi() rejects the whole configuration with -EINVAL, and nv50_hdmi_enable() returns before hdmi->ctrl() runs and before the AVI and VSI infoframes are sent. The result on such a card driving an SCDC-capable HDMI 2.0 sink is that HDMI audio silently stops working. Video is unaffected, and nothing is logged, which makes the failure hard to attribute. SCDC is optional, and the hdmi->scdc() call further down is already guarded against a missing callback. Requesting it on a card that cannot do it need not invalidate the rest of the HDMI configuration, so drop that term from the condition and let the existing guard skip SCDC alone. Fixes: 6c6abab20b99 ("drm/nouveau/disp: add output hdmi config method") Signed-off-by: Giuseppe Ranieri Co-Authored-By: Tano Dzhinski Signed-off-by: Tano Dzhinski Tested-by: Tano Dzhinski Reviewed-by: Lyude Paul Signed-off-by: Lyude Paul Link: https://patch.msgid.link/20260917215114.1136715-1-tano.dzhinski@gmail.com Signed-off-by: Sasha Levin commit 40afa1f7a5c771f85ede8fdc43c1fd6fac5a34a5 Author: Francesco Magazzu Date: Fri Sep 18 15:16:20 2026 +0200 drm/nouveau/clk: don't clobber reclock status when restoring volt/fan [ Upstream commit e5cccdafc855cd5f96f4b51d38114a0360b075d7 ] nvkm_cstate_prog() reuses 'ret' for the voltage and fan-speed restore calls it makes after reprogramming the clocks. Those calls almost always succeed, so the status of the reclock itself is overwritten and the function reports success even when clk->func->calc() or clk->func->prog() failed. The converse is also true: a successful reclock is reported as an error if the final restore call fails, even though that failure is only logged and otherwise ignored. The only consumer of the return value is the error message in nvkm_pstate_work(), so in practice a failing reclock is simply never reported. Nothing else changes, but a function that returns success on failure is a trap for the next caller. Keep the calc/prog status in 'ret' and use a separate local for the restore calls. Fixes: 3eca809b3c05 ("drm/nouveau/clk: cosmetic changes") Signed-off-by: Francesco Magazzu Reviewed-by: Lyude Paul Signed-off-by: Lyude Paul Link: https://patch.msgid.link/20260918131620.405133-5-postadelmaga@gmail.com Signed-off-by: Sasha Levin commit 0d096a7f7d79207490858dde731b16da5337e90c Author: Dan Carpenter Date: Fri Sep 18 15:16:17 2026 +0200 drm/nouveau/clk: fix list cursor use after loop in nvkm_clk_ustate_update [ Upstream commit aff09d9e37e02dc60bde79035ac15b136d602259 ] If list_for_each_entry() exits without hitting a break then "pstate" is not a valid pstate pointer. Introduce a "found" variable instead. The check is reachable from userspace: nvkm_clk_ustate_update() takes the pstate id straight from the 'pstate' debugfs file, so requesting an id that is not in clk->states - or any id at all when the perf tables are broken and the list is empty - makes the pstate->pstate != req test dereference the list head cast to a struct nvkm_pstate, which is an out-of-bounds read. Fixes: 7c8565220697 ("drm/nouveau/clk: implement power state and engine clock control in core") Signed-off-by: Dan Carpenter [Francesco: rebased on drm-misc-next, expanded the commit message] Signed-off-by: Francesco Magazzu Reviewed-by: Lyude Paul Signed-off-by: Lyude Paul Link: https://patch.msgid.link/20260918131620.405133-2-postadelmaga@gmail.com Signed-off-by: Sasha Levin commit f2e9005aadb6784a4c89c0a4213924aa8afce910 Author: Joseph Qi Date: Fri Sep 4 10:37:51 2026 +0800 ocfs2: make ocfs2_calc_xattr_init() return void [ Upstream commit 525c0edc032b3297d0c1056cf1fa20cf1f9e6184 ] ocfs2_calc_xattr_init() used to read the default ACL off the parent inode itself, so it could return an error from ocfs2_xattr_get_nolock(). Commit bd7c05fb4a47 ("ocfs2: fix circular locking dependency in ocfs2_init_acl()") moved that lookup before the transaction starts and deleted the error path, but left the now vestigial 'int ret = 0' declaration and both 'return ret' statements behind, along with an unreachable error branch in ocfs2_mknod(). Drop the leftover variable and convert the return type to void, so the callee states that it always succeeds and the caller no longer carries a check that can never trigger. No functional change. Link: https://lore.kernel.org/20260904023751.3703334-1-joseph.qi@linux.alibaba.com Fixes: bd7c05fb4a47 ("ocfs2: fix circular locking dependency in ocfs2_init_acl()") Signed-off-by: Joseph Qi Signed-off-by: Andrew Morton Reported-by: kernel test robot Closes: https://lore.kernel.org/oe-kbuild-all/202609040247.8B3lmoqX-lkp@intel.com/ Cc: Mark Fasheh Cc: Joel Becker Cc: Junxiao Bi Cc: Changwei Ge Cc: Jun Piao Cc: Heming Zhao Signed-off-by: Sasha Levin commit b17aa4955e3a694c663ebe99626e4c76ca4f7cfb Author: Zeng Heng Date: Fri Sep 11 09:58:59 2026 +0800 arm64: io: Reject non-user protection in ioremap_prot() [ Upstream commit bb756b11ad63832ebee58caf9e8f9381eaecff9f ] Mapping a stack-top page via /dev/mem with PROT_NONE and then reading that process's /proc//cmdline triggers a spurious WARN in ioremap_prot() through generic_access_phys(): WARNING: ./arch/arm64/include/asm/io.h:275 at generic_access_phys Call trace: generic_access_phys+0x1c8/0x228 (P) __access_remote_vm+0x2b4/0x398 access_remote_vm+0x14/0x30 get_mm_cmdline+0xf8/0x2a0 proc_pid_cmdline_read+0x68/0x120 generic_access_phys() passes the protection derived from the user PTE to ioremap_prot(). On arm64, a PROT_NONE mapping is represented by a present-invalid PTE, so pte_present() still returns true and the protection reaches ioremap_prot(). A PROT_NONE mapping does not have PTE_USER, causing the existing WARN_ON_ONCE() in ioremap_prot() to fire even though this is a valid user mapping. Execute-only mappings have the same issue and must not be readable through this path either. ioremap_prot() should therefore reject protection values without PTE_USER without warning. This makes the access fail cleanly for PROT_NONE and execute-only mappings while retaining the existing user-protection contract. Fixes: 8f098037139b ("arm64: io: Extract user memory type in ioremap_prot()") Signed-off-by: Zeng Heng Reviewed-by: Catalin Marinas Signed-off-by: Will Deacon Signed-off-by: Sasha Levin commit 0994c3e3b4e3ecfa824714145786498c9327d8f6 Author: Naman Gulati Date: Sat Sep 12 01:10:51 2026 +0000 netfilter: ctnetlink: fix suspicious RCU usage in expect_iter_name [ Upstream commit 207d591c353201f3bd3e0c89bb7d44a849c8fd59 ] expect_iter_name() is invoked by nf_ct_expect_iterate_net() under spin_lock_bh(&nf_conntrack_expect_lock). It does not hold rcu_read_lock(). When accessing exp->helper with rcu_dereference() in syzbot's report, lockdep warns: ============================= WARNING: suspicious RCU usage syzkaller #0 Not tainted ----------------------------- net/netfilter/nf_conntrack_netlink.c:3393 suspicious rcu_dereference_check() usage! locks held by syz-executor381/5628: 2, last CPU#1: #0: ffffffff9aee42a0 (nfnl_subsys_ctnetlink_exp){+.+.}-{4:4}, at: nfnetlink_rcv_msg+0xa69/0x12b0 #1: ffffffff8ea74d58 (nf_conntrack_expect_lock){+...}-{3:3}, at: nf_ct_expect_iterate_net+0x38/0x180 Call Trace: dump_stack_lvl+0xe8/0x150 lockdep_rcu_suspicious+0x140/0x1d0 expect_iter_name+0xfb/0x100 nf_ct_expect_iterate_net+0xf2/0x180 ctnetlink_del_expect+0x45d/0x640 nfnetlink_rcv_msg+0xcc2/0x12b0 netlink_rcv_skb+0x226/0x4a0 nfnetlink_rcv+0x2b9/0x28c0 netlink_unicast+0x7bd/0x940 netlink_sendmsg+0x813/0xb40 ____sys_sendmsg+0x54e/0x850 ___sys_sendmsg+0x2a5/0x360 __sys_sendmsg+0x2a5/0x360 do_syscall_64+0x166/0x520 entry_SYSCALL_64_after_hwframe+0x77/0x7f Use rcu_dereference_protected() with lockdep_is_held() on nf_conntrack_expect_lock instead, similar to expect_iter_me() in nf_conntrack_helper.c. Fixes: f01794106042 ("netfilter: nf_conntrack_expect: use expect->helper") Reported-by: syzbot+4bd730aede2791e40bdf@syzkaller.appspotmail.com Closes: https://lore.kernel.org/netdev/6aa4a377.f81106d8.2ab401.0024.GAE@google.com/T/#u Signed-off-by: Naman Gulati Signed-off-by: Pablo Neira Ayuso Signed-off-by: Sasha Levin commit f100578039b31d6b3308f86cbb4ecb8f7947f10e Author: Julian Anastasov Date: Fri Sep 11 14:43:15 2026 +0300 ipvs: revalidate ihl before icmp_send [ Upstream commit e290145564886d6a3038810c621f738c1fe9fa51 ] While the outer IP header is already pulled into the skb head, we must be careful and revalidate the embedded headers after reading them from the skb frags to prevent possible out-of-bounds access. One such place reported by Sashiko is ip_vs_in_icmp() where local process can change the ihl field and after pskb_may_pull() we can see larger value. Even if icmp_send() has checks to prevent out-of-bounds access, play safe and add check to drop the packet if the ihl field is changed. As the outer headers are pulled, make sure the transport header is updated too, it was used before commit 7fcc2fe39fed ("net: icmp: avoid invalid transport header access in icmp_send tracepoint") Fixes: f2edb9f7706d ("ipvs: implement passive PMTUD for IPIP packets") Link: https://sashiko.dev/#/patchset/20260806105211.34622-1-ja%40ssi.bg Signed-off-by: Julian Anastasov Signed-off-by: Pablo Neira Ayuso Signed-off-by: Sasha Levin commit 302477c9ffefa0b266111bca102d9208c0f85fb0 Author: Karl Mehltretter Date: Thu Sep 10 22:02:28 2026 +0200 netfilter: nft_synproxy: use the family-aware checksum helper [ Upstream commit a311a898172743558b82f6035ef2aa8c310a4223 ] nft_synproxy_do_eval() verifies the TCP checksum before it switches on skb->protocol. It uses nf_ip_checksum(), which constructs an IPv4 pseudo header and relies on the IPv4 header checksum when folding the whole skb. Neither operation is valid for an IPv6 packet. A correctly checksummed IPv6 segment can therefore fail verification when it reaches the hook as CHECKSUM_NONE or, at NF_INET_LOCAL_IN, CHECKSUM_COMPLETE. nft_synproxy_do_eval() returns NF_DROP before nft_synproxy_eval_v6() can send a SYN-ACK. nft_synproxy_validate() deliberately admits NFPROTO_IPV6 and NFPROTO_INET, and the xtables counterpart ip6t_SYNPROXY.c already calls nf_ip6_checksum(). Use nf_checksum() with nft_pf() so the checksum helper dispatches to the packet family's implementation. Fixes: ad49d86e07a4 ("netfilter: nf_tables: Add synproxy support") Assisted-by: LLM Signed-off-by: Karl Mehltretter Signed-off-by: Pablo Neira Ayuso Signed-off-by: Sasha Levin commit 07a3ec8fdc5a1267c00cd6fd1cf8bbb5cda913c1 Author: Florian Westphal Date: Thu Sep 3 02:41:46 2026 +0200 netfilter: nfnetlink_queue: hold nfnl mutex in event notifier [ Upstream commit 9461613afc59acef44a0071b0dd5075f6e993ffe ] We must serialize the release notifier and the config netlink function. A concurrent thread can issue close() which can call the release function while unrelated socket processes UNBIND request for same portid: Oops: general protection fault, [..] RIP: 0010:__instance_destroy+0x60/0x210 [nfnetlink_queue] Call Trace: nfqnl_recv_config+0x9b0/0xdc0 [nfnetlink_queue] nfnetlink_rcv_msg+0x7c2/0xeb0 ? __pfx_nfnetlink_rcv_msg+0x10/0x10 After this, parallel UNBIND and URELEASE events are impossible. This change isn't nice, but its the shortest fix given instances are not refcounted and the nfnetlink config callback drops the rcu read lock early due to need for sleeping allocations. Fixes: 7af4cc3fa158 ("[NETFILTER]: Add "nfnetlink_queue" netfilter queue handler over nfnetlink") Signed-off-by: Florian Westphal Signed-off-by: Pablo Neira Ayuso Signed-off-by: Sasha Levin commit b0ef9450e93f2adeb77593db14c2ac7c6c8f5060 Author: Jérémy Jean Date: Tue Aug 18 20:00:15 2026 +0000 netfilter: flowtable: publish HW_DEAD after worker is done [ Upstream commit d644b23afe1ef509c9961a6d84a093c2587edf02 ] flow_offload_work_del() sets NF_FLOW_HW_DEAD before the work handler clears NF_FLOW_HW_PENDING. Once a flow is both HW_DYING and HW_DEAD, a concurrent garbage collection pass can remove it and schedule it for RCU freeing. The offload worker holds neither an RCU read lock nor a reference to the flow. If it is preempted after publishing HW_DEAD, the RCU callback can free the flow before the worker resumes and clears HW_PENDING, resulting in a use-after-free. Move HW_DEAD publication to the common worker epilogue after the pending bit is cleared, making it the final flow access by destroy work. Order all preceding flow accesses before publishing the bit that allows garbage collection to free the object. Fixes: 2c8897953f3b ("netfilter: flowtable: Add pending bit for offload work") Assisted-by: Codex:gpt-5 Signed-off-by: Jérémy Jean Signed-off-by: Pablo Neira Ayuso Signed-off-by: Sasha Levin commit 1363c71be27a79b01beeadfb103a1b6fc37a217e Author: Linkui Xiao Date: Wed Sep 16 20:53:16 2026 +0800 ipv4: fib: fix data-race and stale genid check around nh->nh_saddr [ Upstream commit 46bc52d13594848023e681860df8700c8db14354 ] fib_select_multipath() compares nexthop_nh->nh_saddr against the flow source address with no lock held, while fib_info_update_nhc_saddr() stores a new value from another CPU as soon as the preferred source address of the egress device changes. Commit 195374d89368 ("ipv4: fib: annotate races around nh->nh_saddr_genid and nh->nh_saddr") added WRITE_ONCE() on the store side and READ_ONCE() in fib_result_prefsrc() after syzbot reported BUG: KCSAN: data-race in fib_select_path / fib_select_path but it only covered that reader. fib_select_multipath(), reached from fib_select_path(), is a second lockless reader of nh->nh_saddr and was left bare. Moreover, nh_saddr is only meaningful when nh_saddr_genid matches dev_addr_genid, as established by commit 436c3b66ec98 ("ipv4: Invalidate nexthop cache nh_saddr more correctly."). fib_select_multipath() skips that validation, so it can score a nexthop using a stale source address and skew the ECMP selection. Annotate both reads with READ_ONCE() and refresh the cached source address via fib_info_update_nhc_saddr() when the genid does not match, mirroring fib_result_prefsrc(). Fixes: 32607a332cfe ("ipv4: prefer multipath nexthop that matches source address") Signed-off-by: Linkui Xiao Reviewed-by: Ido Schimmel Reviewed-by: Eric Dumazet Link: https://patch.msgid.link/20260916125316.988044-1-xiaolinkui@126.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 5be5ccd3a556394887dbb90d1d4662f706f8dd26 Author: Kumar Kartikeya Dwivedi Date: Fri Sep 18 01:32:18 2026 +0200 libbpf: Reject truncated ldimm64 CO-RE relocations [ Upstream commit b4e875d397da451fb4e9c573ff4b86db53caba05 ] CO-RE relocation of an ldimm64 instruction operates on two instruction slots. A malformed BPF ELF can end a function after the first slot and attach a CO-RE relocation to it. libbpf allocates the instruction array according to the function symbol size, so the shared relocation code would then access beyond the allocation. Reject a terminal ldimm64 in libbpf's relocation loop, where the program length is available, before resolving or applying the relocation. Both resolved and unresolved relocations validate the absent second slot, and unresolved relocation poisoning would additionally write past the array. The in-kernel caller is protected by the verifier's early instruction-stream check before it applies CO-RE relocations. Fixes: eacaaed784e2 ("libbpf: Implement enum value-based CO-RE relocations") Reported-by: Sashiko Signed-off-by: Kumar Kartikeya Dwivedi Link: https://lore.kernel.org/20260914140852.03DA21F0089B@smtp.kernel.org Link: https://patch.msgid.link/20260917233222.2542500-11-memxor@gmail.com Signed-off-by: Eduard Zingerman Signed-off-by: Sasha Levin commit 2d7d04357c41aa49a5ac658fcc207439f2ad6456 Author: Kumar Kartikeya Dwivedi Date: Fri Sep 18 01:32:14 2026 +0200 bpf: Restrict CO-RE poisoning to relocatable instructions [ Upstream commit 394ae398337c5f87e567f6cd63b937fc2b2f6ddc ] CO-RE relocation records can name any instruction offset. When a relocation cannot be resolved, bpf_core_patch_insn() currently poisons its target before checking whether that instruction is a valid relocation target. Malformed metadata can therefore replace jumps, calls, exits, register-source arithmetic, or non-immediate loads instead of failing at the relocation step. Handle poisoning only after the instruction has passed the same class and operand-form checks used for a resolved relocation. Route invalid forms through the existing diagnostic and return a hard error. Keep poisoning supported instructions, including both halves of a plain ldimm64, so an unresolved relocation in dead code remains valid. Extend bpf_core_poison_insn() to poison both halves of ldimm64, and return its status directly from each validated instruction case. This avoids routing the success path through a common label and leaves the helper free to report errors. The shared relocation code applies this restriction to both libbpf and in-kernel CO-RE. Fixes: d7a252708dbc ("libbpf: Improve handling of failed CO-RE relocations") Reported-by: Nicholas Carlini Suggested-by: Nicholas Carlini Signed-off-by: Kumar Kartikeya Dwivedi Acked-by: Eduard Zingerman Link: https://patch.msgid.link/20260917233222.2542500-7-memxor@gmail.com Signed-off-by: Eduard Zingerman Signed-off-by: Sasha Levin commit 2fe1569e32ddff91ccbfe60515c3ef9d4eae5500 Author: Shay Drory Date: Tue Sep 15 14:34:57 2026 +0300 net/mlx5: devcom, Base component size on linked devices [ Upstream commit d09e8f64653c93da5793c16be19330968f2a32e6 ] mlx5_devcom_comp_get_size() returns the component's kref count. That kref is bumped in mlx5_devcom_register_component() under comp_list_lock, before the comp_dev is linked onto comp_dev_list_head under comp->sem. The event broadcast (mlx5_devcom_locked_send_event()) walks that list. Hence, a caller can read the expected size, but send_event won't be sent to all peers. In the SD group registration path, this lets a member broadcast its role-election event over an incomplete list, electing a primary that never completes the group, is never marked ready, and leaves the group with a stale primary. Track the number of linked comp_devs in a dedicated counter, maintained under comp->sem together with the list add/remove, and return it from mlx5_devcom_comp_get_size(). Fixes: 9bb1ac80738a ("net/mlx5: devcom, Add component size getter") Signed-off-by: Shay Drory Reviewed-by: Akiva Goldberger Signed-off-by: Tariq Toukan Link: https://patch.msgid.link/20260915113459.3934760-2-tariqt@nvidia.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 8ebfd34fef750e754624a00c38d5f17ca81fec25 Author: Heyang Tan Date: Mon Sep 14 10:05:21 2026 +0800 octeontx2-af: use seq_file for rsrc_alloc debugfs [ Upstream commit 39c6580765dad6477fb2637f6f616e0d276aae65 ] The rsrc_alloc debugfs reader writes rows directly to userspace without respecting the caller's read count. It also uses the current row length as the userspace stride, which can corrupt output when rows have different widths. Use seq_file to handle userspace buffer sizes, offsets, and partial reads, and write output columns directly to the seq_file buffer. Fixes: 23205e6d06d4 ("octeontx2-af: Dump current resource provisioning status") Signed-off-by: Heyang Tan Reviewed-by: Ratheesh Kannoth Link: https://patch.msgid.link/20260914020521.146-1-thy15333007817@163.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 6127164efa94e56d1571374827430b615db701c4 Author: Kyle Hendry Date: Tue Sep 15 10:39:20 2026 -0700 net: pcs: rzn1-miic: Fix config array initialization [ Upstream commit daf677c2c6449011ee695d55b48b5b2977a36f88 ] Fix memset parameters to initialize the entire DT value array Fixes: f39e968dc168a7bd ("net: pcs: rzn1-miic: Move configuration data to SoC-specific struct") Reviewed-by: Geert Uytterhoeven Signed-off-by: Kyle Hendry Reviewed-by: Lad Prabhakar Link: https://patch.msgid.link/20260915-rzn1-miic-fix-array-v5-1-b7173fd5b97d@reliablecontrols.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit f6696481eedfd5752b4b56eed715d0b0fe48b9c5 Author: Jamal Hadi Salim Date: Wed Sep 16 06:01:14 2026 -0400 net/sched: cls_u32: fix manual hash table handle IDR aliasing [ Upstream commit 0a5f5d9e94dead312d32c366b917c64e552b72f7 ] A u32 hash table created with an explicit handle ('tc filter add ... handle 801: u32 divisor N') keys its IDR entry on the raw handle, while the destroy paths free it under handle2id(handle). The two key domains disagree for handles in the 0x800..0xFFF htid range: handle2id() folds them back into the auto-allocated id space (1..0x7FF). A manual table therefore leaves its raw-keyed IDR entry unreachable on delete (a permanent leak), and its delete can drop the idr entry of an unrelated live auto table. A later auto allocation can then hand out a handle that aliases the live manual table; u32_lookup_ht() first-match routes lookups and TCA_U32_LINK for that htid to the wrong table. Key the divisor-path alloc on handle2id(handle) so allocation and removal share one key domain. A manual handle that maps onto an id already in use is rejected with -ENOSPC, and auto allocation skips ids held by live manual tables. Conditions to recreate: ip link add test0 type dummy tc qdisc add dev test0 clsact tc filter add dev test0 ingress protocol ip pref 1 \ handle 801: u32 divisor 16 tc filter add dev test0 ingress protocol ip pref 2 u32 divisor 16 tc -d filter show dev test0 ingress | grep 'fh 801:' # unpatched: two live tables with handle 0x80100000 (the pref 2 root # hnode is auto-allocated id 1); patched: the auto hnode takes id 2. Also tested with a poc with a live u32 table on the block, add/delete a manual table 'handle 901: u32 divisor 1' twice; unpatched, the re-add fails with -ENOSPC because the raw key leaked on the first delete. Fixes: 73af53d82076 ("net: sched: cls_u32: Fix u32's systematic failure to free IDR entries for hnodes.") Reported-by: Sashiko (gemini + nipa) Closes: https://sashiko.dev/#/patchset/20260822222049.114526-1-jhs@mojatatu.com Reviewed-by: Victor Nogueira Tested-by: hybris Signed-off-by: Jamal Hadi Salim Reviewed-by: Simon Horman Link: https://patch.msgid.link/QDISC-LQFE.v1.20260911041746.1@mojatatu.com Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit f4277b19edbc6abc9e1db7d6f6148608e8cb3c34 Author: Björn Töpel Date: Tue Sep 15 12:49:15 2026 +0200 eth: fbnic: Fix payload page pool error cleanup [ Upstream commit 8e0b235bd918d06f54ba8fddd2c3ddc36ca59c15 ] The payload page pool pointer contains an error pointer when its allocation fails. The cleanup path passes that error pointer to page_pool_destroy() instead of destroying the header page pool. This can dereference the error pointer and leave the header page pool allocated. Destroy the header page pool instead. Fixes: 8a11010fdd96 ("eth: fbnic: allocate unreadable page pool for the payloads") Reported-by: Sashiko Link: https://lore.kernel.org/netdev/178915061000.219967.7726187707862333281@kernel.org/ Signed-off-by: Björn Töpel Reviewed-by: Simon Horman Link: https://patch.msgid.link/20260915104917.3978113-1-bjorn@kernel.org Signed-off-by: Jakub Kicinski Signed-off-by: Sasha Levin commit 6e271f093d15d323f42de26f66556af53fa8e19f Author: Weiming Shi Date: Tue Sep 15 01:02:07 2026 +0800 bpf: Skip unsettled links in link iterator [ Upstream commit 50e80e2bb5e2be8515205b9c496b9640ddefa434 ] bpf_link_prime() inserts a link into link_idr before anon_inode_getfile() succeeds and before bpf_link_settle() publishes the ID in link->id. bpf_link_by_id() treats such an ID-zero link as unsettled, but the link iterator takes a reference without this check. If anon_inode_getfile() then fails, the creator removes the ID and frees its still-private link directly. The iterator is left with a dangling reference and its next bpf_link_put() accesses freed memory. Treat ID-zero entries as transient in bpf_link_get_curr_or_next(), just as bpf_link_by_id() does. BUG: KASAN: slab-use-after-free in bpf_link_put Write of size 8 by task exp/384 Call Trace: bpf_link_put kernel/bpf/syscall.c:3372 bpf_link_seq_next kernel/bpf/link_iter.c:33 bpf_seq_read kernel/bpf/bpf_iter.c:158 vfs_read fs/read_write.c:572 ksys_read fs/read_write.c:716 do_syscall_64 arch/x86/entry/syscall_64.c:84 entry_SYSCALL_64_after_hwframe arch/x86/entry/entry_64.S:121 Kernel panic - not syncing: KASAN: panic_on_warn set ... Fixes: 9f8836127308 ("bpf: Add bpf_link iterator") Reported-by: Xiang Mei Signed-off-by: Weiming Shi Signed-off-by: Andrii Nakryiko Link: https://lore.kernel.org/bpf/20260914170206.170723-2-bestswngs@gmail.com Signed-off-by: Sasha Levin commit b5c0cff0c5331582ab01c4c423731453a8381acc Author: Zhiling Zou Date: Fri Sep 11 00:18:24 2026 +0800 xsk: Use a 32-bit compare in xsk_map_gen_lookup [ Upstream commit 70504de0bb627848667207bec7ccfd647deb8814 ] xsk_map_gen_lookup() loads a u32 key and compares it with max_entries using BPF_JMP_IMM. BPF immediates are sign-extended to 64 bits, so a max_entries value of 0x80000000 or higher becomes a threshold larger than every zero-extended 32-bit key. An out-of-range index then skips the bounds check and the generated lookup reads past xsk_map[]. Compare with BPF_JMP32_IMM so the check stays in 32-bit unsigned range. Fixes: e65650f291ee ("bpf: Implement map_gen_lookup() callback for XSKMAP") Reported-by: Vega Signed-off-by: Zhiling Zou Signed-off-by: Alexei Starovoitov Reviewed-by: Emil Tsalapatis Link: https://patch.msgid.link/7d2cb8e8dfaa9eb8fdff85156987a60960787dc3.1789056660.git.zhilinz@nebusec.ai Signed-off-by: Eduard Zingerman Signed-off-by: Sasha Levin commit e4f3e476cd4066a815ef68cd6caa1828b9c167d5 Author: Julian Sun Date: Tue Sep 15 12:49:12 2026 +0800 fs: avoid repeated scans in evict_inodes() [ Upstream commit 7459c021874246c196f397686e100702f059b9e7 ] We observed hung tasks when users attempted to unmount a filesystem after its disk had been removed while still in use. During device removal, fs_bdev_mark_dead() calls evict_inodes() while holding s_umount. Each time evict_inodes() drops s_inode_list_lock to reschedule, it restarts the walk from the head of s_inodes. With many referenced inodes at the head of the list, these restarts repeatedly scan the same inodes without reclaiming them. This can keep s_umount held for a long time, blocking concurrent umount attempts and triggering hung-task reports. Keep the current inode, already marked I_FREEING, out of the disposal batch until s_inode_list_lock is reacquired. Resume the walk from this inode and dispose of it in a later batch or at the end of the walk. The zero-refcount and state checks under i_lock allow this walker to claim the inode by setting I_FREEING and removing it from the LRU. Other reclaimers skip the inode, leaving this walker responsible for eviction. Only evict() removes it from s_inodes, so keeping it out of the disposal batch ensures that it remains on the list while the lock is dropped. After reacquiring the lock, reading its current next pointer accounts for concurrent removal of following inodes. The existing inode lifetime rules prohibit acquiring a reference to an inode marked I_FREEING or I_WILL_FREE. __iget() requires its caller to hold i_lock and establish that taking a reference is valid. Inode lookup and igrab() check these flags under i_lock when acquiring a reference from zero. ihold() requires an existing reference, which would keep i_count nonzero and prevent this walker from claiming the inode. These rules already allow iput_final() and the inode shrinker to release i_lock after setting I_FREEING and before eviction completes. A temporary __iget() reference would also keep the inode on the list, but its release must preserve last-reference handling. Another user can acquire a reference, update lazy timestamps and drop its reference while the pin is held. If the pin becomes the last reference, dropping it with atomic_dec_and_test() and evicting directly bypasses iput()'s lazytime handling and can lose those timestamp updates. Releasing the pin with iput() preserves that handling, but does not guarantee eviction. fs_bdev_mark_dead() runs with SB_ACTIVE set, so iput() may retain the inode in cache, whereas evict_inodes() must evict eligible zero-reference inodes. The inode may also have been freed when iput() returns, so the walker cannot then use it to force eviction. Using I_FREEING preserves the existing eviction behavior without introducing an additional last-reference transition. The xfstests auto group passed on ext4 and XFS with known unrelated failures excluded. No new issues were observed, and the previously reproducible hung task no longer occurs with this patch. Fixes: ac05fbb40062 ("inode: don't softlockup when evicting inodes") Signed-off-by: Julian Sun Link: https://patch.msgid.link/20260915044912.3183440-1-sunjunchao@bytedance.com Reviewed-by: Jan Kara Signed-off-by: Christian Brauner (Amutable) Signed-off-by: Sasha Levin commit 22225cd612b26240cc2b8ea1bc1e225ea17cc952 Author: Vineeth Vijayan Date: Thu Sep 10 11:32:03 2026 +0200 s390/cio: Guard PMCW field accesses with dnv check [ Upstream commit 9590f4d83880dfb5a81906e48e72779248fbe8f0 ] When PMCW.DNV is 0, no I/O device is associated with the subchannel. However, several code paths access PMCW fields directly from the cached sch->schib without first invoking the update helper. Add explicit DNV validation before accessing PMCW fields from the cached SCHIB to avoid using invalid data. Reported-by: William Bezenah Signed-off-by: Vineeth Vijayan Reviewed-by: Peter Oberparleiter Fixes: 8c58a229688c ("s390/cio: Do not unregister the subchannel based on DNV") Signed-off-by: Heiko Carstens Signed-off-by: Sasha Levin commit 5eae708cca8f4fd9ec09cc65368c67f9efb8939e Author: Vineeth Vijayan Date: Thu Sep 10 11:32:02 2026 +0200 s390/cio: Check pmcw.dnv before pmcw.ena in I/O entry points [ Upstream commit f6f2985eabdb2bfdc82ce90a1ea3ec53ba795f34 ] The device number valid (dnv) bit in the PMCW must be checked before acting on any other PMCW fields for IO-type subchannels. A subchannel with dnv=0 has no valid device number associated, making it meaningless to evaluate the enabled (ena) state or issue any I/O instruction against it. Reported-by: William Bezenah Signed-off-by: Vineeth Vijayan Reviewed-by: Peter Oberparleiter Fixes: 8c58a229688c ("s390/cio: Do not unregister the subchannel based on DNV") Signed-off-by: Heiko Carstens Signed-off-by: Sasha Levin commit 591ba131193ee04cbf452041ae09b21d150706f4 Author: Vineeth Vijayan Date: Thu Sep 10 11:32:01 2026 +0200 s390/cio: Fix cio_update_schib() to not cache invalid schib [ Upstream commit 29d9e5835d89223aa913dcf7b942cc1c148bdd25 ] When pmcw.dnv is 0, the contents of all SCHIB fields are unpredictable. Zero sch->schib in that case to prevent subsequent code from making decisions based on unpredictable data. Reported-by: William Bezenah Signed-off-by: Vineeth Vijayan Reviewed-by: Peter Oberparleiter Fixes: 8c58a229688c ("s390/cio: Do not unregister the subchannel based on DNV") Signed-off-by: Heiko Carstens Signed-off-by: Sasha Levin commit 8bd9905dfbd978f30d1a20253e4dfb37152af7bf Author: Karl Mehltretter Date: Mon Sep 7 07:58:47 2026 +0200 s390/pci/docs: Fix sriov_numvfs attribute name [ Upstream commit 4525a911049543c23885a540a788d13be318a486 ] The attribute is sriov_numvfs (drivers/pci/iov.c); the document names it sriov_numvf, which does not exist. Use sriov_numvfs. Fixes: de267a7c71ba ("s390/pci: Documentation for zPCI") Assisted-by: LLM Signed-off-by: Karl Mehltretter Reviewed-by: Randy Dunlap Signed-off-by: Heiko Carstens Signed-off-by: Sasha Levin commit e163f1dccede7ba298187b8937382bbb8ecbf37c Author: Niklas Schnelle Date: Tue Apr 7 15:24:45 2026 +0200 docs: s390/pci: Improve and update PCI documentation [ Upstream commit 737c4f4a241ca85c597ca2ef1a6f8446bf681ab5 ] Update the s390 specific PCI documentation to better reflect current behavior and terms such as the handling of Isolated VFs via commit 25f39d3dcb48 ("s390/pci: Ignore RID for isolated VFs"). Add a descriptions for /sys/firmware/clp/uid_checking which was added in commit b043a81ce3ee ("s390/pci: Expose firmware provided UID Checking state in sysfs") but missed documentation. Similarly add documentation for the fidparm attribute added by commit 99ad39306a62 ("s390/pci: Expose FIDPARM attribute in sysfs") and add a list of pft values and their names. Finally improve formatting of the different attribute descriptions by adding a separating colon. Reviewed-by: Farhan Ali Acked-by: Randy Dunlap Tested-by: Randy Dunlap Reviewed-by: Matthew Rosato Signed-off-by: Niklas Schnelle Reviewed-by: Gerd Bayer Link: https://lore.kernel.org/r/20260407-uid_slot-v8-1-15ae4409d2ce@linux.ibm.com Signed-off-by: Vasily Gorbik Stable-dep-of: 4525a9110495 ("s390/pci/docs: Fix sriov_numvfs attribute name") Signed-off-by: Sasha Levin commit 6675f801b0d60b0de00d5bfaefe8b494e893125e Author: Bart Van Assche Date: Mon Aug 31 12:27:20 2026 -0700 scsi: megaraid_sas: Protect megasas_get_ctrl_info() in megasas_resume() [ Upstream commit 42d1221d321e55afc7bba9109a77aaf5a817c8a3 ] Protect the megasas_get_ctrl_info() call in megasas_resume() with instance->reset_mutex using scoped_guard(). megasas_get_ctrl_info() may release and reacquire instance->reset_mutex. Hence, calling this function without holding instance->reset_mutex is not safe. Fixes: c3b10a55abc9 ("scsi: megaraid_sas: Update controller info during resume") Cc: Kashyap Desai Cc: Sumit Saxena Cc: Shivasharan S Cc: Chandrakanth patil Signed-off-by: Bart Van Assche Link: https://patch.msgid.link/f06b5ee432b21cf293f0663e15b64f75a84b9fd5.1788204406.git.bvanassche@acm.org Signed-off-by: Martin K. Petersen (Oracle) Signed-off-by: Sasha Levin commit 0117e731435bd299a1ca6b7be9259bc93c5ec4ec Author: Eva Kurchatova Date: Wed Sep 16 23:44:31 2026 +0300 selftests: cgroup: give the O_TMPFILE open in get_temp_fd() a mode [ Upstream commit c774ec8f0a5d02a06d34c27f5a7de7e333b91265 ] O_TMPFILE, like O_CREAT, needs the third argument. Without it glibc refuses the call at compile time as soon as fortification is on: In function 'open', inlined from 'get_temp_fd' at test_memcontrol.c:33:9: /usr/include/bits/fcntl2.h:52:11: error: call to '__open_missing_mode' declared with attribute error: open with O_CREAT or O_TMPFILE in second argument needs 3 arguments The fortify checks take effect only once the compiler optimises, and cgroup/Makefile builds with "-Wall -pthread" alone, so this goes unnoticed in a plain build. Building the tests with the flags distributions commonly use, -O2 -D_FORTIFY_SOURCE=3, loses test_memcontrol entirely. Fixes: 84092dbcf901 ("selftests: cgroup: add memory controller self-tests") Signed-off-by: Eva Kurchatova Signed-off-by: Tejun Heo Signed-off-by: Sasha Levin commit 2189a219397d8e73808a67c4aade837b635705da Author: Lee Jones Date: Tue Sep 15 12:08:22 2026 +0000 Bluetooth: mgmt: Dequeue pending mesh_send_sync entries on cancel [ Upstream commit 71af682ba4692c2ed9ace4c3d4ca462ae368c029 ] In send_cancel(), pending mesh_tx objects are removed from the hdev->mesh_pending list and freed via mesh_send_complete(). However, if a mesh transmission was already queued onto hdev->cmd_sync_work_list via mesh_next(), the queued entry retains a raw pointer to mesh_tx. When hci_cmd_sync_work later processes the entry, it attempts to execute mesh_send_sync and its destroy callback mesh_send_start_complete using the already freed mesh_tx pointer, leading to a use-after-free. Fix this by invoking hci_cmd_sync_dequeue() for mesh_send_sync on the target mesh_tx before completing it. If the entry is found and dequeued, its destroy callback will complete and free the object; otherwise, mesh_send_complete() is called directly. Additionally, ensure the transmission queue advances after cancellation or errors. In mesh_send_start_complete(), call mesh_next() on error unless err is -ECANCELED, because hci_cmd_sync_dequeue() holds hdev->cmd_sync_work_lock and calling mesh_next() synchronously would deadlock. Instead, advance the queue in send_cancel() once the lock is released and if no transmission is in progress. Fixes: b338d91703fa ("Bluetooth: Implement support for Mesh") Signed-off-by: Lee Jones Signed-off-by: Luiz Augusto von Dentz Signed-off-by: Sasha Levin commit 8a36598050cd1e5abb1d5ed42cc224cfcb538c4f Author: Zijun Hu Date: Tue Sep 15 19:17:18 2026 -0700 Bluetooth: btnxpuart: Fix skb leak in nxp_process_fw_dump() [ Upstream commit f2bbb36426581045a8bf7793da5419b9375e4348 ] When CONFIG_DEV_COREDUMP=n, hci_devcd_append() returns -EOPNOTSUPP without freeing its skb argument. This leaks the cloned skb and also prevents nxp_set_ind_reset() from being called to perform recovery. Fix by guarding the hci_devcd_append(hdev, skb_clone(skb, GFP_ATOMIC)) call with IS_ENABLED(CONFIG_DEV_COREDUMP). Fixes: 998e447f443f ("Bluetooth: btnxpuart: Add support for HCI coredump feature") Signed-off-by: Zijun Hu Signed-off-by: Luiz Augusto von Dentz Signed-off-by: Sasha Levin commit 6fdabd654578e5248c4d0f2435901cc149cb2dd8 Author: Christiano Amora Date: Wed Sep 16 10:46:22 2026 -0300 Bluetooth: SMP: reject Security Request over BR/EDR [ Upstream commit f033482d76a9f18080c7a40c5f9c678bd7adc8f3 ] Bose QC Ultra Headphones (dual-mode, same public address on both transports) occasionally send an SMP Security Request on the BR/EDR SMP fixed channel right after the ACL link is encrypted. The kernel handles it as if it were an LE link: smp_cmd_security_req() has no transport check, smp_ltk_encrypt() looks up an LTK with the ACL connection's dst_type, and hci_find_ltk() matches the peer's LE LTK because the LE public address type is stored as ADDR_LE_DEV_PUBLIC (0), the same value as BDADDR_BREDR. HCI_OP_LE_START_ENC is then issued on the ACL handle, the controller rejects it with Invalid HCI Command Parameters, and hci_cs_le_start_enc() disconnects the link with HCI_ERROR_AUTH_FAILURE. The headphones drop within a second of connecting, before any profile is up; a manual reconnect works. btmon (MediaTek MT7922, kernel 7.0.12): > HCI Event: Encryption Change (0x08) plen 4 Status: Success (0x00) Handle: 50 Address: BC:87:FA:47:73:5E (Bose Corporation) Encryption: Enabled with AES-CCM (0x02) > ACL Data RX: Handle 50 flags 0x02 dlen 6 BR/EDR SMP: Security Request (0x0b) len 1 Authentication requirement: No bonding, No MITM, SC (0x08) < HCI Command: LE Start Encryption (0x08|0x0019) plen 28 Handle: 50 Address: BC:87:FA:47:73:5E (Bose Corporation) > HCI Event: Command Status (0x0f) plen 4 LE Start Encryption (0x08|0x0019) ncmd 1 Status: Invalid HCI Command Parameters (0x12) < HCI Command: Disconnect (0x01|0x0006) plen 3 Handle: 50 Address: BC:87:FA:47:73:5E (Bose Corporation) Reason: Authentication Failure (0x05) SMP over BR/EDR is limited to cross-transport key derivation; the Security Request procedure (Core Specification Vol 3, Part H, Section 2.4.6, PDU in Section 3.6.7) has no BR/EDR counterpart. Reply with Pairing Failed / Command Not Supported on a non-LE link, before the PDU is parsed, and keep the connection. The reply is sent directly rather than through smp_failure(): rejecting a command on the wrong transport is not an authentication failure, and MGMT_EV_AUTH_FAILED would make bluetoothd disconnect the device. Tested on the affected host (kernel 7.0.12, MediaTek MT7922, Bose QC Ultra) with the patched module built out of tree: 7 days and 49 reconnects without a drop, against 2 drops in the 3 days before the patch. Every disconnect in that week had a userspace or remote reason. Fixes: b5ae344d4c0f ("Bluetooth: Add full SMP BR/EDR support") Assisted-by: LLM Signed-off-by: Christiano Amora Signed-off-by: Luiz Augusto von Dentz Signed-off-by: Sasha Levin commit 9c051b12ca106aa11bc3a6c24c6e8506ecc0b560 Author: Andre Przywara Date: Mon Sep 14 12:07:44 2026 +0200 pinctrl: sunxi: A523: fix voltage withstand encoding [ Upstream commit ef56085dfd1df3a53ecce8a4ff440cf6d67f430d ] The Allwinner A523 uses the same GPIO voltage "withstand" programming (setting the input level voltage thresholds) as the previous SoCs, but for some odd reason inverts the encoding of 1.8V vs. 3.3V. Add a new bias voltage type to note this difference, and select it for the A523. At the same time also use the newer "CTL" version, which in addition allows to turn off the withstand programming for I/O voltages other than exact 1.8V or 3.3V (for instance for 2.5V sometimes used for Ethernet PHYs). The A523 has that enable register, but didn't use it so far. This fixes eMMC and reportedly Ethernet operation on some A523 boards. Fixes: 648be4cd9517 ("pinctrl: sunxi: Add support for the Allwinner A523") Signed-off-by: Andre Przywara Tested-by: Per Larsson Tested-by: Juan Manuel Lopez Carrillo Reviewed-by: Chen-Yu Tsai Tested-by: Chen-Yu Tsai # Fixes eMMC on Orange Pi 4A Signed-off-by: Linus Walleij Signed-off-by: Sasha Levin commit 0af16d8be3ec7ff28bb56a0da79c9aed66ac2c3b Author: ZHOU Jiaxiang Date: Wed Sep 16 21:58:22 2026 +0800 scsi: sd_zbc: Reject disks with too many zones [ Upstream commit b6ec0f79745967c751c85df373062c8d15e45fc4 ] sd_zbc_read_zones() computes the number of zones with 64-bit arithmetic and stores the result in the unsigned int nr_zones field of struct zoned_disk_info, silently truncating counts that exceed 32 bits. The truncated count is later used to size per-zone resources, while the device may still report more zones than fit. Moreover, sd_zbc_report_zones() counts the reported zones with a signed int zone_idx, which overflows past INT_MAX. Reject devices reporting more than INT_MAX zones at scan time; such a device is not realistic for any medium that exists today, and accepting it produces inconsistent zone bookkeeping. Fixes: 89d947561077 ("sd: Implement support for ZBC devices") Signed-off-by: ZHOU Jiaxiang Reviewed-by: Damien Le Moal Link: https://patch.msgid.link/C41798AB5AA6BF2B+20260916135822.32584-3-me@fxti.xyz Signed-off-by: Martin K. Petersen (Oracle) Signed-off-by: Sasha Levin commit 0cb85377daaefb4ce5ce72975f2deb27ee0e0cad Author: Ran Hongyun Date: Mon Jul 13 19:55:25 2026 +0800 squashfs: Add dictionary size range check to prevent shift-out-of-bounds [ Upstream commit 1f7745fb3580152ca902ef181b605f33cabfb1d0 ] When an abnormal SquashFS image (COMP_OPTS flag is 1 but dictionary size is 0) is mounted, and performs shift operations using dictionarysize, the shift exponent is -1, causing a shift-out-of-bounds. Detail as below: squashfs_comp_opts(msblk, buffer, length) squashfs_xz_comp_opts() if (comp_opts) n = ffs(opts->dict_size) - 1;<----opts->dict_size=0, n=-1 if (opts->dict_size != (1 << n) && opts->dict_size != (1 << n) + (1 << (n + 1))) <----shift-out-of-bounds Fix it by adding a dictionary size range check before the shift operation. Fixes: ff750311d30a ("Squashfs: add compression options support to xz decompressor") Signed-off-by: Ran Hongyun Link: https://patch.msgid.link/20260713115525.2661734-1-ranhongyun1@huawei.com Reviewed-by: Phillip Lougher Reviewed-by: Zhihao Cheng Signed-off-by: Christian Brauner (Amutable) Signed-off-by: Sasha Levin commit 5d54693bc457c38848d3c51e9ffacd14f50ba201 Author: Mark Brown Date: Tue Sep 1 22:47:00 2026 +0100 KVM: arm64: Fix FGT mapping for HFGITR_EL2.nGCSEPP [ Upstream commit 089e4f3c4862ba3f29dff2361caa8084879194fd ] The encoding to trap mapping currently maps a FGT on OP_GCSPOPX to HFGITR_EL2.nGCSEPP but as per DDI0601 2026-06 this FGT controls trapping of GCSPUSHX and GCSPOPCX, and not the separate GCSPOPX instruction. Update the mapping to reflect the architecture. Fixes: 863ac38984a82 ("KVM: arm64: Add missing HFGITR_EL2 FGT entries to nested virt") Reviewed-by: Leonardo Bras Signed-off-by: Mark Brown Reviewed-by: Lorenzo Stoakes (ARM) Link: https://patch.msgid.link/20260901-arm64-gcs-v20-2-f31750bdfadb@kernel.org Signed-off-by: Oliver Upton Signed-off-by: Sasha Levin commit cb68b5589318fabebc7cc4dcc3a5f067daabf76a Author: Karl Mehltretter Date: Sat Aug 29 07:48:55 2026 +0200 KVM: arm64: Return -EINVAL for an empty SMCCC filter range at base 0 [ Upstream commit 64dc6f1db7e620f2e9337bb181f305fb0561da79 ] kvm_smccc_set_filter() only rejects a range if its inclusive end, base + nr_functions - 1, is below base. That catches an empty range (nr_functions == 0) at every nonzero base, but at base 0 the end wraps to U32_MAX and KVM tries to insert [0, U32_MAX], which overlaps the reserved Arm Architecture Calls ranges. KVM_ARM_VM_SMCCC_FILTER then returns -EEXIST instead of the -EINVAL that the smccc_filter selftest expects for an empty range. Reject a zero function count explicitly. Tested with a userspace reproducer on an arm64 VHE host under QEMU TCG: EEXIST before, EINVAL after. Fixes: 821d935c87bc ("KVM: arm64: Introduce support for userspace SMCCC filtering") Assisted-by: LLM Signed-off-by: Karl Mehltretter Reviewed-by: Steffen Eiden Reviewed-by: Fuad Tabba Tested-by: Fuad Tabba Link: https://patch.msgid.link/20260829054856.70549-2-kmehltretter@gmail.com Signed-off-by: Oliver Upton Signed-off-by: Sasha Levin commit 4efda204087145d40a84344edec7e9d3885ad3ab Author: Fuad Tabba Date: Tue Aug 25 09:59:48 2026 +0100 KVM: arm64: Derive GUEST_HAS_SVE from the SVE feature bit at EL2 [ Upstream commit 4f16c5fc8dc4c5596e3777ab9f449a54e3f85fd5 ] pkvm_init_features_from_host() takes KVM_ARCH_FLAG_GUEST_HAS_SVE and KVM_ARM_VCPU_SVE from the host separately, but pkvm_vcpu_init_sve() tests the bit while vcpu_has_sve() reads the flag. A host that sets the flag without the bit gets a vCPU with a NULL sve_state that the world switch loads the guest's SVE state from. Derive the flag from the bit, and drop the protected path's copy of the host's flag, which is dead code since protected VMs are not allowed SVE. Fixes: 41d6028e28bd ("KVM: arm64: Convert the SVE guest vcpu flag to a vm flag") Signed-off-by: Fuad Tabba Link: https://patch.msgid.link/20260825085948.1674721-5-fuad.tabba@linux.dev Signed-off-by: Oliver Upton Signed-off-by: Sasha Levin commit f1269521671465ccbdedc7d1470a9445402bbc7f Author: Fuad Tabba Date: Thu Dec 11 10:47:05 2025 +0000 KVM: arm64: Include VM type when checking VM capabilities in pKVM [ Upstream commit 43a21a0f0c4ab7de755f2cee2ff4700f26fe0bba ] Certain features and capabilities are restricted in protected mode. Most of these features are restricted only for protected VMs, but some are restricted for ALL VMs in protected mode. Extend the pKVM capability check to pass the VM (kvm), and use that when determining supported features. Signed-off-by: Fuad Tabba Link: https://patch.msgid.link/20251211104710.151771-6-tabba@google.com Signed-off-by: Marc Zyngier Stable-dep-of: 4f16c5fc8dc4 ("KVM: arm64: Derive GUEST_HAS_SVE from the SVE feature bit at EL2") Signed-off-by: Sasha Levin commit bbcd7e851d57dcc6abc146df8be6b852153e2188 Author: Fuad Tabba Date: Thu Dec 11 10:47:03 2025 +0000 KVM: arm64: Fix MTE flag initialization for protected VMs [ Upstream commit ebbcaece84738f71b35f32339bdeb8776004e641 ] The function pkvm_init_features_from_host() initializes guest features, propagating them from the host. The logic to propagate KVM_ARCH_FLAG_MTE_ENABLED (Memory Tagging Extension) has a couple of issues. First, the check was in the common path, before the divergence for protected and non-protected VMs. For non-protected VMs, this was unnecessary, as 'kvm->arch.flags' is completely overwritten by host_arch_flags immediately after, which already contains the MTE flag. For protected VMs, this was setting the flag even if the feature is not allowed. Second, the check was reading 'host_kvm->arch.flags' instead of using the local 'host_arch_flags', which is read once from the host flags. Fix these by moving the MTE flag check inside the protected-VM-only path, checking if the feature is allowed, and changing it to use the correct host_arch_flags local variable. This ensures non-protected VMs get the flag via the bulk copy, and protected VMs get it via an explicit check. Fixes: b7f345fbc32a ("KVM: arm64: Fix FEAT_MTE in pKVM") Reviewed-by: Ben Horgan Signed-off-by: Fuad Tabba Link: https://patch.msgid.link/20251211104710.151771-4-tabba@google.com Signed-off-by: Marc Zyngier Stable-dep-of: 4f16c5fc8dc4 ("KVM: arm64: Derive GUEST_HAS_SVE from the SVE feature bit at EL2") Signed-off-by: Sasha Levin commit c142034ad3a84a15449022343a8938ba1d0167b7 Author: Marc Zyngier Date: Thu Nov 20 17:24:55 2025 +0000 KVM: arm64: vgic-v3: Fix GICv3 trapping in protected mode [ Upstream commit 567ebfedb5bd204a8ce6a11695f02730f1bf57f4 ] As we are about to start trapping a bunch of extra things, augment the pKVM trap description with all the registers trapped by ICH_HCR_EL2.TC, making them legal instead of resulting in a UNDEF injection in the guest. While we're at it, ensure that pKVM captures the vgic model so that it can be checked by the emulation code. Tested-by: Fuad Tabba Signed-off-by: Marc Zyngier Tested-by: Mark Brown Link: https://msgid.link/20251120172540.2267180-6-maz@kernel.org Signed-off-by: Oliver Upton Stable-dep-of: 4f16c5fc8dc4 ("KVM: arm64: Derive GUEST_HAS_SVE from the SVE feature bit at EL2") Signed-off-by: Sasha Levin commit bd46b41f412fb81c50bac1a0929ec31320f3a807 Author: Fuad Tabba Date: Tue Aug 25 09:59:46 2026 +0100 KVM: arm64: Do not clear VM-wide SVE feature on vCPU init failure [ Upstream commit a1b3c788ad31837e348075e93dbba3f447492776 ] pkvm_vcpu_init_sve() clears KVM_ARM_VCPU_SVE in kvm->arch.vcpu_features when it fails, but vcpu_has_sve() tests KVM_ARCH_FLAG_GUEST_HAS_SVE, which is left set. Later vCPUs on that VM then skip the SVE setup and register with a NULL sve_state, which the guest's first FP access hands to sve_load_state(). Return the error without touching vcpu_features. Fixes: 5db1bef93342 ("KVM: arm64: Track SVE state in the hypervisor vcpu structure") Signed-off-by: Fuad Tabba Link: https://patch.msgid.link/20260825085948.1674721-3-fuad.tabba@linux.dev Signed-off-by: Oliver Upton Signed-off-by: Sasha Levin commit faf66e55e87498ab4bf88cdf9a629e61b3f2d3bc Author: Fuad Tabba Date: Tue Aug 25 09:59:45 2026 +0100 KVM: arm64: Validate the SVE vector length in pkvm_vcpu_init_sve() [ Upstream commit 2a2eb10795a1e495aebc7f829ccecb72c05b4fd9 ] pkvm_vcpu_init_sve() clamps only the upper bound of the host-provided sve_max_vl, so an invalid vector length reaches sve_state_size_from_vl() and the WARN_ON() there, which is fatal at EL2. The existing !sve_state_size test rejects such a length, but only after the macro has run. Check sve_vl_valid() before deriving the state size. A valid length cannot yield a zero size, so the !sve_state_size test goes with it. Fixes: 5db1bef93342 ("KVM: arm64: Track SVE state in the hypervisor vcpu structure") Reported-by: Stefan Teodorescu Reviewed-by: Marc Zyngier Signed-off-by: Fuad Tabba Link: https://patch.msgid.link/20260825085948.1674721-2-fuad.tabba@linux.dev Signed-off-by: Oliver Upton Signed-off-by: Sasha Levin commit 65bffaf100df8f674764192d84a0e39c2e47be4a Author: Fuad Tabba Date: Fri Aug 21 07:44:44 2026 +0100 KVM: arm64: vgic-its: Skip unreachable devices instead of failing the save [ Upstream commit cc5d96036e01ac330d24b2f0c336d60f82ab4930 ] vgic_its_save_device_tables() aborts with -EINVAL when a device's entry falls outside the device table, which a guest can arrange on its own: an indirect table lets it clear an L1 entry's valid bit without touching GITS_BASER. That fails a save userspace should be able to issue reliably. Skip the device instead, and point the saved DTE chain past it, as commit ad1e686e2378d ("KVM: arm64: vgic-its: Point saved ITEs at the next valid entry") does for ITEs. compute_next_devid_offset() takes the next device off the list whether or not it was saved, so the predecessor would otherwise point at an entry the save never wrote. Restore follows that offset while it stays inside the table being scanned: within an L2 block, or anywhere in a flat table. Both need userspace to remove a memslot under the table, since dropping an L1 entry takes the whole block with it and scan_its_table() stops at the block boundary. Fixes: 57a9a117154c9 ("KVM: arm64: vgic-its: Device table save/restore") Suggested-by: Marc Zyngier Link: https://lore.kernel.org/all/86bjaz5s6v.wl-maz@kernel.org/ Signed-off-by: Fuad Tabba Reviewed-by: Marc Zyngier Link: https://patch.msgid.link/20260821064445.615838-4-fuad.tabba@linux.dev Signed-off-by: Oliver Upton Signed-off-by: Sasha Levin commit 80a6b712dc193993d9f0b63976674bb7ff7989fe Author: Fuad Tabba Date: Fri Aug 21 07:44:42 2026 +0100 KVM: arm64: vgic-its: Free the caches when GITS_BASER changes [ Upstream commit 8cd92f77ae4f5371a7d581f8324c24919670b304 ] A guest that disables the ITS and re-points or shrinks GITS_BASER with VALID still set keeps the devices and collections it mapped against the old table, as KVM frees them only when VALID is cleared. The contents of the table are IMPLEMENTATION DEFINED, so a write that gives GITS_BASER a different address or size may lose whatever the old value described. Free the list whenever the stored value changes, and drop the translation cache with it. The cache is not empty just because the ITS is disabled: its->enabled is written under the cmd_lock, while vgic_its_resolve_lpi() tests it under the its_lock, so an injection can still cache an entry after the ITS was disabled. Hence the invalidation inside the its_lock section. Test for a change rather than a write: its_restore_enable() rewrites GITS_BASER from its probe-time cache on resume, and KVM reports GITS_TYPER.HCC as 0, so nothing re-maps the boot CPU's collection afterwards. Fixes: 36d6961c2b481 ("KVM: arm/arm64: vgic-its: Free caches when GITS_BASER Valid bit is cleared") Suggested-by: Marc Zyngier Link: https://lore.kernel.org/all/87ecg9owwa.wl-maz@kernel.org/ Signed-off-by: Fuad Tabba Reviewed-by: Marc Zyngier Link: https://patch.msgid.link/20260821064445.615838-2-fuad.tabba@linux.dev Signed-off-by: Oliver Upton Signed-off-by: Sasha Levin commit c2230cd149507ff24eb524e8efdfa86ea7e41790 Author: SeungJu Cheon Date: Tue Aug 25 17:37:19 2026 +0900 RISC-V: KVM: Fix perf-backed counter accounting across stop and read [ Upstream commit c7e2cc38c56142cdab25e6f73602a8222bf9479b ] pmu_ctr_read() adds the event count returned by perf_event_read_value() to counter_val, which can accumulate the same count repeatedly across reads. kvm_riscv_vcpu_pmu_ctr_stop() also leaves counter_val stale by not folding the current event count into it. Make reads of perf-backed counters side-effect free, and use perf_event_pause() when stopping a counter to fold the current event count into counter_val while resetting it. This preserves the counter value across stop/start and lets the snapshot path use counter_val directly. Fixes: 0cb74b65d2e5 ("RISC-V: KVM: Implement perf support without sampling") Signed-off-by: SeungJu Cheon Reviewed-by: Anup Patel Link: https://lore.kernel.org/r/20260825083719.643970-4-suunj1331@gmail.com Signed-off-by: Anup Patel Signed-off-by: Sasha Levin commit f6f1d27e0c893d8e3db59ccdffdba87000221395 Author: SeungJu Cheon Date: Tue Aug 25 17:37:18 2026 +0900 RISC-V: KVM: Report snapshot write failure to the guest [ Upstream commit 057dd2639ceae79adced5d8fe52c32d562edcb3a ] If kvm_vcpu_write_guest() fails while updating the PMU snapshot area on counter stop, the guest may receive SBI_SUCCESS without the snapshot being updated, leaving stale data in shared memory. Return SBI_ERR_FAILURE when the snapshot write fails. Fixes: c2f41ddbcdd7 ("RISC-V: KVM: Implement SBI PMU Snapshot feature") Signed-off-by: SeungJu Cheon Reviewed-by: Anup Patel Link: https://lore.kernel.org/r/20260825083719.643970-3-suunj1331@gmail.com Signed-off-by: Anup Patel Signed-off-by: Sasha Levin commit f4363eede3e743cb2a6d6db2fecc1e41099319b0 Author: SeungJu Cheon Date: Tue Aug 25 17:37:17 2026 +0900 RISC-V: KVM: Preserve firmware counter value across stop/start [ Upstream commit 8b3fd1a8b305321171602bfa7c41212441cf69e4 ] Firmware events accumulate in kvpmu->fw_event[].value while running, but counter stop only clears fw_event[].started without saving the value back to pmc->counter_val. A subsequent counter start without SBI_PMU_START_FLAG_SET_INIT_VALUE reloads the stale counter_val into fw_event[].value, losing all events counted so far. Save fw_event[].value into counter_val when actually stopping a running counter, and remove the now redundant synchronization from the snapshot path. Fixes: badc386869e2c ("RISC-V: KVM: Support firmware events") Signed-off-by: SeungJu Cheon Reviewed-by: Anup Patel Link: https://lore.kernel.org/r/20260825083719.643970-2-suunj1331@gmail.com Signed-off-by: Anup Patel Signed-off-by: Sasha Levin commit d6de766d8d26d71b714f548192099112def50a9c Author: Zongmin Zhou Date: Wed Aug 26 15:50:09 2026 +0800 KVM: riscv: Fix NACL hfence entry update order [ Upstream commit b3d346838ec65fac7fd83f5dbcedd13cadfffddb ] The SBI v3.0 specification (section 15.1.2) requires a nested HFENCE entry to be populated as follows: 1) find an unused entry with Config.Pending == 0 2) update the Page_Number and Page_Count words 3) update the Config word with Config.Pending set __kvm_riscv_nacl_hfence() writes the Config word first, so the SBI implementation (or NACL hardware) can observe a pending entry with pnum/pcount values left over from the previous use of that entry, resulting in incorrect TLB flush ranges. Write pnum and pcount first and the Config word last. Since the consumer is an external agent on coherent shared memory, use WRITE_ONCE() to stop the compiler from reordering the stores and smp_wmb() to make the parameter words globally visible before the Pending bit is set. Fixes: d466c19cead5 ("RISC-V: KVM: Add common nested acceleration support") Signed-off-by: Zongmin Zhou Reviewed-by: Anup Patel Link: https://lore.kernel.org/r/20260826075009.68952-1-min_halo@163.com Signed-off-by: Anup Patel Signed-off-by: Sasha Levin commit 5c221f33a518333cb989558858081c59c7305891 Author: Benjamin Tissoires Date: Fri Sep 4 14:53:00 2026 +0200 HID: bpf: fix __hid_bpf_hw_check_params report length [ Upstream commit c4afa4862b878d56e0cc1021298794ac1b45bc49 ] Turns out that USB, I2C and other transport drivers (except uhid which just passes the data) still need to have the report ID in the first byte. Because they expect the first byte to be the report ID or 0, when the report ID is 0, they strip that first byte before forwarding to the device. This means that the transport layer forwards a buffer of size N-1 to the device, which gets rejected. Fixes: 5599f8019661 ("HID: bpf: export hid_hw_output_report as a BPF kfunc") Signed-off-by: Benjamin Tissoires Signed-off-by: Sasha Levin commit 5f9e2ecb725adced013fbf5f94b9f185fad2e7d9 Author: Slawomir Stepien Date: Mon Sep 14 11:50:07 2026 +0200 HID: amd_sfh: Validate PCI BAR size before mapping [ Upstream commit 65bcc5f89704efe5b9d69d4ea2c1002d90c31382 ] The amd_sfh driver maps PCI BAR 2 using pcim_iomap_regions() and subsequently accesses MMIO registers at offsets up to 0x10958 (e.g., AMD_P2C_MSG3 at 0x1068C). However, the driver never validates that the BAR size is large enough to cover these accesses. If the driver is bound to a device with a smaller BAR 2, this leads to an out-of-bounds memory access and a page fault during the probe function. For example, a page fault can occur when reading from privdata->mmio + AMD_P2C_MSG3 in mp2_select_ops(): BUG: unable to handle page fault for address: ffffc9000390368c PGD 100000067 P4D 100000067 PUD 1012c1067 PMD 105b64067 PTE 0 Oops: Oops: 0000 [#1] SMP KASAN NOPTI RIP: 0010:readl arch/x86/include/asm/io.h:59 [inline] RIP: 0010:mp2_select_ops drivers/hid/amd-sfh-hid/amd_sfh_pcie.c:282 [inline] RIP: 0010:amd_mp2_pci_probe+0x337/0x5f0 drivers/hid/amd-sfh-hid/amd_sfh_pcie.c:487 Call Trace: local_pci_probe drivers/pci/pci-driver.c:332 [inline] pci_call_probe drivers/pci/pci-driver.c:394 [inline] __pci_device_probe drivers/pci/pci-driver.c:455 [inline] pci_device_probe+0x431/0xc90 drivers/pci/pci-driver.c:489 Fix this by verifying that the length of BAR 2 is at least 128KB before attempting to map it. Since the maximum accessed offset is 0x10958, and PCI BAR sizes are powers of 2, any legitimate hardware will have a BAR size of at least 128KB. Fixes: 4f567b9f8141 ("SFH: PCIe driver to add support of AMD sensor fusion hub") Assisted-by: Gemini:gemini-3.7-flash Gemini:gemini-3.1-pro-preview syzbot Reported-by: syzbot+4eadd4dfe9e66522bae8@syzkaller.appspotmail.com Closes: https://syzkaller.appspot.com/bug?extid=4eadd4dfe9e66522bae8 Link: https://syzkaller.appspot.com/ai_job?id=3bc1c45c-548f-4ab5-8243-d2c8ec321d6c Signed-off-by: Slawomir Stepien Acked-by: Basavaraj Natikar Link: https://syzkaller.appspot.com/bug?extid=4eadd4dfe9e66522bae8 Signed-off-by: Jiri Kosina Signed-off-by: Sasha Levin commit 905d1a40f34f721d9a5936e02f1a92eb7f0d273f Author: Sean Anderson Date: Mon Aug 17 12:22:11 2026 -0400 pinctrl: meson: Fix typo in s4 group name [ Upstream commit 692f32609a30f75ca3401e25b504bfd06bd5662a ] One of the i2c pin groups has some junk at the end. The name should be i2c2_scl_h1, and indeed that's the name used by i2c2_pins3 in meson-s4.dtsi. Fixes: 775214d389c25 ("pinctrl: meson: add pinctrl driver support for Meson-S4 Soc") Signed-off-by: Sean Anderson Reviewed-by: Neil Armstrong Signed-off-by: Linus Walleij Signed-off-by: Sasha Levin commit b67b4c763f211e7a6391c356fef552e2ea63f08a Author: Donggeun Yoo Date: Mon Sep 7 22:06:23 2026 +0900 bpf, arm64: set up the frame pointer for the exception callback [ Upstream commit ef1fb82f12186dd26153b14d9fbcf4ec98db81b3 ] A program acting as exception boundary saves all callee-saved registers, so build_prologue() takes the exception_cb path and never calls push_callee_regs(). That is the only place find_used_callee_regs() runs, and with it the only place ctx->fp_used is set, so the callback prologue does not emit the mov x25, sp that points BPF_REG_FP at the frame the callback runs on. x25 keeps whatever it held when bpf_throw() was called. If the throw came from a subprogram that uses its own BPF stack, that is the subprogram's frame pointer, and since the subprogram never returns it never restores x25 either. Stack accesses through BPF_REG_FP are rewritten to be stack pointer relative, so those still land in the callback's own frame. Materializing the register does not: a callback that passes the address of a local variable to a helper hands over an address in the dead subprogram's frame. That address is below the callback's stack pointer by then, and the helper's own call chain covers it, so the helper can write over its own return address. 0x1234 below is the value the helper was asked to store: pc : 0x1234 lr : 0x1234 Call trace: 0x1234 (P) bpf_test_run+0x188/0x3e0 bpf_prog_test_run_skb+0x47c/0x998 __sys_bpf+0xbdc/0xdd8 Kernel panic - not syncing: Oops: Fatal exception in interrupt Set ctx->fp_used on the exception callback path so that the existing code further down sets x25 from the stack pointer. The epilogue restores it from the main program's save area along with the other callee-saved registers, as it already does. x86 sets the frame pointer for the callback from the argument it is passed, and powerpc computes it from the stack pointer. Fixes: 5d4fa9ec5643 ("bpf, arm64: Avoid blindly saving/restoring all callee-saved registers") Acked-by: Xu Kuohai Signed-off-by: Donggeun Yoo Link: https://lore.kernel.org/r/20260907130624.611942-2-donggeunyoo.kernel@gmail.com Signed-off-by: Alexei Starovoitov Signed-off-by: Sasha Levin commit ae1ffd04e98beefbe5d83335d506b0bcd5c2a6b3 Author: Geliang Tang Date: Tue Sep 8 17:08:32 2026 +0800 bpf, sockmap: Fix self-redirect copied_seq double-counting [ Upstream commit 490a83d6386eec1d29f470c8d7331677fb46c3b7 ] When a BPF stream_verdict program redirects an skb back to the same socket (self-redirect with BPF_F_INGRESS), sk_psock_verdict_apply() calls tcp_eat_skb() which advances tcp_sk->copied_seq. However, the skb is then delivered to the socket's psock ingress queue and later read by tcp_bpf_recvmsg_parser(), which also advances copied_seq via the copied_from_self accounting path. This double-counting causes copied_seq to advance by 2x the actual data length, triggering: TCP recvmsg seq # bug 2: copied BF2E806, seq BF2E7FD, \ rcvnxt BF2E806, fl 0 WARNING: net/ipv4/tcp.c:2745 at tcp_recvmsg_locked+0x72b/0x2640 Call Trace: tcp_recvmsg+0x10a/0x500 sock_recvmsg+0x168/0x1d0 __sys_recvfrom+0x19a/0x2a0 __x64_sys_recvfrom+0xe4/0x1f0 do_syscall_64+0xf7/0x530 entry_SYSCALL_64_after_hwframe+0x77/0x7f cleanup rbuf bug: copied BF2E806 seq BF2E806 rcvnxt BF2E806 WARNING: net/ipv4/tcp.c:1609 at tcp_cleanup_rbuf+0xf2/0x1c0 Call Trace: tcp_recvmsg_locked+0x8d1/0x2640 tcp_recvmsg+0x10a/0x500 sock_recvmsg+0x168/0x1d0 __sys_recvfrom+0x19a/0x2a0 __x64_sys_recvfrom+0xe4/0x1f0 do_syscall_64+0xf7/0x530 entry_SYSCALL_64_after_hwframe+0x77/0x7f Fix this by converting self-redirect verdict to __SK_PASS at the beginning of sk_psock_verdict_apply(). This bypasses the __SK_REDIRECT case entirely (which calls sk_psock_eat_skb), letting the __SK_PASS path queue the skb to the psock ingress queue. The data is then read via tcp_bpf_recvmsg_parser(), which advances copied_seq exactly once through copied_from_self. Cross-socket redirects continue through __SK_REDIRECT with sk_psock_eat_skb() unchanged. Fixes: e5c6de5fa025 ("bpf, sockmap: Incorrectly handling copied_seq") Suggested-by: Jakub Sitnicki Suggested-by: Jiayuan Chen Signed-off-by: Geliang Tang Reviewed-by: Emil Tsalapatis Reviewed-by: Jiayuan Chen Link: https://lore.kernel.org/r/1a8e797a1b26e2f695aaac22ac644c2862f63466.1788858299.git.tanggeliang@kylinos.cn Signed-off-by: Alexei Starovoitov Signed-off-by: Sasha Levin commit 1d6581b544a88cf4348c2fb60dcb50ad63f012da Author: Pu Lehui Date: Sat Sep 5 02:11:39 2026 +0000 bpf: Fix UAF due to concurrent consumption of ttrace lists in alloc_bulk [ Upstream commit 1c21452d02eec2f008e2c5535820f85adbd7587a ] Syzkaller repeatedly triggered UAF splats related to nodes in waiting_for_gp_ttrace within the bpf memalloc: BUG: KASAN: slab-use-after-free in llist_del_first+0x85/0x110 lib/llist.c:61 Read of size 8 at addr ffff8881572cd080 by task syz.4.470/5112 ... llist_del_first+0x85/0x110 lib/llist.c:61 alloc_bulk+0x193/0x460 kernel/bpf/memalloc.c:229 bpf_mem_refill+0x386/0x560 kernel/bpf/memalloc.c:436 Freed by task 14: ... __free_rcu kernel/bpf/memalloc.c:281 [inline] __free_rcu_tasks_trace+0x48/0xd0 kernel/bpf/memalloc.c:291 rcu_tasks_invoke_cbs+0x1ec/0x3e0 kernel/rcu/tasks.h:571 rcu_tasks_one_gp+0x13d/0x220 kernel/rcu/tasks.h:621 rcu_tasks_kthread+0xf3/0x120 kernel/rcu/tasks.h:651 The reason is that the UAF occurs after the RCU Tasks Trace GP expires: when the __free_rcu() callback runs, there is no synchronization protecting llist_del_all() against concurrent alloc_bulk() operating on waiting_for_gp_ttrace, leading to the race condition below: CPU0 CPU1 __free_rcu (RCU Tasks Trace callback) alloc_bulk llist_del_first(&c->waiting_for_gp_ttrace) entry = smp_load_acquire(&head->first); do { if (entry == NULL) return NULL; free_all(llist_del_all(&c->waiting_for_gp_ttrace)) llist_for_each_safe(pos, t, llnode) free_one(pos); next = READ_ONCE(entry->next); <-- trigger UAF } while (!try_cmpxchg(&head->first, &entry, next)); In addition, there is also a theoretical race condition on the free_by_rcu_ttrace list. This race requires two preconditions: an in-flight Tasks Trace GP keeping c->call_rcu_ttrace_in_progress == 1, and concurrent cross-CPU frees repopulating c->free_by_rcu_ttrace with new nodes. Under these conditions, the following scenario triggers UAF: // CPU0 // irq work is still busy (on PREEMPT_RT) alloc_bulk() llist_del_first(&c->free_by_rcu_ttrace) entry = smp_load_acquire(&head->first); do { if (entry == NULL) return NULL; // CPU1 bpf_mem_alloc_destroy() WRITE_ONCE(c->draining, true) // wait for CPU0 irq_work_sync() // CPU2 do_call_rcu_ttrace(tgt(CPU0)) if (c->draining) { llist_del_all(&c->free_by_rcu_ttrace) free_all() } // CPU0 continue next = READ_ONCE(entry->next); <-- trigger UAF while (!try_cmpxchg(&head->first, &entry, next)); Fix this by introducing a raw spinlock to synchronize the concurrent consumption on waiting_for_gp_ttrace and free_by_rcu_ttrace. Fixes: 04fabf00b4d3 ("bpf: Allow reuse from waiting_for_gp_ttrace list.") Suggested-by: Alexei Starovoitov Suggested-by: Hou Tao Signed-off-by: Pu Lehui Acked-by: Hou Tao Link: https://lore.kernel.org/r/20260905021139.4116529-1-pulehui@huaweicloud.com Signed-off-by: Alexei Starovoitov Signed-off-by: Sasha Levin commit 51a17f0ce19c4ee2a25cf16c93490d509b1ebb0a Author: Kumar Kartikeya Dwivedi Date: Fri Feb 27 14:48:01 2026 -0800 bpf: Register dtor for freeing special fields [ Upstream commit 1df97a7453eec80c1912c2d0360290a3970a7671 ] There is a race window where BPF hash map elements can leak special fields if the program with access to the map value recreates these special fields between the check_and_free_fields done on the map value and its eventual return to the memory allocator. Several ways were explored prior to this patch, most notably [0] tried to use a poison value to reject attempts to recreate special fields for map values that have been logically deleted but still accessible to BPF programs (either while sitting in the free list or when reused). While this approach works well for task work, timers, wq, etc., it is harder to apply the idea to kptrs, which have a similar race and failure mode. Instead, we change bpf_mem_alloc to allow registering destructor for allocated elements, such that when they are returned to the allocator, any special fields created while they were accessible to programs in the mean time will be freed. If these values get reused, we do not free the fields again before handing the element back. The special fields thus may remain initialized while the map value sits in a free list. When bpf_mem_alloc is retired in the future, a similar concept can be introduced to kmalloc_nolock-backed kmem_cache, paired with the existing idea of a constructor. Note that the destructor registration happens in map_check_btf, after the BTF record is populated and (at that point) avaiable for inspection and duplication. Duplication is necessary since the freeing of embedded bpf_mem_alloc can be decoupled from actual map lifetime due to logic introduced to reduce the cost of rcu_barrier()s in mem alloc free path in 9f2c6e96c65e ("bpf: Optimize rcu_barrier usage between hash map and bpf_mem_alloc."). As such, once all callbacks are done, we must also free the duplicated record. To remove dependency on the bpf_map itself, also stash the key size of the map to obtain value from htab_elem long after the map is gone. [0]: https://lore.kernel.org/bpf/20260216131341.1285427-1-mykyta.yatsenko5@gmail.com Fixes: 14a324f6a67e ("bpf: Wire up freeing of referenced kptr") Fixes: 1bfbc267ec91 ("bpf: Enable bpf_timer and bpf_wq in any context") Reported-by: Alexei Starovoitov Tested-by: syzbot@syzkaller.appspotmail.com Signed-off-by: Kumar Kartikeya Dwivedi Link: https://lore.kernel.org/r/20260227224806.646888-2-memxor@gmail.com Signed-off-by: Alexei Starovoitov Stable-dep-of: 1c21452d02ee ("bpf: Fix UAF due to concurrent consumption of ttrace lists in alloc_bulk") Signed-off-by: Sasha Levin commit 6785313d4fac96afb21ef193610c1fa88b1a9e64 Author: Jiayuan Chen Date: Thu Sep 3 18:09:20 2026 +0800 bpf: Fix out-of-bounds read of rtt_min in sock_ops [ Upstream commit 75f8cf22463d82bb1fb0239a3d485fc8f4c8ef03 ] A sockops prog reading skops->rtt_min never checks the sk type: on the tcp_conn_request() path sock_ops->sk is a request_sock (non-full), and the ctx rewrite casts it to a tcp_sock (full) and reads rtt_min past the end of the request_sock, returning dirty adjacent memory. SEC("sockops") int prog(struct bpf_sock_ops *skops) { switch (skops->op) { case BPF_SOCK_OPS_RWND_INIT: leak = skops->rtt_min; /* reads the request_sock OOB */ ... } } For instance one such read returned rtt_min=0xffff8881, the high half of a leaked kernel pointer. Guarding that cast is exactly what SOCK_OPS_GET_FIELD() does -- it checks is_locked_tcp_sock and returns 0 when sock_ops->sk is not a locked full socket. Every other tcp_sock field in sock_ops goes through it; rtt_min is the only one open-coded, so it skips the check. Read rtt_min through SOCK_OPS_GET_FIELD() too. rtt_min is a bit special: it is a struct minmax and we only want the current min, so pass rtt_min.s[0].v. That is equivalent to the old hand-computed offset offsetof(struct tcp_sock, rtt_min) + sizeof_field(struct minmax_sample, t) (s[0] sits at rtt_min + 0 and .v at + sizeof(.t), i.e. what minmax_get() returns), so the loaded field is unchanged and only the full-sock guard is added. The two BUILD_BUG_ON()s that protected the hand-computed offset are no longer needed. Before patch: 0: r1 = *(u64 *)(r1 +0) ; r1 = skops->sk 1: r1 = *(u32 *)(r1 +2324) ; ((tcp_sock *)sk)->rtt_min.s[0].v After patch: 0: *(u64 *)(r1 +56) = r9 1: r9 = *(u8 *)(r1 +50) ; is_locked_tcp_sock 2: if r9 == 0 goto pc+4 ; not a locked full sock -> 0 3: r9 = *(u64 *)(r1 +56) 4: r1 = *(u64 *)(r1 +0) ; r1 = skops->sk 5: r1 = *(u32 *)(r1 +2324) ; rtt_min.s[0].v 6: goto pc+2 7: r9 = *(u64 *)(r1 +56) 8: r1 = 0 Fixes: 44f0e43037d3 ("bpf: Add support for reading sk_state and more") Reported-by: VEGA Signed-off-by: Jiayuan Chen Reviewed-by: Emil Tsalapatis Link: https://lore.kernel.org/r/20260903100921.113374-1-jiayuan.chen@linux.dev Signed-off-by: Alexei Starovoitov Signed-off-by: Sasha Levin commit 44f9bcb5ab2b268bc763ed3ebec83308b4972936 Author: Jose Fernandez (Anthropic) Date: Wed Sep 9 17:51:04 2026 +0000 bpf: Avoid soft lockup in __htab_map_lookup_and_delete_batch() [ Upstream commit 85136bf22404474a815fc0ed26ec0d1cbc1bc3f9 ] __htab_map_lookup_and_delete_batch() has no rescheduling point. The batch count bounds how many entries are copied out, not how many buckets are visited, so one BPF_MAP_LOOKUP_BATCH call can walk the map end to end. The empty-bucket fast path is worse: it stays inside a single rcu_read_lock() / bpf_disable_instrumentation() section for any run of consecutive empty buckets. That holds up on small maps, but it falls apart at scale. On a 144-CPU arm64 host running a CONFIG_PREEMPT_NONE kernel, periodic BPF_MAP_LOOKUP_BATCH calls against an LRU hash map with 16,777,216 buckets held a CPU inside the batch op for 77+ seconds and triggered the soft lockup watchdog. Commit 75134f16e7dd ("bpf: Add schedule points in batch ops") fixed this same problem in the generic batch ops, but not in this htab-native path, which every htab-based hash map variant uses for its lookup[_and_delete] batch ops. Complete that fix here. Leave the critical section after 64 consecutive empty buckets, call cond_resched_tasks_rcu_qs(), and resume at the saved bucket cursor. No locks are held at that point, and resuming from the cursor is already the function's behavior for non-empty buckets. Add the same call to the per-bucket loop after copy_to_user(), where every lock has been dropped. cond_resched_rcu() is not enough here: sleeping with bpf_prog_active elevated makes tracing programs on that CPU silently skip their invocations. Plain cond_resched() is not enough either. It is a no-op under PREEMPT and PREEMPT_LAZY, the only models arm64 and x86 have offered since commit 7dadeaa6e851 ("sched: Further restrict the preemption modes"). It is also never a Tasks RCU quiescent state, in any model: the reschedule counts as a preemption. The walking task stays a holdout and stalls every synchronize_rcu_tasks() caller, ftrace and BPF trampoline teardown included, until the syscall returns [1]. cond_resched_tasks_rcu_qs() is the usual tool for that [2]. It reports the quiescent state at each yield and still reschedules as cond_resched() does on PREEMPT_NONE and PREEMPT_VOLUNTARY kernels. Fixes: 057996380a42 ("bpf: Add batch ops to all htab bpf map") Cc: "Paul E. McKenney" Cc: Rik van Riel Link: https://lore.kernel.org/bpf/20260715215314.44423f47@fangorn/ [1] Link: https://lore.kernel.org/bpf/9d444098-7c03-4163-af12-bd0a79a51443@paulmck-laptop/ [2] Assisted-by: LLM Signed-off-by: Jose Fernandez (Anthropic) Signed-off-by: Josef Bacik Reviewed-by: Rik van Riel Link: https://lore.kernel.org/r/20260909-b4-htab-batch-resched-v2-1-0cb529d8f95a@toxicpanda.com Signed-off-by: Alexei Starovoitov Signed-off-by: Sasha Levin commit 5ad116c23f14963442c0162aa8f5bf3419b39825 Author: Jim Mattson Date: Wed Sep 2 11:47:11 2026 -0700 KVM: x86/pmu: Move Intel PMU global MSRs to intel_is_valid_msr() [ Upstream commit 79a71cc2568f4b5d42284da2aa26f3b4f47ce01b ] Commit c85cdc1cc1ea ("KVM: x86/pmu: Move handling PERF_GLOBAL_CTRL and friends to common x86") moved the existence check for the following Intel PMU MSRs to kvm_pmu_is_valid_msr(): - MSR_CORE_PERF_GLOBAL_STATUS - MSR_CORE_PERF_GLOBAL_CTRL - MSR_CORE_PERF_GLOBAL_OVF_CTRL That commit deemed these MSRs valid whenever pmu->version > 1. It intended to share the check with AMD PerfMonV2 because both vendor implementations require version 2 or greater for global PMU controls. However, as noted in the commit message, AMD uses different MSR indices for its global PMU registers. Commit 4a2771895ca6 ("KVM: x86/svm/pmu: Add AMD PerfMonV2 support") subsequently added AMD PerfMonV2 support and set pmu->version = 2. Because kvm_pmu_is_valid_msr() validated the Intel MSRs whenever pmu->version > 1, KVM incorrectly permitted AMD guests with PerfMonV2 to access these Intel MSRs without a #GP. Move the validation of these Intel MSRs to intel_is_valid_msr() and remove the common switch statement from kvm_pmu_is_valid_msr(). AMD already validates its own global PMU MSRs in amd_is_valid_msr(). Fixes: 4a2771895ca6 ("KVM: x86/svm/pmu: Add AMD PerfMonV2 support") Signed-off-by: Jim Mattson Reviewed-by: Like Xu Reviewed-by: Sandipan Das Link: https://patch.msgid.link/20260902184711.138538-1-jmattson@google.com Signed-off-by: Sean Christopherson Signed-off-by: Sasha Levin commit 08a4d55b26724cc2eecedf9a5c6323b6bb7e7feb Author: Oscar Priego Verdugo Date: Mon Aug 17 06:05:46 2026 -0600 HID: elecom: fix bus type for M-XGL20DLBK [ Upstream commit 8e2a4b458ad25e13422bb059758c30a6562aa9cf ] The M-XGL20DLBK is matched as a USB device by hid-elecom, but its entry in hid_have_special_driver[] uses HID_BLUETOOTH_DEVICE. This prevents the special-driver quirk entry from matching the USB device handled by hid-elecom. Use HID_USB_DEVICE there as well. Fixes: 55633e681afb ("HID: elecom: add support for EX-G M-XGL20DLBK wireless mouse") Signed-off-by: Oscar Priego Verdugo Signed-off-by: Jiri Kosina Signed-off-by: Sasha Levin commit 49688a0d1ac690d67a560e81df03f1e4a84bd098 Author: Lovekesh Solanki Date: Wed Aug 5 01:50:31 2026 +0530 HID: multitouch: Add report ID mismatch quirk for ASUS ROG Z13 Folio [ Upstream commit aaaea79efba5a27cb9e0a5a628046d829d3f2cbb ] Commit e716edafedad ("HID: multitouch: Check to ensure report responses match the request") introduced validating GET_FEATURE responses return the requested report ID. ASUS ROG Z13 Flow (2025) GZ302EA touchpad (USB 0b05:1a30) returns a different report ID for Win8 feature request. Before this check, the response was still processed and allowed device to switch into its full Touchpad Precision mode. After the validation, the response is discarded before hid_report_raw_event() processes it and device remains in fallback mode and no longer exposes ABS_MT_SLOT, ABS_MT_TOOL_TYPE or the multi-finger BTN_TOOL_* capabilities for palm rejection. Add a device quirk to allow the known firmware behavior for this device while preserving report ID validation for all other devices. The device previously matched the generic MT_CLS_WIN_8 entry, so base the new class on MT_CLS_WIN_8 to keep it on the same quirk set as before the regression. MT_QUIRK_CONFIDENCE must be set explicitly: it is normally enabled by the class name check in mt_touch_input_mapping(), which only matches the MT_CLS_WIN_8* names, and it is what makes ABS_MT_TOOL_TYPE available for touchpads. Fixes: e716edafedad ("HID: multitouch: Check to ensure report responses match the request") Signed-off-by: Lovekesh Solanki Reported-by: mayhemandcoffee Closes: https://bugzilla.kernel.org/show_bug.cgi?id=221774 Tested-by: mayhemandcoffee Link: https://bugzilla.kernel.org/show_bug.cgi?id=221774 Signed-off-by: Jiri Kosina Signed-off-by: Sasha Levin commit 55bbebf2829c3f47c5a48fc261740301d0f25f48 Author: Jiayuan Chen Date: Thu Sep 10 19:27:28 2026 +0800 tcp: Skip cond_resched() in inet_csk_listen_stop() under BPF context [ Upstream commit eaab8cab451b9502ce224cd202550375b894a467 ] bpf_sock_destroy() runs from the tcp iterator, under rcu_read_lock(). If the sock is a listener that still has children in its accept queue, tcp_abort() ends up in inet_csk_listen_stop() and the cond_resched() there trips the debug check: BUG: sleeping function called from invalid context at net/ipv4/inet_connection_sock.c:1523 in_atomic(): 0, irqs_disabled(): 0, non_block: 0, pid: 628, name: test_progs preempt_count: 0, expected: 0 RCU nest depth: 1, expected: 0 locks held by test_progs/628: 3, last CPU#3: #0: ffff8881158cee18 (&p->lock){+.+.}-{4:4}, at: bpf_seq_read+0x56/0x1210 #1: ffff8881106bb858 (sk_lock-AF_INET6){+.+.}-{0:0}, at: bpf_iter_tcp_seq_show+0x32b/0x4b0 #2: ffffffffb435af20 (rcu_read_lock){....}-{1:3}, at: bpf_iter_run_prog+0x46b/0xde0 CPU: 3 UID: 0 PID: 628 Comm: test_progs Tainted: G W 7.2.0+ #65 PREEMPT Tainted: [W]=WARN Call Trace: dump_stack_lvl+0xc1/0xf0 dump_stack+0x10/0x20 __might_resched+0x3d2/0x610 inet_csk_listen_stop+0x7b/0xbf0 tcp_abort+0x23b/0x3b0 bpf_sock_destroy+0xfc/0x140 bpf_prog_448133d24601754f_iter_tcp6_server+0x81/0x8a bpf_iter_run_prog+0x538/0xde0 bpf_iter_tcp_seq_show+0x26b/0x4b0 bpf_seq_read+0x424/0x1210 vfs_read+0x197/0xe40 ksys_read+0x119/0x240 __x64_sys_read+0x72/0xc0 x64_sys_call+0x647/0x27e0 do_syscall_64+0xe5/0x610 entry_SYSCALL_64_after_hwframe+0x76/0x7e RIP: 0033:0x7fad39b28aca RSP: 002b:00007ffc381c61c0 EFLAGS: 00000246 ORIG_RAX: 0000000000000000 RAX: ffffffffffffffda RBX: 00007ffc381c6a88 RCX: 00007fad39b28aca RDX: 0000000000000032 RSI: 00007ffc381c6250 RDI: 0000000000000014 RBP: 00007ffc381c61e0 R08: 0000000000000000 R09: 0000000000000000 R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000003 R13: 0000000000000000 R14: 000055f077c1bbb0 R15: 00007fad3a0f3000 The commit that added the kfunc already guards lock_sock() in tcp_abort() and udp_abort() with has_current_bpf_ctx(), but missed the listener path. Do the same for the cond_resched(). The loop runs inside the iterator's rcu_read_lock(), it must not reschedule or report a quiescent state there. Fixes: 4ddbcb886268 ("bpf: Add bpf_sock_destroy kfunc") Signed-off-by: Jiayuan Chen Link: https://lore.kernel.org/r/20260910112736.153710-1-jiayuan.chen@linux.dev Signed-off-by: Alexei Starovoitov Signed-off-by: Sasha Levin commit 91a482a81eabddebf5430cbbdf12027e1a91b767 Author: Jiayuan Chen Date: Thu Sep 10 19:26:26 2026 +0800 bpf: Fix out-of-bounds read of sk_protocol in bpf_sock_destroy() [ Upstream commit 01b245ba016d44861690594e10f67e026ce8552f ] sk_protocol lives in struct sock, not in struct sock_common. A timewait or request sock handed to bpf_sock_destroy() by the tcp iterator is neither, so reading sk->sk_protocol runs past the object: ================================================================== BUG: KASAN: slab-out-of-bounds in bpf_sock_destroy+0xc7/0xe0 Read of size 2 at addr ffff8881047d11b4 by task test_progs/428 Tainted: [W]=WARN Call Trace: dump_stack_lvl+0x91/0xf0 print_report+0xd1/0x630 kasan_report+0xf3/0x130 __asan_report_load2_noabort+0x14/0x30 bpf_sock_destroy+0xc7/0xe0 bpf_prog_c3dd61f9d9cd9f37_iter_tcp6_timewait+0x9f/0xb7 bpf_iter_run_prog+0x538/0xde0 bpf_iter_tcp_seq_show+0x26b/0x4b0 bpf_seq_read+0x424/0x1210 vfs_read+0x197/0xe40 ksys_read+0x119/0x240 __x64_sys_read+0x72/0xc0 x64_sys_call+0x647/0x27e0 do_syscall_64+0xe5/0x610 entry_SYSCALL_64_after_hwframe+0x76/0x7e Only check sk_protocol on full socks. tcp_abort() already knows how to deal with TIME_WAIT and NEW_SYN_RECV socks. Also fix the comment, it never matched the code. Fixes: 4ddbcb886268 ("bpf: Add bpf_sock_destroy kfunc") Reported-by: Xiang Mei (Microsoft) Closes: https://lore.kernel.org/bpf/20260702224519.800135-1-xmei5@asu.edu/ Signed-off-by: Jiayuan Chen Reviewed-by: Kuniyuki Iwashima Link: https://lore.kernel.org/r/20260910112634.152195-1-jiayuan.chen@linux.dev Signed-off-by: Alexei Starovoitov Signed-off-by: Sasha Levin commit 8eb4cd72daad3a9ee63d2281bbdd16ce6d9406b1 Author: Jiayuan Chen Date: Thu Sep 10 20:22:55 2026 +0800 bpf: Fix divide-by-zero in btf_struct_walk() [ Upstream commit b0b3dc66529676228cb938cbcad66920f735c223 ] When an access goes past the struct and the last member is a flexible array, btf_struct_walk() folds the offset back into a single element with (off - moff) % t->size, but never checks that the element type has a size. BTF takes an empty struct, so this in program BTF /* event could be empty */ struct event { #ifdef HAVE_TIMESTAMP __u64 ts; #endif }; struct batch { int nr; struct event events[]; }; divides by zero at prog load time. Getting there needs a PTR_TO_BTF_ID that is not MEM_ALLOC, e.g. a plain read of a local kptr stashed in a map from a sleepable program. Oops: divide error: 0000 [#1] SMP KASAN PTI RIP: 0010:btf_struct_walk+0x53f/0x1570 Call Trace: btf_struct_access+0x42a/0xcd0 check_ptr_to_btf_access+0x4dc/0x1160 check_mem_access+0x3a45/0x8740 check_load_mem+0x36a/0xd10 do_check_common+0x3ef0/0xb210 bpf_check+0x6d3b/0x8580 bpf_prog_load+0xf7c/0x2720 __sys_bpf+0xa83/0x3690 __x64_sys_bpf+0xc7/0x150 x64_sys_call+0x1f3f/0x27e0 do_syscall_64+0xe5/0x610 entry_SYSCALL_64_after_hwframe+0x76/0x7e Reject a zero-sized element type. The fixed array path in the same function already bails out on the same thing: btf_struct_walk() ... /* skip empty array */ if (moff == mtrue_end) continue; msize /= total_nelems; Fixes: 9c5f8a1008a1 ("bpf: Support variable length array in tracing programs") Signed-off-by: Jiayuan Chen Acked-by: Eduard Zingerman Link: https://lore.kernel.org/r/20260910122316.186384-1-jiayuan.chen@linux.dev Signed-off-by: Alexei Starovoitov Signed-off-by: Sasha Levin commit 5334c609bf927a189189c3ad9b795e6a33e064cd Author: Sven Schnelle Date: Wed Sep 9 11:29:53 2026 +0200 selftests/ftrace: Fix unique symbol check in kprobe_non_uniq_symbol.tc [ Upstream commit d22c3e0088e85be8131f7a9283f759cdbb20726d ] The current regex also matches symbols in modules, which makes the test fail on s390 where name_show is present only once in the kernel, but also multiple times in modules: 000001b1401cdc20 t name_show 000001b0c05e6c40 t name_show [mdev] 000001b0c0495f30 t name_show [i2c_core] Fix this by changing the regular expression to only match the function name. Link: https://lore.kernel.org/all/20260909092954.2200558-1-svens@linux.ibm.com/ Fixes: 03b80ff8023a ("selftests/ftrace: Add new test case which checks non unique symbol") Signed-off-by: Sven Schnelle Reviewed-by: Steven Rostedt Signed-off-by: Masami Hiramatsu (Google) Signed-off-by: Sasha Levin commit 37b18688cd18fd86d2f35c212df1611862a26cd5 Author: Weiming Shi Date: Wed Sep 9 12:08:08 2026 +0800 bpf: Disallow bpf_skb_pull_data() for LWT_SEG6LOCAL [ Upstream commit e4a62833adff6ef0fe7c0b90393204fe3c26b5c5 ] An LWT_SEG6LOCAL program can invalidate its cached SRH with bpf_lwt_seg6_adjust_srh() and then call bpf_skb_pull_data(). The latter may reallocate skb->head, leaving the per-CPU SRH pointer dangling. Post-program SRH validation then writes through that pointer. Disallow bpf_skb_pull_data() for LWT_SEG6LOCAL programs so the verifier rejects this unsafe helper combination. Other LWT program types continue to expose the helper through lwt_out_func_proto(). Fixes: 004d4b274e2a ("ipv6: sr: Add seg6local action End.BPF") Reported-by: co+adfca3e91be95776@bugs.sh Suggested-by: Alexei Starovoitov Signed-off-by: Weiming Shi Signed-off-by: Daniel Borkmann Reviewed-by: Emil Tsalapatis Closes: https://lore.kernel.org/all/GCy0KRM2IcQGoJQTjJEU9D0maBxXzEDHuQpq@bugs.sh/ Link: https://lore.kernel.org/bpf/DL9COXZQXX4V.1FN45QO2Q77ZH@gmail.com/ Link: https://lore.kernel.org/bpf/20260909040807.3885815-2-bestswngs@gmail.com Signed-off-by: Sasha Levin commit 69fe51a3bd78d6b3f1fcc16f5a9565c8d5fce23f Author: Daniel Borkmann Date: Mon Sep 7 14:10:24 2026 +0200 bpf: Fix bpf_skb_change_tail wrt csum partial skbs [ Upstream commit 3b55f350c68a0aceff108f47f9d31f47ebffaf7b ] Cilium generates ICMP "frag needed" replies from BPF when a LB DSR packet exceeds the egress MTU. The reply is built by first trimming the packet down to target size via bpf_skb_change_tail(), and then pushing the ICMP error headers in front of it. The trim is rejected for skbs which carry a checksum offload, e.g. TCP packets aggregated by GRO on ingress where tcp_gro_complete() leaves the skb as CHECKSUM_PARTIAL. __bpf_skb_min_len() raises the minimum length to the end of the L4 checksum field, so a trim to 42 bytes bails out with -EINVAL given a min_len of 52 in this case, and due to that the ICMP generator fails. This is not the case if GRO is turned off. Fix this bpf_skb_change_tail() restriction and drop the checksum offload when the new length no longer covers the checksum field. The BPF program rewrites the skb into an ICMP error and computes the checksum itself anyway. Fixes: 5293efe62df8 ("bpf: add bpf_skb_change_tail helper") Reported-by: Tom Hadlaw Reported-by: Yusuke Suzuki Signed-off-by: Daniel Borkmann Link: https://lore.kernel.org/r/20260907121025.1923656-1-daniel@iogearbox.net Signed-off-by: Alexei Starovoitov Signed-off-by: Sasha Levin commit 04bb5b2eb9f1d56110f14d3a4e31fa2c65fb026f Author: Steffen Eiden Date: Wed Aug 12 17:55:20 2026 +0200 s390/uv: Prevent potential out-of-bounds read [ Upstream commit f47190b08b71e8482072978373ee88cb2dfbdaf4 ] When the system has more than 85 secrets, the uv_secret_list struct array only holds up to 85 items per page, resulting in an out of bounds read in find_secret_in_page if the targeted secret is in the next page or not stored at all. Fix this by looping over the number of stored secrets which is the per sub-list count of stored secrets and not the overall count. Fixes: 7c9137af2042 ("s390/uv: Retrieve UV secrets support") Signed-off-by: Steffen Eiden Reviewed-by: Christoph Schlameuss Signed-off-by: Claudio Imbrenda Message-ID: <20260812-uv_secrets_fix-v3-2-a85bd29e0666@linux.ibm.com> Signed-off-by: Sasha Levin commit 018e137ac1da29f252b9b3aa8cf2b602f1402b14 Author: Steffen Eiden Date: Wed Aug 12 17:55:19 2026 +0200 s390/uv: Fix loop condition in uv_find_secrets [ Upstream commit d12ce6bce5ec5175c3581e01c71e7a5abb286d9b ] Systems with more than 85 UV secrets got -ENOENT for any secret past the first page. Fix this by setting the start index at the beginning of the loop in uv_find_secret() and not at the end. First test if there are more secrets left by comparing start_idx with list->next_secret_idx, and then set the start index to the next secret index. Fixes: 7c9137af2042 ("s390/uv: Retrieve UV secrets support") Acked-by: Claudio Imbrenda Reviewed-by: Christoph Schlameuss Signed-off-by: Steffen Eiden Signed-off-by: Claudio Imbrenda Message-ID: <20260812-uv_secrets_fix-v3-1-a85bd29e0666@linux.ibm.com> Signed-off-by: Sasha Levin commit a379ac504fcb8c6fc2aa36203f5489b9581eb2e0 Author: Claudio Imbrenda Date: Wed Aug 12 12:44:33 2026 +0200 KVM: s390: Fix IRQ injection with SIGP Stop and Store Status [ Upstream commit d343407b728a80b74be3c24b59f15e60289ea527 ] When __inject_sigp_stop() is called for a Stop and Store Status operation, if the vCPU is running, the interrupt is marked as pending and the status is stored by the thread performing the KVM_RUN IOCTL. If the vCPU is already stopped, the status is stored immediately. Storing the status means writing into userspace, which might fault, and __inject_sigp_stop() is called from do_inject_vcpu() which in turn is always called holding a spinlock, which is obviously an issue. Fix this by returning -EWOULDBLOCK from __inject_sigp_stop(), and adding a bool flag to indicate whether a store status is needed. The callers of do_inject_vcpu() are modified to pass the pointer to the bool flag; whenever a Store Status operation is needed, the callers can now perform it outside the spinlock. Opportunistically refactor kvm_s390_set_irq_state() to use scoped_guard() and __free(). Fixes: 6cddd432e3da ("KVM: s390: handle stop irqs without action_bits") Signed-off-by: Claudio Imbrenda [ Added Fixes tag while picking -- Claudio ] Message-ID: <20260812104436.109741-7-imbrenda@linux.ibm.com> Signed-off-by: Sasha Levin commit 7c7a6ba6128f2008cc44995d0e9b6ebf4b0f8ff2 Author: Mostafa Saleh Date: Thu Aug 27 20:30:55 2026 +0000 remoteproc: qcom_q6v5_adsp: Fix iommu_unmap() usage [ Upstream commit 0d8e2195bce6f08c1c53c5ef4d7347fe46418101 ] During adsp_map_carveout, the IOVA is computed by combining the physical address and the SID: iova = adsp->mem_phys | (sid << 32); However, adsp_unmap_carveout() uses the physical address and not the IOVA in iommu_unmap(), causing the unmap to fail or leak mappings because the address doesn't match the original IOVA. Cache the constructed IOVA within the qcom_adsp device struct during mapping and use it during unmapping. Fixes: f22eedff28af ("remoteproc: qcom: Add support for memory sandbox") Signed-off-by: Mostafa Saleh Link: https://lore.kernel.org/r/20260827203055.640116-1-smostafa@google.com Signed-off-by: Bjorn Andersson Signed-off-by: Sasha Levin commit 50d844aead62c47fb21f888834444f45aebf0b87 Author: Sean Rhodes Date: Fri Jul 31 22:13:09 2026 +0100 ALSA: hda/realtek: Add StarFighter HDA SSID [ Upstream commit cd401c70df472d3eddd0b6b055726a03c212181a ] Support the new StarFighter HDA SSID while keeping the existing SSID chained to the same quirk until the new match reaches backports. Signed-off-by: Sean Rhodes Signed-off-by: Takashi Iwai Link: https://patch.msgid.link/06865eaedf3de8dff199e9aa7e86cd135572f20f.1785532385.git.sean@starlabs.systems Signed-off-by: Sasha Levin commit 56b321df2863d27a4f9996fe2fc52cf316dbff23 Author: Sean Rhodes Date: Fri Jul 31 22:13:08 2026 +0100 ALSA: hda/realtek: Limit Star Labs internal mic boost [ Upstream commit 186d4adbb40138e7cb7cffc87a81e95630ced123 ] The 30 dB internal mic boost is too high for laptops, especially with fans. Limit Star Labs internal mic boost to 10 dB. Signed-off-by: Sean Rhodes Signed-off-by: Takashi Iwai Link: https://patch.msgid.link/be87292613b24150d6321adac102b4b25d00e9e6.1785532385.git.sean@starlabs.systems Signed-off-by: Sasha Levin commit 2945d3eb73d614c974a653a79e9a7a7e909db5ef Author: Leo Li Date: Wed Sep 23 11:43:30 2026 -0500 drm/amd/display: Atomize IRQ register read/modify/write ops [ Upstream commit 63e19ef3ddab806c472748c825f4dc88dcd994e8 ] [Why] The OTG_GLOBAL_SYNC_STATUS register controls various HW IRQ sources for the output timing generator (OTG). VUPDATE_NO_LOCK is one of them. To enable the IRQ, driver sets the VUPDATE_NO_LOCK_EN bit in the GLOBAL_SYNC_STATUS register. To ack the IRQ after it fires, the driver sets the VUPDATE_NO_LOCK_CLEAR bit in the same GLOBAL_SYNC_STATUS register. The bit sets are done through read/modify/write operations, which are not atomic. Thus, the following race is possible: Thread A: IRQ handler: *HW IRQ fires* # IRQ disable val = read(GLOBAL_SYNC_STATUS) unset(val, VUPDATE_NO_LOCK_EN) write(val, GLOBAL_SYNC_STATUS) # ACK reads VUPDATE_NO_LOCK_EN unset val1 = read(GLOBAL_SYNC_STATUS) set(val1, VUPDATE_NO_LOCK_CLEAR) # IRQ enable val = read(GLOBAL_SYNC_STATUS) set(val, VUPDATE_NO_LOCK_EN) write(val, GLOBAL_SYNC_STATUS) # BAD! clears VUPDATE_NO_LOCK_EN write(val1, GLOBAL_SYNC_STATUS) Regarding the tagged Fixes: change, it appears the change made this race more likely to occur. Since VUPDATE_NO_LOCK is now the sole IRQ source for vblank handling, a single race on high refresh panels can lead to a time out. [How] The GLOBAL_SYNC_STATUS register is only one example, other IRQ control registers also share the same scheme. On top of GLOBAL_SYNC_STATUS, let's clean up those as well. To keep things simple, Let's atomize the IRQ rmw ops via a single driver-wide spinlock. Due to the small scope of this lock, it is unlikely to cause noticeable overhead on top of all the existing locking within the IRQ set/handle paths. Since DM is responsible for locking, wrap dc_interrupt_set/ack with the spinlock in the new amdgpu_dm_irq_set/ack functions. Migrate/drop all references in DM to dc_interrupt_set/ack to use amdgpu_dm_irq_set/ack instead. Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5616 Fixes: c87e6635d2db ("drm/amd/display: consolidate DCN vblank/flip handling onto vupdate_no_lock") Reviewed-by: Mario Limonciello Signed-off-by: Leo Li Signed-off-by: Chenyu Chen Tested-by: Daniel Wheeler Signed-off-by: Alex Deucher (cherry picked from commit 70de0a0216583a53c946155f8c8adedfdca6b4e7) Cc: stable@vger.kernel.org (cherry picked from commit 63e19ef3ddab806c472748c825f4dc88dcd994e8) Modified for unit tests not present in 7.2.y Signed-off-by: Mario Limonciello Signed-off-by: Sasha Levin commit d1079deec828d774c7f415dc90c9dab73d545401 Author: Wenwu Hou Date: Wed Sep 23 16:48:23 2026 +0800 erofs: fix large folio race in erofs_fscache_req_complete This patch is for stable only. Commit c37460cd9b2fc ("erofs: remove fscache backend entirely") upstream removed this code. xas_for_each() iteration can race with reclamation of an unlocked large folio and splitting of its replacement shadow entry. Fix this by advancing past the entire folio before unlocking it. For example: CPU A: EROFS completion Other CPUs ---------------------------------- ----------------------------------- Find F at index 0. Mark F uptodate. Unlock F. Reclaim F. Replace indices 0–3 with a multi-index workingset shadow. Another reader inserts a smaller folio, e.g. order-0 at index 0. Split the large shadow entry: index 0: new folio index 1: shadow index 2: shadow index 3: shadow Find a shadow at index 1. folio_mark_uptodate(folio). This can cause a kernel panic such as: [1030374.432778] [ C31] BUG: unable to handle page fault for address: 00001846af017b01 [1030374.432971] [ C31] #PF: supervisor write access in kernel mode [1030374.432973] [ C31] #PF: error_code(0x0002) - not-present page [1030374.433543] [ C31] PGD 5a44f75067 P4D 5a44f75067 PUD 0 [1030374.433546] [ C31] Oops: 0002 [#1] PREEMPT SMP NOPTI [1030374.433549] [ C31] CPU: 31 PID: 2425624 Comm: node Kdump: loaded Tainted: G OE K 6.6.88-**** [1030374.434154] [ C31] Hardware name: Alibaba Cloud Alibaba Cloud ECS, BIOS ?-20260421_110423-CN.l65g09119.cloud.sqa.na131 04/01/2014 [1030374.434156] [ C31] RIP: 0010:erofs_fscache_req_complete+0xc1/0x1a0 [erofs] [1030374.434764] [ C31] Code: 17 c5 c3 48 89 c7 48 85 c0 0f 84 af 00 00 00 48 81 ff 06 04 00 00 74 1b 48 81 ff 02 04 00 00 0f 84 b9 00 00 00 66 85 ed 75 04 80 0f 08 e8 96 04 24 c3 48 8b 54 24 18 f6 c2 03 0f 95 c0 48 85 [1030374.435044] [ C31] RSP: 0000:ffffb8f2fc973cc0 EFLAGS: 00010046 [1030374.435641] [ C31] [1030374.435642] [ C31] RAX: 00001846af017b01 RBX: 000000000000000f RCX: 0000000000000001 [1030374.436885] [ C31] RDX: 000000000000000c RSI: ffffa02ab337d468 RDI: 00001846af017b01 [1030374.437127] [ C31] RBP: 0000000000000000 R08: ffffffffffffffc0 R09: 0000000000000002 [1030374.437672] [ C31] R10: 0000000000000005 R11: 0000000000000191 R12: ffffa0060ec72300 [1030374.437673] [ C31] R13: 0000000008000000 R14: ffffffffc11b1990 R15: 0000000000007000 [1030374.437677] [ C31] FS: 00007f4928cdec80(0000) GS:ffffa07dc5f80000(0000) knlGS:0000000000000000 [1030374.437678] [ C31] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 [1030374.437680] [ C31] CR2: 00001846af017b01 CR3: 00000063d5356006 CR4: 0000000000770ee0 [1030374.437681] [ C31] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000 [1030374.437682] [ C31] DR3: 0000000000000000 DR6: 00000000fffe07f0 DR7: 0000000000000400 [1030374.437684] [ C31] PKRU: 55555558 [1030374.437684] [ C31] Call Trace: [1030374.437687] [ C31] [1030374.437692] [ C31] erofs_fscache_req_put+0x27/0x40 [erofs] [1030374.438900] [ C31] cachefiles_read_complete+0x48/0x110 [cachefiles] [1030374.440448] [ C31] iomap_dio_bio_end_io+0x128/0x160 [1030374.440456] [ C31] ? __pfx_stripe_end_io+0x10/0x10 [dm_mod] [1030374.440950] [ C31] clone_endio+0x123/0x1f0 [dm_mod] [1030374.441550] [ C31] blk_mq_end_request_batch+0xf4/0x440 [1030374.441556] [ C31] ? nohz_balancer_kick+0x31/0x270 [1030374.441561] [ C31] ? dma_direct_unmap_sg+0x48/0x1d0 [1030374.441565] [ C31] ? dma_pool_free+0x22/0x60 [1030374.441569] [ C31] ? nvme_pci_complete_batch+0xaf/0xc0 [nvme] [1030374.442070] [ C31] nvme_irq+0x6e/0x80 [nvme] [1030374.442422] [ C31] ? __pfx_nvme_pci_complete_batch+0x10/0x10 [nvme] [1030374.442428] [ C31] __handle_irq_event_percpu+0x46/0x1a0 [1030374.442431] [ C31] handle_irq_event+0x37/0x80 [1030374.442433] [ C31] handle_edge_irq+0x93/0x240 [1030374.442436] [ C31] __common_interrupt+0x3b/0xa0 [1030374.442441] [ C31] common_interrupt+0x3f/0xa0 [1030374.442446] [ C31] asm_common_interrupt+0x22/0x40 Fixes: d435d53228dd ("erofs: change to use asynchronous io for fscache readpage/readahead") Signed-off-by: Wenwu Hou Reviewed-by: Gao Xiang Signed-off-by: Sasha Levin commit 14ae74db795ab88269032b40ce700cafae674290 Author: Darrick J. Wong Date: Wed Aug 26 22:31:25 2026 -0700 xfs: don't stash removename operations with unknown ftype [ Upstream commit 865b751e75039fc07838b3200f9740256653da9a ] LOLLM notices that the behavior of xrep_dir_replay_update changes based on the ftype recorded in the stashed removename information. It also notices that the unlink iops sometimes set that ftype to FT_UNKNOWN because the regular directory tree update code paths don't need to know the ftype of the child. Unfortunately, this results in incorrect link counts, which eventually trips link count errors in later phases of xfs_scrub, or in xfs_repair. Fix this by creating a second xfs_name with the type set correctly. Cc: stable@vger.kernel.org # v6.10 Fixes: 8559b21a64d983 ("xfs: implement live updates for directory repairs") Signed-off-by: Darrick J. Wong Assisted-by: LOLLM # finding obvious bugs Reviewed-by: Christoph Hellwig Signed-off-by: Carlos Maiolino Signed-off-by: Sasha Levin