Skip to content

fixed the MMC can not be partitioned - #248

Open
liunan9527 wants to merge 1 commit into
rockchip-linux:develop-4.19from
liunan9527:develop-4.19
Open

liunan9527 wants to merge 1 commit into
rockchip-linux:develop-4.19from
liunan9527:develop-4.19

Conversation

@liunan9527

Copy link
Copy Markdown

fixed the MMC can not be partitioned

scpcom pushed a commit to scpcom/linux that referenced this pull request Dec 28, 2022
…e_event_gen_test_exit()

commit 22ea4ca upstream.

When test_gen_kprobe_cmd() failed after kprobe_event_gen_cmd_end(), it
will goto delete, which will call kprobe_event_delete() and release the
corresponding resource. However, the trace_array in gen_kretprobe_test
will point to the invalid resource. Set gen_kretprobe_test to NULL
after called kprobe_event_delete() to prevent null-ptr-deref.

BUG: kernel NULL pointer dereference, address: 0000000000000070
PGD 0 P4D 0
Oops: 0000 [#1] SMP PTI
CPU: 0 PID: 246 Comm: modprobe Tainted: G        W
6.1.0-rc1-00174-g9522dc5c87da-dirty rockchip-linux#248
Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS
rel-1.15.0-0-g2dd4b9b3f840-prebuilt.qemu.org 04/01/2014
RIP: 0010:__ftrace_set_clr_event_nolock+0x53/0x1b0
Code: e8 82 26 fc ff 49 8b 1e c7 44 24 0c ea ff ff ff 49 39 de 0f 84 3c
01 00 00 c7 44 24 18 00 00 00 00 e8 61 26 fc ff 48 8b 6b 10 <44> 8b 65
70 4c 8b 6d 18 41 f7 c4 00 02 00 00 75 2f
RSP: 0018:ffffc9000159fe00 EFLAGS: 00010293
RAX: 0000000000000000 RBX: ffff88810971d268 RCX: 0000000000000000
RDX: ffff8881080be600 RSI: ffffffff811b48ff RDI: ffff88810971d058
RBP: 0000000000000000 R08: 0000000000000000 R09: 0000000000000001
R10: ffffc9000159fe58 R11: 0000000000000001 R12: ffffffffa0001064
R13: ffffffffa000106c R14: ffff88810971d238 R15: 0000000000000000
FS:  00007f89eeff6540(0000) GS:ffff88813b600000(0000)
knlGS:0000000000000000
CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 0000000000000070 CR3: 000000010599e004 CR4: 0000000000330ef0
DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000
DR3: 0000000000000000 DR6: 00000000fffe0ff0 DR7: 0000000000000400
Call Trace:
 <TASK>
 __ftrace_set_clr_event+0x3e/0x60
 trace_array_set_clr_event+0x35/0x50
 ? 0xffffffffa0000000
 kprobe_event_gen_test_exit+0xcd/0x10b [kprobe_event_gen_test]
 __x64_sys_delete_module+0x206/0x380
 ? lockdep_hardirqs_on_prepare+0xd8/0x190
 ? syscall_enter_from_user_mode+0x1c/0x50
 do_syscall_64+0x3f/0x90
 entry_SYSCALL_64_after_hwframe+0x63/0xcd
RIP: 0033:0x7f89eeb061b7

Link: https://lore.kernel.org/all/20221108015130.28326-3-shangxiaojing@huawei.com/

Fixes: 6483624 ("tracing: Add kprobe event command generation test module")
Signed-off-by: Shang XiaoJing <shangxiaojing@huawei.com>
Cc: stable@vger.kernel.org
Acked-by: Masami Hiramatsu (Google) <mhiramat@kernel.org>
Signed-off-by: Masami Hiramatsu (Google) <mhiramat@kernel.org>
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
scpcom pushed a commit to scpcom/linux that referenced this pull request Jan 31, 2023
…e_event_gen_test_exit()

commit 22ea4ca upstream.

When test_gen_kprobe_cmd() failed after kprobe_event_gen_cmd_end(), it
will goto delete, which will call kprobe_event_delete() and release the
corresponding resource. However, the trace_array in gen_kretprobe_test
will point to the invalid resource. Set gen_kretprobe_test to NULL
after called kprobe_event_delete() to prevent null-ptr-deref.

BUG: kernel NULL pointer dereference, address: 0000000000000070
PGD 0 P4D 0
Oops: 0000 [#1] SMP PTI
CPU: 0 PID: 246 Comm: modprobe Tainted: G        W
6.1.0-rc1-00174-g9522dc5c87da-dirty rockchip-linux#248
Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS
rel-1.15.0-0-g2dd4b9b3f840-prebuilt.qemu.org 04/01/2014
RIP: 0010:__ftrace_set_clr_event_nolock+0x53/0x1b0
Code: e8 82 26 fc ff 49 8b 1e c7 44 24 0c ea ff ff ff 49 39 de 0f 84 3c
01 00 00 c7 44 24 18 00 00 00 00 e8 61 26 fc ff 48 8b 6b 10 <44> 8b 65
70 4c 8b 6d 18 41 f7 c4 00 02 00 00 75 2f
RSP: 0018:ffffc9000159fe00 EFLAGS: 00010293
RAX: 0000000000000000 RBX: ffff88810971d268 RCX: 0000000000000000
RDX: ffff8881080be600 RSI: ffffffff811b48ff RDI: ffff88810971d058
RBP: 0000000000000000 R08: 0000000000000000 R09: 0000000000000001
R10: ffffc9000159fe58 R11: 0000000000000001 R12: ffffffffa0001064
R13: ffffffffa000106c R14: ffff88810971d238 R15: 0000000000000000
FS:  00007f89eeff6540(0000) GS:ffff88813b600000(0000)
knlGS:0000000000000000
CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
CR2: 0000000000000070 CR3: 000000010599e004 CR4: 0000000000330ef0
DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000
DR3: 0000000000000000 DR6: 00000000fffe0ff0 DR7: 0000000000000400
Call Trace:
 <TASK>
 __ftrace_set_clr_event+0x3e/0x60
 trace_array_set_clr_event+0x35/0x50
 ? 0xffffffffa0000000
 kprobe_event_gen_test_exit+0xcd/0x10b [kprobe_event_gen_test]
 __x64_sys_delete_module+0x206/0x380
 ? lockdep_hardirqs_on_prepare+0xd8/0x190
 ? syscall_enter_from_user_mode+0x1c/0x50
 do_syscall_64+0x3f/0x90
 entry_SYSCALL_64_after_hwframe+0x63/0xcd
RIP: 0033:0x7f89eeb061b7

Link: https://lore.kernel.org/all/20221108015130.28326-3-shangxiaojing@huawei.com/

Fixes: 6483624 ("tracing: Add kprobe event command generation test module")
Signed-off-by: Shang XiaoJing <shangxiaojing@huawei.com>
Cc: stable@vger.kernel.org
Acked-by: Masami Hiramatsu (Google) <mhiramat@kernel.org>
Signed-off-by: Masami Hiramatsu (Google) <mhiramat@kernel.org>
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
scpcom pushed a commit to scpcom/linux that referenced this pull request Sep 18, 2026
[ Upstream commit 87d5b9864c8118d26f54de4b66d2bddf2c659272 ]

The admin queue is allocated with blk_mq_alloc_queue() but never
destroyed. nvme_free_ctrl() only drops the last reference and
blk_mq_exit_queue() and blk_sync_queue() never run: the hctx is never
moved to q->unused_hctx_list and the timeout timer and work stay armed on
a queue that is about to be freed which will eventually oops inside
blk_mq_timeout_work().

This can only be triggered when the controller fails to come up and is
then immediately torn down again which is why no one ever ran into this
before.

Let's just copy what the pcie driver does: unquiesce and destroy the admin
queue before nvme_uninit_ctrl().

With this the following WARN followed by a panic no longer happens:

  WARNING: block/blk-mq.c:4390 at blk_mq_release+0x194/0x238, CPU#4: kworker/u34:4/119
  CPU: 4 UID: 0 PID: 119 Comm: kworker/u34:4 Not tainted 7.2.0-rc1-dirty rockchip-linux#248 PREEMPT
  Hardware name: Apple Mac mini (M1, 2020) (DT)
  Workqueue: nvme-wq apple_nvme_remove_dead_ctrl_work
  pstate: 61400005 (nZCv daif +PAN -UAO -TCO +DIT -SSBS BTYPE=--)
  pc : blk_mq_release+0x194/0x238
  lr : blk_mq_release+0x58/0x238
  sp : ffffc000833a3b50
  x29: ffffc000833a3b50 x28: ffff80001d0450f8 x27: ffff800020c95200
  x26: 0000000000000088 x25: 0000000000000000 x24: ffff800020f36805
  x23: 0000000000000000 x22: ffffc00081a86878 x21: ffff800020be9c60
  x20: 0000000000000000 x19: ffff800022501698 x18: 000000000000000a
  x17: 7365757165722066 x16: 666f7265776f7020 x15: 0000000000000000
  x14: 0000000000000028 x13: 0000000000004def x12: 0000000000000003
  x11: 0000000000000000 x10: 0000000000000000 x9 : ffffc000805b4fc8
  x8 : ffffc00081915820 x7 : ffffc00081c4f3c8 x6 : 0000000000000001
  x5 : 0000000000000004 x4 : ffff800022498d80 x3 : ffffc000833a3b14
  x2 : 0000000000000000 x1 : 0000000000000000 x0 : ffff800022501698
  Call trace:
   blk_mq_release+0x194/0x238 (P)
   blk_put_queue+0x8c/0xf0
   nvme_free_ctrl+0x4c/0x260
   device_release+0x44/0x128
   kobject_put+0xa0/0x120
   put_device+0x1c/0x40
   nvme_uninit_ctrl+0x48/0x60
   apple_nvme_remove+0x54/0xb0
   platform_remove+0x28/0x40
   device_remove+0x54/0x98
   device_release_driver_internal+
   device_release_driver+0x20/0x38
   apple_nvme_remove_dead_ctrl_wor
   process_one_work+0x1f4/0x770
   worker_thread+0x1b8/0x360
   kthread+0x140/0x160
   ret_from_fork+0x10/0x20
  irq event stamp: 448
  hardirqs last  enabled at (447):in_unlock_irqrestore+0x74/0x80
  hardirqs last disabled at (448): [<ffffc000811cf5c0>] el1_brk64+0x20/0x60
  softirqs last  enabled at (0): [ess+0xb28/0x2698
  softirqs last disabled at (0): [<0000000000000000>] 0x0
  ---[ end trace 0000000000000000
  Unable to handle kernel NULL pointer dereference at virtual address 0000000000000000
  Mem abort info:
    ESR = 0x0000000096000005
    EC = 0x25: DABT (current EL),
    SET = 0, FnV = 0
    EA = 0, S1PTW = 0
    FSC = 0x05: level 1 translation fault
  Data abort info:
    ISV = 0, ISS = 0x00000005, ISS2 = 0x00000000
    CM = 0, WnR = 0, TnD = 0, TagA
    GCS = 0, Overlay = 0, DirtyBit = 0, Xs = 0
  [0000000000000000] user address
  Internal error: Oops: 0000000096000005 [#1]  SMP
  CPU: 7 UID: 0 PID: 54 Comm: kwor          7.2.0-rc1-dirty #248PREEMPT
  Tainted: [W]=WARN
  Hardware name: Apple Mac mini (M1, 2020) (DT)
  Workqueue: kblockd blk_mq_timeou
  pstate: 01400005 (nzcv daif +PAN -UAO -TCO +DIT -SSBS BTYPE=--)
  pc : percpu_ref_tryget_many.cons
  lr : percpu_ref_tryget_many.constprop.0+0xc0/0x168
  sp : ffffc000829cbce0
  x29: ffffc000829cbce0 x28: ffff800020be9f48 x27: ffff800013e503c0
  x26: 0000000000000108 x25: 000009c05
  x23: 0000000000000000 x22: ffffc000819f5000 x21: ffff800020be9f48
  x20: ffff8001deda4808 x19: ffff8000a
  x17: 00000000580e1fac x16: ffffc00082bbbb7c x15: 0000000000000000
  x14: 0000000000000028 x13: 000000001
  x11: 0000000000000000 x10: 0000000000000000 x9 : ffffc000829cbc20
  x8 : ffffc00081915820 x7 : ffffc0001
  x5 : ffff80001ca77d08 x4 : 0000000000000000 x3 : ffff80001ca77cb8
  x2 : 0000000000000000 x1 : 000000007
  Call trace:
   percpu_ref_tryget_many.constpro
   blk_mq_timeout_work+0x48/0x298
   process_one_work+0x1f4/0x770
   worker_thread+0x1b8/0x360
   kthread+0x140/0x160
   ret_from_fork+0x10/0x20
  Code: 91282000 97ed44b2 17ffffd2
  ---[ end trace 0000000000000000 ]---

Fixes: 5bd2927 ("nvme-apple: Add initial Apple SoC NVMe driver")
Tested-by: Joshua Peisach <jpeisach@ubuntu.com>
Tested-by: Janne Grunau <j@jannau.net>
Tested-by: Nick Chan <towinchenmi@gmail.com>
Signed-off-by: Sven Peter <sven@kernel.org>
Signed-off-by: Sasha Levin <sashal@kernel.org>
scpcom pushed a commit to scpcom/linux that referenced this pull request Sep 18, 2026
[ Upstream commit 87d5b9864c8118d26f54de4b66d2bddf2c659272 ]

The admin queue is allocated with blk_mq_alloc_queue() but never
destroyed. nvme_free_ctrl() only drops the last reference and
blk_mq_exit_queue() and blk_sync_queue() never run: the hctx is never
moved to q->unused_hctx_list and the timeout timer and work stay armed on
a queue that is about to be freed which will eventually oops inside
blk_mq_timeout_work().

This can only be triggered when the controller fails to come up and is
then immediately torn down again which is why no one ever ran into this
before.

Let's just copy what the pcie driver does: unquiesce and destroy the admin
queue before nvme_uninit_ctrl().

With this the following WARN followed by a panic no longer happens:

  WARNING: block/blk-mq.c:4390 at blk_mq_release+0x194/0x238, CPU#4: kworker/u34:4/119
  CPU: 4 UID: 0 PID: 119 Comm: kworker/u34:4 Not tainted 7.2.0-rc1-dirty rockchip-linux#248 PREEMPT
  Hardware name: Apple Mac mini (M1, 2020) (DT)
  Workqueue: nvme-wq apple_nvme_remove_dead_ctrl_work
  pstate: 61400005 (nZCv daif +PAN -UAO -TCO +DIT -SSBS BTYPE=--)
  pc : blk_mq_release+0x194/0x238
  lr : blk_mq_release+0x58/0x238
  sp : ffffc000833a3b50
  x29: ffffc000833a3b50 x28: ffff80001d0450f8 x27: ffff800020c95200
  x26: 0000000000000088 x25: 0000000000000000 x24: ffff800020f36805
  x23: 0000000000000000 x22: ffffc00081a86878 x21: ffff800020be9c60
  x20: 0000000000000000 x19: ffff800022501698 x18: 000000000000000a
  x17: 7365757165722066 x16: 666f7265776f7020 x15: 0000000000000000
  x14: 0000000000000028 x13: 0000000000004def x12: 0000000000000003
  x11: 0000000000000000 x10: 0000000000000000 x9 : ffffc000805b4fc8
  x8 : ffffc00081915820 x7 : ffffc00081c4f3c8 x6 : 0000000000000001
  x5 : 0000000000000004 x4 : ffff800022498d80 x3 : ffffc000833a3b14
  x2 : 0000000000000000 x1 : 0000000000000000 x0 : ffff800022501698
  Call trace:
   blk_mq_release+0x194/0x238 (P)
   blk_put_queue+0x8c/0xf0
   nvme_free_ctrl+0x4c/0x260
   device_release+0x44/0x128
   kobject_put+0xa0/0x120
   put_device+0x1c/0x40
   nvme_uninit_ctrl+0x48/0x60
   apple_nvme_remove+0x54/0xb0
   platform_remove+0x28/0x40
   device_remove+0x54/0x98
   device_release_driver_internal+
   device_release_driver+0x20/0x38
   apple_nvme_remove_dead_ctrl_wor
   process_one_work+0x1f4/0x770
   worker_thread+0x1b8/0x360
   kthread+0x140/0x160
   ret_from_fork+0x10/0x20
  irq event stamp: 448
  hardirqs last  enabled at (447):in_unlock_irqrestore+0x74/0x80
  hardirqs last disabled at (448): [<ffffc000811cf5c0>] el1_brk64+0x20/0x60
  softirqs last  enabled at (0): [ess+0xb28/0x2698
  softirqs last disabled at (0): [<0000000000000000>] 0x0
  ---[ end trace 0000000000000000
  Unable to handle kernel NULL pointer dereference at virtual address 0000000000000000
  Mem abort info:
    ESR = 0x0000000096000005
    EC = 0x25: DABT (current EL),
    SET = 0, FnV = 0
    EA = 0, S1PTW = 0
    FSC = 0x05: level 1 translation fault
  Data abort info:
    ISV = 0, ISS = 0x00000005, ISS2 = 0x00000000
    CM = 0, WnR = 0, TnD = 0, TagA
    GCS = 0, Overlay = 0, DirtyBit = 0, Xs = 0
  [0000000000000000] user address
  Internal error: Oops: 0000000096000005 [#1]  SMP
  CPU: 7 UID: 0 PID: 54 Comm: kwor          7.2.0-rc1-dirty #248PREEMPT
  Tainted: [W]=WARN
  Hardware name: Apple Mac mini (M1, 2020) (DT)
  Workqueue: kblockd blk_mq_timeou
  pstate: 01400005 (nzcv daif +PAN -UAO -TCO +DIT -SSBS BTYPE=--)
  pc : percpu_ref_tryget_many.cons
  lr : percpu_ref_tryget_many.constprop.0+0xc0/0x168
  sp : ffffc000829cbce0
  x29: ffffc000829cbce0 x28: ffff800020be9f48 x27: ffff800013e503c0
  x26: 0000000000000108 x25: 000009c05
  x23: 0000000000000000 x22: ffffc000819f5000 x21: ffff800020be9f48
  x20: ffff8001deda4808 x19: ffff8000a
  x17: 00000000580e1fac x16: ffffc00082bbbb7c x15: 0000000000000000
  x14: 0000000000000028 x13: 000000001
  x11: 0000000000000000 x10: 0000000000000000 x9 : ffffc000829cbc20
  x8 : ffffc00081915820 x7 : ffffc0001
  x5 : ffff80001ca77d08 x4 : 0000000000000000 x3 : ffff80001ca77cb8
  x2 : 0000000000000000 x1 : 000000007
  Call trace:
   percpu_ref_tryget_many.constpro
   blk_mq_timeout_work+0x48/0x298
   process_one_work+0x1f4/0x770
   worker_thread+0x1b8/0x360
   kthread+0x140/0x160
   ret_from_fork+0x10/0x20
  Code: 91282000 97ed44b2 17ffffd2
  ---[ end trace 0000000000000000 ]---

Fixes: 5bd2927 ("nvme-apple: Add initial Apple SoC NVMe driver")
Tested-by: Joshua Peisach <jpeisach@ubuntu.com>
Tested-by: Janne Grunau <j@jannau.net>
Tested-by: Nick Chan <towinchenmi@gmail.com>
Signed-off-by: Sven Peter <sven@kernel.org>
Signed-off-by: Sasha Levin <sashal@kernel.org>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant