[release/2.1] Disable event subscriber during task cleanup #12410

fuweid · 2025-10-24T22:28:00Z

We have individual goroutine for each sandbox container. If there is any error in handler, that goroutine will put event in that backoff queue. So we don't need event subscriber for podsandbox. Otherwise, there will be two goroutines to cleanup sandbox container.

>>>> From EventMonitor
  time="2025-10-23T19:30:59.626254404Z" level=debug msg="Received containerd event timestamp - 2025-10-23 19:30:59.624494674 +0000 UTC, namespace - \"k8s.io\", topic - \"/tasks/exit\""
  time="2025-10-23T19:30:59.626301912Z" level=debug msg="TaskExit event in podsandbox handler container_id:\"22e15114133e4d461ab380654fb76f3e73d3e0323989c422fa17882762979ccf\" id:\"22e15114133e4d461ab380654fb76f3e73d3e0323989c422fa17882762979ccf\" pid:203121 exit_status:137 exited_at:{seconds:1761247859 nanos:624467824}"

>>> If EventMonitor handles task exit well, it will close ttrpc
connection and then waitSandboxExit could encounter ttrpc-closed error

  time="2025-10-23T19:30:59.688031150Z" level=error msg="failed to delete task" error="ttrpc: closed" id=22e15114133e4d461ab380654fb76f3e73d3e0323989c422fa17882762979ccf

If both task.Delete calls fail but the shim has already been shut down, it could trigger a new task.Exit event sent by cleanupAfterDeadShim. This would result in three events in the EventMonitor's backoff queue, which is unnecessary and could cause confusion due to duplicate events.

The worst-case scenario caused by two concurrent task.Delete calls is a shim leak. The timeline for this scenario is as follows:

Timestamp	Component	Action	Result
T1	EventMonitor	Sends `task.Delete`	Marked as Req-1
T2	waitSandboxExit	Sends `task.Delete`	Marked as Req-2
T3	containerd-shim	Handles Req-2	Container transitions from stopped to deleted
T4	containerd-shim	Handles Req-1	Fails - container already deleted Returns error: `cannot delete a deleted process: not found`
T5	EventMonitor	Receives `not found` error	-
T6	EventMonitor	Sends `shim.Shutdown` request	No-op (active container record still exists)
T7	EventMonitor	Closes ttrpc connection	Clean container state dir
T8	containerd-shim	Handles Req-2	Removes container record from memory
T9	waitSandboxExit	Receives error	Error: `ttrpc: closed`
T10	waitSandboxExit	Sends `shim.Shutdown` request	Fails (connection already closed)
T11	waitSandboxExit	Closes ttrpc connection	No-op (already closed)

The containerd-shim is still running because shim.Shutdown was sent at T6 before T8. Because container's state dir is deleted at T7, it's unable to clean it up after containerd restarted.

We should avoid concurrent task.Delete calls here.

I also add subcommand - shutdown - in ctr shim for debug.

Fixed: #12344
Cherry-picked: #12400

(cherry picked from commit 2042e80)

We have individual goroutine for each sandbox container. If there is any error in handler, that goroutine will put event in that backoff queue. So we don't need event subscriber for podsandbox. Otherwise, there will be two goroutines to cleanup sandbox container. ``` >>>> From EventMonitor time="2025-10-23T19:30:59.626254404Z" level=debug msg="Received containerd event timestamp - 2025-10-23 19:30:59.624494674 +0000 UTC, namespace - \"k8s.io\", topic - \"/tasks/exit\"" time="2025-10-23T19:30:59.626301912Z" level=debug msg="TaskExit event in podsandbox handler container_id:\"22e15114133e4d461ab380654fb76f3e73d3e0323989c422fa17882762979ccf\" id:\"22e15114133e4d461ab380654fb76f3e73d3e0323989c422fa17882762979ccf\" pid:203121 exit_status:137 exited_at:{seconds:1761247859 nanos:624467824}" >>> If EventMonitor handles task exit well, it will close ttrpc connection and then waitSandboxExit could encounter ttrpc-closed error time="2025-10-23T19:30:59.688031150Z" level=error msg="failed to delete task" error="ttrpc: closed" id=22e15114133e4d461ab380654fb76f3e73d3e0323989c422fa17882762979ccf ``` If both task.Delete calls fail but the shim has already been shut down, it could trigger a new task.Exit event sent by cleanupAfterDeadShim. This would result in three events in the EventMonitor's backoff queue, which is unnecessary and could cause confusion due to duplicate events. The worst-case scenario caused by two concurrent task.Delete calls is a shim leak. The timeline for this scenario is as follows: | Timestamp | Component | Action | Result | | ------ | ----------- | -------- | -------- | | T1 | EventMonitor | Sends `task.Delete` | Marked as Req-1 | | T2 | waitSandboxExit | Sends `task.Delete` | Marked as Req-2 | | T3 | containerd-shim | Handles Req-2 | Container transitions from stopped to deleted | | T4 | containerd-shim | Handles Req-1 | Fails - container already deleted<br>Returns error: `cannot delete a deleted process: not found` | | T5 | EventMonitor | Receives `not found` error | - | | T6 | EventMonitor | Sends `shim.Shutdown` request | No-op (active container record still exists) | | T7 | EventMonitor | Closes ttrpc connection | Clean container state dir | | T8 | containerd-shim | Handles Req-2 | Removes container record from memory | | T9 | waitSandboxExit | Receives error | Error: `ttrpc: closed` | | T10 | waitSandboxExit | Sends `shim.Shutdown` request | Fails (connection already closed) | | T11 | waitSandboxExit | Closes ttrpc connection | No-op (already closed) | The containerd-shim is still running because shim.Shutdown was sent at T6 before T8. Because container's state dir is deleted at T7, it's unable to clean it up after containerd restarted. We should avoid concurrent task.Delete calls here. I also add subcommand - shutdown - in `ctr shim` for debug. Fixed: containerd#12344 Signed-off-by: Wei Fu <[email protected]> (cherry picked from commit 2042e80) Signed-off-by: Wei Fu <[email protected]>

github-project-automation bot moved this to Needs Triage in Pull Request Review Oct 24, 2025

github-project-automation bot added this to Pull Request Review Oct 24, 2025

k8s-ci-robot added the size/L label Oct 24, 2025

dosubot bot added area/cri Container Runtime Interface (CRI) kind/bug labels Oct 24, 2025

henry118 approved these changes Oct 27, 2025

View reviewed changes

estesp approved these changes Oct 28, 2025

View reviewed changes

github-project-automation bot moved this from Needs Triage to Review In Progress in Pull Request Review Oct 28, 2025

estesp merged commit 477522a into containerd:release/2.1 Oct 28, 2025
145 of 150 checks passed

github-project-automation bot moved this from Review In Progress to Done in Pull Request Review Oct 28, 2025

fuweid deleted the weifu/backport-12400-21 branch October 28, 2025 16:03

austinvazquez mentioned this pull request Nov 5, 2025

[release/2.1] Prepare release notes for v2.1.5 #12483

Merged

dmcgowan changed the title ~~[release/2.1] cri/server/podsandbox: disable event subscriber~~ [release/2.1] Disable event subscriber during task cleanup Nov 5, 2025

dmcgowan added the impact/changelog label Nov 5, 2025

KCSesh mentioned this pull request Nov 7, 2025

Update containerd [1.7/2.0/2.1] verisons to the latest bottlerocket-os/bottlerocket-core-kit#724

Merged

brandond mentioned this pull request Nov 13, 2025

Leftover containerd-shim-runc-v2 processes rancher/rke2#8825

Closed

xhejtman mentioned this pull request Nov 23, 2025

Leftover shims in containerd 2.0 and later #12560

Open

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

[release/2.1] Disable event subscriber during task cleanup #12410

[release/2.1] Disable event subscriber during task cleanup #12410

Uh oh!

fuweid commented Oct 24, 2025 •

edited

Loading

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

5 participants

[release/2.1] Disable event subscriber during task cleanup #12410

[release/2.1] Disable event subscriber during task cleanup #12410

Uh oh!

Conversation

fuweid commented Oct 24, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

5 participants

fuweid commented Oct 24, 2025 •

edited

Loading