Skip to content

Handle graceful shutdown for IoP service quadlets - #926

Open
brantleyr wants to merge 1 commit into
theforeman:masterfrom
brantleyr:graceful-shutdown
Open

brantleyr wants to merge 1 commit into
theforeman:masterfrom
brantleyr:graceful-shutdown

Conversation

@brantleyr

@brantleyr brantleyr commented Oct 8, 2026 •

Copy link
Copy Markdown

Problem

When IoP services are stopped (e.g. systemctl stop foreman.target) several units end up failed: some exit on SIGTERM (143), which systemd treats as a failure, and some are force-killed (137) before they finish.

Change

In the IoP roles (iop_vulnerability, iop_inventory, iop_kafka, iop_ingress):

  • SuccessExitStatus=143 for services that exit on SIGTERM (kafka, host-inventory worker and API).
  • stop_timeout headroom for services that may be mid-work at shutdown: vulnerability consumers and taskomatic (30s), host-inventory (45s), ingress (20s).

Testing

Verified on a Foreman development instance running IoP by generating the quadlets via the containers.podman module and running systemctl stop foreman.target: the generated units carry StopTimeout/SuccessExitStatus and the services stop cleanly (no failed units, no SIGKILL).

Related

Part of the IoP graceful-shutdown fix:

Add SuccessExitStatus=143 for services that exit on SIGTERM (kafka,
host-inventory worker and API) so a clean SIGTERM stop is not reported
as failed, and add stop_timeout headroom for services that may be
mid-work at shutdown (vulnerability consumers and taskomatic, 30s;
host-inventory, 45s; ingress, 20s) so they stop cleanly instead of
being force-killed.
@coderabbitai

coderabbitai Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 370da14c-038a-4569-af28-dd0590dc63c4
📥 Commits

Reviewing files that changed from the base of the PR and between 647887b and 7214062.

📒 Files selected for processing (4)
  • src/roles/iop_ingress/tasks/main.yaml
  • src/roles/iop_inventory/tasks/main.yaml
  • src/roles/iop_kafka/tasks/main.yaml
  • src/roles/iop_vulnerability/tasks/main.yaml

Included review availability: This review used your included allowance. Your plan provides up to 1 included review per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

Container roles now set explicit stop timeouts. Inventory and Kafka Quadlet services treat exit status 143 as successful.

Changes

Graceful Shutdown Configuration

Layer / File(s) Summary
Container stop timeouts
src/roles/iop_ingress/tasks/main.yaml, src/roles/iop_inventory/tasks/main.yaml, src/roles/iop_vulnerability/tasks/main.yaml
Ingress containers use a 20-second timeout. Inventory containers use 45 seconds. Vulnerability containers use 30 seconds. Comments describe the intended shutdown headroom.
Quadlet exit status handling
src/roles/iop_inventory/tasks/main.yaml, src/roles/iop_kafka/tasks/main.yaml
Inventory and Kafka Quadlet services include exit status 143 in SuccessExitStatus.

Priority: ➖ Normal

Estimated code review effort: 2 (Simple) | ~10 minutes

Change: Bug fix

Suggested reviewers: alleny244

Merge Risk: ⚪ Minimal · up to 72140

The configured stop periods are effective with the supported Quadlet generation and fit within systemd’s standard shutdown deadline. No concrete shutdown failure is established, so the change is mergeable.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0…
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description check ✅ Passed The description clearly explains the graceful-shutdown problem, the SuccessExitStatus=143 and stop-timeout changes, and the reported testing results. It is directly related to the changeset.
Title check ✅ Passed The title clearly and concisely summarizes the main change: graceful shutdown handling for IoP service Quadlet units.
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@brantleyr

Copy link
Copy Markdown
Author

@vkrizan @pfreyburg ready for your review.

PartOf=foreman.target
[Service]
Restart=on-failure
SuccessExitStatus=143

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this standard to handle the 143 like this? Should the foreman services benefit from this as well?

cc @ekohl @evgeni

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd prefer not to handle it like this, but I've also certainly seen Foreman things end up in "failed" state with 143.

Is having a sufficiently sized stop_timeout not sufficient to avoid being SIGTERM'ed?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I came across this article that explains the 143 a bit https://www.dash0.com/guides/kubernetes-exit-code-143-a-practical-guide

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Interestingly, https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html#SuccessExitStatus= suggests that SIGTERM should already be considered successful. I'll need to dig that more on Monday.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

143 is not SIGTERM

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You are reviewing container exit codes after a deployment or node drain and see 143. If your monitoring alerts on it, you might think something failed. Exit code 143 is not an error. It is 128 plus signal 15 (SIGTERM), and it means the container’s PID 1 process received SIGTERM and exited voluntarily. This is exactly what docker stop is designed to do.

https://www.netdata.cloud/guides/docker/docker-exit-code-143/

This seems to be some container specific signal reporting.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ah, it's sigterm plus container shenanigans. gotcha.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants