Repository navigation
Conversation
Add SuccessExitStatus=143 for services that exit on SIGTERM (kafka, host-inventory worker and API) so a clean SIGTERM stop is not reported as failed, and add stop_timeout headroom for services that may be mid-work at shutdown (vulnerability consumers and taskomatic, 30s; host-inventory, 45s; ingress, 20s) so they stop cleanly instead of being force-killed.
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (4)
Included review availability: This review used your included allowance. Your plan provides up to 1 included review per hour; 0 remain after this review. 📝 WalkthroughWalkthroughContainer roles now set explicit stop timeouts. Inventory and Kafka Quadlet services treat exit status 143 as successful. ChangesGraceful Shutdown Configuration
Priority: ➖ Normal Estimated code review effort: 2 (Simple) | ~10 minutes Change: Bug fix Suggested reviewers: Merge Risk: ⚪ Minimal · up to The configured stop periods are effective with the supported Quadlet generation and fit within systemd’s standard shutdown deadline. No concrete shutdown failure is established, so the change is mergeable. 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
@vkrizan @pfreyburg ready for your review. |
| PartOf=foreman.target | ||
| [Service] | ||
| Restart=on-failure | ||
| SuccessExitStatus=143 |
There was a problem hiding this comment.
I'd prefer not to handle it like this, but I've also certainly seen Foreman things end up in "failed" state with 143.
Is having a sufficiently sized stop_timeout not sufficient to avoid being SIGTERM'ed?
There was a problem hiding this comment.
I came across this article that explains the 143 a bit https://www.dash0.com/guides/kubernetes-exit-code-143-a-practical-guide
There was a problem hiding this comment.
Interestingly, https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html#SuccessExitStatus= suggests that SIGTERM should already be considered successful. I'll need to dig that more on Monday.
There was a problem hiding this comment.
You are reviewing container exit codes after a deployment or node drain and see 143. If your monitoring alerts on it, you might think something failed. Exit code 143 is not an error. It is 128 plus signal 15 (SIGTERM), and it means the container’s PID 1 process received SIGTERM and exited voluntarily. This is exactly what docker stop is designed to do.
https://www.netdata.cloud/guides/docker/docker-exit-code-143/
This seems to be some container specific signal reporting.
There was a problem hiding this comment.
ah, it's sigterm plus container shenanigans. gotcha.
Problem
When IoP services are stopped (e.g.
systemctl stop foreman.target) several units end upfailed: some exit on SIGTERM (143), which systemd treats as a failure, and some are force-killed (137) before they finish.Change
In the IoP roles (iop_vulnerability, iop_inventory, iop_kafka, iop_ingress):
SuccessExitStatus=143for services that exit on SIGTERM (kafka, host-inventory worker and API).stop_timeoutheadroom for services that may be mid-work at shutdown: vulnerability consumers and taskomatic (30s), host-inventory (45s), ingress (20s).Testing
Verified on a Foreman development instance running IoP by generating the quadlets via the containers.podman module and running
systemctl stop foreman.target: the generated units carry StopTimeout/SuccessExitStatus and the services stop cleanly (no failed units, no SIGKILL).Related
Part of the IoP graceful-shutdown fix: