TL;DR: Give the web application a cheap health endpoint, record job_success , job_failure , and last_run for each notification worker, and use an external heartbeat for failed-job detection. For a logistics notification service, that is the least complex design that catches both visible delivery failures and the nastier case: a worker or cron task that never started. Retain enough evidence to compare a release with its predecessor, but don't keep every payload. Rollback safety comes from pr...