Use a backend metrics dashboard for cron jobs, API failures, and bounded business events, then add a separate heartbeat monitor for every nightly pipeline deadline. The deciding constraint is signal quality: metrics can count completed and failed runs, but they cannot report a job that emitted nothing because it never started. TL;DR: chart successes, failures, duration, backlog, and business-event counts; enrich failure investigation with error records and searchable logs; send one dead-man...