At four on a Sunday morning an internal reconciliation service went down. All four pods were killed for memory inside twenty minutes of each other, came back, and were killed again a few hours later. The service had last been deployed in February. Seven months of continuous uptime, and it was the only thing in our estate that had any. The cause was a leak of about forty megabytes per pod per day. We create an SDK client per request, and each client registers a listener on a static registry th...