Three configuration mistakes, one Friday deploy, and a very long weekend. A candid post-mortem on the container pitfalls that are easy to miss until they are not.
The deploy went out at 4:47pm on a Friday. By 6pm we had our first user reports. By midnight, two of us were staring at container logs trying to understand why memory was climbing without bound.
Mistake one: we were not setting memory limits on our containers. The default is unlimited. Under load, our Node.js process would balloon to fill available RAM, triggering OOM kills on the host — taking out unrelated services in the process.
Mistake two: our health checks were too shallow. We were pinging /health and checking for a 200 response. But /health did not check our database connection or Redis connection — it just confirmed the process was alive. A container could be healthy while being completely unable to serve real requests.
Mistake three: we were using the latest tag in production. A dependency update that had not been tested in staging went straight to prod. It took us four hours to identify this as the root cause.
None of these are exotic problems. They are all in the Docker documentation. We just had not read that part carefully enough.