On Diwali night, the Kubernetes cluster died

The PagerDuty alert went off at 11:30 PM. Outside the window, the Delhi sky was filled with fireworks, and the gulab jamun my mom made was cooling on the plate. Right in the middle of Diwali, the production cluster went down.

The moment I turned on my screen, I knew it was serious. Three nodes had simultaneously fallen into NotReady status, and dozens of pods were stuck in CrashLoopBackOff. The cause was a single autoScaler setting. The HPA threshold we'd put in place for traffic spikes didn't match the holiday traffic pattern. During the Diwali season, payment traffic follows a completely different curve than usual. Failing to account for that was purely our team's mistake.

Hands typing on laptop in dark room with terminal commands during a late-night incident

Incident response in the middle of a festival

I was the on-call engineer. That was fine. The problem was that neither of the two backup engineers could be reached. One was on a family trip to Jaipur, and the other had their phone on silent while watching the fireworks. I posted in the Slack channel, called them myself, and ended up recovering the nodes one by one, hammering away at kubectl alone. Cordon, drain, restart. When the last pod came back to Running status at 2 AM, my hands were shaking.

When I went out to the living room, everyone was asleep. A single candle was still burning on the rangoli. That scene is still vivid in my mind. More than the guilt of having lost an entire festival, the certainty that this would happen again weighed heavier.

India's festival calendar is a blind spot in incident response

Most incident response processes are designed around the Western calendar. Around Christmas and New Year, code freezes go in and on-call rotations get adjusted. But what about Diwali, Holi, Eid, Pongal? At most global companies, they're treated as regular workdays. When engineers at the India office end up on call, they either give up the festival or become unreachable.

This isn't a matter of individual responsibility. It's a system design problem. If you don't incorporate cultural calendars into on-call rotations, you're allowing a structure where incident recovery times will inevitably slow down.

Delhi skyline lit up with Diwali fireworks and festive lights at night

What we changed afterward

At the team retro, I pushed for a few things. During major Indian festival periods, the primary on-call should be assigned to engineers in other regions who aren't celebrating, or at the very least, backup coverage should be doubled. HPA thresholds should be split by season, with holiday traffic profiles managed separately. And deploy freezes should be enforced starting two weeks before a festival. These are simple measures, but until then, no one had proposed them. Uncomfortable topics don't bring themselves up.

This year, I'm off the on-call rotation for Diwali. Even if something goes wrong, the Singapore team is handling first response. I think I'll get to eat my mom's gulab jamun while it's still warm.

Comments