<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>sre on Dhaval Shah</title><link>https://www.dhaval-shah.com/tags/sre/</link><description>Recent content in sre on Dhaval Shah</description><generator>Hugo -- gohugo.io</generator><lastBuildDate>Sat, 25 Jul 2026 01:00:50 +0000</lastBuildDate><atom:link href="https://www.dhaval-shah.com/tags/sre/index.xml" rel="self" type="application/rss+xml"/><item><title>A Black Friday Incident Took 9 Days to Resolve - Here's the Process That Would Have Changed That</title><link>https://www.dhaval-shah.com/sre-gc-ai-review/</link><pubDate>Sat, 25 Jul 2026 01:00:50 +0000</pubDate><guid>https://www.dhaval-shah.com/sre-gc-ai-review/</guid><description>Background An older post on this blog covered the technical details of this incident - the GC types, the flags, the before-and-after metrics. This post covers something different: the process of how the investigation ran, where it lost time, and which decisions (with more structured approach) would have changed.
The incident On Black Friday, an online checkout platform running at over 1000 TPS. CPU spiking to 100%, application crashing. Restarted every 12 hours to keep the business running while the team investigated.</description></item></channel></rss>