All posts

GitHub's August 17 Outage Became a Retry Storm. Check Your Own Retries.

GitHub was degraded for 7h47m on August 17, 2026. A scaling gap started it; retries kept it going. What it means for your CI, your SLA, and your AI clients.

Misha Druzhinin
Telephone operators working a crowded Bell System switchboard in 1943, plugging cords into a wall of jacks.

On Monday, August 17, 2026, GitHub.com was degraded for 7 hours and 47 minutes, from 13:28 to 21:15 UTC. At peak, about 20% of web and API requests failed, and archive and raw-content downloads failed about half the time. Actions, pull requests, issues, webhooks, Copilot, and SAML/OIDC sign-in were all hit, according to GitHub’s incident notes.

The early coverage framed it as a budget story. Devin Culbertson’s TechTimes analysis argued that one outage used up nearly a year’s downtime allowance at three nines. That framing deserves a closer look, and we’ll get to it.

GitHub’s own notes tell a more useful story for engineers. A scaling gap started the outage. Retries kept it alive for hours after GitHub fixed the component that failed. That second part applies to almost every team, especially teams building on LLM APIs. CTO Vlad Fedorov’s post-incident blog post put it plainly: “If you were trying to ship software that day, we let you down.”

This post leans on GitHub’s status-page notes and that August 20 post, plus the sources linked below. GitHub’s August availability report hadn’t been published at the time of writing.

Key Takeaways

  • GitHub corrected the failing component at 16:36 UTC, about three hours in, and most services recovered. The incident stayed open another 4 hours 39 minutes, mostly for Copilot authentication failures.
  • A latent VS Code retry bug pushed Copilot Token Service traffic from 7–9K to 70–100K requests per second, about 10 times normal.
  • Retries multiply across layers. Three attempts at each of five layers is 243 calls for one request. The official Anthropic and OpenAI Python SDKs already make three attempts by default.
  • “Availability” here is three different numbers: GitHub’s weighted status-page uptime, the contractual SLA, and what your team experienced. Only the last one is yours to measure.

What happened on August 17

The outage started in GitHub’s Central US data center when traffic reached a new peak. In Fedorov’s words, “a critical infrastructure component in our Central US data center failed to scale with it.” From there, the status page added services one by one over the next two hours.

  • 13:28 UTC. Impact begins, per GitHub’s resolution note.
  • 13:40–15:21. API Requests, Actions, Webhooks, Issues, Pull Requests, Copilot, Pages, and Git Operations report degradation.
  • 16:36. GitHub reports: “We identified the problematic component and have taken corrective actions.”
  • 16:59. Most services marked mitigated. Copilot isn’t among them.
  • 17:30–20:22. Issues and Git Operations degrade again. API Requests fails again briefly. Sporadic authentication failures continue.
  • 19:13. GitHub has “partially disabled authentication token retries.”
  • 21:02–21:15. Copilot Token Service recovers. The incident is resolved.

The chart below is the part worth staring at. The orange segments are failures reported after the component was fixed.

August 17 impact by service, UTC Timeline of GitHub services reported degraded on August 17, 2026, in UTC. API Requests 13:41 to 16:59, then 18:48 to 19:01. Actions 13:42 to about 18:03. Webhooks 13:44 to 16:59. Issues 13:46 to 16:59, then 17:36 to 20:22. Pull Requests 13:58 to 16:59. Copilot 14:31 to 21:02. Pages 15:10 to 16:59. Git Operations 15:21 to 16:59, then 17:30 to 18:23. GitHub corrected the failing component at 16:36. Most services recovered then. Residual failures, mostly Copilot authentication, lasted until 21:02. GitHub partially disabled authentication token retries at 19:13. Source: GitHub status page. August 17 impact by service (UTC) From each status update reporting degradation to the one reporting recovery 13:00 14:00 15:00 16:00 17:00 18:00 19:00 20:00 21:00 API Requests Actions Webhooks Issues Pull Requests Copilot Pages Git Operations 16:36 component fixed 19:13 retries cut Before 16:36 After 16:36: residual failures
Bars run between status updates, so they include reporting lag. Source: GitHub status page, August 17, 2026. Actions' and Copilot's end times come from GitHub's resolution note.

Most services recovered at 16:36. But the incident stayed open for another 4 hours 39 minutes, more than half its length. GitHub’s notes attribute the slowest part of that tail, Copilot’s recovery, to client retry amplification. They also list scraping attacks on codeload endpoints as a complicating factor.

A scaling gap started it. Retries stretched it out.

GitHub added a root-cause summary to the incident page about a day later. We’ve broken it into the five steps below. Each one is ordinary on its own.

How a capacity limit became a retry storm A chain of five steps from GitHub's August 17 incident notes. One: a new traffic peak reaches the Central US load balancers. Two: an Istio sidecar hits its concurrency limit and does not scale, because the autoscaling policy watched the host service, not the sidecar. Three: failures cascade until four HAProxy nodes exhaust their flow limits. Four: the gateway authentication path degrades, giving about 20 percent web and API errors and about 50 percent errors on archive and raw downloads. Five: optimistic retries add load to the internal load balancers, which feeds back into step three. A second loop on the right: a latent VS Code retry bug multiplies Copilot token requests about tenfold, from 7 to 9 thousand to 70 to 100 thousand requests per second, which feeds back into step four. The fix at the bottom: pause HAProxy on the saturated nodes, cut gateway retries, reject token requests with 403 at the load balancer, then ramp traffic back site by site. 1. New traffic peak in Central US Load balancers approach saturation 2. Istio sidecar hits its concurrency limit Autoscaler watched the host, not the sidecar 3. Four HAProxy nodes exhaust flow limits One failure cascades into the next 4. Gateway auth path degrades ~20% web/API errors, ~50% archive and raw 5. Optimistic retries add load Internal load balancers overload further Client retry bug Latent VS Code bug Copilot token traffic 7–9K to 70–100K RPS ~10x amplification Fix: shed load first, then ramp back Pause saturated HAProxy nodes · cut gateway retries · 403 token requests · ramp per site feedback loop
The August 17 cascade, simplified from GitHub's incident notes. The orange arrows are the loops that kept the outage going after the trigger was fixed. Original diagram; source: GitHub status page.

The first failure was an autoscaling-policy gap, not a code bug. An Istio sidecar reached its concurrency limit and didn’t scale. GitHub calls the policy misconfigured: it “watched host service but not sidecar limits.” So the autoscaler watched one signal while the sidecar saturated.

The failures then cascaded until four HAProxy nodes exhausted their flow limits. That degraded the gateway’s authentication path, causing widespread authentication latency and failures. GitHub’s notes say the problem “was worsened by optimistic retry logic which overloaded internal load balancers.”

GitHub moved some failing traffic from Central US to Northern Virginia, where it was served. Its notes also describe a retry storm in Northern Virginia. A client made it worse. Delayed replies from one internal endpoint triggered “a latent retry bug in VS Code that amplified traffic by approximately 10x.” Copilot Token Service traffic rose from a normal 7–9K requests per second to 70–100K.

Look at how GitHub stopped it. Pausing HAProxy on the saturated nodes produced “immediate broad recovery.” For the retry storm, GitHub temporarily reduced gateway retries with a pull request. It also answered Copilot token requests with a 403 at the load balancers. Then it ramped traffic back up site by site. In short, GitHub had to refuse work so that the work it accepted could succeed.

There’s a lesson in the client part, too. You can patch a server in minutes. You can’t patch every installed editor during an incident. Once a retry bug ships in a client, your fastest lever is the server’s willingness to say no.

Why retries multiply instead of add

Each retrying layer multiplies the load on the layer below it. If every layer makes up to three attempts, two layers make 9 calls and five layers make 243. That 243x figure is the worked example in Marc Brooker’s AWS Builders’ Library article on timeouts and retries. Google’s SRE book gives a 64x version with four attempts per layer. Retries at three layers turn “a single user action” into “64 attempts (4^3) on the database” (Google SRE).

Attempts per user request when every layer retries Bar chart of attempts reaching the bottom of a call stack for one user request, when each layer makes up to 3 attempts. 1 layer: 3 attempts. 2 layers: 9. 3 layers: 27. 4 layers: 81. 5 layers: 243. The 5-layer figure matches the AWS Builders' Library example. Three attempts per layer is also what the Anthropic and OpenAI Python SDKs make by default: one call plus 2 retries. Attempts per request when every layer retries Each layer makes up to 3 attempts (1 call + 2 retries) 0 60 120 180 240 3 1 layer 9 2 layers 27 3 layers 81 4 layers 243 5 layers Layers that retry the same call (SDK, wrapper, queue, agent loop, CI re-run)
Illustrative math: 3 to the power of the number of retrying layers. The 243x case for five layers is the example in the AWS Builders' Library.

The worst part is timing. Retries arrive exactly when the dependency is already struggling. A service that starts failing 20% of requests gets more traffic, not less. That’s the feedback loop in the diagram above.

The standard defenses are well documented, and they’re mostly about restraint:

  • Cap attempts per request. The SRE book’s rule: “If a request has already failed three times, we let the failure bubble up” (Google SRE).
  • Use a retry budget. The same chapter describes a per-client budget. A request is retried only “as long as this ratio is below 10%,” meaning retries to requests.
  • Retry at one layer. Brooker recommends that you “perform retries at a single point in the stack.”
  • Add jitter. Exponential backoff without randomness turns a crowd of clients into synchronized waves.

GitHub’s own follow-up matches this list. Fedorov wrote that GitHub is “applying consistent retry limits, retry budgets, and variable timeouts across service-to-service interactions.”

Your AI stack probably retries at more layers than you think

Teams building on LLM APIs inherit retry layers by default, often without noticing. The official Anthropic Python SDK and OpenAI Python SDK both say certain errors “are automatically retried 2 times by default, with a short exponential backoff.” Both retry connection errors, 408, 409, 429, and 5xx responses. That’s three attempts per call before your code sees an error.

Now count the layers above it. A common production path looks like this:

  1. The SDK call: 3 attempts.
  2. A tenacity or homegrown retry wrapper around it: 3 attempts.
  3. A job queue that redelivers failed jobs: 3 deliveries.
  4. An agent loop where the model reads a tool error and tries the same tool again.
  5. A person or CI job that re-runs the whole thing.

Layers 1 to 3 alone make 27 calls to a provider that’s already returning overload errors. Layer 4 is the newer risk. An agent that sees “503 Service Unavailable” in a tool result can just try again, with no backoff and no budget. The model decides how often to retry, and nobody configured it.

The fix is the same as GitHub’s: one owner for retries, a hard cap, and a budget. Here’s a minimal version with the Anthropic SDK. Disable the SDK’s retries so one layer owns them. Then cap attempts, honor retry-after, and spend from a shared budget. The pattern is the same for OpenAI’s client, which also accepts max_retries=0.

import random
import time

import anthropic

client = anthropic.Anthropic(max_retries=0)  # this wrapper owns retries

RETRYABLE_STATUS = {408, 409, 429}  # plus every 5xx, as the SDK does


def is_retryable(err: Exception) -> bool:
    if isinstance(err, anthropic.APIConnectionError):  # includes timeouts
        return True
    if isinstance(err, anthropic.APIStatusError):
        return err.status_code in RETRYABLE_STATUS or err.status_code >= 500
    return False


class RetryBudget:
    """Retries refill at ~10% of requests (Google SRE's per-client ratio).

    Starts full, so it allows a burst of `cap` retries. Not thread-safe.
    """

    def __init__(self, ratio: float = 0.1, cap: float = 20.0):
        self.ratio, self.cap, self.tokens = ratio, cap, cap

    def on_request(self) -> None:
        self.tokens = min(self.cap, self.tokens + self.ratio)

    def try_spend(self) -> bool:
        if self.tokens >= 1:
            self.tokens -= 1
            return True
        return False


budget = RetryBudget()


def create_message(**kwargs):
    budget.on_request()
    for attempt in range(3):  # at most 3 attempts per request
        try:
            return client.messages.create(**kwargs)
        except Exception as err:
            if not is_retryable(err) or attempt == 2 or not budget.try_spend():
                raise  # let the failure bubble up
            delay = random.uniform(0, 0.5 * 2**attempt)  # full jitter
            response = getattr(err, "response", None)
            retry_after = response.headers.get("retry-after") if response is not None else None
            if retry_after and retry_after.isdigit():
                delay = max(delay, min(float(retry_after), 30.0))  # respect it, capped
            time.sleep(delay)

This sketch is single-threaded; guard the budget with a lock if you share it across threads. When the budget is empty, the wrapper fails fast instead of piling on. For agent tools, return a structured error that says whether to retry, such as {"error": "upstream_unavailable", "retry": false}. Count tool attempts against the same budget, so the model can’t retry without limit.

The shared-fate point matters too. GitHub’s docs say the Copilot cloud agent runs in “its own ephemeral development environment, powered by GitHub Actions.” If your agentic workflow depends on Copilot, Actions, and pull requests, one platform incident can take out all three. On August 17, all three were degraded at once.

“A year’s downtime budget”: which budget?

TechTimes’s arithmetic is right for an annual, time-based view. A year at 99.9% leaves about 8.8 hours of downtime, and this incident ran 7 hours 47 minutes. But “availability” for GitHub is measured three different ways, and they disagree.

Measure How it counts What August 17 does to it
Status-page uptime Per service over 90 days. Major Outage counts 100%, Partial Outage 30%, Degraded Performance 0% Depends on how each period was classified. TechTimes reports Actions fell from 99.39% to 99.33%
Contractual SLA 99.9% per calendar quarter. A minute counts as downtime when “the error rate exceeds five percent” Any minute above 5% errors counts in full; web/API errors peaked near 20%
Your experience Failed runs, queued jobs, blocked merges, on-call time Only you can measure it

The first row comes from GitHub’s April 2026 post on status-page transparency. Under those rules, an hour of Partial Outage counts as 18 minutes of downtime, and Degraded Performance counts as none. We couldn’t verify the 99.33% figure against an archived copy of the status page, so treat it as TechTimes’s reading.

The second row is the one in your contract. The GitHub Online Services SLA (June 2026 version) covers Enterprise Cloud, Actions, and Packages. It measures uptime per calendar quarter, not per year. The third quarter has 92 days, so 99.9% allows about 132 minutes, or 2 hours 12 minutes, of downtime.

Actions was degraded from 13:42 to about 18:03 UTC, or 261 minutes. If its error rate stayed above 5% for more than about 132 of those minutes, one afternoon used up the quarter’s allowance. We don’t know Actions’ minute-by-minute error rate, and GitHub hasn’t published it. Credits are 5%, 10%, or 25% of fees, depending on uptime. They aren’t automatic: you must request them in writing within 30 days after the quarter ends. For the third quarter, that’s by October 30.

The broader point is about your own SLOs. If you promise customers anything that depends on GitHub, a vendor’s number won’t tell you what you lost. Measure it yourself.

When not to over-engineer this

One bad month doesn’t justify a second CI system. A hot-standby CI on another provider costs real engineering time. It also drifts out of date unless you exercise it. GitHub’s July availability report describes ongoing work to move traffic to Azure. Fedorov’s post says Azure now serves roughly 58% of platform load, up from 12% in May. Things may improve.

For most teams, three cheaper controls cover most of the risk:

  • Cache what your builds download. On August 17, archive and raw-content downloads failed about half the time. Builds that fetch dependencies or tools straight from GitHub URLs inherit that failure. A registry or proxy you control takes GitHub out of that path.
  • Write down the degraded-mode path. Decide in advance how you’d ship a hotfix without Actions: a build host you can trigger by hand, a documented manual deploy, a named approver. Self-hosted runners don’t help here, because they get their jobs from the Actions service. A one-page runbook beats a second platform you never test.
  • Fix your own retries. This one costs an afternoon and helps with every dependency, not just GitHub.

Reader questions

Did AI agent traffic cause the outage?

GitHub hasn’t said that. Its notes describe “a new peak in traffic” meeting a scaling policy that watched the wrong signal. Load is growing fast, though. Fedorov’s post says monthly commits grew from 1.4 billion to 2.9 billion since April. The amplification GitHub describes came from server-side retries and a client retry bug.

Should we move off GitHub Actions?

Not on this incident alone. Measure what August cost you in failed runs and delayed releases first. If that cost exceeds what a fallback path would cost to build and maintain, build the fallback. Otherwise, cache dependencies and write the runbook.

Will GitHub issue SLA credits automatically?

No. Under the June 2026 SLA, Enterprise Cloud customers must request credits in writing within 30 days of the end of the quarter. Check which services your contract covers before you file.

Where to start

This week, pick the one external call path your product depends on most. That might be GitHub’s API or an LLM provider. Trace one request through it and list every layer that retries: the SDK, your wrappers, the queue, the agent loop, the CI job. Keep retries at one layer, set a hard cap, and add a budget. Then answer one question: when that dependency returns errors for four hours, does your system back off or pile on?

If you’d like a second pair of eyes on your retry paths or agent tooling, see our audit work or talk to Moonlight AI.