Skip to main content
Writing
11 min read

CI/CD contingency: why self-hosted runners saved no one in the GitHub Actions outage

When GitHub Actions went down for nine hours, teams with self-hosted runners stopped too, and when it came back the queue of lost events didn't come back with it. CI contingency tends to cover the machine that runs the build and forget the layer that knows which build has to run.

  • GitHub Actions
  • CI/CD
  • Self-hosted Runner
  • DevOps
  • Contingency
  • Infrastructure
  • Reliability

Thursday, 6 August 2026, 15:22 UTC. GitHub marks Actions as "degraded performance". Twenty minutes later, availability drops too. From there to mitigation it took almost nine hours, and the incident was only closed at 2:04 the following morning.

Our pipeline here stopped along with the rest of the world for a few hours. And I don't want to turn this into a complaint post, because going down is part of the deal and Actions is still an absurdly good piece of infrastructure for what it costs. What got me was something else: I spent the day watching people say they were fine because they had a self-hosted runner.

They weren't.

Your own runner gives you the machine, not the decision

The official incident text is quite specific on this point: "Both GitHub-hosted and self-hosted runners are affected". And, further down, that those running their own runners might see errors or rate limiting at the moment the runner registers.

That says a lot about where the service actually lives. When we bring up a self-hosted runner, what comes in-house is the compute: the processor that runs pnpm test, the disk that holds the cache, the network that pulls the image. But running the build is the last link in a much longer chain. Before that, someone has to receive the push event, decide that it triggers a workflow, read the YAML, build the job matrix, choose which runner each job goes to, authenticate that runner, hand over the checkout token and then receive the result back to paint the green check on the PR.

None of that happens on your machine. All of that is the service.

In other words: a self-hosted runner doesn't take you out of the dependency, it changes its nature. You stopped depending on GitHub's capacity and started depending on GitHub's dispatch. And dispatch is exactly what broke, which explains why both kinds of runner went down in the same minute.

There was an even more uncomfortable detail for those with their own runners. In the incident updates, GitHub warned that some Actions Runner Controller pods got stuck in an idle state and that recovery needed a human hand: deleting the pods with kubectl or redeploying the ARC application. Whoever invested in their own runners to have more control received, at the worst moment of the day, one more manual task. GitHub itself acknowledged this and said the next versions of the Runner and ARC will bring automatic recovery.

What really hurt came afterwards

Now the point almost nobody commented on, and the most expensive one in the whole story.

In the closing update, GitHub warned that some events that trigger workflows, including push and pull request, were not processed during the incident and cannot be replayed automatically. The recommendation was this: push a new commit, update the PR, or go there and run the workflow by hand.

It is worth reading that again slowly, because the implication is big. It isn't that the builds failed. Failing is great: a failure is red on the screen, it is a notification, it is someone noticing. What happened was quieter: part of the events simply turned into nothing. No run, no pending check, no record. From the outside, that commit looks like it passed.

So the damage didn't end at 2:04 when the incident was closed. It carried on into the next morning, in the form of a question nobody on the team could answer: what exactly was left behind?

Anyone whose contingency plan was "wait for it to pass" found out that waiting brings the service back, but doesn't bring the work back.

The part nobody duplicates

This is where I want to get to.

When we sit down to think about CI contingency, instinct tells us to replicate what is visible and easy to draw on the whiteboard: the machine that runs the build. I buy a runner. I spin up a Jenkins. I keep a GitLab CI in reserve. All of that is money and work invested in the execution layer.

But the execution layer wasn't the one that went down. And if you put the three facts of the incident side by side, they all point to the same place: both kinds of runner went down together because what broke was dispatch; the ARC pods got stuck because they lost their link to whoever commands them; and the lost events didn't come back because their queue exists nowhere else.

Everyone duplicates the part that runs the build. Nobody duplicates the part that knows which build has to run, and that is the one that went down.

That layer is far less glamorous than a runner: it is the event queue, the association between commit and execution, the history of what ran and what didn't. It doesn't show up in any architecture diagram, nobody puts it in the budget, and it is the one thing you really have no copy of.

And it is cheap to rebuild afterwards, if you thought about it beforehand. It doesn't take a second CI, or Kubernetes, or anything expensive. It takes an answerable question: which commits came in and have no execution associated with them?

# Commits that landed on main during the incident window
# and that have no workflow run associated with them.
git log --since="2026-08-06 15:00" --format=%H origin/main | while read sha; do
  runs=$(gh run list --commit "$sha" --json databaseId --jq 'length')
  [ "$runs" -eq 0 ] && echo "no run: $sha"
done

It's six lines. They don't stop CI from going down and they don't speed up GitHub's recovery by a single minute. But they turn "I think everything is fine" into a finite list of commits to reprocess, and that is the difference between closing the incident and thinking you closed it. Does that make sense?

The pipeline that would fit anywhere

That said, there is the half of contingency that is about where to take the build. And here it is worth stating the less obvious version of the problem.

What ties most teams to Actions isn't Actions. It is having written the whole build inside it: thirty YAML steps, fifteen uses: from the marketplace, secrets tied to the vendor's syntax and business logic scattered across workflow if:s. Whoever does it that way doesn't have a bad plan B, they have a rewrite ahead of them. And a rewrite isn't something you do in the middle of an incident.

The design that survives is the opposite: the logic lives in a script, and CI is a thin shell that calls that script.

# Makefile — the real build, which runs the same on your machine and in CI.
ci: lint typecheck test build

lint:
	pnpm biome check .

typecheck:
	pnpm tsc --noEmit

test:
	pnpm vitest run --coverage

build:
	docker build -t $(IMAGE):$(SHA) .

And then the shell, in three different places:

# .github/workflows/ci.yml
- uses: actions/checkout@v4
- run: make ci
# .gitlab-ci.yml
ci:
  script: make ci
// Jenkinsfile
stage('ci') { steps { sh 'make ci' } }

Notice that the core doesn't change. What changes is the line that knows how to check out and the line that knows how to call make. That is why I don't like the "replace Actions with Jenkins" conversation: switching CI shouldn't be a project, it should be an afternoon. And, to be very honest, replacing Actions with a poorly maintained Jenkins is a terrible deal: you trade GitHub's downtime, which has an on-call team, for your own: late security patches, a plugin that breaks on upgrade, a disk filling up, no elastic scale and a single point of failure nobody monitors. Added up over a year, the hours Actions leaves you stranded probably cost less than that.

Jenkins and GitLab CI here are not a recommendation to switch. They are proof that the portable design works.

Portability isn't having two CIs switched on. It is having a pipeline that would fit in any of them.

It is the same reasoning I followed building Lisa, the CLI I maintain: the logic for opening a PR, checking CI and merging is a single one, and there is a thin translation layer per platform for GitHub, GitLab and Bitbucket. When I need to ask whether CI passed, the question is always the same; what changes is whether it becomes a gh pr checks or a call to GitLab's pipeline API. It is a good example precisely because it isn't perfect: on Bitbucket that check returns unknown, because there is no cheap path from the PR URL. The abstraction holds the gap without contaminating the rest.

The ladder, from cheap to expensive

Contingency is a matter of dosage, and I think the best way to decide is to look at it as steps, knowing the price of each one:

  1. Portable pipeline. The build becomes a script, CI becomes a shell. It costs a few hours and you will use it every day, even with no incident at all, because make ci runs the same on your machine.
  2. Knowing what was left behind. The six lines above, or the equivalent in your stack. It costs an afternoon and it is the only step that goes after the damage that remains once the service is back.
  3. Manual deploy, documented and tested. A docker build, a push to the registry, a restart command. The word that carries this step is tested: a runbook nobody has run in the last six months is fiction. Put it on the calendar and run it once a quarter, even with no incident.
  4. Repository mirror. Git is distributed, so the code itself is already on everyone's laptop. What you lose in a bigger outage isn't the code, it is the hub: PRs, issues, pipeline, registry. A mirror on another provider is one more git push and solves code continuity for next to nothing.
  5. A secondary CI that is really switched on. Here the price moves up a level: two sets of secrets, two runner configurations, two places to maintain whenever someone touches the build. It only pays off if step 1 is already done, otherwise you will be maintaining two diverging pipelines.

The first four steps, added together, come to less than a week of work and create nothing new to operate. That is as far as I would go in most of the projects that pass through my hands, and it is where I recommend the overwhelming majority of teams stop.

From the fifth step onwards, a secondary CI, Jenkins in high availability, runners in two clouds, the conversation stops being technical and becomes financial. The criterion is a single question, and it isn't mine to answer: how much does an hour without being able to deploy cost you?

For a SaaS that ships critical bug fixes several times a day, the number is high and step five pays for itself. For a product that does a planned weekly release, nine hours stopped cost a reshuffled schedule and nothing more. If you can't answer that question with a number, the next step isn't to build contingency: it is to find the number. Contingency sized by fear comes out far more expensive than contingency sized by arithmetic.

And to make it clear this isn't about an isolated event: GitHub's own incident history records six incidents touching Actions between 17 July and 7 August. Three weeks. That is the normal rate for any service at this scale, and that is what is worth planning around, not the nine-hour peak, which is the number that scares you and, for exactly that reason, the worst adviser.

Wrapping up

CI/CD contingency is usually discussed as a machine redundancy problem, and that is precisely the easy half. The half that bites is the one you don't see: the state.

If I had to leave this in four sentences:

  1. Your own runner is capacity, not independence. It gives you the processor, but whoever queues, dispatches, authenticates and reports is still on the other side. That is why both kinds of runner went down in the same minute.
  2. The incident doesn't end when the service comes back. An event that never became an execution doesn't recover by itself. Have a way to list what was left behind before you need it.
  3. The build lives in a script, CI is a thin shell. If switching CI is a project, you don't have a plan B, you have a rewrite. And a rewrite isn't done during an incident.
  4. The size of the contingency is the cost of the outage. Do the first four steps, which are cheap and which you use every day. From the fifth onwards, only once you have the number in hand.

Deep down it is the same discomfort I had already described when talking about mobile app releases: we never control when the other side updates, we only control what it receives. There the other side is the user who hasn't installed the new version; here it is the platform that dispatches your jobs. The object changes, the nature of the problem doesn't.

What this outage left me with wasn't an urge to leave GitHub. It was the realisation that I had thought through the whole contingency around where the build would run, and not a single minute around how I would find out what didn't run. The second part costs an afternoon of work. Naturally, it is the one I will use first next time.

And if you got this far without being able to say how much an hour without deploys costs in your case, that is exactly the conversation I like to have. Send me a message and we'll think it through together.

Have a system that needs to survive growth?

That's the kind of decision I help make. If it fits where you are, let's talk.