InfrastructureLesson 4 of 48 min

You are the on-call rota

There is no handover, so every decision about responding is made once and applies always.

Companies pay people to be woken up, in a rota, so that no individual carries every hour. You have the same responsibility and a rota of one. Whatever you decide about answering, you are deciding it for every hour of every week, including the ones you intend to spend asleep.

Worth getting up for

  • Money is moving incorrectly, or not at all
  • Data is being lost or exposed
  • Nobody can log in
  • It is getting worse on its own

Genuinely fine until morning

  • One user hit one strange edge case
  • A page is slow but working
  • Something looks wrong but nothing is wrong
  • You have already stopped the bleeding

Then there is the data. A backup you have never restored is not a backup, it is a belief. Restore one into a scratch environment on a calm afternoon, once. You will find out whether it works, and you will learn how long it takes, which is precisely the number you will want to quote to somebody while it matters.

Real users also bring obligations that have nothing to do with code. Somebody will ask you to delete their account and everything attached to it. If the honest answer is that you do not know everywhere their data went, that is worth fixing now, while there are eleven users and three tables, rather than later, under a deadline, with a lawyer reading over your shoulder.

If this broke at 2am on a Sunday, what have I already decided to do about it?

"Nothing until Monday" is a perfectly respectable answer when it is chosen in advance and published. It is a much worse one when it is discovered in the moment by a user who expected otherwise.

What to remember

  • Write down what you promise to respond to, and publish it.
  • An untested backup is a belief; restore one on a calm day to find out.
  • When nothing was deployed, look outward: expiries, upstreams and quotas.

Terms in this lesson

Field notes

Loaded from a deliberately slow source. The lesson above was already readable while this was still travelling. That is streaming, and it is the same trick a chat interface uses.

The digest that stopped for five weeks

A weekly email job silently stopped running after an infrastructure change. Nobody noticed, because a job that does not run produces no error. It was discovered when a customer asked whether they had been unsubscribed.

Why you alert on missing success

Everyone retried at once

A service had a brief wobble. Every client retried immediately, then again, then again. The retries were far more traffic than the original load, and the service never got a quiet moment to recover. The outage lasted forty minutes longer than the fault did.

The thundering herd

resolved in 899ms · region iad1

Hide field notes toggles a search param the loader reads. With it off, the slow promise is never created, so nothing streams.