InfrastructureLesson 3 of 49 min

Cron and scheduled jobs

Work that happens because of the clock, and the four ways it quietly fails.

A cron job is work triggered by time rather than by a user: send the digest at 8am, expire trials nightly, rebuild the search index every hour. The name comes from a Unix utility from the 1970s and the scheduling syntax has barely changed since.

The schedule is five fields (minute, hour, day of month, month, day of week) where an asterisk means "every". "0 8 * * *" is every day at 08:00. You do not need to memorise this; you need to recognise it and know to double-check it, because a misplaced field is the difference between daily and every minute.

The four classic cron failures
  1. Timezones
    Servers usually run in UTC. "8am" for whom? Daylight saving moves it twice a year.
  2. Overlap
    An hourly job that takes 70 minutes will start again before finishing. Now two copies are running.
  3. Missed runs
    If the machine was down at 8am, most schedulers do not run it late. The day is simply skipped.
  4. Growing work
    "Process yesterday’s signups" is fine at 10 a day and impossible at 100,000.

What happens if this runs twice, or not at all today?

Both will happen eventually. If either answer is "something bad and irreversible", the job needs redesigning before it needs scheduling.

What to remember

  • Cron triggers work by clock, and nobody notices when it stops.
  • Timezones, overlap, missed runs, and growth are the standard failures.
  • Alert on missing success, not on failure.

Terms in this lesson

Field notes

Loaded from a deliberately slow source. The lesson above was already readable while this was still travelling. That is streaming, and it is the same trick a chat interface uses.

The digest that stopped for five weeks

A weekly email job silently stopped running after an infrastructure change. Nobody noticed, because a job that does not run produces no error. It was discovered when a customer asked whether they had been unsubscribed.

Why you alert on missing success

Everyone retried at once

A service had a brief wobble. Every client retried immediately, then again, then again. The retries were far more traffic than the original load, and the service never got a quiet moment to recover. The outage lasted forty minutes longer than the fault did.

The thundering herd

resolved in 900ms · region iad1

Hide field notes toggles a search param the loader reads. With it off, the slow promise is never created, so nothing streams.