InfrastructureLesson 3 of 48 min

Logs, metrics, and traces

Three tools that answer three different questions. Most people only have one.

When something goes wrong, the difference between a five minute fix and a five hour one is whether you can see what happened. Three kinds of visibility exist, and they are not substitutes for each other.

Logs, what happened

  • Individual events with detail
  • Great for one specific failure
  • Terrible for spotting trends
  • Expensive at volume

Metrics, how much, how often

  • Numbers over time: requests, errors, duration
  • Great for "is it worse than yesterday"
  • Cheap to keep for a long time
  • Cannot tell you about one user

The third is tracing: following one request across every service it touches, with timings for each hop. When a page is slow and four systems are involved, a trace tells you which one in seconds. Without it you are guessing across four sets of logs with mismatched clocks.

If this broke right now, what would I actually look at?

If the honest answer is "I would redeploy and hope", that is worth fixing on a calm day rather than discovering on a bad one.

What to remember

  • Logs are events, metrics are trends, traces are one request across systems.
  • A shared request id makes logs searchable instead of scattered.
  • Never log raw request bodies.

Terms in this lesson

Show field notes toggles a search param the loader reads. With it off, the slow promise is never created, so nothing streams.