Logs, metrics, and traces
Three tools that answer three different questions. Most people only have one.
When something goes wrong, the difference between a five minute fix and a five hour one is whether you can see what happened. Three kinds of visibility exist, and they are not substitutes for each other.
Logs, what happened
- Individual events with detail
- Great for one specific failure
- Terrible for spotting trends
- Expensive at volume
Metrics, how much, how often
- Numbers over time: requests, errors, duration
- Great for "is it worse than yesterday"
- Cheap to keep for a long time
- Cannot tell you about one user
The third is tracing: following one request across every service it touches, with timings for each hop. When a page is slow and four systems are involved, a trace tells you which one in seconds. Without it you are guessing across four sets of logs with mismatched clocks.
If this broke right now, what would I actually look at?
If the honest answer is "I would redeploy and hope", that is worth fixing on a calm day rather than discovering on a bad one.
What to remember
- Logs are events, metrics are trends, traces are one request across systems.
- A shared request id makes logs searchable instead of scattered.
- Never log raw request bodies.
Terms in this lesson
Show field notes toggles a search param the loader reads. With it off, the slow promise is never created, so nothing streams.