learninfra · Linux · Networking · Kubernetes · System Design · AI Infrastructure · Exam blueprints · Drills

Observability: Metrics, Logs & Traces

Finding the slow hop in a request that crossed eight services: traces and their critical path, RED metrics and the percentiles averages hide, alerts on symptoms rather than causes, sampling, cardinality, and logs that join up by trace id.

An interactive System Design lesson: 22 steps, about 34 minutes, on a live simulation in your browser.

A customer opens an order page. Behind GET /api/orders/{id} the gateway asks auth to check the session (auth asks redis), then asks orders, which reads orders-db, fetches the customer from users and stock from inventory. The page comes back 200 after 77.3 ms.

If it had taken 3 seconds, which of those services would you blame? Each one's logs show only its own part. A distributed trace shows the whole request: every call is a span (who, what, when it started, how long), all spans share one trace id, and each span names its parent. Drawn as a waterfall, the request is a tree in time.

What you will learn

  1. One request, eight services

    • One request, nine spans: A trace is a tree of spans that share one trace id; each span is one call, with its parent, start and duration.
    • How the trace id travels: Context propagation is the whole trick: each service reads traceparent on the way in and writes its own span id into it on the way out.
    • Where speed-ups pay: Only the critical path decides latency: optimise the longest chain of spans, not the service that looks slowest in isolation.
  2. What only a trace shows

    • N+1 calls: N+1 is a staircase in the waterfall: many small, fast, identical calls in a row that no per-service metric flags.
    • One call for all twenty: Sequential calls add up their round trips; batching N calls into one pays the round trip once.
  3. Metrics: averages lie

    • Rate, errors, duration: RED per service (rate, errors, duration percentiles) answers "is it healthy right now?"; traces answer "why was this one slow?".
    • Two percent of payments: Averages hide tails: watch p99 (and p99.9 at scale), because the slowest 1% is where users and timeouts live.
    • Symptom alerts and cause alerts: Page on symptoms users feel (SLO burn rate), not on causes inside the system; causes go on dashboards to explain the page.
    • Asking the histogram: A percentile is computed from summed buckets, never by averaging percentiles.
    • Drill: p99 in PromQL
  4. Which traces you keep

    • Keep 1% of traces: Head sampling decides before the outcome is known, so it keeps the same 1% of failures as of successes: almost none of the traces you need.
    • Decide at the end: Tail sampling keeps the interesting traces (errors, slow ones) because it decides after seeing the outcome, at the price of buffering everything first.
  5. The label that costs a fortune

    • Add user_id to the metric: Series count is the product of every label's distinct values: one unbounded label (user id, request id, URL) multiplies the whole metric.
    • High-cardinality data goes elsewhere: Metrics are for aggregates with bounded labels; per-user and per-request detail belongs in logs and traces, joined by trace id.
  6. A page at 3 a.m.

    • Recommendations slow down: A timeout with a fallback turns a slow dependency into a degraded feature: the cause alert fires, the symptom stays quiet, and that is the system working as designed.
    • A p99 that never happened: histogram_quantile interpolates inside a bucket, so a p99 can be any value up to the bucket's top: precision comes from where the bucket boundaries are.
    • The work nobody waits for: A timeout protects the caller; only propagated deadlines and cancellation stop the callee from working for nobody.
    • Send the deadline with the call: Propagate deadlines and honour cancellation: work for a caller that has given up should stop when the caller does.
  7. Logs that join up

    • Every log line by trace_id: Put trace_id in every log line: metrics say something is wrong, traces say where, logs say what, and the trace id joins them.
    • Drill: find the slow calls
  8. Recap & playground

    • Cheat sheet
    • Playground