Metrics, logs and traces
The three signals every system gives off, the question each one answers, the four golden signals to watch first, and how to read all of them on your own machine.
By the end of this lesson you'll be able to
0 of 4 completeMonitoring and observability
Monitoring watches for problems you already know about: "alert me if the disk passes 90%" or "if the error rate goes above 1%". Observability is being able to answer questions you didn't plan for, like "why are only users on the Android app in Pune seeing timeouts since 14:02?", without shipping new code to find out.
Both are built on the same three kinds of data, often called the three signals. Each one is good at a different question.
Metrics
A metric is a number measured over time, with labels that say what it describes:
# HELP http_requests_total Total HTTP requests handled. # TYPE http_requests_total counter http_requests_total{method="GET",code="200"} 1027 http_requests_total{method="GET",code="500"} 3
That's the text format Prometheus reads, and you'll write queries against it in the next lesson. The main types:
| Type | Behaves like | Example |
|---|---|---|
| Counter | Only goes up (resets on restart) | Requests served, errors, bytes sent |
| Gauge | Goes up and down | Memory in use, queue length, temperature |
| Histogram | Counts values into buckets | Request duration, response size |
Metrics are small and cheap, so you can keep months of them, graph them and alert on them. What they can't do is tell you about one particular request.
Logs
A log is a record of one event, written by the app or the system when it happens. Plain text logs are easy to read by eye; structured logs (one JSON object per line) are easy for machines to search and filter:
{"time":"2026-03-14T14:02:11Z","level":"error","service":"checkout","msg":"payment declined","order":"A-1042","trace_id":"4bf92f3577b34da6a3ce929d0e0e4736"}
Logs carry the detail metrics leave out: the error message, the order, the user. The cost is volume: a busy service writes gigabytes a day, so logs are usually kept for days or weeks, not months.
Traces
A trace follows one request as it moves through your services. Each step is a span with a start time and a duration, and every span carries the same trace ID:
trace 4bf92f35… checkout request 1,840 ms ├─ api-gateway 12 ms ├─ cart-service 35 ms ├─ payment-service 1,742 ms ← here │ └─ POST bank-api/charge 1,701 ms └─ email-service (async) 22 ms
One look shows that the bank call is the slow part. Traces need the services to pass the trace
ID along with each call (usually in a traceparent HTTP header), which is what OpenTelemetry
sets up for you in the Tracing stage.
Which signal answers which question
| Question | Start with |
|---|---|
| Is something wrong right now? Since when? | Metrics |
| Which service or call is slow or failing? | Traces |
| What exactly happened to this request or user? | Logs |
Notice the trace_id in the log line above. Put the same ID in logs and traces and you can
jump from a slow trace straight to its logs. That link is what turns three separate tools into
observability.
The four golden signals
Google's SRE book says that if you can only measure four things about a user-facing system, measure these:
| Signal | Asks | Example |
|---|---|---|
| Latency | How long do requests take? | p95 checkout time is 480 ms |
| Traffic | How much demand is there? | 1,200 requests per second |
| Errors | How many requests fail? | 0.4% of requests return a 5xx |
| Saturation | How full is the system? | Disk 92% full, database pool 48 of 50 |
Track latency for successful and failed requests separately: a fast error page can make average latency look better while users are having a terrible time.
Read the signals on your own machine
You don't need a monitoring stack to start. Every Linux machine already exposes metrics and logs:
$ uptime # load average over 1, 5 and 15 minutes 14:05:31 up 12 days, 3:41, 1 user, load average: 0.42, 0.37, 0.30 $ free -m # memory, in MB: look at "available", not "free" $ df -h # disk use per filesystem: saturation, at a glance $ journalctl -p err -b -n 20 # logs: the last 20 error-level entries since boot $ curl -o /dev/null -s -w "%{http_code} %{time_total}\n" https://example.com 200 0.254813 # errors and latency for one request, in seconds
uptimeuptimehow long the machine is up, and its load
free -mfreememory use-min megabytes
df -hdfdisk space per filesystem-hhuman-readable sizes
journalctl -p err -b -n 20journalctlread the systemd journal (logs)-pminimum priority, like errerra value the command works on-bsince this boot-nhow many lines20a value the command works on
curl -o /dev/null -s -w "%{http_code} %{time_total}\n" https://example.comcurlmake an HTTP request-owrite the body to this file/dev/nulla path-ssilent: no progress bar-wprint these details after the request"%{http_code} %{time_total}\n"text, kept together by the quoteshttps://example.coma URL
A load average higher than your number of CPU cores (nproc) means work is queueing: that's
saturation too.
Measure the golden signals of a website
Before adding any tools, get a baseline for a site you depend on. Send it 10 requests, record the status code and time for each, then report how many failed and how slow the slowest one was.
~/golden.txtReveal one safe solution
$ for i in $(seq 1 10); do curl -o /dev/null -s -w "%{http_code} %{time_total}\n" https://example.com; done > ~/golden.txt $ cat ~/golden.txt 200 0.254813 200 0.198342 ... $ grep -vc '^200 ' ~/golden.txt # errors: lines that aren't a 200 0 $ sort -k2 -n ~/golden.txt | tail -n 1 # latency: the slowest request 200 0.412907
for i in $(seq 1 10); do curl -o /dev/null -s -w "%{http_code} %{time_total}\n" https://example.com; done > ~/golden.txtforloop over a listia value the command works onina value the command works on$(seq 1 10);the output of another commanddostart of what the loop runs each timecurlmake an HTTP request-owrite the body to this file/dev/nulla path-ssilent: no progress bar-wprint these details after the request"%{http_code} %{time_total}\n"text, kept together by the quoteshttps://example.com;a URLdoneend of the loop>write output to a file, replacing it~/golden.txta path
cat ~/golden.txtcatprint a file~/golden.txta path
grep -vc '^200 ' ~/golden.txtgrepfind lines that match a pattern-vshow lines that do not match-ccount matching lines'^200 'text, kept together by the quotes~/golden.txta path
sort -k2 -n ~/golden.txt | tail -n 1sortsort lines-k2an option-nsort as numbers~/golden.txta path|pipe: send output to the next commandtailshow the last lines-nhow many lines1a value the command works on
Ten requests is the traffic, the non-200 count is the error rate, and the slowest time is a rough worst-case latency. That's three of the four golden signals from one loop. Saturation needs a view from inside the server, which is what Prometheus and node_exporter give you next.
Try it yourself
0 of 5 steps doneRun each task in your own account or terminal, then tick it off.
Quick check
Users say checkout got slow, and checkout calls five services. Which signal tells you fastest which service is slow?
Official reading for this lesson
- GuideMonitoring distributed systems (the four golden signals)Google SRE book
- DocsObservability primerOpenTelemetry
- DocsPrometheus metric typesPrometheus
- DocsPrometheus exposition formatsPrometheus
- GuidePractical monitoring adviceGoogle SRE workbook
- Docsjournalctl(1): reading the systemd journalUbuntu manpages
- Docscurl --write-out variablesEverything curl
Or get every quiz answer right and it completes itself.