DevOps.Academy
CoursesCurriculumPricing

    Loading…

    ↑ ↓ move · ↵ open · Ctrl K anywhere

    ◔My learningSign in
    Join the waitlist
    DevOps.Academy

    Free, hands-on DevOps lessons. Made with care in India.

    Learn

    All coursesLinux roadmapCurriculumMy learning

    Academy

    PricingCertificates

    Legal

    PrivacyTermsContact

    devops.rajeev.pro

    MonitoringLesson 1 of 9
    Your progress · 0%
    On this page
    1. Monitoring and observability
    2. Metrics
    3. Logs
    4. Traces
    5. The four golden signals
    6. Read the signals on your own machine
    ↖ Course roadmap
    Lesson 1 · Observability basics

    Metrics, logs and traces

    The three signals every system gives off, the question each one answers, the four golden signals to watch first, and how to read all of them on your own machine.

    20 min read · intermediate · hands-on

    By the end of this lesson you'll be able to

    0 of 4 complete

    Monitoring and observability

    Monitoring watches for problems you already know about: "alert me if the disk passes 90%" or "if the error rate goes above 1%". Observability is being able to answer questions you didn't plan for, like "why are only users on the Android app in Pune seeing timeouts since 14:02?", without shipping new code to find out.

    Both are built on the same three kinds of data, often called the three signals. Each one is good at a different question.

    Metrics tell you that something is wrong, traces tell you where, logs tell you why.© Diagram: DevOps Academy

    Metrics tell you that something is wrong, traces tell you where, logs tell you why.© Diagram: DevOps Academy

    Metrics

    A metric is a number measured over time, with labels that say what it describes:

    # HELP http_requests_total Total HTTP requests handled.
    # TYPE http_requests_total counter
    http_requests_total{method="GET",code="200"} 1027
    http_requests_total{method="GET",code="500"} 3
    

    That's the text format Prometheus reads, and you'll write queries against it in the next lesson. The main types:

    TypeBehaves likeExample
    CounterOnly goes up (resets on restart)Requests served, errors, bytes sent
    GaugeGoes up and downMemory in use, queue length, temperature
    HistogramCounts values into bucketsRequest duration, response size

    Metrics are small and cheap, so you can keep months of them, graph them and alert on them. What they can't do is tell you about one particular request.

    Logs

    A log is a record of one event, written by the app or the system when it happens. Plain text logs are easy to read by eye; structured logs (one JSON object per line) are easy for machines to search and filter:

    {"time":"2026-03-14T14:02:11Z","level":"error","service":"checkout","msg":"payment declined","order":"A-1042","trace_id":"4bf92f3577b34da6a3ce929d0e0e4736"}
    

    Logs carry the detail metrics leave out: the error message, the order, the user. The cost is volume: a busy service writes gigabytes a day, so logs are usually kept for days or weeks, not months.

    Traces

    A trace follows one request as it moves through your services. Each step is a span with a start time and a duration, and every span carries the same trace ID:

    trace 4bf92f35…   checkout request                       1,840 ms
    ├─ api-gateway                                              12 ms
    ├─ cart-service                                             35 ms
    ├─ payment-service                                       1,742 ms   ← here
    │  └─ POST bank-api/charge                               1,701 ms
    └─ email-service (async)                                    22 ms
    

    One look shows that the bank call is the slow part. Traces need the services to pass the trace ID along with each call (usually in a traceparent HTTP header), which is what OpenTelemetry sets up for you in the Tracing stage.

    Which signal answers which question

    QuestionStart with
    Is something wrong right now? Since when?Metrics
    Which service or call is slow or failing?Traces
    What exactly happened to this request or user?Logs

    Notice the trace_id in the log line above. Put the same ID in logs and traces and you can jump from a slow trace straight to its logs. That link is what turns three separate tools into observability.

    The four golden signals

    Google's SRE book says that if you can only measure four things about a user-facing system, measure these:

    SignalAsksExample
    LatencyHow long do requests take?p95 checkout time is 480 ms
    TrafficHow much demand is there?1,200 requests per second
    ErrorsHow many requests fail?0.4% of requests return a 5xx
    SaturationHow full is the system?Disk 92% full, database pool 48 of 50

    Track latency for successful and failed requests separately: a fast error page can make average latency look better while users are having a terrible time.

    Averages hide pain

    If 99 requests take 100 ms and one takes 10 seconds, the average is about 200 ms and looks fine. Use percentiles: p95 or p99 is the time that 95% or 99% of requests finish within, and that's where unhappy users show up.

    Read the signals on your own machine

    You don't need a monitoring stack to start. Every Linux machine already exposes metrics and logs:

    $ uptime                         # load average over 1, 5 and 15 minutes
     14:05:31 up 12 days,  3:41,  1 user,  load average: 0.42, 0.37, 0.30
    $ free -m                        # memory, in MB: look at "available", not "free"
    $ df -h                          # disk use per filesystem: saturation, at a glance
    $ journalctl -p err -b -n 20     # logs: the last 20 error-level entries since boot
    $ curl -o /dev/null -s -w "%{http_code} %{time_total}\n" https://example.com
    200 0.254813                     # errors and latency for one request, in seconds
    
    uptime
    • uptimehow long the machine is up, and its load
    free -m
    • freememory use
    • -min megabytes
    df -h
    • dfdisk space per filesystem
    • -hhuman-readable sizes
    journalctl -p err -b -n 20
    • journalctlread the systemd journal (logs)
    • -pminimum priority, like err
    • erra value the command works on
    • -bsince this boot
    • -nhow many lines
    • 20a value the command works on
    curl -o /dev/null -s -w "%{http_code} %{time_total}\n" https://example.com
    • curlmake an HTTP request
    • -owrite the body to this file
    • /dev/nulla path
    • -ssilent: no progress bar
    • -wprint these details after the request
    • "%{http_code} %{time_total}\n"text, kept together by the quotes
    • https://example.coma URL

    A load average higher than your number of CPU cores (nproc) means work is queueing: that's saturation too.

    Never log secrets

    Passwords, tokens, card numbers and full personal details must never reach a log. Logs get copied to many systems and kept for weeks, and many more people can read them than your database.

    In production

    Alert on what users feel (latency and errors), not on every CPU spike. A page at 3 a.m. should always mean "users are affected and a human needs to act". Everything else is a dashboard or a ticket.

    ⌘
    Mini mission

    Measure the golden signals of a website

    Before adding any tools, get a baseline for a site you depend on. Send it 10 requests, record the status code and time for each, then report how many failed and how slow the slowest one was.

    Input: https://example.com (or any site you use)Output: ~/golden.txt
    Reveal one safe solution
    $ for i in $(seq 1 10); do curl -o /dev/null -s -w "%{http_code} %{time_total}\n" https://example.com; done > ~/golden.txt
    $ cat ~/golden.txt
    200 0.254813
    200 0.198342
    ...
    $ grep -vc '^200 ' ~/golden.txt            # errors: lines that aren't a 200
    0
    $ sort -k2 -n ~/golden.txt | tail -n 1     # latency: the slowest request
    200 0.412907
    
    for i in $(seq 1 10); do curl -o /dev/null -s -w "%{http_code} %{time_total}\n" https://example.com; done > ~/golden.txt
    • forloop over a list
    • ia value the command works on
    • ina value the command works on
    • $(seq 1 10);the output of another command
    • dostart of what the loop runs each time
    • curlmake an HTTP request
    • -owrite the body to this file
    • /dev/nulla path
    • -ssilent: no progress bar
    • -wprint these details after the request
    • "%{http_code} %{time_total}\n"text, kept together by the quotes
    • https://example.com;a URL
    • doneend of the loop
    • >write output to a file, replacing it
    • ~/golden.txta path
    cat ~/golden.txt
    • catprint a file
    • ~/golden.txta path
    grep -vc '^200 ' ~/golden.txt
    • grepfind lines that match a pattern
    • -vshow lines that do not match
    • -ccount matching lines
    • '^200 'text, kept together by the quotes
    • ~/golden.txta path
    sort -k2 -n ~/golden.txt | tail -n 1
    • sortsort lines
    • -k2an option
    • -nsort as numbers
    • ~/golden.txta path
    • |pipe: send output to the next command
    • tailshow the last lines
    • -nhow many lines
    • 1a value the command works on

    Ten requests is the traffic, the non-200 count is the error rate, and the slowest time is a rough worst-case latency. That's three of the four golden signals from one loop. Saturation needs a view from inside the server, which is what Prometheus and node_exporter give you next.

    Try it yourself

    0 of 5 steps done

    Run each task in your own account or terminal, then tick it off.

    Knowledge check

    Quick check

    Question 1 of 4

    Users say checkout got slow, and checkout calls five services. Which signal tells you fastest which service is slow?

    Go deeper

    Official reading for this lesson

    • GuideMonitoring distributed systems (the four golden signals)Google SRE book ↗
    • DocsObservability primerOpenTelemetry ↗
    • DocsPrometheus metric typesPrometheus ↗
    • DocsPrometheus exposition formatsPrometheus ↗
    • GuidePractical monitoring adviceGoogle SRE workbook ↗
    • Docsjournalctl(1): reading the systemd journalUbuntu manpages ↗
    • Docscurl --write-out variablesEverything curl ↗

    Or get every quiz answer right and it completes itself.

    Next lesson →Prometheus