Lesson 1 of 5

Monitoring and observability

Monitoring means watching for conditions you already know matter: is the service up, is the disk nearly full, are errors above normal? You decide in advance what to check, and you are told when a check fails.

Observability is the ability to work out what is happening inside a system from what it tells you from the outside, including problems nobody predicted. A system is observable when, faced with a new kind of failure, an engineer can ask fresh questions of its data and get answers without first adding new instrumentation or deploying new code.

The two work together:

  • Monitoring checks the questions you thought of in advance and tells you quickly when something is wrong, ideally before users notice.
  • Observability lets you explore questions you did not think of in advance, which matters most in systems made of many services, where one user action can pass through a dozen components.

Neither is a product you buy once. Both depend on systems being built to report useful information about themselves, a practice known as instrumentation, and on people knowing how to use it. A dashboard nobody looks at, or an alert everybody ignores, adds cost without adding safety.

Lesson 2 of 5

Metrics, logs and traces

Systems report on themselves in three main forms, each suited to different questions.

  • Metrics are numbers measured over time: requests per second, error count, response time, memory used. They are usually cheap to store and fast to query, so they suit dashboards and alerts. Avoid labelling them with values that have very many possibilities, such as user or request IDs, or they become slow and costly. They tell you that the error rate jumped at 14:05, but not which request failed or why.
  • Logs are records of individual events, written as they happen: "payment 4821 rejected: card expired". They carry detail and context. Write them in a structured form, with named fields such as time, service, request identifier and outcome, so they can be searched and counted rather than read line by line.
  • Traces follow a single request as it passes through several services, recording each step and how long it took. When a page is slow, a trace shows which of the ten services it touched was responsible. Tracing every request is usually too costly, so most systems keep a sample of traces, often keeping all slow or failed ones.

The three are most useful when they are connected. A common identifier carried with each request, usually called a trace ID, lets you move from a spike on a metrics chart, to the traces of slow requests at that moment, to the log lines those requests wrote.

OpenTelemetry, an open-source project that became a graduated project of the Cloud Native Computing Foundation in May 2026, provides a common, vendor-neutral way to produce, collect and send traces, metrics and logs. Support for traces and metrics is stable in most programming languages; support for logs is stable in some and still maturing in others. Instrumenting with an open standard makes it easier to change monitoring tools later without rewriting your code.

Lesson 3 of 5

What to measure

Measuring everything creates noise. A few well-known approaches help you choose.

  • The four golden signals, from the chapter "Monitoring Distributed Systems" in Google's Site Reliability Engineering book: latency (how long requests take), traffic (how much demand there is), errors (the rate of requests that fail, including those that return a wrong result or break a rule you have set, such as being too slow) and saturation (how full the system is, especially its most constrained resource, such as a queue or a disk).
  • The RED method, proposed by Tom Wilkie, for each service: the rate of requests, the number of errors and the duration of requests, looked at as a distribution rather than just an average.
  • The USE method, devised by Brendan Gregg, for each resource such as a processor, disk or network link: its utilisation (how busy it is), its saturation (how much work is queued waiting for it) and its errors.

Two further habits keep measurement honest:

  • Measure from the user's side. A server can report success while users see failures because of a network, a certificate or a mobile app. Synthetic checks, scripted tests that act like a user at regular intervals, and real-user monitoring, measurements taken from actual users' sessions, show the service as people experience it. Checks from the outside, as a user would see the service, are called black-box monitoring; measurements the system reports about its own internals are white-box monitoring. You need both.
  • Look at the slow tail, not only the average. An average response time of 300 milliseconds can hide the fact that one request in a hundred takes eight seconds. Percentiles, such as the time within which 99% of requests complete, show what the unluckiest users experience.
Lesson 4 of 5

Objectives users care about

To decide whether a service is healthy, agree what "healthy" means in terms users would recognise. Site reliability engineering uses three linked ideas:

  • A service level indicator (SLI) is a measurement of one aspect of the service that users care about, for example the proportion of payment requests that succeed, or the proportion that complete within one second.
  • A service level objective (SLO) is the target for that indicator over a period, for example 99.9% of payment requests succeed over 30 days.
  • The error budget is the gap between the objective and perfection. An objective of 99.9% allows 0.1% of requests to fail in the period. While budget remains, the team can take reasonable risks, such as releasing changes; when it is used up, the priority shifts to reliability.

Good objectives are:

  • Set from the user's point of view: requests that succeed and are fast enough, not servers that are up.
  • Below 100%. No service is perfect, and chasing perfection makes every change frightening and costs far more than users notice.
  • Tied to decisions: an objective that nobody acts on is just a number.

Objectives also make service level agreements (SLAs), the contracts with customers that carry consequences if missed, safer to offer. Set the internal objective tighter than the external promise, so the team is warned before the agreement is at risk.

Lesson 5 of 5

Alerts people trust

An alert should mean that a person needs to act, now or soon. Too many alerts that do not need action, often called alert fatigue, teach people to ignore them, and then the important one is missed.

Good alerting follows a few rules:

  • Alert on symptoms users feel, such as failing or slow requests, or an error budget being used up fast, rather than on every cause, such as one busy processor. The speed at which the budget is being used is called the burn rate: a burn rate of 1 uses up exactly the whole budget by the end of the period, so alert when it stays well above 1.
  • Every alert must be actionable, with a clear owner and a short guide, often called a runbook, saying what to check first.
  • Match the urgency to the impact. Wake someone at night only for something that cannot wait until morning; send the rest to a ticket queue.
  • Review alerts regularly. Remove or fix any alert that fired without needing action, and add one whenever an incident was found by users rather than by monitoring.

Monitoring data needs care too. Logs and traces often contain personal data: names, phone numbers, account numbers, locations. Collect only what you need, mask or leave out sensitive fields, limit who can see the data and delete it when it is no longer needed. In Kenya, the Data Protection Act, 2019 covers personal data in logs just as it covers any other record. Section 25 requires personal data to be limited to what is necessary, and not kept in a form that identifies people for longer than needed. Section 39 limits how long it may be retained. Section 41 requires data protection by design and by default, with measures such as pseudonymisation and encryption.

Knowledge check

Ten questions

Answer all ten questions, then check your answers. You need 9 out of 10 to pass and receive a certificate. If you score less, you will see which answers were right and wrong, and then go through the course again before you retake the check. Your answers, progress and times are kept only in this browser.

Sources

The official documents this course relies on. Laws and guidance change, so check the current version.

  1. Site Reliability Engineering (the book, free online) · Google
  2. Site Reliability Engineering, chapter 6: Monitoring Distributed Systems · Google
  3. Site Reliability Engineering, chapter 4: Service Level Objectives · Google
  4. The Site Reliability Workbook, chapter 5: Alerting on SLOs · Google
  5. The RED Method: How to Instrument Your Services (Tom Wilkie) · Grafana Labs
  6. The USE Method (Brendan Gregg) · brendangregg.com
  7. OpenTelemetry documentation · OpenTelemetry (Cloud Native Computing Foundation)
  8. Cloud Native Computing Foundation announces OpenTelemetry's graduation (21 May 2026) · Cloud Native Computing Foundation
  9. OpenTelemetry status (maturity of traces, metrics and logs) · OpenTelemetry
  10. Data Protection Act, 2019 (No. 24 of 2019) · Kenya Law