Lesson 1 of 5

Services, not servers

People do not use servers, databases or networks. They use services: paying a supplier, checking a customer's identity, sending a statement, logging in to the payroll system. IT service management (ITSM) is the discipline of running technology so that those services are available, dependable and improving, and so that when something goes wrong it is handled in a predictable way.

The shift in thinking is from "is the server up?" to "can the people who depend on this service do their work?" A server can be running perfectly while the service it supports is unusable because a certificate has expired or a connection to another system has failed.

Two references are widely used:

  • ITIL, a widely used body of good-practice guidance for service management, owned by PeopleCert. ITIL 4 (2019) describes service management as a set of practices, such as incident, problem and change management, working together to create value. ITIL (Version 5), released from February 2026, keeps the practices and widens the scope to managing digital products and services. PeopleCert plans to retire the ITIL 4 modules at the end of 2027.
  • ISO/IEC 20000-1:2018 (amended in 2024), the international standard for a service management system, which organisations can be independently certified against.

You do not need to adopt either wholesale. A sensible place to start is a small set of practices, done well: a service desk, incident management, problem management and change control. The rest of this course covers those.

Lesson 2 of 5

The service desk and incidents

An incident is anything unplanned that stops a service working or makes it work worse than it should, such as a system that is down, slow or giving errors. Incident management exists to restore normal service as quickly as possible and limit the impact on the business.

The service desk is the single point of contact where users report problems and ask for help. Not every call is an incident. A service request is a user asking for something normal and expected, such as access, equipment or information. Handle requests through their own simple, often pre-approved, process and keep them out of incident figures.

A good incident process runs the same way every time:

  • Log every incident, however small, with who reported it, when and what they saw.
  • Categorise it, so trends can be seen later.
  • Prioritise it by impact (how many people or how important a service) and urgency (how quickly the harm grows). A payment system down for all customers at month-end outranks a printer fault.
  • Diagnose and escalate: to more specialised teams when the first line cannot fix it, and to management when the impact is serious.
  • Restore service, using a workaround if that is faster than a full fix.
  • Communicate: tell affected users what is happening and when to expect an update, even when there is no news.
  • Close once service is confirmed restored, ideally by the user; if they do not respond, close after an agreed period.

A major incident, one with severe business impact, needs its own procedure: a named incident manager, a dedicated channel, regular updates to leadership and customers, and a review afterwards.

Some incidents are also legal events. Under section 43 of Kenya's Data Protection Act, 2019, where personal data has been accessed or acquired by an unauthorised person and there is a real risk of harm to the people concerned, the data controller must notify the Data Commissioner without delay, within 72 hours of becoming aware of the breach. It must also tell the affected people in writing within a reasonably practical period. A data processor must tell the controller within 48 hours where reasonably practicable. So the incident process should bring in the data protection officer, or whoever handles data protection, early.

Lesson 3 of 5

Problems and root causes

Incident management restores service. Problem management asks why the incident happened and stops it happening again. A problem is whatever lies behind one or more incidents: the reason they happen, whether or not that reason has been found yet.

The two are separate for a reason. During an outage the priority is restoring service, not investigating. Once service is back, problem management takes over:

  • Find problems from major incidents, from repeated similar incidents and from trends, before users notice.
  • Investigate the root cause. Simple techniques go a long way: asking "why?" repeatedly until you reach a cause you can act on, or laying out a timeline of what changed and when.
  • Record known errors: problems you have diagnosed but not yet removed, with the workaround the service desk should use meanwhile.
  • Fix the cause through a controlled change, then confirm the incidents stop.

Reviews work best when they are blameless: they focus on how the system, process and information allowed the mistake, not on who made it. People who fear blame hide information, and hidden information is how the same incident happens twice.

A useful measure of maturity is how many incidents are repeats of something already seen. That number should fall over time.

Lesson 4 of 5

Changing systems safely

Many incidents follow a change: a new release, a configuration edit, a patch, a firewall rule. Change management, which ITIL 4 calls change enablement, exists to let changes happen quickly while keeping that risk under control.

Changes are usually handled in three ways:

  • Standard changes: low-risk, routine and done often, such as adding a user to a group. They are pre-approved and follow a documented procedure.
  • Normal changes: assessed for risk and approved by the right person or group before they are made. The more that could go wrong, the more senior or wider the review.
  • Emergency changes: needed urgently, often to fix a major incident. They follow a faster route but are still recorded and reviewed afterwards.

Every normal change should answer a few questions before it goes ahead:

  • What exactly is changing, and why?
  • What could go wrong, and who would be affected?
  • How was it tested?
  • How will we undo it if it fails, and has that been tried?
  • When is the least disruptive time, and does it clash with other changes or busy periods such as month-end?

Keep a change calendar so everyone can see what is planned, and a record of what each service depends on (systems, certificates, suppliers), so you can judge a change's impact and find an incident's cause faster. Measure results: how often changes cause incidents, and how long it takes to recover. Research by DORA (the DevOps Research and Assessment programme) found that teams making small, frequent, well-tested changes tend to have fewer failed changes and recover faster than teams making large, infrequent ones.

Lesson 5 of 5

Service levels and improvement

To know whether a service is good enough, you need to agree what "good enough" means.

  • A service level agreement (SLA) is a written commitment from a provider to a customer describing what the service will do and how well: for example availability during business hours, how quickly incidents of each priority will be responded to and resolved, and how performance will be reported.
  • Internal agreements between teams, often called operational level agreements, and contracts with suppliers must support the SLA. A promise to customers means little if a supplier it depends on has no matching obligation.

Good service targets are:

  • Meaningful to users: "customers can complete a payment" rather than "server CPU below 80%".
  • Measurable from data you actually collect.
  • Realistic, and agreed rather than imposed.
  • Reviewed regularly with the customer, not only when something breaks.

Measure a few things well: incident volumes and resolution times by priority, repeat incidents, changes that caused incidents, and user satisfaction. A knowledge base of fixes and workarounds, kept current by the service desk, lets more issues be solved at the first contact.

Finally, improve continually. Use the measures and the reviews of incidents and problems to choose a few improvements at a time, make them, and check that they worked. Service management is not a project with an end date; it is how good operations teams work every day.

Knowledge check

Ten questions

Answer all ten questions, then check your answers. You need 9 out of 10 to pass and receive a certificate. If you score less, you will see which answers were right and wrong, and then go through the course again before you retake the check. Your answers, progress and times are kept only in this browser.

Sources

The official documents this course relies on. Laws and guidance change, so check the current version.

  1. ITIL certifications (ITIL 4 and ITIL Version 5) · PeopleCert
  2. ITIL frequently asked questions · PeopleCert
  3. ISO/IEC 20000-1:2018 Information technology — Service management — Part 1: Service management system requirements (with Amd 1:2024) · International Organization for Standardization
  4. Data Protection Act, 2019 (No. 24 of 2019) · Kenya Law
  5. Capabilities: Working in small batches · DORA (Google Cloud)