Declaring an incident and setting severity
Most problems are handled through normal support. A major incident is different: an event with serious impact that needs a coordinated response, now. The first decision is to recognise it and say so.
Organisations define severity levels in advance, so that nobody has to invent them during a crisis. A typical scheme has three to five levels, set by impact, for example:
- Severity 1: a critical service is down or badly degraded for many customers, money or data are at risk, or there is a legal or safety issue.
- Severity 2: a significant service is degraded, or a critical one is at risk of failing.
- Severity 3 and below: limited impact, handled through normal support.
Good practice for declaring incidents:
- Anyone can declare one. The person who notices should not need permission to raise the alarm.
- Declare early. It is cheap to stand down a response that turned out not to be needed, and expensive to start one late.
- Severity can change. Raise or lower it as you learn more.
- The severity sets the response: who is called, how often updates go out and who else must be told.
Prepare before the incident: keep an up-to-date on-call and escalation contact list, including suppliers and cloud and telecoms providers, and run practice drills. Many real incidents start at a third party, and the first delay is often not knowing whom to call.
Some incidents are also security or data protection events. Decide early whether personal data or security are involved, because that brings in other people and, in some cases, legal deadlines.
Roles that keep a response organised
In a serious incident, the biggest risk is often confusion: several people making changes at once, nobody sure who decides, customers told nothing. Clear roles prevent that. The idea comes from the Incident Command System, developed by US fire agencies in the 1970s and now used widely by emergency services, which gives command, operations, planning and public information to separate, named people.
- Incident commander (or incident manager): coordinates the response, sets priorities and makes decisions. They do not usually fix things themselves; their job is to keep the whole picture.
- Operations or technical lead: directs the people investigating and fixing the problem, and makes sure changes are made one at a time and agreed.
- Communications lead: keeps customers, leadership and other teams informed on a regular schedule.
- Scribe or planning lead: records the timeline (what was seen, tried and decided, and when) and handles longer-running tasks such as arranging handovers and follow-up tickets.
- Subject experts, brought in as needed.
In a small team one person may hold several roles, but the roles still need to be named. Other habits help:
- One place to talk: a single incident channel or bridge, so information is not split.
- Hand over explicitly when someone steps out of a role, especially during long incidents.
- Rest: long incidents need shifts, because tired people make mistakes.
Communicating under pressure
During an incident, silence tends to be read as confusion or lack of control, and people fill the gap with guesses. Regular, honest communication keeps trust, even when there is little good news.
- Tell people early that you know about the problem and are working on it.
- Update on a schedule, for example every 30 minutes for the most severe incidents, even if the update is "no change; next update at 14:30".
- Say what you know, what you do not know and what you are doing. Avoid guessing at causes in public.
- Separate audiences: customers need impact and next steps; leadership needs impact, risk and decisions required; engineers need technical detail. One message rarely serves all three.
- Use a status page or agreed channel customers can check, so the service desk is not overwhelmed.
- Confirm the end: announce when service is restored, and what customers should do, if anything.
Some communication is required by law or regulation. In Kenya, for example:
- Personal data breaches. Under the Data Protection Act, 2019 (section 43), where a breach creates a real risk of harm to the people affected, the data controller must notify the Data Commissioner without delay, and within 72 hours of becoming aware of it. The controller must also tell the affected people in writing within a reasonably practical period, with limited exceptions. A data processor, such as an outsourced IT or hosting provider, must tell the controller without delay, and where reasonably practicable within 48 hours.
- Cyber attacks. The Computer Misuse and Cybercrimes Act, 2018 requires owners of designated critical information infrastructure to report threatening incidents to the National Computer and Cybercrimes Co-ordination Committee (section 11).
- Sector rules. The Central Bank of Kenya's 2017 Guidance Note on Cybersecurity asks banks to notify it within 24 hours of a cyber incident that could significantly affect their services, reputation or finances.
Build these into the incident process, with named people responsible, so they are not forgotten under pressure.
Restoring service safely
The first goal of incident response is to restore service, not to find the root cause. Investigation can wait until customers are working again.
- Mitigate first. Roll back a recent change, switch to a standby, remove a failing component, add capacity, or turn off a feature, whichever restores service fastest with least risk. If a security incident is possible, check with the security lead before wiping, rebuilding or rolling back affected systems, as this can destroy evidence.
- Suspect recent changes. Many incidents follow a change, so "what changed?" is one of the most useful first questions.
- One change at a time, agreed through the operations lead, so you know what worked and do not make things worse.
- Record everything in the timeline as you go, including things that did not work.
- Protect evidence where a security incident is possible, so it can be investigated later.
- Confirm recovery from the user's side, not only from dashboards, before standing down.
Then stand down deliberately: confirm service is stable, tell everyone, hand any remaining work to normal support, and schedule the review while memories are fresh. Deeper root-cause work, where needed, continues as a tracked problem record or review action: the investigation waits, but it is not optional.
Learning without blame
A post-incident review (often called a postmortem) turns an incident into improvement. The best-known approach, described in Google's Site Reliability Engineering book, is blameless: it assumes people acted reasonably given what they knew at the time, and asks how the system, tools and processes allowed the failure.
A blameless review still expects people to own their actions and the follow-up work. The reason for it is practical: people who fear blame hide information, and without that information the same incident happens again.
A good review covers:
- What happened: an accurate timeline, including when it started, when it was detected and when service was restored, all in one time zone. The delay before detection is often the most useful measure a review produces.
- Impact: who was affected, how, and for how long.
- Why: the contributing factors, which are usually several, not one "root cause". Ask why the problem happened, why it was not caught sooner and why recovery took as long as it did.
- What went well, so it can be kept.
- Actions: specific improvements, each with an owner and a due date.
The review is only as good as its follow-through. Track the actions like any other work, check that they were done, and look across reviews for patterns: the same weak spot appearing in several incidents is the most valuable finding of all. Share reviews widely, so other teams learn without having the same incident.
Ten questions
Answer all ten questions, then check your answers. You need 9 out of 10 to pass and receive a certificate. If you score less, you will see which answers were right and wrong, and then go through the course again before you retake the check. Your answers, progress and times are kept only in this browser.
Your answers
Your certificate of completion
Enter your name as you want it to appear, then save the certificate as a PDF. In the print window, choose Save as PDF. A certificate is issued once per completion of the course.
Saolix does not record who takes this course, so it cannot verify these certificates. The certificate confirms completion of a free self-paced course and is not an accredited qualification.
Sources
The official documents this course relies on. Laws and guidance change, so check the current version.
- Site Reliability Engineering, chapter 14: Managing Incidents · Google
- Site Reliability Engineering, chapter 15: Postmortem Culture: Learning from Failure · Google
- NIST SP 800-61 Rev. 3: Incident Response Recommendations and Considerations for Cybersecurity Risk Management: A CSF 2.0 Community Profile (April 2025), for security incidents · US National Institute of Standards and Technology
- Data Protection Act, 2019 (No. 24 of 2019) · Kenya Law
- Guidance Note on Cybersecurity for the Banking Sector (2017) · Central Bank of Kenya
- Computer Misuse and Cybercrimes Act, 2018 (No. 5 of 2018) · Kenya Law
- ICS Review Document (the Incident Command System) · US Federal Emergency Management Agency