Lesson 1 of 5

What AIOps means

AIOps, short for artificial intelligence for IT operations, is the use of machine learning and other AI techniques to help run technology services. The term was introduced by the analyst firm Gartner around 2016; by 2017 Gartner was spelling it out as "artificial intelligence for IT operations". It caught on as monitoring tools began producing more alerts and data than people could handle.

AIOps is used for a handful of recurring jobs:

  • Noise reduction: grouping thousands of related alerts into a few meaningful incidents.
  • Anomaly detection: noticing when a measurement behaves unusually, even without a fixed threshold.
  • Correlation and probable cause: linking an incident to a recent change, a failing component or a similar past incident.
  • Forecasting: predicting when capacity will run out.
  • Assistance: helping operators search, summarise, explain and decide, increasingly using language models.
  • Automated response: triggering runbooks to fix known problems.

None of this replaces sound monitoring, clear ownership and good runbooks. AI works on top of them. Start by alerting on symptoms users feel, such as errors, slow responses and failed requests, rather than every internal cause, as Google's SRE guidance recommends; it gives AI far better input. An organisation with poor data and unclear processes will not be rescued by AI; it will get confident-looking answers built on bad foundations.

Lesson 2 of 5

The data it depends on

AI in operations is only as good as the data it learns from and reasons over.

  • Telemetry: metrics, logs and traces from the systems, collected consistently. Gaps and inconsistent formats limit what any model can do.
  • Topology: which services depend on which. Without it, a model cannot tell that fifty alerts share one cause.
  • Change records: what was changed, when and by whom. Many incidents follow a change (Google's Site Reliability Engineering book puts it at roughly 70% of outages), so this is one of the most useful signals.
  • Incident history: past incidents, their causes and their fixes, written clearly enough to learn from.
  • Knowledge: runbooks and documentation the AI can search and cite.

Three data problems recur:

  • Labels: to learn what an incident looks like, a model needs examples marked correctly. If past incidents were closed with "fixed" and no cause, there is little to learn from.
  • Drift: systems change, so a model trained on last year's behaviour may misread this year's. Models need retraining and re-checking.
  • Sensitive data: logs and tickets contain personal data, passwords and customer details. Decide what AI may see, strip what it should not, and apply data protection law as for any other processing. In Kenya that is the Data Protection Act, 2019, including its rules on transferring personal data outside Kenya, which can apply when logs are sent to an AI service hosted abroad.
Lesson 3 of 5

Cutting noise and spotting anomalies

Alert overload is the problem that made AIOps popular. Two techniques do most of the work.

Event correlation groups related alerts. It uses timing (alerts that arrive together), topology (alerts from components that depend on each other) and similarity (alerts that look alike). Good correlation turns an "alert storm" into one incident with a clear starting point. It has its own failure mode: merging two unrelated problems into one incident can hide the second, so check a sample of grouped incidents regularly.

Anomaly detection learns what normal looks like for each measurement, including daily and weekly patterns, and flags departures from it. It can catch problems that fixed thresholds miss, such as traffic falling to half its usual level for the time of day while still sitting inside a fixed limit. It also has well-known weaknesses:

  • Unusual is not the same as bad. A marketing campaign or month-end processing can look like an anomaly.
  • Too sensitive, and it adds noise instead of removing it.
  • It is hard to explain why some models flagged something, which makes people slow to trust them.
  • Slow drifts can be learned as normal. A gradual rise may be absorbed into the baseline and never flagged.

Judge these tools by outcomes: fewer alerts per real incident, incidents detected sooner, fewer missed incidents, fewer incidents wrongly merged, and less time spent by people triaging. Keep the old alerts running alongside the new approach until the evidence shows it works.

Lesson 4 of 5

Language models as operations assistants

Language models add a new kind of help: working with text and questions in plain language. In operations they can:

  • summarise a noisy incident channel for someone joining late;
  • explain an error message, a log pattern or an unfamiliar configuration;
  • search runbooks, past incidents and documentation, and point to the relevant one;
  • draft queries, scripts, status updates and post-incident reports for a person to check;
  • suggest likely causes and next steps, with the evidence behind them.

They also carry risks that matter more in operations than in most places:

  • Plausible but wrong answers (OWASP: Misinformation). A confident, incorrect suggestion during an outage can make it worse. Ask for sources, and treat suggestions as hypotheses to check.
  • Prompt injection (indirect prompt injection). Logs, tickets and web pages can contain text written to redirect the model. An attacker who can write to a log file may try to influence an AI that reads it. There is no complete fix, so assume any text the model reads could be followed as an instruction, and design so that this cannot cause harm: limit what the model can do, and require a person or fixed code to approve actions.
  • Excessive permissions (OWASP: Excessive Agency). An assistant that can run commands can run the wrong ones. Keep assistants read-only unless there is a strong, controlled reason not to.
  • Data exposure (OWASP: Sensitive Information Disclosure). Operational data sent to an outside AI service may include secrets and personal data. Use approved services and strip sensitive content.

The safest pattern is an assistant that reads, explains and proposes, while people and tested automation act.

Lesson 5 of 5

Controls before AI takes action

The step from AI that advises to AI that acts is the important one. Before any AI-driven action runs without a person, put these controls in place:

  • A clear boundary of what the AI may do alone: low-risk, reversible, well-tested actions only. Everything else goes to a person with the evidence.
  • Hard limits in ordinary code, outside the model, that block forbidden actions whatever the model suggests.
  • The same change rules as for people: every action recorded, reversible, and reviewed.
  • Evaluation before and after release: test the AI on past incidents, measure how often it would have been right, and re-test after every model or configuration change.
  • Monitoring of the AI itself: how often its suggestions are accepted, overruled or wrong, and whether that is changing.
  • A stop switch and a named owner.

General AI risk frameworks apply here too. The NIST AI Risk Management Framework organises the work into Govern, Map, Measure and Manage, and is a useful checklist for any organisation adopting AI in operations. Its companion Generative AI Profile (NIST AI 600-1, July 2024) covers risks specific to language models.

The aim is not maximum automation. It is faster, safer operations, where AI takes the tedious work and people stay in control of the decisions that matter.

Knowledge check

Ten questions

Answer all ten questions, then check your answers. You need 9 out of 10 to pass and receive a certificate. If you score less, you will see which answers were right and wrong, and then go through the course again before you retake the check. Your answers, progress and times are kept only in this browser.

Sources

The official documents this course relies on. Laws and guidance change, so check the current version.

  1. Site Reliability Engineering, chapter 6: Monitoring Distributed Systems · Google
  2. Site Reliability Engineering: Introduction · Google
  3. OWASP Top 10 for LLM Applications 2025 · OWASP Gen AI Security Project (OWASP Foundation)
  4. LLM01:2025 Prompt Injection · OWASP Gen AI Security Project
  5. AI Risk Management Framework (AI RMF 1.0) · US National Institute of Standards and Technology
  6. NIST AI 600-1: Artificial Intelligence Risk Management Framework, Generative AI Profile (2024) · US National Institute of Standards and Technology
  7. Data Protection Act, 2019 (No. 24 of 2019) · Kenya Law