What zero-touch operations means
Zero-touch operations is the aim of running technology services so that routine work happens automatically from start to finish, without a person having to touch it. Problems are detected, diagnosed, fixed and checked by automation; people design the system, handle the exceptions and make the decisions that need judgement.
The idea took its best-known form in telecommunications, where networks became too large and changed too quickly to manage by hand, and it builds on earlier research into self-managing ("autonomic") computer systems. In December 2017 the European Telecommunications Standards Institute (ETSI) set up an Industry Specification Group on Zero-touch network and Service Management (ZSM). Its aim is an end-to-end framework in which operational tasks such as delivery, deployment, configuration, assurance and optimisation run automatically, ideally without any manual step. The same thinking now applies to any technology operation: data centres, cloud platforms, applications and service desks.
"Zero-touch" is a direction, not a switch. Few organisations remove people entirely, and in regulated or high-risk work they should not. The practical goal is:
- No manual touch for routine, well-understood work, such as restarting a failed service, adding capacity or resetting an account.
- Fewer, better-informed human decisions for everything else, with the automation gathering the evidence and proposing the action.
- Every action recorded, whether a machine or a person took it.
The building blocks
Zero-touch operations is built as a closed loop: observe, decide, act, then check that the action worked, and repeat. Engineers describe this loop in several ways; a widely cited reference model from autonomic computing is MAPE-K: monitor, analyse, plan and execute, using shared knowledge. This course uses four plain steps: observe, decide (which includes diagnosing the cause), act and verify. Each part of the loop needs its own foundation.
- Observe: good monitoring and observability, so the system knows what is happening. Automation cannot fix what it cannot see.
- Decide: rules, policies and, increasingly, statistical or machine learning models that recognise known situations and choose a response. Clear policies say what the automation is allowed to do and when.
- Act: runbook automation, where the steps an engineer would follow are written as code; infrastructure as code, where servers, networks and settings are defined in files and built automatically; and interfaces (APIs) that let tools make changes safely.
- Verify: after acting, check that the service is healthy again. If it is not, undo the change or hand the problem to a person with everything gathered so far.
Common examples of closed loops include:
- restarting or replacing a failed component and confirming the service recovered;
- adding capacity when demand rises and removing it when demand falls;
- rotating an expiring certificate before it causes an outage;
- resolving routine service requests, such as password resets, without a queue.
Each loop should be small, well understood and tested. A large system is safer built from many simple loops than from one clever one. Every automated fix should still leave a record, and repeated fixes for the same problem should trigger an investigation into the underlying cause, so automation does not hide a defect behind endless restarts.
Guardrails that keep it safe
Automation acts faster than people, which is its value and its danger. A person who makes a mistake usually affects a few systems before noticing; automation can repeat the same mistake across thousands of systems in seconds. Guardrails are what make zero-touch operations trustworthy:
- Limit the blast radius: act on one system or a small group first, check the result, then continue.
- Rate limits: cap how many actions automation can take in a period, so a fault in the automation cannot spread unchecked.
- Approval tiers: let automation act alone only on low-risk, reversible actions; require a person's approval for higher-risk ones, with the evidence presented to them.
- Rollback: every automated change should have a tested way back. Where a change cannot be undone, automation should not make it without a person's approval.
- Sanity checks: automation should stop, not act, when its input is empty, unexpected or far outside normal ranges, and its actions should be safe to repeat. A real case described in Google's Site Reliability Engineering book shows why: an empty list of machines was read as "all machines", and disks across a whole content delivery network were erased.
- A stop switch: people must be able to pause automation quickly, for a single loop or for all of it, and know what will stop working when they do.
- Change records and audit trails: every automated action is logged with what triggered it, what it did and whether it worked, and falls under change management, usually as pre-approved standard changes for well-tested routine actions and the full process for anything else.
- Least privilege: automation accounts get only the permissions their task needs, and their credentials are protected like any administrator's.
Treat automation as production software: test it, version it, monitor it, give it an owner, and make sure it still works when the systems it repairs are down.
Watch for automation fighting automation, for example one tool scaling a service up while another scales it down, or two remediation loops repeatedly undoing each other. Clear ownership of each loop prevents it.
Where to start
Do not start by automating everything. Start with toil. Google's Site Reliability Engineering book describes toil as work tied to running a production service that is manual, repetitive, automatable, tactical (reactive rather than planned), leaves nothing of lasting value behind, and grows in direct proportion to the size of the service. Toil is not simply work people dislike: meetings and administration are overhead, not toil. It is the best first target: frequent, well understood, and tiring for people.
A practical path:
- Measure the work. Look at tickets and alerts to find which tasks happen most often and take the most time.
- Write it down. Turn each task into a clear runbook. If the team describes the steps differently, settle that first; automation would only repeat the disagreement faster.
- Automate the steps, keep the decision. First let automation gather information and prepare the fix, with a person pressing the button.
- Close the loop. Once the automation has been right many times, let it act alone on the low-risk cases, with verification and rollback in place.
- Expand carefully, one loop at a time.
Maturity grows in stages, from manual work, to scripts people run, to automation that suggests actions, to automation that acts with approval, to automation that acts alone within set limits. Different tasks can sit at different stages, and should, according to their risk. Industry bodies have published similar scales; for example, TM Forum describes autonomous network levels from L0 (manual) through L1 (assisted), L2 (partial), L3 (conditional) and L4 (high) to L5 (full autonomy). Scales like this describe a direction of travel, not a target every task must reach.
Measuring and the people side
Measure whether zero-touch operations is working, not just whether automation exists:
- The share of routine events resolved without human touch, and how that changes over time.
- Time to restore service, for automated and manual resolutions.
- Wrong or harmful automated actions: how often automation took an action that was wrong or made things worse. This should be very low and investigated every time.
- Hand-offs: how often automation passed a problem to a person, and whether the evidence it gathered was useful.
- Toil remaining: the hours people still spend on routine manual work. Some teams set an upper limit on toil, for example keeping at least half of engineers' time for improvement work, and treat going over it as a signal to invest in automation.
The people side matters as much as the technology.
- Roles change. Engineers spend less time on repeated fixes and more on designing, testing and improving automation, and on the hard problems automation cannot solve.
- Skills must not fade. If automation handles everything routine, people may lose practice at doing it by hand. Exercises and reviews keep skills current for the day automation fails.
- Trust is earned. Teams trust automation that is transparent about what it did and why, that has been right many times, and that they can stop. They resist automation that acts invisibly.
Zero-touch operations is not about removing people. It is about moving them from repetitive fixes to the work only people can do.
Ten questions
Answer all ten questions, then check your answers. You need 9 out of 10 to pass and receive a certificate. If you score less, you will see which answers were right and wrong, and then go through the course again before you retake the check. Your answers, progress and times are kept only in this browser.
Your answers
Your certificate of completion
Enter your name as you want it to appear, then save the certificate as a PDF. In the print window, choose Save as PDF. A certificate is issued once per completion of the course.
Saolix does not record who takes this course, so it cannot verify these certificates. The certificate confirms completion of a free self-paced course and is not an accredited qualification.
Sources
The official documents this course relies on. Laws and guidance change, so check the current version.
- Zero touch network & Service Management (ZSM) · European Telecommunications Standards Institute
- ETSI launches Zero touch network and Service Management group (December 2017) · European Telecommunications Standards Institute
- Site Reliability Engineering, chapter 7: The Evolution of Automation at Google · Google
- The Site Reliability Workbook, chapter 6: Eliminating Toil · Google
- Autonomous Networks: Empowering Digital Transformation (2019) · TM Forum
- Site Reliability Engineering, chapter 5: Eliminating Toil · Google