The documents operations teams need
In a crisis at three in the morning, the difference between a ten-minute fix and a three-hour outage can be whether a good document exists. Technology operations depends on a small set of document types, each with a distinct job:
- Runbooks: step-by-step instructions for a specific task or problem, such as "restart the payment service" or "respond to a full disk". Some organisations use standard operating procedure for similar documents; a playbook usually covers a wider response involving several people and runbooks.
- Architecture and service descriptions: what each service does, what it depends on, who owns it and how data flows through it.
- Configuration and inventory records: what systems exist, where, in which version and with which settings.
- Knowledge base articles for the service desk: known problems, workarounds and answers to common questions.
- Incident and post-incident reports: what happened, why, and what is being done so it does not happen again.
- Policies and standards: the rules the team works to, such as change approval, access and data retention.
A useful way to think about any document is to ask what the reader is trying to do. The Diátaxis framework, created by Daniele Procida and adopted in hundreds of documentation projects, separates four needs: learning (tutorials), doing a task (how-to guides), looking something up (reference) and understanding (explanation). A runbook is a how-to guide; mixing long explanation into it makes it slower to follow under pressure.
A good runbook states when to use it, what access is needed, each step with the expected result, how to check it worked, how to roll back, and whom to escalate to.
Documentation as code
Documents kept in scattered word-processor files and personal notes drift out of date and are hard to find. Many teams now manage documentation the way they manage software, an approach called docs as code:
- Plain-text formats, such as Markdown, stored in version control alongside the systems they describe.
- Review before publishing: changes to documents go through the same review as code changes, so a second person checks them.
- History: every change is recorded with who made it and when, and the change message explains why, so you can see what a runbook said on the day of an incident.
- Automated checks: broken links, missing sections and out-of-date references can be caught automatically before publishing.
- Published from one source, so everyone reads the same current version.
The biggest benefit is that documentation changes with the system. When an engineer changes a service, the review can ask: "did you update the runbook?" Documentation that lives next to the code, and is checked in the same review, is more likely to be kept current than documentation stored somewhere else.
Generating documentation from the systems
The documentation least likely to drift is documentation produced automatically from the systems themselves, or from the same definitions the systems are built from, and regenerated with every change.
- Interface documentation from specifications: when an API is described in a machine-readable format such as the OpenAPI Specification, reference documentation can be generated from it automatically and kept in step with every release, provided the specification is itself checked against the running service, for example by automated tests.
- Infrastructure as code: when servers, networks and settings are defined in code, inventories and diagrams can be produced from those definitions rather than drawn by hand.
- Configuration and inventory reports exported from the systems of record on a schedule.
- Dependency maps built from what services actually call, observed through monitoring and tracing.
- Change and incident timelines assembled from tickets, alerts and logs.
Generated documents are excellent for what is: facts about the current state. They are poor at why: the reasons for a design, the trade-offs considered, the lesson learned in last year's outage. That still needs people. Short decision records, each explaining one significant design choice, its context and the options rejected, are a lightweight way to capture the why. A good documentation set combines generated reference material with human-written explanation and runbooks, and links between them.
AI-assisted documentation, used safely
Language models can take much of the effort out of documentation. They can:
- draft a runbook from an engineer's notes or a recorded working session;
- summarise an incident from its chat, alerts and timeline into a first draft of the post-incident report;
- explain a configuration file or script in plain language;
- tidy and standardise existing documents into a consistent template;
- answer questions by searching the documentation set.
They also bring real risks, and each needs a control:
- Confident errors. A model can produce steps or facts that look right but are wrong. A wrong command in a runbook is dangerous. Every AI-drafted operational document must be reviewed and tested by a person who knows the system before it is published.
- Leaking secrets. Configuration files, logs and chats often contain passwords, keys and personal data. Remove them before sending anything to an AI service, and use only tools approved for that data.
- Hidden instructions. Text inside logs, tickets or documents can contain instructions that try to redirect the model, a problem known as indirect prompt injection. It cannot be fully prevented, so treat what the model reads as data, do not give the model more access than it needs, and review what it produces.
- Too much access. If an AI tool can run commands or change systems, not just write text, limit what it may do and require a person to approve each action.
- False authority. Mark AI-drafted content until it has been reviewed, so readers know its status.
Used this way, AI speeds up the writing; people remain responsible for the content.
Keeping documents trustworthy
A document people do not trust can be worse than none: they ignore it, or follow it and make things worse. Trust comes from a few habits:
- An owner for every document, named on it, who answers for keeping it right.
- A review date, with overdue documents flagged automatically.
- Test the runbooks. Use them in exercises and during real incidents, and fix every step that was wrong or unclear. A runbook nobody has followed recently should be treated as untested.
- Update after every incident and every change. A post-incident review should always ask which documents failed or were missing.
- Make them findable: one place to search, clear titles, and links from alerts straight to the runbook for that alert.
- Keep them available when systems are down. The runbook for restoring the documentation system should not live only inside the documentation system.
- Control access. Keep secrets out of documents entirely, referring instead to where they are stored, and restrict documents that describe sensitive systems.
- Retire what is obsolete, clearly marked, rather than leaving old instructions for someone to find.
The test of an operations documentation set is simple: could a capable engineer who is new to the team restore a service at night using only what is written down?
Ten questions
Answer all ten questions, then check your answers. You need 9 out of 10 to pass and receive a certificate. If you score less, you will see which answers were right and wrong, and then go through the course again before you retake the check. Your answers, progress and times are kept only in this browser.
Your answers
Your certificate of completion
Enter your name as you want it to appear, then save the certificate as a PDF. In the print window, choose Save as PDF. A certificate is issued once per completion of the course.
Saolix does not record who takes this course, so it cannot verify these certificates. The certificate confirms completion of a free self-paced course and is not an accredited qualification.
Sources
The official documents this course relies on. Laws and guidance change, so check the current version.
- Diátaxis, a systematic approach to technical documentation · Daniele Procida
- Docs as Code · Write the Docs
- OpenAPI Specification (latest) · OpenAPI Initiative
- OWASP Top 10 for LLM Applications 2025 · OWASP Gen AI Security Project (OWASP Foundation)
- LLM01:2025 Prompt Injection · OWASP Gen AI Security Project
- How-to guides · Diátaxis
- About the OpenAPI Initiative · OpenAPI Initiative (Linux Foundation)