Why change needs managing
Every improvement to a system is a change: new features, fixes, security patches, configuration edits, infrastructure updates. Changes are also one of the most common causes of incidents; Google's Site Reliability Engineering book reports that roughly 70% of outages are due to changes in a live system. Change management and release management exist to get the benefits of change without the damage.
- Change management decides whether and when a change should happen, based on its risk, and makes sure it is recorded.
- Release management plans and carries out getting a set of changes into live use: building, testing, deploying and making them available to users.
Two failure modes sit at opposite ends:
- Too little control: changes made directly on live systems, untested and unrecorded, so nobody knows what changed when something breaks.
- Too much friction: slow approval chains and rare, huge releases that bundle hundreds of changes together. Big releases are harder to test, harder to fix and more likely to fail.
Modern practice aims for small, frequent, well-tested and reversible changes, with controls that are automated wherever possible and human attention saved for the changes that really need it.
Assessing and approving changes
Not every change needs the same scrutiny. Most organisations sort changes into types:
- Standard changes: routine, low-risk, pre-approved changes that follow a documented, tested procedure.
- Normal changes: assessed for risk and approved by the right person or group before they are made.
- Emergency changes: needed urgently, usually to fix or prevent an incident; approved through a faster route, still recorded, and reviewed afterwards. A rising number of emergency changes is a warning sign that the normal route is too slow.
Assessing a change means asking what could go wrong, who would be affected, how it was tested, how it will be undone and whether the timing clashes with busy periods or other changes.
Who approves matters. Many organisations use a change advisory board, a meeting that reviews changes. Research on software delivery summarised in the book Accelerate (2018) found that requiring approval from an external body, such as a manager or change advisory board, was associated with longer lead times, less frequent deployments and slower recovery, and had no measurable association with the rate of failed changes. Teams that relied on peer review of changes, supported by automated testing and a deployment pipeline, achieved higher delivery performance.
That does not mean removing control. It means putting control where it works: automated checks on every change, review by people who understand the change, and senior attention reserved for genuinely high-risk changes. In regulated organisations, keep segregation of duties: the person who writes a change should not be the only person who approves it.
Automated pipelines
A deployment pipeline is the automated path a change takes from an engineer's work to live use. Two linked practices are central:
- Continuous integration (CI): every change is merged into a shared codebase often and automatically built and tested, so problems are found within minutes, not weeks.
- Continuous delivery (CD): every change that passes the pipeline is ready to release at any time, so releasing becomes a routine business decision rather than a risky event.
- Continuous deployment goes one step further: every change that passes the pipeline is deployed to production automatically, with no manual release step. Continuous delivery does not require this, and suits regulated settings, and software such as mobile apps, where automatic deployment is not practical.
A good pipeline includes:
- Automated tests at several levels: small unit tests, tests of how components work together, and tests of whole user journeys.
- Security checks, such as scanning for known vulnerabilities in dependencies and for secrets committed by mistake.
- The same process for every environment, so what was tested is what is released.
- An audit trail: who changed what, who reviewed it, which tests ran and what was released when.
Infrastructure and configuration changes should go through a pipeline too. Many serious outages come from a configuration edit that was never tested.
The pipeline itself becomes a powerful control. When every change must pass the same automated checks and review, approval can shift from meetings to evidence.
Releasing with less risk
How a change reaches users affects how much damage it can do. Several release strategies limit the risk:
- Rolling deployment: replace old versions with the new one a few servers at a time, so both versions run side by side during the rollout.
- Blue-green deployment: run two identical environments; release to the idle one, test it, then switch traffic over. Switching back is quick.
- Canary release: send the new version to a small share of users or traffic first, watch closely, and widen only if all is well.
- Feature flags: deploy new code switched off, then turn features on for chosen users, and off again instantly if there is a problem. This separates deploying code from releasing a feature.
Whatever the strategy:
- Before releasing, confirm how you will detect a problem and that the rollback has been tested.
- Watch the release with the measures that matter to users, such as errors and response times, and compare them with before.
- Stop automatically if those measures get worse beyond an agreed limit.
- Avoid releasing at the worst times, such as the start of a busy period, unless the change is urgent.
- Take care with data changes. Changes to database structures are often hard to reverse. Design them to be compatible with both the old and the new version, so the application can be rolled back without undoing the data change.
Measuring delivery, and when it goes wrong
The DORA research programme (DevOps Research and Assessment) has studied software delivery since 2014, mainly through large surveys of technology professionals. Its best-known measures of software delivery are:
- Deployment frequency: how often changes are deployed to production.
- Change lead time: how long a change takes from being committed to version control to running in production.
- Failed deployment recovery time: how long it takes to recover from a deployment that fails and needs immediate intervention. Older reports called a broader version of this "time to restore service".
- Change fail rate: the share of deployments that need immediate intervention after release.
- Deployment rework rate: the share of deployments that are unplanned and happen because of an incident in production. DORA added this fifth measure in 2024.
DORA groups the first three as measures of throughput and the last two as measures of instability. Its research consistently finds that speed and stability are not a trade-off: the best-performing teams deploy frequently and have low failure rates and fast recovery. Use the measures to improve your own team over time, not to rank teams against each other, and do not turn them into targets, or people will start to game them.
When a release does go wrong:
- Roll back or roll forward. Rolling back to the last good version is often the fastest way to restore service, provided the release was designed to be reversible. Rolling forward, fixing forward with a new change, can be right when a rollback is impossible, but it adds risk under pressure.
- Treat it as an incident if users are affected, with the usual roles and communication.
- Review it without blame, and ask what in the pipeline, tests or release strategy would have caught it sooner.
Every failed change is information. Teams that learn from it get faster and safer at the same time.
Ten questions
Answer all ten questions, then check your answers. You need 9 out of 10 to pass and receive a certificate. If you score less, you will see which answers were right and wrong, and then go through the course again before you retake the check. Your answers, progress and times are kept only in this browser.
Your answers
Your certificate of completion
Enter your name as you want it to appear, then save the certificate as a PDF. In the print window, choose Save as PDF. A certificate is issued once per completion of the course.
Saolix does not record who takes this course, so it cannot verify these certificates. The certificate confirms completion of a free self-paced course and is not an accredited qualification.
Sources
The official documents this course relies on. Laws and guidance change, so check the current version.
- DORA's software delivery performance metrics · DORA (Google Cloud)
- History of DORA's software delivery metrics · DORA (Google Cloud)
- Site Reliability Engineering: Introduction · Google
- Capabilities: Streamlining change approval · DORA (Google Cloud)
- Capabilities: Continuous delivery · DORA (Google Cloud)