Teams often assume that shipping faster means shipping less safely. Research on software delivery found the opposite: teams that deploy often, in small changes, with automated checks, tend to have lower failure rates and recover faster.
This guide covers the four DORA delivery metrics, how to design quality gates that give fast feedback, why a written definition of done matters, and a worked example in which a team's change failure rate falls from 15.0% to 8.3% as its deployment frequency rises.
Before You Start
Why Quality Gates and Delivery Metrics Matter
Speed and Stability Are Not a Trade-Off
Research summarized in Accelerate found that teams delivering frequently also tend to have lower failure rates and faster recovery.
Gates Give Fast Feedback
Automated checks that run on every change catch problems while the author still remembers the code.
Manual Approvals Slow Without Protecting
A review board that meets weekly adds waiting. Automated gates and small changes reduce risk more effectively.
Metrics Show Whether Changes Work
Four delivery measures show whether practices are improving speed and stability together.
The Four Delivery Metrics
The DORA (DevOps Research and Assessment) research program identified four measures of software delivery performance, described in Accelerate and in the annual State of DevOps reports:
| Metric | Definition | Type |
|---|---|---|
| Deployment frequency | How often the team deploys to production | Throughput |
| Lead time for changes | Time from a change being committed to running in production | Throughput |
| Change failure rate | Share of deployments that cause a failure needing remediation | Stability |
| Time to restore service | How long it takes to recover from a failure in production (recent DORA reports call this failed deployment recovery time) | Stability |
DORA groups teams into performance tiers, and the definitions and thresholds have been refined over time. Check the current report before comparing yourself with a benchmark, and compare mainly with your own trend.
Designing Quality Gates
- Fast first. Put quick checks, such as build, unit tests, and linting, early, and slower ones later.
- Clear ownership and thresholds. Decide what fails a gate (for example, a critical vulnerability) and what only warns.
- Trustworthy checks. Flaky tests train people to ignore failures. Fix or remove them.
- A written definition of done. Agree what "done" means for a change and for a release, including tests, documentation, and monitoring. See the Definition of Done and Quality Gate Checklist.
- Beware targets on proxies. A code coverage target can be met by tests that assert nothing. Use coverage to find untested areas, not as a goal.
Worked Example: Smaller Batches and Automated Gates
A team measured its delivery over 30 days, then introduced automated integration tests, a security scan, and canary releases, and shipped smaller changes. The figures are illustrative.
| Metric | Before | After |
|---|---|---|
| Deployments in 30 days | 40 | 60 |
| Failed deployments | 6 | 5 |
| Change failure rate | 6 / 40 = 15.0% | 5 / 60 = 8.3% |
| Median lead time for changes | 3 days | 1 day |
| Median time to restore | 2 hours | 35 minutes |
Deployment frequency rose by half, lead time fell by two thirds, and the change failure rate fell from 15.0% to 8.3%, even though the team shipped more often. Restore time improved partly because canary releases limited each failure's impact and rollback was automatic. The team does not claim that the gates alone produced the change: batch size fell at the same time, and one month of data is a small sample. It will keep tracking the four measures and also watch a balancing measure, engineer satisfaction, to make sure the process is not just adding toil.
Record the four measures monthly in the Definition of Done and Quality Gate Checklist workbook, which calculates change failure rate and charts the trend.
Self-Assessment Questions
- Do automated checks run on every change and give results in minutes?
- Do we trust our tests, and do we fix flaky ones?
- Do we have a written definition of done?
- Do we track deployment frequency, lead time, change failure rate, and time to restore?
- Do we compare ourselves with our own trend, not just with benchmarks?
Common Mistakes
Gates That Are Too Slow
A gate that takes hours discourages small changes. Keep feedback fast.
Manual Approval as the Main Control
Approval boards add delay and do not reliably catch defects. Automate the checks and reduce batch size.
Gaming the Metrics
Deploying trivial changes inflates frequency. Look at the four measures together, and at outcomes.
Ignoring Restore Time
Failure will happen. Practice recovery, and make rollback quick.
Quality Gates and Delivery Metrics: Frequently Asked Questions
What are the four DORA metrics?
They are deployment frequency (how often you deploy to production), lead time for changes (time from commit to production), change failure rate (the share of deployments causing a failure that needs remediation), and time to restore service after a production failure. The first two measure throughput and the last two measure stability.
What is a quality gate?
A quality gate is an automated or manual checkpoint in a delivery pipeline that a change must pass before moving on, such as passing unit tests, a security scan below a severity threshold, or integration tests. Good gates give fast, clear feedback and block changes that fail.
Should we set a code coverage target?
Use coverage to discover untested areas, but be careful about setting a numeric target as a goal. When coverage becomes a target, teams can meet it with tests that execute code without checking behavior. Pair it with other measures such as escaped defects and change failure rate.
Sources and Further Reading
- Nicole Forsgren, Jez Humble, and Gene Kim, Accelerate: The Science of Lean Software and DevOps.
- DORA (DevOps Research and Assessment), annual State of DevOps reports.
- Jez Humble and David Farley, Continuous Delivery.
- ISO/IEC 25010, systems and software quality models.