Software and IT are process businesses too. Work flows through queues and reviews, defects escape, systems fail, and changes go out with more or less checking. Ideas from Lean and quality engineering, such as limiting work in progress, measuring where defects are found, running blameless reviews, and building checks into the pipeline, translate well.

This hub groups the site's software and IT guides, tools, and templates with the general quality and improvement material that fits best. Start with the learning path, or jump to the tool you need.

Start with Defect Prevention Open the Flow Metrics Calculator

About This Hub

Educational content only. This hub applies quality and process-improvement methods to software delivery and IT operations. It is not security, legal, compliance, or engineering advice, and it does not replace your organization's policies, the standards and regulations that apply to you, or the judgment of qualified engineers and security professionals. It is written from a quality and operations perspective: the author is not a software engineer, site reliability engineer, or security professional. Have qualified people review changes to production systems and controls, and avoid putting customer data, credentials, or other sensitive information into any tool or template.

Suggested Learning Path

  1. Defect Prevention and Code Review — find defects earlier and measure how well you remove them
  2. Defect Removal Efficiency Calculator — calculate DRE and cost by stage
  3. Flow Metrics and WIP Limits — measure cycle time and forecast with percentiles
  4. Flow Metrics Calculator — see what a WIP limit would do
  5. Incident Management and Blameless Postmortems — learn from outages with SLOs and error budgets
  6. Quality Gates and Delivery Metrics — build fast feedback and track the DORA measures
  7. Definition of Done and Quality Gate Checklist — record gates and delivery metrics

Quality in Software and IT Guides

FMEA

Analyze how a system or change could fail.

Tools and Calculators

Templates

Software Quality and Delivery Terms You Will Meet

Term or frameworkWhat it means for improvement work
Defect removal efficiency (DRE)The share of all defects found before release.
Shift leftMoving checks earlier in the delivery process, closer to where defects are introduced.
Cycle time and lead timeTime from start of work to done, and from request to delivery.
WIP limitA cap on items in progress, to reduce waiting and expose bottlenecks.
SLI, SLO, and error budgetA service measure, its target, and the allowed unreliability (100% minus the SLO).
MTTD, MTTA, and MTTRMean time to detect, acknowledge, and restore, for incidents.
Blameless postmortemA review of an incident that focuses on system and process causes, not blame.
DORA metricsDeployment frequency, lead time for changes, change failure rate, and time to restore service.
Quality gateAn automated or manual checkpoint a change must pass to move on.
Definition of doneThe agreed criteria for a change or release to count as finished.

Body of Knowledge Definitions

Quality in Software and IT Hub: Frequently Asked Questions

Do Lean and quality methods apply to software?

Yes. Limiting work in progress, mapping and shortening flow, measuring defects by where they are found, root cause analysis, and small tested changes all apply to software delivery and IT operations, adapted to the fact that the work is knowledge work and often continuous.

Where should a software team start?

Measure before changing. Record where defects are found to calculate defect removal efficiency, collect cycle times to see the distribution, and review incidents in blameless postmortems with action items. Then choose one improvement, such as smaller changes or an automated gate, and track the effect.

Is this content security or engineering guidance?

No. It is educational material about quality and process-improvement methods, written from a quality and operations perspective. It does not replace your organization's policies, applicable standards, or qualified professional judgment, and changes to production systems and controls should be reviewed by qualified people.

Sources and Further Reading

  • Nicole Forsgren, Jez Humble, and Gene Kim, Accelerate; DORA State of DevOps reports.
  • Betsy Beyer and colleagues (eds.), Site Reliability Engineering; John Allspaw, "Blameless PostMortems and a Just Culture."
  • David J. Anderson, Kanban; Mary and Tom Poppendieck, Lean Software Development.
  • Capers Jones, Applied Software Measurement; Steve McConnell, Code Complete.
  • ASQ Certified Software Quality Engineer Body of Knowledge.