One place for each phase of DMAIC: what it is for, the steps, the tools, the deliverables, the tollgate checklist, and the mistakes to avoid.
DefineMeasureAnalyzeImproveControl
Choose a phase tab below. Each tab opens with an overview, then covers the details needed to run that phase well, with links to the calculators, templates, and guides on this site.
The Define, Measure, and Analyze tabs are available now; the other phases are being added one at a time.
Phase 1 of 5
Define
What problem are we solving, and why does it matter? Define turns a vague concern into a clear, agreed, and measurable project: a problem worth solving, a boundary around it, a customer-based definition of success, a business case, and the people to do the work.
Key question
What problem are we solving, and why does it matter?
Typical duration
About 1 to 4 weeks, depending on project size and how much is already known
Led by
The project leader (Green or Black Belt), with the sponsor and process owner
Primary outputs
Signed project charter, SIPOC, customer requirements (CTQs), business case, team and stakeholder plan
Tollgate decision
Sponsor confirms the project is worth doing, correctly scoped, resourced, and ready for Measure
Most improvement projects that fail were not defeated by difficult statistics. They were defeated at the start: the problem was vague, the scope kept growing, the sponsor was not committed, or nobody agreed what success meant. Define exists to prevent that. It is the phase in which the project earns the right to consume time and money.
What Define must achieve
A problem statement that is specific, measured, and free of causes and solutions
A goal that is measurable, time-bound, and tied to a baseline
A scope that names what is in, what is out, and where the process starts and stops
Customer requirements translated into measurable critical-to-quality (CTQ) characteristics
A business case that finance has reviewed
A sponsor, process owner, and team who have committed time and authority
What Define must not do
Select a solution or assume a cause
Collect large amounts of detailed data (that is Measure)
Map every step of the process in detail
Start a project because of a loud opinion instead of evidence
Proceed without a committed sponsor
Treat the charter as paperwork to be signed once and forgotten
The test of a finished Define. Someone who has never heard of the project should be able to read the charter and answer four questions: what is wrong, how big is it, who is affected, and what will be different when we are done?
Define is also where the improvement method is chosen. Not every problem needs DMAIC. The next sections cover that decision, and each of the steps in order.
The Define Process, Step by Step
The steps overlap and loop back. Customer and scope work often changes the problem statement, which is a sign Define is doing its job.
1
Identify the problem or opportunity
Start from evidence: customer complaints, scrap and rework reports, missed deliveries, audit findings, safety events, cost trends, or a strategic gap. Write down who raised it, what they observed, and what data they used.
Collect the raw signals: complaints, defect and downtime data, cost reports, customer scorecards
Ask the people who raised the concern for examples and dates
Check whether the problem is chronic and recurring or a one-time event
Output: A short list of candidate problems with the evidence behind each
Watch for: Problems described only as opinions (“quality is bad”) and problems already caused by a known, fixable event
2
Select and prioritize the project
Compare candidates against agreed criteria: strategic fit, customer impact, financial impact, feasibility, and sponsor support. Decide whether DMAIC is the right method (see the next section).
Score candidates with a prioritization matrix
Check that the root cause is not already known
Confirm that the project can be completed in roughly three to six months
Output: One selected project, a recorded reason for choosing it, and the method chosen
Watch for: Choosing the project that is easiest to explain, not the one with the best evidence and value
3
Understand the customer and the process at a high level
Identify the customers of the process output, listen to them, and translate what they say into measurable requirements. Draw a SIPOC to agree the boundary and flow.
Gather the voice of the customer from interviews, complaints, surveys, and data
Classify needs with the Kano model
Build a CTQ tree
Draw the SIPOC with the team and the process owner
Output: CTQs with measurable definitions, and a SIPOC
Watch for: Assuming you know what the customer wants without asking, and drawing the process as it should be rather than as it is
4
Quantify the impact and the baseline
Put numbers on the problem: how often, how large, since when, and what it costs. Use the data that already exists. A preliminary baseline is enough; detailed measurement belongs in the next phase.
Pull existing data on the primary metric for a representative period
Estimate the cost of poor quality and the benefit if the goal is met
Ask finance to review the benefit method
Output: A preliminary baseline, a benefit estimate, and agreement on how the benefit will be validated
Watch for: Benefit estimates that count every cost as savings, or that finance has never seen
5
Draft the charter and form the team
Write the charter: problem statement, goal, scope, business case, constraints, risks, timeline, and roles. Name the sponsor, process owner, project leader, and team members, and confirm their time with their managers.
Draft the problem and goal statements and test them with the sponsor
Choose team members who know the process and can act on findings
Analyze stakeholders and plan communication
Output: A draft charter, a named team, a stakeholder analysis, and a communication plan
Watch for: A team assembled by availability rather than knowledge, and a charter written alone by the project leader
6
Hold the tollgate review
Present the charter, SIPOC, customer requirements, baseline, benefit case, and plan to the sponsor and process owner. The aim is a decision: proceed, proceed with conditions, re-scope, or stop.
Use the tollgate checklist in this tab
Record decisions and actions
Obtain signatures
Output: A signed charter and a recorded go, conditional go, re-scope, or stop decision
Watch for: A review that is a formality. A project stopped at this point is a good outcome, because it frees resources for work that is worth doing
Choosing the Right Project and the Right Method
A well-defined project that is the wrong project still wastes months. Selection is therefore part of Define. Use criteria that the sponsor has agreed, score candidate projects against them, and keep a record of why the winner was chosen.
Criterion
Question to ask
How to check
Strategic fit
Does it support a named business objective, customer commitment, or regulatory need?
Ask the sponsor which objective it advances. If there is no answer, expect weak support later.
Customer impact
Does it affect quality, delivery, cost, safety, or service as customers experience it?
Link to complaints, returns, delivery performance, or customer scorecards.
Financial impact
Is the cost of poor quality, or the opportunity, large enough to justify the effort?
Rough estimate first; finance reviews the method.
Feasibility
Can it be done in about three to six months with the people and data available?
Check data availability early. Break large problems into several projects.
Cause not yet known
Is the root cause unclear and does it need analysis?
If the cause and fix are obvious, just do it.
Ownership and support
Is there a sponsor and a process owner who want it solved and can release resources?
No sponsor, no project.
An impact-effort grid helps a team and its sponsor compare candidates quickly. Illustrative candidates are shown; the discussion about where each one belongs is the valuable part. Use the Impact-Effort Matrix Builder or the Project Prioritization Matrix.
Is DMAIC the right method? Match the method to the problem:
If the situation is…
Consider
Why
A safety or quality emergency
Containment first, then a corrective action report such as 8D
Protect people and customers before analysis
The cause and the fix are known and low risk
Just do it, with a short plan
A formal project adds overhead without adding insight
A local problem in one area that a team can solve in days
Writing the Problem Statement and the Goal Statement
The problem statement describes the gap between what is happening and what should be happening, in measurable terms. The goal statement describes the improvement the project will deliver. Both should be clear enough that two people would measure them the same way.
A problem statement answers
What is wrong, and in which process or product?
Where does it occur?
When did it start, and how long has it been going on?
How big is it, in a measured unit?
What is the impact on the customer or the business?
A problem statement does not contain
A cause (“because the operators are careless”)
A solution (“we need a new machine”)
Blame for people or departments
Vague words such as “improve” or “better” with no measure
Weak
Strong
Quality on Line 3 needs improvement.
From weeks 14 to 21, first-pass yield at the Line 3 final test fell from 96.8% to 91.2%, creating about 8.8% rework and shipment risk and costing an estimated $338,000 a year at the current rate.
Our permit process is too slow.
Over the last two quarters, the median time from complete application to permit decision was 41 days against a published standard of 30, and 38% of applications exceeded 45 days.
We have too many medication errors.
In the last 12 months, 27 of 1,900 pediatric doses (1.4%) required intervention because of a weight or unit error, against a target of zero reaching the patient.
We need a new scheduling system.
On-time shipment fell from 97% to 89% over eight weeks, with 72% of late orders tied to schedule changes made within 48 hours of the ship date.
The examples above are illustrative.
Check for hidden solutions. If your statement contains the words “need,” “install,” “implement,” “train,” or “replace,” it probably describes a solution. Rewrite it as the gap the solution is meant to close.
The goal statement follows the SMART pattern: Specific, Measurable, Achievable, Relevant, and Time-bound. It should use the same metric as the problem statement, state the baseline and target, and give a date.
SMART element
Example for the Line 3 project
Specific
First-pass yield at Line 3 final test
Measurable
From 91.2% to at least 96.0%, measured weekly with the same definition of first-pass yield
Achievable
The line ran at 96.8% earlier in the year; the target is below that level
Relevant
Reduces rework cost and shipment risk, supporting the plant's quality and delivery objectives
Time-bound
Sustained for eight consecutive weeks within five months of project start
Set the target with evidence. A goal taken from a best-case month or a competitor's claim may not be achievable. Use demonstrated past performance, customer requirements, or technical limits, and agree that the target may be revisited at the end of Measure when the baseline is firm. For a tool to test wording, see the Project Charter guide.
Scope: Drawing the Boundary, and the SIPOC
Scope says what the project will and will not address. Most projects that stall are too broad, so draw the boundary deliberately and write it down. A boundary has four parts: where the process starts and stops, which products, sites, or customers are included, which causes or solution areas are in bounds, and what is explicitly out.
Scope element
Example (Line 3 final test project)
Process start and end
Starts when assembled units arrive at final test; ends when a unit is recorded as passed or sent to rework
In scope
Final test station, test fixtures, test limits, and the quality of inputs received from assembly
Out of scope
Other production lines, product design changes, supplier changes (any of which may be recommended for later)
Constraints
No capital spending above an agreed limit; no reduction of required customer tests
Assumptions
Test data for the last 12 months can be extracted; production volume stays within 10% of plan
SIPOC stands for Suppliers, Inputs, Process, Outputs, Customers. It is a one-page, high-level picture that shows the boundary and who is connected to it. Keep the process to four to seven steps. Build it with the team and the process owner, and start in the middle: map the process steps first, then outputs, then customers, then inputs and suppliers.
A SIPOC for the Line 3 final test. The process is shown in four to seven steps; detailed mapping is left for Measure. Build your own with the SIPOC Diagram Generator.
Use the SIPOC to challenge the scope. If a supplier or input appears that the team cannot influence, note it as a constraint. If a customer appears that the problem statement ignored, reconsider the problem.
Keep an is / is-not list alongside the scope: what the problem is and is not, where and when it occurs and does not occur. See Is / Is-Not Analysis.
From the Voice of the Customer to Critical-to-Quality Requirements
The customer decides whether the output is good enough, so a project that ignores the customer will optimize the wrong thing. Define converts what customers say, in their words, into specific, measurable characteristics the team can control. The customer may be external, or it may be the next department or process step.
Method
Best for
Cautions
Interviews and site visits
Understanding needs, context, and unspoken problems
Time-consuming; use open questions and listen for needs, not solutions
Surveys
Measuring the importance and satisfaction of known needs across many customers
Poorly worded questions give poor answers; low response rates
Complaints, returns, and warranty data
Finding what already goes wrong
Reflects only customers who complain
Customer scorecards and contractual requirements
Understanding formal targets and penalties
May lag actual experience
Observation and journey mapping
Seeing how the output is used
Needs permission and access
Focus groups
Exploring reactions and ideas
Group dynamics can bias results
Organize what you hear. Group comments with an affinity diagram, then rank importance. The Kano model adds a second dimension: which needs are basic expectations (dissatisfy when absent), which are performance needs (more is better), and which are delighters. See the Voice of the Customer and Kano guide and the Kano Model Analyzer.
Translate needs into CTQs. A critical-to-quality characteristic is a measurable requirement that is closely tied to what the customer values. A CTQ tree moves from the broad need, to a driver, to a specific measure with a target.
A CTQ tree for the Line 3 project. Each branch ends in something that can be measured and that has a target. Build one with the CTQ Tree Builder.
Define the primary metric (Y) operationally. Write exactly how it will be counted: what is a unit, what counts as a pass, what is included and excluded, where and when it is measured, and by whom. Ambiguity here will produce different numbers from different people, which becomes an expensive problem in Measure.
The Business Case: Cost of Poor Quality and Expected Benefits
A project competes for time and attention with other work, so the sponsor needs to know what it is worth. The business case estimates the cost of the problem and the benefit of solving it, in terms finance recognizes.
Cost of poor quality category
What it includes
Examples
Internal failure
Defects found before the customer receives the product
Scrap, rework, retest, downtime from defects, re-inspection
External failure
Defects found after delivery
Returns, warranty, complaints, field repair, lost customers, penalties
Appraisal
Finding defects
Inspection, testing, audits, calibration
Prevention
Stopping defects from happening
Training, process design, preventive maintenance, supplier development
Use the COPQ Estimator and the COPQ overview to structure the estimate. Prevention and appraisal costs are not waste in themselves; the aim is to reduce failures, then reduce the appraisal needed to catch them.
Worked benefit estimate, Line 3 final test. The figures below are illustrative.
Item
Calculation
Result
Annual volume
12,000 units per week × 50 weeks
600,000 units
Rework at current yield
600,000 × 8.8% (91.2% first-pass yield)
52,800 units
Rework cost at current yield
52,800 × $6.40 per unit
$337,920
Rework at target yield
600,000 × 4.0% (96.0% first-pass yield)
24,000 units
Rework cost at target yield
24,000 × $6.40
$153,600
Estimated annual benefit
$337,920 − $153,600
$184,320
Be conservative and transparent. This estimate counts rework labor and materials at a standard cost. It does not include shipment risk, capacity gained, or customer goodwill, and it does not subtract the cost of the project. The assumptions, the method, and the validation owner belong in the charter, and finance should review them.
Hard savings
Reductions in cost that appear in the financial statements, such as less scrap, lower overtime, or fewer purchases. Finance confirms them after the project.
Soft or avoided benefits
Cost avoidance, capacity released, risk reduction, and customer satisfaction. Real, but they are reported separately and should not be mixed with hard savings.
Improvement is done by people, and projects fail from lack of access, authority, and support more often than from lack of technique. Define names the people and confirms what each is committing.
Role
Responsibility
Typical commitment
Sponsor or champion
Owns the business reason for the project, removes barriers, approves the charter and tollgates, and secures resources
Regular reviews; a few hours a month
Process owner
Owns the process being improved, supports data access and changes, and sustains results
Active involvement throughout; responsible for control
Project leader (Green or Black Belt)
Leads the project, the method, the data, and the communication
Part to most of their time, depending on project size
Team members
Bring process knowledge, collect data, test ideas, and implement changes
A few hours a week, agreed with their managers
Finance partner
Reviews the benefit method and validates savings
Reviews at defined points
Coach (Master Black Belt or experienced Black Belt)
Advises on method and statistics and reviews tollgates
As needed
Pick team members for knowledge. Include people who do the work, people who feed it, people who receive its output, and someone who can analyze data. Teams of four to eight work well. Confirm the time with each member's manager, not just the member.
Analyze stakeholders. List the people and groups affected, estimate their influence and their current support, and plan how to engage each. Those with high influence and low support need conversations early. See the Stakeholder Analysis Builder and Stakeholder Analysis.
Clarify who decides what. A RACI matrix shows who is Responsible, Accountable, Consulted, and Informed for key activities.
Activity
Sponsor
Process owner
Project leader
Team
Finance
Approve charter and scope
A
R
R
C
C
Provide process data access
I
A
R
R
I
Validate benefit estimate
I
C
R
I
A
Approve process changes (later phases)
C
A
R
R
I
Tollgate decision
A
C
R
I
I
Build one with the RACI Matrix Builder. Finally, agree a communication plan:
Audience
Message
Channel
Frequency
Owner
Sponsor and process owner
Progress, decisions needed, risks
Short review meeting
Every two weeks
Project leader
Team
Tasks, findings, next steps
Team meeting
Weekly
Project leader
Affected employees
Why, what is happening, what to expect
Huddles and visual board
Monthly or at milestones
Process owner
Finance
Benefit method and results
Review
At tollgates
Project leader
Just Enough Data in Define
Define needs enough data to size the problem and justify the project, and no more. Careful measurement, including checking that the measurement system is reliable, belongs in Measure. If you wait for perfect data before chartering, the project will never start.
Collect in Define
The primary metric for a representative period (often 3 to 12 months)
A rough view of variation over time (a simple run chart)
Obvious stratifications, such as by line, shift, or product
A preliminary look at the baseline. A yield of 91.2% corresponds to a sigma level of about 2.85, and a yield of 96.0% to about 3.25, using the conventional 1.5-sigma shift. The sigma level is a convenient way to compare processes. It is not a goal in itself. Use the Sigma Level and DPMO Suite to convert between yield, defect rates, and sigma level.
Warning on early data. Existing data may be defined differently from what the project needs, may have gaps, or may not be trusted by the people who use it. Note these limitations in the charter and plan to check them in Measure, rather than quietly relying on them.
Look for the shape of the problem. A quick Pareto chart of defect types, or a run chart of the metric over time, often shows whether the issue is sudden (a recent change) or long-standing (a chronic cause). That affects scope and the approach in later phases.
Putting It Together: The Project Charter
The charter is the one-page agreement that records everything above. It is signed by the sponsor and referred to throughout the project. The table shows its elements, with the Line 3 example completed.
Charter element
Line 3 final test example
Project title
Improve first-pass yield at Line 3 final test
Problem statement
From weeks 14 to 21, first-pass yield at final test fell from 96.8% to 91.2%, producing about 8.8% rework and shipment risk.
Goal statement
Raise first-pass yield from 91.2% to at least 96.0%, sustained for eight consecutive weeks within five months of project start.
Business case
Estimated rework cost reduction of about $184,000 a year (illustrative), reduced shipment risk, released test capacity. Method and validation owner: finance.
Customer and CTQs
Packaging and end customer. First-pass yield ≥ 96.0%; fewer than 500 defective units per million at receipt; on-time shipment ≥ 98%.
Scope
In: final test station, fixtures, limits, input quality from assembly. Out: other lines, design changes, supplier changes.
Team and roles
Sponsor: plant manager. Process owner: Line 3 production manager. Leader: Green Belt (quality engineer). Members: test technician, assembly lead, maintenance, finance partner.
Timeline
Define 3 weeks, Measure 4, Analyze 4, Improve 6, Control 4: about 21 weeks (five months).
Constraints and assumptions
No capital spending above the approved limit; 12 months of test data can be extracted; volume within 10% of plan.
Risks and mitigation
Test equipment downtime (schedule around maintenance); seasonal volume (agree a trading calendar); team member availability (confirmed with managers).
Approvals
Sponsor, process owner, project leader, finance reviewer.
The charter is a living agreement. If data later show that the real problem lies elsewhere, change the scope by agreement with the sponsor, record the change, and keep going. Do not let the project drift, and do not hold to a scope the evidence has overturned.
Tollgate Review: Are We Ready for Measure?
A tollgate is a decision point. The sponsor reviews the work and decides whether the project should proceed. Use the checklist below before the review. Your progress is saved in this browser only, and nothing is sent anywhere.
0 of 15 complete
Questions a sponsor should ask:
If this project succeeds, what will be different for customers and for the business, and how will we know?
Is this the right size? Could it be finished in about five or six months?
How was the benefit estimated, and who will validate it?
Who owns the process, and are they committed to sustaining the result?
What are the biggest risks to this project, and what is the plan for each?
Is there anything in the charter that assumes a cause or a solution?
Tollgate outcome
Meaning
Typical next step
Go
All criteria are met
Start Measure
Go with conditions
Minor gaps that can be closed quickly
Record conditions, owners, and dates; review in two weeks
Re-scope
Scope, goal, or business case needs rework
Revise the charter; hold another review
Stop or defer
Not worth doing now, or no committed sponsor
Document why; free the team for other work
Adapting Define to the Situation
Situation
How Define changes
Green Belt project (smaller, local)
A light charter and a short SIPOC; limited VOC; benefit estimate with local finance review; often one to two weeks
Black Belt project (cross-functional, larger value)
Fuller VOC, stakeholder analysis, a formal business case, and more tollgate formality; two to four weeks or more
Lean or kaizen event
A one-page charter or A3 background and problem statement, set before the event, with a narrow scope and a fixed date
Manufacturing
CTQs often relate to specifications, yield, scrap, downtime, and on-time delivery; data is often in production and quality systems
Service and transactional
CTQs relate to time, accuracy, and experience; the boundary of the process is often less visible; data may need to be created
Healthcare and public service
Safety, equity, and legal requirements affect scope and stakeholders; involve patients or residents and staff; see the hub guides for healthcare and government
Lean and Six Sigma together. Lean projects often start from a value stream map that points to waiting and flow problems; Six Sigma projects start from variation and defects. In Define, the same questions apply: what is the problem, who is the customer, and what is the boundary? See Lean Six Sigma Integration.
Common Mistakes and Red Flags
Mistake
What it looks like
How to correct it
The solution is in the problem statement
“We need to install a new fixture.”
Rewrite as the gap; move the idea to a parking lot
Scope is too broad
A single project covers several lines, products, and causes
Split into projects; narrow with Pareto and stratification
No committed sponsor
The sponsor misses reviews and cannot release resources
Escalate; pause until commitment is real
Goal with no baseline
“Reduce defects by 50%” with no starting figure or definition
Establish and state the baseline and the metric definition
Benefit never reviewed by finance
Savings claimed that no one can later find
Agree the method and validation owner at the start
Customer never consulted
CTQs invented by the team
Interview or survey customers, or use customer data
Team chosen by availability
No one who does the work is on the team
Ask who knows the process; negotiate time with managers
Over-collecting data
Weeks spent on detailed measurement before the charter
Use existing data; save detailed work for Measure
Charter as paperwork
Signed once and never revisited
Review at each tollgate and record changes
Define never ends
Endless refinement of the charter
Time-box; sign when the checklist is met
Define Resources on This Site
Guides
Project CharterElements, review checklist, and a worked scrap-reduction charter
It depends on project size and how much is already known. Many Green Belt projects complete Define in one to three weeks, and larger Black Belt projects in two to four. If Define is taking much longer, the usual causes are a problem that is too broad, a sponsor who has not committed, or unclear data access. Time spent here is well spent, because errors in Define carry through the whole project.
What is the difference between a project charter and a project plan?
The charter is the agreement: the problem, goal, scope, business case, team, and high-level timeline that the sponsor approves. A project plan is the detailed working schedule of tasks, owners, and dates that the team maintains. The charter changes only by agreement; the plan changes as the team learns.
Should we include solution ideas in Define?
Capture them, but do not commit to them. Ideas that arise in Define are useful to record in a parking lot and test later, but a charter that names the solution turns the project into an implementation and removes the analysis that DMAIC exists to provide. The problem statement and goal should describe the gap, not the fix.
Do internal process projects need a voice of the customer?
Yes, though the customer may be the next process step, the end user, or an internal department. Every process has a customer who receives its output and decides whether it is good enough. Without that view, teams optimize what is easy to measure and miss what matters to the receiver.
Is a SIPOC enough as a process map in Define?
A SIPOC is enough for Define because its job is to set the boundary and show the high-level flow, inputs, outputs, and customers. Detailed process maps, such as swimlane or value stream maps, belong in the Measure phase, once the scope is agreed.
Who signs the charter?
At a minimum, the sponsor (or champion), the process owner, and the project leader. Finance should review the benefit estimate, and the managers who release team members should confirm the time commitment. A charter signed only by the project leader is a proposal, not an agreement.
Sources and Further Reading
Thomas Pyzdek and Paul Keller, The Six Sigma Handbook, chapters on the Define phase.
Michael L. George, David Rowlands, Mark Price, and John Maxey, The Lean Six Sigma Pocket Toolbook.
T. M. Kubiak and Donald W. Benbow, The Certified Six Sigma Black Belt Handbook (ASQ).
Donald W. Benbow and T. M. Kubiak, The Certified Six Sigma Green Belt Handbook (ASQ).
ISO 13053-1, Quantitative methods in process improvement: Six Sigma, Part 1: DMAIC methodology.
ASQ Certified Six Sigma Green Belt and Black Belt Bodies of Knowledge, Define phase.
Noriaki Kano and colleagues, “Attractive Quality and Must-Be Quality,” 1984.
This content is educational. Financial figures in the examples are illustrative, and your organization's policies and finance team determine how benefits are calculated and reported.
Phase 2 of 5
Measure
What is happening now, and how reliably can we measure it? Measure establishes the current performance of the process with data the team and the sponsor can trust. It defines what to measure, proves the measurement system is good enough, collects the data, and describes the baseline well enough to guide the analysis.
Key question
What is happening now, and how reliably can we measure it?
Typical duration
About 2 to 6 weeks, depending on data availability and the time needed to collect it
Led by
The project leader, with the process owner, team members who do the work, and a quality or data analyst
Primary outputs
Data collection plan, validated measurement system, detailed process map, baseline performance, stratified findings, refined problem and goal
Tollgate decision
Sponsor confirms the baseline is credible and the problem is focused enough to analyze
Core tools
Operational definitions, data collection plan, process maps, MSA and gage R&R, run and control charts, capability and sigma level, Pareto, stratification
What the Measure Phase Is For
Define says what the problem is. Measure says how big it really is, where and when it occurs, and whether the numbers can be believed. Teams that skip this phase tend to analyze guesses; teams that rush it often discover later that their data cannot support a conclusion.
What Measure must achieve
Measures defined so that two people would count the same way (operational definitions)
A measurement system proven good enough for the decisions it will support
A detailed picture of how the process actually works, with candidate input variables (Xs)
A baseline: current performance, its stability, and its capability
The problem broken down by where, when, and what type (stratification)
A confirmed or refined problem statement, goal, and scope
What Measure must not do
Jump to causes or solutions (that is Analyze and Improve)
Collect everything that can be measured “just in case”
Trust existing data or gauges without checking them
Report averages without looking at variation and time order
Treat the baseline as fixed when the data reveal a different problem
Skip the process walk and rely on a map drawn in a conference room
The test of a finished Measure. A skeptical colleague should be able to ask “how do you know?” about any baseline number, and the team can answer: how it was defined, how it was measured, how reliable the measurement is, and how much data supports it.
The Measure Process, Step by Step
Validating the measurement system comes before collecting the bulk of the data. Findings in later steps often send the team back to refine definitions or the plan.
1
Select and define the measures
Translate the CTQ from Define into the primary measure (Y). Add process measures and key input measures (Xs) that may influence it. Write an operational definition for each.
Confirm the primary metric from the charter
List candidate process and input measures
Write definitions: what is counted, how, where, by whom, and when
Decide whether each is continuous or discrete data
Output: Operational definitions for every measure
Watch for: Definitions that differ between shifts or departments, and measures chosen only because the data is easy to get
2
Map the process as it actually runs
Go to the place where the work is done and trace the real flow, including rework loops, waiting, and workarounds. Mark where measurements are or could be taken, and list candidate inputs.
Walk the process from start to end with the people who do it
Record steps, times, handoffs, and decision points
Identify value-adding and non-value-adding steps
List potential Xs with a cause-and-effect matrix or diagram
Output: A detailed process map and a list of candidate Xs
Watch for: A map of the procedure as written instead of the process as performed
3
Build the data collection plan
Decide exactly what data to collect, from where, how, how much, by whom, and over what period, including the factors needed to stratify the data later.
Define the sample size and sampling method
Choose stratification factors up front
Design forms or extracts, and train data collectors
Agree the time period and how special events will be recorded
Output: A written data collection plan and test of the forms
Watch for: Collecting data first and thinking about the questions later
4
Validate the measurement system
Before relying on the data, test whether the measurement process is accurate and precise enough. Use gage R&R for continuous data and attribute agreement analysis for pass/fail or category data.
Choose parts or samples that span the real range
Include the operators who normally measure
Randomize order and blind the samples
Analyze, and improve the system if it is not acceptable
Output: Documented measurement system results and any fixes
Watch for: Skipping the study because “the gauge is calibrated”. Calibration is not the same as capability
5
Collect the data
Follow the plan, record time and context for each data point, and check the data as it arrives so that problems are caught early.
Run a short pilot of the collection
Monitor completeness and plausibility daily or weekly
Record changes, breakdowns, and other events as notes
Output: A clean, time-ordered, annotated data set
Watch for: Changing the process while collecting, which changes the thing you are trying to measure
6
Baseline and stratify
Plot the data in time order, check stability, describe the distribution, and calculate performance and capability. Then break the data down by the factors in the plan to find where the problem concentrates.
Run or control charts, histograms, and box plots
Capability or sigma level, with the method stated
Pareto charts and stratified comparisons
Output: The baseline, with evidence of where and when the problem occurs
Watch for: Calculating capability for a process that is not stable, and reporting only an average
7
Refine the charter and hold the tollgate
Update the problem statement, goal, and scope in light of the data. Present the baseline and findings to the sponsor and process owner and agree the focus for Analyze.
Revise the charter where the data call for it
Use the tollgate checklist in this tab
Record decisions and signatures
Output: An updated charter and a recorded tollgate decision
Watch for: Holding on to the original problem statement after the data show something different
Choosing and Defining the Right Measures
A project needs a small, deliberate set of measures. The primary measure (Y) comes from the CTQ in the charter. Supporting measures describe the process and its inputs, and a balancing measure shows whether improving one thing worsens another.
Measure type
What it describes
Example (Line 3 final test)
Output (Y)
The result the customer experiences
First-pass yield at final test; defects per million units at customer receipt
Process (y)
How well a step performs
Test cycle time; retest rate; first-time connector seating
Input (X)
A condition or factor that may drive the output
Fixture wear; firmware version; material lot; operator; shift
Balancing
A watch on side effects
Throughput per hour; rework cost; test coverage
Y = f(X). Inputs feed the process and shape the output. Measure builds the list of candidate Xs; Analyze tests which of them matter.
Operational definitions. An operational definition says exactly how a measure is obtained, so that two people get the same number. It states what is counted, the unit, the rules for inclusion and exclusion, the method and instrument, the location and time, and who records it.
Element
Example: first-pass yield (FPY) at final test
Definition
Units that pass all final test steps on the first attempt, divided by units that enter final test
Unit of count
One serialized unit; a unit retested after a fixture error is still counted as a first-pass failure
Exclusions
Engineering builds and units pulled for a planned audit
Data source
Test station log, extracted daily at shift end
Recorded by
Test system (automatic); verified weekly by the quality technician
Frequency
Daily, reported weekly
Continuous or discrete data? The data type decides the tools you can use, and continuous data usually carry more information per observation than counts.
Data type
Examples
Typical tools
Continuous (variable)
Time, length, weight, temperature, torque
Histograms, I-MR and Xbar-R charts, capability (Cp, Cpk), gage R&R
Discrete: counts of defects
Defects per unit, errors per form
c and u charts, DPU, DPMO
Discrete: proportions
Percent defective, pass/fail, yield
p and np charts, binomial capability, attribute agreement analysis
Categorical
Defect type, machine, shift
Pareto charts, check sheets, stratification
Prefer continuous data when you can. A measured value (for example, the actual insertion force) shows how close a unit is to the limit. A pass/fail result hides that information and needs far more units to detect the same change.
Mapping the Process and Finding the Candidate Xs
The SIPOC from Define showed the boundary. Measure goes inside it. The aim is to understand how the work really happens, where time and defects arise, and which inputs might matter, without deciding yet which ones do.
Map type
Use it when
Strength
Detailed process map (flowchart)
You need the steps, decisions, and rework loops
Simple and widely understood
Swimlane map
Handoffs between people or departments matter
Shows who does what and where delays occur between roles
Prioritizes steps by severity, occurrence, and detection
Walk it. Observe the process on every shift that matters. Differences between shifts are often findings.
Map what happens, not what should happen. Include rework loops, workarounds, and waiting.
Add data to the map: cycle times, queue sizes, yields, and defects at each step.
Mark value. Classify each step as value-adding, necessary non-value-adding, or waste, from the customer's viewpoint.
From the map to candidate Xs. A cause-and-effect matrix rates each input against the customer requirements and ranks the inputs by their likely influence. A fishbone diagram collects possible causes by category. Both are hypotheses to test in Analyze, not conclusions. See the Cause-and-Effect Matrix and the Fishbone Builder.
The Data Collection Plan
The plan turns questions into data. It should be specific enough that someone else could collect the data and get the same results. Write it before collecting, review it with the process owner, and test the forms in a pilot.
Plan element
Question it answers
Line 3 example
Measure and definition
What exactly are we measuring?
First-pass yield at final test, as defined above
Data type
Continuous, discrete, or categorical?
Discrete (pass/fail), plus failure type (categorical)
Source and method
Where does the data come from, and how is it captured?
Test station log; manual coding of failure type by the technician
Sample plan
How much, how often, and how chosen?
All units for 8 weeks; failure types for every failed unit
Stratification factors
What do we want to compare?
Shift, operator, fixture, product variant, material lot, test station
Who and when
Who collects, and over what period?
Test technicians, every shift, weeks 1 to 8
Verification
How do we check the data?
Weekly audit of 20 records against the physical unit and the test log
Sampling. If you cannot measure everything, choose the sample deliberately. Random sampling gives each item an equal chance. Systematic sampling takes every kth item. Stratified sampling samples separately from groups, such as shifts or lots, so that each is represented. Rational subgrouping collects small groups of consecutive units so that variation within a group reflects short-term noise, and variation between groups reveals shifts. See Sampling Methods.
How much data? The answer depends on the question. Two common calculations:
Purpose
Formula
Example
Estimate a proportion to within ±E
n = z² × p(1 − p) / E²
p = 0.088, E = 0.01, 95% confidence (z = 1.96): n ≈ 3,083 units
Estimate a mean to within ±E
n = (z × σ / E)²
σ = 0.14, E = 0.05, z = 1.96: n ≈ 31 measurements
These are planning aids that assume random samples. Use the Sample Size and Confidence Calculator and build the plan with the Data Collection Plan Builder. For a control chart baseline, a common guideline is 20 to 25 subgroups, and for a continuous-data capability study many practitioners look for at least 100 observations across normal operating variation.
Collect context along with each value. Record the date, time, shift, operator, machine, lot, and any unusual event. The context is what lets you stratify later and explain odd points.
Validating the Measurement System
Every observed value is the true value plus measurement error. If the measurement error is large compared with the process variation or the tolerance, the data cannot reveal real differences. A measurement system analysis (MSA) quantifies the error.
Characteristic
Meaning
Question it answers
Accuracy or bias
Difference between the average measurement and the reference value
Does the system read true?
Linearity
Whether bias changes across the measurement range
Is it equally accurate small and large?
Stability
Whether results are consistent over time
Does it drift?
Repeatability
Variation when the same person measures the same item repeatedly with the same equipment
How much noise is in the equipment?
Reproducibility
Variation between operators, methods, or conditions
Do different people get different answers?
Resolution
Smallest increment the system can detect
Can it see differences that matter?
Gage R&R for continuous data. A typical crossed study uses about 10 parts that span the process range, 2 or 3 operators, and 2 or 3 trials each, measured in random order and blinded where possible. The analysis separates repeatability, reproducibility, and part-to-part variation.
Example results. Total gage R&R is 21.6% of study variation, between the commonly used 10% and 30% guideline lines, so the system is marginal and the decision should be weighed against cost and risk.
Result (% of study variation)
Common guideline
Below 10%
Generally acceptable
10% to 30%
May be acceptable depending on the application, cost of the gage, and cost of errors; decide with the sponsor
Above 30%
Not acceptable; improve the measurement system
Number of distinct categories
At least 5 is commonly recommended; this example gives about 6
Attribute data. For pass/fail or rating data, use an attribute agreement analysis. Several appraisers classify the same set of items, ideally including known reference answers, more than once. The analysis shows agreement within each appraiser, between appraisers, and with the standard, usually summarized with kappa statistics. A commonly used guideline treats kappa above 0.75 as good agreement and below 0.40 as poor. See Attribute Agreement Analysis.
If the system fails. Look for simple causes first: an unclear definition, no fixture, poor lighting, inconsistent technique, worn equipment, or insufficient resolution. Fix, retrain, standardize the method, and repeat the study. Do not proceed to baseline data with a system you do not trust.
Establishing the Baseline: Stability, Capability, and Sigma Level
The baseline describes how the process performs now. Always check stability before calculating capability, because capability indices assume a stable, predictable process.
Plot the data in time order on a run chart or control chart. Look for shifts, trends, cycles, and points outside the limits.
Investigate special causes. Note or correct them; do not blend them silently into the baseline.
Describe the distribution with a histogram and box plot, and check for outliers and multiple peaks.
Calculate performance with the appropriate measures for the data type.
State the method and period so that the baseline can be reproduced.
Continuous data: capability indices. They compare the spread of the process with the specification.
Index
Formula
Meaning
Cp
(USL − LSL) / 6σ
Potential capability if the process were centered
Cpk
min[(USL − μ) / 3σ, (μ − LSL) / 3σ]
Actual capability, accounting for centering
Pp, Ppk
Same formulas using the overall (long-term) standard deviation
Performance over the whole data period, including shifts
Illustrative example: specification 9.5 to 10.5, mean 10.12, standard deviation 0.14. Cp = 1.19 but Cpk = 0.90 because the process is shifted toward the upper limit, which accounts for about 0.3% of output out of specification (assuming normality).
Check the assumptions. The simple indices assume a stable process and, for the percent-out-of-specification estimate, an approximately normal distribution. For non-normal data, use a transformation or a method designed for the distribution. See Process Capability, Non-Normal Capability, and the Process Capability Helper.
Discrete data: defect rates and sigma level.
Measure
Formula
Example
Yield (first-pass)
Good units / units entering
FPY at Line 3 final test = 91.2%
Defects per unit (DPU)
Defects / units
0.15 defects per unit
DPMO
Defects / (units × opportunities) × 1,000,000
150 defects in 1,000 units with 5 opportunities each: 30,000 DPMO
Rolled throughput yield (RTY)
Product of the yields of each step
Steps at 98.5%, 97.5%, and 95.0% give RTY of 91.2%
Sigma level
z-score of the yield, plus 1.5 by convention for the long-term shift
Report the baseline with its context. Write the metric, definition, period, sample size, data source, stability finding, and capability or sigma level together. A baseline number without those details cannot be compared with the result at the end of the project.
Stratification: Where Does the Problem Concentrate?
An average hides structure. Stratification splits the data by a factor that might matter, such as shift, machine, operator, product, supplier, lot, location, time of day, or failure type, and compares the groups. It is the main route from a broad baseline to a focused problem.
Illustrative Pareto chart of 8,448 final-test failures over eight weeks. Connector seating and firmware load errors account for 66% of failures, so the project focuses there. Build one with the Pareto Chart Builder.
Stratified by
Result (illustrative)
What it suggests
Shift 1
7.0% failed
Baseline performance
Shift 2
8.4% failed
Slightly above average
Shift 3
11.0% failed
Noticeably higher: investigate what differs (staffing, procedure, fixture handling, material)
Average
8.8% failed
Matches the overall baseline
Pareto analysis ranks categories and shows the vital few. See the Pareto Analysis guide.
Run charts by group show whether a difference is persistent or a one-off.
Multi-vari charts separate variation within a part, between parts, and over time, to show where the variation lives. See Multi-Vari Studies.
Box plots and histograms by group compare center and spread for continuous data.
Strata are hypotheses, not causes. Shift 3 failing more often does not show that its operators are the cause. It may use a different fixture, material, or supervision pattern. Analyze tests the explanation. Measure only reveals where to look.
Refine the problem. The stratified view often narrows the project considerably: “first-pass yield at final test is 91.2%” becomes “connector seating and firmware load errors, concentrated on shift 3, account for two thirds of failures.”
Data Quality and Common Data Problems
Problem
Symptom
Response
Inconsistent definitions
Different people count differently
Rewrite the operational definition; retrain; re-collect if needed
Missing data
Gaps by time, shift, or operator
Find out why; do not fill with averages without a reason; record the limitation
Outliers
Isolated extreme values
Check the record and context; correct errors; keep genuine events and explain them
Rounding and resolution
Data clustered on a few values
Use a finer scale or a better instrument
Sampling bias
Data collected only when convenient or when someone is watching
Randomize or stratify the sampling; collect on all shifts
Observer effect
Performance improves only while measuring
Collect unobtrusively; extend the period
Transcription errors
Typos and mismatched records
Audit a sample against source; use automatic capture where possible
Data from a changed process
A shift in the data after a known change
Use only data from the current process, or analyze the periods separately
Check the data early and often. Plot it as it arrives. Compare a sample of records with the physical items. Ask the people who collect it what is hard. Problems found in week one cost little; problems found in week eight can cost the schedule.
Revisiting the Charter at the End of Measure
The baseline is the first time the project meets real data. The charter may need to change, and changing it by agreement is a sign of a healthy project.
Charter element
What the Measure data may show
Possible change
Problem statement
The real size or location differs from the assumption
Restate with the confirmed baseline and focus (for example, the top two failure types)
Goal
The original target was set on estimates
Reset to a target justified by the baseline, customer need, or demonstrated best performance
Scope
The problem concentrates in one area, or the process boundary was wrong
Narrow, widen, or split into two projects
Business case
The size of the cost differs from the estimate
Update with finance and confirm the project is still worth doing
The Measure tollgate confirms that the baseline can be trusted and that the problem is focused enough to analyze. Use the checklist before the review. Progress saves in this browser only, and nothing is sent anywhere.
0 of 14 complete
Questions a sponsor should ask:
How do we know the measurements are reliable?
Is the process stable, and what does the baseline tell us about the size of the gap to the goal?
Where does the problem concentrate, and what do we not yet know?
Has anything in the data changed what we should be working on?
What data limitations or risks should we keep in mind?
Do the goal and the business case still hold?
Tollgate outcome
Meaning
Typical next step
Go
The baseline is credible and the focus is clear
Start Analyze
Go with conditions
Small gaps in data or documentation
Record owners and dates; confirm before analysis
Return to Measure
Measurement system or data cannot support conclusions
Fix and re-collect
Re-scope or stop
The data show the problem is different, smaller, or not worth pursuing
Revise the charter or end the project
Adapting Measure to the Situation
Situation
How Measure changes
Manufacturing with automated data
Data may be abundant; focus on validating it, understanding its context, and sampling meaningfully rather than collecting more
Manufacturing with manual inspection
Attribute agreement analysis and clear operational definitions matter most; train and audit inspectors
Service and transactional
Time stamps, error counts, and category data dominate; many “gages” are people, so test agreement; timing data often must be created
Healthcare and public service
Take care with privacy and consent; define events carefully; small volumes may suit run charts better than capability indices; see the healthcare and government hubs
Software and IT
Define events precisely (what counts as an incident or defect); use flow measures such as cycle time and WIP; see the Software and IT hub
Lean flow problems
Add lead time, cycle time, takt, WIP, and value-added ratio; use Little's Law to link them
Small samples or rare events
Plot every event over time; use measures such as time between events; be open about the uncertainty
When data cannot be collected quickly. A short manual data collection with a well-designed form usually beats waiting for a system change. If a proxy measure must be used, state clearly how it relates to the true measure and its limits, and plan to check it later.
Common Mistakes and Red Flags
Mistake
What it looks like
How to correct it
Skipping the measurement system analysis
“The gauge is calibrated, so it is fine”
Run gage R&R or attribute agreement; calibration shows bias, not precision
Weak operational definitions
Different people report different numbers for the same event
Write the definition; test it with two people; retrain
Capability of an unstable process
Cpk reported from data with obvious shifts
Investigate special causes first; report stability alongside capability
Averages without variation
Only the mean is reported
Show the distribution, time order, and spread
Too much data, too little thought
Spreadsheets of unused measures
Collect what answers the planned questions
Too little data
Conclusions from a handful of points
Size the sample; show confidence limits or ranges
No stratification plan
Context not recorded, so the data cannot be split
Decide the factors in advance and record them
Changing the process while measuring
Improvements made before the baseline is complete
Hold changes except for containment; record any change
Ignoring what the data say
The original problem statement is defended against the evidence
Update the charter with the sponsor
Perfectionism
Weeks spent polishing data that is already adequate
Time-box; document limitations; proceed when the checklist is met
Why check the measurement system before collecting baseline data?
Because every number in the project depends on it. If the measurement system adds a large amount of variation, the baseline will be blurred, real differences will be hidden, and later tests may reach wrong conclusions. A measurement system analysis is far cheaper than discovering after months of data that the gauge or the data entry could not be trusted.
How much data do I need for a baseline?
It depends on the data type and the purpose. For a control chart baseline, a common guideline is 20 to 25 subgroups. For a capability study of continuous data, many practitioners look for at least 100 observations collected across normal variation. For a proportion, the sample size follows from the precision you need: estimating a defect rate near 9% to within one percentage point at 95% confidence takes about 3,000 units. Treat these as starting points and justify the choice in the data collection plan.
What if the process is not stable when I measure it?
Then the baseline is a description of an unstable process, and capability indices are not meaningful. Investigate the special causes first, using the time order of the data, and either fix them or document them. A stable baseline is the foundation for the Analyze and Improve phases. If instability is the project, record the pattern as part of the baseline.
Do I need to calculate sigma level for every project?
No. Sigma level is a convenient way to compare processes, but a yield, defect rate, cycle time, or capability index may describe the problem better. Choose the metric that matches the CTQ in the charter. If you report a sigma level, state how it was calculated, including whether the conventional 1.5-sigma shift was used.
Can I proceed if the gage R&R result is above 30%?
Not without action. A result above 30% of study variation usually means the measurement system is not acceptable for the decision. Improve it (training, fixtures, a clearer procedure, a better gage, or a different method) and repeat the study. In some cases a result between 10% and 30% can be acceptable if the cost and the risk are considered, and that judgment should be recorded with the sponsor's agreement.
What if no data exists for the metric?
Create it. Design a simple data collection form, define the metric operationally, train the people who will record it, and collect for a defined period. If waiting is not possible, use a proxy measure and state its limits. A short, well-designed manual collection often provides better data than a large existing database nobody trusts.
Sources and Further Reading
Thomas Pyzdek and Paul Keller, The Six Sigma Handbook, chapters on the Measure phase.
Automotive Industry Action Group (AIAG), Measurement Systems Analysis Reference Manual, for gage R&R guidelines and attribute studies.
Douglas C. Montgomery, Introduction to Statistical Quality Control, on control charts and process capability.
Donald J. Wheeler, Understanding Variation and Making Sense of Data.
T. M. Kubiak and Donald W. Benbow, The Certified Six Sigma Black Belt Handbook (ASQ).
ISO 13053-1 and ISO 22514 (statistical methods in process management: capability and performance).
NIST/SEMATECH e-Handbook of Statistical Methods.
This content is educational. The example data and results are illustrative. Use the acceptance criteria required by your customers, standards, and organization.
Phase 3 of 5
Analyze
Why is the problem happening? Analyze turns the baseline into explanations. It generates and prioritizes hypotheses, tests them with data and with direct observation, and identifies the few verified causes (the vital few Xs) that drive the gap between current and target performance.
Key question
Why is the problem happening?
Typical duration
About 3 to 6 weeks, depending on the number of hypotheses and the time needed to collect data or run trials
Led by
The project leader, with team members who know the process, the process owner, and a quality or data analyst
Primary outputs
Prioritized hypotheses, verified root causes with evidence, quantified contribution of each cause to the gap, and an updated business case
Tollgate decision
Sponsor confirms the causes are credible and sufficient to close the gap, and that the team may move to solutions
Core tools
Fishbone, 5 Whys, cause-and-effect matrix, FMEA, process analysis, hypothesis tests, regression, ANOVA, multi-vari, designed experiments
What the Analyze Phase Is For
Measure shows how large the problem is and where it concentrates. Analyze explains why. The risk in this phase is jumping to a favorite theory, or accepting a plausible story without evidence, and then spending Improve on the wrong fix. The discipline of Analyze is to require evidence for each cause before acting on it.
What Analyze must achieve
A broad, structured list of possible causes, built with the people who know the process
A prioritized short list of hypotheses to test
Evidence for or against each, from data, observation, and small tests
A distinction between symptoms and root causes
The contribution of each verified cause to the gap to the goal
A reviewed business case and a clear brief for the Improve phase
What Analyze must not do
Declare a cause because it is plausible, familiar, or favored by a senior person
Stop at a symptom or at a person
Treat a correlation as proof of cause
Run many tests and report only the significant ones
Implement a full solution before the sponsor has accepted the causes
Continue analyzing past the point of usefulness
The test of a finished Analyze. For each cause in the report, the team can answer: what is the evidence, what is the mechanism, how much of the gap does it explain, and what would we expect to see if we fixed it?
The Analyze Process, Step by Step
Testing and verifying loop: results often generate new hypotheses or eliminate old ones, and the team returns to earlier steps.
1
Review the baseline and the stratified findings
Start from what Measure showed: where and when the problem concentrates, which categories dominate, and what the process map revealed. These are the clues the hypotheses should explain.
Re-read the Pareto charts, run charts, and stratified comparisons
List what is known, what is not, and what changed when the problem began
Check the time of first occurrence against known changes
Output: A short summary of the clues that any explanation must fit
Watch for: Ignoring evidence that does not fit a favored theory
2
Generate hypotheses broadly
Use structured methods and the people who run the process to list every plausible cause, before narrowing. Separate hypotheses about the process (how work flows) from those about data (what the numbers show).
Run a cause-and-effect (fishbone) session
Drill into the top defect types with 5 Whys
Use the process map and FMEA to find where failures can arise
Compare good and bad cases with is / is-not analysis
Output: A full list of hypotheses, grouped by category
Watch for: Letting one voice dominate and recording only the first idea
3
Prioritize the hypotheses
Rank by likelihood and by potential influence on the problem, using the data and process knowledge. Decide what to test first and how.
Use a cause-and-effect matrix or voting to rank
Link each hypothesis to a test or observation
Start with tests that are quick, cheap, and decisive
Output: A ranked test plan
Watch for: Testing only the hypotheses that are easy to test
4
Analyze the process
Look for causes in how the work flows: waiting, rework loops, handoffs, bottlenecks, variation in method, and non-value-adding steps.
Value-added and non-value-added analysis
Cycle-time and queue analysis
Comparison of how different shifts or people do the task
Output: Process-based findings, supported by observation and times
Watch for: Relying on the documented procedure instead of what happens at the workstation
5
Analyze the data
Use graphs first, then statistical tests where needed, to compare groups, show relationships, and separate real effects from random variation.
Choose the test by the data type of the Y and the X
Check assumptions and sample sizes
Report effect sizes and intervals, not only p-values
Output: Evidence for or against each hypothesis
Watch for: Running every possible test and picking the significant ones
6
Verify the causes
Confirm that the apparent cause is real. Observe the mechanism, check that the timing fits, and, where practical, change the factor on a small scale and watch the effect.
Go to the process and see the mechanism
Run a small confirming trial or designed experiment
Check for alternative explanations and confounding
Output: Verified root causes and the evidence for each
Watch for: Accepting a statistical association without a mechanism or a test
7
Quantify the contribution and hold the tollgate
Estimate how much of the gap each verified cause explains and what the business case looks like when they are fixed. Present the causes and evidence to the sponsor and process owner.
Calculate the excess defects or time attributable to each cause
Update the business case with finance
Use the tollgate checklist in this tab
Output: A gap-closure estimate and a recorded tollgate decision
Watch for: A list of causes with no estimate of their size
Generating Hypotheses with Structure
A hypothesis in Analyze is a testable statement of a cause: “Failures at connector seating are higher on fixture 3 because its contact pins are worn.” Good hypotheses are specific, tied to a process step, and testable with data or observation.
A fishbone for Line 3 final-test failures. Each category holds possible causes to test, not conclusions. Build one with the Fishbone Builder.
Symptoms, causes, and root causes. A symptom is what is observed (“connector not seated”). A cause explains it (“contact pins on fixture 3 are worn”). A root cause is the underlying condition that, if corrected, prevents recurrence (“the pin replacement interval was missed because the preventive maintenance task was skipped in a holiday week and has no check”). Keep asking why while the answer is something the team can act on, and verify each link.
Use the people who do the work. The best hypotheses often come from operators, technicians, and maintainers. A hypothesis session with only managers and engineers will miss them.
Prioritizing What to Test
A fishbone can produce fifty causes. Testing all of them is wasteful. Prioritize by how likely each is to be true, how much it could influence the problem, and how easy it is to test.
Hypothesis
Likelihood
Potential influence
Ease of test
Priority
Fixture 3 contact pins worn
High (failures concentrated on one fixture)
High (largest failure type)
Easy (inspect, compare fixtures)
Test first
Firmware v2.3 load timeout
High (timing fits week 14)
High
Easy (compare stations by version)
Test first
Sensor supplier lot B
Medium
Medium
Easy (trace by lot)
Test first
Temperature at the test station
Low (no pattern by time of day)
Low
Medium
Defer
Training on fixtures
Medium
Medium
Hard to isolate
Test after the others
A cause-and-effect matrix (rating each input against the customer requirements and weighting by importance) gives a transparent ranking, and the Impact-Effort Matrix Builder can help balance influence against the effort of testing. An FMEA ranks failure modes by severity, occurrence, and detection; use the FMEA RPN and Action Priority Tool.
Link every hypothesis to a test. For each, write what you will look at, what result would support it, and what result would reject it, before you look. This protects against reading the data to fit the theory.
Process Analysis: Causes in the Flow of Work
Some causes show up in numbers; many are visible only by watching the process. Process analysis looks for causes in how work moves: where it waits, where it loops back, where handoffs lose information, and where methods vary between people.
What to look for
How to analyze it
Examples of findings
Non-value-adding steps
Classify each step from the customer's view as value-adding, necessary non-value-adding, or waste; see the 8 Wastes guide
Duplicate inspections; searching for tools; waiting for approval
Waiting and queues
Measure queue time and work in process; apply Little's Law (lead time = WIP / throughput)
Statistical tests answer narrow questions. Before running any, plot the data and ask what the picture suggests. Choose the technique by the type of data in the outcome (Y) and the factor (X).
A simple guide to choosing an analysis. Nonparametric alternatives exist for data that do not meet the assumptions of the tests shown. The Hypothesis Testing Quick Tester covers common cases.
Before applying a test, check its assumptions: independent observations, adequate sample size, and, for many tests, approximately normal data or similar spread between groups. Use the normality testing and plots for that check.
Hypothesis Testing in Practice
A hypothesis test asks whether the difference in the data is larger than chance would plausibly produce. The null hypothesis states there is no difference or effect. The alternative states that there is. The p-value is the probability of seeing a result at least as extreme as the data if the null hypothesis were true.
Concept
Meaning
Practical note
Alpha (significance level)
The risk of a false alarm that you will accept, commonly 0.05
Set before looking at the data
p-value
Probability of data this extreme if there is no real effect
A small p-value is evidence against the null; it does not measure size or importance
Power
The chance of detecting a real difference of a given size
Usually planned at 0.80 to 0.90; low power can miss real causes
Confidence interval
A range of plausible values for the effect
Shows both direction and size; more informative than p alone
Effect size
How large the difference is in useful units
Decide in advance the smallest difference that matters
Statistical versus practical significance. With very large samples, tiny differences become statistically significant. With small samples, large differences can fail to reach significance. Always ask whether the effect is large enough to matter for the CTQ, and whether the sample was large enough to detect an effect of that size. See the Hypothesis Testing guide for a full treatment and the Sample Size and Confidence Calculator for planning.
Do not hunt for significance. If you run 20 independent tests at alpha 0.05, you should expect about one to appear significant by chance. Decide the tests in advance, link them to hypotheses, and report all of them, including those that did not support the theory.
Relationships Between Variables: Correlation, Regression, and Experiments
When the factor and the outcome are both continuous, a scatter plot shows whether they move together. Regression fits a line (or a more complex model) and estimates how much the outcome changes per unit change in the factor.
Output
How to read it
Caution
Correlation coefficient (r)
Between −1 and +1; shows the strength and direction of a linear relationship
Zero does not mean no relationship; curves and outliers distort it
Slope
Change in Y per unit change in X
Valid only within the range of the data
R-squared
Share of the variation in Y explained by the model
A high value does not prove cause; a low value can still be useful
Residual plots
Show what the model misses
Patterns in residuals mean the model is wrong or incomplete
Multiple regression
Estimates the effect of several Xs at once
Beware of correlated inputs and over-fitting
Correlation is not causation. Two measures can move together because one causes the other, because a third factor drives both, because of reverse causation, or by coincidence. Historical (observational) data can suggest causes but rarely prove them.
Designed experiments establish cause. When you deliberately change factors in a planned pattern and randomize the runs, differences in the outcome can be attributed to the factors and their interactions. Use a designed experiment when historical data cannot separate causes, when factors interact, or when you need to know the effect of a change before making it. See the Design of Experiments guide, the DOE Quick Planner, and Design of Experiments.
Verifying a Cause Before You Act on It
Before a cause goes into the report, put it through several checks. The more of them it passes, the more confidence you can have.
Check
Question
Example
Association
Is the cause present when the effect is present, and absent when it is absent?
Failures concentrate on fixture 3
Timing
Did the cause begin before the effect?
Pin replacement interval was passed in week 13; failures rose in week 14
Dose-response
Does more of the cause give more of the effect?
Failure rate rises with pin wear measured on each fixture
Mechanism
Is there a plausible physical or logical explanation, and can you see it?
Worn pins leave connectors partly unseated
Alternatives
Have other explanations been ruled out, including confounders?
Shift effect disappears once fixture use is accounted for
Reversal or trial
Does changing the cause change the effect?
Replacing pins on fixture 3 returns its rate to the others' level
Replication
Does it hold in other data, places, or periods?
The same pattern appears in earlier weeks
Go and see. Observe the mechanism where it happens. A short visit to the workstation, with a camera or gauge, often turns a statistical result into an understanding. Test small. A confirming trial on one fixture or one station is cheap and decisive, and it is part of verification, not the Improve phase's full implementation.
Worked Example: Line 3 Final-Test Yield
This example continues the project from Define and Measure. The baseline was a first-pass yield of 91.2% against a goal of 96.0%. The Pareto chart showed that connector seating (40% of failures), firmware load errors (26%), and intermittent sensor failures (16%) accounted for 82% of the 8,448 failures in eight weeks. All figures are illustrative.
The run chart shows a step change in week 14, not a gradual decline. That points to something that changed, which is a strong clue for hypotheses.
Hypotheses and tests. The team lists changes made around week 14 and builds a fishbone for each of the three failure types. It prioritizes the hypotheses that fit the timing and tests each in turn.
Hypothesis 1: contact pin wear on one fixture. Plotting connector-seating failures by fixture shows a clear outlier.
Fixture 3 fails at 7.2% against about 2.3% for the others (chi-square test p < 0.0001 across 24,000 units per fixture). The pattern is a lead; the mechanism and a trial come next.
Inspection shows the contact pins on fixture 3 are beyond their wear limit, and maintenance records show the pin replacement interval was passed in week 13, when a preventive maintenance task was skipped during a holiday week. A trial replacing the pins on fixture 3 returns its failure rate to 2.4% within three days.
The shift puzzle. In Measure, shift 3 had the highest failure rate (11.0%). The fixture finding explains why: fixture 3 ran 45% of shift 3's volume, against 20% and 10% on shifts 2 and 1. The apparent shift effect was largely a fixture effect. This is a common example of confounding, and it shows why a stratified difference is a hypothesis and not a cause.
Hypothesis 2: the firmware change. Three of the six test stations moved to firmware v2.3 in week 14, and three remained on v2.2 for logistical reasons. Over the same weeks, shifts, and products, firmware load failures were about 4.1% on v2.3 stations and about 0.5% on v2.2 stations. The station logs show load timeouts, and rolling one station back to v2.2 reduces its rate to 0.6%.
Hypothesis 3: sensor lot. Traceability shows that 30,000 of the 96,000 units used sensors from a new supplier lot, introduced in week 14. Intermittent sensor failures were about 3.9% on lot B units and about 0.3% on all others. Incoming tests of lot B parts reproduce the intermittent behavior.
Quantifying How Much Each Cause Explains
Finding causes is not enough; the team must show that they are large enough to reach the goal. For each verified cause, estimate the excess defects it produces, which is the observed defects minus the number expected without it.
Verified cause
Excess failures (8 weeks)
Points of FPY
Evidence
Fixture 3 contact pin wear
1,178
1.23
7.2% vs 2.3% on other fixtures; worn pins measured; pin replacement trial returned the rate to 2.4%
Firmware v2.3 load timeouts
1,716
1.79
4.1% on v2.3 stations vs 0.5% on v2.2 stations in the same weeks; logs show timeouts; rollback trial returned 0.6%
Sensor lot B
1,108
1.15
3.9% on lot B vs 0.3% on other lots; incoming test reproduces the fault
Total explained
4,002
4.17
FPY would be about 95.4% if all three were eliminated
Three verified causes account for 4.17 of the 4.8 points needed. The remaining 0.63 points will need further work on smaller categories such as solder defects.
Update the business case. Each percentage point of yield at 600,000 units a year and $6.40 per reworked unit is worth about $38,400. The three causes are therefore worth about $160,000 a year, close to the $184,000 target in the charter; the remainder depends on the smaller causes. Review the figures with finance and decide with the sponsor whether to pursue the remaining gap in Improve.
If the verified causes explain too little. Either the goal needs to be reconsidered, more causes need to be found, or the project may need to be split. Do not let the team declare victory with causes that account for a small part of the gap.
Analytical Pitfalls to Avoid
Pitfall
What happens
Protection
Confirmation bias
The team looks only for evidence that supports a favored cause
Write what would disprove each hypothesis; have someone argue the opposite
Confounding
A third factor explains an apparent effect (shift vs fixture)
Stratify by other factors; use designed experiments
Simpson's paradox
A trend in each group reverses when groups are combined
Look at data both combined and split by key factors
Data dredging
Many tests until something is significant
Decide tests in advance; adjust for multiple tests; replicate
Survivorship and selection bias
Only the units that were recorded or kept are analyzed
Check how data were selected and what is missing
Over-fitting
A complex model fits noise
Keep models simple; check with new data
Extrapolation
Predictions outside the range of the data
Restrict conclusions to the observed range; test beyond it
Authority
The most senior opinion is accepted without evidence
Ask for the evidence from everyone, politely and consistently
Analysis paralysis
Endless analysis with no decision
Time-box; stop when the main part of the gap is explained
Tollgate Review: Are We Ready for Improve?
The Analyze tollgate confirms that the causes are credible and large enough to justify solutions. Use the checklist before the review. Progress saves in this browser only, and nothing is sent anywhere.
0 of 12 complete
Questions a sponsor should ask:
What is the evidence for each cause, and how was it checked?
How much of the gap does each cause explain, and how much remains?
Could something else explain the same pattern?
Have we seen the mechanism with our own eyes?
What did we test that turned out not to be a cause?
Do the findings change the goal or the business case?
Tollgate outcome
Meaning
Typical next step
Go
Causes are verified and explain enough of the gap
Start Improve
Go with conditions
A cause needs a small confirming trial or a data check
Complete early in Improve; record owners and dates
Return to Analyze
Causes explain too little, or evidence is weak
More hypotheses, data, or a designed experiment
Re-scope or stop
The causes are outside the project's reach or not worth fixing
Revise the charter or end the project
Adapting Analyze to the Situation
Situation
How Analyze changes
Manufacturing with rich data
Stratified comparisons, ANOVA, regression, and multi-vari charts; confirm with trials or designed experiments
Low-volume or high-mix production
Pool data carefully; use rate-based measures; rely more on process observation and failure analysis of individual events
Service and transactional
Process analysis of waiting, handoffs, and rework dominates; use time stamps and error categories; interview staff and customers
Healthcare and public service
Combine data with observation and staff and patient or resident accounts; take care with privacy and equity; see the healthcare and government hubs
Software and IT
Use incident and defect data, timelines, and postmortem methods; see the Software and IT hub
Small samples or rare events
Case-by-case root cause analysis (5 Whys, fault trees), with timelines; avoid over-interpreting small differences
Lean flow problems
Value stream analysis, bottleneck analysis, and queue theory show where lead time accumulates
Qualitative evidence counts. Interviews, observations, photographs, and failed-part examinations are legitimate evidence. Record them with dates and sources, and use them with the data, not instead of it.
Common Mistakes and Red Flags
Mistake
What it looks like
How to correct it
Jumping to a solution
The team starts designing fixes before causes are verified
Park the ideas; require evidence for each cause
Stopping at the first “why”
The cause is a symptom or a person
Continue until the answer is a process or system condition the team can change
One cause fits all
A single cause is claimed to explain everything
Quantify what it explains; look for others
Correlation treated as cause
A scatter plot is presented as proof
Look for the mechanism; test by changing the factor
Significance without size
A tiny but significant difference is reported as important
Report effect size and compare with the CTQ
No test plan
Analyses are run until something appears
Plan tests in advance, linked to hypotheses
Ignoring process observation
All analysis is done from a desk
Go to the process and see it
Not quantifying contribution
A list of causes with no sizes
Estimate excess defects or time for each cause
Analysis never ends
More data and more tests with no decision
Time-box and stop when the main part of the gap is explained
How do I know I have found the root cause and not a symptom?
Ask whether fixing it would stop the problem, and whether you can show it with evidence. A root cause is a condition that, if changed, would prevent the effect, and it is supported by more than one kind of evidence: data showing the association, a plausible mechanism, timing that fits, and ideally a small test in which changing the cause changes the result. Asking “why” until the answer is something the team can act on, and checking each link with evidence, helps you avoid stopping at a symptom.
What if the data show a correlation but I cannot explain why?
Treat it as a lead, not a finding. Correlation can come from a hidden third factor, a coincidence in timing, or reverse causation. Look for a mechanism by observing the process, check whether the relationship holds in other data, and if possible test it by deliberately changing the factor on a small scale or with a designed experiment.
Do I always need statistical tests?
No. Many causes are clear from a good chart, a process observation, or a simple comparison, and a Pareto or run chart may be enough. Use hypothesis tests when the difference is not obvious, when the decision is costly, or when you must show that an apparent difference is unlikely to be chance. Statistical significance alone is never enough; the size of the effect and the mechanism matter as well.
What if there are many possible causes?
Prioritize. Use process knowledge, a cause-and-effect matrix, FMEA, and the stratified data from Measure to rank hypotheses, and test the most likely and most influential first. Most problems are driven by a few causes, so quantify how much of the gap each verified cause explains and stop adding hypotheses when the main part of the gap is accounted for.
When should I use a designed experiment instead of historical data?
When historical data cannot separate causes because factors move together, when you need to know the effect of changing a factor to a level the process has not used, when several factors may interact, or when it is safe and affordable to run trials. Designed experiments are often the strongest way to establish cause, and they are most often used in Analyze to confirm and in Improve to optimize.
Can I start working on solutions during Analyze?
Record ideas as they arise, but do not commit to them until causes are verified. Small trials used to test a cause (for example, replacing a worn part on one fixture to see whether the failure rate drops) are part of the analysis. Full implementation belongs in Improve, after the sponsor has agreed the causes at the tollgate.
Sources and Further Reading
Thomas Pyzdek and Paul Keller, The Six Sigma Handbook, chapters on the Analyze phase.
Douglas C. Montgomery, Design and Analysis of Experiments, and Introduction to Statistical Quality Control.
Douglas C. Montgomery and George C. Runger, Applied Statistics and Probability for Engineers.
George E. P. Box, J. Stuart Hunter, and William G. Hunter, Statistics for Experimenters.
Donald J. Wheeler, Understanding Variation.
T. M. Kubiak and Donald W. Benbow, The Certified Six Sigma Black Belt Handbook (ASQ).
Judea Pearl and Dana Mackenzie, The Book of Why, for a modern treatment of causation.
NIST/SEMATECH e-Handbook of Statistical Methods.
This content is educational. The example data and results are illustrative. Use qualified statistical and engineering judgment for decisions that affect safety, regulated products, or large investments.
Phase
Improve
What changes will remove the causes and improve performance?
This tab is being built. The Improve toolbox will follow the same layout as the Define tab: an overview, step-by-step guidance, the core tools, a tollgate checklist, common mistakes, and links to the site's calculators and templates. In the meantime, see the DMAIC Roadmap guide and the DMAIC overview.
Phase
Control
How will we hold the gains and prevent backsliding?
This tab is being built. The Control toolbox will follow the same layout as the Define tab: an overview, step-by-step guidance, the core tools, a tollgate checklist, common mistakes, and links to the site's calculators and templates. In the meantime, see the DMAIC Roadmap guide and the DMAIC overview.