Choose a phase tab below. Each tab opens with an overview, then covers the details needed to run that phase well, with links to the calculators, templates, and guides on this site. The Define, Measure, and Analyze tabs are available now; the other phases are being added one at a time.

Phase 1 of 5

Define

What problem are we solving, and why does it matter? Define turns a vague concern into a clear, agreed, and measurable project: a problem worth solving, a boundary around it, a customer-based definition of success, a business case, and the people to do the work.

Key question
What problem are we solving, and why does it matter?
Typical duration
About 1 to 4 weeks, depending on project size and how much is already known
Led by
The project leader (Green or Black Belt), with the sponsor and process owner
Primary outputs
Signed project charter, SIPOC, customer requirements (CTQs), business case, team and stakeholder plan
Tollgate decision
Sponsor confirms the project is worth doing, correctly scoped, resourced, and ready for Measure
Core tools
Project charter, SIPOC, VOC and CTQ tree, Kano model, COPQ, stakeholder analysis, RACI, impact-effort matrix

What the Define Phase Is For

Most improvement projects that fail were not defeated by difficult statistics. They were defeated at the start: the problem was vague, the scope kept growing, the sponsor was not committed, or nobody agreed what success meant. Define exists to prevent that. It is the phase in which the project earns the right to consume time and money.

What Define must achieve

  • A problem statement that is specific, measured, and free of causes and solutions
  • A goal that is measurable, time-bound, and tied to a baseline
  • A scope that names what is in, what is out, and where the process starts and stops
  • Customer requirements translated into measurable critical-to-quality (CTQ) characteristics
  • A business case that finance has reviewed
  • A sponsor, process owner, and team who have committed time and authority

What Define must not do

  • Select a solution or assume a cause
  • Collect large amounts of detailed data (that is Measure)
  • Map every step of the process in detail
  • Start a project because of a loud opinion instead of evidence
  • Proceed without a committed sponsor
  • Treat the charter as paperwork to be signed once and forgotten
The test of a finished Define. Someone who has never heard of the project should be able to read the charter and answer four questions: what is wrong, how big is it, who is affected, and what will be different when we are done?

Define is also where the improvement method is chosen. Not every problem needs DMAIC. The next sections cover that decision, and each of the steps in order.

The Define Process, Step by Step

Identify Problem and opportunity Select Right project, right method Understand Customer and process Quantify Impact and baseline Charter Scope, goal, team Tollgate Sponsor decision
The steps overlap and loop back. Customer and scope work often changes the problem statement, which is a sign Define is doing its job.
  1. Identify the problem or opportunity

    Start from evidence: customer complaints, scrap and rework reports, missed deliveries, audit findings, safety events, cost trends, or a strategic gap. Write down who raised it, what they observed, and what data they used.

    • Collect the raw signals: complaints, defect and downtime data, cost reports, customer scorecards
    • Ask the people who raised the concern for examples and dates
    • Check whether the problem is chronic and recurring or a one-time event

    Output: A short list of candidate problems with the evidence behind each

    Watch for: Problems described only as opinions (“quality is bad”) and problems already caused by a known, fixable event

  2. Select and prioritize the project

    Compare candidates against agreed criteria: strategic fit, customer impact, financial impact, feasibility, and sponsor support. Decide whether DMAIC is the right method (see the next section).

    • Score candidates with a prioritization matrix
    • Check that the root cause is not already known
    • Confirm that the project can be completed in roughly three to six months

    Output: One selected project, a recorded reason for choosing it, and the method chosen

    Watch for: Choosing the project that is easiest to explain, not the one with the best evidence and value

  3. Understand the customer and the process at a high level

    Identify the customers of the process output, listen to them, and translate what they say into measurable requirements. Draw a SIPOC to agree the boundary and flow.

    • Gather the voice of the customer from interviews, complaints, surveys, and data
    • Classify needs with the Kano model
    • Build a CTQ tree
    • Draw the SIPOC with the team and the process owner

    Output: CTQs with measurable definitions, and a SIPOC

    Watch for: Assuming you know what the customer wants without asking, and drawing the process as it should be rather than as it is

  4. Quantify the impact and the baseline

    Put numbers on the problem: how often, how large, since when, and what it costs. Use the data that already exists. A preliminary baseline is enough; detailed measurement belongs in the next phase.

    • Pull existing data on the primary metric for a representative period
    • Estimate the cost of poor quality and the benefit if the goal is met
    • Ask finance to review the benefit method

    Output: A preliminary baseline, a benefit estimate, and agreement on how the benefit will be validated

    Watch for: Benefit estimates that count every cost as savings, or that finance has never seen

  5. Draft the charter and form the team

    Write the charter: problem statement, goal, scope, business case, constraints, risks, timeline, and roles. Name the sponsor, process owner, project leader, and team members, and confirm their time with their managers.

    • Draft the problem and goal statements and test them with the sponsor
    • Choose team members who know the process and can act on findings
    • Analyze stakeholders and plan communication

    Output: A draft charter, a named team, a stakeholder analysis, and a communication plan

    Watch for: A team assembled by availability rather than knowledge, and a charter written alone by the project leader

  6. Hold the tollgate review

    Present the charter, SIPOC, customer requirements, baseline, benefit case, and plan to the sponsor and process owner. The aim is a decision: proceed, proceed with conditions, re-scope, or stop.

    • Use the tollgate checklist in this tab
    • Record decisions and actions
    • Obtain signatures

    Output: A signed charter and a recorded go, conditional go, re-scope, or stop decision

    Watch for: A review that is a formality. A project stopped at this point is a good outcome, because it frees resources for work that is worth doing

Choosing the Right Project and the Right Method

A well-defined project that is the wrong project still wastes months. Selection is therefore part of Define. Use criteria that the sponsor has agreed, score candidate projects against them, and keep a record of why the winner was chosen.

CriterionQuestion to askHow to check
Strategic fitDoes it support a named business objective, customer commitment, or regulatory need?Ask the sponsor which objective it advances. If there is no answer, expect weak support later.
Customer impactDoes it affect quality, delivery, cost, safety, or service as customers experience it?Link to complaints, returns, delivery performance, or customer scorecards.
Financial impactIs the cost of poor quality, or the opportunity, large enough to justify the effort?Rough estimate first; finance reviews the method.
FeasibilityCan it be done in about three to six months with the people and data available?Check data availability early. Break large problems into several projects.
Cause not yet knownIs the root cause unclear and does it need analysis?If the cause and fix are obvious, just do it.
Ownership and supportIs there a sponsor and a process owner who want it solved and can release resources?No sponsor, no project.
Quick wins High impact, low effort Major projects High impact, high effort Fill-ins Low impact, low effort Thankless tasks Low impact, high effort Effort to deliver (low to high) Impact (high to low) Line 3 yield Permit backlog Label change New ERP module Break-room redesign
An impact-effort grid helps a team and its sponsor compare candidates quickly. Illustrative candidates are shown; the discussion about where each one belongs is the valuable part. Use the Impact-Effort Matrix Builder or the Project Prioritization Matrix.

Is DMAIC the right method? Match the method to the problem:

If the situation is…ConsiderWhy
A safety or quality emergencyContainment first, then a corrective action report such as 8DProtect people and customers before analysis
The cause and the fix are known and low riskJust do it, with a short planA formal project adds overhead without adding insight
A local problem in one area that a team can solve in daysA kaizen event or A3Faster and lighter than DMAIC
A chronic problem, unknown cause, measurable impactDMAICNeeds data, analysis, and verified improvement
A new product or process to design from scratchDesign for Six Sigma (DMADV)There is no existing process to improve

Writing the Problem Statement and the Goal Statement

The problem statement describes the gap between what is happening and what should be happening, in measurable terms. The goal statement describes the improvement the project will deliver. Both should be clear enough that two people would measure them the same way.

A problem statement answers

  • What is wrong, and in which process or product?
  • Where does it occur?
  • When did it start, and how long has it been going on?
  • How big is it, in a measured unit?
  • What is the impact on the customer or the business?

A problem statement does not contain

  • A cause (“because the operators are careless”)
  • A solution (“we need a new machine”)
  • Blame for people or departments
  • Vague words such as “improve” or “better” with no measure
WeakStrong
Quality on Line 3 needs improvement.From weeks 14 to 21, first-pass yield at the Line 3 final test fell from 96.8% to 91.2%, creating about 8.8% rework and shipment risk and costing an estimated $338,000 a year at the current rate.
Our permit process is too slow.Over the last two quarters, the median time from complete application to permit decision was 41 days against a published standard of 30, and 38% of applications exceeded 45 days.
We have too many medication errors.In the last 12 months, 27 of 1,900 pediatric doses (1.4%) required intervention because of a weight or unit error, against a target of zero reaching the patient.
We need a new scheduling system.On-time shipment fell from 97% to 89% over eight weeks, with 72% of late orders tied to schedule changes made within 48 hours of the ship date.

The examples above are illustrative.

Check for hidden solutions. If your statement contains the words “need,” “install,” “implement,” “train,” or “replace,” it probably describes a solution. Rewrite it as the gap the solution is meant to close.

The goal statement follows the SMART pattern: Specific, Measurable, Achievable, Relevant, and Time-bound. It should use the same metric as the problem statement, state the baseline and target, and give a date.

SMART elementExample for the Line 3 project
SpecificFirst-pass yield at Line 3 final test
MeasurableFrom 91.2% to at least 96.0%, measured weekly with the same definition of first-pass yield
AchievableThe line ran at 96.8% earlier in the year; the target is below that level
RelevantReduces rework cost and shipment risk, supporting the plant's quality and delivery objectives
Time-boundSustained for eight consecutive weeks within five months of project start

Set the target with evidence. A goal taken from a best-case month or a competitor's claim may not be achievable. Use demonstrated past performance, customer requirements, or technical limits, and agree that the target may be revisited at the end of Measure when the baseline is firm. For a tool to test wording, see the Project Charter guide.

Scope: Drawing the Boundary, and the SIPOC

Scope says what the project will and will not address. Most projects that stall are too broad, so draw the boundary deliberately and write it down. A boundary has four parts: where the process starts and stops, which products, sites, or customers are included, which causes or solution areas are in bounds, and what is explicitly out.

Scope elementExample (Line 3 final test project)
Process start and endStarts when assembled units arrive at final test; ends when a unit is recorded as passed or sent to rework
In scopeFinal test station, test fixtures, test limits, and the quality of inputs received from assembly
Out of scopeOther production lines, product design changes, supplier changes (any of which may be recommended for later)
ConstraintsNo capital spending above an agreed limit; no reduction of required customer tests
AssumptionsTest data for the last 12 months can be extracted; production volume stays within 10% of plan

SIPOC stands for Suppliers, Inputs, Process, Outputs, Customers. It is a one-page, high-level picture that shows the boundary and who is connected to it. Keep the process to four to seven steps. Build it with the team and the process owner, and start in the middle: map the process steps first, then outputs, then customers, then inputs and suppliers.

S Suppliers Line 3 assembly Component vendors Test engineering Calibration lab I Inputs Assembled units Test fixtures Test software and limits Calibrated instruments P Process 1. Receive units 2. Load and connect 3. Run test sequence 4. Evaluate results 5. Sort pass and fail 6. Record data O Outputs Passed units Failed units for rework Test records C Customers Packaging and shipping End customer Quality and engineering
A SIPOC for the Line 3 final test. The process is shown in four to seven steps; detailed mapping is left for Measure. Build your own with the SIPOC Diagram Generator.
Use the SIPOC to challenge the scope. If a supplier or input appears that the team cannot influence, note it as a constraint. If a customer appears that the problem statement ignored, reconsider the problem.

Keep an is / is-not list alongside the scope: what the problem is and is not, where and when it occurs and does not occur. See Is / Is-Not Analysis.

From the Voice of the Customer to Critical-to-Quality Requirements

The customer decides whether the output is good enough, so a project that ignores the customer will optimize the wrong thing. Define converts what customers say, in their words, into specific, measurable characteristics the team can control. The customer may be external, or it may be the next department or process step.

MethodBest forCautions
Interviews and site visitsUnderstanding needs, context, and unspoken problemsTime-consuming; use open questions and listen for needs, not solutions
SurveysMeasuring the importance and satisfaction of known needs across many customersPoorly worded questions give poor answers; low response rates
Complaints, returns, and warranty dataFinding what already goes wrongReflects only customers who complain
Customer scorecards and contractual requirementsUnderstanding formal targets and penaltiesMay lag actual experience
Observation and journey mappingSeeing how the output is usedNeeds permission and access
Focus groupsExploring reactions and ideasGroup dynamics can bias results

Organize what you hear. Group comments with an affinity diagram, then rank importance. The Kano model adds a second dimension: which needs are basic expectations (dissatisfy when absent), which are performance needs (more is better), and which are delighters. See the Voice of the Customer and Kano guide and the Kano Model Analyzer.

Translate needs into CTQs. A critical-to-quality characteristic is a measurable requirement that is closely tied to what the customer values. A CTQ tree moves from the broad need, to a driver, to a specific measure with a target.

Reliable products, shipped when promised Customer need Product passes final test first time First-pass yield at final test of at least 96.0% No defects reach the customer Fewer than 500 defective units per million at customer receipt Delivered when promised Shipped within one day of the promised date for at least 98% of orders DriverCritical to quality (measurable)
A CTQ tree for the Line 3 project. Each branch ends in something that can be measured and that has a target. Build one with the CTQ Tree Builder.

Define the primary metric (Y) operationally. Write exactly how it will be counted: what is a unit, what counts as a pass, what is included and excluded, where and when it is measured, and by whom. Ambiguity here will produce different numbers from different people, which becomes an expensive problem in Measure.

The Business Case: Cost of Poor Quality and Expected Benefits

A project competes for time and attention with other work, so the sponsor needs to know what it is worth. The business case estimates the cost of the problem and the benefit of solving it, in terms finance recognizes.

Cost of poor quality categoryWhat it includesExamples
Internal failureDefects found before the customer receives the productScrap, rework, retest, downtime from defects, re-inspection
External failureDefects found after deliveryReturns, warranty, complaints, field repair, lost customers, penalties
AppraisalFinding defectsInspection, testing, audits, calibration
PreventionStopping defects from happeningTraining, process design, preventive maintenance, supplier development

Use the COPQ Estimator and the COPQ overview to structure the estimate. Prevention and appraisal costs are not waste in themselves; the aim is to reduce failures, then reduce the appraisal needed to catch them.

Worked benefit estimate, Line 3 final test. The figures below are illustrative.

ItemCalculationResult
Annual volume12,000 units per week × 50 weeks600,000 units
Rework at current yield600,000 × 8.8% (91.2% first-pass yield)52,800 units
Rework cost at current yield52,800 × $6.40 per unit$337,920
Rework at target yield600,000 × 4.0% (96.0% first-pass yield)24,000 units
Rework cost at target yield24,000 × $6.40$153,600
Estimated annual benefit$337,920 − $153,600$184,320
Be conservative and transparent. This estimate counts rework labor and materials at a standard cost. It does not include shipment risk, capacity gained, or customer goodwill, and it does not subtract the cost of the project. The assumptions, the method, and the validation owner belong in the charter, and finance should review them.

Hard savings

Reductions in cost that appear in the financial statements, such as less scrap, lower overtime, or fewer purchases. Finance confirms them after the project.

Soft or avoided benefits

Cost avoidance, capacity released, risk reduction, and customer satisfaction. Real, but they are reported separately and should not be mixed with hard savings.

For an investment view, use the Project ROI Calculator and the Kaizen Savings ROI Calculator.

Team, Roles, and Stakeholders

Improvement is done by people, and projects fail from lack of access, authority, and support more often than from lack of technique. Define names the people and confirms what each is committing.

RoleResponsibilityTypical commitment
Sponsor or championOwns the business reason for the project, removes barriers, approves the charter and tollgates, and secures resourcesRegular reviews; a few hours a month
Process ownerOwns the process being improved, supports data access and changes, and sustains resultsActive involvement throughout; responsible for control
Project leader (Green or Black Belt)Leads the project, the method, the data, and the communicationPart to most of their time, depending on project size
Team membersBring process knowledge, collect data, test ideas, and implement changesA few hours a week, agreed with their managers
Finance partnerReviews the benefit method and validates savingsReviews at defined points
Coach (Master Black Belt or experienced Black Belt)Advises on method and statistics and reviews tollgatesAs needed

Pick team members for knowledge. Include people who do the work, people who feed it, people who receive its output, and someone who can analyze data. Teams of four to eight work well. Confirm the time with each member's manager, not just the member.

Analyze stakeholders. List the people and groups affected, estimate their influence and their current support, and plan how to engage each. Those with high influence and low support need conversations early. See the Stakeholder Analysis Builder and Stakeholder Analysis.

Clarify who decides what. A RACI matrix shows who is Responsible, Accountable, Consulted, and Informed for key activities.

ActivitySponsorProcess ownerProject leaderTeamFinance
Approve charter and scopeARRCC
Provide process data accessIARRI
Validate benefit estimateICRIA
Approve process changes (later phases)CARRI
Tollgate decisionACRII

Build one with the RACI Matrix Builder. Finally, agree a communication plan:

AudienceMessageChannelFrequencyOwner
Sponsor and process ownerProgress, decisions needed, risksShort review meetingEvery two weeksProject leader
TeamTasks, findings, next stepsTeam meetingWeeklyProject leader
Affected employeesWhy, what is happening, what to expectHuddles and visual boardMonthly or at milestonesProcess owner
FinanceBenefit method and resultsReviewAt tollgatesProject leader

Just Enough Data in Define

Define needs enough data to size the problem and justify the project, and no more. Careful measurement, including checking that the measurement system is reliable, belongs in Measure. If you wait for perfect data before chartering, the project will never start.

Collect in Define

  • The primary metric for a representative period (often 3 to 12 months)
  • A rough view of variation over time (a simple run chart)
  • Obvious stratifications, such as by line, shift, or product
  • Cost data to size the business case
  • Confirmation that the data can be obtained

Leave for Measure

  • Measurement system analysis (MSA and gage R&R)
  • A formal data collection plan
  • Detailed process maps and time studies
  • Process capability studies
  • Statistical tests

A preliminary look at the baseline. A yield of 91.2% corresponds to a sigma level of about 2.85, and a yield of 96.0% to about 3.25, using the conventional 1.5-sigma shift. The sigma level is a convenient way to compare processes. It is not a goal in itself. Use the Sigma Level and DPMO Suite to convert between yield, defect rates, and sigma level.

Warning on early data. Existing data may be defined differently from what the project needs, may have gaps, or may not be trusted by the people who use it. Note these limitations in the charter and plan to check them in Measure, rather than quietly relying on them.

Look for the shape of the problem. A quick Pareto chart of defect types, or a run chart of the metric over time, often shows whether the issue is sudden (a recent change) or long-standing (a chronic cause). That affects scope and the approach in later phases.

Putting It Together: The Project Charter

The charter is the one-page agreement that records everything above. It is signed by the sponsor and referred to throughout the project. The table shows its elements, with the Line 3 example completed.

Charter elementLine 3 final test example
Project titleImprove first-pass yield at Line 3 final test
Problem statementFrom weeks 14 to 21, first-pass yield at final test fell from 96.8% to 91.2%, producing about 8.8% rework and shipment risk.
Goal statementRaise first-pass yield from 91.2% to at least 96.0%, sustained for eight consecutive weeks within five months of project start.
Business caseEstimated rework cost reduction of about $184,000 a year (illustrative), reduced shipment risk, released test capacity. Method and validation owner: finance.
Customer and CTQsPackaging and end customer. First-pass yield ≥ 96.0%; fewer than 500 defective units per million at receipt; on-time shipment ≥ 98%.
ScopeIn: final test station, fixtures, limits, input quality from assembly. Out: other lines, design changes, supplier changes.
Team and rolesSponsor: plant manager. Process owner: Line 3 production manager. Leader: Green Belt (quality engineer). Members: test technician, assembly lead, maintenance, finance partner.
TimelineDefine 3 weeks, Measure 4, Analyze 4, Improve 6, Control 4: about 21 weeks (five months).
Constraints and assumptionsNo capital spending above the approved limit; 12 months of test data can be extracted; volume within 10% of plan.
Risks and mitigationTest equipment downtime (schedule around maintenance); seasonal volume (agree a trading calendar); team member availability (confirmed with managers).
ApprovalsSponsor, process owner, project leader, finance reviewer.

See the Project Charter guide for how to draft, review, and keep a charter alive. The DMAIC project files and the Kaizen Event Charter Template provide formats to start from.

The charter is a living agreement. If data later show that the real problem lies elsewhere, change the scope by agreement with the sponsor, record the change, and keep going. Do not let the project drift, and do not hold to a scope the evidence has overturned.

Tollgate Review: Are We Ready for Measure?

A tollgate is a decision point. The sponsor reviews the work and decides whether the project should proceed. Use the checklist below before the review. Your progress is saved in this browser only, and nothing is sent anywhere.

0 of 15 complete
Problem and value
Scope and process
Customer
People and plan

Questions a sponsor should ask:

  • If this project succeeds, what will be different for customers and for the business, and how will we know?
  • Is this the right size? Could it be finished in about five or six months?
  • How was the benefit estimated, and who will validate it?
  • Who owns the process, and are they committed to sustaining the result?
  • What are the biggest risks to this project, and what is the plan for each?
  • Is there anything in the charter that assumes a cause or a solution?
Tollgate outcomeMeaningTypical next step
GoAll criteria are metStart Measure
Go with conditionsMinor gaps that can be closed quicklyRecord conditions, owners, and dates; review in two weeks
Re-scopeScope, goal, or business case needs reworkRevise the charter; hold another review
Stop or deferNot worth doing now, or no committed sponsorDocument why; free the team for other work

Adapting Define to the Situation

SituationHow Define changes
Green Belt project (smaller, local)A light charter and a short SIPOC; limited VOC; benefit estimate with local finance review; often one to two weeks
Black Belt project (cross-functional, larger value)Fuller VOC, stakeholder analysis, a formal business case, and more tollgate formality; two to four weeks or more
Lean or kaizen eventA one-page charter or A3 background and problem statement, set before the event, with a narrow scope and a fixed date
ManufacturingCTQs often relate to specifications, yield, scrap, downtime, and on-time delivery; data is often in production and quality systems
Service and transactionalCTQs relate to time, accuracy, and experience; the boundary of the process is often less visible; data may need to be created
Healthcare and public serviceSafety, equity, and legal requirements affect scope and stakeholders; involve patients or residents and staff; see the hub guides for healthcare and government

Lean and Six Sigma together. Lean projects often start from a value stream map that points to waiting and flow problems; Six Sigma projects start from variation and defects. In Define, the same questions apply: what is the problem, who is the customer, and what is the boundary? See Lean Six Sigma Integration.

Common Mistakes and Red Flags

MistakeWhat it looks likeHow to correct it
The solution is in the problem statement“We need to install a new fixture.”Rewrite as the gap; move the idea to a parking lot
Scope is too broadA single project covers several lines, products, and causesSplit into projects; narrow with Pareto and stratification
No committed sponsorThe sponsor misses reviews and cannot release resourcesEscalate; pause until commitment is real
Goal with no baseline“Reduce defects by 50%” with no starting figure or definitionEstablish and state the baseline and the metric definition
Benefit never reviewed by financeSavings claimed that no one can later findAgree the method and validation owner at the start
Customer never consultedCTQs invented by the teamInterview or survey customers, or use customer data
Team chosen by availabilityNo one who does the work is on the teamAsk who knows the process; negotiate time with managers
Over-collecting dataWeeks spent on detailed measurement before the charterUse existing data; save detailed work for Measure
Charter as paperworkSigned once and never revisitedReview at each tollgate and record changes
Define never endsEndless refinement of the charterTime-box; sign when the checklist is met

Define Resources on This Site

Guides

Tools

Templates

Body of Knowledge

Define Phase Frequently Asked Questions

How long should the Define phase take?

It depends on project size and how much is already known. Many Green Belt projects complete Define in one to three weeks, and larger Black Belt projects in two to four. If Define is taking much longer, the usual causes are a problem that is too broad, a sponsor who has not committed, or unclear data access. Time spent here is well spent, because errors in Define carry through the whole project.

What is the difference between a project charter and a project plan?

The charter is the agreement: the problem, goal, scope, business case, team, and high-level timeline that the sponsor approves. A project plan is the detailed working schedule of tasks, owners, and dates that the team maintains. The charter changes only by agreement; the plan changes as the team learns.

Should we include solution ideas in Define?

Capture them, but do not commit to them. Ideas that arise in Define are useful to record in a parking lot and test later, but a charter that names the solution turns the project into an implementation and removes the analysis that DMAIC exists to provide. The problem statement and goal should describe the gap, not the fix.

Do internal process projects need a voice of the customer?

Yes, though the customer may be the next process step, the end user, or an internal department. Every process has a customer who receives its output and decides whether it is good enough. Without that view, teams optimize what is easy to measure and miss what matters to the receiver.

Is a SIPOC enough as a process map in Define?

A SIPOC is enough for Define because its job is to set the boundary and show the high-level flow, inputs, outputs, and customers. Detailed process maps, such as swimlane or value stream maps, belong in the Measure phase, once the scope is agreed.

Who signs the charter?

At a minimum, the sponsor (or champion), the process owner, and the project leader. Finance should review the benefit estimate, and the managers who release team members should confirm the time commitment. A charter signed only by the project leader is a proposal, not an agreement.

Sources and Further Reading

  • Thomas Pyzdek and Paul Keller, The Six Sigma Handbook, chapters on the Define phase.
  • Michael L. George, David Rowlands, Mark Price, and John Maxey, The Lean Six Sigma Pocket Toolbook.
  • T. M. Kubiak and Donald W. Benbow, The Certified Six Sigma Black Belt Handbook (ASQ).
  • Donald W. Benbow and T. M. Kubiak, The Certified Six Sigma Green Belt Handbook (ASQ).
  • ISO 13053-1, Quantitative methods in process improvement: Six Sigma, Part 1: DMAIC methodology.
  • ASQ Certified Six Sigma Green Belt and Black Belt Bodies of Knowledge, Define phase.
  • Noriaki Kano and colleagues, “Attractive Quality and Must-Be Quality,” 1984.

This content is educational. Financial figures in the examples are illustrative, and your organization's policies and finance team determine how benefits are calculated and reported.

Phase 2 of 5

Measure

What is happening now, and how reliably can we measure it? Measure establishes the current performance of the process with data the team and the sponsor can trust. It defines what to measure, proves the measurement system is good enough, collects the data, and describes the baseline well enough to guide the analysis.

Key question
What is happening now, and how reliably can we measure it?
Typical duration
About 2 to 6 weeks, depending on data availability and the time needed to collect it
Led by
The project leader, with the process owner, team members who do the work, and a quality or data analyst
Primary outputs
Data collection plan, validated measurement system, detailed process map, baseline performance, stratified findings, refined problem and goal
Tollgate decision
Sponsor confirms the baseline is credible and the problem is focused enough to analyze
Core tools
Operational definitions, data collection plan, process maps, MSA and gage R&R, run and control charts, capability and sigma level, Pareto, stratification

What the Measure Phase Is For

Define says what the problem is. Measure says how big it really is, where and when it occurs, and whether the numbers can be believed. Teams that skip this phase tend to analyze guesses; teams that rush it often discover later that their data cannot support a conclusion.

What Measure must achieve

  • Measures defined so that two people would count the same way (operational definitions)
  • A measurement system proven good enough for the decisions it will support
  • A detailed picture of how the process actually works, with candidate input variables (Xs)
  • A baseline: current performance, its stability, and its capability
  • The problem broken down by where, when, and what type (stratification)
  • A confirmed or refined problem statement, goal, and scope

What Measure must not do

  • Jump to causes or solutions (that is Analyze and Improve)
  • Collect everything that can be measured “just in case”
  • Trust existing data or gauges without checking them
  • Report averages without looking at variation and time order
  • Treat the baseline as fixed when the data reveal a different problem
  • Skip the process walk and rely on a map drawn in a conference room
The test of a finished Measure. A skeptical colleague should be able to ask “how do you know?” about any baseline number, and the team can answer: how it was defined, how it was measured, how reliable the measurement is, and how much data supports it.

The Measure Process, Step by Step

Select measures Y, process, and inputs Map the process As it really runs Plan the data What, how, who, when Validate Measurement system Collect Following the plan Baseline Stability, capability, strata Tollgate Sponsor decision
Validating the measurement system comes before collecting the bulk of the data. Findings in later steps often send the team back to refine definitions or the plan.
  1. Select and define the measures

    Translate the CTQ from Define into the primary measure (Y). Add process measures and key input measures (Xs) that may influence it. Write an operational definition for each.

    • Confirm the primary metric from the charter
    • List candidate process and input measures
    • Write definitions: what is counted, how, where, by whom, and when
    • Decide whether each is continuous or discrete data

    Output: Operational definitions for every measure

    Watch for: Definitions that differ between shifts or departments, and measures chosen only because the data is easy to get

  2. Map the process as it actually runs

    Go to the place where the work is done and trace the real flow, including rework loops, waiting, and workarounds. Mark where measurements are or could be taken, and list candidate inputs.

    • Walk the process from start to end with the people who do it
    • Record steps, times, handoffs, and decision points
    • Identify value-adding and non-value-adding steps
    • List potential Xs with a cause-and-effect matrix or diagram

    Output: A detailed process map and a list of candidate Xs

    Watch for: A map of the procedure as written instead of the process as performed

  3. Build the data collection plan

    Decide exactly what data to collect, from where, how, how much, by whom, and over what period, including the factors needed to stratify the data later.

    • Define the sample size and sampling method
    • Choose stratification factors up front
    • Design forms or extracts, and train data collectors
    • Agree the time period and how special events will be recorded

    Output: A written data collection plan and test of the forms

    Watch for: Collecting data first and thinking about the questions later

  4. Validate the measurement system

    Before relying on the data, test whether the measurement process is accurate and precise enough. Use gage R&R for continuous data and attribute agreement analysis for pass/fail or category data.

    • Choose parts or samples that span the real range
    • Include the operators who normally measure
    • Randomize order and blind the samples
    • Analyze, and improve the system if it is not acceptable

    Output: Documented measurement system results and any fixes

    Watch for: Skipping the study because “the gauge is calibrated”. Calibration is not the same as capability

  5. Collect the data

    Follow the plan, record time and context for each data point, and check the data as it arrives so that problems are caught early.

    • Run a short pilot of the collection
    • Monitor completeness and plausibility daily or weekly
    • Record changes, breakdowns, and other events as notes

    Output: A clean, time-ordered, annotated data set

    Watch for: Changing the process while collecting, which changes the thing you are trying to measure

  6. Baseline and stratify

    Plot the data in time order, check stability, describe the distribution, and calculate performance and capability. Then break the data down by the factors in the plan to find where the problem concentrates.

    • Run or control charts, histograms, and box plots
    • Capability or sigma level, with the method stated
    • Pareto charts and stratified comparisons

    Output: The baseline, with evidence of where and when the problem occurs

    Watch for: Calculating capability for a process that is not stable, and reporting only an average

  7. Refine the charter and hold the tollgate

    Update the problem statement, goal, and scope in light of the data. Present the baseline and findings to the sponsor and process owner and agree the focus for Analyze.

    • Revise the charter where the data call for it
    • Use the tollgate checklist in this tab
    • Record decisions and signatures

    Output: An updated charter and a recorded tollgate decision

    Watch for: Holding on to the original problem statement after the data show something different

Choosing and Defining the Right Measures

A project needs a small, deliberate set of measures. The primary measure (Y) comes from the CTQ in the charter. Supporting measures describe the process and its inputs, and a balancing measure shows whether improving one thing worsens another.

Measure typeWhat it describesExample (Line 3 final test)
Output (Y)The result the customer experiencesFirst-pass yield at final test; defects per million units at customer receipt
Process (y)How well a step performsTest cycle time; retest rate; first-time connector seating
Input (X)A condition or factor that may drive the outputFixture wear; firmware version; material lot; operator; shift
BalancingA watch on side effectsThroughput per hour; rework cost; test coverage
The processY = f(X1, X2, ...) Material lot and supplier Operator and shift Fixture condition and fit Test limits and firmware version Ambient temperature Handling procedure followed Y: first-pass yield at final test Controllable (X) Noise Procedural
Y = f(X). Inputs feed the process and shape the output. Measure builds the list of candidate Xs; Analyze tests which of them matter.

Operational definitions. An operational definition says exactly how a measure is obtained, so that two people get the same number. It states what is counted, the unit, the rules for inclusion and exclusion, the method and instrument, the location and time, and who records it.

ElementExample: first-pass yield (FPY) at final test
DefinitionUnits that pass all final test steps on the first attempt, divided by units that enter final test
Unit of countOne serialized unit; a unit retested after a fixture error is still counted as a first-pass failure
ExclusionsEngineering builds and units pulled for a planned audit
Data sourceTest station log, extracted daily at shift end
Recorded byTest system (automatic); verified weekly by the quality technician
FrequencyDaily, reported weekly

Continuous or discrete data? The data type decides the tools you can use, and continuous data usually carry more information per observation than counts.

Data typeExamplesTypical tools
Continuous (variable)Time, length, weight, temperature, torqueHistograms, I-MR and Xbar-R charts, capability (Cp, Cpk), gage R&R
Discrete: counts of defectsDefects per unit, errors per formc and u charts, DPU, DPMO
Discrete: proportionsPercent defective, pass/fail, yieldp and np charts, binomial capability, attribute agreement analysis
CategoricalDefect type, machine, shiftPareto charts, check sheets, stratification
Prefer continuous data when you can. A measured value (for example, the actual insertion force) shows how close a unit is to the limit. A pass/fail result hides that information and needs far more units to detect the same change.

Mapping the Process and Finding the Candidate Xs

The SIPOC from Define showed the boundary. Measure goes inside it. The aim is to understand how the work really happens, where time and defects arise, and which inputs might matter, without deciding yet which ones do.

Map typeUse it whenStrength
Detailed process map (flowchart)You need the steps, decisions, and rework loopsSimple and widely understood
Swimlane mapHandoffs between people or departments matterShows who does what and where delays occur between roles
Value stream mapFlow, waiting, and inventory are centralShows lead time against value-adding time; see the Value Stream Mapping guide
Spaghetti diagramMovement of people or material is a concernMakes travel and backtracking visible; see the Spaghetti Diagrams guide
Process FMEAThe risk of failure by step is the questionPrioritizes steps by severity, occurrence, and detection
  • Walk it. Observe the process on every shift that matters. Differences between shifts are often findings.
  • Map what happens, not what should happen. Include rework loops, workarounds, and waiting.
  • Add data to the map: cycle times, queue sizes, yields, and defects at each step.
  • Mark value. Classify each step as value-adding, necessary non-value-adding, or waste, from the customer's viewpoint.

From the map to candidate Xs. A cause-and-effect matrix rates each input against the customer requirements and ranks the inputs by their likely influence. A fishbone diagram collects possible causes by category. Both are hypotheses to test in Analyze, not conclusions. See the Cause-and-Effect Matrix and the Fishbone Builder.

The Data Collection Plan

The plan turns questions into data. It should be specific enough that someone else could collect the data and get the same results. Write it before collecting, review it with the process owner, and test the forms in a pilot.

Plan elementQuestion it answersLine 3 example
Measure and definitionWhat exactly are we measuring?First-pass yield at final test, as defined above
Data typeContinuous, discrete, or categorical?Discrete (pass/fail), plus failure type (categorical)
Source and methodWhere does the data come from, and how is it captured?Test station log; manual coding of failure type by the technician
Sample planHow much, how often, and how chosen?All units for 8 weeks; failure types for every failed unit
Stratification factorsWhat do we want to compare?Shift, operator, fixture, product variant, material lot, test station
Who and whenWho collects, and over what period?Test technicians, every shift, weeks 1 to 8
VerificationHow do we check the data?Weekly audit of 20 records against the physical unit and the test log

Sampling. If you cannot measure everything, choose the sample deliberately. Random sampling gives each item an equal chance. Systematic sampling takes every kth item. Stratified sampling samples separately from groups, such as shifts or lots, so that each is represented. Rational subgrouping collects small groups of consecutive units so that variation within a group reflects short-term noise, and variation between groups reveals shifts. See Sampling Methods.

How much data? The answer depends on the question. Two common calculations:

PurposeFormulaExample
Estimate a proportion to within ±En = z² × p(1 − p) / E²p = 0.088, E = 0.01, 95% confidence (z = 1.96): n ≈ 3,083 units
Estimate a mean to within ±En = (z × σ / E)²σ = 0.14, E = 0.05, z = 1.96: n ≈ 31 measurements

These are planning aids that assume random samples. Use the Sample Size and Confidence Calculator and build the plan with the Data Collection Plan Builder. For a control chart baseline, a common guideline is 20 to 25 subgroups, and for a continuous-data capability study many practitioners look for at least 100 observations across normal operating variation.

Collect context along with each value. Record the date, time, shift, operator, machine, lot, and any unusual event. The context is what lets you stratify later and explain odd points.

Validating the Measurement System

Every observed value is the true value plus measurement error. If the measurement error is large compared with the process variation or the tolerance, the data cannot reveal real differences. A measurement system analysis (MSA) quantifies the error.

CharacteristicMeaningQuestion it answers
Accuracy or biasDifference between the average measurement and the reference valueDoes the system read true?
LinearityWhether bias changes across the measurement rangeIs it equally accurate small and large?
StabilityWhether results are consistent over timeDoes it drift?
RepeatabilityVariation when the same person measures the same item repeatedly with the same equipmentHow much noise is in the equipment?
ReproducibilityVariation between operators, methods, or conditionsDo different people get different answers?
ResolutionSmallest increment the system can detectCan it see differences that matter?

Gage R&R for continuous data. A typical crossed study uses about 10 parts that span the process range, 2 or 3 operators, and 2 or 3 trials each, measured in random order and blinded where possible. The analysis separates repeatability, reproducibility, and part-to-part variation.

Repeatability (equipment) 18% Reproducibility (operators) 12% Total gage R&R 21.6% Part-to-part variation 97.6% 10% 30% Percent of total study variation
Example results. Total gage R&R is 21.6% of study variation, between the commonly used 10% and 30% guideline lines, so the system is marginal and the decision should be weighed against cost and risk.
Result (% of study variation)Common guideline
Below 10%Generally acceptable
10% to 30%May be acceptable depending on the application, cost of the gage, and cost of errors; decide with the sponsor
Above 30%Not acceptable; improve the measurement system
Number of distinct categoriesAt least 5 is commonly recommended; this example gives about 6

These guidelines are from widely used automotive industry MSA practice; use the criteria your customers or standards require. Run a study with the Gauge R&R Calculator and record it in the Gage R&R Study Template. See the Measurement System Analysis guide.

Attribute data. For pass/fail or rating data, use an attribute agreement analysis. Several appraisers classify the same set of items, ideally including known reference answers, more than once. The analysis shows agreement within each appraiser, between appraisers, and with the standard, usually summarized with kappa statistics. A commonly used guideline treats kappa above 0.75 as good agreement and below 0.40 as poor. See Attribute Agreement Analysis.

If the system fails. Look for simple causes first: an unclear definition, no fixture, poor lighting, inconsistent technique, worn equipment, or insufficient resolution. Fix, retrain, standardize the method, and repeat the study. Do not proceed to baseline data with a system you do not trust.

Establishing the Baseline: Stability, Capability, and Sigma Level

The baseline describes how the process performs now. Always check stability before calculating capability, because capability indices assume a stable, predictable process.

  1. Plot the data in time order on a run chart or control chart. Look for shifts, trends, cycles, and points outside the limits.
  2. Investigate special causes. Note or correct them; do not blend them silently into the baseline.
  3. Describe the distribution with a histogram and box plot, and check for outliers and multiple peaks.
  4. Calculate performance with the appropriate measures for the data type.
  5. State the method and period so that the baseline can be reproduced.

Choose the control chart by data type with the Control Chart Selector, and calculate limits with the Control Limits Generator. See the SPC Control Charts guide.

Continuous data: capability indices. They compare the spread of the process with the specification.

IndexFormulaMeaning
Cp(USL − LSL) / 6σPotential capability if the process were centered
Cpkmin[(USL − μ) / 3σ, (μ − LSL) / 3σ]Actual capability, accounting for centering
Pp, PpkSame formulas using the overall (long-term) standard deviationPerformance over the whole data period, including shifts
LSL 9.5 USL 10.5 Mean 10.12 Process output (shaded area is outside the specification) About 0.3% beyond the USL
Illustrative example: specification 9.5 to 10.5, mean 10.12, standard deviation 0.14. Cp = 1.19 but Cpk = 0.90 because the process is shifted toward the upper limit, which accounts for about 0.3% of output out of specification (assuming normality).

Check the assumptions. The simple indices assume a stable process and, for the percent-out-of-specification estimate, an approximately normal distribution. For non-normal data, use a transformation or a method designed for the distribution. See Process Capability, Non-Normal Capability, and the Process Capability Helper.

Discrete data: defect rates and sigma level.

MeasureFormulaExample
Yield (first-pass)Good units / units enteringFPY at Line 3 final test = 91.2%
Defects per unit (DPU)Defects / units0.15 defects per unit
DPMODefects / (units × opportunities) × 1,000,000150 defects in 1,000 units with 5 opportunities each: 30,000 DPMO
Rolled throughput yield (RTY)Product of the yields of each stepSteps at 98.5%, 97.5%, and 95.0% give RTY of 91.2%
Sigma levelz-score of the yield, plus 1.5 by convention for the long-term shift91.2% yield is about 2.85 sigma

Convert between yield, DPMO, and sigma level with the Sigma Level and DPMO Suite and the Rolled Throughput Yield Calculator. See the Sigma Level, DPMO, and RTY guide.

Report the baseline with its context. Write the metric, definition, period, sample size, data source, stability finding, and capability or sigma level together. A baseline number without those details cannot be compared with the result at the end of the project.

Stratification: Where Does the Problem Concentrate?

An average hides structure. Stratification splits the data by a factor that might matter, such as shift, machine, operator, product, supplier, lot, location, time of day, or failure type, and compares the groups. It is the main route from a broad baseline to a focused problem.

0% 25% 50% 75% 100% Connector seating 3,380 Firmware load error 2,196 Intermittent sensor 1,352 Solder defect 760 Cosmetic 422 Other 338 66% 82% 91% 96% 100%
Illustrative Pareto chart of 8,448 final-test failures over eight weeks. Connector seating and firmware load errors account for 66% of failures, so the project focuses there. Build one with the Pareto Chart Builder.
Stratified byResult (illustrative)What it suggests
Shift 17.0% failedBaseline performance
Shift 28.4% failedSlightly above average
Shift 311.0% failedNoticeably higher: investigate what differs (staffing, procedure, fixture handling, material)
Average8.8% failedMatches the overall baseline
  • Pareto analysis ranks categories and shows the vital few. See the Pareto Analysis guide.
  • Run charts by group show whether a difference is persistent or a one-off.
  • Multi-vari charts separate variation within a part, between parts, and over time, to show where the variation lives. See Multi-Vari Studies.
  • Box plots and histograms by group compare center and spread for continuous data.
Strata are hypotheses, not causes. Shift 3 failing more often does not show that its operators are the cause. It may use a different fixture, material, or supervision pattern. Analyze tests the explanation. Measure only reveals where to look.

Refine the problem. The stratified view often narrows the project considerably: “first-pass yield at final test is 91.2%” becomes “connector seating and firmware load errors, concentrated on shift 3, account for two thirds of failures.”

Data Quality and Common Data Problems

ProblemSymptomResponse
Inconsistent definitionsDifferent people count differentlyRewrite the operational definition; retrain; re-collect if needed
Missing dataGaps by time, shift, or operatorFind out why; do not fill with averages without a reason; record the limitation
OutliersIsolated extreme valuesCheck the record and context; correct errors; keep genuine events and explain them
Rounding and resolutionData clustered on a few valuesUse a finer scale or a better instrument
Sampling biasData collected only when convenient or when someone is watchingRandomize or stratify the sampling; collect on all shifts
Observer effectPerformance improves only while measuringCollect unobtrusively; extend the period
Transcription errorsTypos and mismatched recordsAudit a sample against source; use automatic capture where possible
Data from a changed processA shift in the data after a known changeUse only data from the current process, or analyze the periods separately

Check the data early and often. Plot it as it arrives. Compare a sample of records with the physical items. Ask the people who collect it what is hard. Problems found in week one cost little; problems found in week eight can cost the schedule.

Revisiting the Charter at the End of Measure

The baseline is the first time the project meets real data. The charter may need to change, and changing it by agreement is a sign of a healthy project.

Charter elementWhat the Measure data may showPossible change
Problem statementThe real size or location differs from the assumptionRestate with the confirmed baseline and focus (for example, the top two failure types)
GoalThe original target was set on estimatesReset to a target justified by the baseline, customer need, or demonstrated best performance
ScopeThe problem concentrates in one area, or the process boundary was wrongNarrow, widen, or split into two projects
Business caseThe size of the cost differs from the estimateUpdate with finance and confirm the project is still worth doing
Team and timelineNew skills or data neededAdd members; adjust the schedule

See the Project Charter guide on keeping a charter alive, and return to the Define tab for the original elements.

Tollgate Review: Are We Ready for Analyze?

The Measure tollgate confirms that the baseline can be trusted and that the problem is focused enough to analyze. Use the checklist before the review. Progress saves in this browser only, and nothing is sent anywhere.

0 of 14 complete
Measures and definitions
Process understanding
Data and measurement system
Baseline and findings
Approval

Questions a sponsor should ask:

  • How do we know the measurements are reliable?
  • Is the process stable, and what does the baseline tell us about the size of the gap to the goal?
  • Where does the problem concentrate, and what do we not yet know?
  • Has anything in the data changed what we should be working on?
  • What data limitations or risks should we keep in mind?
  • Do the goal and the business case still hold?
Tollgate outcomeMeaningTypical next step
GoThe baseline is credible and the focus is clearStart Analyze
Go with conditionsSmall gaps in data or documentationRecord owners and dates; confirm before analysis
Return to MeasureMeasurement system or data cannot support conclusionsFix and re-collect
Re-scope or stopThe data show the problem is different, smaller, or not worth pursuingRevise the charter or end the project

Adapting Measure to the Situation

SituationHow Measure changes
Manufacturing with automated dataData may be abundant; focus on validating it, understanding its context, and sampling meaningfully rather than collecting more
Manufacturing with manual inspectionAttribute agreement analysis and clear operational definitions matter most; train and audit inspectors
Service and transactionalTime stamps, error counts, and category data dominate; many “gages” are people, so test agreement; timing data often must be created
Healthcare and public serviceTake care with privacy and consent; define events carefully; small volumes may suit run charts better than capability indices; see the healthcare and government hubs
Software and ITDefine events precisely (what counts as an incident or defect); use flow measures such as cycle time and WIP; see the Software and IT hub
Lean flow problemsAdd lead time, cycle time, takt, WIP, and value-added ratio; use Little's Law to link them
Small samples or rare eventsPlot every event over time; use measures such as time between events; be open about the uncertainty

When data cannot be collected quickly. A short manual data collection with a well-designed form usually beats waiting for a system change. If a proxy measure must be used, state clearly how it relates to the true measure and its limits, and plan to check it later.

Common Mistakes and Red Flags

MistakeWhat it looks likeHow to correct it
Skipping the measurement system analysis“The gauge is calibrated, so it is fine”Run gage R&R or attribute agreement; calibration shows bias, not precision
Weak operational definitionsDifferent people report different numbers for the same eventWrite the definition; test it with two people; retrain
Capability of an unstable processCpk reported from data with obvious shiftsInvestigate special causes first; report stability alongside capability
Averages without variationOnly the mean is reportedShow the distribution, time order, and spread
Too much data, too little thoughtSpreadsheets of unused measuresCollect what answers the planned questions
Too little dataConclusions from a handful of pointsSize the sample; show confidence limits or ranges
No stratification planContext not recorded, so the data cannot be splitDecide the factors in advance and record them
Changing the process while measuringImprovements made before the baseline is completeHold changes except for containment; record any change
Ignoring what the data sayThe original problem statement is defended against the evidenceUpdate the charter with the sponsor
PerfectionismWeeks spent polishing data that is already adequateTime-box; document limitations; proceed when the checklist is met

Measure Resources on This Site

Guides

Tools

Templates

Body of Knowledge

Measure Phase Frequently Asked Questions

Why check the measurement system before collecting baseline data?

Because every number in the project depends on it. If the measurement system adds a large amount of variation, the baseline will be blurred, real differences will be hidden, and later tests may reach wrong conclusions. A measurement system analysis is far cheaper than discovering after months of data that the gauge or the data entry could not be trusted.

How much data do I need for a baseline?

It depends on the data type and the purpose. For a control chart baseline, a common guideline is 20 to 25 subgroups. For a capability study of continuous data, many practitioners look for at least 100 observations collected across normal variation. For a proportion, the sample size follows from the precision you need: estimating a defect rate near 9% to within one percentage point at 95% confidence takes about 3,000 units. Treat these as starting points and justify the choice in the data collection plan.

What if the process is not stable when I measure it?

Then the baseline is a description of an unstable process, and capability indices are not meaningful. Investigate the special causes first, using the time order of the data, and either fix them or document them. A stable baseline is the foundation for the Analyze and Improve phases. If instability is the project, record the pattern as part of the baseline.

Do I need to calculate sigma level for every project?

No. Sigma level is a convenient way to compare processes, but a yield, defect rate, cycle time, or capability index may describe the problem better. Choose the metric that matches the CTQ in the charter. If you report a sigma level, state how it was calculated, including whether the conventional 1.5-sigma shift was used.

Can I proceed if the gage R&R result is above 30%?

Not without action. A result above 30% of study variation usually means the measurement system is not acceptable for the decision. Improve it (training, fixtures, a clearer procedure, a better gage, or a different method) and repeat the study. In some cases a result between 10% and 30% can be acceptable if the cost and the risk are considered, and that judgment should be recorded with the sponsor's agreement.

What if no data exists for the metric?

Create it. Design a simple data collection form, define the metric operationally, train the people who will record it, and collect for a defined period. If waiting is not possible, use a proxy measure and state its limits. A short, well-designed manual collection often provides better data than a large existing database nobody trusts.

Sources and Further Reading

  • Thomas Pyzdek and Paul Keller, The Six Sigma Handbook, chapters on the Measure phase.
  • Automotive Industry Action Group (AIAG), Measurement Systems Analysis Reference Manual, for gage R&R guidelines and attribute studies.
  • Douglas C. Montgomery, Introduction to Statistical Quality Control, on control charts and process capability.
  • Donald J. Wheeler, Understanding Variation and Making Sense of Data.
  • T. M. Kubiak and Donald W. Benbow, The Certified Six Sigma Black Belt Handbook (ASQ).
  • ISO 13053-1 and ISO 22514 (statistical methods in process management: capability and performance).
  • NIST/SEMATECH e-Handbook of Statistical Methods.

This content is educational. The example data and results are illustrative. Use the acceptance criteria required by your customers, standards, and organization.

Phase 3 of 5

Analyze

Why is the problem happening? Analyze turns the baseline into explanations. It generates and prioritizes hypotheses, tests them with data and with direct observation, and identifies the few verified causes (the vital few Xs) that drive the gap between current and target performance.

Key question
Why is the problem happening?
Typical duration
About 3 to 6 weeks, depending on the number of hypotheses and the time needed to collect data or run trials
Led by
The project leader, with team members who know the process, the process owner, and a quality or data analyst
Primary outputs
Prioritized hypotheses, verified root causes with evidence, quantified contribution of each cause to the gap, and an updated business case
Tollgate decision
Sponsor confirms the causes are credible and sufficient to close the gap, and that the team may move to solutions
Core tools
Fishbone, 5 Whys, cause-and-effect matrix, FMEA, process analysis, hypothesis tests, regression, ANOVA, multi-vari, designed experiments

What the Analyze Phase Is For

Measure shows how large the problem is and where it concentrates. Analyze explains why. The risk in this phase is jumping to a favorite theory, or accepting a plausible story without evidence, and then spending Improve on the wrong fix. The discipline of Analyze is to require evidence for each cause before acting on it.

What Analyze must achieve

  • A broad, structured list of possible causes, built with the people who know the process
  • A prioritized short list of hypotheses to test
  • Evidence for or against each, from data, observation, and small tests
  • A distinction between symptoms and root causes
  • The contribution of each verified cause to the gap to the goal
  • A reviewed business case and a clear brief for the Improve phase

What Analyze must not do

  • Declare a cause because it is plausible, familiar, or favored by a senior person
  • Stop at a symptom or at a person
  • Treat a correlation as proof of cause
  • Run many tests and report only the significant ones
  • Implement a full solution before the sponsor has accepted the causes
  • Continue analyzing past the point of usefulness
The test of a finished Analyze. For each cause in the report, the team can answer: what is the evidence, what is the mechanism, how much of the gap does it explain, and what would we expect to see if we fixed it?

The Analyze Process, Step by Step

Review Baseline and strata Generate Hypotheses Prioritize Likely and influential Test Data and observation Verify Mechanism and trial Quantify Gap explained Tollgate Sponsor decision
Testing and verifying loop: results often generate new hypotheses or eliminate old ones, and the team returns to earlier steps.
  1. Review the baseline and the stratified findings

    Start from what Measure showed: where and when the problem concentrates, which categories dominate, and what the process map revealed. These are the clues the hypotheses should explain.

    • Re-read the Pareto charts, run charts, and stratified comparisons
    • List what is known, what is not, and what changed when the problem began
    • Check the time of first occurrence against known changes

    Output: A short summary of the clues that any explanation must fit

    Watch for: Ignoring evidence that does not fit a favored theory

  2. Generate hypotheses broadly

    Use structured methods and the people who run the process to list every plausible cause, before narrowing. Separate hypotheses about the process (how work flows) from those about data (what the numbers show).

    • Run a cause-and-effect (fishbone) session
    • Drill into the top defect types with 5 Whys
    • Use the process map and FMEA to find where failures can arise
    • Compare good and bad cases with is / is-not analysis

    Output: A full list of hypotheses, grouped by category

    Watch for: Letting one voice dominate and recording only the first idea

  3. Prioritize the hypotheses

    Rank by likelihood and by potential influence on the problem, using the data and process knowledge. Decide what to test first and how.

    • Use a cause-and-effect matrix or voting to rank
    • Link each hypothesis to a test or observation
    • Start with tests that are quick, cheap, and decisive

    Output: A ranked test plan

    Watch for: Testing only the hypotheses that are easy to test

  4. Analyze the process

    Look for causes in how the work flows: waiting, rework loops, handoffs, bottlenecks, variation in method, and non-value-adding steps.

    • Value-added and non-value-added analysis
    • Cycle-time and queue analysis
    • Comparison of how different shifts or people do the task

    Output: Process-based findings, supported by observation and times

    Watch for: Relying on the documented procedure instead of what happens at the workstation

  5. Analyze the data

    Use graphs first, then statistical tests where needed, to compare groups, show relationships, and separate real effects from random variation.

    • Choose the test by the data type of the Y and the X
    • Check assumptions and sample sizes
    • Report effect sizes and intervals, not only p-values

    Output: Evidence for or against each hypothesis

    Watch for: Running every possible test and picking the significant ones

  6. Verify the causes

    Confirm that the apparent cause is real. Observe the mechanism, check that the timing fits, and, where practical, change the factor on a small scale and watch the effect.

    • Go to the process and see the mechanism
    • Run a small confirming trial or designed experiment
    • Check for alternative explanations and confounding

    Output: Verified root causes and the evidence for each

    Watch for: Accepting a statistical association without a mechanism or a test

  7. Quantify the contribution and hold the tollgate

    Estimate how much of the gap each verified cause explains and what the business case looks like when they are fixed. Present the causes and evidence to the sponsor and process owner.

    • Calculate the excess defects or time attributable to each cause
    • Update the business case with finance
    • Use the tollgate checklist in this tab

    Output: A gap-closure estimate and a recorded tollgate decision

    Watch for: A list of causes with no estimate of their size

Generating Hypotheses with Structure

A hypothesis in Analyze is a testable statement of a cause: “Failures at connector seating are higher on fixture 3 because its contact pins are worn.” Good hypotheses are specific, tied to a process step, and testable with data or observation.

TechniqueUse it toNotes
Brainstorming and affinity groupingCollect many ideas quickly and group themInclude operators, maintenance, and suppliers; see Brainstorming Methods
Fishbone (cause-and-effect) diagramOrganize possible causes by category (machine, method, material, measurement, people, environment)Build it for one specific effect, such as one failure type
5 WhysDrill from a symptom toward a cause that can be acted onSupport each answer with evidence; see the 5 Whys guide
Cause-and-effect matrixRank inputs by their influence on customer requirementsCarried forward from Measure and updated with data
FMEAFind ways the process can fail and rank the risksComplements data analysis; see the FMEA guide
Is / is-not analysisCompare where, when, and what the problem is and is notDifferences point to causes; see Is / Is-Not Analysis
Interrelationship diagramShow how causes influence each otherUseful when causes are tangled
Process and value stream analysisFind waiting, rework, handoffs, and bottlenecksStarts from the maps built in Measure
Final-test failures Machine Fixture contact pin wear Test station calibration Connector press force Method Firmware load procedure Handling between steps Retest rules Material Sensor supplier lot Connector variation Solder paste age Measurement Test limits too tight Fixture contact resistance Data entry of fail codes People Training on fixtures Shift handover Staffing on shift 3 Environment Temperature at station Static discharge control Lighting at inspection
A fishbone for Line 3 final-test failures. Each category holds possible causes to test, not conclusions. Build one with the Fishbone Builder.

Symptoms, causes, and root causes. A symptom is what is observed (“connector not seated”). A cause explains it (“contact pins on fixture 3 are worn”). A root cause is the underlying condition that, if corrected, prevents recurrence (“the pin replacement interval was missed because the preventive maintenance task was skipped in a holiday week and has no check”). Keep asking why while the answer is something the team can act on, and verify each link.

Use the people who do the work. The best hypotheses often come from operators, technicians, and maintainers. A hypothesis session with only managers and engineers will miss them.

Prioritizing What to Test

A fishbone can produce fifty causes. Testing all of them is wasteful. Prioritize by how likely each is to be true, how much it could influence the problem, and how easy it is to test.

HypothesisLikelihoodPotential influenceEase of testPriority
Fixture 3 contact pins wornHigh (failures concentrated on one fixture)High (largest failure type)Easy (inspect, compare fixtures)Test first
Firmware v2.3 load timeoutHigh (timing fits week 14)HighEasy (compare stations by version)Test first
Sensor supplier lot BMediumMediumEasy (trace by lot)Test first
Temperature at the test stationLow (no pattern by time of day)LowMediumDefer
Training on fixturesMediumMediumHard to isolateTest after the others

A cause-and-effect matrix (rating each input against the customer requirements and weighting by importance) gives a transparent ranking, and the Impact-Effort Matrix Builder can help balance influence against the effort of testing. An FMEA ranks failure modes by severity, occurrence, and detection; use the FMEA RPN and Action Priority Tool.

Link every hypothesis to a test. For each, write what you will look at, what result would support it, and what result would reject it, before you look. This protects against reading the data to fit the theory.

Process Analysis: Causes in the Flow of Work

Some causes show up in numbers; many are visible only by watching the process. Process analysis looks for causes in how work moves: where it waits, where it loops back, where handoffs lose information, and where methods vary between people.

What to look forHow to analyze itExamples of findings
Non-value-adding stepsClassify each step from the customer's view as value-adding, necessary non-value-adding, or waste; see the 8 Wastes guideDuplicate inspections; searching for tools; waiting for approval
Waiting and queuesMeasure queue time and work in process; apply Little's Law (lead time = WIP / throughput)Work waiting longer than it is processed
BottlenecksCompare step capacities and cycle times; see the Theory of Constraints guideOne station limiting output
Rework loopsCount how often units return and what triggers itUnits failing at test, returned, and retested
Method variationObserve different people and shifts doing the same task and compareDifferent handling or setup between shifts
HandoffsFollow the information and material between rolesMissing information at shift change
MovementDraw travel paths; see the Spaghetti Diagrams guideLong walks for tools or material

Use the Value-Added Flow Analysis tool and the Cycle Time and Takt Gap Analyzer for the numbers. See also the Standard Work guide when method variation is the finding.

Data Analysis: Start with Pictures, Then Test

Statistical tests answer narrow questions. Before running any, plot the data and ask what the picture suggests. Choose the technique by the type of data in the outcome (Y) and the factor (X).

Continuous Y, categorical X t-test (2 groups), ANOVA (3 or more) Box plots, means, spread by group Continuous Y, continuous X Scatter plot, correlation, regression Fitted line, R-squared, residuals Discrete Y, categorical X Chi-square, 2-proportion test Bar charts and rate by group Discrete Y, continuous X Logistic regression Probability of failure by level of X Type of the factor (X): categorical (left) or continuous (right) Type of the outcome (Y): continuous (top) or discrete (bottom)
A simple guide to choosing an analysis. Nonparametric alternatives exist for data that do not meet the assumptions of the tests shown. The Hypothesis Testing Quick Tester covers common cases.
QuestionGraphical tool firstTypical test
Is the mean different between two groups?Box plots or dot plots by groupTwo-sample t-test (or Mann-Whitney if not normal)
Do three or more groups differ?Box plots by groupOne-way ANOVA (or Kruskal-Wallis); see ANOVA
Is the spread different?Box plots; compare standard deviationsTest for equal variances (Levene's test)
Do rates differ between groups?Bar chart of rates with intervalsTwo-proportion test or chi-square
Do two variables move together?Scatter plotCorrelation and regression; see Regression Analysis
Is the process different before and after a change?Run or control chart with the change markedControl chart rules; two-sample test with care for time order
Where does the variation come from?Multi-vari chartVariance components; see Multi-Vari Studies

Before applying a test, check its assumptions: independent observations, adequate sample size, and, for many tests, approximately normal data or similar spread between groups. Use the normality testing and plots for that check.

Hypothesis Testing in Practice

A hypothesis test asks whether the difference in the data is larger than chance would plausibly produce. The null hypothesis states there is no difference or effect. The alternative states that there is. The p-value is the probability of seeing a result at least as extreme as the data if the null hypothesis were true.

ConceptMeaningPractical note
Alpha (significance level)The risk of a false alarm that you will accept, commonly 0.05Set before looking at the data
p-valueProbability of data this extreme if there is no real effectA small p-value is evidence against the null; it does not measure size or importance
PowerThe chance of detecting a real difference of a given sizeUsually planned at 0.80 to 0.90; low power can miss real causes
Confidence intervalA range of plausible values for the effectShows both direction and size; more informative than p alone
Effect sizeHow large the difference is in useful unitsDecide in advance the smallest difference that matters

Statistical versus practical significance. With very large samples, tiny differences become statistically significant. With small samples, large differences can fail to reach significance. Always ask whether the effect is large enough to matter for the CTQ, and whether the sample was large enough to detect an effect of that size. See the Hypothesis Testing guide for a full treatment and the Sample Size and Confidence Calculator for planning.

Do not hunt for significance. If you run 20 independent tests at alpha 0.05, you should expect about one to appear significant by chance. Decide the tests in advance, link them to hypotheses, and report all of them, including those that did not support the theory.

Relationships Between Variables: Correlation, Regression, and Experiments

When the factor and the outcome are both continuous, a scatter plot shows whether they move together. Regression fits a line (or a more complex model) and estimates how much the outcome changes per unit change in the factor.

OutputHow to read itCaution
Correlation coefficient (r)Between −1 and +1; shows the strength and direction of a linear relationshipZero does not mean no relationship; curves and outliers distort it
SlopeChange in Y per unit change in XValid only within the range of the data
R-squaredShare of the variation in Y explained by the modelA high value does not prove cause; a low value can still be useful
Residual plotsShow what the model missesPatterns in residuals mean the model is wrong or incomplete
Multiple regressionEstimates the effect of several Xs at onceBeware of correlated inputs and over-fitting

Correlation is not causation. Two measures can move together because one causes the other, because a third factor drives both, because of reverse causation, or by coincidence. Historical (observational) data can suggest causes but rarely prove them.

Designed experiments establish cause. When you deliberately change factors in a planned pattern and randomize the runs, differences in the outcome can be attributed to the factors and their interactions. Use a designed experiment when historical data cannot separate causes, when factors interact, or when you need to know the effect of a change before making it. See the Design of Experiments guide, the DOE Quick Planner, and Design of Experiments.

Verifying a Cause Before You Act on It

Before a cause goes into the report, put it through several checks. The more of them it passes, the more confidence you can have.

CheckQuestionExample
AssociationIs the cause present when the effect is present, and absent when it is absent?Failures concentrate on fixture 3
TimingDid the cause begin before the effect?Pin replacement interval was passed in week 13; failures rose in week 14
Dose-responseDoes more of the cause give more of the effect?Failure rate rises with pin wear measured on each fixture
MechanismIs there a plausible physical or logical explanation, and can you see it?Worn pins leave connectors partly unseated
AlternativesHave other explanations been ruled out, including confounders?Shift effect disappears once fixture use is accounted for
Reversal or trialDoes changing the cause change the effect?Replacing pins on fixture 3 returns its rate to the others' level
ReplicationDoes it hold in other data, places, or periods?The same pattern appears in earlier weeks

Go and see. Observe the mechanism where it happens. A short visit to the workstation, with a camera or gauge, often turns a statistical result into an understanding. Test small. A confirming trial on one fixture or one station is cheap and decisive, and it is part of verification, not the Improve phase's full implementation.

Worked Example: Line 3 Final-Test Yield

This example continues the project from Define and Measure. The baseline was a first-pass yield of 91.2% against a goal of 96.0%. The Pareto chart showed that connector seating (40% of failures), firmware load errors (26%), and intermittent sensor failures (16%) accounted for 82% of the 8,448 failures in eight weeks. All figures are illustrative.

88% 90% 92% 94% 96% 98% 1 2 4 6 8 10 12 14 16 18 20 21 Week Week 14: new firmware and sensor lot
The run chart shows a step change in week 14, not a gradual decline. That points to something that changed, which is a strong clue for hypotheses.

Hypotheses and tests. The team lists changes made around week 14 and builds a fishbone for each of the three failure types. It prioritizes the hypotheses that fit the timing and tests each in turn.

Hypothesis 1: contact pin wear on one fixture. Plotting connector-seating failures by fixture shows a clear outlier.

Fixture 1 2.2% Fixture 2 2.4% Fixture 3 7.2% Fixture 4 2.3% All fixtures 3.5% Connector-seating failure rate (%)
Fixture 3 fails at 7.2% against about 2.3% for the others (chi-square test p < 0.0001 across 24,000 units per fixture). The pattern is a lead; the mechanism and a trial come next.

Inspection shows the contact pins on fixture 3 are beyond their wear limit, and maintenance records show the pin replacement interval was passed in week 13, when a preventive maintenance task was skipped during a holiday week. A trial replacing the pins on fixture 3 returns its failure rate to 2.4% within three days.

The shift puzzle. In Measure, shift 3 had the highest failure rate (11.0%). The fixture finding explains why: fixture 3 ran 45% of shift 3's volume, against 20% and 10% on shifts 2 and 1. The apparent shift effect was largely a fixture effect. This is a common example of confounding, and it shows why a stratified difference is a hypothesis and not a cause.

Hypothesis 2: the firmware change. Three of the six test stations moved to firmware v2.3 in week 14, and three remained on v2.2 for logistical reasons. Over the same weeks, shifts, and products, firmware load failures were about 4.1% on v2.3 stations and about 0.5% on v2.2 stations. The station logs show load timeouts, and rolling one station back to v2.2 reduces its rate to 0.6%.

Hypothesis 3: sensor lot. Traceability shows that 30,000 of the 96,000 units used sensors from a new supplier lot, introduced in week 14. Intermittent sensor failures were about 3.9% on lot B units and about 0.3% on all others. Incoming tests of lot B parts reproduce the intermittent behavior.

Quantifying How Much Each Cause Explains

Finding causes is not enough; the team must show that they are large enough to reach the goal. For each verified cause, estimate the excess defects it produces, which is the observed defects minus the number expected without it.

Verified causeExcess failures (8 weeks)Points of FPYEvidence
Fixture 3 contact pin wear1,1781.237.2% vs 2.3% on other fixtures; worn pins measured; pin replacement trial returned the rate to 2.4%
Firmware v2.3 load timeouts1,7161.794.1% on v2.3 stations vs 0.5% on v2.2 stations in the same weeks; logs show timeouts; rollback trial returned 0.6%
Sensor lot B1,1081.153.9% on lot B vs 0.3% on other lots; incoming test reproduces the fault
Total explained4,0024.17FPY would be about 95.4% if all three were eliminated
Total gap (91.2% to 96.0%) 4.80 points Fixture 3 pin wear 1.23 Firmware v2.3 load timeouts 1.79 Sensor lot B 1.15 Not yet explained 0.63 Percentage points of first-pass yield
Three verified causes account for 4.17 of the 4.8 points needed. The remaining 0.63 points will need further work on smaller categories such as solder defects.

Update the business case. Each percentage point of yield at 600,000 units a year and $6.40 per reworked unit is worth about $38,400. The three causes are therefore worth about $160,000 a year, close to the $184,000 target in the charter; the remainder depends on the smaller causes. Review the figures with finance and decide with the sponsor whether to pursue the remaining gap in Improve.

If the verified causes explain too little. Either the goal needs to be reconsidered, more causes need to be found, or the project may need to be split. Do not let the team declare victory with causes that account for a small part of the gap.

Analytical Pitfalls to Avoid

PitfallWhat happensProtection
Confirmation biasThe team looks only for evidence that supports a favored causeWrite what would disprove each hypothesis; have someone argue the opposite
ConfoundingA third factor explains an apparent effect (shift vs fixture)Stratify by other factors; use designed experiments
Simpson's paradoxA trend in each group reverses when groups are combinedLook at data both combined and split by key factors
Data dredgingMany tests until something is significantDecide tests in advance; adjust for multiple tests; replicate
Survivorship and selection biasOnly the units that were recorded or kept are analyzedCheck how data were selected and what is missing
Over-fittingA complex model fits noiseKeep models simple; check with new data
ExtrapolationPredictions outside the range of the dataRestrict conclusions to the observed range; test beyond it
AuthorityThe most senior opinion is accepted without evidenceAsk for the evidence from everyone, politely and consistently
Analysis paralysisEndless analysis with no decisionTime-box; stop when the main part of the gap is explained

Tollgate Review: Are We Ready for Improve?

The Analyze tollgate confirms that the causes are credible and large enough to justify solutions. Use the checklist before the review. Progress saves in this browser only, and nothing is sent anywhere.

0 of 12 complete
Hypotheses
Evidence
Verification
Quantification and decision

Questions a sponsor should ask:

  • What is the evidence for each cause, and how was it checked?
  • How much of the gap does each cause explain, and how much remains?
  • Could something else explain the same pattern?
  • Have we seen the mechanism with our own eyes?
  • What did we test that turned out not to be a cause?
  • Do the findings change the goal or the business case?
Tollgate outcomeMeaningTypical next step
GoCauses are verified and explain enough of the gapStart Improve
Go with conditionsA cause needs a small confirming trial or a data checkComplete early in Improve; record owners and dates
Return to AnalyzeCauses explain too little, or evidence is weakMore hypotheses, data, or a designed experiment
Re-scope or stopThe causes are outside the project's reach or not worth fixingRevise the charter or end the project

Adapting Analyze to the Situation

SituationHow Analyze changes
Manufacturing with rich dataStratified comparisons, ANOVA, regression, and multi-vari charts; confirm with trials or designed experiments
Low-volume or high-mix productionPool data carefully; use rate-based measures; rely more on process observation and failure analysis of individual events
Service and transactionalProcess analysis of waiting, handoffs, and rework dominates; use time stamps and error categories; interview staff and customers
Healthcare and public serviceCombine data with observation and staff and patient or resident accounts; take care with privacy and equity; see the healthcare and government hubs
Software and ITUse incident and defect data, timelines, and postmortem methods; see the Software and IT hub
Small samples or rare eventsCase-by-case root cause analysis (5 Whys, fault trees), with timelines; avoid over-interpreting small differences
Lean flow problemsValue stream analysis, bottleneck analysis, and queue theory show where lead time accumulates
Recurring failures in equipmentFailure analysis, MTBF and MTTR, and reliability methods; see the MTBF, MTTR, and Availability guide

Qualitative evidence counts. Interviews, observations, photographs, and failed-part examinations are legitimate evidence. Record them with dates and sources, and use them with the data, not instead of it.

Common Mistakes and Red Flags

MistakeWhat it looks likeHow to correct it
Jumping to a solutionThe team starts designing fixes before causes are verifiedPark the ideas; require evidence for each cause
Stopping at the first “why”The cause is a symptom or a personContinue until the answer is a process or system condition the team can change
One cause fits allA single cause is claimed to explain everythingQuantify what it explains; look for others
Correlation treated as causeA scatter plot is presented as proofLook for the mechanism; test by changing the factor
Significance without sizeA tiny but significant difference is reported as importantReport effect size and compare with the CTQ
No test planAnalyses are run until something appearsPlan tests in advance, linked to hypotheses
Ignoring process observationAll analysis is done from a deskGo to the process and see it
Not quantifying contributionA list of causes with no sizesEstimate excess defects or time for each cause
Analysis never endsMore data and more tests with no decisionTime-box and stop when the main part of the gap is explained
Hiding disconfirming evidenceResults that do not fit the theory are droppedReport all findings, and change the theory

Analyze Resources on This Site

Guides

Tools

Templates

Body of Knowledge

Analyze Phase Frequently Asked Questions

How do I know I have found the root cause and not a symptom?

Ask whether fixing it would stop the problem, and whether you can show it with evidence. A root cause is a condition that, if changed, would prevent the effect, and it is supported by more than one kind of evidence: data showing the association, a plausible mechanism, timing that fits, and ideally a small test in which changing the cause changes the result. Asking “why” until the answer is something the team can act on, and checking each link with evidence, helps you avoid stopping at a symptom.

What if the data show a correlation but I cannot explain why?

Treat it as a lead, not a finding. Correlation can come from a hidden third factor, a coincidence in timing, or reverse causation. Look for a mechanism by observing the process, check whether the relationship holds in other data, and if possible test it by deliberately changing the factor on a small scale or with a designed experiment.

Do I always need statistical tests?

No. Many causes are clear from a good chart, a process observation, or a simple comparison, and a Pareto or run chart may be enough. Use hypothesis tests when the difference is not obvious, when the decision is costly, or when you must show that an apparent difference is unlikely to be chance. Statistical significance alone is never enough; the size of the effect and the mechanism matter as well.

What if there are many possible causes?

Prioritize. Use process knowledge, a cause-and-effect matrix, FMEA, and the stratified data from Measure to rank hypotheses, and test the most likely and most influential first. Most problems are driven by a few causes, so quantify how much of the gap each verified cause explains and stop adding hypotheses when the main part of the gap is accounted for.

When should I use a designed experiment instead of historical data?

When historical data cannot separate causes because factors move together, when you need to know the effect of changing a factor to a level the process has not used, when several factors may interact, or when it is safe and affordable to run trials. Designed experiments are often the strongest way to establish cause, and they are most often used in Analyze to confirm and in Improve to optimize.

Can I start working on solutions during Analyze?

Record ideas as they arise, but do not commit to them until causes are verified. Small trials used to test a cause (for example, replacing a worn part on one fixture to see whether the failure rate drops) are part of the analysis. Full implementation belongs in Improve, after the sponsor has agreed the causes at the tollgate.

Sources and Further Reading

  • Thomas Pyzdek and Paul Keller, The Six Sigma Handbook, chapters on the Analyze phase.
  • Douglas C. Montgomery, Design and Analysis of Experiments, and Introduction to Statistical Quality Control.
  • Douglas C. Montgomery and George C. Runger, Applied Statistics and Probability for Engineers.
  • George E. P. Box, J. Stuart Hunter, and William G. Hunter, Statistics for Experimenters.
  • Donald J. Wheeler, Understanding Variation.
  • T. M. Kubiak and Donald W. Benbow, The Certified Six Sigma Black Belt Handbook (ASQ).
  • Judea Pearl and Dana Mackenzie, The Book of Why, for a modern treatment of causation.
  • NIST/SEMATECH e-Handbook of Statistical Methods.

This content is educational. The example data and results are illustrative. Use qualified statistical and engineering judgment for decisions that affect safety, regulated products, or large investments.

Phase

Improve

What changes will remove the causes and improve performance?

This tab is being built. The Improve toolbox will follow the same layout as the Define tab: an overview, step-by-step guidance, the core tools, a tollgate checklist, common mistakes, and links to the site's calculators and templates. In the meantime, see the DMAIC Roadmap guide and the DMAIC overview.

Phase

Control

How will we hold the gains and prevent backsliding?

This tab is being built. The Control toolbox will follow the same layout as the Define tab: an overview, step-by-step guidance, the core tools, a tollgate checklist, common mistakes, and links to the site's calculators and templates. In the meantime, see the DMAIC Roadmap guide and the DMAIC overview.