A Lean Six Sigma Black Belt turns a costly, stubborn business problem into a verified, sustained result. The role combines statistical analysis, process and flow knowledge, project management, and the leadership needed to move sponsors, teams, and process owners.
This pocket guide is the Black Belt companion to the Green Belt Pocket Guide. It covers the full DMAIC path at Black Belt depth, with decision tables, formulas, a complete worked case study, and a glossary you can print.
Black Belt Roadmap at a Glance
Quick Links
1. What Is a Lean Six Sigma Black Belt?
A Black Belt is a full-time, or nearly full-time, improvement leader who takes on problems that are too large, too technical, or too cross-functional for a local team. Where a Green Belt improves the process next to their own desk, a Black Belt works across departments, sites, suppliers, and sometimes customers, and is expected to deliver a verified financial result.
The role has three parts that have to be balanced. The technical part is statistics, measurement, experimentation, and process analysis. The leadership part is sponsor management, team facilitation, and change leadership. The business part is choosing problems that matter and proving the money. Black Belts who are strong in only one of the three tend to produce clever analyses that never get implemented, or enthusiastic projects that cannot prove their results.
| Responsibility | What It Means | Practical Output |
|---|---|---|
| Lead complex DMAIC projects | Own cross-functional projects, typically 4 to 6 months, with a charter, a team, and a sponsor. | Charter, tollgate reviews, verified results, handoff to the process owner. |
| Apply advanced analytics | Choose the right statistical method for the data type and the question, and defend the conclusion. | Measurement-system studies, hypothesis tests, regression models, designed experiments. |
| Coach Green and Yellow Belts | Review their projects, teach tools in context, and keep their scope realistic. | A pipeline of finished Green Belt projects with sound analysis. |
| Manage stakeholders and change | Build sponsor commitment, surface resistance early, and prepare the process owner to own the result. | Stakeholder plan, communication plan, training and rollout plan. |
| Validate and report benefits | Work with finance to separate hard savings from soft benefits and track them after the project closes. | Signed-off benefit statement and a 12-month tracking record. |
| Support deployment | Help leaders select projects, set standards for the program, and share what worked. | Project portfolio input, lessons learned, reusable templates. |
2. Choosing and Chartering Black Belt Projects
Most failed Black Belt projects fail at selection, not at analysis. A project that is too small wastes a scarce resource. A project that is too large never finishes. A project without a committed sponsor loses its team the first time the quarter gets busy.
| Criterion | Test to Apply | Red Flag |
|---|---|---|
| Business impact | Annualized benefit that finance will accept, usually at least $250K for a Black Belt project. | Benefit cannot be tied to a cost, revenue, or risk line. |
| Measurable Y | One primary metric with an operational definition, a data source, and a baseline or a way to get one. | The goal is a feeling, such as "better communication". |
| Cause unknown | The root cause is not already known. If it is, implement the fix instead of running DMAIC. | A solution is already chosen and the project is there to justify it. |
| Scope fits the timeline | Completable in 4 to 6 months with the team available. | Several unrelated problems share one charter. |
| Sponsor and owner | A named sponsor with budget authority and a process owner who will run the new process. | No one will own the result after closure. |
| Data availability | Data exists, or can be collected in weeks, not quarters. | The first three months would go to building a measurement system. |
Problem statement
State what is wrong, where, how big, and since when, without a cause or a solution. Example: first-pass test failures on Line 2 rose from 2.1% to 4.2% over six months, costing about $38 per failed board.
Goal statement
Make it specific, measurable, and time-bound, and tie it to the baseline: reduce first-pass failures from 4.2% to 1.5% or less within six months.
Scope and boundaries
Name the process start and stop points and list what is explicitly out of scope. Scope creep is the most common reason Black Belt projects slip.
Business case
Show the baseline cost, the target cost, and the benefit calculation, and get finance to review it before the kickoff, not after the closeout.
Team and roles
Name the process owner, the sponsor, a finance partner, and two to six subject-matter experts. Agree on the time commitment in writing.
Milestones and tollgates
Set a date for each phase review. Tollgates are decision points, not status meetings: the sponsor either approves the next phase or redirects the project.
For a full charter walk-through, see the Project Charter guide.
3. DMAIC at Black Belt Depth
Black Belts use the same five phases as Green Belts, but each phase has a higher evidence bar. The tollgate question is not "did we do the tools?" but "is the evidence strong enough to justify the next investment?" The table lists the Black Belt toolset and the exit criteria a sponsor should expect at each tollgate.
| Phase | Black Belt Tools | Exit Criteria |
|---|---|---|
| Define | Charter, VOC, Kano, CTQ tree, SIPOC, COPQ, stakeholder analysis, high-level process map. | Approved charter, CTQs tied to customer requirements, baseline cost estimate, named sponsor and owner. |
| Measure | Data collection plan, operational definitions, Gage R&R, attribute agreement analysis, sampling plan, normality checks, capability and sigma baseline, detailed process map or VSM. | A measurement system shown to be adequate, a statistically valid baseline, and a stratified view of where the defects occur. |
| Analyze | Multi-vari studies, stratification, hypothesis tests, ANOVA, nonparametric tests, correlation, simple and multiple regression, FMEA for cause screening. | Verified root causes with statistical evidence and an estimate of how much of the gap each cause explains. |
| Improve | Designed experiments, solution selection matrix, FMEA on the new process, simulation or pilot, cost-benefit, implementation plan. | A piloted solution with confirmed results, a risk assessment, and approval to roll out. |
| Control | SPC charts, process capability on the new process, control plan, reaction plan, standard work, training, audits, benefit tracking. | Process owner accepts the process, the control plan is live, and benefits are signed off by finance. |
The Black Belt is also responsible for the thing a tollgate cannot show: whether the people who run the process believe in the result. Evidence changes minds slowly, so involve operators in data collection, show them the charts, and let them test the proposed changes. For the full phase-by-phase reference, open the DMAIC Toolbox.
4. Define: Customers, CTQs, and the Cost of Poor Quality
Define turns a complaint into a measurable project. The Black Belt's job in this phase is to make sure the team is solving a problem the customer and the business actually care about, and to put a defensible number on it.
Voice of the customer
Collect requirements from interviews, complaints, returns, surveys, and service data. Separate what customers say from what they need. See the VOC and Kano guide.
Kano classification
Must-be requirements cause dissatisfaction when missing and no delight when present. Performance requirements scale with satisfaction. Delighters are unexpected. Priorities differ for each.
CTQ tree
Break a broad need into a measurable requirement: need (fast delivery), driver (order-to-ship time), CTQ (ships within 24 hours of order, 98% of the time).
Cost of poor quality
Add up internal failure (scrap, rework), external failure (returns, warranty, penalties), appraisal (inspection, testing), and prevention costs. The visible part is usually the smaller part.
Stakeholder analysis
Rate each stakeholder on influence and support. Plan different actions for champions, neutral parties with high influence, and likely opponents.
SIPOC and high-level map
Agree on the process boundary and the handoffs before collecting data, so the data covers the process the project will actually change.
Use the SIPOC Diagram Generator for the boundary map and the COPQ Estimator to size the cost of poor quality.
5. Measure: Trustworthy Data and an Honest Baseline
Measure answers two questions in order: can we trust the data, and how is the process really performing? Skipping the first question is the most expensive shortcut in Six Sigma, because every later conclusion rests on the numbers.
| Data Type | Examples | Typical Summary | Typical Chart |
|---|---|---|---|
| Continuous | Time, length, temperature, weight, cost. | Mean, median, standard deviation. | Histogram, box plot, I-MR, X-bar R. |
| Discrete count | Defects per unit, calls per hour. | Rate per unit, DPU. | c chart, u chart, Pareto. |
| Binary (pass/fail) | Defective or not, on time or late. | Proportion, DPMO. | p chart, np chart, Pareto. |
| Ordinal or categorical | Severity 1 to 5, defect type, shift. | Counts, percentages, mode. | Bar chart, Pareto, stacked bar. |
Measurement System Analysis
Every measured value is the true value plus measurement error. A measurement system analysis estimates how much of the variation you see comes from the gauge and the people using it. A crossed Gage R&R study for continuous data uses 10 parts that span the real process range, 3 operators, and 2 to 3 repeats, measured in random order, and separates repeatability (same person, same gauge) from reproducibility (different people).
| Result | Acceptable | Marginal | Unacceptable |
|---|---|---|---|
| % Study Variation (or % of tolerance) | Under 10% | 10% to 30% | Over 30% |
| Number of distinct categories (ndc) | 10 or more is ideal | 5 to 9 is adequate | Under 5 cannot distinguish parts |
| What to do | Proceed | Decide by risk, cost, and the use of the data | Improve the gauge or method before collecting data |
For pass/fail or rating data, run an attribute agreement analysis: several appraisers classify the same set of parts, ideally twice, against a known standard. Agreement is summarized with Cohen's or Fleiss' kappa. By common practice, kappa of 0.9 or more is excellent, 0.7 to 0.9 is acceptable, and under 0.7 means the classification rules need work. Visual inspection is the usual culprit.
Sampling
Sample enough to see the effect you care about. For a mean, the sample size per group is n = 2(Zα/2 + Zβ)²σ² / δ² for a two-sample comparison, where δ is the smallest shift worth detecting. With α = 0.05, power 0.80, and δ = σ, that is about 16 per group. Use the Sample Size and Confidence Calculator for other cases. Sample rationally: stratify by shift, machine, or lot so the sample can show the differences the project needs to explain.
Baseline Capability
Capability compares what the process does with what the customer needs. Cp = (USL − LSL) / 6σ shows potential if the process were centered. Cpk = min[(USL − μ) / 3σ, (μ − LSL) / 3σ] shows actual capability including centering. Use short-term σ for Cp and Cpk and overall σ for Pp and Ppk. Check normality first (probability plot or Anderson-Darling test). A non-normal distribution calls for a transformation, a non-normal capability method, or a defect-based metric instead.
| Cpk | Approx. Sigma (Z + 1.5 shift) | Defects per Million (approx.) | Reading |
|---|---|---|---|
| 0.67 | 3.5 | 22,750 (one side) | Not capable. Expect frequent defects. |
| 1.00 | 4.5 | 1,350 (one side) | Marginal. Little room for drift. |
| 1.33 | 5.5 | 32 (one side) | Common minimum for stable, controlled processes. |
| 1.67 | 6.5 | 0.3 (one side) | Strong capability, with room for drift. |
Run a quick check with the Process Capability Helper.
6. Analyze: Verifying Root Causes with Statistics
Analyze converts a list of suspected causes into a short list of verified ones. The Black Belt's discipline is to match the test to the data and the question, to decide the acceptable risk before looking at results, and to separate statistical significance from practical importance.
How to read a p-value. The p-value is the probability of seeing a result at least this extreme if the null hypothesis were true. It is not the probability that the null hypothesis is true, and it says nothing about the size of the effect. Pair every p-value with an effect size and a confidence interval, and ask whether the effect is large enough to matter to the business.
| Question | Data | Parametric Test | Nonparametric or Alternative |
|---|---|---|---|
| Is the mean different from a target? | Continuous, one sample | 1-sample t | 1-sample Wilcoxon signed rank |
| Do two groups have different means? | Continuous, two independent groups | 2-sample t (Welch) | Mann-Whitney |
| Did the same units change? | Continuous, paired | Paired t | Wilcoxon signed rank |
| Do three or more groups differ? | Continuous, one factor | One-way ANOVA | Kruskal-Wallis, Mood's median |
| Do groups have different spread? | Continuous, two or more groups | F test (2 groups, normal), Bartlett | Levene's test (robust) |
| Is a proportion different? | Binary, one or two samples | 1- or 2-proportion test | Fisher's exact (small counts) |
| Are two categories related? | Counts in a table | Chi-square test of association | Fisher's exact |
| Does X predict a continuous Y? | Continuous X and Y | Simple or multiple regression | Transform, or use a non-linear model |
| Does X predict a pass/fail Y? | Binary Y | Binary logistic regression | Chi-square for a single category X |
Check assumptions first
t-tests and ANOVA assume independent observations, roughly normal residuals, and (for ANOVA) similar variances. Check residual plots, not just the raw data.
ANOVA
Compares group means by splitting variation into between-group and within-group parts. A significant F says that at least one mean differs; follow with Tukey or Fisher comparisons to say which.
Regression
Estimates how the average Y changes with X. Read R-squared and adjusted R-squared, then check residuals for patterns, the p-values of each term, and VIF for multicollinearity (above 5 to 10 is a warning).
Correlation is not cause
A strong correlation can come from a lurking variable or from the way data was collected. Confirm important causes by changing the factor on purpose, which is what a designed experiment does.
Multi-vari and stratification
Plot the data by time, machine, operator, and lot before testing. Patterns that jump out of a multi-vari chart tell you which hypotheses are worth testing.
Practical significance
With enough data, a trivial difference becomes statistically significant. State the smallest difference that would change a decision, and design the sample to detect it.
The Hypothesis Testing Quick Tester and the Hypothesis Testing guide cover the mechanics, and the One-Way ANOVA page in Stat Dojo covers the model.
7. Improve: Designed Experiments and Solution Selection
Analyze tells you which factors matter. Improve finds the settings or design that deliver the result and proves it before rollout. The strongest Improve tool a Black Belt has is the designed experiment, because it changes several factors at once, on purpose, and estimates both their effects and their interactions.
Why Not Change One Factor at a Time?
One-factor-at-a-time testing needs many more runs for the same precision and cannot see interactions, where the effect of one factor depends on the level of another. A full factorial design with k factors at two levels needs 2k runs and uses every run to estimate every effect.
| Term | Meaning | Why It Matters |
|---|---|---|
| Factor and level | An input you set (temperature) and the values you test (235 and 245 °C). | Choose levels wide enough to see an effect but still safe to run. |
| Response | The output you measure (first-pass failure rate). | Must be measurable with an adequate measurement system. |
| Main effect | Average change in the response when a factor moves from low to high. | Effect = mean of runs at high minus mean of runs at low. |
| Interaction | The effect of one factor depends on the level of another. | A significant interaction means you cannot set the factors independently. |
| Randomization | Running the experiment in random order. | Protects against drift, warm-up, and shift effects being mistaken for factor effects. |
| Replication | Repeating runs to estimate pure error. | Needed to test significance and detect smaller effects. |
| Blocking | Grouping runs by a known nuisance source such as shift or material lot. | Removes that source of variation from the comparison. |
| Center points | Runs at the midpoint of all numeric factors. | Test for curvature that a two-level design cannot model. |
| Resolution | How strongly the effects are confounded in a fractional design. | Resolution III confounds main effects with two-factor interactions; IV keeps main effects clear; V keeps main effects and two-factor interactions clear. |
When there are five or more factors, a fractional factorial design (2k−p) screens them in a fraction of the runs, at the cost of confounding some effects. Use screening to find the vital few, then run a full factorial or a response surface design on those. The DOE Quick Planner and the Design of Experiments guide cover design choices, and the case study below walks through a full 23 example with the effects calculated.
Selecting and Proving the Solution
Criteria-based selection
Score candidate solutions against weighted criteria such as impact on the Y, cost, time to implement, risk, and fit with the culture. A Pugh matrix compares options against a baseline design.
FMEA on the new process
Score severity, occurrence, and detection for each way the new process can fail, and act on the highest risks before rollout. See the FMEA tool.
Mistake-proofing
Design the error out where you can. Poka-yoke beats inspection and training. See Mistake-Proofing.
Pilot before rollout
Run the change in one area, with enough volume to confirm the result against the baseline with a hypothesis test, and write down what you learned.
Confirmation run
After an experiment predicts the best settings, run those settings and check that the observed result falls within the prediction interval.
Implementation plan
Assign owners, dates, training, and resources. A solution that is not implemented has a benefit of zero.
8. Control: Holding the Gain
Most improvements decay. Control is how the Black Belt makes sure the new process stays the new process after the team disbands. It has four parts: monitoring with the right chart, a documented standard, a reaction plan, and an owner who looks at the data on a schedule.
| Data | Subgroup | Chart | Use When |
|---|---|---|---|
| Continuous | Individual values (n = 1) | I-MR (individuals and moving range) | Slow processes, batch results, one reading per period. |
| Continuous | 2 to 9 per subgroup | X-bar and R | Frequent subgroups from a stable process. |
| Continuous | 10 or more per subgroup | X-bar and S | Large subgroups, where S is more efficient than R. |
| Continuous, small shifts | Any | EWMA or CUSUM | You need to detect a shift of 0.5 to 1.5 standard deviations quickly. |
| Defectives (pass/fail) | Varying sample size | p chart | Proportion of units that fail. |
| Defectives (pass/fail) | Constant sample size | np chart | Number of units that fail. |
| Defects (counts) | Constant opportunity | c chart | Number of defects per unit or area. |
| Defects (counts) | Varying opportunity | u chart | Defects per unit when the unit size changes. |
| Subgroup n | A2 | D3 | D4 | d2 |
|---|---|---|---|---|
| 2 | 1.880 | 0 | 3.267 | 1.128 |
| 3 | 1.023 | 0 | 2.574 | 1.693 |
| 4 | 0.729 | 0 | 2.282 | 2.059 |
| 5 | 0.577 | 0 | 2.114 | 2.326 |
For an I-MR chart, the limits are X-bar ± 2.66 × MR-bar for individuals and 3.267 × MR-bar for the upper limit on the moving range. For a p chart, the limits are p-bar ± 3√[p-bar(1 − p-bar)/n].
Rules for special causes
A point beyond 3 sigma; nine in a row on one side of the center line; six in a row rising or falling; 14 alternating up and down; two of three beyond 2 sigma on the same side; four of five beyond 1 sigma. Use the rules that match the risk, since each added rule raises the false-alarm rate.
Control versus specification
Control limits describe what the process does. Specification limits describe what the customer needs. Never draw spec limits on a chart of averages, and never use control limits to judge conformance.
Control plan
For each critical characteristic: the measurement method, sample size and frequency, the owner, the control method, the specification, and the reaction plan. Keep it to one living document.
Reaction plan
Say exactly what the operator does when a point goes out of control, who is called, and what gets contained. A chart without a reaction plan is decoration.
Standard work and training
Update the SOP, the work instruction, and the training record. If the new method lives only in the project team's heads, it will not survive the next staffing change. See Standard Work.
Audits and benefit tracking
Schedule process audits for the first 6 to 12 months and track the financial benefit with finance until it appears in the budget.
Choose a chart with the Control Chart Selector and read the SPC Control Charts guide for interpretation.
9. Lean Flow Tools Every Black Belt Uses
Lean Six Sigma means Black Belts work on flow as well as variation. A process can be statistically capable and still slow, because work spends most of its life waiting. Flow tools find that waiting and the policies that cause it.
| Metric | Formula | Worked Example |
|---|---|---|
| Process cycle efficiency | Value-added time / total lead time | 40 min of value-added work in a 10-hour lead time (600 min) is 6.7%. Most of the time is queue time. |
| Takt time | Available time per shift / customer demand per shift | 450 min available and 90 units demanded gives a takt of 5.0 min per unit. |
| Little's Law | WIP = throughput × lead time | 120 orders in process and a throughput of 40 per day gives an average lead time of 3 days. |
| Rolled throughput yield | Product of the yields at each step | 98% × 95% × 97% × 99% = 89.4%, although each step looks good. |
| OEE | Availability × performance × quality | 90% × 95% × 98% = 83.8%. |
| Kanban cards | N = D × L × (1 + S) / C | Demand 100 per day, lead time 0.5 day, 10% safety factor, container size 10: N = 5.5, rounded up to 6. |
Value stream map
Map the current state with real data from the floor: process time, queue time, changeover, defects, inventory. Then design a future state with fewer handoffs and smaller batches. See Value Stream Mapping.
Pull and kanban
Replace push schedules with signals from the next process. Pull caps WIP, which shortens lead time through Little's Law. See Kanban Pull Systems.
Theory of Constraints
Find the bottleneck, exploit it, subordinate everything else to it, then elevate it. An hour lost at the constraint is an hour lost for the whole system. See Theory of Constraints.
Quick changeover
Cut setup time to allow smaller batches. Convert internal setup steps to external ones first. See SMED.
Waste and takt
Use the eight wastes to organize observations and takt time to set the pace the process must meet. See 8 Wastes and Takt, Cycle, and Lead Time.
Standard work and 5S
Lean gains hold when the work is standardized and the workplace is organized. See 5S.
10. Leading Teams, Sponsors, and Change
Black Belts are chosen for analytical strength, and they succeed or fail on leadership. Technically correct recommendations are routinely rejected by the people who have to live with them, so treat the human side as a workstream with a plan, not a soft skill.
| Stakeholder | What They Need | How to Engage |
|---|---|---|
| Executive sponsor | A clear business result and no surprises. | Short tollgate reviews with a decision requested, and early warning when a milestone is at risk. |
| Process owner | To keep running the process after the project, and to trust the new method. | Include in every phase, give them the data, and have them present the controls. |
| Operators and front-line staff | To know why the change is happening and that their knowledge is valued. | Collect data with them, test ideas with them, and credit them in the results. |
| Finance partner | Benefit calculations they can defend. | Agree on baseline, method, and timing at kickoff; review at each tollgate. |
| Skeptics and opponents | A hearing, evidence, and a face-saving way to change position. | Listen first, test their concerns with data, and give them a role in the pilot. |
Run the team as a team
Teams move from forming to storming to norming to performing. Expect friction in the first weeks, set ground rules and decision methods early, and keep meetings short and visual. See Conflict Resolution in a Team.
Plan for resistance
People resist when they do not understand why, fear loss, or do not believe the change works. ADKAR (awareness, desire, knowledge, ability, reinforcement) helps locate which step is missing. See Change Management for Improvement.
Coach Green Belts well
Review their data before their conclusions, ask what would change their mind, and let them present to the sponsor. Teach the tool when the project needs it, not before.
Communicate in the audience's terms
Executives want cost, risk, and timing. Operators want to know what changes in their day. Show each group the chart that answers their question.
Hold tollgates honestly
Present the evidence, the open risks, and a clear request. Hiding problems until the next tollgate removes the sponsor's chance to help.
Close the loop
Share results, thank the team, and record lessons learned so the next project starts from what this one learned. See Leadership Principles.
11. Financial Validation
A Black Belt project is a business investment, and the benefit has to survive a finance review. Agree on how benefits are counted before the project starts, and report them in the same form finance uses.
| Type | Examples | Treatment |
|---|---|---|
| Hard savings | Scrap and rework reductions, lower material cost, avoided penalties, reduced overtime, lower warranty cost. | Appears in the budget or the P&L. Count in the project result after finance validates it. |
| Cost avoidance | Preventing a cost that was projected, such as avoiding new equipment by raising capacity. | Count only when the avoided spend was in an approved plan, and label it. |
| Revenue growth | Capacity released and sold, better on-time delivery winning orders, lower churn. | Count at contribution margin, not revenue, and only with sales confirmation. |
| Soft benefits | Faster response, better morale, reduced risk, less stress, improved safety. | Report separately and do not add to the dollar total. |
| Category | Examples |
|---|---|
| Prevention | Training, process design, planning, mistake-proofing, preventive maintenance. |
| Appraisal | Inspection, testing, audits, calibration. |
| Internal failure | Scrap, rework, retest, downtime, re-inspection. |
| External failure | Returns, warranty, complaints, penalties, recalls, lost customers. |
Simple payback and return. Payback in months = investment / annual benefit × 12. First-year return = (annual benefit − investment) / investment. For multi-year decisions, discount the cash flows to net present value, with finance supplying the discount rate. Use the Project ROI Calculator and the Kaizen Savings ROI Calculator for quick checks.
12. Case Study: Cutting Solder Failures on a Circuit Board Line
This worked example is an illustrative composite, not a report from a single company. It shows how the phases connect, with calculations you can reproduce. A mid-size electronics manufacturer builds about 240,000 boards a year on one surface-mount line. First-pass failures at in-circuit test, most of them from solder defects, had risen to 4.2%, and each failed board costs about $38 in diagnosis, rework, and retest.
| Element | Detail |
|---|---|
| Problem | First-pass in-circuit test failures from solder defects were 4.2% (252 of 6,000 boards sampled over six weeks), up from 2.1% a year earlier. Each failure costs about $38. |
| Goal | Reduce first-pass failures to 1.5% or less within six months and hold it for three months. |
| Scope | Reflow soldering and automated optical inspection (AOI) on Line 2. Board design and paste supplier changes were out of scope. |
| Baseline cost | 4.2% × 240,000 boards × $38 = about $383,000 per year. |
| Sponsor / owner | Operations manager (sponsor); Line 2 production supervisor (process owner). |
Measure
The team first ran an attribute agreement analysis on the AOI classification. Four inspectors classified the same 60 boards twice against a master set: kappa was 0.62, well under the 0.7 threshold, because inspectors disagreed on when a thin fillet counts as insufficient solder. After the team wrote photo-based acceptance criteria and retrained, kappa rose to 0.88. Only then did they collect baseline data.
The baseline was 42,000 defective boards per million, a sigma level of about 3.2 using the 1.5 shift. Time above liquidus, a key reflow characteristic with a specification of 45 to 90 seconds, had a mean of 66 s and a standard deviation of 7 s, so Cpk was (66 − 45) / (3 × 7) = 1.00. The process was centered but too variable.
Analyze
Stratifying the failures by oven, shift, and paste lot showed that the failure rate differed between the three ovens, and a one-way ANOVA on peak temperature confirmed that oven 2 ran about 6 °C cooler at the board than the others. Paste lot and shift showed no meaningful difference. Regression of failure rate on peak temperature and conveyor speed was suggestive but could not separate the two factors, because in daily production they were changed together. A designed experiment was needed.
Improve: a 23 Factorial Experiment
The team ran eight randomized runs of 1,500 boards each: peak temperature (A: 235 or 245 °C), conveyor speed (B: 70 or 85 cm/min), and nitrogen atmosphere (C: off or on). Measurement used the retrained AOI criteria.
| Run | A: Peak temp | B: Speed | C: Nitrogen | First-pass failure (%) |
|---|---|---|---|---|
| 1 | − | − | − | 4.4 |
| 2 | + | − | − | 3.6 |
| 3 | − | + | − | 4.0 |
| 4 | + | + | − | 1.5 |
| 5 | − | − | + | 4.2 |
| 6 | + | − | + | 3.5 |
| 7 | − | + | + | 3.8 |
| 8 | + | + | + | 1.3 |
Each effect is the average response at the high level minus the average at the low level. For peak temperature, the high-level runs (2, 4, 6, 8) average 2.475% and the low-level runs (1, 3, 5, 7) average 4.100%, so the effect of A is −1.63 points.
| Term | Effect | Reading |
|---|---|---|
| A: Peak temperature | −1.63 | Large. Higher peak temperature lowers failures. |
| B: Conveyor speed | −1.28 | Large. Higher speed lowers failures. |
| AB interaction | −0.88 | Large. The benefit of speed depends on temperature. |
| C: Nitrogen | −0.18 | Small. Not worth the nitrogen cost. |
| AC, BC, ABC | 0.03, −0.03, −0.03 | Negligible. |
The recommended setting was 245 °C and 85 cm/min with nitrogen off, predicted at about 1.5%. A confirmation pilot of 6,000 boards at those settings gave 78 failures, or 1.3%. A two-proportion test of 252 of 6,000 against 78 of 6,000 gave z of about 9.7 and a p-value far below 0.001, and the time-above-liquidus standard deviation fell to 4.8 s, so Cpk rose to (67 − 45) / (3 × 4.8) = 1.53.
Control
The team wrote the new oven profile into the process standard, added the oven-to-oven temperature check to the daily start-up, and put a p chart on weekly first-pass failures with limits of 1.80% and 0.80% around a center line of 1.30% (about 4,600 boards per week). The reaction plan sends a signal to the shift supervisor, who checks oven profiles, then paste and stencil, before restarting the line.
| Metric | Baseline | After rollout | Change |
|---|---|---|---|
| First-pass failure rate | 4.2% | 1.3% | 69% reduction |
| Defective boards per million | 42,000 | 13,000 | Sigma level from 3.2 to 3.7 |
| Time above liquidus Cpk | 1.00 | 1.53 | Capable with margin |
| Annual failure cost | $383,040 | $118,560 | $264,480 saved per year |
| Project investment | — | $61,000 | Profiling equipment, training, team time |
| Payback and first-year return | — | About 2.8 months; 334% | First-year net benefit $203,480 |
13. Mistakes to Avoid
| Mistake | Why It Hurts | Risk |
|---|---|---|
| Choosing a solution before analysis | The project becomes a justification exercise, and the real cause stays in place. | High |
| Skipping the measurement system study | Measurement error can hide a real effect or create a false one. | High |
| Treating p-values as proof of importance | A tiny effect can be significant with a large sample. A real effect can be missed with a small one. | High |
| Using the wrong test | Assumptions fail, and the conclusion is confident and wrong. | Medium |
| One-factor-at-a-time experiments | Interactions are missed, and the optimum is not found. | Medium |
| Over-analysis | Months of modeling after the evidence is already strong enough to act. | Medium |
| Ignoring the human side | Resistance kills good solutions at rollout. | High |
| Counting soft savings as hard | Finance rejects the numbers, and the program loses credibility. | High |
| Weak sponsor engagement | The team loses priority and resources when business gets busy. | High |
| No control plan or owner | Gains erode within months. The process drifts back. | Critical |
| Letting scope grow | The project never finishes, and the benefit never arrives. | High |
14. Black Belt Compared with Other Belts, and Your First 90 Days
| Attribute | Green Belt | Black Belt | Master Black Belt |
|---|---|---|---|
| Focus | Departmental projects | Cross-functional programs | Enterprise strategy and coaching |
| Time on improvement | 25% to 50% | 75% to 100% | 100% |
| Typical project length | 3 to 4 months | 4 to 6 months | Portfolio and mentoring |
| Statistical depth | Hypothesis tests, regression, SPC | DOE, advanced modeling, multivariate and nonparametric methods | Methodology design and innovation |
| Typical benefit per project | $50K to $250K | $250K to $1M or more | Portfolio impact |
| Coaches | Yellow Belts and team members | Green Belts | Black Belts |
Certification routes include the ASQ Certified Six Sigma Black Belt and programs run by training providers and employers. Requirements differ and change, usually including project experience as well as an exam, so check the current requirements with the certifying body. The Black Belt Body of Knowledge entry lists the knowledge areas, and the Master Black Belt Competencies entry describes the next level.
A Practical First 90 Days
- Days 1 to 30. Meet the sponsor and the finance partner. Agree on how benefits will be counted. Confirm the charter, scope, team, and tollgate dates. Walk the process and talk to the people who run it.
- Days 31 to 60. Finish the measurement-system study, collect baseline data, and present a baseline that the process owner accepts as true. Hold the Define and Measure tollgates.
- Days 61 to 90. Begin Analyze with stratified data and a short list of hypotheses. Agree on the experiments or tests needed to verify them and book the line time now, not when the analysis is finished.
15. Formula Sheet
| Topic | Formula | Notes |
|---|---|---|
| Sample mean and standard deviation | x-bar = Σx / n; s = √[Σ(x − x-bar)² / (n − 1)] | Use n − 1 for samples. |
| Z score | Z = (x − μ) / σ | Converts a value to standard deviations from the mean. |
| Cp | (USL − LSL) / 6σ | Potential capability; ignores centering. |
| Cpk | min[(USL − μ) / 3σ, (μ − LSL) / 3σ] | Actual capability; use short-term σ. Pp and Ppk use overall σ. |
| DPMO | defects / (units × opportunities per unit) × 1,000,000 | Sigma level = NORMSINV(1 − DPMO/106) + 1.5 under the usual shift convention. |
| Rolled throughput yield | RTY = Y1 × Y2 × ... × Yn | Probability that a unit passes every step first time. |
| Confidence interval for a mean | x-bar ± tα/2, n−1 × s / √n | Wider with higher confidence or smaller n. |
| Sample size, two means | n per group = 2(Zα/2 + Zβ)²σ² / δ² | About 16 per group for δ = σ, α = 0.05, power 0.80. |
| Gage R&R | GRR variance = repeatability + reproducibility; ndc = 1.41 × (part SD / GRR SD) | Truncate ndc to a whole number. |
| Effect in a 2k design | Mean(high) − Mean(low) | Interaction effect uses the product of the factor columns. |
| R-squared | 1 − SSerror / SStotal | Adjusted R-squared penalizes extra terms. |
| X-bar R limits | X-double-bar ± A2 × R-bar; UCLR = D4 × R-bar; LCLR = D3 × R-bar | Constants depend on subgroup size. |
| I-MR limits | X-bar ± 2.66 × MR-bar; UCLMR = 3.267 × MR-bar | Moving range of two consecutive points. |
| p chart limits | p-bar ± 3√[p-bar(1 − p-bar) / n] | Limits vary with n when the subgroup size varies. |
| Takt time | Available time / customer demand | Pace the process must match. |
| Little's Law | WIP = throughput × lead time | Valid on average for a stable system. |
| OEE | Availability × performance × quality | A single losses-based measure of equipment effectiveness. |
| Payback (months) | Investment / annual benefit × 12 | A simple screen; use NPV for long horizons. |
16. Quick Reference Glossary
| Term | Definition |
|---|---|
| ANOVA | Analysis of variance; compares means across three or more groups by splitting variation into between-group and within-group parts. |
| Attribute agreement analysis | A study of how consistently appraisers classify the same items, summarized with kappa. |
| Alpha (α) | The accepted probability of a Type I error, commonly 0.05. |
| Beta (β) | The probability of a Type II error; power is 1 − β. |
| Blocking | Grouping experimental runs by a known nuisance source of variation. |
| Champion / Sponsor | A leader who removes barriers, secures resources, and approves tollgates. |
| Confounding | When the effects of two factors cannot be separated by the design. |
| COPQ | Cost of poor quality: the cost of failure, appraisal, and the prevention needed to avoid failure. |
| Cp / Cpk | Capability indices comparing process spread, and spread plus centering, with specification limits. |
| CTQ | Critical to quality: a measurable requirement derived from the voice of the customer. |
| DMAIC | Define, Measure, Analyze, Improve, Control. |
| DOE | Design of experiments; a planned set of runs that changes factors on purpose to estimate their effects. |
| DPMO | Defects per million opportunities. |
| FMEA | Failure mode and effects analysis; ranks failure modes by severity, occurrence, and detection. |
| Gage R&R | A study of repeatability and reproducibility of a measurement system. |
| Hypothesis (null and alternative) | The null states no difference or effect; the alternative states that one exists. |
| Interaction | A situation where the effect of one factor depends on the level of another. |
| Kappa | A measure of agreement between appraisers beyond what chance would produce. |
| Lead time | Total elapsed time from request to delivery, including waiting. |
| Little's Law | Average WIP equals throughput times average lead time. |
| MSA | Measurement system analysis. |
| Multicollinearity | Strong correlation between predictors in a regression, which makes coefficients unstable. |
| Multi-vari chart | A chart that shows variation by source, such as within a unit, between units, and over time. |
| Normality | The assumption that data follow a normal distribution; many tests assume normal residuals. |
| Operational definition | A precise description of how a metric is measured, so everyone gets the same result. |
| Pareto chart | Ranked bars of causes or categories with a cumulative line. |
| PCE | Process cycle efficiency: value-added time divided by total lead time. |
| Poka-yoke | Mistake-proofing; design that prevents or immediately detects an error. |
| Power | The probability of detecting a real effect of a stated size. |
| p-value | The probability of a result at least as extreme as observed if the null hypothesis were true. |
| Randomization | Running experimental trials in random order to protect against hidden trends. |
| Replication | Repeating runs to estimate experimental error. |
| Resolution | How strongly effects are confounded in a fractional factorial design. |
| RTY | Rolled throughput yield; the product of step yields. |
| SIPOC | Suppliers, inputs, process, outputs, customers. |
| SPC | Statistical process control; monitoring a process with control charts. |
| Takt time | Available time divided by customer demand. |
| Tollgate | A formal review at the end of a phase to approve, redirect, or stop the project. |
| Type I / Type II error | A false alarm; a missed real effect. |
| VOC | Voice of the customer. |
| VSM | Value stream map. |
Lean Six Sigma Black Belt Pocket Guide: Frequently Asked Questions
What does a Lean Six Sigma Black Belt do that a Green Belt does not?
A Black Belt leads larger, cross-functional projects, uses advanced statistics such as designed experiments and multivariate analysis, coaches Green Belts, and works on improvement most or all of the time. A Green Belt usually leads smaller departmental projects while keeping a regular job.
How long does a Black Belt project usually take?
Most run four to six months from charter to handoff. Projects that run much longer usually have a scope that is too large, and projects that finish much faster often have a known cause that did not need a full DMAIC.
What is an acceptable Gage R&R result?
Under AIAG guidelines, a measurement system with under 10% of study variation is acceptable, 10% to 30% is marginal and depends on risk and cost, and over 30% is not acceptable. The number of distinct categories should be at least 5.
Why use a designed experiment instead of changing one factor at a time?
A designed experiment changes several factors together in a planned pattern, so every run contributes to every effect estimate and interactions become visible. One-factor-at-a-time testing needs more runs for the same precision and cannot detect interactions.
What is the 1.5 sigma shift?
It is a convention that assumes a process mean drifts by up to 1.5 standard deviations over the long term, which is why 6 sigma is quoted as 3.4 defects per million. Always state whether a quoted sigma level includes the shift.
Sources and Further Reading
- Automotive Industry Action Group, Measurement Systems Analysis Reference Manual, 4th ed. (Gage R&R and attribute agreement guidelines).
- Montgomery, D. C., Introduction to Statistical Quality Control, Wiley (control charts, capability, constants).
- Montgomery, D. C., Design and Analysis of Experiments, Wiley (factorial and fractional factorial designs, resolution).
- NIST/SEMATECH, e-Handbook of Statistical Methods, nist.gov (hypothesis tests, ANOVA, regression, control charts).
- Pyzdek, T. and Keller, P., The Six Sigma Handbook, McGraw-Hill.
- ISO 13053-1 and ISO 13053-2, Quantitative methods in process improvement: Six Sigma.
- Womack, J. P. and Jones, D. T., Lean Thinking, Free Press (value stream and flow concepts).
- Hopp, W. J. and Spearman, M. L., Factory Physics, Waveland Press (Little's Law and flow).
- American Society for Quality, Certified Six Sigma Black Belt (CSSBB) certification information, asq.org.