One place for each stage of a designed experiment: what it is for, the steps, the tools, the checklist, and the mistakes to avoid.
FrameDesignRunAnalyzeOptimize
Choose a stage tab below. Each tab opens with an overview, then covers the details needed to do that stage well, with links to the calculators, templates, and guides on this site.
Available now: Frame, Design, Run, Analyze.
The Five Stages at a Glance
A designed experiment changes several inputs on purpose, in a planned pattern, and measures the output, so that every run teaches something about every factor. The work divides into five stages. One example runs through all five tabs: an ultrasonic weld on a sensor housing, where 4.2% of housings leak at final test. All figures in the examples are illustrative.
The stages build on each other. Large problems usually need a sequence of small experiments, with the Optimize stage of one pointing to the Frame stage of the next.
Where an experiment sits in a project: usually the Improve phase
Step 1 of 5
Frame
What do we need to learn, what will we measure, and which factors and ranges are worth testing? Frame turns a vague wish to “try some settings” into a clear question, a response you can trust, a short list of factors with sensible ranges, and a budget. Most of what makes an experiment succeed is decided here, before a single run.
Key question
What do we need to learn, and what must be true before an experiment can answer it?
Typical duration
A few days to two weeks, mostly spent on measurement checks, baseline data, and talking to process experts
Led by
A project leader or Black Belt with the process engineer, operators, and a measurement owner
Starts from
A problem, a gap to a target, or suspected causes from the Analyze phase of a project
Primary outputs
An experiment frame: objective, response, measurement check, baseline, factors and levels, constraints, and approval
Hands off to
Design: a frame that tells the team which design they need and how many runs they can afford
Gate decision
The sponsor and process owner agree the question, the budget, and that the response can be measured well enough
Core tools
Fishbone, P-diagram, Gage R&R, histogram and control chart of baseline data, factor and level table, risk review
What the Frame Step Is For
A designed experiment is a question put to the process. The design, the runs, and the analysis can only answer the question that was asked, with the response that was measured, over the ranges that were tried. Frame is the step that makes sure all three are right before time, materials, and line time are spent.
What Frame must achieve
An objective written as a question and a decision the answer will drive
One primary response that is measured in numbers, with a direction and a unit
A measurement system that is good enough, checked with data
A baseline: where the response is now, how much it varies, and whether it is stable
A short list of factors, each with a reason, a type, and a low and a high level
Constraints and a budget in runs, material, and time
Approval from the people who own the process and the cost
What Frame must not do
Choose a design before the question is clear
Pick a pass/fail response when a measurement is available
Skip the measurement check because the gauge is “what we always use”
Test only factors the team already believes in
Set levels so close that nothing could change
Hide the cost, so the experiment is cancelled halfway
The test of a finished frame. Could an engineer who was not in the discussion read the one-page frame and know what you are trying to learn, what will be measured and how well, which settings will change and by how much, what will be held constant, and what the experiment is allowed to cost? If so, the frame is ready for the Design tab.
The Frame Step, Step by Step
Frame moves from a problem to a priced, approved experiment frame. The steps overlap: a poor measurement check will send you back to the response, and a budget limit will send you back to the factor list.
1
Confirm a designed experiment is the right tool
Compare your situation with the cases in the table below. If the cause is known, fix it. If only one factor is in question, test it directly. If you cannot set the inputs on purpose, use observed data.
Check that you can control the factors you want to test
Check that the response can be measured
Check that the process is stable enough to learn from
Output: A decision: run an experiment, or use another method
Watch for: Reaching for a DOE because it looks rigorous, not because the question needs it
2
State the objective as a question and a decision
Write what you want to learn and what you will do with the answer. Choose the type of experiment: screen many factors, characterize a few, optimize, make the process robust, or confirm a result.
One sentence for the question
One sentence for the decision it will drive
Name the sponsor who will make that decision
Output: An objective and an experiment type
Watch for: An objective like “understand the process”, which no result can fulfil
3
Choose the response and check how well you can measure it
Pick one primary response, preferably continuous, with a unit, a direction (larger, smaller, on target), and an exact method of measurement. Run a measurement system study.
Define how, when, and by whom the response is measured
Check resolution against the process variation
Run a Gage R&R, or the nested version for a destructive test
Output: A defined response and a measurement noise estimate
Watch for: Using a pass/fail judgment, or a gauge that no one has ever checked
4
Establish the baseline
Collect or find data on the response under current settings. Plot it in time order and as a histogram. Estimate the mean, the variation from run to run, and whether the process is stable.
Use at least 30 recent observations
Look for trends, shifts, and mixtures
Record how the response relates to the specification
Output: Baseline mean and standard deviation, and a stability check
Watch for: Experimenting on an unstable process, where drift will look like an effect
5
List the candidate factors
Brainstorm every input that could plausibly move the response. Use a fishbone diagram, a P-diagram, the process map, past data, and the people who run the process.
Include operators, maintenance, and engineers
Look at what changed when the problem began
Record how each factor could be set, and how precisely
Output: A long list of candidate factors
Watch for: A list drawn only from the engineer’s favorite theories
6
Classify the factors and choose which to test
Sort factors into controllable factors you can set, noise factors you cannot control, and factors to hold constant. Choose the controllable factors to put in the experiment, and decide how to handle the rest.
Hold constant what you can
Block or record what you cannot hold
Prefer more factors at wide ranges over fewer at narrow ranges
Output: The factors to vary, the factors to hold, and a plan for noise
Watch for: Letting a noise factor, such as the material lot, change halfway through
7
Set the levels
Choose a low and a high level for each factor. Make them wide enough to move the response, and safe enough that every combination can be run. Keep the current setting between them if possible so center points can be added.
Ask the process experts for the widest safe range
Check the extreme combinations for safety and feasibility
Record exactly how each level will be set and verified
Output: A table of factors, units, and levels
Watch for: Levels so close together that a real effect disappears into the noise
8
Set the constraints and the budget
List how many runs you can afford in material, line time, test time, and money, and what the problem is costing. Check safety, quality, and regulatory approvals.
Count the runs and the units per run
Compare the cost with the value of the answer
Get approval for the line time and the scrap
Output: A run and cost budget, and approvals
Watch for: Finding out halfway that you cannot afford the runs you planned
9
Write the experiment frame and get it signed off
Put everything on one page. Review it with the sponsor, the process owner, and the measurement owner. Agree what a successful result would look like and what will be done with it.
One page, readable by someone who was not there
Sign-off before the Design step starts
Keep it with the project record
Output: A one-page experiment frame, approved
Watch for: Starting the design while the sponsor still has a different question in mind
Is a Designed Experiment the Right Tool?
A designed experiment is the strongest way to find cause and effect, but it costs runs, time, and disruption. Check the situation first.
Drift looks like an effect and ruins the comparison
Ingredients that must sum to 100%
A mixture design (see the Design tab)
Ordinary factorial designs do not fit proportions
Each run is very expensive or very slow
A fractional or sequential design, or a smaller number of factors
Make every run count; use screening before optimizing
The factors cannot be randomized (a hard-to-change setting)
A split-plot design, with expert help
Standard randomization would cost too many setup changes
DOE inside DMAIC. In a project, the experiment usually belongs in the Improve phase, after the Analyze phase has narrowed the suspects. See the DMAIC Toolbox. A DOE can also be used in Analyze to test a short list of suspected causes at once.
Writing the Objective
The objective has two parts: what you want to learn, and the decision it will drive. Without the second part, the experiment produces numbers but not action.
Weak objective
Better objective
Understand the welding process.
Find which of six weld settings have a real effect on burst pressure, so the team can set them to cut leak failures from 4.2% to under 0.5%.
Optimize the process.
Find settings that raise mean burst pressure by at least 20 kPa without increasing its variation, and confirm them on a second resin lot.
See if temperature matters.
Decide whether to add temperature control to the oven by finding out whether a 10 °C change moves yield by more than 2 points.
Type
Question it answers
Typical situation
Common designs
Screening
Which of many factors matter?
Five or more suspects, little prior knowledge
Fractional factorial, Plackett-Burman
Characterization
How do a few factors and their interactions affect the response?
Two to five important factors
Full factorial, with replicates and center points
Optimization
What settings give the best response?
Two or three key factors, curvature suspected
Response surface (central composite, Box-Behnken)
Robustness
How do I make the response insensitive to noise?
Variation from material, environment, or use
Designs with noise factors (robust design, Taguchi)
Confirmation
Does the predicted best setting really work?
After any of the above
A few runs at the chosen setting, with an interval
Large problems are usually solved with a sequence of small experiments: screen, then characterize, then optimize, then confirm. Plan to spend no more than about a quarter of the available runs on the first experiment, so something is left to follow up. The Design tab describes how to choose.
Choosing the Response
The response is the output you will measure in every run. It is the most important choice in the frame, and the one that most often goes wrong.
Type
Example
Information per run
Use
Continuous measurement
Burst pressure (kPa), thickness, time, yield %
High
Always the first choice
Count
Defects per unit
Medium
When a measurement does not exist; needs more runs
Ordinal rating
Appearance on a 1 to 5 scale
Low to medium
Only with a defined scale and trained raters
Pass/fail
Leak or no leak
Very low
Last resort; needs very many units per run
Why it matters, with numbers. The weld leak rate is 4.2%. To see the leak rate fall to 2% with 80% power would need about 975 housings per group. The burst pressure behind the leaks varies with a standard deviation of 22.4 kPa, so seeing a 20 kPa change takes only about 20 housings per group. The measurement is about 49 times cheaper to experiment on.
Choose a response that is close to the cause, and measured directly, such as burst pressure, not the later symptom, such as field returns.
Use one primary response. Add a few secondary responses you will watch (cost, cycle time, appearance), and decide how to trade them off.
Write the operational definition: the method, the instrument, the units, the number of readings, and who measures.
Give it a direction: larger is better, smaller is better, or on target between limits.
Measure the response in the same way for every run, and ideally with the same person and instrument.
Checking the Measurement System
Every run produces a number, and that number is the true value plus the measurement error. If the error is large compared with the effects you want to find, the experiment fails however well it is run. Check before you start.
Check
Question
How
Target
Resolution
Can the gauge show differences much smaller than the process variation?
Compare the smallest reading step with the process standard deviation
Step no more than one tenth of the process standard deviation
Repeatability and reproducibility
How much of the observed variation is the gauge and the people?
Under 10% of study variation is good; 10% to 30% is marginal; plan to average repeats
Bias and stability
Is the gauge accurate, and does it drift?
Measure a reference or a check standard over days
No trend, within the allowed error
Calibration
Is it calibrated and traceable?
Check the record
In date, with a record
Operator method
Does everyone measure the same way?
Write and train a measurement procedure
A written method that all operators follow
What to carry forward. The Design tab needs the size of the measurement error to decide how many runs are needed and whether to average repeated readings. Record it now. If the measurement noise is a large share of the run-to-run variation, plan to average several readings per run.
A marginal gauge is not a reason to cancel, but it is a reason to plan for it: more repeats per run, a more careful procedure, or a better instrument for the experiment even if it is not the routine one.
Establishing the Baseline
The baseline tells you where the response is today, how much it varies when nothing is changed on purpose, and whether the process is stable enough to experiment on. It also supplies the noise estimate that sets how many runs you need.
Gather data from the current process over a representative period: at least 30 observations, from different shifts and material lots if possible.
Plot in time order (run chart or individuals chart) and look for shifts and drift. A stable process has no special causes.
Plot a histogram and compare with the specification. Look for two humps, which suggest a mixture of sources.
Estimate the standard deviation. This is the run-to-run noise that effects must stand out from.
Decide the smallest effect that is worth finding. Effects smaller than this do not change the business result, and they are costly to chase.
Cast the net wide at first. The factor you did not think of is often the one that matters. Ask the people who run the process, look at what changed when the problem started, and use more than one method.
A fishbone diagram for the weld example. Each bone is a category of inputs; the other bones carry more causes than fit on the figure.
Not every factor belongs in the experiment. Sort the long list into three groups, then choose.
A parameter diagram. Blue circles are factors to vary, gold are noise, and gray are procedures to hold the same.
Group
Meaning
What to do
Controllable factors
Settings you can choose and hold
Candidates to vary in the experiment
Noise factors
Inputs you cannot or do not want to control: material lot, ambient conditions, wear
Hold constant for the experiment, block on them, record them, or include them deliberately in a robustness study
Held constant
Everything else that matters
Fix at one level and write down how
Factor
Group
Decision
Reason
Weld amplitude
Controllable
Vary
Strong suspected effect; easy to set
Weld force
Controllable
Vary
Suspected interaction with weld time
Weld time
Controllable
Vary
Directly controls energy delivered
Hold time
Controllable
Vary
Suspected effect on joint cooling
Resin drying time
Controllable
Vary
Moisture suspected; this is a new idea the team has not tested
Clamp pressure
Controllable
Vary
Fixture alignment suspected
Resin lot
Noise
Hold: one lot for the experiment
Confirm on a second lot afterwards
Horn wear
Noise
Hold: new horn, record the cycle count
Wear changes the energy over time
Room humidity
Noise
Record every run
Covariate for the analysis
Loading method
Procedure
Hold: one trained operator
Same method every run
Burst tester and operator
Measurement
Hold: same tester, same person
Reduce measurement variation
Hold constant what you can, and write it down: lot, operator, equipment, and method.
Block what you cannot hold, for example two resin lots, with each lot run as a block (see the Design tab).
Record what you cannot control, such as humidity, so you can check afterwards whether it mattered.
Include a factor if you have a reason to suspect it, even if you doubt it. A surprise factor is the best result an experiment can give.
Do not drop a factor only because it is expensive to set. Raise the cost with the sponsor.
Setting the Levels
The low and high levels define the region the experiment explores. They determine whether an effect shows above the noise, and whether any run can be made at all.
Levels centered on the current setting, so that center points can be added later to check for curvature.
Rule
Why
Bold, but safe. Wide enough that the effect should exceed the noise, narrow enough that every combination makes usable product
Narrow ranges are the main reason real effects are missed; unsafe ranges ruin the experiment
Ask the experts for the extremes they would run on purpose
They know where the process falls over
Check the corners. Every combination of lows and highs must be possible and safe
A 2k design runs every corner, including the odd ones
Center on the current setting when possible
Allows center points to check for curvature
Use real, settable values, with a way to verify them
“High” must mean the same on every run
Quantitative factors (amplitude) use two levels first; categorical factors (supplier, machine) use the levels that exist
Two levels suit screening; more levels add runs quickly
Factor
Unit
Low (−1)
Current
High (+1)
Basis for the range
Weld amplitude
%
70
80
90
Supplier’s range for this horn is 60 to 100; 70 and 90 stay clear of the edges
Weld force
N
400
500
600
Below 400 N the joint does not seat; above 600 N the part flash increases
Weld time
s
0.30
0.40
0.50
Process engineer’s widest settings that give acceptable appearance
Hold time
s
0.5
1.0
1.5
Cycle time limit allows 1.5 s
Resin drying time
h
2
4
6
The dryer cycle runs from 2 to 6 hours in practice
Clamp pressure
kPa
300
375
450
Fixture rating is 500; keep a margin
Constraints and Budget
Every experiment lives within limits. Write them down before the design, because they decide how many runs you can have. Then compare the cost with what the answer is worth.
Constraint
Question
Example
Material
How many units can be used or scrapped?
240 housings and the matching parts are available
Test time
How long does each run and each measurement take?
3 minutes per housing for the burst test
Line time
How much production time can be given up?
2 line days
Capacity
How many housings per day can be welded and tested?
About 32 per day, so up to 64 housings in all
Money
Material, labor, and downtime
$2,432 of housings, plus 40 hours of engineering at $75
Safety and approval
Does the range affect safety, regulation, or the customer?
Quality and safety owners approve the six ranges; no shipment of experimental units
Randomization
Can the run order be randomized?
Yes: the settings can be changed between welds in about 10 minutes
The value of the answer. 130 leak failures a month at $38 each cost about $59,280 a year. Cutting the rate from 4.2% to 0.5% would save about $52,212 a year. The experiment would cost about $5,432, which the saving would repay in roughly 5 weeks.
Plan the runs you can afford, not the runs you wish for. The budget goes to the Design tab, where it sets the number of runs and the number of units per run. If the budget is too small to detect the effect you care about, say so now, and choose between more budget, fewer factors, or a smaller goal.
Worked Example: The Weld Experiment Frame
A plant makes a plastic sensor housing that is joined by ultrasonic welding. Each month about 3,100 housings are tested for leaks at the end of the line, and 130 fail (4.2%), costing $38 each. The team decided to run an experiment on the weld settings. All figures are illustrative, and the same example runs through all five tabs.
Burst pressure of 60 recent housings. 3 fell below the lower limit of 265 kPa (shown in red).
Frame element
What the team wrote
Objective
Find which of six weld settings affect burst pressure, and the settings that raise it by at least 20 kPa without raising its variation, so that leak failures fall from 4.2% to under 0.5%. The sponsor will decide whether to change the work instruction and the machine program.
Experiment type
Screening first, then a smaller characterization and confirmation (a sequence of experiments)
Response
Burst pressure in kPa, measured by destructive test, larger is better, lower specification limit 265 kPa
Measurement check
Nested Gage R&R on the burst tester: measurement standard deviation 4.0 kPa, which is 18% of the total standard deviation (3.2% of the variance): marginal but acceptable for screening if readings are repeated. Resolution 1 kPa against a standard deviation of 22 kPa: adequate
Baseline
60 housings: mean 305 kPa, standard deviation 22.4 kPa, range 261 to 367, 3 below the limit. The individuals chart shows no points beyond the limits (Shapiro-Wilk p = 0.59), so the process is stable and the data look normal. Estimated fraction below the limit 3.6%, consistent with the 4.2% leak rate
Goal
Raise the mean to about 322 kPa at the same variation, which puts the lower limit 2.5 standard deviations away and cuts the expected failure rate to about 0.5%
Smallest effect worth finding
20 kPa, about one standard deviation of weld-to-weld variation
Factors to vary
Amplitude, force, weld time, hold time, resin drying time, clamp pressure, at the levels in the table above
Held constant
One resin lot, new horn (cycle count recorded), same operator and burst tester, same loading method
Recorded, not controlled
Room humidity and temperature
Budget
Up to 64 housings over 2 line days, cost about $5,432, against a saving of about $52,212 a year (5-week payback)
Risks
Out-of-range welds could damage the horn: test the extreme corners first with two housings; no experimental units ship
Sign-off
Sponsor (plant manager), process owner (manufacturing engineering), quality, and the test lab lead
What the frame tells Design. Six controllable factors, a noise level of 22 kPa per housing, a smallest effect of 20 kPa, and a budget of 64 housings. One housing per run would leave the effect hard to see, so averaging several housings per run will be considered in Design.
On the schedule. The team spent two days on the measurement study and baseline, a day with the operators and the process engineer on factors and ranges, and a half day on the one-page frame and sign-off. About three and a half days of preparation, against a line time of two days.
Frame Check: Are We Ready to Design?
Before starting the Design tab, confirm the frame is complete. Use this checklist; progress saves in this browser only, and nothing is sent anywhere.
0 of 15 complete
Questions a sponsor or a coach can ask:
What do you want to learn, and what will you do with the answer?
How do you know the response can be measured well?
What does the process do today, and is it stable?
Which factors did you consider, and why did you keep these?
How wide are the levels, and who agreed they are safe?
What will it cost, and what is the answer worth?
Outcome
Meaning
Next step
Ready
The frame is complete and approved
Start the Design tab
Ready with conditions
Minor gaps, for example one range still to be confirmed
Close them before the design is final
Not ready
The measurement is poor, the process is unstable, or the question is unclear
Fix those first
Wrong tool
The cause is known, the factors cannot be set, or one factor is in question
Use a PDCA cycle, a multi-vari study, or a direct test instead
Adapting Frame to the Situation
Situation
How Frame changes
Manufacturing process
Controllable settings are plentiful; the usual danger is noise from material lots and equipment wear. Block on lots and record equipment state
Chemical and process industries
Responses are often continuous and run times long; check stability and the cost of a failed run; consider mixtures
Transactional and service work
Factors may be scripts, staffing, or sequence; responses are times or error rates; randomize across customers or days; get approval for customer-facing changes
Software and online services
Controlled online experiments (A/B and multivariate tests) need a clear metric, enough traffic, and protection against interference between tests; see the Software and IT hub
Healthcare
Ethics and patient safety come first; use protocols and approvals; many questions call for small tests, not factorial designs; see the healthcare hub
Food, agriculture, and chemistry labs
Biological variation and batch effects are large; block on batch, replicate generously
Where the factors cannot be set
Do not force a DOE. Use a multi-vari study or regression on observed data, and use the findings to plan a later experiment
Scale the framing to the risk and cost. A cheap, fast, reversible experiment needs a half-page frame. An experiment that stops a line, uses scarce material, or affects customers needs the full review.
Common Mistakes and Red Flags
Mistake
What it looks like
How to correct it
Choosing the design first
“Let’s run a Box-Behnken” with no objective written
State the question, then choose
Pass/fail response
Counting leaks, with 12 runs of 20 units
Find the measurement behind the pass/fail
Unchecked gauge
No one knows the measurement error
Run a Gage R&R before the experiment
Unstable baseline
The response drifts or has shifts before any change
When is a designed experiment the right tool, and when is it overkill?
Use a designed experiment when several inputs might affect an output, you can set those inputs on purpose, and you need to know which matter, by how much, and whether they interact. It is overkill when the cause is already known and the fix is obvious, when only one factor is in question (a two-sample test or a PDCA cycle is enough), or when the inputs cannot be controlled (use a multi-vari study or regression on observed data instead).
Why spend so much time framing before choosing a design?
Because the design only answers the question you asked. A perfectly executed experiment on the wrong response, a noisy gauge, or factor ranges that are too narrow gives a clean, useless answer. Most failed experiments fail in the framing: the response could not be measured well, an important factor was left out, or the levels were too timid to move the output.
Should the response be continuous or pass/fail?
Continuous whenever you can get it. A measurement carries far more information than a pass/fail judgment of the same part, so it needs far fewer runs to see the same change. In the worked example, detecting a drop in leak rate from 4.2% to 2% would take about 975 housings per group, while detecting a 20 kPa change in burst pressure takes about 20.
How many factors should I include in the first experiment?
As many as you have good reason to suspect, up to what your budget can screen. Five to eight factors is typical for a screening design; fewer than four can usually go straight to a full factorial. Leaving out an important factor wastes the experiment, while including a few extra costs little in a fractional design.
How wide should the factor levels be?
Wide enough that the effect, if there is one, will exceed the noise, and narrow enough that every combination is safe and produces usable output. Levels set too close together are the most common reason a real effect is missed. Ask the process experts: what is the lowest and highest setting you would run on purpose without scrapping the lot?
What if I cannot measure the response well?
Fix the measurement first. If measurement variation is a large share of what you see, effects will be buried. Run a measurement system study before the experiment. If the gauge is only marginal, plan to take repeated readings per run and average them, which the Design tab covers.
Sources and Further Reading
Douglas C. Montgomery, Design and Analysis of Experiments, Wiley, on guidelines for designing experiments.
George E. P. Box, J. Stuart Hunter, and William G. Hunter, Statistics for Experimenters, Wiley.
Mark J. Anderson and Patrick J. Whitcomb, DOE Simplified, Productivity Press.
NIST/SEMATECH, e-Handbook of Statistical Methods, Process Improvement: Design of Experiments (itl.nist.gov/div898/handbook).
Automotive Industry Action Group, Measurement Systems Analysis Reference Manual, 4th ed.
Minitab Support, documentation for Create Factorial Design and Gage R&R (support.minitab.com).
This content is educational. The example data and results are illustrative. Follow your organization's quality system, change control, and safety requirements.
Step 2 of 5
Design
Which design answers the question within the budget, and how many runs, replicates, and housings per run do we need? Design turns the approved frame into a plan: which runs to make, in what order, how many units to test at each setting, and how well the plan will detect the effect that matters. Check the plan on paper before spending a single housing.
Key question
Which design finds the effect that matters, within the budget, with the smallest risk of a misleading result?
Typical duration
A few days: design options, power check, review with the sponsor
Led by
The project leader or Black Belt, with the process engineer and, if available, a statistician
Starts from
The approved frame: objective, response, noise, smallest effect, factors and levels, budget
Primary outputs
A design with its run count, replicates, blocks, randomization plan, aliasing, and a power check
Hands off to
Run: a design that is ready to turn into a randomized run sheet
Gate decision
The sponsor and process owner accept the design and what it can and cannot tell them
Core tools
Design finder, resolution and alias table, power calculation, DOE Quick Planner, Minitab Create Factorial Design
What the Design Step Is For
The frame says what to learn. The design says how to learn it with the fewest runs that still give a trustworthy answer. A good design makes every run count toward every effect, protects the result from drift and from things you did not control, and says in advance which effects it can and cannot separate.
What Design must achieve
A design that fits the objective and the number of factors
A run count and a number of units per run that the budget allows
A power check: would an effect of the size that matters be found?
Known aliasing: which effects are mixed up with which
A plan for replication, blocking, and center points
A randomization plan, and a plan for hard-to-change factors
A written summary the sponsor can approve
What Design must not do
Choose the design that matches last time’s experiment, not this question
Skip the power check and hope the runs will be enough
Mistake several units at one setup for replicates
Fix the number of runs by what feels convenient
Run in standard order, so drift is mixed with the factors
Hide the aliasing from the people who will read the results
The test of a finished design. Could someone read the design summary and tell you how many setups there will be, which effects can be estimated cleanly, which are mixed up with others, how big an effect the experiment can find, what it will cost, and what the plan is for noise and drift? If so, the design is ready for the Run tab.
The Design Step, Step by Step
Design moves from the objective to an approved plan. The power check often sends you back to change the number of runs or the units per run.
1
Choose the design family
Match the design to the objective and the number of factors: screening, characterization, optimization, robustness, or a mixture.
Use the design finder below
Plan a sequence if needed: screen, then characterize, then optimize
Check that every combination of the chosen levels is safe to run
Output: A design family
Watch for: Using a response-surface design to screen, or a screening design to optimize
2
Choose the number of runs and the resolution
List the full factorial and the fractions that fit the budget. Choose the resolution that keeps the effects you care about clear of each other.
Prefer resolution V or higher if interactions matter
Accept resolution IV for pure screening, with a plan for the follow-up
Avoid resolution III unless you must
Output: A design with a known alias structure
Watch for: A fraction so small that main effects are mixed up with the interactions you suspect
3
Decide the setups and the units per setup
Separate two numbers: how many setups (runs) and how many units to make at each. Use the noise components to see what each choice buys.
Setup-to-setup variation is not reduced by making more units at one setup
More setups usually beat more units per setup
Count the setup time
Output: Setups, units per setup, and total units
Watch for: Treating several units from one setup as replicates
4
Check the power
Use the run-level standard deviation and the number of runs to find the probability of detecting the smallest effect that matters. Compare the options.
Use the smallest effect worth finding from the frame
Aim for at least 80% power
If power is low, change runs, units, or the goal, not the arithmetic
Output: A power figure for each option, and a choice
Watch for: Running the experiment hoping the effect is large
5
Add blocks, center points, and a random order
Block on anything you know will change during the experiment, such as the day. Add center points if curvature is a concern and the budget allows. Randomize the order within blocks.
Confound blocks with high-order interactions
Record the center points as a separate type of run
Plan for any hard-to-change factor
Output: Blocks, center points, and a randomization plan
Watch for: Running one block, then the other, in standard order
6
Summarize and get approval
Write the design on one page: factors and levels, runs, units, blocks, aliasing, power, cost, and time. Review it with the sponsor and process owner and note what the experiment cannot tell them.
State the aliasing in plain language
State the smallest effect it can find
Agree what the next experiment will be if it finds something
Output: A one-page design summary, approved
Watch for: Hiding a limitation that will surface when the results arrive
Finding the Design Family
Start from the objective written in the frame. Large problems usually pass through the families from left to right.
Family
When
Runs (typical)
Limits
Screening (fractional factorial, Plackett-Burman)
Five or more candidate factors; find the few that matter
8 to 32
Aliasing; no curvature unless center points are added
Characterization (full factorial or high-resolution fraction)
Two to five important factors and their interactions
8 to 32, plus replicates
Two levels cannot show curvature
Optimization (response surface)
Two or three key factors where the best setting is inside the region
13 to 30
Needs the important factors already known
Robustness (control factors crossed with noise)
Make the output insensitive to material, environment, or use
A two-level factorial sets each factor at a low and a high level, written −1 and +1, and runs every combination. It estimates every main effect and every interaction. Each run contributes to every estimate, which is why it needs far fewer runs than changing one factor at a time.
Factors (k)
Runs in a full factorial (2k)
Number of effects to estimate
2
4
2 main + 1 two-factor + higher order
3
8
3 main + 3 two-factor + higher order
4
16
4 main + 6 two-factor + higher order
5
32
5 main + 10 two-factor + higher order
6
64
6 main + 15 two-factor + higher order
7
128
7 main + 21 two-factor + higher order
The number of runs doubles with every factor, so beyond four or five factors a full factorial is rarely worth it: most of the runs only estimate three-way and higher interactions, which are usually negligible.
Orthogonal and balanced. Each factor is at each level in half the runs, and all effects can be estimated independently.
Coded levels (−1 and +1) put every factor on the same scale, and make the calculation of effects simple. See Analyzing Designed Experiments.
Interactions come free. A significant interaction means the effect of one factor depends on the level of another.
Fractional Factorials, Aliasing, and Resolution
A fractional factorial runs a selected part of the full factorial: half, a quarter, an eighth. The price is aliasing: some effects can no longer be told apart, because the same pattern of −1 and +1 estimates both. The figure shows the idea with three factors.
Running only the highlighted corners (a half fraction) halves the runs. The third factor is set by multiplying the first two, so its effect cannot be separated from the A × B interaction.
The design is described by its generators (here C = AB) and its defining relation (I = ABC). Every effect is aliased with the effect obtained by multiplying it by the defining relation. The resolution is the length of the shortest word in the defining relation, and it tells you what is mixed up with what:
Resolution
What is aliased
Use it for
III
Main effects are aliased with two-factor interactions
Screening many factors when interactions are assumed small
IV
Main effects are clear of two-factor interactions; two-factor interactions are aliased with each other
Screening, when you want trustworthy main effects
V
Main effects and two-factor interactions are clear of each other (aliased with three-factor interactions or higher)
Characterization when interactions matter
VI and higher
Two-factor interactions are aliased only with four-factor interactions or higher
As good as a full factorial for practical purposes
Factors
Full factorial runs
Half fraction
Quarter fraction
Eighth fraction
3
8
23−1 III (4 runs)
4
16
24−1 IV (8 runs)
5
32
25−1 V (16 runs)
25−2 III (8 runs)
6
64
26−1 VI (32 runs)
26−2 IV (16 runs)
26−3 III (8 runs)
7
128
27−1 VII (64 runs)
27−2 IV (32 runs); 27−3 IV (16 runs)
27−4 III (8 runs)
8
256
28−2 V (64 runs)
28−3 IV (32 runs); 28−4 IV (16 runs)
Assumption behind every fraction. Higher-order interactions (three-way and above) are small. That is usually true, and it is the reason fractions work. If it is not true for your process, the aliased effects will mislead you, so confirm important findings with a follow-up run.
Runs, Replicates, and Units per Run
Two different numbers decide how much information an experiment buys: the number of setups (runs, each with the factors freshly set), and the number of units made or measured at each setup. They are not interchangeable.
Term
Meaning
Reduces which noise?
Setup (run)
One complete setting of all the factors, made fresh
—
Replicate
A complete repeat of a setup: reset, run, and measure again
Setup-to-setup and within-setup noise
Repeat / subsample
Several units made or measured at one setting
Within-setup noise only
Measurement repeat
The same unit measured more than once
Measurement noise only
In the weld example, a short run of 12 consecutive housings at one setting gave a standard deviation of 18 kPa. The overall weld-to-weld standard deviation is 22.4 kPa, which means setup-to-setup variation accounts for the rest: √(22.4² − 18²) = 13.4 kPa. The standard deviation of a run that averages m housings is √(13.4² + 18²/m).
Averaging more housings helps less and less. Beyond about three, the setup-to-setup floor dominates.
The practical rule. Averaging several units at one setup cannot reduce setup-to-setup noise. More setups can. When setup changes are cheap enough, use more runs with fewer units each. When they are expensive, accept more units per run, but compute the run-level noise honestly, and do not use the units as replicates in the analysis.
How Many Runs: Checking the Power
Power is the probability that the experiment finds an effect of a given size when it is really there. It depends on the effect size, the noise of a run, the number of runs, and the error degrees of freedom. For a two-level design:
Standard error of an effect = 2 σrun / √N | t = effect / standard error | power = P(|t| exceeds the critical t, given the true effect)
Smallest effect worth finding (from the frame): 20 kPa.
Run-level standard deviation: 18.5 kPa when two housings are averaged at each setup.
Number of runs N: 32 setups.
Standard error = 2 × 18.5 / √32 = 6.52 kPa.
Error degrees of freedom: 9 (from the three-factor interactions, which are assumed negligible).
Power for a 20 kPa effect = 78%. The smallest effect found with 80% power is about 20.5 kPa.
In Minitab use Stat > Power and Sample Size > 2-Level Factorial Design, which gives the same answer for the number of corner points, replicates, and center points. See Sample Size and Power.
Low power is not a small problem. A real, useful effect will often be missed, and the team will wrongly conclude that the factor does not matter.
If power is low, add runs (or replicates), reduce noise, accept a larger smallest-effect, or drop a factor. Do not just run it and see.
The error estimate must be real. It comes from replicates, from center points, or from higher-order interactions assumed to be zero. A design with none of these cannot test its own effects, and relies on a normal plot of effects.
Comparing the Options
With a budget of 64 housings and 960 line minutes (2 days), the team compared five ways to run six two-level factors. A setup change takes about 10 minutes and each housing about 3 minutes to weld and test.
Design
Setups
Housings per setup
Total housings
Run SD (kPa)
SE of an effect
Power for 20 kPa
Line minutes
A
2^(6-2) IV
16
4
64
16.1
8.1
n/a
352 of 960
B
2^(6-2) IV
32
2
64
18.5
6.5
82%
512 of 960
C
2^(6-1) VI
32
2
64
18.5
6.5
78%
512 of 960
D
2^6
64
1
64
22.4
5.6
94%
832 of 960
E
2^(6-1) VI
32
1
32
22.4
7.9
61%
416 of 960
The same 64 housings can be spread over 16, 32, or 64 setups. More setups gain power, up to the limit of the line time.
Option
Strengths
Weaknesses
A: 16 runs, 4 housings each
Cheapest in time; main effects clear of two-factor interactions
No error estimate at all, so no test of significance; the 15 two-factor interactions are aliased in 7 groups
B: 16 runs replicated, 2 each
A true error estimate from replicates; blocks by replicate
Two-factor interactions are still aliased with each other
C: half fraction, 32 runs
All six main effects and all 15 two-factor interactions can be estimated clearly; blocks into two days
No replicates; the error comes from assuming three-factor interactions are negligible; power just under 80%
D: full factorial, 64 runs
Highest power; every effect estimated
Uses 64 housings and about 87% of the line time; no slack for a failed run, and nothing left for follow-up
E: half fraction, 1 housing each
Half the housings
Power only about 60%; likely to miss real effects
Choice. The team chose Option C: the half fraction with 32 setups and two housings at each. Interactions between weld force and weld time are suspected, so a resolution VI design that keeps all two-factor interactions clear is worth more than the extra margin of Option B. Option D has more power but uses all the housings and most of the line time. Power for a 20 kPa effect is 78%, and effects of about 21 kPa or more would be found with 80% power. That is just at the smallest effect the sponsor asked for, which is stated openly in the design summary.
Blocks, Center Points, and Randomization
Three more decisions protect the experiment from things the factors cannot explain.
Device
What it does
How to use it
Blocking
Separates a known source of variation (day, batch, lot, operator) from the factor effects
Divide the runs into blocks, each containing a balanced part of the design. The block effect is confounded with a high-order interaction you are willing to give up
Center points
Runs at the middle of every factor range: they estimate pure error and test for curvature
Add three to five; they cost little. A significant curvature test means a response-surface experiment is the next step
Randomization
Spreads uncontrolled drift evenly across all the settings
Randomize the run order within each block. Record the actual order. For a hard-to-change factor, talk to a statistician about a split-plot design
Replication
Gives a true error estimate and more power
Reset all settings between replicates. Never treat several units from one setup as replicates
In the example, the 2 line days are the two blocks. Dividing the 32 runs by the sign of the A × B × C column puts 16 runs in each day, with the A × B × C interaction (which is aliased with D × E × F) confounded with the day-to-day difference. The team accepts that: three-factor interactions are expected to be small, and the day effect is then removed from every main effect and two-factor interaction.
Do not run the experiment in standard order. Standard order changes one factor slowly and others fast, so any drift lines up with the slowly changing factor. Randomize within blocks.
Other Designs You Will Meet
Design
Use it when
Notes
Plackett-Burman
Screening up to 11 factors in 12 runs (or 19 in 20, 23 in 24)
Resolution III; use only when interactions are believed small
Definitive screening designs
Screening continuous factors when curvature is possible, with software support
Three levels per factor; main effects are clear of two-factor interactions; needs software
Central composite
Optimizing two to five continuous factors, with curvature
Optimizing three or more factors without running the corners
Avoids extreme combinations
Mixture designs
Recipes where the proportions sum to 100%
Factors are not independent
Split-plot designs
Some factors are hard to change, so full randomization is impractical
Needs a different analysis; ask for help
Taguchi and robust designs
Make the output insensitive to noise
Control factors crossed with noise factors; see Taguchi Methods
Optimal designs
Irregular regions, unusual run counts, or constraints
Computer-generated; understand what the software optimized
Worked Example: The Weld Design
The weld team took the approved frame to the design step. Their inputs: six factors at two levels, a noise level of 22.4 kPa per housing (of which 18 kPa is within a setup), a smallest effect of 20 kPa, a budget of 64 housings, and 960 line minutes. They compared the five options above and chose a half fraction. All figures are illustrative.
Design element
What the team wrote
Design
26−1 fractional factorial, resolution VI, 32 setups; generator F = ABCDE, defining relation I = ABCDEF
Factors (coded −1 / +1)
A amplitude 70 / 90 %; B force 400 / 600 N; C weld time 0.3 / 0.5 s; D hold time 0.5 / 1.5 s; E drying time 2 / 6 h; F clamp pressure 300 / 450 kPa
Units per setup
2 housings welded at each setup; the run response is the average of the two burst pressures (run standard deviation about 18.5 kPa)
Total housings
64 (the approved budget)
Blocks
Two days of 16 setups each, blocked on A × B × C
Replicates and center points
None in this experiment. Four center setups (+4 housings, about $152) are requested as an option and would add a curvature check
Aliasing
Main effects are aliased with five-factor interactions, and two-factor interactions with four-factor interactions: all 6 main effects and 15 two-factor interactions are estimable and clear of each other
Power
78% for a 20 kPa effect; effects of 21 kPa or more are found with 80% power
Time
32 setups × 10 minutes + 64 housings × 3 minutes = 512 minutes, against 960 available, leaving time to repeat a bad setup
Randomization
Random order within each day. Settings are fully reset between setups
Held constant
As in the frame: one resin lot, a new horn with the cycle count recorded, one operator, the same burst tester. Humidity is recorded
Limits the sponsor accepted
No replicates, so the error depends on three-factor interactions being small; power is just below 80%; no curvature test unless the four center setups are approved
The first eight runs of the design in standard order (before randomization), showing how factor F is generated from A to E, and which day each run belongs to:
Std run
A
B
C
D
E
F = ABCDE
Day (block)
1
-1
-1
-1
-1
-1
-1
1
2
+1
-1
-1
-1
-1
+1
2
3
-1
+1
-1
-1
-1
+1
2
4
+1
+1
-1
-1
-1
-1
1
5
-1
-1
+1
-1
-1
+1
2
6
+1
-1
+1
-1
-1
-1
1
7
-1
+1
+1
-1
-1
-1
1
8
+1
+1
+1
-1
-1
+1
2
Some of the aliasing, to show what the half fraction does and does not cost:
Effect
Aliased with
A (Weld amplitude)
BCDEF
B (Weld force)
ACDEF
C (Weld time)
ABDEF
D (Hold time)
ABCEF
E (Resin drying time)
ABCDF
F (Clamp pressure)
ABCDE
A × B (Weld amplitude × Weld force)
CDEF
A × C (Weld amplitude × Weld time)
BDEF
B × C (Weld force × Weld time)
ADEF
D × E (Hold time × Resin drying time)
ABCF
C × F (Weld time × Clamp pressure)
ABDE
On the schedule. Two days to compare the options and the power, a day to build the design in Minitab and check it, and a half day to review with the sponsor and process owner. The 32 setups will take 8.5 hours of line time in the Run step.
Design Check: Are We Ready to Run?
Before starting the Run tab, confirm the design is complete. Use this checklist; progress saves in this browser only, and nothing is sent anywhere.
0 of 13 complete
Questions a sponsor or a coach can ask:
Why this design and not a smaller or a larger one?
How many setups, and how many units at each?
What is the smallest effect it would find, and how sure are we?
Which effects are mixed up with which?
What happens if a run goes wrong?
What will we do if it finds something, and if it finds nothing?
Outcome
Meaning
Next step
Ready
The design is complete, has adequate power, and is approved
Start the Run tab
Ready with conditions
A minor gap, such as a missing center-point approval
Close it before the first run
Not ready
Power is too low, or the aliasing hides the effects you care about
Change runs, units, or the design
Wrong size
The budget cannot reach the smallest effect, even with the best option
Take the choices to the sponsor: more budget, fewer factors, or a larger smallest effect
Adapting Design to the Situation
Situation
How Design changes
Cheap, fast runs (a few minutes each)
Use larger designs and full replicates; the cost of extra runs is small compared with the risk of an inconclusive result
Expensive or destructive runs
Use fractions, plan the power carefully, run sequentially, and consider smaller experiments in a sequence
Hard-to-change factors (furnace temperature, a tool change)
Use a split-plot design, with the hard-to-change factor in the whole plots; ask for statistical help
Many factors, very few runs
Plackett-Burman or a definitive screening design, then confirm; accept aliasing
Curvature expected
Include center points in a screening design, then move to a response surface design
Transactional and service work
Runs are often cheap; randomize across customers, days, or agents; protect against customers who see more than one condition
Software and online services
Controlled online experiments allow many runs and large samples; the design questions are about interference and the metrics
Healthcare and other regulated settings
Ethics and approvals decide what can be randomized; use protocols and review; many questions suit small sequential tests
Use software for the arithmetic, not for the thinking. The DOE Quick Planner, Minitab, and other packages generate designs and aliasing tables in seconds. The choices described in this tab, such as the budget, the noise, and the smallest effect, are yours.
Common Mistakes and Red Flags
Mistake
What it looks like
How to correct it
Several units counted as replicates
Error looks tiny and everything is significant
Replicate by resetting the setup; model the setup noise
No power check
The experiment is run and nothing is significant
Compute power before the run; revisit the budget if low
A fraction that hides the interaction you suspect
Main effect is aliased with the interaction you care about
Choose a higher resolution or fold the design later
Standard run order
Drift lines up with one factor
Randomize, within blocks
No blocking when conditions change
A day effect appears as a factor effect
Block on day, lot, or operator
Too many factors at too many levels
A giant design that cannot be afforded
Screen with two levels, then refine
No error estimate
No replicates, no center points, no negligible interactions
What is the difference between a full factorial and a fractional factorial?
A full factorial runs every combination of the factor levels: 2k runs for k two-level factors. A fractional factorial runs a carefully chosen part of them, such as half or a quarter. It needs fewer runs, but some effects are mixed up (aliased) with others. The design's resolution tells you which effects are mixed up.
What does resolution mean?
It describes which effects are aliased. In a resolution III design main effects are mixed up with two-factor interactions. In resolution IV, main effects are clear of two-factor interactions, but two-factor interactions are mixed up with each other. In resolution V, main effects and two-factor interactions are all clear of each other. The higher the resolution, the more runs it needs.
Are replicates the same as repeated measurements?
No. A replicate is a complete new run: the settings are reset and the whole process is repeated. Repeated measurements or several units made at one setting only capture the variation within that setting. Using them as if they were replicates makes the error look too small and finds effects that are not real. Averaging several units per run is fine, but the error for testing effects must still reflect setup-to-setup variation.
How many runs do I need?
Enough that an effect of the size that matters would stand out of the noise with high probability, usually 80% power. The standard error of an effect in a two-level design is 2σ/√N, where σ is the standard deviation of a run and N the number of runs. Use that, or the Power and Sample Size menu in Minitab, to check the options before you commit.
Should I always add center points?
Add them when curvature is a realistic concern and you plan to stay in the factorial region. They cost few runs, estimate pure error, and test for curvature. They cannot tell you which factor is curved. If the budget is tight, a screening design can skip them and the next experiment can add them.
Why randomize the run order?
Randomizing spreads uncontrolled drift (tool wear, temperature, operator fatigue, material changes) across all the factor settings, so it does not masquerade as a factor effect. If a factor is very hard to change, ask for help with a split-plot design instead of giving up randomization.
When do I use a Plackett-Burman design?
When you need to screen a lot of factors in very few runs and you are willing to assume interactions are small. They are resolution III, so main effects are mixed up with two-factor interactions. A fractional factorial of resolution IV or higher is usually the better choice when you can afford the runs.
Sources and Further Reading
Douglas C. Montgomery, Design and Analysis of Experiments, Wiley, on two-level factorial and fractional factorial designs, resolution, blocking, and power.
George E. P. Box, J. Stuart Hunter, and William G. Hunter, Statistics for Experimenters, Wiley.
Mark J. Anderson and Patrick J. Whitcomb, DOE Simplified, Productivity Press.
NIST/SEMATECH, e-Handbook of Statistical Methods, Process Improvement: Design of Experiments (itl.nist.gov/div898/handbook).
R. L. Plackett and J. P. Burman, “The Design of Optimum Multifactorial Experiments,” Biometrika, 1946.
Minitab Support, documentation for Create Factorial Design and Power and Sample Size for 2-Level Factorial Design (support.minitab.com).
This content is educational. The example data and results are illustrative. Follow your organization's quality system, change control, and safety requirements.
Step 3 of 5
Run
How do we run the experiment exactly as planned, and record what really happened? The best design is worth nothing if the runs are made at the wrong settings, in the wrong order, or recorded badly. Run is the discipline step: follow the plan, protect the data, and write down everything that does not go to plan.
Key question
Is every run made at the planned settings, in the planned order, and recorded so that we can trust the data?
Typical duration
A day to a week of preparation, then the runs themselves; the weld example needs two line days
Led by
The experiment owner, with a setter, an operator, a tester, and a recorder, plus the process engineer on call
Starts from
The approved design: factors, levels, runs, blocks, units per setup, and the randomization plan
Primary outputs
A completed run sheet, the actual settings, the results, noise records, a log of events, and a checked data file
Hands off to
Analyze: clean data with a complete record of what really happened
Gate decision
The experiment owner confirms all runs are complete or accounted for, and the data are checked and locked
Core tools
DOE Run Sheet Generator, pilot run, readback of settings, run log, check standard, time-order plot
What the Run Step Is For
An experiment is a controlled change, and the control is easy to lose. A setting that is a little off, a run done out of order, a lot of material that changed halfway, or a result copied wrongly can each turn a clean design into an ambiguous data set. The Run step protects the plan from the plant.
What Run must achieve
Every setup made at the planned settings, verified by a second person
The runs made in the randomized order, within blocks
Everything that was supposed to be held constant, held constant
The actual settings, the results, and the noise records written down
Every deviation noticed, recorded, and decided on at the time
A data file that is checked, backed up, and traceable to the run sheet
What Run must not do
Reorder runs to save setup time
Quietly fix a run and record it as planned
Let people who know the settings judge the units
Collect results on loose paper that is entered later from memory
Change anything else about the process during the experiment
Start analyzing before the data are checked and locked
The test of a finished Run. Could someone who was not there read the run sheet, the log, and the data file and tell exactly what was done at every setup, in what order, by whom, with which material, and what went wrong? If so, the data can be trusted, even if some runs did not go to plan.
The Run Step, Step by Step
Run moves from a readiness review to a locked data set. Most trouble comes from skipped readiness checks and unrecorded deviations.
1
Review readiness
Walk through the checklist: material set aside, equipment checked, gauges calibrated, people briefed, approvals in place, time booked.
Set aside enough of one lot for every run and the pilot
Confirm the instrument calibration and the check standard
Book the line, the operators, and the lab
Output: A readiness sign-off
Watch for: Starting without enough material of the one lot
2
Generate and print the run sheet
Build the randomized sheet with the planned design, blocks, and units per setup, and print enough copies. Keep the seed with the records.
Use the DOE Run Sheet Generator
Check the first rows by hand
Give each setup a run number that goes on every unit
Output: A randomized run sheet, with the seed recorded
Watch for: Re-sorting the sheet by setting to “save time”
3
Brief the team
Explain why the order matters, who does what, how to record, and what to do when something goes wrong. Make clear that reporting a problem is the right thing.
Name the setter, the welder, the tester, and the recorder
Agree who can stop the experiment
Give them the deviation rules
Output: A team that knows the rules and the roles
Watch for: A team that tries to be helpful by fixing things quietly
4
Run a pilot
Before the first real run, put a few units through the whole process, including the extreme corners. Check that every setting can be reached, that no combination damages the equipment, and that the measurement chain works.
Run the all-low and all-high corners
Test the measurement and the recording form
Time a setup change
Output: A go decision, and a corrected plan if needed
Watch for: Discovering on the first real run that a corner cannot be run
5
Execute setup by setup
For each run in order: reset every factor, have a second person read the settings back, make the units, test them with only the run number visible, and record the actual settings and results immediately.
Do not skip the reset, even when the next setting is the same
Keep everything else the same
Record before moving on
Output: Completed runs, one at a time
Watch for: Batching the recording until the end of the day
6
Record and handle deviations
Write down anything that did not go to plan, when it happened, and what was decided. Use the deviation rules, and call the experiment owner when the rules do not cover the case.
One log, in time order
Decisions made at the time, with initials
Never overwrite a record
Output: A deviation log
Watch for: A fix that is not written down
7
Check the data and close out
Enter the data from the sheet, check it against the paper, plot the results in the actual run order, and look for drift and obvious errors. Lock the data file and keep the originals.
Double-check the entries
Plot by run order and by block
Keep the paper and the file together
Output: A checked, locked data set
Watch for: Analyzing a file that has not been checked against the sheet
Readiness: What Must Be True Before Run 1
Area
Check
Why
Weld example
Material
All units for the experiment and the pilot set aside from one lot, labeled, and stored the same way
A lot change mid-experiment looks like a factor effect
One resin lot; 64 housings plus the pilot housings set aside; resin dried in one batch except for the drying-time factor
Equipment
In good order; wear parts new or at a known state; counters reset
Wear changes the energy delivered
New horn installed; cycle counter reset to zero
Settings
Every factor can be set to every planned level and read back
A level that cannot be reached cannot be run
All amplitude, force, time, hold, and clamp settings verified on the control panel; dryer times verified
Measurement
Calibrated; check standard measured; method written
Poor measurement hides the effects
Burst tester calibrated that morning; same tester and tester operator for all units
People
Briefed and trained; roles assigned; backups named
Surprises cause mistakes
One setter, one welder, one tester, one recorder; engineer on call
Approvals
Safety, quality, and the owner of the product have approved the ranges
Unsafe or unapproved runs are not allowed
Quality and safety approved the six ranges; experimental units are scrapped, not shipped
Time and space
Booked, with room for repeats
A rushed experiment loses discipline
Two line days booked; planned line time well below the time available
Paperwork
Run sheets, log, data form, and labels printed
Writing on scraps invites errors
32 setups on two day sheets, with columns for results and notes
The Run Sheet and the Random Order
The run sheet turns the design into a list of setups in a random order, with space to record what actually happened. Use the DOE Run Sheet Generator: choose the design, enter the factors and levels, set the blocks, replicates, center points, and units per setup, and it produces a printable, randomized sheet. It also reports the resolution and the aliasing, so you can check that it is the design you approved.
Field
What it is for
Run number
The order in which to run the setups. Mark every unit with it
Block (day)
Which block the setup belongs to; finish one block before starting the next
Design row
The row of the underlying design, for the analysis
Factor settings
The planned level of every factor, with the coded level
Setting verified
Initials of the second person who read the settings back
Result columns
One per unit made at the setup
Notes
Anything unusual: time, observation, name
Randomize more than the run order. Also randomize which material goes to which setup, and the order in which units are measured, where practical.
Keep the seed. The same seed gives the same order, so the sheet can be rebuilt if it is lost.
Do not sort the sheet by a factor to make changeovers easier. If a factor is very hard to change, tell the design team: a split-plot design may be needed.
Spread center points through the order, so they can show drift.
Roles and Communication
Role
Does
Does not
Experiment owner
Owns the plan, decides on deviations, signs off the data
Run the setups
Setter
Resets and sets all factors for each run, reads them to the verifier
Make or test units
Verifier
Reads the settings independently and initials the sheet before the first unit
Sign without reading
Operator
Makes the units for the run, with the same method every time
Adjust anything
Tester
Measures each unit, seeing only the run number
Know or guess the setting
Recorder
Writes actual settings, results, noise records, and events in time order
Rely on memory
Process engineer (on call)
Helps if equipment misbehaves
Change settings on the fly
One person may play more than one role in a small team, but the verifier should not be the setter, and the tester should not see the settings. Agree before starting who may stop the experiment, and make it clear that anyone may call a stop.
Executing a Run: The Rules
Do
Do not
Why
Reset every factor to the planned level before each setup, even if it is already there
Skip the reset when two setups look alike
A setup that is not reset is not a new setup and does not give true replication
Have a second person read each setting back
Rely on the setter’s memory
Mis-set levels are the commonest error in experiments
Make all units for a setup before changing anything
Interrupt a setup to do something else
Interruptions change the conditions
Keep everything else the same: lot, operator, method, timing
Improve the method during the experiment
A change halfway ruins the comparison
Label each unit with the run number at once
Label later from memory
Mixed-up units cannot be sorted out afterwards
Record the actual setting, the result, and the time immediately
Write on a scrap and copy it up later
Transcription from memory is the commonest data error
Test units with only the run number visible
Let the tester see the setting
Expectations bias judgments and readings
Stop and ask when something is unexpected
Press on to keep to the schedule
A deviation noticed early is cheap; one found after the analysis is not
What to Record
Record more than you think you need. You can ignore extra information later, but you cannot recover what you did not write down.
Record
Content
Example
Run record
Run number, block, design row, date, start and end time, initials
Run 7, block 1, row 1, 08:42, initials
Actual settings
The value read from the machine for every factor, whether or not it equals the plan
Amplitude 90, force 600, and so on
Results
One value per unit, with units and the tester’s initials
Burst pressure 312 kPa
Noise and covariates
Things you cannot control but can measure: humidity, temperature, horn cycle count, lot, operator
Humidity 46% on day 1
Events log
Anything unusual, in time order, with the decision made and who made it
Dryer timer showed 5.4 h
Unit traceability
Run number on the unit, retained samples
Failed housings kept in a labeled bag
Handling Deviations
Things will go wrong. What matters is that each one is noticed, written down, and handled by a rule agreed in advance, not by whoever happens to be on the line. The table gives defaults; the experiment owner makes the final call.
Event
Default response
Why
A setting is wrong, and no unit has been made
Correct it, have it verified again, note the time lost
No data are affected
A setting is found wrong after units were made
Record the actual setting; either repeat the setup at the end of the block or analyze with the actual value
The data are real but not the planned run
A factor cannot be set to the planned level
Stop. Contact the experiment owner. Do not substitute silently
The design may need to change
Equipment fault or stoppage
Record the time and what was done; check the equipment state before resuming; repeat the setup if units were affected
Faults can change conditions without being obvious
A unit is lost or damaged
Record the reason; keep the other units; repeat the setup if too few remain
Lost units lower precision, and hiding them hides the cause
A result looks extreme
Do not discard. Check the measurement and the record, note the findings, and keep the value unless an error is proven
Extreme values may be the finding
Material runs out or a new lot is needed
Stop, and record. Start a new block with the new lot, if the design allows
A lot change is a block, not a footnote
The schedule is slipping
Do not drop runs or reorder. Ask the owner for more time
A partly completed design is much harder to analyze
Never quietly fix a run. A correction that is not recorded makes the record untrue, and the next person to read it will believe a run went as planned when it did not.
Keeping the Measurement Honest
Same instrument, same method, same person for the whole experiment, where possible.
Check standard. Measure a reference at the start, in the middle, and at the end of each day. A drift in the check standard means the gauge changed, not the process.
Blind the tester. Only the run number on the unit; the setting stays with the setter.
Measure promptly, and in the same conditions: parts can change with time, temperature, or humidity.
Record raw readings, not rounded or converted ones, and note the units.
For destructive tests, keep a record of which unit was tested and in what order. The unit is gone, so the record is all there is.
Checking the Data Before Analysis
Enter the data from the sheet into the file, with a second person checking a sample or the whole file against the paper.
Check plausibility: sort each column and look at the smallest and largest values for typing errors and unit mix-ups.
Plot the results in the order they were run. Look for drift, steps, and the effect of blocks. A trend with run order is a warning that something changed.
Compare the actual settings with the planned settings, and mark any differences.
Reconcile the units: units made, units tested, and units recorded must agree, and each missing one must be explained in the log.
Lock the file and keep a copy of the paper. Every analysis starts from this version.
The weld team ran the design from the Design step: 32 setups, two housings at each, over two line days. All figures are illustrative.
Step
What happened
Readiness
One resin lot set aside for all 64 housings and four pilot housings; new horn installed with its cycle counter at zero; burst tester calibrated; the team briefed on the deviation rules; quality and safety approvals confirmed
Run sheet
Generated with the DOE Run Sheet Generator: half fraction of six factors, two blocks (days), two units per setup, seed 2026. The first setup on day 1 is design row 30, and the second is row 25. Day 1 holds the 16 setups with A × B × C = −1
Pilot (afternoon before)
The all-low and all-high corners, two housings each: both welded cleanly, the horn showed no damage, and the burst tester and the recording form worked. Pilot housings were taken from the spare stock, and were not used in the analysis. 32 minutes
Day 1
16 setups. Humidity 46%. One near miss and one lost housing (see the log)
Day 2
16 setups. Humidity 58%. One setup pulled and rerun at the end of the day
Data
63 of 64 housings gave a valid result. The horn counter read 68 at the end (0 at the start, four pilot welds and 64 experimental housings). The data were entered from the sheets, checked against the paper, and locked
The first eight setups in the order they were run, with the settings read back and the results:
Run
Day
Design row
Amplitude (%)
Force (N)
Weld time (s)
Hold time (s)
Drying (h)
Clamp (kPa)
Results (kPa)
Run average
1
1
30
90
400
0.5
1.5
6
300
293, 300
296.5
2
1
25
70
400
0.3
1.5
6
300
272, 312
292.0
3
1
12
90
600
0.3
1.5
2
450
302, 284
293.0
4
1
23
70
600
0.5
0.5
6
450
327, 324
325.5
5
1
7
70
600
0.5
0.5
2
300
308, 331
319.5
6
1
28
90
600
0.3
1.5
6
300
354, 330
342.0
7
1
1
70
400
0.3
0.5
2
300
245, 254
249.5
8
1
22
90
400
0.5
0.5
6
450
330, 314
322.0
The events log, in time order:
When
Event
Decision
Effect on the data
Day 1, run 7
Setter set amplitude 70 %; the verifier read back 72 % on the panel
Reset, verified again, then welded. 5 minutes lost
None: caught before any unit was made
Day 1, run 14
One of the two housings cracked while being unloaded from the fixture (design row 20)
Kept the other housing; no repeat, because the lost housing was handling damage, not a weld result
The setup has one valid result: the run average is a single value (313 kPa)
Day 2, run 25
The dryer timer for the resin showed 5.4 hours instead of the planned 6 for design row 27
Pulled the setup before welding; dried 40 more minutes; reran it as the last setup of day 2
Moved from position 25 to position 32; planned time 512 minutes became 557
Both days
Humidity recorded at the start of each day
46% on day 1 and 58% on day 2
Recorded as a covariate; the day block also carries it
Run averages in the order they were run. No drift or step is visible; the gold points are day 2.The experiment took 557 minutes against 960 available, leaving room for the lost time.
Data check
Result
Setups completed
32 of 32
Valid housings
63 of 64 (one lost to handling)
Burst pressure, all valid housings
mean 304.7 kPa, standard deviation 29.0, range 245 to 366
Day averages (of setup averages)
day 1 303.8 kPa and day 2 306.0 kPa: no important day-to-day difference
Drift with run order
None visible on the run chart
Entries checked against the paper
100% of entries, by a second person
On the schedule. A day of preparation and the pilot, two line days for the runs, and a half day to check, lock, and hand over the data. The data go to the Analyze tab exactly as recorded, with the log.
Run Check: Is the Data Ready for Analysis?
Before starting the Analyze tab, confirm that the experiment is complete and the data are trustworthy. Use this checklist; progress saves in this browser only, and nothing is sent anywhere.
0 of 17 complete
Questions an experiment owner or coach can ask:
Did every run follow the sheet? If not, what changed and who decided?
Who verified the settings, and where is the record?
What was held constant, and did it stay constant?
Where are the original sheets, and who checked the data against them?
Does the run chart show any drift or any block effect?
Is there anything we did not write down?
Outcome
Meaning
Next step
Ready
All runs complete, deviations logged, data checked and locked
Start the Analyze tab
Ready with conditions
One or two documented deviations that the analysis can handle
State them in the analysis and check whether they matter
Not ready
Missing runs, unlogged changes, or unchecked entries
Repeat the runs or fix the records first
Invalid
A major change occurred (new lot, broken equipment) that was not blocked
Treat the experiment as incomplete: repeat it, or analyze with the change as a block
Adapting Run to the Situation
Situation
How Run changes
Manufacturing line
Protect the experiment from the normal production flow: label and segregate units, and hold material, operator, and equipment constant
Long processes (days per run)
Plan the schedule carefully; record the conditions through the run; consider a smaller design
Batch and chemical processes
A run is a batch: record the batch conditions; randomize batch order; watch for carry-over between batches
Transactional and service work
Randomize across customers, days, or agents; keep the script or the method fixed; protect customers who see more than one condition
Software and online services
Randomization is automatic, but check that the assignment really is random, that logging works, and that overlapping experiments do not interact
Healthcare and regulated work
Follow the approved protocol exactly; document every deviation formally; follow consent and ethics rules
Outdoor and field work
Weather and site conditions are noise: record them, and block on them where you can
Scale the discipline to the risk. A cheap, quick, internal experiment needs a simple sheet and a log. A costly or regulated one needs signatures, formal deviation reports, and retained samples.
Common Mistakes and Red Flags
Mistake
What it looks like
How to correct it
Standard-order running
All the high settings of one factor done together
Follow the randomized sheet, within blocks
No reset between similar setups
Two setups ran back to back with no change
Reset and verify every setup
Unverified settings
A mis-set level found after the analysis
A second person reads every setting
Unrecorded fixes
Data look strange and nobody remembers why
Log every event, with the decision and the name
The lot changed halfway
A shift in the response with the second half of the runs
One lot, or a new lot as a block
Tester knows the settings
Results move in the expected direction too neatly
Label by run number only
Recording after the fact
Entries copied from memory or scraps
Record at the machine, immediately
Skipping the pilot
A corner of the design cannot be run on day 1
Run the extreme corners first
Dropping runs to save time
A design with holes that cannot be analyzed cleanly
Ask for more time instead
Analyzing before checking
A typing error becomes a finding
Check the data against the paper and plot them first
Why must the runs be done in the randomized order?
Randomizing spreads drift, such as tool wear, temperature, material changes, and operator fatigue, across all the factor settings. If you run in a convenient order, any drift lines up with some factor and looks like an effect of that factor. If a particular order is truly impossible, say so in the design step so that a split-plot or blocked plan can be used, and do not reorder on the fly.
What should I do if I notice a mistake in the middle of the experiment?
Stop, record exactly what happened, and decide with the experiment owner. If the mistake happened before any unit was made, correct it and carry on. If units were made at the wrong setting, record the actual setting, and either repeat the run at the end or analyze it with the actual values. Never quietly fix a run and pretend it went as planned.
Should the people who measure the response know the settings?
Where possible, no. Label units with a run number only, and keep the settings sheet with the person who set the machine. A measurement made by someone who expects a larger value can be nudged toward it. For machine-read measurements this matters less, but the labeling is still good practice.
What if a unit is lost or a measurement is clearly wrong?
Record it in the log with the reason, keep the other units from that setup, and do not replace or discard it in the data file without a note. A lost unit is a deviation to report, not a problem to hide. If the loss is serious, repeat the setup at the end of the block.
Do I record settings or just the results?
Record both. Write down the setting you actually read on the machine for every factor, even when it equals the plan, and record the noise factors that you chose not to control: humidity, resin lot, horn cycle count, operator, time. Those records are what allow you to explain a surprise afterwards.
How many pilot runs should I do?
Enough to test the extremes and the measurement chain: the combinations of all-low and all-high settings, and a few housings through the whole process from setup to recorded result. Pilot units are not part of the experiment, and they are paid for outside its budget.
Sources and Further Reading
Douglas C. Montgomery, Design and Analysis of Experiments, Wiley, on guidelines for conducting experiments.
George E. P. Box, J. Stuart Hunter, and William G. Hunter, Statistics for Experimenters, Wiley.
Mark J. Anderson and Patrick J. Whitcomb, DOE Simplified, Productivity Press.
NIST/SEMATECH, e-Handbook of Statistical Methods, Process Improvement: Design of Experiments (itl.nist.gov/div898/handbook).
Automotive Industry Action Group, Measurement Systems Analysis Reference Manual, 4th ed.
This content is educational. The example data and results are illustrative. Follow your organization's quality system, change control, and safety requirements.
Step 4 of 5
Analyze
Which factors and interactions really matter, how big are they, and can we trust the model? Analysis turns 32 setup averages into a short list of effects you can believe, a model that predicts the response, and a clear statement of what the experiment could not tell you. The aim is a defensible conclusion, not a long table.
Key question
Which factors and interactions are real, how big are they, and does the model hold up?
Typical duration
A day or two: effects, plots, a reduced model, checks, and a short write-up
Led by
The project leader or Black Belt, with the process engineer to judge whether the effects make physical sense
Starts from
The locked data file, the run log, and the design summary with its aliasing
Primary outputs
Effects with a measure of certainty, a checked reduced model, plots, and a statement of limits
Hands off to
Optimize: a model that can predict, and a list of which settings matter and which do not
Gate decision
The model passes its checks, and the team agrees the effects make engineering sense
Core tools
Pareto and normal plot of effects, ANOVA, main effect and interaction plots, residual plots, Minitab Analyze Factorial Design
What the Analyze Step Is For
The Run step produced numbers. The Analyze step decides what they mean. In a factorial experiment each effect is a simple difference between averages, so the arithmetic is easy; the hard part is judging which differences are bigger than the noise, whether the model describes the process, and whether the answer is useful.
What Analyze must achieve
A look at the raw data and the log before any model
Effects for every factor and interaction the design can estimate
A fair test of which effects are real, with an error estimate that fits the design
Plots that show the findings without a table
A reduced model that respects hierarchy and passes its checks
A comparison of every effect with the smallest effect worth finding
A clear statement of what the experiment could not tell us
What Analyze must not do
Treat several units at one setup as independent replicates
Test every term at 5% and report whichever passes
Drop a parent main effect while keeping its interaction
Skip the residual checks because the p-values look good
Interpret main effects alone when an interaction is large
Present a prediction without its uncertainty
The test of a finished analysis. Could someone read one page and tell which factors matter and by how much, which do not, how sure we are, what assumptions the conclusion rests on, and what we still do not know? If so, the analysis is ready for the Optimize step.
The Analyze Step, Step by Step
Analysis moves from the data to a conclusion. A failed residual check sends you back to the model.
1
Look at the data first
Read the log, plot the results in run order and by block, and note every deviation. Decide the unit of analysis: one value per setup.
Check for drift, steps, and block differences
Treat events from the log as part of the data
Average the units at each setup
Output: One response per setup, and a list of known issues
Watch for: Analyzing every housing as if it were independent
2
Estimate all the effects
Fit the full model: every main effect and every two-factor interaction the design can separate. The effect of a term is the average response at its high level minus the average at its low level.
Use coded levels so effects are comparable
Estimate the block effect separately
Keep the aliasing in mind when you read the list
Output: A table of effects for all terms
Watch for: Fitting a model with terms the design cannot separate
3
Judge which effects are real
Compare each effect with an error estimate. Use the error from higher-order interactions, Lenth's pseudo standard error, or replicates, and look at a normal plot of the effects.
Use two methods and see whether they agree
Look at the plot as well as the p-values
Treat a borderline effect as a question for the next experiment
Output: A short list of effects that stand out of the noise
Watch for: Using p-values from a model with no real error estimate
4
Plot the findings
Draw the main effects and the interaction plots for the important terms. Plots show the direction and the size, and they show the interaction in a way a table cannot.
Lines for main effects; two lines for each interaction
Same vertical scale on every plot
Mark the specification or the target
Output: Plots that tell the story
Watch for: Reading a main effect that is part of a large interaction
5
Reduce the model
Keep the real terms, plus their parents. Fit the reduced model, look at the ANOVA, the coefficients, and R-squared, adjusted and predicted.
Keep hierarchy
Use the reduced model's error to test and to predict
Keep the block in the model if the design had blocks
Output: A reduced model with coefficients and an ANOVA table
Watch for: Removing terms until only the significant ones are left, ignoring hierarchy
6
Check the model
Plot the residuals against the fitted values, in run order, and on a normal probability plot. Look at outliers and influence.
Random scatter, no funnel
No trend in run order
Points near the line in the normal plot
Output: A checked model, or a reason to change it
Watch for: Trusting the p-values of a model that fails its checks
7
State the practical meaning and the limits
Compare each effect, with its interval, with the smallest effect worth finding. Say what the experiment could not tell you: curvature, factors held fixed, the range tested.
Intervals, not only p-values
Say which settings do not matter and so can be chosen for cost or convenience
List what the next experiment should check
Output: A one-page conclusion
Watch for: Reporting effects with no sense of their size or uncertainty
Start With the Data, Not the Model
Read the log. Every deviation affects how you analyze: a lost unit, a setup that moved in the order, a block change.
Plot the results in the order they were run, and by block. A step between blocks, or a trend with run order, is a warning that something changed.
Choose the unit of analysis. The experimental unit is the setup, not the housing. Use one value per setup, normally the average of its valid units. A setup with one unit, such as the one with a cracked housing, is simply noisier, and a weighted analysis can check that it matters.
Know what you cannot test. With no replicates and no center points, the design cannot give a pure-error estimate or test curvature, so those assumptions need care.
With coded levels, the effect of a factor is the average response at its high level minus the average at its low level. For weld force (B), the 16 setups at 600 N average 320.1 kPa and the 16 at 400 N average 289.7 kPa, so the effect is +30.4 kPa. Interactions are estimated in the same way, from the product of the two coded columns. The sum of squares for a term is N × effect² / 4: for force, 32 × 30.4² / 4 = 7,412.
The design is a resolution VI half fraction, so every main effect and every two-factor interaction has its own column: six main effects plus 15 two-factor interactions, which is 21 effects, with 9 degrees of freedom left from the three-factor interactions (A × B × C is confounded with the day). The +2.2 kPa difference between days is small.
Term
Effect (kPa)
Standard error
t
p-value
Beyond the noise?
B Weld force
+30.4
5.4
+5.66
< 0.001
Yes
E Resin drying time
+20.9
5.4
+3.89
0.004
Yes
C Weld time
+20.9
5.4
+3.88
0.004
Yes
B × C
-15.7
5.4
-2.92
0.017
Yes
A Weld amplitude
-6.3
5.4
-1.17
0.271
No
B × F
-5.0
5.4
-0.93
0.377
No
D × E
+4.8
5.4
+0.89
0.394
No
A × B
-3.1
5.4
-0.58
0.576
No
D × F
-3.1
5.4
-0.57
0.583
No
A × C
-2.9
5.4
-0.55
0.598
No
How the standard error was found. The 9 distinct three-factor interactions (the tenth, A × B × C, is the block) are assumed to be zero, so their variation is error: mean square 231, standard error of an effect = √(4 × 231 / 32) = 5.38 kPa, with 9 degrees of freedom. An effect is significant at 5% when it exceeds 2.26 × 5.38 = 12.2 kPa.
Which Effects Are Real?
With no replicates there is no direct measure of pure error, so the experiment estimates it indirectly. Use more than one method; where they agree you can be more confident.
Method
How it works
In this experiment
Pooled higher-order interactions
Treat three-factor and higher interactions as zero, so their variation is error
Error from 9 degrees of freedom: standard error 5.4 kPa; line at 12.2 kPa; real effects: B, E, C, BC
Lenth's method
Estimates the noise from the median size of the effects, ignoring the large ones
Pseudo standard error 3.8; margin of error 8.9 kPa; real effects: B, C, E, BC
Normal plot of effects
Effects from noise fall on a straight line; real ones fall away from it
The same four effects stand clear of the line (see the plot)
Replicates or center points
A direct pure-error estimate
Not available in this design
The Pareto chart ranks the effects. Bars beyond the dashed line are larger than the noise; the sign is beside each bar.
Both methods agree. Weld force (B), drying time (E), weld time (C), and the force by weld time interaction (B × C) stand out. Every other effect is smaller than 9 kPa, which is under both cut-offs for noise. The smallest real effect, B × C, has p = 0.017: a result worth confirming rather than leaning on.
Plotting the Findings
A normal plot of the effects. Small effects, which are noise, line up on the dashed line; the labeled points do not.Main effects. Steeper lines are bigger effects. The gray lines are factors whose effect is within the noise.The interaction plot. The lines are not parallel, which is what an interaction looks like.
Force and drying time raise burst pressure from low to high by 30 and 21 kPa. Longer drying probably helps by removing moisture from the resin, a mechanism the team had not tested before.
The force by weld time interaction means the factors trade off. At low force, a longer weld time raises burst pressure by 37 kPa; at high force it adds only 5. Enough force makes weld time nearly irrelevant, so reading the weld time main effect alone would mislead.
Amplitude, hold time, and clamp pressure had no effect that stands out of the noise, so they can be set for cost, cycle time, or convenience, within the tested range.
The Reduced Model
Keep the real effects, B, C, E and B × C (whose parents B and C are already there), and the day block. Fitted to the 32 setup averages, the reduced model has 26 degrees of freedom for error, so its tests and its intervals use all the information in the experiment.
Source
DF
Sum of squares
Mean square
F
p-value
Model
5
16,414
3,283
25.7
< 0.001
Day (block)
1
40
40
0.3
0.578
B Weld force
1
7,412
7,412
58.1
< 0.001
C Weld time
1
3,486
3,486
27.3
< 0.001
E Resin drying time
1
3,507
3,507
27.5
< 0.001
B × C
1
1,969
1,969
15.4
< 0.001
Error
26
3,317
128
Total
31
19,731
Term
Effect
Coefficient
SE
t
p-value
95% interval for the effect
Constant
304.88
2.00
152.69
< 0.001
Day (block)
+2.2
1.12
2.00
0.56
0.578
-6.0 to +10.5
B Weld force
+30.4
15.22
2.00
7.62
< 0.001
+22.2 to +38.6
C Weld time
+20.9
10.44
2.00
5.23
< 0.001
+12.7 to +29.1
E Resin drying time
+20.9
10.47
2.00
5.24
< 0.001
+12.7 to +29.1
B × C
-15.7
-7.84
2.00
-3.93
< 0.001
-23.9 to -7.5
Model in coded units (−1 at the low level, +1 at the high level): Burst pressure = 304.9 +1.1 Day + 15.2 B + 10.4 C + 10.5 E −7.8 B×C. Standard deviation of the residuals S = 11.3 kPa; R² = 83%, adjusted 80%, predicted 75%. The predicted R² is close to the ordinary one, so the model is not just fitting the noise.
Noise came in lower than planned. The Design step assumed a run-level standard deviation of about 18.5 kPa. The residual standard deviation here is 11.3. The planning figure was cautious, which is the safe way to plan: the experiment ended up with better power than the 78% promised, and the effects are more certain than the design promised.
Sensitivity check. One setup has a single valid housing, and so is noisier. Refitting with each setup weighted by the inverse of its variance changes no effect by more than 0.3 kPa, so the simple analysis stands.
Checking the Model
Residual plots: random scatter, no trend with run order, and a straight line in the normal plot are what you want.
Check
What to look for
Result
Residuals versus fitted
No curve, no funnel
Random scatter. Spread in the upper half of the fitted values is 1.24 times that in the lower half (Levene's p = 0.28)
Residuals versus run order
No trend, no step at the day change
Correlation with run order -0.01; Durbin-Watson 2.08 (2 means no serial correlation)
Normal plot of residuals
Points near the line
Close to the line (Shapiro-Wilk p = 0.84)
Outliers and influence
Standardized residual under about 2.5; Cook's distance well below 1
Largest standardized residual 2.19; largest Cook's distance 0.19
Curvature
A significant difference between center-point average and factorial average
Cannot be tested: there are no center points. Check in the next experiment
The frame set the smallest effect worth finding at 20 kPa. Compare every effect, with its interval, against it.
Effect
Size (kPa)
95% interval
Compared with 20 kPa
Meaning
B Weld force
+30.4
+22.2 to +38.6
1.5 times
Clear of the 20 kPa line even at the lower end
E Resin drying time
+20.9
+12.7 to +29.1
1.0 times
Real; the lower end of the interval is below 20 kPa
C Weld time
+20.9
+12.7 to +29.1
1.0 times
Real; the lower end of the interval is below 20 kPa
B × C
-15.7
-23.9 to -7.5
0.8 times
Real; the lower end of the interval is below 20 kPa
Factors with no real effect do not prove that the factor has no influence. The interval tells you how large an effect could still be hiding:
Effect
Size (kPa)
Interval from the pooled error
Meaning
A Weld amplitude
-6.3
-18.5 to +5.9
Under 20 kPa either way: set for cost or convenience
D Hold time
+0.2
-11.9 to +12.4
Under 20 kPa either way: set for cost or convenience
F Clamp pressure
-1.2
-13.4 to +11.0
Under 20 kPa either way: set for cost or convenience
Weld force is the main lever: 30 kPa, about 1.5 times the smallest effect that matters.
Drying time and weld time each give about 21 kPa, right at the size that matters, with the lower end of the interval below it. Both look worth using, and both are worth confirming.
The interaction has a practical use. At high force, weld time barely matters, so the team can shorten the weld time to save cycle time without losing strength.
Amplitude, hold time, and clamp pressure can be set to whatever is cheapest, because even their upper limits sit below 20 kPa. Their ranges were safe, and they were held within them.
When the Results Look Odd
Symptom
Likely cause
What to do
Nothing is significant
Noise larger than planned, narrow ranges, low power, or a lost run
Check the measurement and the ranges; report the intervals; plan the next experiment with more runs or bolder ranges
Everything is significant
Error estimate too small, for example units treated as replicates
Analyze one value per setup; check the error degrees of freedom
A large interaction hides a main effect
The factor helps in one condition and hurts in another
Read the interaction plot, not the main effect
Residuals curve with the fitted values
Curvature, or a missing term
Add center points and axial runs in a follow-up; consider a transformation
Residuals fan out
Variation grows with the response
Consider a log or other transformation, or analyze the variation separately
A trend with run order
Drift: wear, temperature, material
Add run order or the covariate to the model; repeat if it is large
One point is far from the rest
A recording error, a lost or damaged unit, a real event
Check the log and the records; do not delete without a reason; analyze with and without it
An effect is aliased with one you did not expect
A fraction of low resolution, or an alias you forgot
Check the alias table; fold over or run the other fraction to separate them
Effects do not make engineering sense
A mislabeled column, a coding error, or a surprise
Check the coding first, then discuss with the process engineer; confirm the surprise before using it
Enter one row per setup with the average response and the coded −1 and +1 columns for every factor and the block.
Effects:=AVERAGEIF(col, 1, response) - AVERAGEIF(col, -1, response) for each factor. Make an interaction column by multiplying two coded columns.
Model:Data > Data Analysis > Regression with the response as Y and the block, B, C, E, and B×C columns as X. The coefficients are half the effects, and the output has the ANOVA table, S, and R-squared.
Error for the full model: square and sum the three-factor effects (× N/4), divide by their count, and use the formula for the standard error given above.
Plots: a scatter of the residuals against the fitted values, and a line chart of the cell averages for the interaction. Excel has no normal plot of effects; sort the effects and plot them against =NORM.S.INV((i-0.5)/n).
Minitab
Stat > DOE > Factorial > Analyze Factorial Design. Choose the response (one value per setup), and in Terms include the factors and two-factor interactions, with the block.
In Graphs, choose a Pareto of effects, a normal plot of effects, and the four-in-one residual plots. With no error degrees of freedom Minitab uses Lenth's method for the reference line.
Look at the Pareto and the normal plot, then use Terms or Stepwise to fit the reduced model, keeping parents of interactions.
Stat > DOE > Factorial > Factorial Plots for the main effects and interaction plots.
Read the session window: the Analysis of Variance table, the Model Summary, and the Coded Coefficients, where Minitab gives the effect as twice the coefficient.
Typed excerpt of the reduced-model output, simplified from the calculated example. Menu names follow recent versions of Minitab Statistical Software and can differ slightly in older releases.
Worked Example: Analyzing the Weld Experiment
The team took the locked data from the Run step: 32 setup averages, from 63 valid housings, over two days. All figures are illustrative.
Step
What the team found
First look
The run chart showed no drift. Day averages differ by +2.2 kPa, which is small. The setup with one valid housing is noisier but a weighted check changes no effect by more than 0.3 kPa. The analysis uses one value per setup
Effects
Of 21 effects, four stand out: B weld force +30.4, E drying time +20.9, C weld time +20.9, and B × C -15.7 kPa. The largest of the rest is A at -6.3
Which are real
The error from 9 three-factor interactions gives a line at 12.2 kPa; Lenth's margin is 8.9. Both methods and the normal plot pick the same four
Plots
Main effects: force, drying time, and weld time raise the burst pressure. The interaction plot shows that at 600 N the weld time barely matters
Reduced model
Day, B, C, E, B × C. R² 83%, adjusted 80%, predicted 75%; S = 11.3 kPa, below the 18 planned
Checks
Residuals are random against the fitted values and the run order, and are close to normal (Shapiro-Wilk p = 0.84). No point has a Cook's distance above 0.19
Practical meaning
Force (30 kPa) is the main lever; drying time and weld time give about 21 each. Amplitude, hold time, and clamp pressure are within the noise and can be set for cost
Limits
No center points, so curvature is untested; drying time is at the edge of its tested range, so more may help but is unproven; one resin lot only; the aliasing is clean, but the three-factor interactions were assumed to be negligible
What goes to Optimize. A model with four active terms, a residual standard deviation of 11.3 kPa, and a short list of settings that do not matter. The next question is: what settings give the target of 322 kPa on average, with a good margin, and does the prediction hold in a confirmation run?
Analyze Check: Do We Understand the Results?
Before starting the Optimize tab, confirm that the analysis is complete and the model can be trusted. Use this checklist; progress saves in this browser only, and nothing is sent anywhere.
0 of 12 complete
Questions a sponsor or a coach can ask:
Which factors matter, and by how much?
Which do not, and how sure are you?
Is any effect large enough to use, and is it certain enough?
Does the interaction change the advice?
What did the residual plots show?
What could the experiment not tell us?
Outcome
Meaning
Next step
Ready
The model passes its checks, and the effects make sense
Start the Optimize tab
Ready with conditions
A weak point, such as one borderline effect or untested curvature
Carry it as a question into the confirmation
Not ready
A failed residual check, or an unexplained pattern
Fix the model, or collect more data
Inconclusive
Nothing stands out of the noise, or the intervals are too wide
Plan a follow-up with more power or bolder ranges
Adapting Analyze to the Situation
Situation
How Analyze changes
Replicated designs
Use the replicates for pure error and a lack-of-fit test; use a full ANOVA with all terms and drop the non-significant ones
Designs with center points
Add the curvature test: the center-point average against the factorial average
With several units per setup, analyze the standard deviation or its log as a second response; with few units it is very noisy
Many responses
Analyze each, then find settings that satisfy all of them with the optimizer
Unbalanced or incomplete designs
Use regression with the actual data, check the aliasing and the variance of each effect, and be careful with orthogonality
Software and online experiments
Use the metric's own noise; consider multiple-comparison corrections when many metrics are tested
Healthcare and regulated settings
Follow the pre-specified analysis plan, and document any departure
Decide the analysis before seeing the data. A pre-planned analysis, with the model terms and the error estimate agreed in advance, protects against finding what you hoped to find.
Common Mistakes and Red Flags
Mistake
What it looks like
How to correct it
Units as replicates
Dozens of significant terms and a tiny error
One value per setup; use the setup as the experimental unit
No real error estimate
A table of p-values from a model with no degrees of freedom left
Use pooled higher-order terms, Lenth's method, or replicates
Ignoring hierarchy
An interaction in the model without its parent main effects
Keep the parents
Reading main effects through an interaction
Advice contradicted by the interaction plot
Interpret the interaction plot first
Skipping the residual checks
Strange patterns appear later in the confirmation
Plot the residuals before trusting any p-value
Deleting an unusual run without a reason
A clean model that was cleaned by hand
Check the log; analyze with and without it; report
Reporting p-values only
No sense of size
Give effects with intervals and compare with the smallest effect that matters
Over-reading a borderline effect
A decision rests on p = 0.04
Confirm with runs before changing the process
Extrapolating
Settings outside the tested ranges
Stay inside the region; run a new experiment to go beyond it
Over-fitting
R-squared climbs and predicted R-squared falls
Reduce the model; compare adjusted and predicted R-squared
How do I tell which effects are real when there are no replicates?
Use an error estimate that does not need replicates. In a fractional factorial, the three-factor and higher interactions are usually negligible, so their pooled sum of squares estimates the error. Lenth's method estimates it from the effects themselves, and a normal probability plot of the effects shows the real ones standing apart from the line of small ones. Using more than one method and seeing them agree is reassuring.
What is the difference between statistical and practical significance here?
A statistically significant effect is unlikely to be noise. A practically significant effect is big enough to matter for the decision. Compare each effect, and its interval, with the smallest effect worth finding that you set in the frame. An effect can be real and too small to use, or too uncertain to rule out as important.
Should I analyze each unit or the average of the units at each setup?
Analyze one value per setup, usually the average of its units. Units made at one setup share the same setup, so they are not independent evidence about the factors. Treating them as independent makes the error look too small and finds effects that are not there. The Design step planned for this: the setup is the experimental unit.
Why keep a non-significant main effect in the model when its interaction is significant?
Because the model should respect hierarchy: if an interaction is in the model, its parent main effects should be too. Dropping a parent distorts the interaction and makes the model depend on how the factors were coded. In the example, both parents of the force by time interaction are significant anyway.
What if nothing is significant?
Check the possible reasons before concluding that nothing matters: measurement noise, factor ranges that were too narrow, an unstable process, a lost run, or too little power for an effect of the size that matters. Report the intervals. A well-run experiment that finds nothing larger than the smallest effect worth finding is still a result: it tells you where not to look.
Can I trust the model outside the region I tested?
No. A two-level model describes a flat surface, and the process may curve beyond or between the levels. Without center points the design cannot even test for curvature inside the region. Use the model to choose settings within the tested ranges, and confirm with runs before changing the process.
Sources and Further Reading
Douglas C. Montgomery, Design and Analysis of Experiments, Wiley, on the analysis of two-level factorial designs.
George E. P. Box, J. Stuart Hunter, and William G. Hunter, Statistics for Experimenters, Wiley.
Russell V. Lenth, “Quick and Easy Analysis of Unreplicated Factorials,” Technometrics, 1989.
Cuthbert Daniel, “Use of Half-Normal Plots in Interpreting Factorial Two-Level Experiments,” Technometrics, 1959.
NIST/SEMATECH, e-Handbook of Statistical Methods, Process Improvement: Design of Experiments (itl.nist.gov/div898/handbook).
Minitab Support, “Methods and formulas for Analyze Factorial Design” (support.minitab.com).
This content is educational. The example data and results are illustrative. Follow your organization's quality system, change control, and safety requirements.
Step
Optimize
What settings should we use, does the prediction hold, and what do we do next?
This tab is being built. It will follow the same layout as the Frame tab: an overview, step-by-step guidance, the core tools, a checklist, common mistakes, and links to the site's calculators and templates. In the meantime, see the Design of Experiments guide and Analyzing Designed Experiments.