Choose a stage tab below. Each tab opens with an overview, then covers the details needed to do that stage well, with links to the calculators, templates, and guides on this site. Available now: Frame, Design, Run, Analyze.

The Five Stages at a Glance

A designed experiment changes several inputs on purpose, in a planned pattern, and measures the output, so that every run teaches something about every factor. The work divides into five stages. One example runs through all five tabs: an ultrasonic weld on a sensor housing, where 4.2% of housings leak at final test. All figures in the examples are illustrative.

Frame Question, response, factors, budget Design Design, runs, randomization Run Execute and record Analyze Effects, model, checks Optimize Settings and confirmation
The stages build on each other. Large problems usually need a sequence of small experiments, with the Optimize stage of one pointing to the Frame stage of the next.
StageKey questionMain outputsGate decision
F · FrameWhat do we need to learn, what will we measure, and which factors and ranges are worth testing?Objective, response, measurement check, baseline, factors and levels, budgetSponsor and process owner agree the question, the budget, and that the response can be measured
D · DesignWhich design answers the question within the budget, and how many runs and replicates do we need?A design, run count, replicates, center points, randomization, and blocking planThe design is approved and has enough power for the effect that matters
R · RunHow do we run the experiment exactly as planned and record what really happened?A randomized run sheet, completed runs, a log of settings, deviations, and noise factorsAll runs are complete, or deviations are documented and understood
A · AnalyzeWhich factors and interactions matter, and can we trust the model?Effects, ANOVA, plots, a checked and reduced modelThe model passes its residual checks and the effects make engineering sense
O · OptimizeWhat settings should we use, does the prediction hold, and what do we do next?Best settings, confirmation runs, next experiment or a handoff to ControlThe confirmed settings are adopted, or the next experiment is planned

How this toolbox fits with the rest of the site

ResourceUse it for
Design of Experiments guideThe concepts: factors, effects, interactions, and why DOE beats one-factor-at-a-time
DOE Quick PlannerA quick design, run count, and run matrix
Analyzing Designed Experiments (Stat Dojo)The statistics of the analysis, worked by hand, in Excel, and in Minitab
Minitab GuideThe software workflow for creating and analyzing a design
DMAIC ToolboxWhere an experiment sits in a project: usually the Improve phase

Step 1 of 5

Frame

What do we need to learn, what will we measure, and which factors and ranges are worth testing? Frame turns a vague wish to “try some settings” into a clear question, a response you can trust, a short list of factors with sensible ranges, and a budget. Most of what makes an experiment succeed is decided here, before a single run.

Key question
What do we need to learn, and what must be true before an experiment can answer it?
Typical duration
A few days to two weeks, mostly spent on measurement checks, baseline data, and talking to process experts
Led by
A project leader or Black Belt with the process engineer, operators, and a measurement owner
Starts from
A problem, a gap to a target, or suspected causes from the Analyze phase of a project
Primary outputs
An experiment frame: objective, response, measurement check, baseline, factors and levels, constraints, and approval
Hands off to
Design: a frame that tells the team which design they need and how many runs they can afford
Gate decision
The sponsor and process owner agree the question, the budget, and that the response can be measured well enough
Core tools
Fishbone, P-diagram, Gage R&R, histogram and control chart of baseline data, factor and level table, risk review

What the Frame Step Is For

A designed experiment is a question put to the process. The design, the runs, and the analysis can only answer the question that was asked, with the response that was measured, over the ranges that were tried. Frame is the step that makes sure all three are right before time, materials, and line time are spent.

What Frame must achieve

  • An objective written as a question and a decision the answer will drive
  • One primary response that is measured in numbers, with a direction and a unit
  • A measurement system that is good enough, checked with data
  • A baseline: where the response is now, how much it varies, and whether it is stable
  • A short list of factors, each with a reason, a type, and a low and a high level
  • Constraints and a budget in runs, material, and time
  • Approval from the people who own the process and the cost

What Frame must not do

  • Choose a design before the question is clear
  • Pick a pass/fail response when a measurement is available
  • Skip the measurement check because the gauge is “what we always use”
  • Test only factors the team already believes in
  • Set levels so close that nothing could change
  • Hide the cost, so the experiment is cancelled halfway
The test of a finished frame. Could an engineer who was not in the discussion read the one-page frame and know what you are trying to learn, what will be measured and how well, which settings will change and by how much, what will be held constant, and what the experiment is allowed to cost? If so, the frame is ready for the Design tab.

The Frame Step, Step by Step

Right tool Is DOE needed? Objective Question and decision Response Measure and check it Baseline Where are we now? Factors List, classify, choose Levels and budget Ranges, runs, cost
Frame moves from a problem to a priced, approved experiment frame. The steps overlap: a poor measurement check will send you back to the response, and a budget limit will send you back to the factor list.
  1. Confirm a designed experiment is the right tool

    Compare your situation with the cases in the table below. If the cause is known, fix it. If only one factor is in question, test it directly. If you cannot set the inputs on purpose, use observed data.

    • Check that you can control the factors you want to test
    • Check that the response can be measured
    • Check that the process is stable enough to learn from

    Output: A decision: run an experiment, or use another method

    Watch for: Reaching for a DOE because it looks rigorous, not because the question needs it

  2. State the objective as a question and a decision

    Write what you want to learn and what you will do with the answer. Choose the type of experiment: screen many factors, characterize a few, optimize, make the process robust, or confirm a result.

    • One sentence for the question
    • One sentence for the decision it will drive
    • Name the sponsor who will make that decision

    Output: An objective and an experiment type

    Watch for: An objective like “understand the process”, which no result can fulfil

  3. Choose the response and check how well you can measure it

    Pick one primary response, preferably continuous, with a unit, a direction (larger, smaller, on target), and an exact method of measurement. Run a measurement system study.

    • Define how, when, and by whom the response is measured
    • Check resolution against the process variation
    • Run a Gage R&R, or the nested version for a destructive test

    Output: A defined response and a measurement noise estimate

    Watch for: Using a pass/fail judgment, or a gauge that no one has ever checked

  4. Establish the baseline

    Collect or find data on the response under current settings. Plot it in time order and as a histogram. Estimate the mean, the variation from run to run, and whether the process is stable.

    • Use at least 30 recent observations
    • Look for trends, shifts, and mixtures
    • Record how the response relates to the specification

    Output: Baseline mean and standard deviation, and a stability check

    Watch for: Experimenting on an unstable process, where drift will look like an effect

  5. List the candidate factors

    Brainstorm every input that could plausibly move the response. Use a fishbone diagram, a P-diagram, the process map, past data, and the people who run the process.

    • Include operators, maintenance, and engineers
    • Look at what changed when the problem began
    • Record how each factor could be set, and how precisely

    Output: A long list of candidate factors

    Watch for: A list drawn only from the engineer’s favorite theories

  6. Classify the factors and choose which to test

    Sort factors into controllable factors you can set, noise factors you cannot control, and factors to hold constant. Choose the controllable factors to put in the experiment, and decide how to handle the rest.

    • Hold constant what you can
    • Block or record what you cannot hold
    • Prefer more factors at wide ranges over fewer at narrow ranges

    Output: The factors to vary, the factors to hold, and a plan for noise

    Watch for: Letting a noise factor, such as the material lot, change halfway through

  7. Set the levels

    Choose a low and a high level for each factor. Make them wide enough to move the response, and safe enough that every combination can be run. Keep the current setting between them if possible so center points can be added.

    • Ask the process experts for the widest safe range
    • Check the extreme combinations for safety and feasibility
    • Record exactly how each level will be set and verified

    Output: A table of factors, units, and levels

    Watch for: Levels so close together that a real effect disappears into the noise

  8. Set the constraints and the budget

    List how many runs you can afford in material, line time, test time, and money, and what the problem is costing. Check safety, quality, and regulatory approvals.

    • Count the runs and the units per run
    • Compare the cost with the value of the answer
    • Get approval for the line time and the scrap

    Output: A run and cost budget, and approvals

    Watch for: Finding out halfway that you cannot afford the runs you planned

  9. Write the experiment frame and get it signed off

    Put everything on one page. Review it with the sponsor, the process owner, and the measurement owner. Agree what a successful result would look like and what will be done with it.

    • One page, readable by someone who was not there
    • Sign-off before the Design step starts
    • Keep it with the project record

    Output: A one-page experiment frame, approved

    Watch for: Starting the design while the sponsor still has a different question in mind

Is a Designed Experiment the Right Tool?

A designed experiment is the strongest way to find cause and effect, but it costs runs, time, and disruption. Check the situation first.

SituationBetter approachWhy
The cause is known and the fix is obviousFix it, then verify with a PDCA cycleNo question left for an experiment to answer
One factor, two settingsA two-sample test or a PDCA cycle (see t-Tests)A designed experiment adds nothing for a single factor
Many suspected factors, and you can set themDesigned experimentFinds main effects and interactions with the fewest runs
Many suspected factors, but you can only observe the processMulti-vari study, or regression on historical data (see the Multi-Vari Studies entry)Observed data are confounded; the experiment needs control over the factors
The measurement system is poorFix the gauge first (see Measurement System Analysis)Noise from the gauge hides every effect
The process is unstableFind and remove the special causes firstDrift looks like an effect and ruins the comparison
Ingredients that must sum to 100%A mixture design (see the Design tab)Ordinary factorial designs do not fit proportions
Each run is very expensive or very slowA fractional or sequential design, or a smaller number of factorsMake every run count; use screening before optimizing
The factors cannot be randomized (a hard-to-change setting)A split-plot design, with expert helpStandard randomization would cost too many setup changes
DOE inside DMAIC. In a project, the experiment usually belongs in the Improve phase, after the Analyze phase has narrowed the suspects. See the DMAIC Toolbox. A DOE can also be used in Analyze to test a short list of suspected causes at once.

Writing the Objective

The objective has two parts: what you want to learn, and the decision it will drive. Without the second part, the experiment produces numbers but not action.

Weak objectiveBetter objective
Understand the welding process.Find which of six weld settings have a real effect on burst pressure, so the team can set them to cut leak failures from 4.2% to under 0.5%.
Optimize the process.Find settings that raise mean burst pressure by at least 20 kPa without increasing its variation, and confirm them on a second resin lot.
See if temperature matters.Decide whether to add temperature control to the oven by finding out whether a 10 °C change moves yield by more than 2 points.
TypeQuestion it answersTypical situationCommon designs
ScreeningWhich of many factors matter?Five or more suspects, little prior knowledgeFractional factorial, Plackett-Burman
CharacterizationHow do a few factors and their interactions affect the response?Two to five important factorsFull factorial, with replicates and center points
OptimizationWhat settings give the best response?Two or three key factors, curvature suspectedResponse surface (central composite, Box-Behnken)
RobustnessHow do I make the response insensitive to noise?Variation from material, environment, or useDesigns with noise factors (robust design, Taguchi)
ConfirmationDoes the predicted best setting really work?After any of the aboveA few runs at the chosen setting, with an interval

Large problems are usually solved with a sequence of small experiments: screen, then characterize, then optimize, then confirm. Plan to spend no more than about a quarter of the available runs on the first experiment, so something is left to follow up. The Design tab describes how to choose.

Choosing the Response

The response is the output you will measure in every run. It is the most important choice in the frame, and the one that most often goes wrong.

TypeExampleInformation per runUse
Continuous measurementBurst pressure (kPa), thickness, time, yield %HighAlways the first choice
CountDefects per unitMediumWhen a measurement does not exist; needs more runs
Ordinal ratingAppearance on a 1 to 5 scaleLow to mediumOnly with a defined scale and trained raters
Pass/failLeak or no leakVery lowLast resort; needs very many units per run

Why it matters, with numbers. The weld leak rate is 4.2%. To see the leak rate fall to 2% with 80% power would need about 975 housings per group. The burst pressure behind the leaks varies with a standard deviation of 22.4 kPa, so seeing a 20 kPa change takes only about 20 housings per group. The measurement is about 49 times cheaper to experiment on.

  • Choose a response that is close to the cause, and measured directly, such as burst pressure, not the later symptom, such as field returns.
  • Use one primary response. Add a few secondary responses you will watch (cost, cycle time, appearance), and decide how to trade them off.
  • Write the operational definition: the method, the instrument, the units, the number of readings, and who measures.
  • Give it a direction: larger is better, smaller is better, or on target between limits.
  • Measure the response in the same way for every run, and ideally with the same person and instrument.

Checking the Measurement System

Every run produces a number, and that number is the true value plus the measurement error. If the error is large compared with the effects you want to find, the experiment fails however well it is run. Check before you start.

CheckQuestionHowTarget
ResolutionCan the gauge show differences much smaller than the process variation?Compare the smallest reading step with the process standard deviationStep no more than one tenth of the process standard deviation
Repeatability and reproducibilityHow much of the observed variation is the gauge and the people?Gage R&R; use the nested method if the test destroys the part (MSA for destructive testing)Under 10% of study variation is good; 10% to 30% is marginal; plan to average repeats
Bias and stabilityIs the gauge accurate, and does it drift?Measure a reference or a check standard over daysNo trend, within the allowed error
CalibrationIs it calibrated and traceable?Check the recordIn date, with a record
Operator methodDoes everyone measure the same way?Write and train a measurement procedureA written method that all operators follow

What to carry forward. The Design tab needs the size of the measurement error to decide how many runs are needed and whether to average repeated readings. Record it now. If the measurement noise is a large share of the run-to-run variation, plan to average several readings per run.

A marginal gauge is not a reason to cancel, but it is a reason to plan for it: more repeats per run, a more careful procedure, or a better instrument for the experiment even if it is not the routine one.

Establishing the Baseline

The baseline tells you where the response is today, how much it varies when nothing is changed on purpose, and whether the process is stable enough to experiment on. It also supplies the noise estimate that sets how many runs you need.

  • Gather data from the current process over a representative period: at least 30 observations, from different shifts and material lots if possible.
  • Plot in time order (run chart or individuals chart) and look for shifts and drift. A stable process has no special causes.
  • Plot a histogram and compare with the specification. Look for two humps, which suggest a mixture of sources.
  • Estimate the standard deviation. This is the run-to-run noise that effects must stand out from.
  • Decide the smallest effect that is worth finding. Effects smaller than this do not change the business result, and they are costly to chase.

See Descriptive Statistics, Control Chart Theory, and Graphical Analysis for the methods.

Listing the Candidate Factors

Cast the net wide at first. The factor you did not think of is often the one that matters. Ask the people who run the process, look at what changed when the problem started, and use more than one method.

Low burst pressure and leaks Machine Weld amplitude Weld force Horn wear Method Weld time Hold time Clamp pressure Material Resin lot Moisture Part flash Measurement Tester rate Fixture alignment Calibration Environment Room humidity Ambient temperature Part temperature People Operator loading Setup practice Shift
A fishbone diagram for the weld example. Each bone is a category of inputs; the other bones carry more causes than fit on the figure.
MethodWhat it contributesTips
Fishbone diagramOrganizes ideas by category so none is missedCover machine, method, material, measurement, environment, and people
P-diagram (parameter diagram)Separates controllable inputs, noise, and the responseDraw it with the team; see the figure below
Process mapShows every step where a factor could actMap the real process, with the people who run it
Historical data and past experimentsShows which factors have mattered beforeCheck maintenance logs, change records, and earlier studies
Engineering knowledgePhysical reasoning about what should matterWrite down the reason for each factor
Multi-vari studyFinds where the variation lives (shift, position, lot)Use it before an experiment to narrow the list
BrainstormingMany ideas quicklyCollect first, judge later

Classifying and Choosing the Factors

Not every factor belongs in the experiment. Sort the long list into three groups, then choose.

The processY = f(X1, X2, ...) Weld amplitude Weld force Weld time Hold time Resin drying time Clamp pressure Resin lot Horn wear Room humidity Loading method Burst pressure (kPa) Controllable (X) Noise Procedural
A parameter diagram. Blue circles are factors to vary, gold are noise, and gray are procedures to hold the same.
GroupMeaningWhat to do
Controllable factorsSettings you can choose and holdCandidates to vary in the experiment
Noise factorsInputs you cannot or do not want to control: material lot, ambient conditions, wearHold constant for the experiment, block on them, record them, or include them deliberately in a robustness study
Held constantEverything else that mattersFix at one level and write down how
FactorGroupDecisionReason
Weld amplitudeControllableVaryStrong suspected effect; easy to set
Weld forceControllableVarySuspected interaction with weld time
Weld timeControllableVaryDirectly controls energy delivered
Hold timeControllableVarySuspected effect on joint cooling
Resin drying timeControllableVaryMoisture suspected; this is a new idea the team has not tested
Clamp pressureControllableVaryFixture alignment suspected
Resin lotNoiseHold: one lot for the experimentConfirm on a second lot afterwards
Horn wearNoiseHold: new horn, record the cycle countWear changes the energy over time
Room humidityNoiseRecord every runCovariate for the analysis
Loading methodProcedureHold: one trained operatorSame method every run
Burst tester and operatorMeasurementHold: same tester, same personReduce measurement variation
  • Hold constant what you can, and write it down: lot, operator, equipment, and method.
  • Block what you cannot hold, for example two resin lots, with each lot run as a block (see the Design tab).
  • Record what you cannot control, such as humidity, so you can check afterwards whether it mattered.
  • Include a factor if you have a reason to suspect it, even if you doubt it. A surprise factor is the best result an experiment can give.
  • Do not drop a factor only because it is expensive to set. Raise the cost with the sponsor.

Setting the Levels

The low and high levels define the region the experiment explores. They determine whether an effect shows above the noise, and whether any run can be made at all.

Low (-1)Current (center)High (+1) Weld amplitude 70 %80 %90 % Weld force 400 N500 N600 N Weld time 0.3 s0.4 s0.5 s Hold time 0.5 s1 s1.5 s Resin drying time 2 h4 h6 h Clamp pressure 300 kPa375 kPa450 kPa
Levels centered on the current setting, so that center points can be added later to check for curvature.
RuleWhy
Bold, but safe. Wide enough that the effect should exceed the noise, narrow enough that every combination makes usable productNarrow ranges are the main reason real effects are missed; unsafe ranges ruin the experiment
Ask the experts for the extremes they would run on purposeThey know where the process falls over
Check the corners. Every combination of lows and highs must be possible and safeA 2k design runs every corner, including the odd ones
Center on the current setting when possibleAllows center points to check for curvature
Use real, settable values, with a way to verify them“High” must mean the same on every run
Quantitative factors (amplitude) use two levels first; categorical factors (supplier, machine) use the levels that existTwo levels suit screening; more levels add runs quickly
FactorUnitLow (−1)CurrentHigh (+1)Basis for the range
Weld amplitude%708090Supplier’s range for this horn is 60 to 100; 70 and 90 stay clear of the edges
Weld forceN400500600Below 400 N the joint does not seat; above 600 N the part flash increases
Weld times0.300.400.50Process engineer’s widest settings that give acceptable appearance
Hold times0.51.01.5Cycle time limit allows 1.5 s
Resin drying timeh246The dryer cycle runs from 2 to 6 hours in practice
Clamp pressurekPa300375450Fixture rating is 500; keep a margin

Constraints and Budget

Every experiment lives within limits. Write them down before the design, because they decide how many runs you can have. Then compare the cost with what the answer is worth.

ConstraintQuestionExample
MaterialHow many units can be used or scrapped?240 housings and the matching parts are available
Test timeHow long does each run and each measurement take?3 minutes per housing for the burst test
Line timeHow much production time can be given up?2 line days
CapacityHow many housings per day can be welded and tested?About 32 per day, so up to 64 housings in all
MoneyMaterial, labor, and downtime$2,432 of housings, plus 40 hours of engineering at $75
Safety and approvalDoes the range affect safety, regulation, or the customer?Quality and safety owners approve the six ranges; no shipment of experimental units
RandomizationCan the run order be randomized?Yes: the settings can be changed between welds in about 10 minutes

The value of the answer. 130 leak failures a month at $38 each cost about $59,280 a year. Cutting the rate from 4.2% to 0.5% would save about $52,212 a year. The experiment would cost about $5,432, which the saving would repay in roughly 5 weeks.

Plan the runs you can afford, not the runs you wish for. The budget goes to the Design tab, where it sets the number of runs and the number of units per run. If the budget is too small to detect the effect you care about, say so now, and choose between more budget, fewer factors, or a smaller goal.

Worked Example: The Weld Experiment Frame

A plant makes a plastic sensor housing that is joined by ultrasonic welding. Each month about 3,100 housings are tested for leaks at the end of the line, and 130 fail (4.2%), costing $38 each. The team decided to run an experiment on the weld settings. All figures are illustrative, and the same example runs through all five tabs.

0 4 8 12 16 250 270 290 310 330 350 370 Lower limit 265 Mean 305 Burst pressure (kPa) Housings
Burst pressure of 60 recent housings. 3 fell below the lower limit of 265 kPa (shown in red).
Frame elementWhat the team wrote
ObjectiveFind which of six weld settings affect burst pressure, and the settings that raise it by at least 20 kPa without raising its variation, so that leak failures fall from 4.2% to under 0.5%. The sponsor will decide whether to change the work instruction and the machine program.
Experiment typeScreening first, then a smaller characterization and confirmation (a sequence of experiments)
ResponseBurst pressure in kPa, measured by destructive test, larger is better, lower specification limit 265 kPa
Measurement checkNested Gage R&R on the burst tester: measurement standard deviation 4.0 kPa, which is 18% of the total standard deviation (3.2% of the variance): marginal but acceptable for screening if readings are repeated. Resolution 1 kPa against a standard deviation of 22 kPa: adequate
Baseline60 housings: mean 305 kPa, standard deviation 22.4 kPa, range 261 to 367, 3 below the limit. The individuals chart shows no points beyond the limits (Shapiro-Wilk p = 0.59), so the process is stable and the data look normal. Estimated fraction below the limit 3.6%, consistent with the 4.2% leak rate
GoalRaise the mean to about 322 kPa at the same variation, which puts the lower limit 2.5 standard deviations away and cuts the expected failure rate to about 0.5%
Smallest effect worth finding20 kPa, about one standard deviation of weld-to-weld variation
Factors to varyAmplitude, force, weld time, hold time, resin drying time, clamp pressure, at the levels in the table above
Held constantOne resin lot, new horn (cycle count recorded), same operator and burst tester, same loading method
Recorded, not controlledRoom humidity and temperature
BudgetUp to 64 housings over 2 line days, cost about $5,432, against a saving of about $52,212 a year (5-week payback)
RisksOut-of-range welds could damage the horn: test the extreme corners first with two housings; no experimental units ship
Sign-offSponsor (plant manager), process owner (manufacturing engineering), quality, and the test lab lead

What the frame tells Design. Six controllable factors, a noise level of 22 kPa per housing, a smallest effect of 20 kPa, and a budget of 64 housings. One housing per run would leave the effect hard to see, so averaging several housings per run will be considered in Design.

On the schedule. The team spent two days on the measurement study and baseline, a day with the operators and the process engineer on factors and ranges, and a half day on the one-page frame and sign-off. About three and a half days of preparation, against a line time of two days.

Frame Check: Are We Ready to Design?

Before starting the Design tab, confirm the frame is complete. Use this checklist; progress saves in this browser only, and nothing is sent anywhere.

0 of 15 complete
Objective and response
Measurement and baseline
Factors and levels
Budget and approval

Questions a sponsor or a coach can ask:

  • What do you want to learn, and what will you do with the answer?
  • How do you know the response can be measured well?
  • What does the process do today, and is it stable?
  • Which factors did you consider, and why did you keep these?
  • How wide are the levels, and who agreed they are safe?
  • What will it cost, and what is the answer worth?
OutcomeMeaningNext step
ReadyThe frame is complete and approvedStart the Design tab
Ready with conditionsMinor gaps, for example one range still to be confirmedClose them before the design is final
Not readyThe measurement is poor, the process is unstable, or the question is unclearFix those first
Wrong toolThe cause is known, the factors cannot be set, or one factor is in questionUse a PDCA cycle, a multi-vari study, or a direct test instead

Adapting Frame to the Situation

SituationHow Frame changes
Manufacturing processControllable settings are plentiful; the usual danger is noise from material lots and equipment wear. Block on lots and record equipment state
Chemical and process industriesResponses are often continuous and run times long; check stability and the cost of a failed run; consider mixtures
Transactional and service workFactors may be scripts, staffing, or sequence; responses are times or error rates; randomize across customers or days; get approval for customer-facing changes
Software and online servicesControlled online experiments (A/B and multivariate tests) need a clear metric, enough traffic, and protection against interference between tests; see the Software and IT hub
HealthcareEthics and patient safety come first; use protocols and approvals; many questions call for small tests, not factorial designs; see the healthcare hub
Food, agriculture, and chemistry labsBiological variation and batch effects are large; block on batch, replicate generously
Where the factors cannot be setDo not force a DOE. Use a multi-vari study or regression on observed data, and use the findings to plan a later experiment

Scale the framing to the risk and cost. A cheap, fast, reversible experiment needs a half-page frame. An experiment that stops a line, uses scarce material, or affects customers needs the full review.

Common Mistakes and Red Flags

MistakeWhat it looks likeHow to correct it
Choosing the design first“Let’s run a Box-Behnken” with no objective writtenState the question, then choose
Pass/fail responseCounting leaks, with 12 runs of 20 unitsFind the measurement behind the pass/fail
Unchecked gaugeNo one knows the measurement errorRun a Gage R&R before the experiment
Unstable baselineThe response drifts or has shifts before any changeStabilize the process first
Factor list from one personOnly the engineer’s theories appearInclude operators, maintenance, and the lab
Timid levelsLow and high differ by a tiny amountAsk for the widest safe range
Noise left to chanceTwo material lots change mid-experimentHold, block, or record noise factors
No budgetThe experiment stops when the material runs outCount runs and units before the design
No decision attachedResults are reported but nothing changesName the decision and the decider
Too many questions in one experimentFive responses, none rankedChoose a primary response and rank the rest

Frame Resources on This Site

Guides

Tools

Stat Dojo

Body of Knowledge

Frame Step Frequently Asked Questions

When is a designed experiment the right tool, and when is it overkill?

Use a designed experiment when several inputs might affect an output, you can set those inputs on purpose, and you need to know which matter, by how much, and whether they interact. It is overkill when the cause is already known and the fix is obvious, when only one factor is in question (a two-sample test or a PDCA cycle is enough), or when the inputs cannot be controlled (use a multi-vari study or regression on observed data instead).

Why spend so much time framing before choosing a design?

Because the design only answers the question you asked. A perfectly executed experiment on the wrong response, a noisy gauge, or factor ranges that are too narrow gives a clean, useless answer. Most failed experiments fail in the framing: the response could not be measured well, an important factor was left out, or the levels were too timid to move the output.

Should the response be continuous or pass/fail?

Continuous whenever you can get it. A measurement carries far more information than a pass/fail judgment of the same part, so it needs far fewer runs to see the same change. In the worked example, detecting a drop in leak rate from 4.2% to 2% would take about 975 housings per group, while detecting a 20 kPa change in burst pressure takes about 20.

How many factors should I include in the first experiment?

As many as you have good reason to suspect, up to what your budget can screen. Five to eight factors is typical for a screening design; fewer than four can usually go straight to a full factorial. Leaving out an important factor wastes the experiment, while including a few extra costs little in a fractional design.

How wide should the factor levels be?

Wide enough that the effect, if there is one, will exceed the noise, and narrow enough that every combination is safe and produces usable output. Levels set too close together are the most common reason a real effect is missed. Ask the process experts: what is the lowest and highest setting you would run on purpose without scrapping the lot?

What if I cannot measure the response well?

Fix the measurement first. If measurement variation is a large share of what you see, effects will be buried. Run a measurement system study before the experiment. If the gauge is only marginal, plan to take repeated readings per run and average them, which the Design tab covers.

Sources and Further Reading

  • Douglas C. Montgomery, Design and Analysis of Experiments, Wiley, on guidelines for designing experiments.
  • George E. P. Box, J. Stuart Hunter, and William G. Hunter, Statistics for Experimenters, Wiley.
  • Mark J. Anderson and Patrick J. Whitcomb, DOE Simplified, Productivity Press.
  • NIST/SEMATECH, e-Handbook of Statistical Methods, Process Improvement: Design of Experiments (itl.nist.gov/div898/handbook).
  • Automotive Industry Action Group, Measurement Systems Analysis Reference Manual, 4th ed.
  • Minitab Support, documentation for Create Factorial Design and Gage R&R (support.minitab.com).

This content is educational. The example data and results are illustrative. Follow your organization's quality system, change control, and safety requirements.

Step 2 of 5

Design

Which design answers the question within the budget, and how many runs, replicates, and housings per run do we need? Design turns the approved frame into a plan: which runs to make, in what order, how many units to test at each setting, and how well the plan will detect the effect that matters. Check the plan on paper before spending a single housing.

Key question
Which design finds the effect that matters, within the budget, with the smallest risk of a misleading result?
Typical duration
A few days: design options, power check, review with the sponsor
Led by
The project leader or Black Belt, with the process engineer and, if available, a statistician
Starts from
The approved frame: objective, response, noise, smallest effect, factors and levels, budget
Primary outputs
A design with its run count, replicates, blocks, randomization plan, aliasing, and a power check
Hands off to
Run: a design that is ready to turn into a randomized run sheet
Gate decision
The sponsor and process owner accept the design and what it can and cannot tell them
Core tools
Design finder, resolution and alias table, power calculation, DOE Quick Planner, Minitab Create Factorial Design

What the Design Step Is For

The frame says what to learn. The design says how to learn it with the fewest runs that still give a trustworthy answer. A good design makes every run count toward every effect, protects the result from drift and from things you did not control, and says in advance which effects it can and cannot separate.

What Design must achieve

  • A design that fits the objective and the number of factors
  • A run count and a number of units per run that the budget allows
  • A power check: would an effect of the size that matters be found?
  • Known aliasing: which effects are mixed up with which
  • A plan for replication, blocking, and center points
  • A randomization plan, and a plan for hard-to-change factors
  • A written summary the sponsor can approve

What Design must not do

  • Choose the design that matches last time’s experiment, not this question
  • Skip the power check and hope the runs will be enough
  • Mistake several units at one setup for replicates
  • Fix the number of runs by what feels convenient
  • Run in standard order, so drift is mixed with the factors
  • Hide the aliasing from the people who will read the results
The test of a finished design. Could someone read the design summary and tell you how many setups there will be, which effects can be estimated cleanly, which are mixed up with others, how big an effect the experiment can find, what it will cost, and what the plan is for noise and drift? If so, the design is ready for the Run tab.

The Design Step, Step by Step

Family Match to the objective Runs Full or fraction, resolution Units per run Replicates versus repeats Power Can it find the effect? Protect Blocks, center, random order Summary Review and approve
Design moves from the objective to an approved plan. The power check often sends you back to change the number of runs or the units per run.
  1. Choose the design family

    Match the design to the objective and the number of factors: screening, characterization, optimization, robustness, or a mixture.

    • Use the design finder below
    • Plan a sequence if needed: screen, then characterize, then optimize
    • Check that every combination of the chosen levels is safe to run

    Output: A design family

    Watch for: Using a response-surface design to screen, or a screening design to optimize

  2. Choose the number of runs and the resolution

    List the full factorial and the fractions that fit the budget. Choose the resolution that keeps the effects you care about clear of each other.

    • Prefer resolution V or higher if interactions matter
    • Accept resolution IV for pure screening, with a plan for the follow-up
    • Avoid resolution III unless you must

    Output: A design with a known alias structure

    Watch for: A fraction so small that main effects are mixed up with the interactions you suspect

  3. Decide the setups and the units per setup

    Separate two numbers: how many setups (runs) and how many units to make at each. Use the noise components to see what each choice buys.

    • Setup-to-setup variation is not reduced by making more units at one setup
    • More setups usually beat more units per setup
    • Count the setup time

    Output: Setups, units per setup, and total units

    Watch for: Treating several units from one setup as replicates

  4. Check the power

    Use the run-level standard deviation and the number of runs to find the probability of detecting the smallest effect that matters. Compare the options.

    • Use the smallest effect worth finding from the frame
    • Aim for at least 80% power
    • If power is low, change runs, units, or the goal, not the arithmetic

    Output: A power figure for each option, and a choice

    Watch for: Running the experiment hoping the effect is large

  5. Add blocks, center points, and a random order

    Block on anything you know will change during the experiment, such as the day. Add center points if curvature is a concern and the budget allows. Randomize the order within blocks.

    • Confound blocks with high-order interactions
    • Record the center points as a separate type of run
    • Plan for any hard-to-change factor

    Output: Blocks, center points, and a randomization plan

    Watch for: Running one block, then the other, in standard order

  6. Summarize and get approval

    Write the design on one page: factors and levels, runs, units, blocks, aliasing, power, cost, and time. Review it with the sponsor and process owner and note what the experiment cannot tell them.

    • State the aliasing in plain language
    • State the smallest effect it can find
    • Agree what the next experiment will be if it finds something

    Output: A one-page design summary, approved

    Watch for: Hiding a limitation that will surface when the results arrive

Finding the Design Family

What is the objective? Screening Five or more factors, little prior knowledge Fractional factorial or Plackett-Burman Characterization Two to five important factors, interactions matter Full factorial or half fraction, with replicates Optimization Two or three key factors, curvature likely Response surface: central composite or Box-Behnken Robustness Variation from noise you cannot control Control factors crossed with noise factors Recipes Ingredients that sum to 100% Mixture design
Start from the objective written in the frame. Large problems usually pass through the families from left to right.
FamilyWhenRuns (typical)Limits
Screening (fractional factorial, Plackett-Burman)Five or more candidate factors; find the few that matter8 to 32Aliasing; no curvature unless center points are added
Characterization (full factorial or high-resolution fraction)Two to five important factors and their interactions8 to 32, plus replicatesTwo levels cannot show curvature
Optimization (response surface)Two or three key factors where the best setting is inside the region13 to 30Needs the important factors already known
Robustness (control factors crossed with noise)Make the output insensitive to material, environment, or use16 to 64Needs a way to vary the noise on purpose
MixtureRecipes where the ingredients sum to 100%10 to 30Constraints on the ingredient ranges

The rest of this tab concentrates on the two-level factorial family, which does most of the work in screening and characterization. See Response Surface Methodology, Taguchi Methods, and Robust Design for the others.

Two-Level Factorial Designs

A two-level factorial sets each factor at a low and a high level, written −1 and +1, and runs every combination. It estimates every main effect and every interaction. Each run contributes to every estimate, which is why it needs far fewer runs than changing one factor at a time.

Factors (k)Runs in a full factorial (2k)Number of effects to estimate
242 main + 1 two-factor + higher order
383 main + 3 two-factor + higher order
4164 main + 6 two-factor + higher order
5325 main + 10 two-factor + higher order
6646 main + 15 two-factor + higher order
71287 main + 21 two-factor + higher order

The number of runs doubles with every factor, so beyond four or five factors a full factorial is rarely worth it: most of the runs only estimate three-way and higher interactions, which are usually negligible.

  • Orthogonal and balanced. Each factor is at each level in half the runs, and all effects can be estimated independently.
  • Coded levels (−1 and +1) put every factor on the same scale, and make the calculation of effects simple. See Analyzing Designed Experiments.
  • Interactions come free. A significant interaction means the effect of one factor depends on the level of another.

Fractional Factorials, Aliasing, and Resolution

A fractional factorial runs a selected part of the full factorial: half, a quarter, an eighth. The price is aliasing: some effects can no longer be told apart, because the same pattern of −1 and +1 estimates both. The figure shows the idea with three factors.

(−,−,−) (−,−,+) (−,+,−) (−,+,+) (+,−,−) (+,−,+) (+,+,−) (+,+,+) A: low to high B: low to high C: low to high Run in the half fraction (4 runs) Not run (the other 4 corners) Rule: C = A × B Four runs, three factors. C is mixed up (aliased) with the A × B interaction.
Running only the highlighted corners (a half fraction) halves the runs. The third factor is set by multiplying the first two, so its effect cannot be separated from the A × B interaction.

The design is described by its generators (here C = AB) and its defining relation (I = ABC). Every effect is aliased with the effect obtained by multiplying it by the defining relation. The resolution is the length of the shortest word in the defining relation, and it tells you what is mixed up with what:

ResolutionWhat is aliasedUse it for
IIIMain effects are aliased with two-factor interactionsScreening many factors when interactions are assumed small
IVMain effects are clear of two-factor interactions; two-factor interactions are aliased with each otherScreening, when you want trustworthy main effects
VMain effects and two-factor interactions are clear of each other (aliased with three-factor interactions or higher)Characterization when interactions matter
VI and higherTwo-factor interactions are aliased only with four-factor interactions or higherAs good as a full factorial for practical purposes
FactorsFull factorial runsHalf fractionQuarter fractionEighth fraction
3823−1 III (4 runs)
41624−1 IV (8 runs)
53225−1 V (16 runs)25−2 III (8 runs)
66426−1 VI (32 runs)26−2 IV (16 runs)26−3 III (8 runs)
712827−1 VII (64 runs)27−2 IV (32 runs); 27−3 IV (16 runs)27−4 III (8 runs)
825628−2 V (64 runs)28−3 IV (32 runs); 28−4 IV (16 runs)
Assumption behind every fraction. Higher-order interactions (three-way and above) are small. That is usually true, and it is the reason fractions work. If it is not true for your process, the aliased effects will mislead you, so confirm important findings with a follow-up run.

Runs, Replicates, and Units per Run

Two different numbers decide how much information an experiment buys: the number of setups (runs, each with the factors freshly set), and the number of units made or measured at each setup. They are not interchangeable.

TermMeaningReduces which noise?
Setup (run)One complete setting of all the factors, made fresh—
ReplicateA complete repeat of a setup: reset, run, and measure againSetup-to-setup and within-setup noise
Repeat / subsampleSeveral units made or measured at one settingWithin-setup noise only
Measurement repeatThe same unit measured more than onceMeasurement noise only

In the weld example, a short run of 12 consecutive housings at one setting gave a standard deviation of 18 kPa. The overall weld-to-weld standard deviation is 22.4 kPa, which means setup-to-setup variation accounts for the rest: √(22.4² − 18²) = 13.4 kPa. The standard deviation of a run that averages m housings is √(13.4² + 18²/m).

0 5 10 15 20 25 Floor: setup-to-setup variation 13.4 22.41 18.52 16.93 16.14 15.65 15.26 15.07 14.88 Housings welded and averaged at each setup Standard deviation of a run (kPa)
Averaging more housings helps less and less. Beyond about three, the setup-to-setup floor dominates.
The practical rule. Averaging several units at one setup cannot reduce setup-to-setup noise. More setups can. When setup changes are cheap enough, use more runs with fewer units each. When they are expensive, accept more units per run, but compute the run-level noise honestly, and do not use the units as replicates in the analysis.

How Many Runs: Checking the Power

Power is the probability that the experiment finds an effect of a given size when it is really there. It depends on the effect size, the noise of a run, the number of runs, and the error degrees of freedom. For a two-level design:

Standard error of an effect = 2 σrun / √N   |   t = effect / standard error   |   power = P(|t| exceeds the critical t, given the true effect)
  1. Smallest effect worth finding (from the frame): 20 kPa.
  2. Run-level standard deviation: 18.5 kPa when two housings are averaged at each setup.
  3. Number of runs N: 32 setups.
  4. Standard error = 2 × 18.5 / √32 = 6.52 kPa.
  5. Error degrees of freedom: 9 (from the three-factor interactions, which are assumed negligible).
  6. Power for a 20 kPa effect = 78%. The smallest effect found with 80% power is about 20.5 kPa.

In Minitab use Stat > Power and Sample Size > 2-Level Factorial Design, which gives the same answer for the number of corner points, replicates, and center points. See Sample Size and Power.

  • Low power is not a small problem. A real, useful effect will often be missed, and the team will wrongly conclude that the factor does not matter.
  • If power is low, add runs (or replicates), reduce noise, accept a larger smallest-effect, or drop a factor. Do not just run it and see.
  • The error estimate must be real. It comes from replicates, from center points, or from higher-order interactions assumed to be zero. A design with none of these cannot test its own effects, and relies on a normal plot of effects.

Comparing the Options

With a budget of 64 housings and 960 line minutes (2 days), the team compared five ways to run six two-level factors. A setup change takes about 10 minutes and each housing about 3 minutes to weld and test.

DesignSetupsHousings per setupTotal housingsRun SD (kPa)SE of an effectPower for 20 kPaLine minutes
A2^(6-2) IV1646416.18.1n/a352 of 960
B2^(6-2) IV3226418.56.582%512 of 960
C2^(6-1) VI3226418.56.578%512 of 960
D2^66416422.45.694%832 of 960
E2^(6-1) VI3213222.47.961%416 of 960
B: 2^(6-2) IV, 32 setups, 2 per setup 82% C: 2^(6-1) VI, 32 setups, 2 per setup 78% D: 2^6, 64 setups, 1 per setup 94% E: 2^(6-1) VI, 32 setups, 1 per setup 61% 80% Power to detect a 20 kPa effect (reference line at 80%)
The same 64 housings can be spread over 16, 32, or 64 setups. More setups gain power, up to the limit of the line time.
OptionStrengthsWeaknesses
A: 16 runs, 4 housings eachCheapest in time; main effects clear of two-factor interactionsNo error estimate at all, so no test of significance; the 15 two-factor interactions are aliased in 7 groups
B: 16 runs replicated, 2 eachA true error estimate from replicates; blocks by replicateTwo-factor interactions are still aliased with each other
C: half fraction, 32 runsAll six main effects and all 15 two-factor interactions can be estimated clearly; blocks into two daysNo replicates; the error comes from assuming three-factor interactions are negligible; power just under 80%
D: full factorial, 64 runsHighest power; every effect estimatedUses 64 housings and about 87% of the line time; no slack for a failed run, and nothing left for follow-up
E: half fraction, 1 housing eachHalf the housingsPower only about 60%; likely to miss real effects
Choice. The team chose Option C: the half fraction with 32 setups and two housings at each. Interactions between weld force and weld time are suspected, so a resolution VI design that keeps all two-factor interactions clear is worth more than the extra margin of Option B. Option D has more power but uses all the housings and most of the line time. Power for a 20 kPa effect is 78%, and effects of about 21 kPa or more would be found with 80% power. That is just at the smallest effect the sponsor asked for, which is stated openly in the design summary.

Blocks, Center Points, and Randomization

Three more decisions protect the experiment from things the factors cannot explain.

DeviceWhat it doesHow to use it
BlockingSeparates a known source of variation (day, batch, lot, operator) from the factor effectsDivide the runs into blocks, each containing a balanced part of the design. The block effect is confounded with a high-order interaction you are willing to give up
Center pointsRuns at the middle of every factor range: they estimate pure error and test for curvatureAdd three to five; they cost little. A significant curvature test means a response-surface experiment is the next step
RandomizationSpreads uncontrolled drift evenly across all the settingsRandomize the run order within each block. Record the actual order. For a hard-to-change factor, talk to a statistician about a split-plot design
ReplicationGives a true error estimate and more powerReset all settings between replicates. Never treat several units from one setup as replicates

In the example, the 2 line days are the two blocks. Dividing the 32 runs by the sign of the A × B × C column puts 16 runs in each day, with the A × B × C interaction (which is aliased with D × E × F) confounded with the day-to-day difference. The team accepts that: three-factor interactions are expected to be small, and the day effect is then removed from every main effect and two-factor interaction.

Do not run the experiment in standard order. Standard order changes one factor slowly and others fast, so any drift lines up with the slowly changing factor. Randomize within blocks.

Other Designs You Will Meet

DesignUse it whenNotes
Plackett-BurmanScreening up to 11 factors in 12 runs (or 19 in 20, 23 in 24)Resolution III; use only when interactions are believed small
Definitive screening designsScreening continuous factors when curvature is possible, with software supportThree levels per factor; main effects are clear of two-factor interactions; needs software
Central compositeOptimizing two to five continuous factors, with curvatureAdds axial points to a factorial; see Response Surface Methodology
Box-BehnkenOptimizing three or more factors without running the cornersAvoids extreme combinations
Mixture designsRecipes where the proportions sum to 100%Factors are not independent
Split-plot designsSome factors are hard to change, so full randomization is impracticalNeeds a different analysis; ask for help
Taguchi and robust designsMake the output insensitive to noiseControl factors crossed with noise factors; see Taguchi Methods
Optimal designsIrregular regions, unusual run counts, or constraintsComputer-generated; understand what the software optimized

Worked Example: The Weld Design

The weld team took the approved frame to the design step. Their inputs: six factors at two levels, a noise level of 22.4 kPa per housing (of which 18 kPa is within a setup), a smallest effect of 20 kPa, a budget of 64 housings, and 960 line minutes. They compared the five options above and chose a half fraction. All figures are illustrative.

Design elementWhat the team wrote
Design26−1 fractional factorial, resolution VI, 32 setups; generator F = ABCDE, defining relation I = ABCDEF
Factors (coded −1 / +1)A amplitude 70 / 90 %; B force 400 / 600 N; C weld time 0.3 / 0.5 s; D hold time 0.5 / 1.5 s; E drying time 2 / 6 h; F clamp pressure 300 / 450 kPa
Units per setup2 housings welded at each setup; the run response is the average of the two burst pressures (run standard deviation about 18.5 kPa)
Total housings64 (the approved budget)
BlocksTwo days of 16 setups each, blocked on A × B × C
Replicates and center pointsNone in this experiment. Four center setups (+4 housings, about $152) are requested as an option and would add a curvature check
AliasingMain effects are aliased with five-factor interactions, and two-factor interactions with four-factor interactions: all 6 main effects and 15 two-factor interactions are estimable and clear of each other
Power78% for a 20 kPa effect; effects of 21 kPa or more are found with 80% power
Time32 setups × 10 minutes + 64 housings × 3 minutes = 512 minutes, against 960 available, leaving time to repeat a bad setup
RandomizationRandom order within each day. Settings are fully reset between setups
Held constantAs in the frame: one resin lot, a new horn with the cycle count recorded, one operator, the same burst tester. Humidity is recorded
Limits the sponsor acceptedNo replicates, so the error depends on three-factor interactions being small; power is just below 80%; no curvature test unless the four center setups are approved

The first eight runs of the design in standard order (before randomization), showing how factor F is generated from A to E, and which day each run belongs to:

Std runABCDEF = ABCDEDay (block)
1-1-1-1-1-1-11
2+1-1-1-1-1+12
3-1+1-1-1-1+12
4+1+1-1-1-1-11
5-1-1+1-1-1+12
6+1-1+1-1-1-11
7-1+1+1-1-1-11
8+1+1+1-1-1+12

Some of the aliasing, to show what the half fraction does and does not cost:

EffectAliased with
A (Weld amplitude)BCDEF
B (Weld force)ACDEF
C (Weld time)ABDEF
D (Hold time)ABCEF
E (Resin drying time)ABCDF
F (Clamp pressure)ABCDE
A × B (Weld amplitude × Weld force)CDEF
A × C (Weld amplitude × Weld time)BDEF
B × C (Weld force × Weld time)ADEF
D × E (Hold time × Resin drying time)ABCF
C × F (Weld time × Clamp pressure)ABDE
On the schedule. Two days to compare the options and the power, a day to build the design in Minitab and check it, and a half day to review with the sponsor and process owner. The 32 setups will take 8.5 hours of line time in the Run step.

Design Check: Are We Ready to Run?

Before starting the Run tab, confirm the design is complete. Use this checklist; progress saves in this browser only, and nothing is sent anywhere.

0 of 13 complete
Design choice
Runs, units, and power
Protection and approval

Questions a sponsor or a coach can ask:

  • Why this design and not a smaller or a larger one?
  • How many setups, and how many units at each?
  • What is the smallest effect it would find, and how sure are we?
  • Which effects are mixed up with which?
  • What happens if a run goes wrong?
  • What will we do if it finds something, and if it finds nothing?
OutcomeMeaningNext step
ReadyThe design is complete, has adequate power, and is approvedStart the Run tab
Ready with conditionsA minor gap, such as a missing center-point approvalClose it before the first run
Not readyPower is too low, or the aliasing hides the effects you care aboutChange runs, units, or the design
Wrong sizeThe budget cannot reach the smallest effect, even with the best optionTake the choices to the sponsor: more budget, fewer factors, or a larger smallest effect

Adapting Design to the Situation

SituationHow Design changes
Cheap, fast runs (a few minutes each)Use larger designs and full replicates; the cost of extra runs is small compared with the risk of an inconclusive result
Expensive or destructive runsUse fractions, plan the power carefully, run sequentially, and consider smaller experiments in a sequence
Hard-to-change factors (furnace temperature, a tool change)Use a split-plot design, with the hard-to-change factor in the whole plots; ask for statistical help
Many factors, very few runsPlackett-Burman or a definitive screening design, then confirm; accept aliasing
Curvature expectedInclude center points in a screening design, then move to a response surface design
Transactional and service workRuns are often cheap; randomize across customers, days, or agents; protect against customers who see more than one condition
Software and online servicesControlled online experiments allow many runs and large samples; the design questions are about interference and the metrics
Healthcare and other regulated settingsEthics and approvals decide what can be randomized; use protocols and review; many questions suit small sequential tests

Use software for the arithmetic, not for the thinking. The DOE Quick Planner, Minitab, and other packages generate designs and aliasing tables in seconds. The choices described in this tab, such as the budget, the noise, and the smallest effect, are yours.

Common Mistakes and Red Flags

MistakeWhat it looks likeHow to correct it
Several units counted as replicatesError looks tiny and everything is significantReplicate by resetting the setup; model the setup noise
No power checkThe experiment is run and nothing is significantCompute power before the run; revisit the budget if low
A fraction that hides the interaction you suspectMain effect is aliased with the interaction you care aboutChoose a higher resolution or fold the design later
Standard run orderDrift lines up with one factorRandomize, within blocks
No blocking when conditions changeA day effect appears as a factor effectBlock on day, lot, or operator
Too many factors at too many levelsA giant design that cannot be affordedScreen with two levels, then refine
No error estimateNo replicates, no center points, no negligible interactionsAdd replicates or center points
Overreaching with a screening designFitting a curved model to a two-level designAdd center points and axial runs in a follow-up
Design chosen by software defaultNo one can say why this designWrite down the choice and the alternatives
Ignoring setup timeThe plan needs more minutes than the line hasCount setup, test, and slack time

Design Resources on This Site

Guides

Tools

Stat Dojo

Body of Knowledge

Design Step Frequently Asked Questions

What is the difference between a full factorial and a fractional factorial?

A full factorial runs every combination of the factor levels: 2k runs for k two-level factors. A fractional factorial runs a carefully chosen part of them, such as half or a quarter. It needs fewer runs, but some effects are mixed up (aliased) with others. The design's resolution tells you which effects are mixed up.

What does resolution mean?

It describes which effects are aliased. In a resolution III design main effects are mixed up with two-factor interactions. In resolution IV, main effects are clear of two-factor interactions, but two-factor interactions are mixed up with each other. In resolution V, main effects and two-factor interactions are all clear of each other. The higher the resolution, the more runs it needs.

Are replicates the same as repeated measurements?

No. A replicate is a complete new run: the settings are reset and the whole process is repeated. Repeated measurements or several units made at one setting only capture the variation within that setting. Using them as if they were replicates makes the error look too small and finds effects that are not real. Averaging several units per run is fine, but the error for testing effects must still reflect setup-to-setup variation.

How many runs do I need?

Enough that an effect of the size that matters would stand out of the noise with high probability, usually 80% power. The standard error of an effect in a two-level design is 2σ/√N, where σ is the standard deviation of a run and N the number of runs. Use that, or the Power and Sample Size menu in Minitab, to check the options before you commit.

Should I always add center points?

Add them when curvature is a realistic concern and you plan to stay in the factorial region. They cost few runs, estimate pure error, and test for curvature. They cannot tell you which factor is curved. If the budget is tight, a screening design can skip them and the next experiment can add them.

Why randomize the run order?

Randomizing spreads uncontrolled drift (tool wear, temperature, operator fatigue, material changes) across all the factor settings, so it does not masquerade as a factor effect. If a factor is very hard to change, ask for help with a split-plot design instead of giving up randomization.

When do I use a Plackett-Burman design?

When you need to screen a lot of factors in very few runs and you are willing to assume interactions are small. They are resolution III, so main effects are mixed up with two-factor interactions. A fractional factorial of resolution IV or higher is usually the better choice when you can afford the runs.

Sources and Further Reading

  • Douglas C. Montgomery, Design and Analysis of Experiments, Wiley, on two-level factorial and fractional factorial designs, resolution, blocking, and power.
  • George E. P. Box, J. Stuart Hunter, and William G. Hunter, Statistics for Experimenters, Wiley.
  • Mark J. Anderson and Patrick J. Whitcomb, DOE Simplified, Productivity Press.
  • NIST/SEMATECH, e-Handbook of Statistical Methods, Process Improvement: Design of Experiments (itl.nist.gov/div898/handbook).
  • R. L. Plackett and J. P. Burman, “The Design of Optimum Multifactorial Experiments,” Biometrika, 1946.
  • Minitab Support, documentation for Create Factorial Design and Power and Sample Size for 2-Level Factorial Design (support.minitab.com).

This content is educational. The example data and results are illustrative. Follow your organization's quality system, change control, and safety requirements.

Step 3 of 5

Run

How do we run the experiment exactly as planned, and record what really happened? The best design is worth nothing if the runs are made at the wrong settings, in the wrong order, or recorded badly. Run is the discipline step: follow the plan, protect the data, and write down everything that does not go to plan.

Key question
Is every run made at the planned settings, in the planned order, and recorded so that we can trust the data?
Typical duration
A day to a week of preparation, then the runs themselves; the weld example needs two line days
Led by
The experiment owner, with a setter, an operator, a tester, and a recorder, plus the process engineer on call
Starts from
The approved design: factors, levels, runs, blocks, units per setup, and the randomization plan
Primary outputs
A completed run sheet, the actual settings, the results, noise records, a log of events, and a checked data file
Hands off to
Analyze: clean data with a complete record of what really happened
Gate decision
The experiment owner confirms all runs are complete or accounted for, and the data are checked and locked
Core tools
DOE Run Sheet Generator, pilot run, readback of settings, run log, check standard, time-order plot

What the Run Step Is For

An experiment is a controlled change, and the control is easy to lose. A setting that is a little off, a run done out of order, a lot of material that changed halfway, or a result copied wrongly can each turn a clean design into an ambiguous data set. The Run step protects the plan from the plant.

What Run must achieve

  • Every setup made at the planned settings, verified by a second person
  • The runs made in the randomized order, within blocks
  • Everything that was supposed to be held constant, held constant
  • The actual settings, the results, and the noise records written down
  • Every deviation noticed, recorded, and decided on at the time
  • A data file that is checked, backed up, and traceable to the run sheet

What Run must not do

  • Reorder runs to save setup time
  • Quietly fix a run and record it as planned
  • Let people who know the settings judge the units
  • Collect results on loose paper that is entered later from memory
  • Change anything else about the process during the experiment
  • Start analyzing before the data are checked and locked
The test of a finished Run. Could someone who was not there read the run sheet, the log, and the data file and tell exactly what was done at every setup, in what order, by whom, with which material, and what went wrong? If so, the data can be trusted, even if some runs did not go to plan.

The Run Step, Step by Step

Readiness Material, people, gauges Run sheet Randomized, printed Pilot Extremes and the chain Execute Setup, verify, run, test Log Settings, results, events Check data Plot, review, lock
Run moves from a readiness review to a locked data set. Most trouble comes from skipped readiness checks and unrecorded deviations.
  1. Review readiness

    Walk through the checklist: material set aside, equipment checked, gauges calibrated, people briefed, approvals in place, time booked.

    • Set aside enough of one lot for every run and the pilot
    • Confirm the instrument calibration and the check standard
    • Book the line, the operators, and the lab

    Output: A readiness sign-off

    Watch for: Starting without enough material of the one lot

  2. Generate and print the run sheet

    Build the randomized sheet with the planned design, blocks, and units per setup, and print enough copies. Keep the seed with the records.

    • Use the DOE Run Sheet Generator
    • Check the first rows by hand
    • Give each setup a run number that goes on every unit

    Output: A randomized run sheet, with the seed recorded

    Watch for: Re-sorting the sheet by setting to “save time”

  3. Brief the team

    Explain why the order matters, who does what, how to record, and what to do when something goes wrong. Make clear that reporting a problem is the right thing.

    • Name the setter, the welder, the tester, and the recorder
    • Agree who can stop the experiment
    • Give them the deviation rules

    Output: A team that knows the rules and the roles

    Watch for: A team that tries to be helpful by fixing things quietly

  4. Run a pilot

    Before the first real run, put a few units through the whole process, including the extreme corners. Check that every setting can be reached, that no combination damages the equipment, and that the measurement chain works.

    • Run the all-low and all-high corners
    • Test the measurement and the recording form
    • Time a setup change

    Output: A go decision, and a corrected plan if needed

    Watch for: Discovering on the first real run that a corner cannot be run

  5. Execute setup by setup

    For each run in order: reset every factor, have a second person read the settings back, make the units, test them with only the run number visible, and record the actual settings and results immediately.

    • Do not skip the reset, even when the next setting is the same
    • Keep everything else the same
    • Record before moving on

    Output: Completed runs, one at a time

    Watch for: Batching the recording until the end of the day

  6. Record and handle deviations

    Write down anything that did not go to plan, when it happened, and what was decided. Use the deviation rules, and call the experiment owner when the rules do not cover the case.

    • One log, in time order
    • Decisions made at the time, with initials
    • Never overwrite a record

    Output: A deviation log

    Watch for: A fix that is not written down

  7. Check the data and close out

    Enter the data from the sheet, check it against the paper, plot the results in the actual run order, and look for drift and obvious errors. Lock the data file and keep the originals.

    • Double-check the entries
    • Plot by run order and by block
    • Keep the paper and the file together

    Output: A checked, locked data set

    Watch for: Analyzing a file that has not been checked against the sheet

Readiness: What Must Be True Before Run 1

AreaCheckWhyWeld example
MaterialAll units for the experiment and the pilot set aside from one lot, labeled, and stored the same wayA lot change mid-experiment looks like a factor effectOne resin lot; 64 housings plus the pilot housings set aside; resin dried in one batch except for the drying-time factor
EquipmentIn good order; wear parts new or at a known state; counters resetWear changes the energy deliveredNew horn installed; cycle counter reset to zero
SettingsEvery factor can be set to every planned level and read backA level that cannot be reached cannot be runAll amplitude, force, time, hold, and clamp settings verified on the control panel; dryer times verified
MeasurementCalibrated; check standard measured; method writtenPoor measurement hides the effectsBurst tester calibrated that morning; same tester and tester operator for all units
PeopleBriefed and trained; roles assigned; backups namedSurprises cause mistakesOne setter, one welder, one tester, one recorder; engineer on call
ApprovalsSafety, quality, and the owner of the product have approved the rangesUnsafe or unapproved runs are not allowedQuality and safety approved the six ranges; experimental units are scrapped, not shipped
Time and spaceBooked, with room for repeatsA rushed experiment loses disciplineTwo line days booked; planned line time well below the time available
PaperworkRun sheets, log, data form, and labels printedWriting on scraps invites errors32 setups on two day sheets, with columns for results and notes

The Run Sheet and the Random Order

The run sheet turns the design into a list of setups in a random order, with space to record what actually happened. Use the DOE Run Sheet Generator: choose the design, enter the factors and levels, set the blocks, replicates, center points, and units per setup, and it produces a printable, randomized sheet. It also reports the resolution and the aliasing, so you can check that it is the design you approved.

FieldWhat it is for
Run numberThe order in which to run the setups. Mark every unit with it
Block (day)Which block the setup belongs to; finish one block before starting the next
Design rowThe row of the underlying design, for the analysis
Factor settingsThe planned level of every factor, with the coded level
Setting verifiedInitials of the second person who read the settings back
Result columnsOne per unit made at the setup
NotesAnything unusual: time, observation, name
  • Randomize more than the run order. Also randomize which material goes to which setup, and the order in which units are measured, where practical.
  • Keep the seed. The same seed gives the same order, so the sheet can be rebuilt if it is lost.
  • Do not sort the sheet by a factor to make changeovers easier. If a factor is very hard to change, tell the design team: a split-plot design may be needed.
  • Spread center points through the order, so they can show drift.

Roles and Communication

RoleDoesDoes not
Experiment ownerOwns the plan, decides on deviations, signs off the dataRun the setups
SetterResets and sets all factors for each run, reads them to the verifierMake or test units
VerifierReads the settings independently and initials the sheet before the first unitSign without reading
OperatorMakes the units for the run, with the same method every timeAdjust anything
TesterMeasures each unit, seeing only the run numberKnow or guess the setting
RecorderWrites actual settings, results, noise records, and events in time orderRely on memory
Process engineer (on call)Helps if equipment misbehavesChange settings on the fly

One person may play more than one role in a small team, but the verifier should not be the setter, and the tester should not see the settings. Agree before starting who may stop the experiment, and make it clear that anyone may call a stop.

Executing a Run: The Rules

DoDo notWhy
Reset every factor to the planned level before each setup, even if it is already thereSkip the reset when two setups look alikeA setup that is not reset is not a new setup and does not give true replication
Have a second person read each setting backRely on the setter’s memoryMis-set levels are the commonest error in experiments
Make all units for a setup before changing anythingInterrupt a setup to do something elseInterruptions change the conditions
Keep everything else the same: lot, operator, method, timingImprove the method during the experimentA change halfway ruins the comparison
Label each unit with the run number at onceLabel later from memoryMixed-up units cannot be sorted out afterwards
Record the actual setting, the result, and the time immediatelyWrite on a scrap and copy it up laterTranscription from memory is the commonest data error
Test units with only the run number visibleLet the tester see the settingExpectations bias judgments and readings
Stop and ask when something is unexpectedPress on to keep to the scheduleA deviation noticed early is cheap; one found after the analysis is not

What to Record

Record more than you think you need. You can ignore extra information later, but you cannot recover what you did not write down.

RecordContentExample
Run recordRun number, block, design row, date, start and end time, initialsRun 7, block 1, row 1, 08:42, initials
Actual settingsThe value read from the machine for every factor, whether or not it equals the planAmplitude 90, force 600, and so on
ResultsOne value per unit, with units and the tester’s initialsBurst pressure 312 kPa
Noise and covariatesThings you cannot control but can measure: humidity, temperature, horn cycle count, lot, operatorHumidity 46% on day 1
Events logAnything unusual, in time order, with the decision made and who made itDryer timer showed 5.4 h
Unit traceabilityRun number on the unit, retained samplesFailed housings kept in a labeled bag

Handling Deviations

Things will go wrong. What matters is that each one is noticed, written down, and handled by a rule agreed in advance, not by whoever happens to be on the line. The table gives defaults; the experiment owner makes the final call.

EventDefault responseWhy
A setting is wrong, and no unit has been madeCorrect it, have it verified again, note the time lostNo data are affected
A setting is found wrong after units were madeRecord the actual setting; either repeat the setup at the end of the block or analyze with the actual valueThe data are real but not the planned run
A factor cannot be set to the planned levelStop. Contact the experiment owner. Do not substitute silentlyThe design may need to change
Equipment fault or stoppageRecord the time and what was done; check the equipment state before resuming; repeat the setup if units were affectedFaults can change conditions without being obvious
A unit is lost or damagedRecord the reason; keep the other units; repeat the setup if too few remainLost units lower precision, and hiding them hides the cause
A result looks extremeDo not discard. Check the measurement and the record, note the findings, and keep the value unless an error is provenExtreme values may be the finding
Material runs out or a new lot is neededStop, and record. Start a new block with the new lot, if the design allowsA lot change is a block, not a footnote
The schedule is slippingDo not drop runs or reorder. Ask the owner for more timeA partly completed design is much harder to analyze
Never quietly fix a run. A correction that is not recorded makes the record untrue, and the next person to read it will believe a run went as planned when it did not.

Keeping the Measurement Honest

  • Same instrument, same method, same person for the whole experiment, where possible.
  • Check standard. Measure a reference at the start, in the middle, and at the end of each day. A drift in the check standard means the gauge changed, not the process.
  • Blind the tester. Only the run number on the unit; the setting stays with the setter.
  • Measure promptly, and in the same conditions: parts can change with time, temperature, or humidity.
  • Record raw readings, not rounded or converted ones, and note the units.
  • For destructive tests, keep a record of which unit was tested and in what order. The unit is gone, so the record is all there is.

Checking the Data Before Analysis

  1. Enter the data from the sheet into the file, with a second person checking a sample or the whole file against the paper.
  2. Check plausibility: sort each column and look at the smallest and largest values for typing errors and unit mix-ups.
  3. Plot the results in the order they were run. Look for drift, steps, and the effect of blocks. A trend with run order is a warning that something changed.
  4. Compare the actual settings with the planned settings, and mark any differences.
  5. Reconcile the units: units made, units tested, and units recorded must agree, and each missing one must be explained in the log.
  6. Lock the file and keep a copy of the paper. Every analysis starts from this version.

See Graphical Analysis for the run chart and Descriptive Statistics for summaries.

Worked Example: Running the Weld Experiment

The weld team ran the design from the Design step: 32 setups, two housings at each, over two line days. All figures are illustrative.

StepWhat happened
ReadinessOne resin lot set aside for all 64 housings and four pilot housings; new horn installed with its cycle counter at zero; burst tester calibrated; the team briefed on the deviation rules; quality and safety approvals confirmed
Run sheetGenerated with the DOE Run Sheet Generator: half fraction of six factors, two blocks (days), two units per setup, seed 2026. The first setup on day 1 is design row 30, and the second is row 25. Day 1 holds the 16 setups with A × B × C = −1
Pilot (afternoon before)The all-low and all-high corners, two housings each: both welded cleanly, the horn showed no damage, and the burst tester and the recording form worked. Pilot housings were taken from the spare stock, and were not used in the analysis. 32 minutes
Day 116 setups. Humidity 46%. One near miss and one lost housing (see the log)
Day 216 setups. Humidity 58%. One setup pulled and rerun at the end of the day
Data63 of 64 housings gave a valid result. The horn counter read 68 at the end (0 at the start, four pilot welds and 64 experimental housings). The data were entered from the sheets, checked against the paper, and locked

The first eight setups in the order they were run, with the settings read back and the results:

RunDayDesign rowAmplitude (%)Force (N)Weld time (s)Hold time (s)Drying (h)Clamp (kPa)Results (kPa)Run average
1130904000.51.56300293, 300296.5
2125704000.31.56300272, 312292.0
3112906000.31.52450302, 284293.0
4123706000.50.56450327, 324325.5
517706000.50.52300308, 331319.5
6128906000.31.56300354, 330342.0
711704000.30.52300245, 254249.5
8122904000.50.56450330, 314322.0

The events log, in time order:

WhenEventDecisionEffect on the data
Day 1, run 7Setter set amplitude 70 %; the verifier read back 72 % on the panelReset, verified again, then welded. 5 minutes lostNone: caught before any unit was made
Day 1, run 14One of the two housings cracked while being unloaded from the fixture (design row 20)Kept the other housing; no repeat, because the lost housing was handling damage, not a weld resultThe setup has one valid result: the run average is a single value (313 kPa)
Day 2, run 25The dryer timer for the resin showed 5.4 hours instead of the planned 6 for design row 27Pulled the setup before welding; dried 40 more minutes; reran it as the last setup of day 2Moved from position 25 to position 32; planned time 512 minutes became 557
Both daysHumidity recorded at the start of each day46% on day 1 and 58% on day 2Recorded as a covariate; the day block also carries it
240 260 280 300 320 340 360 380 1 2 4 6 8 10 12 14 16 18 20 22 24 26 28 30 32 Setup number in the order run (average of the units tested) Day 2
Run averages in the order they were run. No drift or step is visible; the gold points are day 2.
Planned 512 min Actual 557 min Available 960 min Line minutes over the two days
The experiment took 557 minutes against 960 available, leaving room for the lost time.
Data checkResult
Setups completed32 of 32
Valid housings63 of 64 (one lost to handling)
Burst pressure, all valid housingsmean 304.7 kPa, standard deviation 29.0, range 245 to 366
Day averages (of setup averages)day 1 303.8 kPa and day 2 306.0 kPa: no important day-to-day difference
Drift with run orderNone visible on the run chart
Entries checked against the paper100% of entries, by a second person
On the schedule. A day of preparation and the pilot, two line days for the runs, and a half day to check, lock, and hand over the data. The data go to the Analyze tab exactly as recorded, with the log.

Run Check: Is the Data Ready for Analysis?

Before starting the Analyze tab, confirm that the experiment is complete and the data are trustworthy. Use this checklist; progress saves in this browser only, and nothing is sent anywhere.

0 of 17 complete
Before the first run
During the runs
After the runs

Questions an experiment owner or coach can ask:

  • Did every run follow the sheet? If not, what changed and who decided?
  • Who verified the settings, and where is the record?
  • What was held constant, and did it stay constant?
  • Where are the original sheets, and who checked the data against them?
  • Does the run chart show any drift or any block effect?
  • Is there anything we did not write down?
OutcomeMeaningNext step
ReadyAll runs complete, deviations logged, data checked and lockedStart the Analyze tab
Ready with conditionsOne or two documented deviations that the analysis can handleState them in the analysis and check whether they matter
Not readyMissing runs, unlogged changes, or unchecked entriesRepeat the runs or fix the records first
InvalidA major change occurred (new lot, broken equipment) that was not blockedTreat the experiment as incomplete: repeat it, or analyze with the change as a block

Adapting Run to the Situation

SituationHow Run changes
Manufacturing lineProtect the experiment from the normal production flow: label and segregate units, and hold material, operator, and equipment constant
Long processes (days per run)Plan the schedule carefully; record the conditions through the run; consider a smaller design
Batch and chemical processesA run is a batch: record the batch conditions; randomize batch order; watch for carry-over between batches
Transactional and service workRandomize across customers, days, or agents; keep the script or the method fixed; protect customers who see more than one condition
Software and online servicesRandomization is automatic, but check that the assignment really is random, that logging works, and that overlapping experiments do not interact
Healthcare and regulated workFollow the approved protocol exactly; document every deviation formally; follow consent and ethics rules
Outdoor and field workWeather and site conditions are noise: record them, and block on them where you can

Scale the discipline to the risk. A cheap, quick, internal experiment needs a simple sheet and a log. A costly or regulated one needs signatures, formal deviation reports, and retained samples.

Common Mistakes and Red Flags

MistakeWhat it looks likeHow to correct it
Standard-order runningAll the high settings of one factor done togetherFollow the randomized sheet, within blocks
No reset between similar setupsTwo setups ran back to back with no changeReset and verify every setup
Unverified settingsA mis-set level found after the analysisA second person reads every setting
Unrecorded fixesData look strange and nobody remembers whyLog every event, with the decision and the name
The lot changed halfwayA shift in the response with the second half of the runsOne lot, or a new lot as a block
Tester knows the settingsResults move in the expected direction too neatlyLabel by run number only
Recording after the factEntries copied from memory or scrapsRecord at the machine, immediately
Skipping the pilotA corner of the design cannot be run on day 1Run the extreme corners first
Dropping runs to save timeA design with holes that cannot be analyzed cleanlyAsk for more time instead
Analyzing before checkingA typing error becomes a findingCheck the data against the paper and plot them first

Run Resources on This Site

Guides

Tools

Templates

Body of Knowledge

Run Step Frequently Asked Questions

Why must the runs be done in the randomized order?

Randomizing spreads drift, such as tool wear, temperature, material changes, and operator fatigue, across all the factor settings. If you run in a convenient order, any drift lines up with some factor and looks like an effect of that factor. If a particular order is truly impossible, say so in the design step so that a split-plot or blocked plan can be used, and do not reorder on the fly.

What should I do if I notice a mistake in the middle of the experiment?

Stop, record exactly what happened, and decide with the experiment owner. If the mistake happened before any unit was made, correct it and carry on. If units were made at the wrong setting, record the actual setting, and either repeat the run at the end or analyze it with the actual values. Never quietly fix a run and pretend it went as planned.

Should the people who measure the response know the settings?

Where possible, no. Label units with a run number only, and keep the settings sheet with the person who set the machine. A measurement made by someone who expects a larger value can be nudged toward it. For machine-read measurements this matters less, but the labeling is still good practice.

What if a unit is lost or a measurement is clearly wrong?

Record it in the log with the reason, keep the other units from that setup, and do not replace or discard it in the data file without a note. A lost unit is a deviation to report, not a problem to hide. If the loss is serious, repeat the setup at the end of the block.

Do I record settings or just the results?

Record both. Write down the setting you actually read on the machine for every factor, even when it equals the plan, and record the noise factors that you chose not to control: humidity, resin lot, horn cycle count, operator, time. Those records are what allow you to explain a surprise afterwards.

How many pilot runs should I do?

Enough to test the extremes and the measurement chain: the combinations of all-low and all-high settings, and a few housings through the whole process from setup to recorded result. Pilot units are not part of the experiment, and they are paid for outside its budget.

Sources and Further Reading

  • Douglas C. Montgomery, Design and Analysis of Experiments, Wiley, on guidelines for conducting experiments.
  • George E. P. Box, J. Stuart Hunter, and William G. Hunter, Statistics for Experimenters, Wiley.
  • Mark J. Anderson and Patrick J. Whitcomb, DOE Simplified, Productivity Press.
  • NIST/SEMATECH, e-Handbook of Statistical Methods, Process Improvement: Design of Experiments (itl.nist.gov/div898/handbook).
  • Automotive Industry Action Group, Measurement Systems Analysis Reference Manual, 4th ed.

This content is educational. The example data and results are illustrative. Follow your organization's quality system, change control, and safety requirements.

Step 4 of 5

Analyze

Which factors and interactions really matter, how big are they, and can we trust the model? Analysis turns 32 setup averages into a short list of effects you can believe, a model that predicts the response, and a clear statement of what the experiment could not tell you. The aim is a defensible conclusion, not a long table.

Key question
Which factors and interactions are real, how big are they, and does the model hold up?
Typical duration
A day or two: effects, plots, a reduced model, checks, and a short write-up
Led by
The project leader or Black Belt, with the process engineer to judge whether the effects make physical sense
Starts from
The locked data file, the run log, and the design summary with its aliasing
Primary outputs
Effects with a measure of certainty, a checked reduced model, plots, and a statement of limits
Hands off to
Optimize: a model that can predict, and a list of which settings matter and which do not
Gate decision
The model passes its checks, and the team agrees the effects make engineering sense
Core tools
Pareto and normal plot of effects, ANOVA, main effect and interaction plots, residual plots, Minitab Analyze Factorial Design

What the Analyze Step Is For

The Run step produced numbers. The Analyze step decides what they mean. In a factorial experiment each effect is a simple difference between averages, so the arithmetic is easy; the hard part is judging which differences are bigger than the noise, whether the model describes the process, and whether the answer is useful.

What Analyze must achieve

  • A look at the raw data and the log before any model
  • Effects for every factor and interaction the design can estimate
  • A fair test of which effects are real, with an error estimate that fits the design
  • Plots that show the findings without a table
  • A reduced model that respects hierarchy and passes its checks
  • A comparison of every effect with the smallest effect worth finding
  • A clear statement of what the experiment could not tell us

What Analyze must not do

  • Treat several units at one setup as independent replicates
  • Test every term at 5% and report whichever passes
  • Drop a parent main effect while keeping its interaction
  • Skip the residual checks because the p-values look good
  • Interpret main effects alone when an interaction is large
  • Present a prediction without its uncertainty
The test of a finished analysis. Could someone read one page and tell which factors matter and by how much, which do not, how sure we are, what assumptions the conclusion rests on, and what we still do not know? If so, the analysis is ready for the Optimize step.

The Analyze Step, Step by Step

First look Data, log, blocks Effects Every term Which are real Error and plots Reduced model Hierarchy, ANOVA Check Residuals Conclude Meaning and limits
Analysis moves from the data to a conclusion. A failed residual check sends you back to the model.
  1. Look at the data first

    Read the log, plot the results in run order and by block, and note every deviation. Decide the unit of analysis: one value per setup.

    • Check for drift, steps, and block differences
    • Treat events from the log as part of the data
    • Average the units at each setup

    Output: One response per setup, and a list of known issues

    Watch for: Analyzing every housing as if it were independent

  2. Estimate all the effects

    Fit the full model: every main effect and every two-factor interaction the design can separate. The effect of a term is the average response at its high level minus the average at its low level.

    • Use coded levels so effects are comparable
    • Estimate the block effect separately
    • Keep the aliasing in mind when you read the list

    Output: A table of effects for all terms

    Watch for: Fitting a model with terms the design cannot separate

  3. Judge which effects are real

    Compare each effect with an error estimate. Use the error from higher-order interactions, Lenth's pseudo standard error, or replicates, and look at a normal plot of the effects.

    • Use two methods and see whether they agree
    • Look at the plot as well as the p-values
    • Treat a borderline effect as a question for the next experiment

    Output: A short list of effects that stand out of the noise

    Watch for: Using p-values from a model with no real error estimate

  4. Plot the findings

    Draw the main effects and the interaction plots for the important terms. Plots show the direction and the size, and they show the interaction in a way a table cannot.

    • Lines for main effects; two lines for each interaction
    • Same vertical scale on every plot
    • Mark the specification or the target

    Output: Plots that tell the story

    Watch for: Reading a main effect that is part of a large interaction

  5. Reduce the model

    Keep the real terms, plus their parents. Fit the reduced model, look at the ANOVA, the coefficients, and R-squared, adjusted and predicted.

    • Keep hierarchy
    • Use the reduced model's error to test and to predict
    • Keep the block in the model if the design had blocks

    Output: A reduced model with coefficients and an ANOVA table

    Watch for: Removing terms until only the significant ones are left, ignoring hierarchy

  6. Check the model

    Plot the residuals against the fitted values, in run order, and on a normal probability plot. Look at outliers and influence.

    • Random scatter, no funnel
    • No trend in run order
    • Points near the line in the normal plot

    Output: A checked model, or a reason to change it

    Watch for: Trusting the p-values of a model that fails its checks

  7. State the practical meaning and the limits

    Compare each effect, with its interval, with the smallest effect worth finding. Say what the experiment could not tell you: curvature, factors held fixed, the range tested.

    • Intervals, not only p-values
    • Say which settings do not matter and so can be chosen for cost or convenience
    • List what the next experiment should check

    Output: A one-page conclusion

    Watch for: Reporting effects with no sense of their size or uncertainty

Start With the Data, Not the Model

  • Read the log. Every deviation affects how you analyze: a lost unit, a setup that moved in the order, a block change.
  • Plot the results in the order they were run, and by block. A step between blocks, or a trend with run order, is a warning that something changed.
  • Choose the unit of analysis. The experimental unit is the setup, not the housing. Use one value per setup, normally the average of its valid units. A setup with one unit, such as the one with a cracked housing, is simply noisier, and a weighted analysis can check that it matters.
  • Know what you cannot test. With no replicates and no center points, the design cannot give a pure-error estimate or test curvature, so those assumptions need care.

See Graphical Analysis for the run chart and the box plot by block.

Estimating the Effects

With coded levels, the effect of a factor is the average response at its high level minus the average at its low level. For weld force (B), the 16 setups at 600 N average 320.1 kPa and the 16 at 400 N average 289.7 kPa, so the effect is +30.4 kPa. Interactions are estimated in the same way, from the product of the two coded columns. The sum of squares for a term is N × effect² / 4: for force, 32 × 30.4² / 4 = 7,412.

The design is a resolution VI half fraction, so every main effect and every two-factor interaction has its own column: six main effects plus 15 two-factor interactions, which is 21 effects, with 9 degrees of freedom left from the three-factor interactions (A × B × C is confounded with the day). The +2.2 kPa difference between days is small.

TermEffect (kPa)Standard errortp-valueBeyond the noise?
B Weld force+30.45.4+5.66< 0.001Yes
E Resin drying time+20.95.4+3.890.004Yes
C Weld time+20.95.4+3.880.004Yes
B × C-15.75.4-2.920.017Yes
A Weld amplitude-6.35.4-1.170.271No
B × F-5.05.4-0.930.377No
D × E+4.85.4+0.890.394No
A × B-3.15.4-0.580.576No
D × F-3.15.4-0.570.583No
A × C-2.95.4-0.550.598No

How the standard error was found. The 9 distinct three-factor interactions (the tenth, A × B × C, is the block) are assumed to be zero, so their variation is error: mean square 231, standard error of an effect = √(4 × 231 / 32) = 5.38 kPa, with 9 degrees of freedom. An effect is significant at 5% when it exceeds 2.26 × 5.38 = 12.2 kPa.

Which Effects Are Real?

With no replicates there is no direct measure of pure error, so the experiment estimates it indirectly. Use more than one method; where they agree you can be more confident.

MethodHow it worksIn this experiment
Pooled higher-order interactionsTreat three-factor and higher interactions as zero, so their variation is errorError from 9 degrees of freedom: standard error 5.4 kPa; line at 12.2 kPa; real effects: B, E, C, BC
Lenth's methodEstimates the noise from the median size of the effects, ignoring the large onesPseudo standard error 3.8; margin of error 8.9 kPa; real effects: B, C, E, BC
Normal plot of effectsEffects from noise fall on a straight line; real ones fall away from itThe same four effects stand clear of the line (see the plot)
Replicates or center pointsA direct pure-error estimateNot available in this design
B Weld force +30.4 E Resin drying time +20.9 C Weld time +20.9 B x C -15.7 A Weld amplitude -6.3 B x F -5.0 D x E +4.8 A x B -3.1 D x F -3.1 A x C -2.9 5% line: 12.2 kPa Size of the effect on burst pressure (kPa, absolute value; sign beside the bar)
The Pareto chart ranks the effects. Bars beyond the dashed line are larger than the noise; the sign is beside each bar.
Both methods agree. Weld force (B), drying time (E), weld time (C), and the force by weld time interaction (B × C) stand out. Every other effect is smaller than 9 kPa, which is under both cut-offs for noise. The smallest real effect, B × C, has p = 0.017: a result worth confirming rather than leaning on.

Plotting the Findings

-2 -1 0 1 2 -20 -10 0 10 20 30 BC C E B Effect (kPa) Normal score
A normal plot of the effects. Small effects, which are noise, line up on the dashed line; the labeled points do not.
280 290 300 310 320 330 Low levelHigh level A Weld amplitude (-6) B Weld force (+30) C Weld time (+21) D Hold time (+0) E Resin drying time (+21) F Clamp pressure (-1) Average burst pressure (kPa)
Main effects. Steeper lines are bigger effects. The gray lines are factors whose effect is within the noise.
260 270 280 290 300 310 320 330 Weld time 0.3 s Weld time 0.5 s Weld time Average burst pressure (kPa) Force 400 N Force 600 N
The interaction plot. The lines are not parallel, which is what an interaction looks like.
  • Force and drying time raise burst pressure from low to high by 30 and 21 kPa. Longer drying probably helps by removing moisture from the resin, a mechanism the team had not tested before.
  • The force by weld time interaction means the factors trade off. At low force, a longer weld time raises burst pressure by 37 kPa; at high force it adds only 5. Enough force makes weld time nearly irrelevant, so reading the weld time main effect alone would mislead.
  • Amplitude, hold time, and clamp pressure had no effect that stands out of the noise, so they can be set for cost, cycle time, or convenience, within the tested range.

The Reduced Model

Keep the real effects, B, C, E and B × C (whose parents B and C are already there), and the day block. Fitted to the 32 setup averages, the reduced model has 26 degrees of freedom for error, so its tests and its intervals use all the information in the experiment.

SourceDFSum of squaresMean squareFp-value
Model516,4143,28325.7< 0.001
  Day (block)140400.30.578
  B Weld force17,4127,41258.1< 0.001
  C Weld time13,4863,48627.3< 0.001
  E Resin drying time13,5073,50727.5< 0.001
  B × C11,9691,96915.4< 0.001
Error263,317128
Total3119,731
TermEffectCoefficientSEtp-value95% interval for the effect
Constant304.882.00152.69< 0.001
Day (block)+2.21.122.000.560.578-6.0 to +10.5
B Weld force+30.415.222.007.62< 0.001+22.2 to +38.6
C Weld time+20.910.442.005.23< 0.001+12.7 to +29.1
E Resin drying time+20.910.472.005.24< 0.001+12.7 to +29.1
B × C-15.7-7.842.00-3.93< 0.001-23.9 to -7.5
Model in coded units (−1 at the low level, +1 at the high level): Burst pressure = 304.9 +1.1 Day + 15.2 B + 10.4 C + 10.5 E −7.8 B×C. Standard deviation of the residuals S = 11.3 kPa; R² = 83%, adjusted 80%, predicted 75%. The predicted R² is close to the ordinary one, so the model is not just fitting the noise.

Noise came in lower than planned. The Design step assumed a run-level standard deviation of about 18.5 kPa. The residual standard deviation here is 11.3. The planning figure was cautious, which is the safe way to plan: the experiment ended up with better power than the 78% promised, and the effects are more certain than the design promised.

Sensitivity check. One setup has a single valid housing, and so is noisier. Refitting with each setup weighted by the inverse of its variance changes no effect by more than 0.3 kPa, so the simple analysis stands.

Checking the Model

Residuals versus fitted values Residuals versus run order Normal probability plot of residuals Fitted burst pressure (kPa) Day 2 Setup number in the order run Normal score
Residual plots: random scatter, no trend with run order, and a straight line in the normal plot are what you want.
CheckWhat to look forResult
Residuals versus fittedNo curve, no funnelRandom scatter. Spread in the upper half of the fitted values is 1.24 times that in the lower half (Levene's p = 0.28)
Residuals versus run orderNo trend, no step at the day changeCorrelation with run order -0.01; Durbin-Watson 2.08 (2 means no serial correlation)
Normal plot of residualsPoints near the lineClose to the line (Shapiro-Wilk p = 0.84)
Outliers and influenceStandardized residual under about 2.5; Cook's distance well below 1Largest standardized residual 2.19; largest Cook's distance 0.19
CurvatureA significant difference between center-point average and factorial averageCannot be tested: there are no center points. Check in the next experiment

See Residual Analysis and Model Checking for how to read these plots and what to do when they show a problem.

Statistical Versus Practical Meaning

The frame set the smallest effect worth finding at 20 kPa. Compare every effect, with its interval, against it.

EffectSize (kPa)95% intervalCompared with 20 kPaMeaning
B Weld force+30.4+22.2 to +38.61.5 timesClear of the 20 kPa line even at the lower end
E Resin drying time+20.9+12.7 to +29.11.0 timesReal; the lower end of the interval is below 20 kPa
C Weld time+20.9+12.7 to +29.11.0 timesReal; the lower end of the interval is below 20 kPa
B × C-15.7-23.9 to -7.50.8 timesReal; the lower end of the interval is below 20 kPa

Factors with no real effect do not prove that the factor has no influence. The interval tells you how large an effect could still be hiding:

EffectSize (kPa)Interval from the pooled errorMeaning
A Weld amplitude-6.3-18.5 to +5.9Under 20 kPa either way: set for cost or convenience
D Hold time+0.2-11.9 to +12.4Under 20 kPa either way: set for cost or convenience
F Clamp pressure-1.2-13.4 to +11.0Under 20 kPa either way: set for cost or convenience
  • Weld force is the main lever: 30 kPa, about 1.5 times the smallest effect that matters.
  • Drying time and weld time each give about 21 kPa, right at the size that matters, with the lower end of the interval below it. Both look worth using, and both are worth confirming.
  • The interaction has a practical use. At high force, weld time barely matters, so the team can shorten the weld time to save cycle time without losing strength.
  • Amplitude, hold time, and clamp pressure can be set to whatever is cheapest, because even their upper limits sit below 20 kPa. Their ranges were safe, and they were held within them.

When the Results Look Odd

SymptomLikely causeWhat to do
Nothing is significantNoise larger than planned, narrow ranges, low power, or a lost runCheck the measurement and the ranges; report the intervals; plan the next experiment with more runs or bolder ranges
Everything is significantError estimate too small, for example units treated as replicatesAnalyze one value per setup; check the error degrees of freedom
A large interaction hides a main effectThe factor helps in one condition and hurts in anotherRead the interaction plot, not the main effect
Residuals curve with the fitted valuesCurvature, or a missing termAdd center points and axial runs in a follow-up; consider a transformation
Residuals fan outVariation grows with the responseConsider a log or other transformation, or analyze the variation separately
A trend with run orderDrift: wear, temperature, materialAdd run order or the covariate to the model; repeat if it is large
One point is far from the restA recording error, a lost or damaged unit, a real eventCheck the log and the records; do not delete without a reason; analyze with and without it
An effect is aliased with one you did not expectA fraction of low resolution, or an alias you forgotCheck the alias table; fold over or run the other fraction to separate them
Effects do not make engineering senseA mislabeled column, a coding error, or a surpriseCheck the coding first, then discuss with the process engineer; confirm the surprise before using it

Running the Analysis in Excel and Minitab

The statistics are explained, worked by hand, in Analyzing Designed Experiments in Stat Dojo, and the software workflow is in the Minitab Guide. In summary:

Excel

  1. Enter one row per setup with the average response and the coded −1 and +1 columns for every factor and the block.
  2. Effects: =AVERAGEIF(col, 1, response) - AVERAGEIF(col, -1, response) for each factor. Make an interaction column by multiplying two coded columns.
  3. Model: Data > Data Analysis > Regression with the response as Y and the block, B, C, E, and B×C columns as X. The coefficients are half the effects, and the output has the ANOVA table, S, and R-squared.
  4. Error for the full model: square and sum the three-factor effects (× N/4), divide by their count, and use the formula for the standard error given above.
  5. Plots: a scatter of the residuals against the fitted values, and a line chart of the cell averages for the interaction. Excel has no normal plot of effects; sort the effects and plot them against =NORM.S.INV((i-0.5)/n).

Minitab

  1. Stat > DOE > Factorial > Analyze Factorial Design. Choose the response (one value per setup), and in Terms include the factors and two-factor interactions, with the block.
  2. In Graphs, choose a Pareto of effects, a normal plot of effects, and the four-in-one residual plots. With no error degrees of freedom Minitab uses Lenth's method for the reference line.
  3. Look at the Pareto and the normal plot, then use Terms or Stepwise to fit the reduced model, keeping parents of interactions.
  4. Stat > DOE > Factorial > Factorial Plots for the main effects and interaction plots.
  5. Read the session window: the Analysis of Variance table, the Model Summary, and the Coded Coefficients, where Minitab gives the effect as twice the coefficient.
Factorial Regression: Burst versus Blocks, Force, WeldTime, Drying

Analysis of Variance

Source                  DF   Adj SS   Adj MS  F-Value  P-Value
Model                    5   16414.0    3282.8    25.73    0.000
  Blocks                  1      40.5      40.5     0.32    0.578
  Linear                  3   14404.7    4801.6    37.64    0.000
    Force                 1    7411.5    7411.5    58.09    0.000
    WeldTime              1    3486.1    3486.1    27.33    0.000
    Drying                1    3507.0    3507.0    27.49    0.000
  2-Way Interactions      1    1968.8    1968.8    15.43    0.001
    Force*WeldTime        1    1968.8    1968.8    15.43    0.001
Error                   26    3317.0     127.6
Total                   31   19731.0

Model Summary

      S    R-sq  R-sq(adj)  R-sq(pred)
11.2950  83.19%     79.96%      74.53%

Coded Coefficients

Term              Effect    Coef  SE Coef  T-Value  P-Value   VIF
Constant                   304.88   2.00   152.69    0.000
Blocks                       1.12   2.00     0.56    0.578
Force               30.44    15.22   2.00     7.62    0.000  1.00
WeldTime            20.88    10.44   2.00     5.23    0.000  1.00
Drying              20.94    10.47   2.00     5.24    0.000  1.00
Force*WeldTime     -15.69    -7.84   2.00    -3.93    0.001  1.00

Typed excerpt of the reduced-model output, simplified from the calculated example. Menu names follow recent versions of Minitab Statistical Software and can differ slightly in older releases.

Worked Example: Analyzing the Weld Experiment

The team took the locked data from the Run step: 32 setup averages, from 63 valid housings, over two days. All figures are illustrative.

StepWhat the team found
First lookThe run chart showed no drift. Day averages differ by +2.2 kPa, which is small. The setup with one valid housing is noisier but a weighted check changes no effect by more than 0.3 kPa. The analysis uses one value per setup
EffectsOf 21 effects, four stand out: B weld force +30.4, E drying time +20.9, C weld time +20.9, and B × C -15.7 kPa. The largest of the rest is A at -6.3
Which are realThe error from 9 three-factor interactions gives a line at 12.2 kPa; Lenth's margin is 8.9. Both methods and the normal plot pick the same four
PlotsMain effects: force, drying time, and weld time raise the burst pressure. The interaction plot shows that at 600 N the weld time barely matters
Reduced modelDay, B, C, E, B × C. R² 83%, adjusted 80%, predicted 75%; S = 11.3 kPa, below the 18 planned
ChecksResiduals are random against the fitted values and the run order, and are close to normal (Shapiro-Wilk p = 0.84). No point has a Cook's distance above 0.19
Practical meaningForce (30 kPa) is the main lever; drying time and weld time give about 21 each. Amplitude, hold time, and clamp pressure are within the noise and can be set for cost
LimitsNo center points, so curvature is untested; drying time is at the edge of its tested range, so more may help but is unproven; one resin lot only; the aliasing is clean, but the three-factor interactions were assumed to be negligible
What goes to Optimize. A model with four active terms, a residual standard deviation of 11.3 kPa, and a short list of settings that do not matter. The next question is: what settings give the target of 322 kPa on average, with a good margin, and does the prediction hold in a confirmation run?

Analyze Check: Do We Understand the Results?

Before starting the Optimize tab, confirm that the analysis is complete and the model can be trusted. Use this checklist; progress saves in this browser only, and nothing is sent anywhere.

0 of 12 complete
Data and unit of analysis
Effects and model
Checks and meaning

Questions a sponsor or a coach can ask:

  • Which factors matter, and by how much?
  • Which do not, and how sure are you?
  • Is any effect large enough to use, and is it certain enough?
  • Does the interaction change the advice?
  • What did the residual plots show?
  • What could the experiment not tell us?
OutcomeMeaningNext step
ReadyThe model passes its checks, and the effects make senseStart the Optimize tab
Ready with conditionsA weak point, such as one borderline effect or untested curvatureCarry it as a question into the confirmation
Not readyA failed residual check, or an unexplained patternFix the model, or collect more data
InconclusiveNothing stands out of the noise, or the intervals are too widePlan a follow-up with more power or bolder ranges

Adapting Analyze to the Situation

SituationHow Analyze changes
Replicated designsUse the replicates for pure error and a lack-of-fit test; use a full ANOVA with all terms and drop the non-significant ones
Designs with center pointsAdd the curvature test: the center-point average against the factorial average
Non-normal or proportional dataConsider a transformation, or a model for counts or proportions (see Binary Logistic Regression)
Variation as a responseWith several units per setup, analyze the standard deviation or its log as a second response; with few units it is very noisy
Many responsesAnalyze each, then find settings that satisfy all of them with the optimizer
Unbalanced or incomplete designsUse regression with the actual data, check the aliasing and the variance of each effect, and be careful with orthogonality
Software and online experimentsUse the metric's own noise; consider multiple-comparison corrections when many metrics are tested
Healthcare and regulated settingsFollow the pre-specified analysis plan, and document any departure

Decide the analysis before seeing the data. A pre-planned analysis, with the model terms and the error estimate agreed in advance, protects against finding what you hoped to find.

Common Mistakes and Red Flags

MistakeWhat it looks likeHow to correct it
Units as replicatesDozens of significant terms and a tiny errorOne value per setup; use the setup as the experimental unit
No real error estimateA table of p-values from a model with no degrees of freedom leftUse pooled higher-order terms, Lenth's method, or replicates
Ignoring hierarchyAn interaction in the model without its parent main effectsKeep the parents
Reading main effects through an interactionAdvice contradicted by the interaction plotInterpret the interaction plot first
Skipping the residual checksStrange patterns appear later in the confirmationPlot the residuals before trusting any p-value
Deleting an unusual run without a reasonA clean model that was cleaned by handCheck the log; analyze with and without it; report
Reporting p-values onlyNo sense of sizeGive effects with intervals and compare with the smallest effect that matters
Over-reading a borderline effectA decision rests on p = 0.04Confirm with runs before changing the process
ExtrapolatingSettings outside the tested rangesStay inside the region; run a new experiment to go beyond it
Over-fittingR-squared climbs and predicted R-squared fallsReduce the model; compare adjusted and predicted R-squared

Analyze Resources on This Site

Guides

Tools

Stat Dojo

Body of Knowledge

Analyze Step Frequently Asked Questions

How do I tell which effects are real when there are no replicates?

Use an error estimate that does not need replicates. In a fractional factorial, the three-factor and higher interactions are usually negligible, so their pooled sum of squares estimates the error. Lenth's method estimates it from the effects themselves, and a normal probability plot of the effects shows the real ones standing apart from the line of small ones. Using more than one method and seeing them agree is reassuring.

What is the difference between statistical and practical significance here?

A statistically significant effect is unlikely to be noise. A practically significant effect is big enough to matter for the decision. Compare each effect, and its interval, with the smallest effect worth finding that you set in the frame. An effect can be real and too small to use, or too uncertain to rule out as important.

Should I analyze each unit or the average of the units at each setup?

Analyze one value per setup, usually the average of its units. Units made at one setup share the same setup, so they are not independent evidence about the factors. Treating them as independent makes the error look too small and finds effects that are not there. The Design step planned for this: the setup is the experimental unit.

Why keep a non-significant main effect in the model when its interaction is significant?

Because the model should respect hierarchy: if an interaction is in the model, its parent main effects should be too. Dropping a parent distorts the interaction and makes the model depend on how the factors were coded. In the example, both parents of the force by time interaction are significant anyway.

What if nothing is significant?

Check the possible reasons before concluding that nothing matters: measurement noise, factor ranges that were too narrow, an unstable process, a lost run, or too little power for an effect of the size that matters. Report the intervals. A well-run experiment that finds nothing larger than the smallest effect worth finding is still a result: it tells you where not to look.

Can I trust the model outside the region I tested?

No. A two-level model describes a flat surface, and the process may curve beyond or between the levels. Without center points the design cannot even test for curvature inside the region. Use the model to choose settings within the tested ranges, and confirm with runs before changing the process.

Sources and Further Reading

  • Douglas C. Montgomery, Design and Analysis of Experiments, Wiley, on the analysis of two-level factorial designs.
  • George E. P. Box, J. Stuart Hunter, and William G. Hunter, Statistics for Experimenters, Wiley.
  • Russell V. Lenth, “Quick and Easy Analysis of Unreplicated Factorials,” Technometrics, 1989.
  • Cuthbert Daniel, “Use of Half-Normal Plots in Interpreting Factorial Two-Level Experiments,” Technometrics, 1959.
  • NIST/SEMATECH, e-Handbook of Statistical Methods, Process Improvement: Design of Experiments (itl.nist.gov/div898/handbook).
  • Minitab Support, “Methods and formulas for Analyze Factorial Design” (support.minitab.com).

This content is educational. The example data and results are illustrative. Follow your organization's quality system, change control, and safety requirements.

Step

Optimize

What settings should we use, does the prediction hold, and what do we do next?

This tab is being built. It will follow the same layout as the Frame tab: an overview, step-by-step guidance, the core tools, a checklist, common mistakes, and links to the site's calculators and templates. In the meantime, see the Design of Experiments guide and Analyzing Designed Experiments.