How to Run a Design of Experiments: A 5-Phase Checklist from Plan to Confirmation

July 2026 14 min read Bioprocess Engineering

Key Takeaways

Contents

  1. The DOE methodology in one picture
  2. Phase 1 — Plan
  3. Phase 2 — Design
  4. Phase 3 — Conduct
  5. Phase 4 — Analyze
  6. Phase 5 — Confirm
  7. The barriers that sink DOE projects
  8. Your DOE project checklist
  9. Frequently Asked Questions

Most people learning how to run a design of experiments are handed a statistics course when what they need is a project plan. The arithmetic is the easy part and modern software does it for you. What actually decides whether a DOE study changes a process is a sequence of unglamorous decisions made before any vessel is inoculated: what exactly are we measuring, can we measure it well enough, which factors are in, how wide are the ranges, and what will we do with the answer. This guide walks the full DOE methodology in five phases, with the numbers and the failure modes at each step, and ends with a DOE project checklist you can work through before committing a campaign. It assumes no statistical background. If you are earlier than that, start with design of experiments for beginners first.

The DOE methodology in one picture

Antony divides the methodology into four phases: planning, designing, conducting, and analysing. A fifth, confirmation, belongs on the end, and the whole thing is a loop rather than a line. Each pass through it should be smaller and sharper than the last.

Framing it this way is the single most useful shift for anyone learning how to run a design of experiments for the first time. The question is not "which statistical test do I use", it is "which phase am I in, and what does that phase owe the next one".

The loop matters as much as the phases. Antony is explicit that experiments should be conducted iteratively so that information gained from one feeds the next, and that it is better to run several smaller sequential experiments than one large one that consumes the entire budget. A first pass separates the vital few factors from the trivial many. A second characterises them. A third maps the optimum. Confirmation closes each pass.

1 · PLAN response, factors, levels, interactions → factor list 2 · DESIGN choose design, randomize, block → design matrix 3 · CONDUCT run to the matrix, monitor, record → data sheet 4 · ANALYZE effects, ANOVA, reduce the model → predicted optimum 5 · CONFIRM run the optimum, test the interval → verified setting PASS → implement and control chart it FAIL, or a sharper question → iterate screen → characterise → optimize, ≤25% of budget in pass 1 The five phases of a DOE project Antony's four phases, plus confirmation, run as a loop rather than a line
Figure 1. The five phases and what each one hands to the next. The dashed return path is the point: a DOE project is a sequence of small experiments, not one big one.

Figure 1 shows five boxes in sequence. Phase 1 Plan produces a factor list from the response, factors, levels and interactions. Phase 2 Design produces a design matrix by choosing the design, randomizing and blocking. Phase 3 Conduct produces a data sheet by running to the matrix, monitoring and recording. Phase 4 Analyze produces a predicted optimum from effects, ANOVA and model reduction. Phase 5 Confirm produces a verified setting by running the optimum and testing the interval. A pass branch leads to implementing and control charting the process. A dashed fail or sharper-question branch loops back from Confirm to Plan, labelled screen then characterise then optimize, with no more than 25 per cent of budget in the first pass.

Table 1. The five steps in design of experiments, what each produces, and how each one fails.
PhaseCore decisionArtefact producedCommonest failure
1. PlanWhat are we measuring and what might move it?Problem statement, response + measurement plan, factor and level tableA key factor left out, or a response nobody can measure precisely
2. DesignWhich design, how many runs, what is confounded?Randomized design matrix with run orderResolution too low, so the effect you find is aliased with one you did not model
3. ConductCan we execute this exactly as written?Completed data sheet with actual setpoints recordedRun order quietly re-sorted for convenience; setpoints not achieved
4. AnalyzeWhich effects are real and what do they predict?Reduced model, ANOVA table, predicted optimumKeeping every term, or reading significance off the wrong error term
5. ConfirmDoes the prediction hold at the bench?Confirmation runs judged against a prediction intervalSkipped entirely, or judged against a confidence interval
Table 1. Each phase hands a concrete artefact to the next. If a phase cannot produce its artefact, that is the signal to stop rather than to push on.

Phase 1 — Plan

Planning defines the problem, the response and how it will be measured, the factors, their levels, and which interactions matter. It is the phase that decides whether the other four are worth doing. Antony's planning phase has six steps, and each maps onto a bioprocess decision.

Problem recognition and formulation. Write a statement containing a specific, measurable objective. "Improve the process" is not one. "Raise harvest titer from 2.8 to at least 3.5 g/L without dropping monomer below 96 per cent" is. Form the team at this point: process, analytical, manufacturing, and someone who will have to accept the recommendation.

Select the response. Use a continuous, quantitatively measurable characteristic, and define the measurement system before any runs. Antony is blunt that many DOE programmes fail because their responses cannot be measured quantitatively, and the arithmetic is unforgiving.

Why pass/fail responses are unaffordable

Antony's example: a process with a 0.5% defect rate yields about 5 defects per 1000 parts, so a 16-trial experiment needs roughly 16 × 1000 = 16,000 parts to carry usable information.

The bioprocess parallel is worse, because runs are not parts. Suppose the response is "did the batch contaminate", and contamination runs at 1%:

Nobody has that. The fix is not a bigger design, it is a better response: measure a continuous surrogate such as titer in g/L, monomer in per cent, or time-to-turbidity in hours, and treat the binary outcome as a downstream consequence.

Select and classify the process variables. Use process knowledge, historical data, cause-and-effect analysis and brainstorming to build the factor list, then split it into controllable factors and noise factors. Noise factors are the ones you cannot or will not control in production, such as ambient conditions, raw material lot, and operator. They do not disappear because you ignored them, and the three DOE principles exist to handle them: blocking for the ones you can identify and group, randomization and replication for the ones you cannot.

Determine the levels. Two levels are generally enough for quantitative factors in early experimentation. Use three or more only where you expect a non-linear response and need to quantify curvature, or add centre points to a two-level design to test for it cheaply. Qualitative factors such as medium lot or clone often need more than two levels by their nature.

List the interactions of interest. The number of two-factor interactions grows quadratically with the factor count:

N = n(n − 1) / 2

Four factors give 6 two-factor interactions, six give 15, and eight give 28. You cannot estimate all of them cheaply, so the planning phase has to say which ones you actually care about. In cell culture the usual suspects are temperature by feed rate, pH by osmolality, and feed rate by seed density. Everything else can be deliberately confounded in a first pass.

Figure 2. Two-factor interactions and full-factorial run count against the number of factors. Both curves are the argument for screening first. Hover any point for exact values.

Turn the factor list into a real design

Enter your factors and levels, see the run count and confounding before you commit, and export a randomized run order.

Open the DOE generator

Phase 2 — Design

The designing phase picks the design, fixes its size, and settles what will be confounded with what. Antony's guidance is to have the design matrix ready for the team before execution begins, showing every factor setting and the order runs will be performed in.

Design size follows from the number of factors, the number of levels, the interactions you decided to keep, and the budget. That last constraint is the one people skip, so make it explicit before choosing: it is easier to accept a resolution IV screen when you can see it costs 8 runs against a 40-run budget. A design selection guide covers the choice in detail, but the short version is that many factors call for a Plackett-Burman or fractional factorial screen, few factors call for a full factorial, and an optimum needs a response surface design.

Budgeting a 40-run bioreactor campaign

Antony advises investing no more than 25 per cent of the experimental budget in the first phase. With 40 available bench runs and five candidate factors:

Total committed 32 of 40. Compare that with spending all 40 on a single 25 full factorial replicated once: it answers one question, has nothing left for confirmation, and cannot recover from a single lost vessel.

Figure 3. The 40-run budget split across four sequential passes with contingency, against the all-in-one-experiment alternative. Sequential spending keeps a reserve and buys a confirmation.

The other job of this phase is to face the confounding structure honestly. A resolution III design aliases main effects with two-factor interactions, which is acceptable in a screen and dangerous anywhere else. Write down the alias structure before you run, so that when a factor comes out significant you already know what else it could have been.

Finally, decide the randomization strategy now. If a factor is expensive to reset, say so at this point and use a split-plot design rather than pretending you will fully randomize and then quietly grouping the runs during execution.

Phase 3 — Conduct

The conducting phase executes the matrix exactly as written and records what actually happened. It is the least intellectually interesting phase and the one that silently destroys the most studies.

Antony lists three practical prerequisites before execution: a location free of external noise sources, availability of the materials, operators and equipment for the whole campaign, and a cost-benefit check confirming the experiment is worth more than it costs. In bioprocess terms that means confirming the medium lot will last the whole study, the analytical method has capacity for every sample, and the suite is booked end to end rather than in fragments.

During execution, three rules do most of the work:

That last point is worth expanding. If the design called for pH 7.1 and the controller held 7.04, record 7.04. Coded designs assume the levels were hit; when they were not, an analysis using actual values is more honest and often more sensitive. It also catches the case where two nominally different levels were in practice the same, which turns a real factor into a wasted column.

Run order discipline deserves the same treatment. The randomized order exists to stop time-related drift, such as a slowly ageing seed train or a drifting analyzer, from loading onto a factor. Re-sorting the runs so all the low-temperature ones happen first is exactly the trap that randomization defends against, and no analysis can undo it after the fact.

Get a run order you can hand to the suite

Build the design, randomize the trial order, add replicates and centre points, then export the sheet.

Build a design of experiments

Phase 4 — Analyze

The analysing phase answers four questions: which factors move the mean, which move the variability, what settings give the optimum, and whether further improvement is possible. Those are Antony's four objectives for the phase, and the last one is the one that gets forgotten.

The working sequence is consistent regardless of software. Compute the effects and plot them, decide which are real, drop the rest, refit, and check the residuals. Main effects plots, interaction plots, cube plots and a Pareto of effects do most of the interpretive work; the normal probability plot of effects separates signal from noise without a formal test, and the normal probability plot of residuals checks the model assumptions. A full walkthrough of what each output means lives in reading DOE results.

Three cautions carry most of the risk in this phase.

  1. Reduce the model. Keeping every term because it is in the output inflates the standard errors and produces a prediction that fits the design points and forecasts badly. Drop non-significant high-order terms, respecting hierarchy: keep a main effect if it appears in a retained interaction.
  2. Use the right error term. If the design was blocked or split into whole plots, the residual your software prints by default may not be the correct denominator for every factor.
  3. Do not extrapolate. A model fitted between pH 6.9 and 7.1 says nothing about pH 7.4. If the optimum sits on a boundary, that is a signal to run another pass shifted in that direction, not to predict beyond the edge.

Interpretation is a separate skill from analysis, and Antony flags this as a real gap: many software systems stress data analysis without properly addressing data interpretation, leaving engineers who have run the statistics unsure how to use the result. The practical antidote is to state the conclusion in process language before writing it up. "Feed rate is the dominant factor; raising it from 3 to 5 per cent per day is worth about 0.6 g/L, and its benefit depends on pH" is a usable sentence. A table of p-values is not.

Phase 5 — Confirm

A fitted model predicts a response at settings nobody has actually run, so the last phase is to run them. This is the phase most often skipped, usually because the model looked convincing.

Confirmation runs sit outside the design and never feed back into the model fit. You pick the settings the reduced model recommends, run several independent replicates there, and compare the measured mean against a prediction interval computed in advance. Antony recommends between 4 and 20 confirmatory runs depending on how expensive each one is, and treats a conclusive result as the trigger for improvement action while an inconclusive one calls for further investigation before anything is implemented.

Two details decide whether the test is meaningful. The interval must be a prediction interval rather than a confidence interval, because a confirmation run is a future observation and carries run-to-run variability on top of the model's own uncertainty. And the interval must be written into the protocol before the runs happen, otherwise it is not a test. The mechanics, including the leverage term and how many runs actually buy you anything, are worked through in DOE confirmation runs.

A failed confirmation is a result, not an embarrassment. The usual causes, in rough order of how cheaply you can check them, are execution errors, measurement error, unmodelled curvature between the design points, a factor that was never in the study, and inadequate control of noise factors. Each one points at a specific next pass around the loop.

The barriers that sink DOE projects

Antony groups the obstacles to effective DOE as educational, management, cultural, communication, and tooling barriers, and none of them is statistical. Recognising which one you are facing is more useful than another course on ANOVA.

Table 2. Antony's barriers to DOE, translated into bioprocess development.
BarrierHow it shows up in a development groupWhat helps
EducationalFear of statistics; training that taught probability theory rather than DOE, so people default to one factor at a timeRun one small real study end to end. Competence follows a completed project, not a course
ManagementPressure for a quick answer; preference for home-grown fixes that deliver a short-term resultShow the run-count comparison against OFAT and commit to a first pass under 25% of budget
Cultural"That will not work for our process"; reluctance to plan before runningDemonstrate a successful application from a comparable process, then repeat it in house
CommunicationStatisticians choose implausible factor ranges; engineers misread interactions; nobody owns the measurement systemPut process, analytical and statistical people in the same planning session, once, before the design is fixed
ToolingSoftware analyses the data but gives no guidance on choosing an approach or interpreting the resultDecide the design and the acceptance criteria before opening any software
Table 2. Four of the five barriers are organisational. The fifth is a limitation of the tools, not of the method.

The educational barrier has a measurable signature worth naming. Antony observes that DOE training courses and textbooks often spend 70 to 80 per cent of their time on the analysis of experiments, while successful application requires a mixture of statistical, planning, engineering, communication and teamwork skills. The imbalance is the problem: the phase that receives the least attention in training is the one where projects are actually lost. If your team has run a DOE that produced a clean ANOVA table and changed nothing, the fault is almost never in the arithmetic.

Your DOE project checklist

Work through this before booking the suite. If you cannot answer an item, that is the next thing to do, not a reason to start running.

Before the design (Phase 1)

Before execution (Phase 2 to 3)

After the runs (Phase 4 to 5)

None of this is statistics. Almost every item is a planning or execution decision, which is the real answer to how to run a design of experiments: it is project management with an experimental design at the centre. Build the design in a free design of experiments calculator, keep each pass small, confirm before you implement, and let what you learn set the next question. That loop is the whole DOE methodology, and it works in a bioreactor suite for the same reason it works on a moulding machine.

Frequently Asked Questions

What are the steps in design of experiments?

Antony's DOE methodology divides the work into four phases: planning, designing, conducting, and analysing. In practice a fifth phase belongs on the end, confirmation, because a fitted model is a prediction until it is tested at the bench. Planning defines the problem, the response and its measurement system, the factors, their levels, and which interactions matter. Designing picks the design and fixes the run order. Conducting executes it with discipline. Analysing extracts and reduces the model. Confirming tests the predicted optimum against a prediction interval.

Which phase of a DOE project matters most?

Planning. Antony notes that DOE training courses and textbooks often spend 70 to 80 per cent of their time on the analysis of experiments, while success depends on a mixture of statistical, planning, engineering, communication and teamwork skills. No analysis recovers from a response that cannot be measured accurately, a factor left out of the study, or ranges set too narrow to move the response. Every hour spent in planning is cheaper than a run.

How much of the budget should the first DOE experiment use?

No more than 25 per cent. Antony advises against investing more than a quarter of the experimental budget in the first phase, because DOE is iterative and information from each experiment feeds the next. A worked 40-run bioreactor budget might spend 8 runs on screening (20 per cent), 12 on a factorial with centre points (30 per cent), 8 on augmenting to a response surface (20 per cent), 4 on confirmation (10 per cent), and hold 8 runs (20 per cent) back for contamination and repeats.

Why should a DOE response be continuous rather than pass or fail?

Because attribute data needs enormous sample sizes to carry the same information. Antony's example: a process with a 0.5 per cent defect rate produces about 5 defects per 1000 parts, so a 16-trial experiment would require roughly 16,000 parts. The bioprocess parallel is the same. If contamination runs at 1 per cent, seeing about 5 events per condition needs around 500 runs per condition, or 8,000 fermentations for a 16-run design. Measure titer in g/L or monomer in per cent instead.

How many interactions will a DOE have to consider?

The number of two-factor interactions is n(n − 1)/2, where n is the number of factors. Four factors give 6 two-factor interactions, six factors give 15, and eight factors give 28. This is why the planning phase asks explicitly which interactions are of interest, and why a screening design that deliberately confounds interactions is the right first move when the factor list is long.

What makes DOE projects fail in practice?

Antony groups the barriers as educational, management, cultural, communication, and tooling. In bioprocess development they usually surface as a specific handful: fear of statistics keeping people on one-factor-at-a-time, managers wanting a quick answer before the design finishes, factor ranges set too narrow to move the response, a measurement system noisier than the effects being chased, and software that analyses the data but never tells you what the result means.

Related Tools

References

  1. Antony, J. (2003). Design of Experiments for Engineers and Scientists. Butterworth-Heinemann/Elsevier. ISBN 0-7506-4709-4. (Ch. 4 systematic methodology, §4.2 barriers, §4.3.1–4.3.4 the four phases; ch. 8 practical tips, §8.1.5 continuous responses, §8.1.7 iterative experimentation and the 25% budget rule, §8.1.12 confirmatory runs.)
  2. Montgomery, D.C. (2017). Design and Analysis of Experiments, 9th ed. Wiley. ISBN 978-1119113478. (Guidelines for designing an experiment, and the sequential nature of experimentation.)
  3. Box, G.E.P., Hunter, J.S. & Hunter, W.G. (2005). Statistics for Experimenters, 2nd ed. Wiley. ISBN 978-0471718130. (Sequential assembly of designs and iterative learning.)
  4. Mandenius, C.-F. & Brundin, A. (2008). Bioprocess optimization using design-of-experiments methodology. Biotechnology Progress, 24(6), 1191–1203. doi:10.1002/btpr.67
  5. Politis, S.N., Colombo, P., Colombo, G. & Rekkas, D.M. (2017). Design of experiments (DoE) in pharmaceutical development. Drug Development and Industrial Pharmacy, 43(6), 889–901. doi:10.1080/03639045.2017.1291672

Resources & Further Reading