Two habits separate a design of experiments that survives scrutiny from one that quietly lies to you, and neither of them is a design type. Randomization and replication in DOE cost almost nothing to plan and are the first things dropped when the schedule tightens, which is exactly why so many screening studies end up chasing an effect that was never there. Randomization is your insurance against the noise you cannot name; replication is the only way to know how big your noise is in the first place. This guide shows what each one buys with worked numbers, untangles the replication vs repetition confusion that costs labs real bioreactor time, and sizes an experiment from signal-to-noise. The run-order shuffling itself is a solved problem, so you can hand it to a free design of experiments calculator and spend your attention on the parts that need judgement.
What randomization and replication in DOE actually do
Randomization protects your effect estimates from bias; replication tells you how much of what is left is noise. They solve different halves of the same problem, and doing one without the other leaves a hole. Along with blocking in design of experiments, they form the three principles Antony sets out in Design of Experiments for Engineers and Scientists as the ways to reduce or remove experimental bias. Large experimental bias, he notes, can produce wrong optimal settings or mask the effect of genuinely significant factors, so an opportunity for process understanding is simply lost.
| Principle | Problem it solves | Mechanism | What it costs |
|---|---|---|---|
| Randomization | Unknown, time-related noise biasing an effect | Shuffle the physical run order | Extra setup changes between runs |
| Replication | No estimate of experimental error | Reset and rerun trial conditions | Runs, material and time |
| Blocking | A known nuisance source inflating error | Group homogeneous runs, remove between-block variation | One high-order interaction |
The order in which you apply them matters and is easy to remember: block against what you can name, randomize against what you cannot, and replicate enough that the error bar you end up with is real. The rest of this guide takes the first two apart.
Randomization in design of experiments
Randomization in design of experiments means performing the trials in a random sequence rather than the order in which they are logically listed, so that uncontrolled variation is spread evenly across every factor instead of piling onto one. That is the whole mechanism. It does not remove noise, it removes the noise's ability to correlate with a factor.
Antony frames the reasoning bluntly: we live in a non-stationary world where noise factors never stay still. Machine parts wear, calibration slips, ambient temperature and humidity move, raw material varies lot to lot, and operator behaviour changes across a shift. None of these appear in your design matrix. All of them can move your response. By properly randomizing, you help average out the effects of noise factors that may be present in the process, and you ensure that all levels of a factor have an equal chance of being affected by them.
The quality engineer Dorian Shainin called randomization the experimenter's insurance policy, and warned that failure to randomize the trial conditions mitigates the statistical validity of an experiment. The insurance framing is exact: you pay a small, certain premium in setup time to avoid a rare but catastrophic loss, namely a whole study whose headline conclusion is an artefact.
There is one honest boundary condition worth stating. If you genuinely believe your process is stable, you do not strictly need to randomize. If you believe it is unstable, you must. And if it is so unstable that randomization would make the experiment impossible to execute, the correct response is not to run it unrandomized. It is not to run it at all yet, and instead to use process control methods to bring the process into a state of statistical control first.
Why randomize experiments? A drift that fakes an effect
Because in standard run order, a slow linear drift lines up almost perfectly with the slowest-changing factor column, and the entire drift is charged to that factor as if it were a real effect. This is not a subtle statistical worry. It is arithmetic, and it is worth watching happen once.
Take a 2³ screen: three factors, eight runs, executed back to back in the standard (Yates) order where A alternates every run, B every two runs and C every four. Now suppose something in the system declines slowly and steadily over the campaign. A dissolved-oxygen probe slowly fouls, a media lot ages on the bench, a feed pump head wears in. Say the effect is a modest −0.06 g/L of titer per run, so by run 8 the system is delivering 0.42 g/L less than at run 1 for reasons that have nothing to do with your factors.
Where does the drift land? Split the eight drift values by each column's signs and take the difference of the means. That is exactly how a two-level effect is estimated, so whatever the drift contributes there is added to the real effect.
| Effect column | Alternates every | Contamination in standard order (g/L) | Contamination if randomized (g/L) |
|---|---|---|---|
| A | 1 run | −0.06 | 0.00 ± 0.10 |
| B | 2 runs | −0.12 | 0.00 ± 0.10 |
| C | 4 runs | −0.24 | 0.00 ± 0.10 |
| AB, AC, BC, ABC | — | 0.00 | 0.00 ± 0.10 |
Read the third column carefully. In standard order the contamination is not random, it is deterministic: factor C is guaranteed to be handed −0.24 g/L of pure artefact, every time, and you have no way to detect it from the data because the interaction columns all come back clean and the design looks perfectly balanced. If C's true effect were +0.25 g/L, you would report +0.01 and conclude the factor does nothing. If its true effect were zero, you would report a 0.24 g/L effect and go optimize it.
Worked example: the drift that killed a feed-rate factor
An E. coli fed-batch team screens three factors in a 2³: induction OD600 (A), induction temperature (B), and feed rate (C). Eight runs, one bioreactor, eight consecutive days, executed in standard order. The DO probe slowly fouls across the campaign, costing 0.06 g/L per run.
- Run order: C is at its low level for runs 1 to 4 (mean drift −0.09 g/L) and at its high level for runs 5 to 8 (mean drift −0.33 g/L).
- Contamination: −0.33 − (−0.09) = −0.24 g/L charged to the feed-rate effect.
- True feed-rate effect: +0.30 g/L. Reported effect: +0.30 − 0.24 = +0.06 g/L, well inside the noise. Feed rate is dropped from the model as insignificant.
- Randomize instead (say C at its high level on runs 2, 3, 6 and 8): contamination falls to −0.03 g/L and the reported effect is +0.27 g/L. The factor survives.
- The cost of the fix: reordering a run sheet. The cost of skipping it: eight bioreactor runs and a real process lever thrown away.
Antony gives the same failure in a different industry. All the low levels of factor A are run first and all the high levels afterwards; during the experiment the workplace humidity changes by 50 per cent and moves the response. The analysis declares factor A statistically significant. In reality factor A is inert, and the humidity change caused the apparent effect. Randomization would have prevented the confusion. This is the same mechanism that makes one-factor-at-a-time experimentation so fragile: OFAT is standard order taken to its extreme, with every level of a factor run in one contiguous block of time.
Get a randomized run sheet, not a standard-order one
Build a factorial, fractional or screening design and get the physical run order shuffled for you, with replicates and blocks, free in the browser.
Replication: honest error bars for your effects
Replication means repeating an entire experiment, or a portion of it, under a fresh setup, and it buys two things: an estimate of the experimental error, and a more precise estimate of every factor and interaction effect. Without it you are estimating effects with no yardstick to judge them against.
The first property is the one people underrate. With a single unreplicated run per condition, there is no direct measurement of how much the response moves when nothing changes. You can back an error estimate out of high-order interactions or a normal probability plot, but you are assuming those interactions are zero rather than measuring anything. As Antony puts it, if the number of replicates is one, you cannot make satisfactory conclusions about the effect of either factors or interactions, because any apparent effect could be the result of experimental error. Replication turns that assumption into a measurement, which is the same "pure error" that center points in DOE provide at the middle of the design space. It is also the standard deviation that sets the width of the interval your DOE confirmation runs are eventually judged against, so under-replicating early makes the final verification harder, not easier.
The second property is precision. Increasing the number of replicates decreases the error variance, which shrinks the standard deviation used to estimate factor effects. Concretely, the standard error of a two-level effect estimate scales as 1/√N, so quadrupling the runs halves the error bar on every effect at once. That is a good deal only if the runs are honest replicates, which brings us to the distinction that trips up most teams.
Replication vs repetition (the costly confusion)
Replication requires resetting each trial condition from scratch; repetition takes several measurements under the same setup. The variation due to setup cannot be captured by repetition, so repetition produces an error estimate that is too small. Many process engineers are unsure of this difference, and it is the single most common way a DOE analysis ends up over-confident.
The consequence is asymmetric and unforgiving. Using repetition as if it were replication does not make your answer slightly optimistic, it inflates every F ratio by the ratio of the two variances. In Figure 2 that is (0.146 / 0.025)² ≈ 34, so an effect that is genuinely indistinguishable from noise can come back with a p value of less than 0.001. You then confirm nothing, scale up a setting that does not matter, and blame biology.
| Replication | Repetition | |
|---|---|---|
| What is reset | The entire trial condition, from scratch | Nothing. Same setup throughout |
| Variation captured | Setup + process + measurement | Measurement only |
| Error estimate | Valid experimental error | Systematically too small |
| Bioprocess example | Two independent bioreactor runs at the same setpoints | Two vials assayed from one bioreactor run |
| Analysis role | Denominator of the F test | Average first, then treat as one observation |
| Cost | High: reset time, material, schedule | Low: one extra assay |
Repetition is not useless, it is just not replication. Averaging several assay samples per run reduces measurement noise on each observation, which is genuinely worth doing when your assay is the weak link. The rule is simply that repeated measurements get averaged into a single response value for that run, and only independently set-up runs count toward the degrees of freedom in your error term. If you are unsure which of your numbers are which, the DOE output is where the mistake shows up: an implausibly small residual mean square alongside a long list of significant effects is the signature.
How much replication do you need?
Do not start from a replicate count. Start from the smallest effect you care about and the noise you already have, then let the arithmetic tell you the total run count. The common question before an experiment is how many runs are required to identify a significant effect given the current process variation, and it has a usable answer.
Antony gives a rule of thumb for total run count as a function of the signal-to-noise ratio:
where N is the total number of experimental runs, r is the number of levels per factor, Δ is the size of the effect you want to detect, and σ is the noise level. The derivation targets roughly 90 per cent confidence of finding an active effect of size Δ. For two-level factors this collapses to N = 64 × (σ/Δ)², which is easy to do in your head.
The signal is the change in response you want to detect, and you have to decide the smallest change worth acting on. The noise is the random variation the response shows under standard operating conditions, estimated either from a control chart (σ = R̄/d2) or from the root mean square error in the ANOVA table of a previous designed experiment.
| Signal-to-noise (Δ/σ) | Minimum total runs, N | Detectable effect at σ = 0.20 g/L | A design that fits |
|---|---|---|---|
| 1.0 | 64 | 0.20 g/L | 25 full factorial × 2 replicates |
| 1.4 | 32 | 0.28 g/L | 25−1 × 2 replicates |
| 2.0 | 16 | 0.40 g/L | 24 full factorial, unreplicated |
| 2.8 | 8 | 0.56 g/L | 23 or an 8-run screening design |
Figure 3. Run count against signal-to-noise for two-level factors. The curve is an inverse square, which is why small effects are so expensive and why halving your process noise is worth four times as much as doubling your run budget.
Worked example: sizing a CHO titer screen
A CHO process has a run-to-run titer standard deviation of σ = 0.20 g/L, measured as the RMSE of last quarter's designed experiment. Five factors are candidates. Two levels each, so r = 2 and N = 64 × (σ/Δ)².
- Detect Δ = 0.40 g/L: N = 64 × (0.20/0.40)² = 64 × 0.25 = 16 runs. A 25−1 fractional factorial with no replication does it.
- Detect Δ = 0.30 g/L: N = 64 × (0.20/0.30)² = 28 runs, so round up to a 32-run plan (16-run design × 2 replicates).
- Detect Δ = 0.20 g/L: N = 64 × 1 = 64 runs. Halving the target effect from 0.40 to 0.20 g/L quadrupled the bioreactor time.
- The alternative: cut σ from 0.20 to 0.14 g/L by fixing the assay and standardizing the inoculum, and 0.20 g/L becomes detectable in 32 runs instead of 64. Reducing noise beats buying runs.
If you cannot afford the runs the formula asks for, the honest options are to accept a larger detectable effect, reduce the noise, or reduce the factor count. Running the smaller experiment anyway and hoping is not one of them.
Two practical constraints sit on top of the arithmetic. Replication can substantially increase the calendar time of a study, and if material is expensive it can dominate the budget, so its use has to be justified in time and cost terms like anything else. The counterweight is that any bias or experimental error associated with setup changes gets evenly distributed across the runs when you replicate properly. For the full treatment of run counts across every design family, see how many experiments a DOE needs.
Lay out the replicates and the run order together
Pick a design, set the number of replicates and center points, and export a randomized run sheet you can take straight to the bench.
Randomization, blocking and restricted randomization
Randomization is not always fully achievable, and the correct response is to restrict it deliberately rather than abandon it quietly. Bioprocessing is full of setpoints that are slow or expensive to change. Temperature in a chemical process, as Antony notes, may be a hard-to-change factor that makes complete randomization almost impossible.
Under those circumstances it may be desirable to change the levels of that factor less frequently than the others, which is what restricted randomization means: group the runs by the hard-to-change factor's level, and randomize freely inside each group. The structure that results is a split-plot design, and it carries two error terms rather than one. The hard-to-change factor is tested against the between-group error, the easy factors against the within-group error. Analysing it as if it had been fully randomized understates the error on the hard-to-change factor and will overstate its significance, which is the most common analysis error in bioreactor DOE. The topic gets a fuller treatment in DOE for cell culture and fermentation.
Before deciding how far to randomize, Antony suggests running through five questions:
- What is the cost associated with a change of factor levels?
- Have we incorporated any noise factors into the experimental layout?
- What is the setup time between trials?
- How many factors in the experiment are expensive or difficult to control?
- Where do we assign factors whose levels are difficult to change from one level to another?
Blocking sits alongside both habits rather than competing with them. Assign runs to blocks first so each block is a homogeneous set (one media lot, one day, one seed expansion), then randomize the sequence within each block, then decide how many complete replicates of the block structure you can afford. Where the blocks are defined by inoculum batches, planning the expansion so each batch cleanly covers a block is a scheduling problem a seed train planner can lay out for you before the design is locked.
Randomize your run order free
None of this requires software you have to buy. Choose the design, set the replicate count from the signal-to-noise arithmetic above, add blocks if you have known nuisance seams, and let the tool emit the shuffled run order. The DOE generator and randomizer handles the layout, replication and randomization in the browser, and pairs naturally with a fed-batch calculator when the factor levels you are randomizing are feed rates and induction points that need to be worked out first.
If anyone on the team still asks why randomize experiments at all when the run sheet is already balanced, Table 2 is the answer in one line: balance on paper does not protect you from drift in time. The discipline is worth more than the tooling. Print the randomized order, follow it even when a different order would be more convenient, record the actual execution sequence in case it slips, and never let a repeated assay masquerade as a replicated run. Those three habits protect an experiment better than any upgrade in design sophistication.
Frequently Asked Questions
What are randomization and replication in DOE?
Randomization and replication are two of the three basic principles of design of experiments, alongside blocking. Randomization means performing the experimental trials in a random order rather than the order in which they are logically listed, so that slow drifts and unknown noise factors cannot line up with any one factor column. Replication means repeating an entire experiment or a portion of it under a fresh setup, which gives an estimate of the experimental error and a more precise estimate of every factor and interaction effect.
Why should you randomize experiments?
Because time-related noise is real and invisible. Probes drift, media ages, ambient conditions move, and operators get faster. If runs are done in standard order, the slowest-changing factor column lines up almost perfectly with that drift, so the drift is silently added to that factor's effect estimate. In the worked example on this page, a 0.06 g/L per run decline loads a fake 0.24 g/L effect onto factor C in standard order. Randomizing makes the expected contamination zero for every column and pushes what remains into the residual, where it widens the error bar honestly instead of manufacturing a confident wrong answer.
What is the difference between replication and repetition?
Replication requires resetting each trial condition from scratch, so the setup is torn down and rebuilt between measurements. Repetition takes several measurements under the same setup. Repetition therefore cannot capture the variation caused by setup, and it produces an error estimate that is too small, which makes trivial effects look statistically significant. In bioprocessing, two vials assayed from the same bioreactor run are repetition; two independent bioreactor runs at the same factor settings are replication.
How many replicates does a DOE need?
Size the whole experiment against the signal-to-noise ratio rather than picking a replicate count first. A common rule of thumb for two-level factors is N = (4r)² × (σ/Δ)², where r is the number of levels, Δ is the smallest effect you want to detect and σ is the run-to-run standard deviation. At two levels this gives 64 runs to detect an effect the size of one standard deviation, 16 runs at twice the standard deviation, and 8 runs at 2.8 times it. Divide the total run count by the size of your base design to get the number of replicates.
Can you skip randomization if a factor is hard to change?
You do not skip it, you restrict it. When a factor such as bioreactor temperature is expensive or slow to reset, complete randomization may be impractical, so restricted randomization is used: the hard-to-change factor level is changed less frequently, and the remaining runs are randomized within each of those groups. That structure is a split-plot design and must be analysed with two error terms. Treating a restricted-randomization experiment as if it had been fully randomized understates the error on the hard-to-change factor and overstates its significance.
Does replication help if my process is unstable?
Replication measures instability, it does not cure it. If the process is so unstable that randomization would make the experiment meaningless, the right move is to bring the process into statistical control first using process control methods, then run the DOE. Replicating an out-of-control process simply buys a very precise estimate of a very large error, and the run count needed to detect a useful effect grows with the square of the noise.
Related Tools
- DOE Experiment Generator — Build a design, set replicates and center points, and export a randomized run order.
- Fed-Batch Calculator — Work out the feed rates and induction points that become the factor levels you randomize.
- Seed Train Planner — Plan inoculum expansions so each seed batch forms a clean block before you randomize within it.
References
- Antony, J. (2003). Design of Experiments for Engineers and Scientists. Butterworth-Heinemann/Elsevier. ISBN 0-7506-4709-4. (§2.2.1 randomization; §2.2.2 replication; §8.1.8 randomize the trial order; §8.1.9 replicate to dampen noise, Eq. 8.1 and Table 8.2.)
- Box, G.E.P., Hunter, J.S. & Hunter, W.G. (2005). Statistics for Experimenters, 2nd ed. Wiley. ISBN 978-0471718130. (Ch. 3–4, randomization and replication.)
- Montgomery, D.C. (2017). Design and Analysis of Experiments, 9th ed. Wiley. ISBN 978-1119113478. (Ch. 1 and 5, the three principles and replication in factorials.)
- Ganju, J. & Lucas, J.M. (1997). Bias in test statistics when restrictions in randomization are caused by factors. Communications in Statistics – Theory and Methods, 26(1), 47–63. doi:10.1080/03610929708831901
- NIST/SEMATECH (2012). e-Handbook of Statistical Methods, Section 5.1.3: What are the steps of DOE? itl.nist.gov