Blocking in Design of Experiments: Removing Nuisance Variation

July 2026 12 min read Bioprocess Engineering

Key Takeaways

Contents

  1. What is blocking (and why)?
  2. Nuisance factors vs factors of interest
  3. Randomized complete block design
  4. Blocking a factorial (confounding a block)
  5. Blocks in bioprocess DOE
  6. Block, then randomize
  7. Set up blocks free
  8. Frequently Asked Questions

Your experiment will not finish in one sitting. Half the runs go on Monday's media lot, half on Tuesday's; two bioreactors do the work, not one; two analysts read the plates. Every one of those seams adds variation that has nothing to do with the factors you are studying, and if you ignore it, that variation lands in your error term and buries the effects you came to find. Blocking in design of experiments is the fix: you group the runs so each known nuisance source is contained, measured, and subtracted out. This guide explains what blocking does, how a randomized block design works, and the neat trick for blocking a factorial without sacrificing the effects that matter, with a worked bioprocess example. To lay out and randomize the runs, you can use a free design of experiments calculator.

What is blocking (and why)?

Blocking is a method of eliminating the effect of extraneous variation due to nuisance factors, thereby improving the efficiency of an experimental design. The idea, as Antony states it in Design of Experiments for Engineers and Scientists, is to arrange similar experimental runs into blocks (or groups), where a block is a set of relatively homogeneous experimental conditions. Observations collected under the same conditions, same day, same shift, same lot, are said to be in the same block. Crucially, the variability between blocks is then removed from the experimental error, which increases the precision of the experiment.

That last sentence is the whole payoff. Total variation in your data splits into three buckets: the factor effects you want, the block-to-block nuisance you can name, and the leftover random error. Ignore the middle bucket and it merges into error, inflating your noise floor. Name it as a block and it is pulled out cleanly, shrinking the error against which every effect is tested. Smaller error means smaller effects become statistically detectable, from the same runs.

How blocking moves nuisance variation out of the residual error Without blocking Factor effects Residual error (hides day-to-day nuisance) block it With blocking Factor effects Block (day) Residual (smaller)
Figure 1. Blocking does not remove variation from the experiment, it relabels the known nuisance part so it stops inflating the residual error. Effects are then tested against a smaller noise floor.

Nuisance factors vs factors of interest

A factor of interest is a variable you want to estimate the effect of; a nuisance factor is a known source of variation you want to remove but not study. Temperature, pH, and feed rate are usually factors of interest. The calendar day, the specific bioreactor, the raw-material lot, and the operator are usually nuisance factors, they move the response, but you would never write a report conclusion about "the effect of Tuesday."

The distinction decides how you treat the variable. Antony gives two clean examples. In the first, a metallurgist studies four factors at two levels in an eight-run experiment but can only run four trials per day, so each day is treated as a separate block to reduce day-to-day variation. In the second, a chemical process needs two batches of raw material for the full run set, so each batch of raw material is treated as a block to minimize batch-to-batch variability. In both cases the block is a known, unavoidable seam, not something under study.

Randomized complete block design

A randomized complete block design (RCBD) runs every treatment once inside every block, then randomizes the order within each block. "Complete" means each block holds a full replicate of all treatments; "randomized" means the sequence inside the block is shuffled to defend against sub-trends. Because every treatment appears in every block, the block-to-block shift cancels out of every treatment-vs-treatment comparison.

Picture comparing three media formulations across three days, one full set per day. Whatever makes Wednesday run high or low affects all three formulations equally on Wednesday, so it drops out when you compare formulation A to B to C. The RCBD is the simplest blocked design and the right choice whenever a single nuisance factor, day, batch, plate, can be held constant across a complete set of runs.

Blocking a factorial (confounding a block)

When a factorial has more runs than fit in one homogeneous block, you split it across blocks by deliberately confounding the block difference with a high-order interaction you are willing to give up. This is the elegant core of blocking a factorial design, and it is worth seeing once in full.

Take a 2³ factorial: three factors, eight runs. Suppose only four runs fit on one media lot, so you need two blocks of four. You cannot just cut the eight runs in half arbitrarily, that would tangle the block with your main effects. Instead you use the highest-order interaction, ABC, as the block generator: runs where the ABC column is minus go in block 1, runs where it is plus go in block 2.

Table 1. Blocking a 2³ factorial by confounding the block with the ABC interaction. Note A, B, C each balance (two +, two −) within each block.
RunABCABCBlock
11
4++1
6++1
7++1
2++2
3++2
5++2
8++++2

The magic is in the last two columns. Because the block boundary follows the ABC sign, only the ABC interaction is now confounded with the block, you can no longer separate "a real three-factor interaction" from "a day-to-day shift." But that is a cheap price: three-factor interactions are rarely important and hardest to interpret anyway. Meanwhile A, B, C and all three two-factor interactions (AB, AC, BC) stay perfectly balanced within each block, so they are estimated as cleanly as if there were no block at all.

This is the one rule Antony underscores for choosing your block generator: ensure the block generators are not confounded with the main effects and two-factor interaction effects. Sacrifice a high-order interaction, never a main effect. Standard tables (Box, Hunter and Hunter, 1978) list the recommended block generators, block sizes, and resulting resolutions for larger designs. The mechanics are the same family of ideas as aliasing in a fractional factorial design, applied to the block instead of to a fraction, and it builds directly on the effects and interactions covered in the full factorial design guide.

Worked example: two media lots, one clean answer

A CHO titer study runs a 2³ on glucose (A), temperature (B), and feed rate (C). Only four bioreactor runs fit per media lot, and lot 2 happens to run +0.5 g/L higher across the board, a pure nuisance shift. Split the eight runs by the ABC generator (Table 1) so lot 1 = block 1 and lot 2 = block 2.

Figure 2. Illustrative residual error for the worked 2³ example. Blocking the media lot out of the error term shrinks the noise floor every effect is tested against.

Build and split the design automatically

Generate a factorial or screening design, assign blocks, and get a randomized run order per block, free in the browser, no confounding math by hand.

Open the free DOE generator →

Blocks in bioprocess DOE (day, bioreactor, media lot)

Bioprocessing is unusually rich in nuisance factors, which makes blocking one of the highest-value habits in a bioprocess DOE. The usual suspects to block on:

Table 2. Common bioprocess nuisance factors and how to block them.
Nuisance factorWhy it variesHow to block it
Day / campaignAmbient, calibration driftOne complete replicate per day
Bioreactor / well positionUnit-to-unit, edge effectsBalance treatments across units
Raw-material / media lotSupplier lot chemistryTreat each lot as a block
Analyst / assay plateOperator, plate-to-plateBalance across operators & plates
Seed-train batchInoculum age & densityGroup runs from one inoculum

The seed-train case is especially common: a single expansion feeds only so many bioreactors, so runs from different inoculum batches carry a built-in seam. Planning that expansion so each seed batch forms a clean block, rather than straddling your factor combinations, is exactly the kind of logistics a seed train planner helps you lay out. And when a nuisance factor is not merely a batch effect but a genuinely hard-to-change setpoint (reactor temperature you cannot reset every run), blocking shades into a related design discussed in DOE for cell culture and fermentation.

Block, then randomize

Blocking and randomization are partners, not alternatives: block against the nuisance you can name, randomize against the nuisance you cannot. Blocking removes a known source of variation; randomization insures against unknown, lurking ones, machine ageing, humidity drift, a reagent slowly degrading. You want both.

The order matters. First assign runs to blocks using the generator, then randomize the run sequence within each block. That way the block structure stays intact (each block is still a homogeneous set) while the order inside it is shuffled so any within-block trend cannot line up with a factor. Randomizing across blocks would destroy the very homogeneity you built. A good DOE tool does both steps for you, emitting a per-block randomized run order you can take straight to the bench.

Set up blocks free

You do not need to work out block generators by hand. Build your factorial or screening design in a free tool, tell it how many blocks you need, and it will confound the block with the appropriate high-order interaction and hand back a randomized run order per block. Use the DOE generator and randomizer for the layout, and plan the physical batches, seed expansions, and lots that define your blocks alongside it. The result is an experiment whose real effects are protected from the day, the lot, and the operator, before you have run a single flask.

Frequently Asked Questions

What is blocking in design of experiments?

Blocking is a method of eliminating the effect of extraneous variation caused by nuisance factors, by grouping experimental runs into blocks of relatively homogeneous conditions. A block is a set of runs done under conditions as similar as possible (same day, same raw-material lot, same operator). The variation between blocks is estimated separately and removed from the experimental error, which increases the precision of the experiment so that real factor effects are easier to detect.

What is the difference between a blocking factor and a factor of interest?

A factor of interest is a variable you are studying and want to estimate the effect of, such as temperature or feed rate. A nuisance (blocking) factor is a known, controllable source of variation that you are not interested in but that would otherwise inflate your error, such as the day, the bioreactor unit, the media lot, or the operator. You block on the nuisance factor to remove its variation; you never draw conclusions about the block itself.

What is a randomized complete block design?

A randomized complete block design (RCBD) runs every treatment once within every block, and randomizes the run order inside each block. Because each block contains a full set of treatments, the nuisance variation between blocks is balanced out of every treatment comparison. It is the simplest and most common blocked design, used when one nuisance factor (like day or batch) can be held constant within each group of runs.

How do you block a factorial design without losing effects?

When a factorial has too many runs to fit in one block, you split it across blocks by deliberately confounding the block difference with a high-order interaction you are willing to sacrifice, usually the highest-order interaction (for a 2³, the three-factor ABC interaction). The block generator is that interaction: runs where it is minus go in one block, plus in the other. Main effects and low-order interactions stay orthogonal to the block, so they are estimated cleanly. The rule is to never confound the block with a main effect or an important two-factor interaction.

Should I block or randomize?

Do both. Block what you can control and know about (day, batch, operator, machine); randomize the run order within each block to defend against the nuisance factors you do not know about. The classic advice is: block against what you can, randomize against what you cannot. Blocking removes a known source of variation from the error; randomization insures against unknown, lurking sources.

Related Tools

References

  1. Antony, J. (2003). Design of Experiments for Engineers and Scientists. Butterworth-Heinemann/Elsevier. ISBN 0-7506-4709-4. (§2.2.3 blocking; §8.1.10 blocking strategy; §2.4 confounding.)
  2. Box, G.E.P., Hunter, W.G. & Hunter, J.S. (1978). Statistics for Experimenters. Wiley. ISBN 978-0471093152.
  3. Montgomery, D.C. (2017). Design and Analysis of Experiments, 9th ed. Wiley. ISBN 978-1119113478. (Ch. 7, blocking and confounding in factorials.)
  4. NIST/SEMATECH (2012). e-Handbook of Statistical Methods, Section 5.3.3.3: How can we block a design? itl.nist.gov

Resources & Further Reading