Nanobody and Antibody Fragment Production: VHH, Fab, and scFv Expression Systems, Purification, and Bioprocess Development

September 2026 18 min read Bioprocess Engineering

Key Takeaways

Contents

  1. Antibody Fragment Formats Compared
  2. Expression Systems for Nanobody Production
  3. E. coli Expression: Periplasm vs Cytoplasm
  4. Pichia pastoris: The High-Yield Secretion Platform
  5. The Disulfide Bond Problem
  6. Why Do scFv Constructs Aggregate?
  7. Purification Without Protein A
  8. Bioprocess Scale-Up Considerations
  9. Frequently Asked Questions

Nanobody production has emerged as one of the fastest-growing areas of biopharmaceutical manufacturing. With over 200 VHH-based candidates in clinical pipelines and the commercial success of caplacizumab (the first approved nanobody drug), process development scientists face a practical question: which expression system and purification strategy best fits their antibody fragment format? This guide covers the bioprocess engineering of VHH nanobodies, Fab fragments, and scFv constructs, from host selection through downstream processing to manufacturing scale-up.

Unlike conventional monoclonal antibodies (~150 kDa) that require mammalian expression systems for proper folding and glycosylation, antibody fragments are small enough (12–50 kDa) to fold correctly in microbial hosts. This fundamental advantage makes nanobody production significantly cheaper and faster than full-length antibody manufacturing. However, each fragment format presents its own bioprocess challenges, particularly around disulfide bond formation and aggregation.

Antibody Fragment Formats Compared

Five antibody fragment formats dominate current pipelines, each with distinct size, disulfide bond requirements, and expression complexity. The choice of format directly determines which expression system and purification strategy is feasible.

Table 1. Antibody Fragment Format Comparison
Format Size (kDa) Disulfide Bonds Valency Serum Half-Life Key Applications
VHH / Nanobody12–151 (intradomain)Monovalent1–2 hDiagnostics, imaging, intracellular targets
Fab~502 (1 per domain + 1 interchain)Monovalent12–20 hTherapeutics (ranibizumab), biosimilars
scFv25–302 (intradomain)Monovalent2–4 hCAR-T targeting domains, bispecific building blocks
Diabody50–602 (intradomain)Bivalent3–5 hImaging (short linker forces dimerization)
Minibody~804Bivalent5–12 hImmuno-PET imaging
Fragment formats ranked by increasing molecular weight. All lack the Fc region required for Protein A binding.

VHH nanobodies are the simplest fragment to produce because they contain only a single immunoglobulin variable domain with one intradomain disulfide bond. Their small size (2.5 nm diameter, 4 nm height) enables tissue penetration that full-length antibodies cannot achieve, making them particularly valuable for solid tumor imaging and intracellular targeting applications.

Fab fragments retain both the variable and first constant domains of heavy and light chains, giving them higher thermal stability and longer serum half-life than VHH. However, the interchain disulfide bond between heavy and light chain constant domains adds expression complexity: both chains must be co-expressed and correctly paired.

Antibody Fragment Expression System Decision Tree Which fragment format? VHH / Nanobody Fab scFv Need > 500 mg/L? Yes Pichia pastoris 0.1–3 g/L secreted No E. coli SHuffle 200–300 mg/L cyto Need glycosylation? Yes CHO / HEK293 0.5–2 g/L No E. coli periplasm 50–500 mg/L Aggregation-prone? Yes Refolding from IB 50–200 mg/L net No E. coli periplasm 10–100 mg/L Purification Strategy (No Protein A) IMAC (His-tag) 80–90% yield, 80% purity Protein L (kappa) 75–85% yield, 90% purity CEX (tag-free) 70–80% yield, 70% purity Anti-VHH affinity 85–95% yield, 95% purity All fragments lack Fc → Protein A is NOT an option Cost ranking ($/g): E. coli cyto $40–100 (cheapest) Pichia $30–80 E. coli peri $200–500 (lower yield)
Figure 1. Expression system decision tree for antibody fragments. Fragment format, glycosylation requirement, and aggregation propensity determine the optimal host. All formats require purification strategies that bypass Protein A.

Decision tree showing that VHH nanobodies can use E. coli SHuffle cytoplasm (200-300 mg/L) or Pichia pastoris secretion (0.1-3 g/L) depending on yield needs; Fab fragments use E. coli periplasm (50-500 mg/L) or CHO/HEK293 (0.5-2 g/L) if glycosylation is needed; scFv uses E. coli periplasm (10-100 mg/L) or inclusion body refolding (50-200 mg/L net) depending on aggregation propensity. Purification options include IMAC, Protein L, CEX, and anti-VHH affinity.

Expression Systems for Nanobody Production

The choice of expression system for nanobody production is primarily driven by yield requirements and cost targets. E. coli remains the workhorse for early-stage discovery and research-grade material, while Pichia pastoris is increasingly adopted for clinical and commercial manufacturing where gram-per-liter yields are essential.

Table 2. Expression System Comparison for VHH Nanobody Production
Host System Compartment Typical Yield Disulfide Bonds Timeline Est. Cost ($/g)
E. coli BL21(DE3) periplasmPeriplasm10–100 mg/LOxidizing (DsbA/DsbC)3–5 days200–500
E. coli SHuffle T7 ExpressCytoplasm200–300 mg/LOxidizing (trxB/gor mutant)3–5 days40–100
E. coli Origami 2(DE3)Cytoplasm50–100 mg/LOxidizing (trxB/gor mutant)3–5 days100–250
Pichia pastoris (AOX1)Secreted0.1–3 g/LER oxidizing5–8 days30–80
CHO (transient)Secreted50–200 mg/LER oxidizing7–14 days1,000–3,000
Cell-free (PURE/S30)In vitro0.1–5 mg/mLRequires additive4–24 hours5,000–15,000
Yield ranges represent literature-reported values across multiple VHH sequences under optimized conditions. Cost estimates are for purified material at research scale (1–10 L).
Figure 2. VHH nanobody yield (mg/L) and estimated production cost ($/g) across six expression systems. Pichia pastoris achieves the best combination of high yield and low cost at manufacturing scale.

E. coli Expression: Periplasm vs Cytoplasm

E. coli is the default starting point for nanobody production because of its fast doubling time (20–30 min), well-established molecular biology tools, and low media cost. The critical decision is whether to target the periplasm or the cytoplasm, and this choice hinges on how the single intradomain disulfide bond in the VHH will form.

Periplasmic expression

Periplasmic expression uses an N-terminal signal peptide (pelB, OmpA, or DsbA) to translocate the nascent VHH across the inner membrane via the Sec or SRP pathway. The periplasm provides an oxidizing environment where the DsbA/DsbC foldase system catalyzes disulfide bond formation. Typical yields range from 10–100 mg/L, limited by the small volume of the periplasmic space (~8–16% of total cell volume) and Sec translocon capacity.

Cytoplasmic expression in engineered strains

E. coli SHuffle T7 Express is a strain engineered with mutations in the thioredoxin reductase (trxB) and glutathione reductase (gor) genes, making the cytoplasm sufficiently oxidizing for disulfide bond formation. It also constitutively expresses DsbC in the cytoplasm, a disulfide bond isomerase that corrects mispaired disulfides. This eliminates the translocation bottleneck, and yields of 200–300 mg/L soluble VHH are routinely achieved.

Worked Example: VHH Expression in SHuffle T7 Express

Setup: Anti-GFP VHH cloned into pET-28b (T7 promoter, N-terminal His6 tag), transformed into SHuffle T7 Express.

Volumetric productivity = 180 mg / (1 L × 24 h) = 7.5 mg/L/h
Compare: E. coli BL21 periplasmic = 50 mg / (1 L × 24 h) = 2.1 mg/L/h

Pichia pastoris: The High-Yield Secretion Platform

Pichia pastoris (Komagataella phaffii) achieves 6–14 times higher nanobody titers than E. coli periplasmic expression by secreting functional VHH directly into the culture medium, bypassing the periplasmic bottleneck entirely. The methanol-inducible AOX1 promoter combined with the alpha-mating factor (αMF) secretion signal is the standard configuration.

Secretion into the culture medium offers three process advantages over intracellular E. coli expression:

  1. Simplified harvest: Cell removal by centrifugation yields a clarified supernatant containing the VHH as the dominant protein (70–90% of total secreted protein), eliminating the need for cell lysis.
  2. Correct disulfide bonds: VHH passes through the ER where protein disulfide isomerase (PDI) catalyzes disulfide bond formation under oxidizing conditions.
  3. Product homogeneity: Pichia-derived VHH elutes as a single peak in SEC, compared to the pre-peak and post-peak shoulders often seen with E. coli material (aggregates and degradation products).

The key process parameters for Pichia nanobody production are methanol feed rate (typically 3–6 g/L/h), dissolved oxygen (>20% air saturation to support methanol oxidation), and pH (5.5–6.0 for protease minimization). Fed-batch fermentation over 72–120 hours with a glycerol growth phase (24–36 h to 100–150 g/L wet cell weight) followed by a methanol induction phase (48–84 h) typically achieves 0.5–1.5 g/L VHH in research-scale bioreactors, with optimized processes reaching 3 g/L.

Methanol-free expression systems using the GAP promoter or engineered AOX1 variants are gaining traction, particularly for GMP manufacturing where flammable methanol handling adds facility complexity and cost. GAP-driven constitutive expression simplifies the process to a single growth phase but typically yields 30–50% less than methanol-induced systems.

The Disulfide Bond Problem

Every antibody fragment requires at least one intradomain disulfide bond for structural stability, and this single requirement drives most of the complexity in fragment expression. A VHH nanobody has one disulfide bond connecting the B and F beta-strands of the immunoglobulin fold (Cys22–Cys92 in Kabat numbering). Without this bond, the VHH unfolds and aggregates within minutes at 37 °C.

Table 3. Disulfide Bond Requirements and Engineering Solutions
Fragment Required Disulfide Bonds E. coli Cytoplasm (wild-type) E. coli SHuffle/CyDisCo Periplasm/Secretion
VHH1 intradomainFails: reducingWorks: 200–300 mg/LWorks: 10–100 mg/L
Fab2 intradomain + 1 interchainFailsDifficult: chain pairingWorks: 50–500 mg/L
scFv2 intradomainFailsVariable: 50–200 mg/LVariable: 10–100 mg/L
VHH with extra Cys1 intradomain + 1 extraFailsRequires optimizationWorks, but extra bond may mispair
The reducing E. coli cytoplasm (glutathione redox potential approximately −240 mV) cannot form disulfide bonds. SHuffle (trxB/gor double mutant) shifts the redox potential to approximately −205 mV, sufficient for most single-domain fragments.

An alternative to engineered strains is the CyDisCo system, which co-expresses Erv1p (a sulfhydryl oxidase from yeast mitochondria) and human PDI in the E. coli cytoplasm. CyDisCo can be used in any E. coli strain, including high-density fermentation workhorses like BL21(DE3), and has achieved 100–400 mg/L soluble VHH depending on the construct.

Why Do scFv Constructs Aggregate?

scFv aggregation is the single largest barrier to scalable antibody fragment production. Unlike VHH nanobodies (which are inherently stable as single domains), scFv constructs link two variable domains (VH and VL) with a flexible peptide linker, creating multiple opportunities for misfolding.

Three mechanisms drive scFv aggregation in E. coli:

  1. Intermolecular domain swapping. The VH domain of one scFv molecule pairs with the VL domain of a neighboring molecule instead of its own, forming dimers, trimers, and higher-order oligomers. A (Gly4Ser)3 linker (15 residues) minimizes this by keeping the two domains in close proximity, but shorter linkers (5–10 residues) deliberately promote dimerization to create diabodies.
  2. Hydrophobic exposure during folding. The VH–VL interface is normally buried, but during folding in the crowded periplasm (protein concentration 200–400 mg/mL), partially folded intermediates expose hydrophobic patches that nucleate aggregation.
  3. Disulfide scrambling. Each scFv has two intradomain disulfide bonds (one in VH, one in VL). If one bond forms before the domain is fully folded, it can lock the domain in a non-native conformation that exposes aggregation-prone regions.

Practical mitigation strategies include lowering induction temperature to 16–20 °C (slows folding rate, allowing each domain to reach native conformation before the next intermediate appears), co-expressing periplasmic chaperones (Skp, FkpA, SurA), and engineering more soluble frameworks (e.g., the Herceptin scFv framework has been optimized for E. coli solubility across hundreds of CDR grafts).

When periplasmic expression fails, deliberate inclusion body (IB) formation followed by in vitro refolding can recover 30–60% of functional scFv. The refolding protocol typically uses 6–8 M urea solubilization, rapid dilution into a redox-controlled refolding buffer (0.5 mM oxidized / 5 mM reduced glutathione), and 0.4 M L-arginine to suppress aggregation during the folding process.

Purification Without Protein A

The absence of an Fc region in all antibody fragments means that Protein A chromatography, the universal mAb capture step, is not available. This fundamentally changes the downstream processing strategy, requiring format-specific alternatives that are typically less selective and more expensive per gram of product.

Figure 3. Purification strategy comparison for antibody fragments lacking Fc. IMAC provides the best yield but moderate purity; Protein L delivers high purity for kappa chain-bearing fragments; CEX is the only option for tag-free nanobodies at scale.

IMAC (immobilized metal affinity chromatography)

IMAC capture of His6-tagged fragments is the fastest path from cell lysate to purified product. Ni-NTA or Co-TALON resins bind the polyhistidine tag with high affinity (Kd ~10-13 M), and step elution with 250–300 mM imidazole typically yields 80–90% recovery at 75–85% purity. A subsequent SEC polishing step removes aggregates and co-purifying E. coli proteins that contain natural histidine clusters.

For therapeutic applications, the His-tag must be removed. A TEV or HRV 3C protease cleavage site between the tag and the VHH adds a processing step but leaves a native N-terminus. Subtractive IMAC (passing the cleaved pool over Ni-NTA to capture uncleaved material and the His-tagged protease) typically achieves >99% tag removal.

Protein L for kappa chain fragments

Protein L from Peptostreptococcus magnus binds the framework region of kappa (κ) light chains without interfering with antigen binding. This makes it suitable for Fab and scFv fragments containing a kappa variable domain, but not for VHH nanobodies (which have no light chain) or lambda-bearing fragments. Typical step yield is 75–85% with 85–95% purity in a single step.

Ion exchange chromatography for tag-free nanobodies

CEX (cation exchange) at pH 4.5–5.5 is the primary capture option for tag-free nanobody production at manufacturing scale. Most VHH nanobodies have a pI of 6–9, so they bind CEX resins at mildly acidic pH while most E. coli host cell proteins (average pI ~5.5) flow through. A salt gradient elution (0–500 mM NaCl) resolves the VHH from remaining contaminants, typically yielding 70–80% recovery at 65–75% purity. A polishing AEX (anion exchange) step in flow-through mode at pH 7–8 removes endotoxin, DNA, and acidic HCPs to achieve >95% purity.

Worked Example: Two-Step Tag-Free VHH Purification

Starting material: 500 mL clarified Pichia culture supernatant, 1.2 g/L VHH (600 mg total), pH adjusted to 5.0

Overall yield = 0.80 × 0.90 = 72% (430 mg from 600 mg input)
Cost of resin per batch = ~$180 (CEX) + ~$120 (AEX) = $300
Resin cost per gram VHH = $300 / 0.43 g = $698/g (first cycle; amortizes over 50–100 cycles)

Bioprocess Scale-Up Considerations

Scaling nanobody production from shake flasks to bioreactors introduces oxygen transfer, heat dissipation, and feed strategy challenges that differ between E. coli and Pichia platforms. The key parameters to control during scale-up depend on the host organism.

E. coli scale-up

High-cell-density E. coli fermentation targeting 50–100 g/L dry cell weight requires careful glucose feeding to avoid acetate overflow (critical threshold: specific glucose uptake rate <1.0 g/g/h). For SHuffle strains, the additional metabolic burden of maintaining an oxidizing cytoplasm makes them more sensitive to oxygen limitation. Maintain DO >30% air saturation during the induction phase, and consider reducing the induction temperature to 16–20 °C in a controlled ramp rather than a step change to avoid cold shock that reduces the proportion of soluble VHH.

Pichia scale-up

Methanol induction at bioreactor scale requires a calibrated methanol sensor (or off-gas CO2 tracking) to maintain residual methanol at 0.5–2 g/L. Methanol concentrations above 5 g/L are toxic, while concentrations below 0.1 g/L starve the AOX1-driven expression system. The oxygen demand during methanol oxidation is exceptionally high (1.5 mol O2 per mol methanol consumed), requiring kLa values of 200–400 h-1 at production scale.

Proteolytic degradation of secreted VHH by Pichia vacuolar proteases (PEP4, PRB1) is the primary scale-up yield loss. Strategies include using protease-deficient strains (GS115 Δpep4 Δprb1), maintaining pH below 6.0, adding casamino acids (0.5–1% w/v) as competitive protease substrates, and minimizing culture time by increasing methanol feed rate.

E. coli Expression Optimizer

Optimize your nanobody expression: strain selection (BL21, SHuffle, Origami), IPTG induction, and soluble vs inclusion body strategies.

Open Calculator

Chromatography Calculator

Size your IMAC, CEX, and AEX columns for antibody fragment purification. Calculate bed volumes, flow rates, and buffer requirements.

Open Calculator

Buffer Calculator

Prepare IMAC binding/elution buffers, ion exchange equilibration buffers, and SEC running buffers with precise pH and conductivity targets.

Open Calculator

Related tools

Frequently Asked Questions

What is the difference between a nanobody and a conventional antibody?

A nanobody (VHH) is a single-domain antibody fragment of approximately 12–15 kDa derived from camelid heavy-chain-only antibodies (HCAb). Unlike conventional antibodies (~150 kDa) composed of two heavy and two light chains requiring complex post-translational processing, nanobodies consist of only one variable domain yet retain full antigen-binding specificity and affinities in the low nanomolar range. Their small size enables tissue penetration, intracellular targeting, and simple microbial expression that conventional antibodies cannot achieve.

Can nanobodies be expressed in E. coli?

Yes. Nanobodies are among the simplest antibody formats to express in E. coli. Periplasmic expression using pelB or OmpA signal peptides yields 10–100 mg/L with proper disulfide bond formation via the DsbA/DsbC system. Cytoplasmic expression in engineered strains such as SHuffle T7 Express achieves 200–300 mg/L soluble VHH by enabling disulfide bond formation in the cytoplasm through trxB and gor mutations.

Why can't Protein A be used to purify nanobodies?

Protein A binds to the CH2–CH3 interface of the Fc region of immunoglobulins. Nanobodies (VHH), Fab fragments, and scFv constructs lack the Fc region entirely, so Protein A has no binding site. Alternatives include IMAC for His-tagged constructs (80–90% yield), Protein L for kappa light chain-bearing Fab and scFv (75–85% yield), and ion exchange chromatography for tag-free nanobodies (70–80% yield).

What yield can you expect from Pichia pastoris nanobody expression?

Pichia pastoris secretes functional nanobodies into the culture medium at 0.1–3 g/L in optimized fed-batch fermentation. The AOX1 promoter with alpha-mating factor secretion signal drives high-level extracellular production, and the secreted product is more homogeneous than E. coli-derived material (single SEC peak vs multiple peaks). Methanol-free systems using the GAP promoter achieve 30–50% lower titers but eliminate the need for flammable methanol handling.

What is the main challenge with scFv production in E. coli?

Aggregation. scFv constructs link two variable domains (VH and VL) with a flexible linker, requiring two intradomain disulfide bonds for stability. In the reducing E. coli cytoplasm, these bonds fail to form. Even in the oxidizing periplasm, the high local concentration of partially folded chains promotes intermolecular domain swapping (VH of one molecule pairs with VL of another) rather than correct intramolecular folding, forming dimers, trimers, and insoluble aggregates.

References

  1. de Marco A. (2020). Recombinant expression of nanobodies and nanobody-derived immunoreagents. Protein Expression and Purification, 172, 105645. doi:10.1016/j.pep.2020.105645
  2. Zheng Y., Li B., Zhao S., Liu J., Li D. (2024). A Universal Strategy for the Efficient Expression of Nanobodies in Pichia pastoris. Fermentation, 10(1), 37. doi:10.3390/fermentation10010037
  3. Li X., Liu W., Li H., Wang X., Zhao Y. (2022). Capture and purification of an untagged nanobody by mixed weak cation chromatography and cation exchange chromatography. Protein Expression and Purification, 191, 106030. doi:10.1016/j.pep.2021.106030
  4. Sarker A., Rathore A.S., Gupta R.D. (2019). Evaluation of scFv protein recovery from E. coli by in vitro refolding and mild solubilization process. Microbial Cell Factories, 18(1), 5. doi:10.1186/s12934-019-1053-9
  5. Duggan S. (2018). Caplacizumab: First Global Approval. Drugs, 78(15), 1639–1642. doi:10.1007/s40265-018-0989-0

Resources & Further Reading