Codon Optimization for Recombinant Protein Expression: CAI, Harmonization & Host-Specific Strategies

August 2026 · Updated September 2026 22 min read Bioprocess Engineering

Key Takeaways

Contents

  1. What Is Codon Optimization?
  2. How to Calculate the Codon Adaptation Index (CAI)
  3. Rare Codons and Their Effect on Expression
  4. Codon Optimization vs Codon Harmonization
  5. Host-Specific Codon Optimization Strategies
  6. Sequence Design Constraints Beyond Codon Usage
  7. Codon Optimization Workflow: From Sequence to Expression
  8. When Should You Use Codon Harmonization Instead of Full Optimization?
  9. How Reliable Is Codon Optimization for Improving Protein Yield?
  10. Quality Control Measures for Codon-Optimized Genes
  11. Frequently Asked Questions

What Is Codon Optimization?

Codon optimization is the redesign of a gene's nucleotide sequence using synonymous codon substitutions to match the codon usage preferences of the expression host, increasing translation efficiency and recombinant protein yield. Because 18 of the 20 standard amino acids are encoded by two to six synonymous codons, the same protein can be specified by astronomically different DNA sequences, and the choice of synonymous codons has a measurable effect on expression level, protein folding, and solubility.

The principle is straightforward: each organism preferentially uses a subset of synonymous codons that correspond to its most abundant tRNAs. When a heterologous gene contains codons that are rare in the host, translation slows at those positions. In extreme cases, the ribosome stalls, frameshifts, or terminates prematurely, reducing both yield and quality of the expressed protein.

Codon optimization strategies fall on a spectrum from aggressive (replacing every codon with the single most-used synonym, maximizing the codon adaptation index) to conservative (codon harmonization, which preserves the relative translational speed profile of the native gene). The right strategy depends on the protein, the host, and whether yield or correct folding is the primary goal.

How to Calculate the Codon Adaptation Index (CAI)

The codon adaptation index (CAI) quantifies how closely a gene's codon usage matches the preferred codon usage of highly expressed genes in a target organism. CAI ranges from 0 to 1.0, where 1.0 means every codon is the most-used synonym for its amino acid. A CAI above 0.8 generally correlates with high expression in E. coli, while native highly expressed genes (ribosomal proteins, elongation factors) typically have CAI values of 0.70 to 0.85.

The calculation proceeds in two steps:

  1. Relative adaptiveness (wi): For each codon i encoding a given amino acid, compute wi = fi / fmax, where fi is the frequency of codon i in the reference set and fmax is the frequency of the most-used codon for that amino acid.
  2. Geometric mean: CAI = (w1 × w2 × … × wL)1/L, where L is the number of codons in the gene (excluding Met and Trp, which have single codons).

Worked Example: CAI Calculation for a Short Peptide in E. coli

Consider a 6-codon sequence encoding Met-Ala-Arg-Leu-Ile-Lys:

ATG GCG CGT CTG ATT AAA

Using E. coli K-12 codon usage tables (Kazusa database):

CAI = (1.00 × 1.00 × 1.00 × 1.00 × 1.00)1/5 = 1.00

If we instead used AGG (Arg, rare in E. coli, w = 0.05) and ATA (Ile, rare, w = 0.07):

ATG GCG AGG CTG ATA AAA

CAI = (1.00 × 0.05 × 1.00 × 0.07 × 1.00)1/5 = (0.0035)0.2 = 0.33

This 3-fold CAI drop from just two rare codons illustrates why even a few poorly chosen codons can drag down the entire gene's translational fitness.

Table 1. CAI values of native highly expressed genes across expression hosts
Host Organism Native HEG CAI Range Recommended Target CAI Diminishing Returns Above
E. coli K-12 / BL21 0.70 – 0.85 0.80 – 0.90 0.95
CHO-K1 / HEK293 0.72 – 0.82 0.75 – 0.85 0.90
Pichia pastoris (K. phaffii) 0.65 – 0.80 0.75 – 0.85 0.90
S. cerevisiae 0.70 – 0.82 0.75 – 0.85 0.90
Sf9 / Hi5 (insect cells) 0.60 – 0.75 0.70 – 0.80 0.85

Rare Codons and Their Effect on Expression

Rare codons reduce expression by starving the ribosome of the cognate aminoacyl-tRNA, causing translational pauses that cascade into three failure modes: ribosome stalling and drop-off (reduced yield), translational frameshifting (incorrect protein), and misincorporation of wrong amino acids at stalled positions (reduced quality). In E. coli, six codons account for the vast majority of expression problems.

Table 2. The six rarest codons in E. coli and their effects on recombinant expression
Codon Amino Acid Usage in E. coli (%) Cognate tRNA Expression Effect
AGG Arg 1.2 tRNAArg4 Frameshifting, premature termination
AGA Arg 2.1 tRNAArg4 Frameshifting at AGG-AGA clusters
AUA Ile 4.3 tRNAIle2 Misincorporation (Met for Ile)
CUA Leu 3.6 tRNALeu3 Ribosome stalling, reduced yield
CGA Arg 3.3 tRNAArg5 Wobble-decoded, slow translation
CCC Pro 4.4 tRNAPro3 Polyproline stalling (with EFP)

Clusters of two or more consecutive rare codons are far more damaging than isolated rare codons scattered through the sequence. A single AGA in a 1,000-codon gene is negligible. Two consecutive AGG-AGA codons can cause a −1 frameshift that produces a truncated, non-functional protein. Genes with more than 5% total rare codon content from these six codons routinely show 5 to 50-fold reduced expression relative to optimized variants.

The BL21(DE3) Rosetta and Rosetta 2 strains co-express rare tRNAs (argU, argW, ileX, glyT, leuW, proL) from the pRARE plasmid, partially rescuing expression of rare-codon-rich genes. However, tRNA co-expression adds metabolic burden and can be lost under selection pressure. Codon optimization of the gene itself is the more robust solution for production-scale expression.

Codon Optimization vs Codon Harmonization

Full codon optimization replaces every codon with the single most frequently used synonym in the target host, maximizing the CAI score and overall translation speed. Codon harmonization, by contrast, preserves the relative translational speed profile of the native gene: codons that are rare in the source organism are replaced with equivalently rare codons in the host, maintaining the natural pauses that allow co-translational domain folding.

The choice between these strategies has measurable consequences. Full optimization typically produces the highest total protein yield but increases the risk of misfolding and inclusion body formation for complex proteins. Harmonization often yields less total protein but a higher fraction of soluble, correctly folded product.

Table 3. Codon optimization vs codon harmonization: when to use each strategy
Parameter Full Optimization Codon Harmonization
CAI achieved 0.90 – 1.00 0.60 – 0.80
Total protein yield Highest Moderate
Soluble fraction Variable (0 – 80%) Higher (40 – 95%)
Inclusion body risk Higher Lower
Best for Small proteins < 30 kDa, IBs acceptable Multi-domain proteins > 50 kDa, activity critical
Typical yield improvement 5 – 100-fold vs native 4 – 1,000-fold vs native (soluble)

Angov et al. (2008) demonstrated codon harmonization on three Plasmodium falciparum proteins expressed in E. coli. The harmonized genes produced 4 to 1,000-fold higher expression than native sequences, with the proteins being soluble and reacting with conformation-specific antibodies, confirming correct folding. Mignon et al. (2018) showed that harmonization outperformed full optimization for a 67-kDa multi-domain protein, yielding 2.5-fold more soluble protein despite lower total expression.

Host-Specific Codon Optimization Strategies

Codon usage bias differs dramatically across expression hosts, and a gene optimized for one organism can fail in another. The most striking differences involve arginine, isoleucine, and leucine codons, where preferred and rare codons are essentially swapped between E. coli and mammalian cells.

Gene Sequence Source organism Target Host Codon usage table Strategy Selection Full CAI optimization Codon harmonization Design Check GC 40-60% No repeats >8 bp No splice sites Gene Synthesis + Cloning into vector Expression Test SDS-PAGE, Western, activity Soluble vs insoluble fraction Yield OK? Scale-Up Production Yes No Misfolded? Try harmonization Low mRNA? Check 5' UTR structure Low yield? tRNA co-expression or re-optimize
Figure 1. Codon optimization workflow from gene sequence to production, with iterative troubleshooting loops for misfolding, low mRNA levels, and low yield.
Workflow diagram showing six sequential steps: gene sequence input, target host codon table selection, strategy selection between full CAI optimization and codon harmonization, sequence design constraints check including GC content and repeat sequences, gene synthesis and cloning, and expression testing. A decision diamond checks whether yield is acceptable; if yes the process proceeds to scale-up production; if no, the process iterates through troubleshooting for misfolding, low mRNA, or low yield.

E. coli

E. coli has the strongest codon bias of any common expression host. Highly expressed genes use a restricted set of about 25 preferred codons out of 61 sense codons. The key optimization rules for E. coli are:

CHO and HEK293 (Mammalian)

Mammalian cells have more balanced tRNA pools than E. coli, so codon usage bias has a smaller effect on expression. The primary gains from optimization come from mRNA-level features rather than translational efficiency:

Pichia pastoris (Komagataella phaffii)

Pichia has an intermediate codon bias. The AOX1 promoter drives methanol-induced expression, and protein secretion adds its own constraints beyond codon usage. Optimization improves expression 2.3 to 2.6-fold on average:

Figure 2. Relative protein yield by codon optimization strategy across three expression hosts. Values normalized to native (wild-type) sequence = 1.0. Based on compiled literature data (Gustafsson et al. 2004, Angov et al. 2008, Mignon et al. 2018).

Sequence Design Constraints Beyond Codon Usage

Maximizing CAI alone is insufficient for reliable expression. Several sequence-level features must be controlled simultaneously during gene design to avoid mRNA instability, aberrant processing, or synthesis failures.

Figure 3. Usage frequency (%) of the most problematic amino acid codons across three expression hosts. A codon preferred in one host may be rare in another, explaining why a gene optimized for E. coli can fail in CHO or Pichia. Source: Kazusa Codon Usage Database.

Codon Optimization Workflow: From Sequence to Expression

A systematic codon optimization workflow reduces the need for iterative troubleshooting by addressing all known sequence-level bottlenecks before gene synthesis. The six-step process below is applicable to any expression host.

  1. Analyze the native gene: Calculate CAI against the target host, identify rare codon clusters (two or more consecutive rare codons), and note GC content distribution along the sequence.
  2. Choose the optimization strategy: For proteins smaller than 30 kDa with no known folding issues, use full CAI optimization. For multi-domain proteins, membrane proteins, or proteins with complex disulfide patterns, use codon harmonization. When in doubt, order both variants.
  3. Apply sequence constraints: After codon substitution, scan for restriction sites, repeats, splice sites (mammalian), and mRNA secondary structure at the 5' end. Adjust codons locally to fix any issues without dropping CAI below 0.75.
  4. Verify computationally: Recalculate CAI, check GC content in 50-bp windows, and predict mRNA secondary structure (RNAfold, mfold). Ensure the optimized sequence still encodes the exact same amino acid sequence.
  5. Synthesize, clone, and express: Order the gene from a synthesis vendor (IDT, Twist, GenScript). Clone into the expression vector. Test expression at small scale (shake flask or 24-well plates) before committing to bioreactor runs.
  6. Iterate if needed: If expression is low, check mRNA levels (qPCR). If protein is insoluble, try harmonization or lower induction temperature. If frameshifting is suspected, sequence the expressed product by mass spectrometry.

Worked Example: Optimizing a Human Cytokine for E. coli Expression

A 522-bp human IL-6 gene (174 amino acids, 22 kDa) is to be expressed in E. coli BL21(DE3).

Step 1 – Analyze: Native human IL-6 has CAI = 0.68 against E. coli. It contains 12 rare codons (7 AGG/AGA, 3 AUA, 2 CUA), including one AGG-AGA cluster at positions 94-95.

Step 2 – Strategy: IL-6 is a small, single-domain protein with three disulfide bonds. Full optimization is appropriate; the small size reduces folding risk.

Step 3 – Optimize:

Step 4 – Verify: Optimized CAI = 0.87. GC content = 52% (range 42-61% in 50-bp windows). No repeats above 8 bp. Amino acid sequence identical.

Result: Optimized gene expressed at 180 mg/L in shake flask vs 12 mg/L for native sequence (15-fold improvement). Protein was soluble (85%) and biologically active.

E. coli Expression Optimizer

Optimize your E. coli expression conditions: strain, promoter, IPTG concentration, and temperature. Works with codon-optimized and native sequences.

Open Calculator

When Should You Use Codon Harmonization Instead of Full Optimization?

Codon harmonization should be the default strategy when the expressed protein must be soluble and correctly folded, especially for complex targets where inclusion body formation would require expensive refolding. The decision depends on three factors: protein size, domain architecture, and whether activity matters more than total yield.

Choose harmonization when:

Choose full optimization when:

A practical middle path is the hybrid strategy: optimize the majority of the sequence for high CAI but deliberately retain or introduce rare codons at domain boundaries (interdomain linker regions) to create translational pauses. This approach has been shown to combine the high mRNA levels of full optimization with the folding fidelity of harmonization.

mRNA Yield Calculator

Estimate mRNA yield from in vitro transcription reactions and scale up your manufacturing process.

Open Calculator

How Reliable Is Codon Optimization for Improving Protein Yield?

Codon optimization is not a deterministic yield multiplier. Across the published record, synonymous redesign of the same gene produces outcomes ranging from more than 100-fold improvement to no measurable change to a measurable loss of yield, and the direction of the effect is frequently not predictable from codon usage metrics alone. Treating a high CAI score as a guarantee of high expression is the single most common mistake in codon optimization.

The clearest demonstration comes from synonymous-variant libraries. Kudla and colleagues expressed 154 synonymous variants of the same GFP gene in E. coli and measured a 250-fold spread in expression between the best and worst variants. Crucially, CAI showed no significant correlation with that spread: the dominant predictor was the folding free energy of the mRNA in the region around the start codon. Variants with stable 5' structure expressed poorly regardless of how favorable their codon usage was. Later large-scale library work reinforced the point from the other direction, finding that much of the codon effect on protein output in E. coli acts through steady-state mRNA level rather than through elongation speed.

The practical consequence is that CAI is one input to a multi-objective design problem, not the objective itself. A sequence at CAI 0.85 with clean 5' structure will usually outperform a sequence at CAI 0.95 with a −18 kcal/mol hairpin over the ribosome binding site.

Table 4. Why an optimized gene can express worse than the native sequence: failure modes, diagnostic signature, and fix
Failure Mode Diagnostic Signature Corrective Action
Stable mRNA structure at the 5' end Low protein, normal or high mRNA Redesign first 30–50 nt to ΔG above −10 kcal/mol
Loss of co-translational folding pauses High total protein, low soluble fraction Harmonize, or reintroduce rare codons at domain boundaries
tRNA pool depletion at very high CAI Yield falls as CAI rises above ~0.95; growth defect on induction Cap CAI at 0.80–0.90; distribute synonyms rather than using one per residue
Introduced cryptic element (splice site, polyA, TATA, terminator) Low mRNA, or truncated transcript on northern/RT-PCR Re-scan and silently mutate the offending motif
Altered mRNA stability from GC shift Low mRNA with no structural or motif explanation Return global GC toward the host's native range (40–60%)
Misincorporation or frameshift at a starved codon Correct yield, wrong intact mass or reduced activity Remove residual rare codons; confirm by LC-MS (see below)

Two of these modes are invisible unless mRNA and protein are measured separately, which is why a single yield number is a poor troubleshooting input. Measure both. A protein drop with unchanged transcript points at translation initiation, elongation, or folding. A protein drop that tracks a transcript drop points at mRNA stability or a cryptic regulatory element introduced during redesign. RT-qPCR against a housekeeping reference plus a densitometry-quantified western, run on the native and optimized constructs side by side under identical induction conditions, separates the two in one experiment.

The reliability of codon optimization also varies systematically by host. In E. coli the effect size is largest and the variance is largest: gains of 5 to 50-fold are common for rare-codon-rich genes, but so are null results. In CHO and HEK293 the tRNA pools are more balanced, so the ceiling is lower — typically 1.5 to 3-fold — and most of that gain comes from removing cryptic splice sites and correcting GC content rather than from codon usage as such. For therapeutic proteins there is a further consideration: regulators and the literature both note that synonymous redesign can alter folding kinetics and, through that, higher-order structure and immunogenicity, so a yield gain is not on its own sufficient evidence that the optimized construct is the better one.

The design implication is to stop ordering a single optimized variant. Gene synthesis is now a small fraction of the cost of a failed expression campaign, so order three to four variants — native, fully optimized, harmonized, and one with 5' structure explicitly relaxed — and screen them in parallel at small scale before committing to a bioreactor run. A four-variant screen answers in one round what iterative single-variant redesign takes three rounds to answer.

Quality Control Measures for Codon-Optimized Genes

Codon optimization changes the DNA sequence while leaving the encoded protein sequence identical — in principle. Quality control exists because that identity has to be demonstrated rather than assumed, at four points: before the gene is ordered, after it is synthesized, after it is cloned, and on the expressed product. Errors caught at the first checkpoint cost nothing; the same error caught at the fourth has consumed a synthesis order, a cloning campaign, and an expression run.

The highest-value check is also the cheapest and the most often skipped: back-translate the final optimized sequence and confirm it encodes the exact target protein. Off-by-one errors, a silently introduced internal stop, and frame mismatches at fusion-tag junctions all survive every codon-usage metric, because CAI, GC content, and repeat scans are all computed without reference to the protein you meant to make.

Table 5. Quality control checkpoints for a codon-optimized construct, from sequence design to expressed product
Stage Check Method Acceptance Criterion
Pre-synthesis Protein identity Back-translation and alignment to target 100% identity; ORF and tags in frame; no internal stop
Pre-synthesis Synthesizability GC scan in 50-bp windows; repeat scan Global GC 40–60%; no window below 30% or above 70%; no direct repeat above 8 bp
Pre-synthesis Translation initiation RNAfold or mfold on the first 30–50 nt ΔG above −10 kcal/mol
Pre-synthesis Cryptic elements Splice-site, polyA, TATA, terminator and restriction-site scans No high-confidence prediction retained; no unintended cloning sites
Post-synthesis Synthesis fidelity Vendor certificate plus Sanger (both strands) or NGS for constructs above ~1 kb 100% match to the ordered sequence
Post-cloning Construct integrity Junction sequencing across both vector–insert boundaries Reading frame continuous through promoter, tags and terminator
Expression Bottleneck localization RT-qPCR plus quantified western, native vs optimized, n ≥ 3 Both reported; soluble fraction stated separately from total
Product Primary structure Intact mass LC-MS; peptide mapping LC-MS/MS Observed mass within a few Da of theoretical; sequence coverage above 95%
Product Higher-order structure and function SEC for aggregate content; potency or activity assay Monomer content and specific activity comparable to the reference material

Intact mass measurement deserves particular emphasis because it is the only routine assay that detects the quality failures codon optimization is capable of causing. Because synonymous redesign should not change the protein, any mass discrepancy is a real finding rather than assay noise. The two documented substitutions in E. coli have distinctive signatures: methionine misincorporated for isoleucine at residual AUA codons shifts the intact mass by roughly +18 Da per event, and lysine misincorporated for arginine at AGA or AGG codons shifts it by roughly −28 Da per event. A −1 frameshift at an AGG-AGA cluster produces a truncated species instead, detectable as a major peak at an unexpected lower mass. Peptide mapping then localizes the substitution to a specific position, which usually points straight at the rare codon that was left in.

Yield and identity are not the whole of quality. Because the mechanism by which optimization can go wrong is disrupted co-translational folding, conformational assays belong in the QC panel and not only in late characterization: size-exclusion chromatography for aggregate content and a functional potency assay will catch a correctly-sequenced protein that folded wrongly, which mass spectrometry by definition will not. For any assay carried into development, the usual expectations apply on specificity, accuracy and precision under analytical method validation.

Regulatory expectations: ICH Q5B

For a recombinant therapeutic, construct quality control is not discretionary. ICH Q5B (Analysis of the Expression Construct in Cells Used for Production of r-DNA Derived Protein Products) requires the nucleotide sequence of the coding region of the expression construct, together with the flanking control regions, to be provided and verified — and this applies to the sequence you actually built, which for an optimized gene is the redesigned sequence, not the native gene it was derived from. The construct must also be shown to be intact in the master cell bank and stable through to the limit of in vitro cell age.

Two practical consequences follow. First, the optimized sequence becomes a regulatory commitment early, so the four-variant screen described above should happen before cell line development begins, not after. Changing the coding sequence afterwards means a new construct, a new cell bank, and a comparability exercise. Second, keep the design record: the native sequence, the codon optimization parameters and software version, the in-silico scan results, and the verification data. Reconstructing why a particular synonymous choice was made, two years later and without notes, is a routine and entirely avoidable problem.

Frequently Asked Questions

What is the difference between codon optimization and codon harmonization?

Codon optimization replaces every codon with the single most-used synonym in the host, maximizing CAI and translation speed. Codon harmonization preserves the relative translational speed profile of the native gene, matching rare codons in the source organism with equivalently rare codons in the host to maintain co-translational folding pauses. Harmonization typically yields lower total protein but higher soluble, correctly folded product for complex multi-domain proteins.

What CAI value should I target for high expression in E. coli?

A CAI of 0.8 or above generally correlates with high expression in E. coli. Highly expressed native E. coli genes (ribosomal proteins, elongation factors) have CAI values of 0.70 to 0.85. Pushing CAI above 0.95 by using a single codon per amino acid can deplete tRNA pools and reduce yield. A practical target is CAI 0.80 to 0.90 with no more than 5% rare codons.

Can codon optimization cause protein misfolding?

Yes. Removing all translational pauses by maximizing codon usage can disrupt co-translational folding, leading to aggregation and inclusion body formation. This is especially problematic for large multi-domain proteins (greater than 50 kDa) and proteins with complex disulfide patterns. If you observe increased inclusion body formation after codon optimization, try codon harmonization or reintroduce a small number of rare codons at domain boundaries.

Which rare codons cause the most problems in E. coli expression?

The six rarest codons in E. coli are AGG, AGA (Arg), AUA (Ile), CUA (Leu), CGA (Arg), and CCC (Pro). Clusters of two or more consecutive rare codons are especially problematic, causing ribosome stalling, frameshifting, and premature termination. Genes with more than 5% rare codon content from these six codons often show 5 to 50-fold reduced expression.

Should I codon-optimize genes for CHO cell expression?

The benefit of codon optimization in CHO cells is smaller than in E. coli because mammalian tRNA pools are more balanced. However, optimizing GC content to 45 to 65%, removing cryptic splice sites, and eliminating internal TATA boxes or polyadenylation signals typically improves expression 1.5 to 3-fold. Full CAI maximization beyond 0.85 rarely adds further benefit in CHO.

Does codon optimization always increase protein yield?

No. Synonymous redesign of the same gene can improve yield more than 100-fold, do nothing, or reduce yield, and CAI alone predicts the outcome poorly. In a library of 154 synonymous GFP variants expressed in E. coli, expression varied 250-fold and the dominant predictor was the folding free energy of the mRNA near the start codon, not codon usage. Expect the largest and most variable effects in E. coli (5 to 50-fold for rare-codon-rich genes) and a smaller ceiling in CHO and HEK293 (typically 1.5 to 3-fold). Order three to four variants and screen them in parallel rather than committing to a single optimized sequence.

Why did my codon-optimized gene express worse than the native sequence?

Six mechanisms account for most yield losses: stable mRNA secondary structure introduced at the 5' end, loss of co-translational folding pauses causing insoluble product, tRNA pool depletion at CAI above about 0.95, a cryptic regulatory element (splice site, polyadenylation signal, TATA box or terminator) created during redesign, altered mRNA stability from a large GC shift, and misincorporation or frameshifting at a residual rare codon. Measure mRNA by RT-qPCR and protein by quantified western on the same samples: a protein drop with unchanged transcript points at translation initiation or folding, while a protein drop that tracks a transcript drop points at mRNA stability or a cryptic element.

What quality control checks should I run on a codon-optimized gene?

Check at four stages. Before ordering, back-translate the optimized sequence and confirm 100% identity to the target protein with tags in frame, then scan GC content in 50-bp windows, repeats, restriction sites, cryptic splice and polyadenylation signals, and the folding free energy of the first 30 to 50 nucleotides. After synthesis, verify against the ordered sequence by Sanger sequencing of both strands or NGS. After cloning, sequence across both vector-insert junctions. On the expressed product, confirm intact mass by LC-MS, map peptides by LC-MS/MS, and run SEC and a potency assay to detect misfolding that mass spectrometry cannot see.

How do you detect amino acid misincorporation caused by rare codons?

Intact mass analysis by LC-MS is the routine assay. Because codon optimization should not change the encoded protein, any mass discrepancy is a real finding rather than assay noise. In E. coli, methionine misincorporated for isoleucine at residual AUA codons shifts the intact mass by roughly +18 Da per event, and lysine misincorporated for arginine at AGA or AGG codons shifts it by roughly -28 Da per event. A -1 frameshift at an AGG-AGA cluster instead produces a truncated species at an unexpected lower mass. Peptide mapping by LC-MS/MS then localizes the substitution to a specific residue, which usually identifies the rare codon that was left in.

What does ICH Q5B require for a codon-optimized construct?

ICH Q5B requires the nucleotide sequence of the coding region of the expression construct and its flanking control regions to be provided and verified, and the construct to be shown intact in the master cell bank and stable to the limit of in vitro cell age. For an optimized gene this applies to the redesigned sequence actually built, not the native gene it was derived from. Because the coding sequence becomes a regulatory commitment early, screen optimized variants before cell line development starts: changing it afterwards means a new construct, a new cell bank, and a comparability exercise.

Related Tools

References

  1. Plotkin JB, Kudla G. Synonymous but not the same: the causes and consequences of codon bias. Nature Reviews Genetics. 2011;12:32-42. doi:10.1038/nrg2899
  2. Gustafsson C, Govindarajan S, Minshull J. Codon bias and heterologous protein expression. Trends in Biotechnology. 2004;22:346-353. doi:10.1016/j.tibtech.2004.04.006
  3. Mignon C, Mariano N, Stadthagen G et al. Codon harmonization – going beyond the speed limit for protein expression. FEBS Letters. 2018;592:1554-1564. doi:10.1002/1873-3468.13046
  4. Angov E, Hillier CJ, Kincaid RL, Lyon JA. Heterologous protein expression is enhanced by harmonizing the codon usage frequencies of the target gene with those of the expression host. PLoS ONE. 2008;3:e2189. doi:10.1371/journal.pone.0002189
  5. Mauro VP. Codon optimization in the production of recombinant biotherapeutics: potential risks and considerations. BioDrugs. 2018;32:69-81. doi:10.1007/s40259-018-0261-x
  6. Kudla G, Murray AW, Tollervey D, Plotkin JB. Coding-sequence determinants of gene expression in Escherichia coli. Science. 2009;324:255-258. doi:10.1126/science.1170160
  7. Boël G, Letso R, Neely H et al. Codon influence on protein expression in E. coli correlates with mRNA levels. Nature. 2016;529:358-363. doi:10.1038/nature16509
  8. ICH. Q5B: Quality of Biotechnological Products — Analysis of the Expression Construct in Cells Used for Production of r-DNA Derived Protein Products. 1995. ICH Q5B Guideline

Resources & Further Reading