SeqBench

Codon Usage Tables by Organism

The genetic code is redundant: most amino acids are encoded by several synonymous codons, and organisms use those alternatives unevenly. This codon usage table shows, for E. coli, human, yeast, CHO, Pichia, insect (Sf9), Arabidopsis and zebrafish, what fraction of each amino acid's codons is each synonymous codon. The most-used — the preferred codon — is highlighted for every organism.

Amino acid1-letter3-letterCodonE. coliHumanYeastCHOPichiaInsectArabidopsisZebrafish
AlanineAAlaGCT16%27%38%32%45%35%43%32%
AlanineAAlaGCC27%40%22%37%26%29%16%30%
AlanineAAlaGCA21%23%29%23%23%18%27%25%
AlanineAAlaGCG36%11%11%7%6%18%14%13%
ArginineRArgCGT38%8%14%11%16%24%17%13%
ArginineRArgCGC40%18%6%18%5%24%7%18%
ArginineRArgCGA6%11%7%14%10%8%12%12%
ArginineRArgCGG10%20%4%19%5%6%9%12%
ArginineRArgAGA4%21%48%19%48%19%35%26%
ArginineRArgAGG2%21%21%19%16%19%20%19%
AsparagineNAsnAAT45%47%59%45%48%32%52%40%
AsparagineNAsnAAC55%53%41%55%52%68%48%60%
Aspartic acidDAspGAT63%46%65%47%58%40%68%47%
Aspartic acidDAspGAC37%54%35%53%42%60%32%53%
CysteineCCysTGT44%46%63%47%64%39%60%50%
CysteineCCysTGC56%54%37%53%36%61%40%50%
Glutamic acidEGluGAA69%42%70%41%56%45%52%36%
Glutamic acidEGluGAG31%58%30%59%44%55%48%64%
GlutamineQGlnCAA35%27%69%24%61%43%56%26%
GlutamineQGlnCAG65%74%31%76%39%57%44%74%
GlycineGGlyGGT34%16%47%20%44%34%34%22%
GlycineGGlyGGC41%34%19%34%14%31%14%28%
GlycineGGlyGGA11%25%22%25%33%28%37%34%
GlycineGGlyGGG15%25%12%21%10%7%16%16%
HistidineHHisCAT57%42%64%44%56%36%61%42%
HistidineHHisCAC43%58%36%56%44%64%39%58%
IsoleucineIIleATT51%36%46%35%50%30%41%34%
IsoleucineIIleATC42%47%26%51%31%54%35%50%
IsoleucineIIleATA7%17%27%14%18%16%24%16%
LeucineLLeuTTA13%8%28%6%16%10%14%8%
LeucineLLeuTTG13%13%29%14%33%20%22%13%
LeucineLLeuCTT10%13%13%13%17%12%26%14%
LeucineLLeuCTC10%20%6%19%8%21%17%18%
LeucineLLeuCTA4%7%14%8%11%9%11%7%
LeucineLLeuCTG50%40%11%39%15%30%11%41%
LysineKLysAAA76%43%58%39%47%35%49%49%
LysineKLysAAG24%57%42%61%53%65%51%51%
MethionineMMetATG100%100%100%100%100%100%100%100%
PhenylalanineFPheTTT57%46%59%47%54%27%51%47%
PhenylalanineFPheTTC43%54%41%53%46%73%49%53%
ProlinePProCCT16%29%31%31%35%29%38%31%
ProlinePProCCC12%32%15%32%15%28%11%24%
ProlinePProCCA19%28%42%29%42%28%33%30%
ProlinePProCCG53%11%12%8%9%16%18%15%
SerineSSerTCT15%19%26%22%29%17%28%20%
SerineSSerTCC15%22%16%22%20%21%13%18%
SerineSSerTCA12%15%21%14%18%17%20%16%
SerineSSerTCG15%5%10%5%9%13%10%7%
SerineSSerAGT15%15%16%15%15%14%16%16%
SerineSSerAGC28%24%11%22%9%18%13%22%
ThreonineTThrACT16%25%35%26%40%27%34%26%
ThreonineTThrACC44%36%22%37%26%32%20%29%
ThreonineTThrACA13%28%30%29%24%23%31%31%
ThreonineTThrACG27%11%14%8%11%18%15%13%
TryptophanWTrpTGG100%100%100%100%100%100%100%100%
TyrosineYTyrTAT57%44%56%44%47%29%52%43%
TyrosineYTyrTAC43%56%44%56%53%71%48%57%
ValineVValGTT26%18%39%18%42%20%40%22%
ValineVValGTC22%24%21%24%23%29%19%23%
ValineVValGTA15%12%21%12%15%17%15%11%
ValineVValGTG37%46%19%46%19%34%26%44%
Stop*StopTAA64%30%47%26%51%63%36%36%
Stop*StopTAG7%24%23%24%29%18%20%18%
Stop*StopTGA29%47%30%50%20%18%44%46%

green= preferred codon (highest fraction) for that organism· Values are fractions among synonymous codons · Each amino acid sums to ~100%

What codon usage means

Because the genetic code is degenerate, leucine has six codons, isoleucine three, and only methionine and tryptophan have one each. The numbers above are not raw counts: each value is the fraction of that codon among the synonymous codons for the same amino acid, so every amino-acid group sums to roughly 1.0 (100%). This relative measure — similar in spirit to RSCU (relative synonymous codon usage) — is what reveals codon bias: the systematic preference an organism shows for some synonymous codons over others, driven by its tRNA pool, genome composition and selection on highly expressed genes.

Preferred and rare codons

For each amino acid and organism, the codon with the highest fraction is the preferred (optimal) codon— shown in green above. It is the codon a host's most highly expressed genes tend to use, and the one a codon optimizer selects when rewriting a gene. At the other end, codons with a low fraction are rare codons: their matching tRNAs are often scarce, so clusters of them can slow or stall translation, reduce protein yield and occasionally promote misfolding. When a gene is moved into a new host, swapping rare codons for the host's preferred synonyms — without changing the encoded protein — is the core of codon optimization.

About these values

These fractions are reference approximations derived from the Kazusa Codon Usage Database for each organism and are intended for guidance. Published codon tables differ between sources, releases and the exact gene set they are computed from, so for critical work — synthesizing a gene, troubleshooting low expression — you should verify the values against your specific expression system rather than treat any single table as definitive.

Frequently asked questions

What is a codon usage table?

A codon usage table lists, for each amino acid, the synonymous codons that encode it and how often each one is actually used in a given organism's genes. Because the genetic code is redundant — most amino acids have two to six codons — organisms use those alternatives unevenly, and the table captures that preference.

What does the fraction or percentage mean?

Each value is the fraction of that codon among the synonymous codons for the same amino acid, so the codons for any one amino acid sum to about 1.0 (100%). For example, if Leucine's CTG shows 50% in E. coli, half of all leucine codons in E. coli genes are CTG. The number is a relative share within an amino acid, not the codon's frequency across the whole genome.

What is the preferred (optimal) codon?

The preferred codon is the synonymous codon with the highest usage fraction for an amino acid in that organism — the cells highlighted in green in the table. It is the codon a highly expressed gene is most likely to use, and the one a codon optimizer will pick when rewriting a gene for that host.

What are rare codons and why avoid them?

Rare codons are synonymous codons with a low usage fraction in the target organism, often because the matching tRNA is scarce. A run of rare codons can slow or stall the ribosome, lower protein yield and sometimes cause misfolding — so they are usually avoided when expressing a gene in a heterologous host.

Why do E. coli and human codon usage differ?

Codon bias is shaped by each organism's tRNA pool, genome GC content and evolutionary history, so different species favor different synonymous codons. A gene that is well optimized for human cells can contain codons that are rare in E. coli, which is why a sequence often needs codon optimization before it is expressed in a new host.

See also

Sources

  1. 1
    Codon Usage Database (CUTG: Codon Usage Tabulated from GenBank)
    Nakamura Y (maintainer) · Kazusa DNA Research Institute, Department of Plant Gene Research · Data source: NCBI-GenBank Flat File Release 160.0 [15 June 2007]; 35,799 organisms / 3,027,973 CDS
    Every percentage in the eight-organism table — all 64 codon rows across E. coli, Human, Yeast, CHO, Pichia, Insect, Arabidopsis and Zebrafish (CODON_USAGE in src/lib/bio/codon-usage.ts) — and therefore also the green preferred-codon highlighting derived from them, and the 'About these values' section's attribution to Kazusa.
  2. 2
    Codon usage tabulated from international DNA sequence databases: status for the year 2000
    Nakamura Y, Gojobori T, Ikemura T · Nucleic Acids Res 28(1):292 · 2000
    The citable reference for the database itself — this is what to cite alongside the URL when the page says the values are 'derived from the Kazusa Codon Usage Database'. It backs the provenance claim, not any individual percentage.
  3. 3
    Codon usage table: Escherichia coli W3110 [gbbct] (Kazusa species 316407)
    Kazusa Codon Usage Database · Dataset: 4,332 CDS / 1,372,057 codons, from NCBI-GenBank Flat File Release 160.0 [15 June 2007]
    The 'E. coli' column specifically — all 64 values. Named separately because the repo deliberately switched to this dataset (see the 2026-08-05 audit note in src/lib/bio/codon-usage.ts) after finding the more obvious species=83333 'E. coli K12' entry is a 14-CDS / 5,122-codon dataset whose amber stop TAG reads 0.0. Anyone re-deriving the column needs this exact species id, not 'E. coli'.

Related tools and references