Protein Tags Reference Table
A reference table of the common protein tags used in recombinant expression: affinity tags for purification, epitope tags for detection, and solubility / fusion partners for folding. Amino acid sequences are shown in MONOSPACE; large fusion proteins list see notes with their approximate mass and source instead of a full sequence.
| Tag | Type | AA sequence | Length | Mass (kDa) | Purification / detection | Cleavable by | Notes |
|---|---|---|---|---|---|---|---|
| His6 (6×His) | Affinity | HHHHHH | 6 | 0.84 | Ni-NTA / Co (TALON) IMAC resin; anti-His mAb | — | Most common purification tag. Binds immobilized Ni2+/Co2+; elute with imidazole or low pH. Tag is not itself protease-cleavable — cleavage depends on a separate site (e.g. TEV, thrombin) engineered between tag and target. Codons often CAT/CAC mix. |
| His8 (8×His) | Affinity | HHHHHHHH | 8 | 1.1 | Ni-NTA / Co IMAC resin; anti-His mAb | — | Higher-avidity variant of His6; tighter IMAC binding, useful for low-expression or membrane proteins. Otherwise same handling as His6. |
| FLAG | Affinity | DYKDDDDK | 8 | 1.01 | Anti-FLAG M1/M2/M5 mAb; anti-FLAG (M2) affinity gel; elute with FLAG peptide or low pH | Enterokinase (cleaves after DDDDK↓) | Sigma/Millipore trademark. Hydrophilic, often surface-exposed. Enterokinase cleaves at the C-terminal DDDDK, leaving native N-terminus when FLAG is N-terminal. M1 antibody requires free N-terminal Asp (Ca2+-dependent). |
| 3×FLAG | Affinity | DYKDHDGDYKDHDIDYKDDDDK | 22 | 2.73 | Anti-FLAG M2 mAb / M2 affinity gel; 3×FLAG peptide elution | Enterokinase (at terminal DDDDK↓) | Sigma trademark. Tandem repeat gives ~10× more sensitive detection than single FLAG. ⚠ Exact junction residues (DHDGD / DHDID) are from the Sigma 3×FLAG design; sequence above is the standard published form. |
| Strep-tag II | Affinity | WSHPQFEK | 8 | 1.06 | Strep-Tactin / Strep-Tactin XT resin; StrepMAB; gentle elution with desthiobiotin/biotin | — | IBA Lifesciences trademark. Kd ~1 µM to Strep-Tactin. Very mild, physiological elution → good for intact complexes. Twin-Strep-tag (two copies, WSHPQFEK...GGGSGGGSGGSA...WSHPQFEK) binds much tighter; use it for demanding preps. |
| S-tag | Affinity | KETAAAKFERQHMDS | 15 | 1.75 | S-protein resin / S-protein–HRP (RNase S system); quantitative S-Tag assay | — | Novagen/Merck. Derived from the S-peptide of RNase A; binds S-protein to reconstitute RNase activity, enabling sensitive quantitation. Not designed as a standalone cleavage site. |
| HA (hemagglutinin) | Epitope | YPYDVPDYA | 9 | 1.1 | Anti-HA mAb (12CA5, 3F10, HA.11); anti-HA agarose | — | From influenza hemagglutinin HA1 (residues ~98–106). Widely used for Western/IP/IF; well-tolerated at N- or C-terminus. |
| c-Myc | Epitope | EQKLISEEDL | 10 | 1.2 | Anti-Myc mAb (9E10); anti-Myc agarose | — | From human c-Myc (residues 410–419). Classic detection/IP tag; 9E10 is the standard antibody. Often combined with His6. |
| V5 | Epitope | GKPIPNPLLGLDST | 14 | 1.42 | Anti-V5 mAb; anti-V5 agarose | — | Derived from the P/V proteins of simian virus SV5 (paramyxovirus). Common in Invitrogen/Thermo vectors (e.g. pcDNA). Low background, good for mammalian expression. |
| T7-tag | Epitope | MASMTGGQQMG | 11 | 1.1 | Anti-T7 mAb; T7-Tag antibody agarose | — | From the N-terminus (leader) of T7 gene 10 capsid protein. Novagen/Merck pET-system detection tag. ⚠ Length varies by vector: an 11-residue form (MASMTGGQQMG) is standard; some report the first ~11–13 residues. |
| GST (glutathione S-transferase) | Solubility | see notes | 218 | 26 | Glutathione-Sepharose / GSH resin (affinity); anti-GST mAb; elute with reduced glutathione | Thrombin or PreScission/HRV-3C (site depends on vector, e.g. pGEX) | Schistosoma japonicum GST, ~26 kDa (~218 aa). Dual affinity + moderate solubility enhancer. Forms dimers — can be a drawback for oligomerization studies. Full sequence: UniProt P08515 / Addgene pGEX vectors. Cleavage site (LVPR↓GS thrombin or LEVLFQ↓GP 3C) is vector-specific. |
| MBP (maltose-binding protein) | Solubility | see notes | 396 | 43.4 | Amylose resin (affinity); anti-MBP mAb; elute with maltose | Factor Xa, TEV, or PreScission/3C (site depends on vector, e.g. pMAL) | E. coli MalE. The 396 aa / 43.4 kDa pair here is the UniProt P0AEX9 PRECURSOR, and the two figures now describe the same molecule: the row previously read 396 aa beside 42.5 kDa, which is neither the precursor (43.4) nor the mature periplasmic chain 27–396 (370 aa, 40.7 kDa — run either through Protein Properties and it reproduces). ~42.5 kDa is NEB's figure for the MBP moiety its pMAL vectors express, a third construct again. Strong solubility enhancer plus amylose affinity. ⚠ length_aa ~366–396 depending on signal-peptide/linker variant used. |
| SUMO (Smt3) | Solubility | see notes | 98 | 11 | Typically paired with an N-terminal His6 for Ni-NTA capture; no intrinsic affinity resin | SUMO protease (Ulp1 / SENP), cleaves after C-terminal di-Gly | Yeast Smt3 (~11 kDa) is the common form (LifeSensors "Champion SUMO"). Enhances solubility/expression; Ulp1 recognizes tertiary structure and cleaves after the C-terminal Gly-Gly, leaving a NATIVE N-terminus (any residue except Pro). Full sequence: UniProt Q12306 (Smt3) / Addgene pET-SUMO. |
| NusA | Solubility | see notes | 495 | 55 | No intrinsic affinity resin — used with a co-tag (e.g. His6) for purification | TEV or other vector-defined protease site | E. coli transcription factor NusA (~55 kDa). Very effective solubility enhancer, especially for toxic/aggregation-prone targets, but large — high metabolic burden and reduced molar yield. Full sequence: UniProt P0AFF6 / Novagen pET-44 (NusA·Tag). |
| Thioredoxin (Trx / TrxA) | Solubility | see notes | 109 | 12 | No intrinsic affinity resin — pair with His6/His-patch (ThioFusion); anti-Trx available | Enterokinase / thrombin / TEV (vector-dependent) | E. coli TrxA (~11.7 kDa, 109 aa). Compact, highly soluble; best for small targets (<30 kDa) without disulfides in the reducing cytoplasm. Full sequence: UniProt P0AA25 / Invitrogen pTrxFus, pET-32 (Trx·Tag). |
| HaloTag | Self-labeling | see notes | 297 | 33 | HaloTag ligands (chloroalkane): fluorophores, biotin, HaloLink resin — COVALENT capture | TEV (in Promega HaloTag vectors, a TEV site flanks the tag) | Promega. Engineered haloalkane dehalogenase (~33 kDa, ~297 aa) that forms an IRREVERSIBLE covalent bond to chloroalkane ligands → very stable pulldowns/labeling and one-step covalent immobilization. ⚠ mass cited variously as 33–34 kDa. Not a classic epitope/affinity peptide. |
| SNAP-tag | Self-labeling | see notes | 182 | 20 | O6-benzylguanine (BG) ligands: fluorophores, biotin, resin — COVALENT self-labeling | Vector-dependent (TEV/3C sites offered in some constructs) | NEB. Engineered human O6-alkylguanine-DNA alkyltransferase (hAGT, ~20 kDa, ~182 aa) that covalently reacts with benzylguanine substrates. CLIP-tag is a companion that reacts with benzylcytosine (orthogonal labeling). ⚠ mass ~19.4–20 kDa depending on construct. |
Choosing a tag
Start with the smallest tag that solves your problem. A His6 tag is the standard workhorse for purification; add an epitope tag like HA, c-Myc or V5 when you need antibody-based detection. If the target is insoluble or aggregation-prone, reach for a solubility partner (MBP, SUMO, NusA or Trx) — but remember these are large and reduce molar yield. Once purified, check the predicted mass and pI of your tagged construct with the protein molecular weight and pI guide.
Removing a tag
Most peptide tags do not cleave themselves — cleavage depends on a separate protease site engineered between the tag and your protein. Common choices are TEV, thrombin, PreScission/HRV-3C, enterokinase (for FLAG) and SUMO protease (which leaves a native N-terminus). The Cleavable by column above lists the protease each vector typically pairs with.
Frequently asked questions
What is a protein tag?
A protein tag is a short peptide or a whole fusion protein genetically added to your target's N- or C-terminus. Tags let you purify (affinity tags), detect (epitope tags), or improve the folding and solubility (solubility tags) of a recombinant protein.
Which protein tag should I use for purification?
His6 (6×His) is the default first choice: small, cheap Ni-NTA/Co IMAC purification that works under native or denaturing conditions. Use FLAG or Strep-tag II when you need very mild, specific elution (e.g. for intact complexes), and GST or MBP when the target also needs a solubility boost.
What is the difference between an affinity tag and an epitope tag?
Affinity tags (His6, Strep-tag II, GST, MBP) bind a resin or ligand so you can capture and elute the protein. Epitope tags (HA, c-Myc, V5, T7) are recognized by well-characterized antibodies and are used mainly for Western blot, immunoprecipitation and immunofluorescence. Several tags (FLAG, S-tag) do both.
Do protein tags need to be removed?
Not always — small tags like His6 are usually left on. When the tag interferes with activity, crystallization or immunogenicity, engineer a protease site (TEV, thrombin, PreScission/HRV-3C, enterokinase, or SUMO protease) between the tag and target so it can be cleaved off after purification.
How big is a His tag?
A 6×His tag is six histidine residues, about 0.84 kDa. His8 (eight histidines, ~1.1 kDa) binds IMAC resin more tightly and is useful for low-expression or membrane proteins.
Why are GST, MBP and SUMO shown without a full sequence?
These are whole proteins (hundreds of residues) fused as solubility/affinity partners, so the table lists "see notes" plus an approximate mass and the UniProt/vector source instead of inlining the full sequence. Retrieve the exact sequence from the cited UniProt accession or vector map.
Learn more
Sources
- 1UniProtKB reviewed (Swiss-Prot) entries for the five fusion partners: P08515, P0AEX9, Q12306, P0AFF6, P0AA25UniProt Consortium (EMBL-EBI / SIB / PIR) · 2026The Length and Mass (kDa) columns for all five 'see notes' fusion rows. I fetched each accession's flat file individually (https://rest.uniprot.org/uniprotkb/<acc>.txt): P08515 GST26_SCHJA Schistosoma japonicum GST = 218 aa, 25,499 Da (page: 218 aa, 26 kDa — matches); P0AEX9 MALE_ECOLI = 396 aa, 43,388 Da, SIGNAL 1-26, mature CHAIN 27-396 = 370 aa (page: 396 aa — matches the precursor, and the page's own caveat '~366-396 depending on signal-peptide/linker variant' is correct); Q12306 SUMO_YEAST Smt3 = 101 aa, 11,597 Da, PROPEP 99-101 removed leaving mature CHAIN 2-98 ending at Gly98 (page: 98 aa, 11 kDa — this is the MATURE form, and the page's note that Ulp1 'cleaves after the C-terminal Gly-Gly' is exactly the 98/101 boundary UniProt annotates); P0AFF6 NUSA_ECOLI = 495 aa, 54,871 Da (page: 495 aa, 55 kDa — matches); P0AA25 THIO_ECOLI TrxA = 109 aa, 11,807 Da, INIT_MET removed, CHAIN 2-109 (page: 109 aa, 12 kDa, note '~11.7 kDa' — matches to rounding). Also backs every accession the Notes column cites. ONE MISMATCH to fix or footnote: the page's MBP mass of 42.5 kDa matches neither the UniProt precursor (43.4 kDa) nor the mature MalE (~40.7 kDa) — see `unsourced`.
- 2A Short Polypeptide Marker Sequence Useful for Recombinant Protein Identification and PurificationHopp TP, Prickett KS, Price VL, Libby RT, March CJ, Cerretti DP, Urdal DL, Conlon PJ · Bio/Technology (now Nature Biotechnology) 6:1204-1210 · 1988The FLAG row in full: the eight-residue sequence DYKDDDDK (Asp-Tyr-Lys-Asp-Asp-Asp-Asp-Lys), length_aa 8, the 'hydrophilic, often surface-exposed' note, and the Cleavable by entry 'Enterokinase (cleaves after DDDDK)'. It is also the parent design the 3xFLAG row's terminal DYKDDDDK repeat derives from (but NOT the 3xFLAG junction residues — see `unsourced`).
- 3Genetic Approach to Facilitate Purification of Recombinant Proteins with a Novel Metal Chelate AdsorbentHochuli E, Bannwarth W, Döbeli H, Gentz R, Stüber D · Bio/Technology (now Nature Biotechnology) 6:1321-1325 · 1988The His6 and His8 rows' underlying principle: a genetically fused poly-histidine peptide captured on a nitrilotriacetate (NTA) metal-chelate adsorbent — i.e. the 'Ni-NTA / Co (TALON) IMAC resin' purification column and the notes' 'Binds immobilized Ni2+/Co2+'. It does NOT back the specific 6x vs 8x choice, the imidazole/low-pH elution recipes, or the His8 'higher avidity / useful for membrane proteins' claim, all of which are vendor practice.
- 4Purification of a RAS-responsive adenylyl cyclase complex from Saccharomyces cerevisiae by use of an epitope addition methodField J, Nikawa J, Broek D, MacDonald B, Rodgers L, Wilson IA, Lerner RA, Wigler M · Molecular and Cellular Biology 8(5):2159-2165 (PMID 2455217, PMC363397) · 1988The HA (hemagglutinin) row's origin as a TAG: the first use of the influenza-HA peptide epitope as a genetically added purification/detection handle ('epitope addition'), which is what the row's use in Western/IP/IF rests on. The nine-residue sequence YPYDVPDYA itself and its HA1 ~98-106 numbering are better attributed to the paper this one builds on — Wilson IA, Niman HL, Houghten RA, Cherenson AR, Connolly ML, Lerner RA, 'The structure of an antigenic determinant in a protein', Cell 37(3):767-778, 1984, PMID 6204768, https://doi.org/10.1016/0092-8674(84)90412-4 — whose abstract I verified states that the anti-peptide monoclonals recognise 'one specific nine amino acid sequence' in influenza hemagglutinin. Wilson and Lerner are co-authors on both, so the pairing is the real lineage, not a guess.
- 5Isolation of monoclonal antibodies specific for human c-myc proto-oncogene productEvan GI, Lewis GK, Ramsay G, Bishop JM · Molecular and Cellular Biology 5(12):3610-3616 (PMID 3915782, PMC369192) · 1985The c-Myc row: the origin of the 9E10 monoclonal named in the Purification/detection column, and of the synthetic c-myc peptide immunogen from which the EQKLISEEDL tag is taken. The verified abstract states the antibodies were raised 'from mice immunized with synthetic peptide immunogens whose sequences are derived from that of the human c-myc gene product'.
- 6Identification of an epitope on the P and V proteins of simian virus 5 that distinguishes between two isolates with different biological characteristicsSouthern JA, Young DF, Heaney F, Baumgärtner WK, Randall RE · Journal of General Virology 72(7):1551-1557 · 1991The V5 row: the Pk epitope shared by the P and V proteins of simian virus 5 (a paramyxovirus), from which the 14-residue GKPIPNPLLGLDST tag derives — i.e. the row's sequence, length_aa 14, and the note 'Derived from the P/V proteins of simian virus SV5'. It does NOT back the 'common in Invitrogen/Thermo vectors (e.g. pcDNA)' or 'low background' claims, which are vendor practice.
- 7The Strep-tag system for one-step purification and high-affinity detection or capturing of proteinsSchmidt TGM, Skerra A · Nature Protocols 2(6):1528-1535 (PMID 17571060) · 2007Most of the Strep-tag II row, and it is the only source here whose abstract contains the tag sequence literally: 'The Strep-tag II is an eight-residue minimal peptide sequence (Trp-Ser-His-Pro-Gln-Phe-Glu-Lys)' = WSHPQFEK, confirming the sequence and length_aa 8; plus affinity chromatography 'on a matrix carrying an engineered streptavidin (Strep-Tactin)' and the StrepMAB antibodies, i.e. the Purification/detection column. Note Schmidt is at IBA GmbH, which is the trademark holder the Notes column names. It does NOT back the 'Kd ~1 uM to Strep-Tactin' figure or the Twin-Strep-tag linker sequence — see `unsourced`.
- 8HaloTag: a novel protein labeling technology for cell imaging and protein analysisLos GV, Encell LP, McDougall MG, et al. (Promega Corporation) · ACS Chemical Biology 3(6):373-382 (PMID 18533659) · 2008The HaloTag row's mechanism claims, verbatim from the verified abstract: 'a modified haloalkane dehalogenase designed to covalently bind to synthetic ligands', ligands comprising 'a chloroalkane linker attached to a variety of useful molecules, such as fluorescent dyes, affinity handles, or solid surfaces', and bond formation that 'is essentially irreversible' — which is the row's COVALENT capture claim and the 'IRREVERSIBLE covalent bond' note, including protein immobilization. It does NOT back the 33 kDa / 297 aa figures in the Mass and Length columns (see `unsourced`).
- 9A general method for the covalent labeling of fusion proteins with small molecules in vivoKeppler A, Gendreizig S, Gronemeyer T, Pick H, Vogel H, Johnsson K · Nature Biotechnology 21(1):86-89 (PMID 12469133) · 2003The SNAP-tag row's mechanism: the founding paper for covalent self-labeling of fusion proteins via human O6-alkylguanine-DNA alkyltransferase (hAGT) reacting with O6-benzylguanine substrates — i.e. the row's 'Engineered human O6-alkylguanine-DNA alkyltransferase (hAGT)... covalently reacts with benzylguanine substrates' note and the BG-ligand Purification/detection column. The engineering of the hAGT mutants that made it practical is the companion paper Juillerat A et al., Chem Biol 10(4):313-317, 2003, PMID 12725859, https://doi.org/10.1016/s1074-5521(03)00068-1 (verified), if you want the directed-evolution half. Neither backs the 20 kDa / 182 aa figures (see `unsourced`).
Related tools and references
Tools
Compute molecular weight, isoelectric point, extinction coefficient and composition.
Digest a protein with trypsin, Lys-C, chymotrypsin and more, and get peptide masses.
Back-translate a protein to DNA using most-frequent or degenerate IUPAC codons.
Nearby reference tables
One-letter and three-letter amino acid codes with key properties.
Side-chain and terminal pKa values and how they set the isoelectric point.
Kyte-Doolittle, Hopp-Woods and Eisenberg values for all 20 amino acids, with the window sizes and transmembrane threshold each scale is read with.