SeqBench

Protein Tags Reference Table

A reference table of the common protein tags used in recombinant expression: affinity tags for purification, epitope tags for detection, and solubility / fusion partners for folding. Amino acid sequences are shown in MONOSPACE; large fusion proteins list see notes with their approximate mass and source instead of a full sequence.

TagTypeAA sequenceLengthMass (kDa)Purification / detectionCleavable byNotes
His6 (6×His)AffinityHHHHHH60.84Ni-NTA / Co (TALON) IMAC resin; anti-His mAb—Most common purification tag. Binds immobilized Ni2+/Co2+; elute with imidazole or low pH. Tag is not itself protease-cleavable — cleavage depends on a separate site (e.g. TEV, thrombin) engineered between tag and target. Codons often CAT/CAC mix.
His8 (8×His)AffinityHHHHHHHH81.1Ni-NTA / Co IMAC resin; anti-His mAb—Higher-avidity variant of His6; tighter IMAC binding, useful for low-expression or membrane proteins. Otherwise same handling as His6.
FLAGAffinityDYKDDDDK81.01Anti-FLAG M1/M2/M5 mAb; anti-FLAG (M2) affinity gel; elute with FLAG peptide or low pHEnterokinase (cleaves after DDDDK↓)Sigma/Millipore trademark. Hydrophilic, often surface-exposed. Enterokinase cleaves at the C-terminal DDDDK, leaving native N-terminus when FLAG is N-terminal. M1 antibody requires free N-terminal Asp (Ca2+-dependent).
3×FLAGAffinityDYKDHDGDYKDHDIDYKDDDDK222.73Anti-FLAG M2 mAb / M2 affinity gel; 3×FLAG peptide elutionEnterokinase (at terminal DDDDK↓)Sigma trademark. Tandem repeat gives ~10× more sensitive detection than single FLAG. ⚠ Exact junction residues (DHDGD / DHDID) are from the Sigma 3×FLAG design; sequence above is the standard published form.
Strep-tag IIAffinityWSHPQFEK81.06Strep-Tactin / Strep-Tactin XT resin; StrepMAB; gentle elution with desthiobiotin/biotin—IBA Lifesciences trademark. Kd ~1 µM to Strep-Tactin. Very mild, physiological elution → good for intact complexes. Twin-Strep-tag (two copies, WSHPQFEK...GGGSGGGSGGSA...WSHPQFEK) binds much tighter; use it for demanding preps.
S-tagAffinityKETAAAKFERQHMDS151.75S-protein resin / S-protein–HRP (RNase S system); quantitative S-Tag assay—Novagen/Merck. Derived from the S-peptide of RNase A; binds S-protein to reconstitute RNase activity, enabling sensitive quantitation. Not designed as a standalone cleavage site.
HA (hemagglutinin)EpitopeYPYDVPDYA91.1Anti-HA mAb (12CA5, 3F10, HA.11); anti-HA agarose—From influenza hemagglutinin HA1 (residues ~98–106). Widely used for Western/IP/IF; well-tolerated at N- or C-terminus.
c-MycEpitopeEQKLISEEDL101.2Anti-Myc mAb (9E10); anti-Myc agarose—From human c-Myc (residues 410–419). Classic detection/IP tag; 9E10 is the standard antibody. Often combined with His6.
V5EpitopeGKPIPNPLLGLDST141.42Anti-V5 mAb; anti-V5 agarose—Derived from the P/V proteins of simian virus SV5 (paramyxovirus). Common in Invitrogen/Thermo vectors (e.g. pcDNA). Low background, good for mammalian expression.
T7-tagEpitopeMASMTGGQQMG111.1Anti-T7 mAb; T7-Tag antibody agarose—From the N-terminus (leader) of T7 gene 10 capsid protein. Novagen/Merck pET-system detection tag. ⚠ Length varies by vector: an 11-residue form (MASMTGGQQMG) is standard; some report the first ~11–13 residues.
GST (glutathione S-transferase)Solubilitysee notes21826Glutathione-Sepharose / GSH resin (affinity); anti-GST mAb; elute with reduced glutathioneThrombin or PreScission/HRV-3C (site depends on vector, e.g. pGEX)Schistosoma japonicum GST, ~26 kDa (~218 aa). Dual affinity + moderate solubility enhancer. Forms dimers — can be a drawback for oligomerization studies. Full sequence: UniProt P08515 / Addgene pGEX vectors. Cleavage site (LVPR↓GS thrombin or LEVLFQ↓GP 3C) is vector-specific.
MBP (maltose-binding protein)Solubilitysee notes39643.4Amylose resin (affinity); anti-MBP mAb; elute with maltoseFactor Xa, TEV, or PreScission/3C (site depends on vector, e.g. pMAL)E. coli MalE. The 396 aa / 43.4 kDa pair here is the UniProt P0AEX9 PRECURSOR, and the two figures now describe the same molecule: the row previously read 396 aa beside 42.5 kDa, which is neither the precursor (43.4) nor the mature periplasmic chain 27–396 (370 aa, 40.7 kDa — run either through Protein Properties and it reproduces). ~42.5 kDa is NEB's figure for the MBP moiety its pMAL vectors express, a third construct again. Strong solubility enhancer plus amylose affinity. ⚠ length_aa ~366–396 depending on signal-peptide/linker variant used.
SUMO (Smt3)Solubilitysee notes9811Typically paired with an N-terminal His6 for Ni-NTA capture; no intrinsic affinity resinSUMO protease (Ulp1 / SENP), cleaves after C-terminal di-GlyYeast Smt3 (~11 kDa) is the common form (LifeSensors "Champion SUMO"). Enhances solubility/expression; Ulp1 recognizes tertiary structure and cleaves after the C-terminal Gly-Gly, leaving a NATIVE N-terminus (any residue except Pro). Full sequence: UniProt Q12306 (Smt3) / Addgene pET-SUMO.
NusASolubilitysee notes49555No intrinsic affinity resin — used with a co-tag (e.g. His6) for purificationTEV or other vector-defined protease siteE. coli transcription factor NusA (~55 kDa). Very effective solubility enhancer, especially for toxic/aggregation-prone targets, but large — high metabolic burden and reduced molar yield. Full sequence: UniProt P0AFF6 / Novagen pET-44 (NusA·Tag).
Thioredoxin (Trx / TrxA)Solubilitysee notes10912No intrinsic affinity resin — pair with His6/His-patch (ThioFusion); anti-Trx availableEnterokinase / thrombin / TEV (vector-dependent)E. coli TrxA (~11.7 kDa, 109 aa). Compact, highly soluble; best for small targets (<30 kDa) without disulfides in the reducing cytoplasm. Full sequence: UniProt P0AA25 / Invitrogen pTrxFus, pET-32 (Trx·Tag).
HaloTagSelf-labelingsee notes29733HaloTag ligands (chloroalkane): fluorophores, biotin, HaloLink resin — COVALENT captureTEV (in Promega HaloTag vectors, a TEV site flanks the tag)Promega. Engineered haloalkane dehalogenase (~33 kDa, ~297 aa) that forms an IRREVERSIBLE covalent bond to chloroalkane ligands → very stable pulldowns/labeling and one-step covalent immobilization. ⚠ mass cited variously as 33–34 kDa. Not a classic epitope/affinity peptide.
SNAP-tagSelf-labelingsee notes18220O6-benzylguanine (BG) ligands: fluorophores, biotin, resin — COVALENT self-labelingVector-dependent (TEV/3C sites offered in some constructs)NEB. Engineered human O6-alkylguanine-DNA alkyltransferase (hAGT, ~20 kDa, ~182 aa) that covalently reacts with benzylguanine substrates. CLIP-tag is a companion that reacts with benzylcytosine (orthogonal labeling). ⚠ mass ~19.4–20 kDa depending on construct.

Choosing a tag

Start with the smallest tag that solves your problem. A His6 tag is the standard workhorse for purification; add an epitope tag like HA, c-Myc or V5 when you need antibody-based detection. If the target is insoluble or aggregation-prone, reach for a solubility partner (MBP, SUMO, NusA or Trx) — but remember these are large and reduce molar yield. Once purified, check the predicted mass and pI of your tagged construct with the protein molecular weight and pI guide.

Removing a tag

Most peptide tags do not cleave themselves — cleavage depends on a separate protease site engineered between the tag and your protein. Common choices are TEV, thrombin, PreScission/HRV-3C, enterokinase (for FLAG) and SUMO protease (which leaves a native N-terminus). The Cleavable by column above lists the protease each vector typically pairs with.

Frequently asked questions

What is a protein tag?

A protein tag is a short peptide or a whole fusion protein genetically added to your target's N- or C-terminus. Tags let you purify (affinity tags), detect (epitope tags), or improve the folding and solubility (solubility tags) of a recombinant protein.

Which protein tag should I use for purification?

His6 (6×His) is the default first choice: small, cheap Ni-NTA/Co IMAC purification that works under native or denaturing conditions. Use FLAG or Strep-tag II when you need very mild, specific elution (e.g. for intact complexes), and GST or MBP when the target also needs a solubility boost.

What is the difference between an affinity tag and an epitope tag?

Affinity tags (His6, Strep-tag II, GST, MBP) bind a resin or ligand so you can capture and elute the protein. Epitope tags (HA, c-Myc, V5, T7) are recognized by well-characterized antibodies and are used mainly for Western blot, immunoprecipitation and immunofluorescence. Several tags (FLAG, S-tag) do both.

Do protein tags need to be removed?

Not always — small tags like His6 are usually left on. When the tag interferes with activity, crystallization or immunogenicity, engineer a protease site (TEV, thrombin, PreScission/HRV-3C, enterokinase, or SUMO protease) between the tag and target so it can be cleaved off after purification.

How big is a His tag?

A 6×His tag is six histidine residues, about 0.84 kDa. His8 (eight histidines, ~1.1 kDa) binds IMAC resin more tightly and is useful for low-expression or membrane proteins.

Why are GST, MBP and SUMO shown without a full sequence?

These are whole proteins (hundreds of residues) fused as solubility/affinity partners, so the table lists "see notes" plus an approximate mass and the UniProt/vector source instead of inlining the full sequence. Retrieve the exact sequence from the cited UniProt accession or vector map.

Learn more

Sources

  1. 1
    UniProtKB reviewed (Swiss-Prot) entries for the five fusion partners: P08515, P0AEX9, Q12306, P0AFF6, P0AA25
    UniProt Consortium (EMBL-EBI / SIB / PIR) · 2026
    The Length and Mass (kDa) columns for all five 'see notes' fusion rows. I fetched each accession's flat file individually (https://rest.uniprot.org/uniprotkb/<acc>.txt): P08515 GST26_SCHJA Schistosoma japonicum GST = 218 aa, 25,499 Da (page: 218 aa, 26 kDa — matches); P0AEX9 MALE_ECOLI = 396 aa, 43,388 Da, SIGNAL 1-26, mature CHAIN 27-396 = 370 aa (page: 396 aa — matches the precursor, and the page's own caveat '~366-396 depending on signal-peptide/linker variant' is correct); Q12306 SUMO_YEAST Smt3 = 101 aa, 11,597 Da, PROPEP 99-101 removed leaving mature CHAIN 2-98 ending at Gly98 (page: 98 aa, 11 kDa — this is the MATURE form, and the page's note that Ulp1 'cleaves after the C-terminal Gly-Gly' is exactly the 98/101 boundary UniProt annotates); P0AFF6 NUSA_ECOLI = 495 aa, 54,871 Da (page: 495 aa, 55 kDa — matches); P0AA25 THIO_ECOLI TrxA = 109 aa, 11,807 Da, INIT_MET removed, CHAIN 2-109 (page: 109 aa, 12 kDa, note '~11.7 kDa' — matches to rounding). Also backs every accession the Notes column cites. ONE MISMATCH to fix or footnote: the page's MBP mass of 42.5 kDa matches neither the UniProt precursor (43.4 kDa) nor the mature MalE (~40.7 kDa) — see `unsourced`.
  2. 2
    A Short Polypeptide Marker Sequence Useful for Recombinant Protein Identification and Purification
    Hopp TP, Prickett KS, Price VL, Libby RT, March CJ, Cerretti DP, Urdal DL, Conlon PJ · Bio/Technology (now Nature Biotechnology) 6:1204-1210 · 1988
    The FLAG row in full: the eight-residue sequence DYKDDDDK (Asp-Tyr-Lys-Asp-Asp-Asp-Asp-Lys), length_aa 8, the 'hydrophilic, often surface-exposed' note, and the Cleavable by entry 'Enterokinase (cleaves after DDDDK)'. It is also the parent design the 3xFLAG row's terminal DYKDDDDK repeat derives from (but NOT the 3xFLAG junction residues — see `unsourced`).
  3. 3
    Genetic Approach to Facilitate Purification of Recombinant Proteins with a Novel Metal Chelate Adsorbent
    Hochuli E, Bannwarth W, Döbeli H, Gentz R, Stüber D · Bio/Technology (now Nature Biotechnology) 6:1321-1325 · 1988
    The His6 and His8 rows' underlying principle: a genetically fused poly-histidine peptide captured on a nitrilotriacetate (NTA) metal-chelate adsorbent — i.e. the 'Ni-NTA / Co (TALON) IMAC resin' purification column and the notes' 'Binds immobilized Ni2+/Co2+'. It does NOT back the specific 6x vs 8x choice, the imidazole/low-pH elution recipes, or the His8 'higher avidity / useful for membrane proteins' claim, all of which are vendor practice.
  4. 4
    Purification of a RAS-responsive adenylyl cyclase complex from Saccharomyces cerevisiae by use of an epitope addition method
    Field J, Nikawa J, Broek D, MacDonald B, Rodgers L, Wilson IA, Lerner RA, Wigler M · Molecular and Cellular Biology 8(5):2159-2165 (PMID 2455217, PMC363397) · 1988
    The HA (hemagglutinin) row's origin as a TAG: the first use of the influenza-HA peptide epitope as a genetically added purification/detection handle ('epitope addition'), which is what the row's use in Western/IP/IF rests on. The nine-residue sequence YPYDVPDYA itself and its HA1 ~98-106 numbering are better attributed to the paper this one builds on — Wilson IA, Niman HL, Houghten RA, Cherenson AR, Connolly ML, Lerner RA, 'The structure of an antigenic determinant in a protein', Cell 37(3):767-778, 1984, PMID 6204768, https://doi.org/10.1016/0092-8674(84)90412-4 — whose abstract I verified states that the anti-peptide monoclonals recognise 'one specific nine amino acid sequence' in influenza hemagglutinin. Wilson and Lerner are co-authors on both, so the pairing is the real lineage, not a guess.
  5. 5
    Isolation of monoclonal antibodies specific for human c-myc proto-oncogene product
    Evan GI, Lewis GK, Ramsay G, Bishop JM · Molecular and Cellular Biology 5(12):3610-3616 (PMID 3915782, PMC369192) · 1985
    The c-Myc row: the origin of the 9E10 monoclonal named in the Purification/detection column, and of the synthetic c-myc peptide immunogen from which the EQKLISEEDL tag is taken. The verified abstract states the antibodies were raised 'from mice immunized with synthetic peptide immunogens whose sequences are derived from that of the human c-myc gene product'.
  6. 6
    Identification of an epitope on the P and V proteins of simian virus 5 that distinguishes between two isolates with different biological characteristics
    Southern JA, Young DF, Heaney F, Baumgärtner WK, Randall RE · Journal of General Virology 72(7):1551-1557 · 1991
    The V5 row: the Pk epitope shared by the P and V proteins of simian virus 5 (a paramyxovirus), from which the 14-residue GKPIPNPLLGLDST tag derives — i.e. the row's sequence, length_aa 14, and the note 'Derived from the P/V proteins of simian virus SV5'. It does NOT back the 'common in Invitrogen/Thermo vectors (e.g. pcDNA)' or 'low background' claims, which are vendor practice.
  7. 7
    The Strep-tag system for one-step purification and high-affinity detection or capturing of proteins
    Schmidt TGM, Skerra A · Nature Protocols 2(6):1528-1535 (PMID 17571060) · 2007
    Most of the Strep-tag II row, and it is the only source here whose abstract contains the tag sequence literally: 'The Strep-tag II is an eight-residue minimal peptide sequence (Trp-Ser-His-Pro-Gln-Phe-Glu-Lys)' = WSHPQFEK, confirming the sequence and length_aa 8; plus affinity chromatography 'on a matrix carrying an engineered streptavidin (Strep-Tactin)' and the StrepMAB antibodies, i.e. the Purification/detection column. Note Schmidt is at IBA GmbH, which is the trademark holder the Notes column names. It does NOT back the 'Kd ~1 uM to Strep-Tactin' figure or the Twin-Strep-tag linker sequence — see `unsourced`.
  8. 8
    HaloTag: a novel protein labeling technology for cell imaging and protein analysis
    Los GV, Encell LP, McDougall MG, et al. (Promega Corporation) · ACS Chemical Biology 3(6):373-382 (PMID 18533659) · 2008
    The HaloTag row's mechanism claims, verbatim from the verified abstract: 'a modified haloalkane dehalogenase designed to covalently bind to synthetic ligands', ligands comprising 'a chloroalkane linker attached to a variety of useful molecules, such as fluorescent dyes, affinity handles, or solid surfaces', and bond formation that 'is essentially irreversible' — which is the row's COVALENT capture claim and the 'IRREVERSIBLE covalent bond' note, including protein immobilization. It does NOT back the 33 kDa / 297 aa figures in the Mass and Length columns (see `unsourced`).
  9. 9
    A general method for the covalent labeling of fusion proteins with small molecules in vivo
    Keppler A, Gendreizig S, Gronemeyer T, Pick H, Vogel H, Johnsson K · Nature Biotechnology 21(1):86-89 (PMID 12469133) · 2003
    The SNAP-tag row's mechanism: the founding paper for covalent self-labeling of fusion proteins via human O6-alkylguanine-DNA alkyltransferase (hAGT) reacting with O6-benzylguanine substrates — i.e. the row's 'Engineered human O6-alkylguanine-DNA alkyltransferase (hAGT)... covalently reacts with benzylguanine substrates' note and the BG-ligand Purification/detection column. The engineering of the hAGT mutants that made it practical is the companion paper Juillerat A et al., Chem Biol 10(4):313-317, 2003, PMID 12725859, https://doi.org/10.1016/s1074-5521(03)00068-1 (verified), if you want the directed-evolution half. Neither backs the 20 kDa / 182 aa figures (see `unsourced`).

Related tools and references