SeqBench

Protein Molecular Weight and pI Calculator

6 min read · Updated June 10, 2026

A chain of eight amino acid residues with two acidic (negative), two basic (positive), and four neutral side chains, shown alongside a 0-14 pH scale marking an illustrative isoelectric point (pI) of about 7.8.Charge distribution along a protein chain++acidic (Asp/Glu)basic (Lys/Arg)neutral0714pHpI ≈ 7.8Illustrative pI, not precisely calculated

A protein sequence quietly encodes a handful of physical properties you need every week in the lab: how heavy the protein is, what pH it carries no net charge at, and how strongly it absorbs UV light. This guide explains where molecular weight, isoelectric point and the extinction coefficient come from and how to put them to work.

Molecular weight

A protein's molecular weight is the sum of its amino-acid residue masses plus one water molecule (the residues are joined by losing water at each peptide bond). Using average isotopic masses gives the figure quoted by tools like ExPASy ProtParam, usually reported in daltons (Da) or kilodaltons (kDa).

Molecular weight tells you where a band should run on an SDS-PAGE gel, how to convert between mass and moles of protein, and roughly how big the molecule is. Bear in mind that post-translational modifications, bound cofactors and disulfide bonds can shift the real mass away from the sequence-based estimate.

Theoretical isoelectric point (pI)

The isoelectric point is the pH at which the protein carries no net charge. Below its pI a protein is net positive; above it, net negative. It is calculated by adding up the charge contributions of the ionisable groups — the N- and C-termini and the side chains of Asp, Glu, Cys, Tyr, His, Lys and Arg — at each pH using their pKa values, then finding the pH where the total is zero.

The pI guides ion-exchange chromatography (which resin and buffer pH to use), predicts where a protein may aggregate (solubility is often lowest near the pI), and helps interpret 2D gels and isoelectric focusing.

Extinction coefficient and protein concentration

Proteins absorb UV light at 280 nm mainly because of tryptophan and tyrosine (and, to a small degree, cystines). The molar extinction coefficient ε₂₈₀ can be estimated directly from the counts of these residues using the Edelhoch/Pace method: ε = 5500 × nTrp + 1490 × nTyr + 125 × n(cystine).

With ε you can turn an A280 reading into a concentration via Beer's law (A = ε × c × path length). The related A(0.1%) value — the absorbance of a 1 g/L solution in a 1 cm cuvette — lets you read concentration in mg/mL straight off the spectrophotometer: concentration (mg/mL) = A280 / A(0.1%).

GRAVY and other quick descriptors

The GRAVY score (grand average of hydropathy) averages Kyte–Doolittle hydropathy values across the sequence; a positive value suggests a more hydrophobic protein and a negative value a more hydrophilic one. Together with amino-acid composition and the counts of positively and negatively charged residues, these descriptors give a fast first impression of a protein before any wet-lab work.

Worked example: molecular weight, extinction coefficient and concentration from a sequence

To see where these numbers actually come from, work through a short 10-residue peptide: MAWYHKDECR (Met-Ala-Trp-Tyr-His-Lys-Asp-Glu-Cys-Arg).

  1. Sum each residue's average mass: Met 131.193 + Ala 71.079 + Trp 186.213 + Tyr 163.176 + His 137.141 + Lys 128.174 + Asp 115.089 + Glu 129.116 + Cys 103.139 + Arg 156.188 = 1320.51 Da
  2. Add one water molecule for the whole chain: 1320.51 + 18.02 ≈ 1338.52 Da (≈1.34 kDa) — the peptide's molecular weight
  3. Count Trp, Tyr and cystines for the extinction coefficient: 1 Trp, 1 Tyr, 0 cystines (the lone Cys has no partner to pair with) → ε₂₈₀ = 5500×1 + 1490×1 + 125×0 = 6990 M⁻¹cm⁻¹
  4. Convert to A(0.1%): 6990 / 1338.52 ≈ 5.22 — the absorbance a 1 mg/mL solution of this peptide would show in a 1 cm cuvette

How pI is actually calculated — and why two tools can disagree

The pI calculation described above is really a search. At any given pH, every ionisable group — each terminus, plus the Asp, Glu, Cys, Tyr, His, Lys and Arg side chains — is mostly protonated or mostly deprotonated according to the Henderson–Hasselbalch equation, contributing a fractional +1, -1 or 0 to the net charge. A pI calculator scans (or bisects) across pH values until that sum crosses zero; the crossing point is the theoretical pI.

The logic is easiest to see with a single free amino acid, which has only a few ionisable groups. Aspartic acid has three: the α-carboxyl (pKa ≈ 1.88), the side-chain carboxyl (pKa ≈ 3.65), and the α-amino group (pKa ≈ 9.60). Below pH 1.88, both carboxyls are protonated and the amino group is protonated, giving a net charge of +1. Between pH 1.88 and 3.65, the more acidic α-carboxyl has deprotonated (-1) while the side-chain carboxyl is still neutral and the amino group is still +1 — net charge exactly 0. Above pH 3.65 both carboxyls are deprotonated (-2) against the still-protonated amino group (+1), giving -1, and above pH 9.60 the amino group deprotonates too, giving -2.

Because the neutral species sits in the pH range bounded by the two lowest pKa values, its pI is simply their average: (1.88 + 3.65) / 2 = 2.77 — the standard textbook pI of aspartic acid. A full-length protein has many more ionisable groups and rarely has one single pH range that's exactly neutral, so instead of averaging two pKas, software sums the charge contribution of every group at a trial pH and narrows in on wherever the total hits zero. Same underlying logic, just automated across dozens of groups instead of three.

That's also why two calculators can report noticeably different pI values for the same protein: they may use different pKa tables (the EMBOSS, ExPASy/Bjellqvist and Sillero–Ribeiro sets are all in common use and disagree by a few tenths of a pH unit per group). Neither is 'wrong' — both are self-consistent estimates from the sequence alone, not measurements. A protein's experimentally measured pI, from isoelectric focusing or a 2D gel, can differ further still, because folding, bound ligands and modifications such as phosphorylation or glycosylation shift the effective behaviour of buried or modified groups in ways a sequence-only calculation can't see.

Common mistakes when calculating protein weight, pI or concentration

  • Feeding in the full ORF translation — including a signal peptide, propeptide or affinity tag — when you actually want the mass and pI of the mature, processed protein. Trim the sequence to match whatever form you're actually measuring.
  • Plugging an A280 reading into Beer's law with the wrong extinction coefficient. The molar ε₂₈₀ (M⁻¹cm⁻¹) needs a molar concentration; A(0.1%) is the one you divide directly into mg/mL. Mixing the two up gives a concentration off by orders of magnitude.
  • Counting every cysteine as a cystine in the ε₂₈₀ formula. The 125 M⁻¹cm⁻¹ term applies per disulfide bond (a pair of linked Cys), not per free thiol — a protein with 4 cysteines forming 2 disulfides contributes 2 × 125, not 4 × 125.
  • Treating a small difference in pI — between two tools, or between a calculated and an experimentally measured value — as a meaningful biological finding, rather than the expected result of different pKa tables or of folding/PTM effects the calculation can't see.
  • Assuming a bound cofactor or prosthetic group (heme, FAD, a metal centre) doesn't affect A280 just because it isn't part of the amino-acid sequence. Many of these groups absorb near 280 nm too, so a sequence-only ε₂₈₀ can under- or overestimate concentration for a holoprotein.
  • Also worth remembering: all of the figures above are 'average mass' — each residue's mass averaged over its natural isotope mix, which is what ExPASy ProtParam and most protein calculators report by default, and the right figure for pipetting and dilutions. Mass spectrometry instead usually reports monoisotopic mass, built from only the single most abundant isotope of each atom. The two are close for a short peptide but drift further apart as a protein gets larger, so comparing an observed MS peak against the wrong kind of calculated mass is a common source of 'my protein doesn't match its predicted mass' confusion.

Frequently asked questions

Why does my protein run at a different size on a gel than its calculated weight?

SDS-PAGE migration depends on more than mass — highly charged, glycosylated or unusually shaped proteins can run faster or slower. The sequence-based molecular weight is the mass of the unmodified polypeptide.

How do I get concentration from A280?

Divide your A280 reading by the A(0.1%) value to get mg/mL (for a 1 cm path length), or use the molar extinction coefficient with Beer's law to get molar concentration. Choose the reduced or cystine extinction value depending on whether disulfides are formed.

Why do two protein calculators give slightly different pI or molecular weight for the same sequence?

Small differences are normal, not errors. pI differs between tools mainly because they use different pKa tables — EMBOSS, ExPASy/Bjellqvist and Sillero–Ribeiro are all in common use and disagree by a few tenths of a pH unit per ionisable group. Molecular weight differs mainly when tools disagree on average vs. monoisotopic mass, or on whether an initiator methionine or a signal peptide/tag was included in the sequence you actually pasted in.

Does a protein molecular weight or pI calculator account for my signal peptide or purification tag?

Only if it's part of the sequence you paste in — the calculation runs on exactly the residues you provide, nothing more. Trim off the signal peptide and any tag first if you want the mass and pI of the mature, secreted or cleaved protein; leave them in if you want the properties of the full expressed construct as it would run before cleavage.

Do disulfide bonds change a protein's calculated molecular weight?

Yes, slightly. Forming each disulfide bond oxidises two cysteine thiols into one S–S bond, removing two hydrogen atoms — about 2 Da lighter than the fully reduced mass a plain sequence calculation gives you. Most sequence-based tools report the fully reduced mass by default, so a protein with several intact disulfides will actually run and weigh a few daltons lighter than the calculator's number.

Related references

Related tools

Related guides