SeqBench

Promoter Calculator — Sigma-70 Promoter Strength and Library Design

Find every sigma-70 promoter in a region on both strands, with the free-energy terms behind each transcription rate — then build a library of promoters spanning a range.

🔒 Nothing you paste is logged or stored on our servers

Predicted, not measured
How good is it?
R^2 = 0.45 and 0.60 against the two INDEPENDENT in vivo datasets the authors tested (Hossain et al., 4,350 promoters, Spearman rho = 0.69; Urtecho et al., 10,898 promoters, rho = 0.67). The widely quoted R^2 = 0.80 is a held-out tenth of the authors' OWN in vitro transcription data and is not the number to plan against: a promoter in a cell is the in vivo case, where between a third and a half of the variance is unexplained.
Only valid for:
sigma-70 (housekeeping) promoters in E. coli. NOT valid for another sigma factor, another organism, a promoter under activator or repressor control, or anything about mRNA stability or translation — rbs_predict is the translation half, and neither speaks to the other.
Fitted on:
4,673 E. coli sigma-70 promoters (90% of a filtered library) measured by in vitro transcription, fitted as a biophysical model with 346 ridge-regression coefficients over the -35 box, the -10 box, the spacer, the discriminator, the extended -10 and the initial transcribed region. La Fleur, Hossain & Salis, Nat Commun 13:5159 (2022).

Two questions about promoters have different answers and the same starting point. The first is about a promoter you already have: how strong is it, and why — which box is doing the work, is the spacer the wrong length, is the discriminator costing you. The second is about a set you do not have yet: give me eight promoters spanning three orders of magnitude so I can titrate this pathway. This page does both. The scan reads both strands, which matters more than it sounds: a promoter pointing backwards inside a cassette transcribes antisense RNA against everything upstream of it and knocks expression down in a way that looks like a cloning problem, and no forward-only tool will ever show it to you. Each hit comes with the free-energy decomposition the model is built from, so a weak promoter tells you which term to change rather than only that it is weak. The rates are a MODEL's estimate, not a measurement — R² = 0.45 to 0.60 against the independent in vivo datasets the authors tested, which is enough to rank these promoters against each other and not enough to plan a yield around, and the page says so where the numbers are.

Paste a region to see the sigma-70 promoters in it, on both strands.

Design a non-repetitive promoter library

Scoring a promoter you already have is one question; getting a SET of them that span a range is the other, and it has a constraint nobody expects — a library whose members share long stretches recombines with itself once it is integrated, losing members silently. This builds promoters that span a predicted-rate range AND share no long stretch with each other.

Sigma-70 promoters that span a range of predicted rates AND share no long stretch with each other — the second property is what stops a library recombining with itself once it is integrated, and it is invisible in any per-promoter check.

How to use the Promoter Calculator tool

  1. 1Paste the region — a promoter and its context, a 5' UTR, or a whole cassette you want swept.
  2. 2Read the table: TSS, strand, the two boxes, the spacer length and the estimated rate.
  3. 3Check the antisense row count. A reverse promoter inside a construct is the finding people most often miss.
  4. 4For a set rather than a score, use the library designer below — it spans a rate range AND keeps the members from recombining with each other.

Frequently asked questions

How accurate is the transcription rate?

The model is a 346-coefficient biophysical fit over the −35 box, the −10 box, the spacer, the discriminator, the extended −10 and the initial transcribed region (La Fleur, Hossain & Salis, Nat Commun 2022). Its R² is 0.80 on held-out data from the authors' own IN VITRO transcription experiments, and 0.45 to 0.60 against the two INDEPENDENT IN VIVO datasets they tested. A promoter in a cell is the in vivo case, so plan against 0.45–0.60: the ranking between two promoters is usable, the absolute number is not.

Why does it scan the reverse strand?

Because that is where the expensive surprise lives. A sigma-70 promoter pointing backwards inside your cassette transcribes antisense RNA across everything upstream of it, which reads as poor expression, a bad clone or a plasmid problem — and it is completely invisible to a forward-only scan. Coordinates are reported on the forward strand with the strand named, so a hit can be found in your own map.

It found no promoter. Does that mean there is none?

It means no SIGMA-70 promoter, in E. coli. A T7 promoter, a promoter read by another sigma factor, a eukaryotic promoter, or one that needs an activator bound are all outside what this model reads. It is also worth saying the obvious: a sequence can carry a perfectly good promoter that this model scores low, which is what an R² of 0.45–0.60 means in practice.

Why does a promoter library need to be non-repetitive?

Because a library of twelve promoters that share long stretches is twelve recombination substrates in one molecule. Once it is integrated, members are lost silently — and your screen comes back with some levels missing, which reads as biology. The library designer selects members that span a predicted-rate range AND share no more than a length you choose, and it reports the rungs it could not fill rather than padding them out.

Can I get the free-energy terms from the API?

Yes. promoter_predict returns every term — total, binding, UP element, −10, −35, spacer, discriminator, extended −10 and initial transcribed region — over the REST API, the MCP server and batch, and a missing term is reported as null rather than as zero, because those are different facts.

More