RBS Designer — Translation Initiation Rate Prediction & 5' UTR Design
Predict the translation initiation rate at every start codon, and design a 5' UTR to hit a target expression level, with OSTIR and ViennaRNA.
Predicted, not measuredSpearman ρ = 0.39 on two 5' UTR datasets it was not fitted to
- How good is it?
- Spearman ρ = 0.39 against measured expression on two 5' UTR datasets it was not fitted to (Gilliot & Gorochowski, Nucleic Acids Res 2024;52(13):e58); reproduced in this repo at ρ = 0.509 on an independent leave-one-context-out split over 194,636 measurements, and at a mean ρ = 0.62 over fifteen further measurement systems it was not fitted to — six species, four reporters, three readouts — on which every learned model tried here scored BELOW it (scripts/rbs-eval) — above the published figure because 65,841 rows of that compilation record no measurement at all (fluorescence mean and s.d. both exactly 0) and are excluded here rather than scored as the weakest expression observed. The widely quoted 53% within 2-fold / 91% within 10-fold are calibration residuals on the fitting set, not held-out validation.
- Only valid for:
- translation INITIATION only, in E. coli-like Gram-negative hosts (the model is parameterized on the E. coli anti-Shine-Dalgarno sequence). Rankings within one construct context; the absolute value has no units and no meaning.
- Fitted on:
- A thermodynamic model whose coefficients OSTIR refitted against ViennaRNA's Turner 2004 energies over the 132 measurements of Salis et al. 2009 (Nat Biotechnol 27:946).
Paste a bacterial mRNA and get the predicted translation initiation rate at every start codon it contains, with the complete free-energy breakdown behind each number: how well the Shine-Dalgarno region hybridizes to the 16S rRNA, what it costs to unfold the mRNA structure occluding the site, the penalty for non-optimal SD-to-start spacing, the standby-site term, and initiator-tRNA binding. Switch to design mode and give it a coding sequence instead: it builds a spread of Shine-Dalgarno cores and SD-to-start spacings, scores every one against your own CDS — which matters, because the rate depends on how the RBS interacts with that specific CDS's 5' folding and can't be read off a parts table — and ranks them. Supply a target rate to rank by closeness instead of raw strength, or your existing 5' UTR to get a measured baseline and a fold-change for every candidate. The predictions come from OSTIR with ViennaRNA free energies, not from a lookup table of published part strengths.
Design, edit, and verify complete constructsSeqStudio combines sequence editing and annotation, plasmid maps, primer/cloning/CRISPR design, Sanger verification, and GenBank/SnapGene files in one workspace.
How to use the RBS Designer tool
- 1Pick a mode: "Predict rate" to score an existing mRNA, or "Design an RBS" to generate and rank candidates for a CDS.
- 2Paste your sequence — an mRNA (5' UTR plus the start of the CDS) to predict, or a CDS beginning at its own start codon to design. DNA and RNA spellings are both accepted.
- 3Optionally add your current 5' UTR for a baseline fold-change, a target rate to design toward, your promoter's real transcribed leader, or a non-E. coli anti-Shine-Dalgarno sequence.
- 4Read the ranked candidates and the ΔG breakdown, and copy the top candidate's full 5' UTR to clone.
Frequently asked questions
Which model produces these numbers?
OSTIR (Roots, Lukasiewicz & Barrick, Journal of Open Source Software 2021), which continues the last open-source release of the Salis lab's RBS Calculator and re-fitted the thermodynamic model's coefficients against ViennaRNA's energy parameters. Because of that refit, OSTIR values are not interchangeable with numbers from RBS Calculator v2 — don't mix the two in one comparison. All free energies are computed by the ViennaRNA Package.
What are the units of the predicted rate?
There aren't any — it's an arbitrary scale. Ratios between two predictions are the meaningful quantity ('this candidate is about 8x the one I have now'), and the absolute number is not a protein concentration and can't be converted into one. That is a property of the model, not a limitation of this implementation.
How accurate is it?
There are two different numbers here and the difference matters. The figures usually quoted for OSTIR — 53% of measurements within 2-fold and 91% within 10-fold — are calibration residuals: they say how closely the fitted model reproduces the 132 measurements of Salis et al. 2009 that its own coefficients were derived from, which is not the same as how it performs on a sequence it has never seen. The number for that is a rank correlation of Spearman ρ = 0.39 against measured expression on two 5' UTR datasets it was not fitted to (Gilliot & Gorochowski, Nucleic Acids Res 2024). So: use it to rank candidates within one construct context, treat two candidates that score close together as indistinguishable, and don't read the absolute value as an expression level. We also measured what that rank correlation means for an actual choice between two candidates, which is the thing you are using it for: over fourteen published libraries, hosts and reporters spanning six bacterial species — and stated against the ratio of the two PREDICTED rates, which is the thing you can actually see — the stronger of a pair is ranked first 57% of the time when the two predictions are within 2-fold of each other, 72% at 3–5-fold, and 87% above 100-fold, where 50% is a coin toss. Below 1.2-fold apart the prediction carries essentially no information (53%). Individual systems vary widely around those means, so read them as a guide to when a gap stops being trustworthy rather than as a per-prediction probability. A series of designs sitting within about threefold of one another is therefore not reliably rankable — build it and measure it.
How are the design candidates generated?
By combining a spread of Shine-Dalgarno cores (from full complementarity to the anti-SD down to minimal) with SD-to-start spacings across the biologically relevant range, then scoring every combination with OSTIR in the context of your actual CDS. No strength is asserted for any candidate sequence in advance — the ranking comes entirely from the model. The spacer is poly-A by construction so that varying spacing doesn't also introduce new secondary structure; check the returned sequence if your cloning strategy needs a particular site inside the UTR.
Why does it ask for a 5' leader, and what happens if I don't give one?
The standby-site term depends on the sequence upstream of the RBS, so the model needs some 5' context. If you don't supply one, a 20 nt unstructured poly-A leader is assumed and the result says so. For a construct-specific number, paste the real transcribed leader your promoter produces.
Does it work for organisms other than E. coli?
Partially. The model is parameterized on E. coli, but you can supply a different anti-Shine-Dalgarno sequence (the 16S rRNA 3' end) for your host, which is the single largest host-specific term. Everything else in the model stays E. coli-derived, so treat a non-E. coli prediction as a ranking heuristic rather than a calibrated rate.
What does this NOT tell me?
It models translation initiation only. It says nothing about transcription, elongation, mRNA stability, protein folding or toxicity, codon usage, or whether the designed UTR introduces a restriction site, cryptic promoter or out-of-frame upstream start codon in your final construct — check the returned sequence for those separately. The Construct QC Linter covers several of them.
Can I run this from code?
Yes — both rbs_predict and rbs_design are callable from the REST API and the MCP server. Because each run does multi-second RNA folding on a shared service, they are rate limited; the error tells you how long to wait if you hit it.
More
Related tools
Find every sigma-70 promoter in a region on both strands, with the free-energy terms behind each transcription rate — then build a library of promoters spanning a range.
Find every stretch of a transcript that can pair over the Shine-Dalgarno sequence or the start codon, with the length and melting temperature of each duplex — the mechanism behind every translational switch, and the first thing to check when a construct is silent.
Score a coding sequence's codon usage against an expression host, before optimizing.