CRISPR On-Target Score - Rule Set 3 Guide Efficiency Online
Rank SpCas9 guides with Rule Set 3, the 2022 successor to Doench 2014 - and pick the tracrRNA you will actually use, because the model treats it as a feature.
🔒 Nothing you paste is logged or stored on our servers
Predicted, not measured
- How good is it?
- Held out six datasets (23,629 context sequences) from training; Rule Set 3 (Sequence) had the highest Spearman correlation on three of the six. On a separate tiling library generated for the paper, with every spacer any model had seen removed, it significantly outperformed all other models (p < 0.002) when the correct tracrRNA was given. Calibration on that independent set: of the lowest-scoring guides, 87.7% / 74.9% / 82.0% landed in the bottom two activity quintiles (Hsu / Chen / DeWeirdt tracrRNA), and of the highest-scoring, 77.0% / 69.0% / 75.5% landed in the top two. No single held-out Spearman is quoted here because the paper reports it per gene as a distribution rather than as one figure, and inventing a headline number from a figure would be the fit-residual mistake in a different costume.
- Only valid for:
- SpCas9 with an NGG PAM and a 20 nt spacer, scored over a 30-mer (4 nt upstream + spacer + PAM + 3 nt downstream) — nothing is padded, so a guide without that flank is skipped rather than estimated. Knockout activity in mammalian pooled screens with a U6/Pol III promoter; the authors expect it to generalise less well to Pol II-transcribed sgRNAs, and in vitro transcribed sgRNAs (as used in zebrafish) are known to be poorly predicted by models trained this way. The tracrRNA must be the one you will actually use. This is Rule Set 3 (Sequence) only: the Sequence+Target model, which adds 6.7% to the Spearman correlation on average, needs per-gene conservation and protein-domain lookups and is not implemented here.
- Fitted on:
- 46,526 unique 30-mer context sequences from seven pooled CRISPR-Cas9 screening datasets in mammalian cells (45% using the Chen 2013 tracrRNA), fitted as a LightGBM gradient-boosting regressor over the Rule Set 2 feature set plus nucleotide-run length, sgRNA:DNA melting temperature, folded-spacer minimum free energy, and tracrRNA identity. Activity is z-scored dropout in a viability screen — i.e. the likelihood of disrupting protein function, not a cutting rate.
Most on-target scores you will find online are Doench 2014 Rule Set 1 or CRISPRscan, both from before 2016. Rule Set 3 is the current model from the same lab: a gradient-boosting ensemble fitted on 46,526 guides across seven pooled screens, and the first one to treat the tracrRNA as a feature rather than an afterthought. That last point is the one to act on. The scaffold in lentiCRISPRv2 (Hsu 2013) and the sgRNA(F+E) scaffold used by the Sanger libraries (Chen 2013) give measurably different activity for the same spacer, and specifying the wrong one makes the prediction worse - so the selector here is not cosmetic. The number returned is a z-score, which ranks guides against each other; it is not a percentage of edited alleles and this page will not pretend otherwise. Doench 2014 and CRISPRscan are shown beside it rather than blended in, because a composite of three models is a fourth model nobody validated.
0 bases. A guide needs 4 bases before its spacer and 3 after its PAM to be scored — guides without that flank are skipped rather than scored on a padded window.
Not cosmetic — Rule Set 3 models the scaffold as a feature, and the paper measures its accuracy dropping when the wrong one is given.
Paste a target region and score it to rank the guides in it.
How to use the CRISPR On-Target (Rule Set 3) tool
- 1Paste the region you intend to cut, with at least 4 bases before and 3 after every guide you care about - guides without that flank are skipped rather than estimated on a padded window.
- 2Choose the tracrRNA your vector actually uses. Hsu2013 for lentiCRISPRv2 and most published libraries; Chen2013 for sgRNA(F+E).
- 3Read the ranking, not the absolute number. The spread between the best and worst guide is what tells you whether the ranking is worth acting on.
- 4Check the guide you pick for off-targets separately - this model says nothing about them.
Frequently asked questions
How is this different from the Doench 2014 score on the guide designer?
Rule Set 3 is the same lab's current model, fitted on far more data and on a feature set that includes the tracrRNA, nucleotide-run length, sgRNA:DNA melting temperature and folded-spacer free energy. Doench 2014 Rule Set 1 is a position-weight table from 1,841 guides. Both are shown here so you can see when they disagree, which happens often enough to be worth looking at.
Why does the tracrRNA change the score?
Because it changes the activity. The Hsu 2013 scaffold has a run of four thymidines that can terminate Pol III transcription early, so spacers ending in T are measurably less active with it than with the Chen scaffold. Rule Set 3 learned that interaction, and the paper measures its accuracy dropping when the wrong tracrRNA is specified.
What does the number mean?
A z-scored activity in the model's own training distribution - higher is better, and the useful comparison is between guides in the same list. On the paper's independent validation set, 87.7% of the lowest-scoring guides landed in the bottom two activity quintiles and 77.0% of the highest-scoring landed in the top two (Hsu tracrRNA). It is not a percentage of alleles edited.
Does it work for Cas12a, SaCas9 or a non-NGG PAM?
No. Rule Set 3 was fitted on SpCas9 with an NGG PAM and a 20 nt spacer, and nothing else is scored. The guide designer will still find candidates for the other nucleases; it just cannot score them with this model.
Is this the full Rule Set 3?
It is Rule Set 3 (Sequence). The published model has a second half, Rule Set 3 (Sequence + Target), that adds conservation and protein-domain features around the cut site and improves the correlation by about 6.7% on average. That half needs per-gene lookups against Ensembl and the UCSC browser and is not implemented here.
More
Related tools
Pick a high-fidelity set of 4-base fusion sites for an N-part assembly, pinning the overhangs your vector already commits you to.
Choose the smallest set of sequencing primers that can tell your candidate molecules apart, and find out which pairs nothing will separate - before you pay for the reads.
Calculate GC%, AT% and per-base composition for DNA or RNA.