SeqBench

Base Editing Quantification from Sanger — CBE and ABE Efficiency Without NGS

Load an unedited control trace and your edited pool's trace and get per-position C→T or A→G percentages, their z-scores against your own run's background, and the detection limit those numbers sit on.

🔒 Nothing you paste is logged or stored

Predicted, not measured
How good is it?
None is published for this implementation. Every run instead reports what it rests on: the background mean, sd, robust (MAD-based) sd, outlier count and n, and a noiseFloorPercent that says how much apparent editing the background alone reaches at the significance threshold — null, rather than 0, when the run has no null with any spread to derive a limit from. This implementation's recovery of known synthetic mixtures (to under a percentage point) is deliberately NOT offered as validation — it tests the arithmetic and the coordinate handling, not whether the linear mixture model fits a real capillary trace. Two properties ARE characterised. The bias is directional and one-sided: the (1 − b) rescale assumes a fully converted position would read as alt fraction 1.0, which real chemistry does not reach, so percentages run low by roughly the crosstalk fraction (order 5-10% relative at typical 3% bleed). The variance is not constant across positions: dividing by (1 − b) amplifies the error in a and b by 1/(1 − b), so the variance of the estimate goes as 1/(1 − b)². Positions whose control alt fraction b exceeds 0.25 are therefore reported but NOT quantified, which bounds that amplification at 1.33x on everything the tool does quantify.
Only valid for:
A pool edited by a cytosine or adenine base editor, read on the same amplicon and chemistry as an unedited control that starts within ±40 bases of it, where the editing is a SUBSTITUTION. Not valid for indels — a base editor also makes them, and an indel-bearing allele shifts the downstream trace so that it degrades the fit at every window position rather than showing up anywhere. Not valid where the control read already carries the alt base at a window position (a pre-existing variant, or the wrong control), since then there is nothing left to correct against — that condition is DETECTED rather than only described: a window position whose control alt fraction exceeds 0.25 is reported with an excludedReason and an editedPercent of 0 instead of a number, warned about, and failed by the hard control-supports-the-window-base-calls gate check. Reported percentages are pool-level: Sanger sees the superposition, so which allele carries which combination of bystander edits is not recoverable from it at all.
Fitted on:
Nothing. There are no fitted coefficients: the background subtracted at each position is the caller's OWN control trace at that same position, and the significance threshold is a z-score against a null built from the caller's own read (every control position outside the editing window carrying the same base). The model — that a pool trace at a substituted position is a linear mixture of the unedited and converted peak profiles, so the corrected alt-channel fraction is the edited fraction — is the one EditR uses (Kluesner et al., The CRISPR Journal 2018).

A base editor does not cut, so an edited pool reads as a mixed peak at one position rather than as a shifted trace — which is why an indel decomposition returns a confident zero on a CBE or ABE experiment. This tool reads the mixture instead: at every editable position in the activity window it takes the edited trace's alt-channel fraction a, subtracts the control's own alt fraction b at that same position as background, and rescales by (1 − b) to give the percentage of the pool carrying the conversion — the intended edit and its bystanders alike. Significance is a z-score against a null built from your own run (every control position outside the window carrying the same base), so the threshold follows your chemistry rather than a hardcoded cut-off, and each run reports the noise floor that null implies: the percentage the background alone reaches, below which a small number is not a small edit.

Both traces must be the same amplicon on the same chemistry; the control is what the background is subtracted from.

In standard deviations of this sample's own background. Raising it raises the reported noise floor with it.

Edited position p matches control position p + offset.

How to use the Base Editing Quantification tool

  1. 1Load the unedited control trace and the edited pool's trace, both .ab1 — or click Load example for a synthetic BE4max pair with two known conversion levels.
  2. 2Pick the base editor (BE3, BE4max, ABE7.10 or ABE8e) and paste the protospacer to have its activity window located in the control read; a spacer found on the reverse strand is handled, and a CBE's C→T then reads as G→A.
  3. 3Or switch to Explicit window and give the window in control coordinates plus the conversion as it is read on that strand, for an editor this registry does not list.
  4. 4Name the intended edit — a protospacer position, or a control position — so the target is labelled by your design rather than by whichever position came back highest.
  5. 5Read the per-position percentages against this run's own detection limit, and check the gate first: it reports whether the background rested on enough positions, whether the two reads agree outside the window, and whether the control supports its own base calls.

Frequently asked questions

How can a Sanger trace measure base editing at all?

Because the trace of a pool is every allele superposed. At a position the editor converted, part of the pool still reads as the original base and part reads as the new one, so the peak is mixed and the fraction of the signal in the converted base's channel is the fraction of the pool carrying the edit. Two corrections make that usable. The control's own signal in that same channel is subtracted first, because dye crosstalk is position- and context-dependent and a single global constant would be wrong per position; then the difference is divided by (1 − b), since an unedited allele already contributes b to that channel. This is the model EditR introduced (Kluesner et al., The CRISPR Journal 2018).

Why can't an indel decomposition measure a base edit?

Because a base editor makes a mixed base, not a shift. TIDE-style decomposition works by fitting the edited trace onto shifted copies of the control, and a substitution shifts nothing — it contributes to no shifted column at all, so the fit comes back at roughly zero per cent edited with a high R², which looks like a confident negative rather than a mismatch of method. The two tools are complements on the same pair of traces: this one for the conversions, the CRISPR Editing Efficiency tool for the indel fraction a base editor also produces through nicking and repair.

How accurate are the percentages?

No accuracy figure against amplicon sequencing is published for this implementation, and none is quoted here. Every run instead reports what the number rests on: the background's mean, sd, MAD-based robust sd, outlier count and n, plus the noise floor those imply. Two properties are characterised rather than corrected. The bias is one-sided and low — the (1 − b) rescale assumes a fully converted position would read as alt fraction 1.0, which real chemistry does not reach, so percentages run low by roughly the crosstalk fraction (order 5–10% relative at a typical 3% bleed). The variance is not constant across positions: dividing by (1 − b) amplifies the error in a and b, which is why a position whose control alt fraction exceeds 0.25 is not quantified at all — that cap bounds the amplification at 1.33× on everything that is. This implementation does recover known synthetic mixtures closely, and that is deliberately NOT offered as validation: it tests the arithmetic and the coordinate handling, not whether a linear mixture model fits a real capillary trace.

What is the noise floor, and why does every run have its own?

It is the editing percentage the background alone reaches at your significance threshold — this run's detection limit. It is derived from your own two reads rather than from a table, because crosstalk depends on the chemistry, the instrument, the dye lot and how hot the run was. Anything below it is not a small edit; it is indistinguishable from crosstalk. When the null has no spread at all — synthetic data, or the same file given twice — no floor is reported and no z-score is computed, because a stated detection limit of 0% is the one number a reader would use to dismiss a small result.

Can it tell me which allele carries which bystander edits?

No, and nothing that reads a Sanger trace can. The trace is the superposition of the whole pool, so per-position percentages are all it supports: two positions at 40% could be one 40% double-edited allele or two disjoint 40% populations, and no arithmetic on this data separates them. The bystander total the tool reports sums per-position percentages and is a burden indicator rather than a share of the pool — it can legitimately exceed 100%. Haplotypes need amplicon sequencing.

What if my control already reads as edited at a window position?

That position is reported but not quantified, and it fails a hard gate check. Above a control alt fraction of 0.25 there is too little of (1 − b) left for the rescale to mean anything — at b = 0.998 a 0.1-point difference becomes half the pool — and the usual causes are a pre-existing variant in your line or the wrong control trace, neither of which is background to subtract. Its reported 0% is the absence of an estimate, not a measurement of no editing, and the run warns you so.

What does the control trace have to be?

The same amplicon on the same chemistry, unedited, starting within 40 bases of the edited read — the offset between the two is found from the base calls outside the window and reported, along with the base-call identity at the offset actually used. Below 90% identity the gate fails: a mismatched control makes its alt-channel value the wrong background to subtract. The null also needs at least 10 control positions outside the window carrying the base the editor acts on, so a very short read over a base-poor amplicon produces percentages whose significance the run cannot test — that too is a gate failure rather than a silent pass.

Does the gate tell me whether the edit worked?

No, deliberately. Every check is a fact about the run — enough background positions, the same amplicon, the window actually read, the control supporting its own base calls — and none of them touches how much editing there was. Whether an efficiency is high enough is a judgement about your experiment, and gating on the magnitude of an estimated allele fraction that carries a documented one-sided bias and a per-run detection limit would be putting a verdict on exactly the number that cannot support one.

More