Synthesis Complexity Checker — What a Gene Synthesis Vendor's Filter Will Measure
Measure the repeats, GC extremes, GC transitions, homopolymer runs and hairpins that make a fragment hard to synthesise — each against the threshold vendors publish, with no score and no probability.
🔒 Nothing you paste is logged or stored on our servers
A synthesis order that comes back delayed or cancelled usually fails on something measurable, and the measurement is cheap to make before you pay for it. This reports every determinant the major vendors' design guides name — repeated stretches with all their coordinates, the lowest and highest GC windows, the largest GC transition between adjacent windows, homopolymer runs, and inverted repeats with the melting temperature of each stem — and puts each one next to the threshold it is being compared against, so a reader who disagrees with a threshold can still use the number. Repetition leads the output because the published feature-importance analysis of DNA synthesis outcomes puts it first among all causes, and because it is the one you can actually fix: the Non-Repetitive Parts designer builds replacements that share nothing. What this deliberately does not give you is a success probability. The published classifier for that question was trained on 1,076 sequences of which only 303 were real orders — the rest are 373 negative controls designed to violate vendor filters and 400 positive-control subsequences of fragments already known to have worked — split by class rather than by source, so its headline accuracy establishes that deliberately-broken sequence separates from known-good sequence. That is not the question you have about your own gene, and its labels are one vendor's turnaround times from before 2020, which age in the direction of condemning sequences that are now fine.
Paste the sequence you are about to order to see what a vendor's filter will measure.
How to use the Synthesis Complexity Checker tool
- 1Paste the fragment exactly as you would order it — raw sequence, FASTA or GenBank.
- 2Press Measure. Every measurement is reported with the published vendor threshold beside it.
- 3Start with the repeats: they are the largest single cause of synthesis failure and the only class you can systematically design out.
- 4Feed the repeated regions to the Non-Repetitive Parts designer to build replacements, then re-measure.
Frequently asked questions
Why is there no synthesis success score?
Because the number would not mean what it looks like. The published model for this is a classifier trained on a set that is about 72% synthetic control sequences — 373 designed on purpose to violate vendor filters and 400 taken from fragments already known to have synthesised — with a split by class rather than by source, so its held-out set has the same composition. Its accuracy therefore establishes that deliberately-broken sequence separates from known-good sequence, which is not the question a real order asks. Its labels are also one vendor's turnaround times from before 2020, and vendors have improved since, so the number would age towards false alarms. The measurements underneath it are exact and are what this tool reports.
Do vendors not already screen my sequence for free?
Yes — Twist, IDT and GenScript all run a complexity check before you pay, and for an ordinary gene that check is the authority. This is useful earlier: while you are still choosing between designs, when you want to know WHICH feature is the problem rather than that there is one, and when a sequence is refused by more than one vendor and you need to see what they are all seeing.
What counts as a repeat here?
An exact stretch of at least the length you set (twenty bases by default, which is where vendor repeat filters sit) that appears more than once, on either strand. Every occurrence is reported together rather than as pairs, because a stretch appearing five times is one thing to fix. The percentage of the sequence covered by repeats is reported too, since a short repeat many times over is a different problem from one long one.
Are the hairpins a folding prediction?
No. They are inverted repeats — stretches that are reverse-complementary to something nearby — with the nearest-neighbour melting temperature of each stem. That is arithmetic over the letters. Whether a given stem actually forms at a given temperature is what the RNA Folding tool answers, and that tool is marked as predictive precisely because this one is not.
Where do the thresholds come from?
They are the limits the major synthesis vendors publish in their own design guidelines, which agree with each other to within a few points. They are conventions, not a fitted decision boundary, and they are reported alongside every measurement so you can disagree with a threshold without having to disagree with the number. Different vendors and different products genuinely differ.
More
Related tools
Locate the exact direct repeats that let a construct recombine away the DNA between them, and build the shortened molecule you would actually recover.
308 publicly deposited vectors with their full GenBank feature tables, and 617 parts harvested from them.
Digest level-0 part plasmids with a Type IIS enzyme and let the overhangs decide the assembly order.