Case study · Selection analysis

Three targets, twenty-three days.

A research group wrote to us with sequencing from three aptamer selections. All three had enriched. None had resolved into anything they could confidently order. This is what we did with it.

End to end
23 days
Targets
3
Sequenced rounds
6
Reads analysed
5.38 M
The problem

A selection that works still leaves you guessing.

Three campaigns against different targets, all built on the same starting library, sequenced at two rounds each. The pools had clearly enriched. What the group could not do was turn that into a decision: which molecules to synthesise, in what order, and on what grounds.

This is the ordinary failure mode of a selection that succeeded. Abundance ranking hands you the winner and tells you nothing about why it won, whether it is one molecule or one molecule's sequencing noise, or whether the runner-up is a genuine alternative or an artefact. The gap is not in the wet lab. It is in the analysis.

How it ran

Same-day first pass, then the rest of the campaign.

DAY 0

First dataset in

Family sheets went back the same day — together with a correction to one of the construct parameters the group had given us.

DAY 4

First report out

Delivered more than three weeks ahead of the deadline they had set us.

DAY 7

They sent the rest

Within three days of first delivery the group committed every remaining dataset, and invited us to co-author the resulting publication.

DAY 23

Combined report

All three targets in one analysis: body, per-target annexes, a chain-by-chain appendix, and a machine-readable table of every family member.

What we actually did

Nine steps, each of which changed a number.

Verified the library from the reads

Construct parameters are checked against the data, not taken from the brief. Here the brief was wrong — and every downstream figure would have inherited the error.

Established strand orientation

Which read is the aptamer and which its complement, derived from the selection chemistry and confirmed independently. Not assumed from how the files were named.

Used paired-read agreement as the filter

Agreement between the two reads of a pair is a stricter test than any quality-score threshold, and it costs nothing.

Collapsed error clouds before grouping

A dominant molecule surrounded by its own sequencing variants otherwise reads as a family of ten. Collapse first, and every chain in a family is a different molecule.

Separated families from motif groups

One is the unit you order synthesis against; the other is the unit a biological argument rests on. Conflating them inflates both.

Measured against a matched background

A fold-enrichment is only as good as its denominator. We use the complete singleton fraction of the earliest rounds, and state the choice so the figures can be recomputed rather than believed.

Tested significance, then corrected it

A handful of chains sharing a rare element gives a spectacular ratio and almost no evidence. Every reported family survives a test corrected for the whole search.

Resolved structure instead of eyeballing it

Base-pairing partners are resolved computationally. Reading a dot-bracket string by eye is how confident structural claims turn out to be wrong.

Looked for contamination on purpose

Index bleed-through between multiplexed samples was found, measured and identified as an artefact — before it could be read as a biological result.

2
independent pipelines, not one. We run a second, separately written analysis over the same data and compare. On this project the comparison caught a real flaw in one of them: a grouping step was letting short, general motifs absorb the chains of longer, more specific ones — so a single clear result was being reported as three vague ones. A single pipeline always agrees with itself. Two do not have to.
The deliverable

Built to be checked, not trusted.

A report a client cannot audit is one they have to take on faith — the wrong basis for spending a synthesis budget.

  • Every threshold it applies is stated in the report
  • Checksums for each input file, so any figure traces back
  • Rebuilds from raw sequencing with a single command
  • A regression check confirms a rebuild reproduces the delivered figures exactly
  • Where the analysis is limited, the report gives the number rather than the impression
What shipped
Combined report — three targets, 32 pages
Per-target annexes — family sheets with folding
Appendix — chain by chain, element marked
Machine-readable table — every member of every family
What it does not claim

Sequencing gives you an order. Not an answer.

Nothing in a deliverable like this has been shown to bind. Selection sequencing produces a priority order for synthesis and testing — not affinities, not specificity, and not a validated binder. We say so in every report, because the alternative is a client spending months discovering it.

Supported by

Who backs this work.

Partner incubator
Boundless Accelerator
Commercialization mentorship, evidence reviews and introductions.
Academic host
Humber Polytechnic
Office of Research & Innovation, applied-research support.
Programme
Labs4 TRL
Technology Readiness Level-Up, Ontario Hub.
Submitted
Mitacs
Business Strategy Internship, Winter 2027 cohort.

Send us the pool you're stuck on.

Start a free first look →How Rescue works
AptaPilot — case study, client de-identified← All services