Developability
Developability first: the constraints we set before generating a single aptamer
Most oligonucleotide binders fail for reasons that have nothing to do with affinity. Those reasons are knowable in advance, which means they belong in the design specification rather than in the post-mortem.
A nanomolar aptamer that is degraded within minutes in serum, cleared through the kidney in an hour, triggers an innate immune response, or costs an unreasonable amount per gram to manufacture is not a lead. It is an interesting result. The distance between those two things is developability, and in nucleic acid therapeutics it is unusually tractable: the liabilities are chemical and structural, they are well characterized, and they can be written down as constraints before any sequence is proposed.
This is the part of our workflow that runs before generation. A target definition arrives, and the first artifact we produce is not a candidate set but a constraint specification: what chemistry the molecule will be built from, what it must fold into, what it must not contain, and how long it is allowed to be. Everything downstream is conditioned on that document.
Chemistry comes before sequence
Unmodified RNA has a serum half-life measured in seconds to minutes, dominated by ribonuclease attack at the 2′-hydroxyl. The standard countermeasures are well established. Substituting pyrimidines with 2′-fluoro or 2′-O-methyl groups removes the nucleophile and sharply increases nuclease resistance, while also raising duplex thermal stability. Locked nucleic acid residues push stability further and are useful at defined positions, though over-substitution rigidifies the fold and frequently destroys binding. Phosphorothioate linkages slow exonuclease digestion but introduce stereochemical heterogeneity and a well-documented tendency toward nonspecific protein binding, so we treat them as a targeted tool rather than a backbone-wide default. A 3′ inverted deoxythymidine cap is cheap, chemically simple and blocks 3′ exonuclease processing.
The important consequence is that chemistry is not a finishing step. A 2′-fluoro pyrimidine pool folds differently from an all-natural RNA pool of the same sequence, because the substitution changes sugar pucker preference and base-pair stability. Selecting in one chemistry and manufacturing in another is a known route to losing activity. Equally, folding predictions trained predominantly on natural RNA thermodynamics will systematically mis-rank modified sequences unless the parameter set is adjusted, which is a limitation we carry explicitly in our scoring rather than one we hope averages out.
Clearance sets the size of the problem
An unconjugated 30-mer is far below the glomerular filtration threshold and is cleared renally within an hour or two. For most systemic indications that is disqualifying on its own. The established remedy is conjugation to increase hydrodynamic radius, most familiarly a branched poly(ethylene glycol) group. Cholesterol and albumin-binding conjugates offer alternative pharmacokinetic profiles.
Conjugation is not free, and its cost is paid in design. The attachment point has to be somewhere that does not contact the target, which means the binding interface must be known or at least modeled before the handle is placed. A conjugate large enough to prevent filtration can also sterically occlude a shallow epitope. We therefore treat the conjugation site as a design variable with a protected-region constraint, not as a decision for process chemistry to make later.
Motifs the immune system recognizes
Oligonucleotides are exactly the class of molecule that innate immunity evolved to detect. Unmethylated CpG dinucleotides in a DNA backbone are agonists for Toll-like receptor 9; single-stranded uridine-rich RNA is recognized by Toll-like receptors 7 and 8; double-stranded regions of sufficient length engage Toll-like receptor 3 and cytosolic sensors. For a therapeutic intended to be dosed repeatedly, these are liabilities rather than features.
They are also sequence motifs, which makes them cheap to exclude at generation time. Our constraint sets prohibit unmethylated CpG in DNA designs, cap contiguous uridine runs in RNA designs, and bound the length of perfectly paired duplex regions. The same applies to G-quadruplex-forming runs, which we exclude by default: four or more guanine tracts separated by short loops give a fold that is thermally stable, structurally polymorphic, difficult to characterize batch-to-batch, and frequently promiscuous in its protein binding. There are aptamers whose function genuinely depends on a quadruplex, and in those programs we relax the constraint deliberately and document why.
Folding robustness, not folding score
A minimum free energy structure is a single point estimate of something that is really a distribution. Two sequences can share an identical predicted fold while differing completely in how much of their ensemble actually occupies it. We score the ensemble defect, which measures expected deviation from the target structure across the Boltzmann ensemble, in preference to minimum free energy alone, and we require the intended fold to remain dominant across a range of magnesium concentrations rather than at a single modeled condition.
This last point is responsible for a great deal of irreproducibility in the field. Aptamer folding is strongly dependent on divalent cations, and a candidate whose structure is only dominant at 5 millimolar magnesium may be largely unfolded in the 1 to 2 millimolar free magnesium range of plasma. We also require a defined refolding protocol, because a molecule that reaches its active conformation only after heating and slow cooling is a molecule whose assay results depend on handling.
Length is a cost function
Solid-phase synthesis yield falls with every coupling step, so cost and purity degrade with length in a way that is close to exponential at manufacturing scale. A 25-mer and a 60-mer are not variations on a theme; they are different cost-of-goods conversations. Shorter is also easier to characterize analytically. We put an explicit length penalty into candidate scoring and, where a program allows it, we run truncation series on confirmed binders early rather than at the end, because the minimal binding element is almost always shorter than the sequence selection delivered.
The specification, in one table
| Constraint | Why it matters | How we encode it |
|---|---|---|
| Nuclease resistance | Unmodified RNA survives minutes in serum | Fixed modification pattern before generation; folding scored in that chemistry |
| Renal clearance | Small oligos filter within hours | Conjugation handle placed outside a protected interface region |
| Innate immune recognition | CpG, U-rich and long duplex motifs are receptor agonists | Hard motif exclusions and run-length caps at sampling time |
| Structural polymorphism | Quadruplexes and alternative folds confound characterization | G-tract exclusion by default; ensemble dominance requirement |
| Fold robustness | Activity that depends on buffer is not activity | Ensemble defect across a magnesium range, not minimum free energy |
| Manufacturability | Yield and purity fall with length | Length penalty in scoring; early truncation series |
| Specificity | Matrix and off-target binding masquerade as hits | Counter-screen panel defined before the first synthesis batch |
Writing this specification is not glamorous work, and it is the single highest-leverage step in the loop. Every constraint above removes a region of sequence space that generation would otherwise spend capacity exploring, and removes a failure that wet-lab time would otherwise be spent discovering. Scientific agents assemble the draft specification from the target context and the relevant literature, and a human scientist signs it off before anything is generated, because the places where constraints should be relaxed are exactly the places where judgement is required.
Further reading
- Keefe, Pai and Ellington, Nature Reviews Drug Discovery 2010, on aptamer therapeutics and the chemical modifications used to stabilize them.
- Ng and colleagues, Nature Reviews Drug Discovery 2006, on pegaptanib, including the role of PEGylation in extending residence time.
- Zhou and Rossi, Nature Reviews Drug Discovery 2017, for a current view of aptamer developability and clinical translation.
- Vaught and colleagues, Journal of the American Chemical Society 2010, on selection with 5-position modified pyrimidines and the effect of expanded chemistry on binding.
- Lorenz and colleagues, Algorithms for Molecular Biology 2011, ViennaRNA 2.0, for partition function and ensemble calculations.
- Work on innate immune recognition of nucleic acids by Toll-like receptors 3, 7, 8 and 9, which underlies the motif exclusions described above.