Therapeutic Oligonucleotide Design Platform — Technical Whitepaper
Modality by modality: molecular form, chemistry, delivery, deliverables, and industry benchmarking
Bioneer BioFoundryCenter · Published 2026 · Document version 2.0 · Contact geneorder@bioneer.co.kr
0. What This Document Is
This is a public technical whitepaper. It is not a product manual. It describes, modality by modality, the biological mechanism by which a therapeutic oligonucleotide acts, the demands that mechanism places on design, chemistry and delivery, and the computation and data that meet those demands. The unit of description is the molecule, not the program. What a customer ultimately holds is not software but a sequence, a chemistry specification and a delivery specification — and understanding what determines those three is the substance of a development decision.
Every modality is therefore described under the same eight headings: (1) the form and definition of the molecule, (2) the molecular biology of its action, (3) the primary literature and structural evidence that established that mechanism, (4) clinical and regulatory precedent and the surrounding patent landscape, (5) whether chemical modification is required and why, (6) whether a delivery vehicle is required and why, (7) what the design deliverables consist of, and (8) what this modality cannot solve in principle. Never omitting the eighth heading is a principle of this document. A technical document that does not state a modality's boundaries leaves the reader to discover them as a failed experiment, at a cost many times that of the design itself.
Figures cited here fall into three classes. Values reported in the public literature are given with author, journal and year. Values measured on in-house benchmarks are given with the dataset they were measured on. Values whose basis is not established are not given at all — the fact that they are undetermined is stated instead. Refusing to fill that third class with estimates is the only thing that makes the first two credible.
1. What Kinds of Oligonucleotide Can Be Designed
Therapeutic oligonucleotides are not one technology but eleven distinct operating principles sharing a common chemical foundation. Two nucleic acids of similar length can behave entirely differently: one cleaves its target RNA, another occupies a site without cleaving it, another raises rather than lowers gene expression, another stimulates an immune receptor, and another folds into a ligand that binds a protein. This chapter lays out that whole landscape first; each subsequent chapter expands one row of the table.
1.1 Eleven roles by operating principle
|
# |
Role |
Common name |
What the molecule does to its target |
Type of target |
|---|---|---|---|---|
|
I |
Catalytic silencing (RISC cleavage) |
siRNA · RNAi (RNA interference) |
Enzymatically cleaves the target RNA, then is reused on the next target |
Mature cytoplasmic mRNA |
|
II |
Catalytic silencing (RNase H1 cleavage) |
ASO gapmer (antisense oligonucleotide) |
Forms a DNA:RNA heteroduplex that recruits a host nuclease to cleave the RNA strand |
Nuclear pre-mRNA and cytoplasmic mRNA |
|
III |
Splice switching (non-cleaving steric block) |
SSO (splice-switching oligonucleotide) |
Physically masks a splicing regulatory element, changing isoform choice |
Intron/exon boundary elements of pre-mRNA |
|
IV |
Noncoding RNA silencing |
ncRNA-directed siRNA / ASO |
Removes an RNA that makes no protein, by mechanism I or II |
lncRNA · circRNA · snoRNA · translation-enhancing elements |
|
V |
Transcriptional activation |
saRNA · RNAa (small activating RNA) |
Binds the promoter region and increases transcription of that gene |
The genomic region around the transcription start site and the RNA made there |
|
VI |
miRNA axis modulation |
miRNA mimic · anti-miR (antagomir) |
Replaces the function of an endogenous miRNA (mimic) or sequesters it (antagonist) |
A mature miRNA and its target mRNA set |
|
VII |
Site-directed base editing |
ADAR-recruiting editing oligo (AIMer · EON · arRNA) |
Recruits an endogenous editing enzyme to convert a specific adenosine to inosine |
A single adenosine in a mature mRNA |
|
VIII |
Premature stop codon readthrough |
suppressor tRNA (ACE-tRNA · sup-tRNA) |
Reads a stop codon as an amino acid, restoring full-length protein synthesis |
The premature termination codon of a nonsense-mutant mRNA |
|
IX |
Innate immune receptor modulation |
CpG-ODN · immunomodulatory oligo (TLR agonist/antagonist) |
Stimulates or suppresses an endosomal Toll-like receptor |
TLR7 · TLR8 · TLR9 proteins |
|
X |
Targeted delivery (conjugate) |
aptamer conjugate (ApDC · AOC) |
A folded nucleic acid binds a surface receptor and carries a payload into that cell |
A cell-surface protein |
|
XI |
Structure-based ligand candidate generation |
de novo aptamer scaffold |
Generates a candidate pool by folding fitness; affinity is determined experimentally |
Any protein target (decided at the experimental stage) |
Roles I and II both cleave the target, but the catalyst differs. In I the Argonaute protein is itself the slicer; in II the oligonucleotide builds a DNA:RNA heteroduplex that the host's RNase H1 recognises and cuts. That difference propagates straight into design: I must be double-stranded (or, if single-stranded, must satisfy the RISC loading requirements), while II must carry a central DNA window. Role III must not be cleaved at all, so the presence of that same DNA window makes the design fail. Within one class of single-stranded molecule, 'a central DNA gap is mandatory' and 'a central DNA gap is forbidden' stand opposed — and that single fact is the entire reason II and III are different modalities.
Role V runs in the opposite direction. Where I through IV reduce the target, V increases it. When the therapeutic hypothesis is not 'there is too much of this protein' but 'there is too little' — a lost tumour suppressor, a haploinsufficient gene, an insufficient compensatory pathway — silencing mechanisms cannot be the answer in principle. Roles VII and VIII change neither quantity nor occupancy but sequence: they apply where the problem is the quality of the transcript rather than its amount, that is, a point mutation or a premature stop.
Roles IX through XI do not target RNA at all. IX acts on a protein receptor as a ligand, X uses a folded nucleic acid as a targeting device in place of an antibody, and XI generates candidates for that device. They belong to the same technology family because the material and the manufacturing route are identical — the same solid-phase phosphoramidite synthesis and the same repertoire of chemical modifications.
1.2 The form of the molecule — single-stranded or double-stranded
Form is not a technical detail. A double-stranded molecule requires two strands to be synthesised and annealed, doubling synthesis cost and quality-control items; it requires thermodynamic control over which strand is loaded as the active guide; and the strand that is not loaded becomes an independent off-target burden. A single-stranded molecule has none of those problems, but it has no partner strand to protect it, so its chemical modification burden is higher and its target affinity has to be secured chemically.
|
Role |
Form |
Typical length |
Strand composition |
What the form imposes |
|---|---|---|---|---|
|
I siRNA / RNAi — catalytic silencing (RISC) |
double-stranded |
21 bp core + 2 nt overhang ×2 = 23 nt per strand |
guide (antisense) + passenger (sense) |
Strand-selection asymmetry must be secured; passenger off-targets must be managed |
|
I variant: ss-siRNA (single-strand RISC) |
single-stranded |
~20 nt |
guide only |
A 5'-phosphate mimic is an absolute requirement; duplex chemistry does not transfer |
|
II ASO gapmer — catalytic silencing (RNase H1) |
single-stranded |
16–20 nt |
wing–DNA gap–wing (typically 5-10-5) |
Central DNA gap mandatory; wings carry the affinity |
|
III SSO — splice switching |
single-stranded |
18–25 nt |
uniform modification, no gap |
DNA residues forbidden; uniform high affinity mandatory |
|
IV ncRNA-directed silencing |
double- or single-stranded |
as I or II |
follows the chosen mechanism |
Target class dictates which mechanism applies |
|
V saRNA / RNAa — transcriptional activation |
double-stranded |
21 bp core + overhangs |
guide + passenger |
No cleavage requirement; nuclear localisation and chromatin access are what matter |
|
VI-a miRNA mimic |
double-stranded |
mature miRNA length (typically 21–23 nt) |
guide = the mature miRNA sequence |
Sequence is fixed in advance; design freedom exists only in chemistry and passenger |
|
VI-b anti-miR / antagomir |
single-stranded |
8 nt (seed-directed) to full length |
single strand |
Must not recruit RNase H, so modification must be uniform |
|
VII ADAR-recruiting editing (AIMer / EON) |
single-stranded |
~20–40 nt (short) / 50–100 nt (long) |
single strand |
A C mismatch opposite the target A is fixed; the central window must stay lightly modified |
|
VIII suppressor tRNA (ACE-tRNA) |
folded tRNA (~72–76 nt) |
mature tRNA length |
single chain, cloverleaf fold |
Aminoacylation identity elements invariant; tertiary structure must be preserved |
|
IX CpG-ODN / TLR ligand |
single-stranded DNA or RNA |
18–30 nt (class dependent) |
single strand or palindromic dimer |
The phosphorothioate backbone is part of the function; motif placement is decisive |
|
X aptamer conjugate (ApDC / AOC) |
folded single strand + payload |
aptamer 25–90 nt + linker + payload |
aptamer–linker–payload |
The fold is the prerequisite for binding; payload attachment must not break it |
|
XI aptamer scaffold pool |
folded single strand |
20–100 nt (per constraints) |
single strand |
Structural diversity is the goal, not a single optimum |
The '21 bp core + 2 nt overhang' entry is this platform's double-stranded standard. The overhang bases are not an arbitrary thymidine dimer but the real bases from the target sequence context, and the antisense overhang is the reverse complement of the two upstream nucleotides. The rationale is given in 2.6.
1.3 Which modalities require chemical modification and which do not
Chemical modification is not a uniform procedure applied to every oligonucleotide. What it is for differs by modality, and in some cases modification destroys function. Its purposes are four: (a) nuclease resistance, (b) target binding affinity, (c) suppression of innate immune stimulation, and (d) mediating cellular uptake itself. The fourth applies only where there is no delivery vehicle; where there is one, it becomes a liability instead.
|
Role |
Chemical modification |
For what |
Where modification is forbidden or limited |
|---|---|---|---|
|
I siRNA / RNAi (RISC) |
required |
Plasma stability, immune suppression, seed off-target mitigation, hepatic conjugation |
Around the position opposing the AGO2 cleavage site, only modifications that do not block cleavage |
|
I variant: ss-siRNA |
required, under different rules |
A 5'-phosphate mimic is the precondition for activity; systemic stability |
Excessive 2'-O-alkyl or bicyclic substitution lowers activity; duplex templates cannot be reused |
|
II ASO gapmer (RNase H1) |
required, and structurally enforced |
Wing affinity plus preservation of catalytic recognition in the gap |
The central gap must remain 2'-deoxy — a 2'-modification there abolishes cleavage |
|
III SSO (splice switching) |
required, uniform across every position |
Maximum affinity, nuclease resistance, and avoidance of RNase H |
No DNA residues — even one risks unintended cleavage |
|
IV ncRNA-directed silencing |
required |
as I or II |
Inherits the constraints of the chosen mechanism |
|
V saRNA / RNAa |
required, minus the cleavage rule |
Stability and immune suppression. No cleavage is needed, so catalytic-site protection does not apply |
Bulky additions that impede nuclear localisation warrant caution |
|
VI-a miRNA mimic |
required |
Preserving the natural function of the mature miRNA while gaining stability |
Over-modification of the seed region (positions 2–8) distorts target recognition |
|
VI-b anti-miR / antagomir |
required, uniform and fully backbone-modified |
Maximum sequestration affinity while avoiding RNase H |
Uniformity is mandatory, so partial-modification designs do not exist |
|
VII ADAR-recruiting editing (AIMer) |
required, but the centre must stay light |
Stabilisation of the termini and flanks |
The central window around the orphan cytidine must retain unmodified 2'-OH and phosphodiester linkages — modification there directly lowers editing efficiency |
|
VIII suppressor tRNA (ACE-tRNA) |
not required in principle |
Stabilisation is discussed only where a synthetic RNA is dosed directly; the standard route is an expression cassette |
Artificial modifications that perturb native modification sites or tertiary structure are avoided |
|
IX CpG-ODN / TLR ligand |
required — the modification is the function |
The phosphorothioate backbone provides both stability and uptake; 2'-O-methyl is an active component of antagonist design |
In agonist design, heavy 2'-modification around the CpG motif abolishes recognition |
|
X aptamer conjugate (ApDC / AOC) |
required |
Serum nuclease resistance (2'-fluoro, 2'-O-methyl), delayed renal clearance (PEGylation) |
Residues involved in binding must retain their original chemistry |
|
XI aptamer scaffold pool |
optional |
Candidates are evaluated unmodified for folding; stabilisation is applied after selection |
Fixing chemistry during exploration shrinks the search space |
The two rows that deserve most attention are VII and VIII. VII is the only modality in which modification is mandatory and yet one specific window must remain unmodified — stability and activity collide directly on the same axis. VIII is the only modality in which modification is not required at all, because the molecule is not a synthetic ligand but a functional RNA that must fold inside the cell, be charged with an amino acid, and enter the ribosome. An artificial modification that perturbs any one of those three processes disables it.
1.4 Which modalities require a delivery vehicle and which do not
Delivery splits into two questions: does the molecule get into the cell, and does it reach the right tissue. Nucleic acids are large and strongly anionic, so they do not cross membranes by free diffusion. Some uptake mechanism is therefore mandatory, and there are only three options: encapsulate in a lipid particle, attach a receptor-binding ligand, or make the backbone chemistry itself bind proteins and drive endocytosis. The third is carrier-free (gymnotic) uptake, and the protein-binding character of the phosphorothioate backbone is what makes it work.
|
Role |
Delivery vehicle |
Route in practice |
Why a vehicle is unnecessary or unsuitable |
|---|---|---|---|
|
I siRNA / RNAi (RISC) |
required |
GalNAc conjugation for hepatic targets; lipid nanoparticles or alternative vehicles elsewhere |
A duplex cannot carry a fully phosphorothioate backbone, so carrier-free uptake barely applies |
|
I variant: ss-siRNA |
optional |
Lipid conjugation or carrier-free dosing are both reported |
A single strand can carry a high PS fraction, but potency is traded away in doing so |
|
II ASO gapmer (RNase H1) |
optional — often unnecessary |
Systemically via GalNAc conjugation or carrier-free; intrathecally for the central nervous system |
A phosphorothioate single strand is the canonical case where carrier-free uptake works |
|
III SSO (splice switching) |
usually unnecessary |
Intrathecal for CNS; systemic dosing with tissue distribution for muscle |
Several approved products are dosed as naked oligonucleotides |
|
IV ncRNA-directed silencing |
follows the mechanism |
as I or II |
Subcellular localisation of the target drives the choice |
|
V saRNA / RNAa |
required |
Lipid nanoparticles or conjugates |
The molecule must reach the nucleus, so cytoplasmic delivery alone is insufficient |
|
VI-a miRNA mimic |
required |
Lipid nanoparticles or conjugates |
Duplex, so the same constraint as I |
|
VI-b anti-miR / antagomir |
optional |
A fully phosphorothioate single strand supports carrier-free uptake; GalNAc conjugation for liver |
Same logic as II |
|
VII ADAR-recruiting editing (AIMer) |
required |
Lipid nanoparticles or conjugates — the unmodified central window lowers stability, so protection matters |
Chemical protection cannot be placed at the centre, so the carrier's role is comparatively larger |
|
VIII suppressor tRNA (ACE-tRNA) |
required |
An expression cassette in a lipid particle or viral vector, or synthetic tRNA in a particle |
A folded functional RNA does not undergo carrier-free uptake |
|
IX CpG-ODN / TLR ligand |
unnecessary |
Local or subcutaneous administration; DNA CpG is co-formulated with adjuvants such as aluminium salts |
The target is an endosomal receptor, so endocytosis is itself arrival, and the PS backbone drives it |
|
X aptamer conjugate (ApDC / AOC) |
unnecessary — the molecule is the vehicle |
Intravenous or local administration followed by receptor-mediated internalisation |
The aptamer is the targeting device; adding a carrier dilutes that targeting |
|
XI aptamer scaffold pool |
not applicable |
Laboratory-stage output |
Not dosed in vivo |
1.5 The chemistry × delivery map
Overlaying the two preceding tables yields four quadrants. This map is used directly to allocate effort early in a programme: in the quadrants that need a vehicle, no amount of sequence quality relieves the delivery bottleneck, and in the quadrants where chemistry is the function, a sequence without a fixed chemistry specification means nothing.
|
Delivery vehicle required |
Vehicle unnecessary or optional |
|
|---|---|---|
|
Chemical modification required |
I siRNA/RNAi · V saRNA/RNAa · VI-a miRNA mimic · VII ADAR editing (AIMer) |
II ASO gapmer · III SSO · VI-b anti-miR · IX CpG-ODN/TLR ligand · X aptamer conjugate (ApDC/AOC) |
|
Chemical modification unnecessary or optional |
VIII suppressor tRNA (ACE-tRNA) |
XI aptamer scaffold pool (pre-in vivo stage) |
The upper-left quadrant is where both problems must be solved, and is therefore the hardest — except for hepatic targets, where GalNAc conjugation is a mature answer that collapses the practical difficulty. The upper-right quadrant is where chemistry is delivery: a single parameter, phosphorothioate fraction, moves uptake, stability and toxicity simultaneously, so locating its productive window is the core of design. The lower-left is where the molecule is a functional RNA rather than a synthetic oligonucleotide, and the lower-right is a stage before in vivo administration.
1.7 Terminology map — the many names for the same molecule
This field calls the same molecule by different names in the literature, in regulatory documents and in industry. The table below maps the role numbers used in this document to their common names and to the synonyms and adjacent terms most often encountered, so that a reader knows which term to search on in literature or patent work.
|
Role |
This document |
Most widely used name |
Synonyms and adjacent terms |
|---|---|---|---|
|
I |
Double-stranded catalytic silencing |
siRNA (small interfering RNA), RNAi |
double-stranded siRNA · duplex siRNA · GalNAc-siRNA · ESC / ESC+ siRNA · RNA interference · RISC-mediated silencing · shRNA (expressed analogue) |
|
I variant |
Single-stranded catalytic silencing |
ss-siRNA (single-stranded siRNA) |
ss-RNAi · single-stranded RNAi · 5'-VP guide |
|
II |
Single-stranded catalytic silencing |
ASO gapmer (gapmer antisense oligonucleotide) |
antisense oligonucleotide · AON · RNase H gapmer · 2'-MOE gapmer · cEt gapmer · LNA gapmer · GalNAc-ASO · third-generation antisense |
|
III |
Splice switching |
SSO (splice-switching oligonucleotide) |
SSA · splice-switching antisense · exon-skipping oligonucleotide · PMO (morpholino) · PPMO (peptide-conjugated PMO) · steric-block antisense · TOSS / TOES |
|
IV |
Noncoding RNA silencing |
ncRNA-directed siRNA / ASO |
lncRNA-targeting antisense · circRNA-targeting siRNA · anti-lncRNA · snoRNA-directed oligo · SINEUP |
|
V |
Transcriptional activation |
saRNA (small activating RNA), RNAa |
RNA activation · promoter-targeted RNA · pshRNA · NAT-targeting antisense (AntagoNAT) |
|
VI-a |
miRNA mimicry |
miRNA mimic |
miRNA replacement therapy · synthetic miRNA duplex |
|
VI-b |
miRNA antagonism |
anti-miR / antagomir |
antimiR · miRNA inhibitor · anti-miRNA oligonucleotide (AMO) · tiny LNA · miRNA sponge (expressed analogue) · mixmer |
|
VII |
Site-directed base editing |
ADAR-recruiting editing oligonucleotide |
AIMer · EON (editing oligonucleotide) · Axiomer · arRNA · RESTORE · LEAPER · ADAR-recruiting guide · A-to-I RNA editing |
|
VIII |
Premature stop codon readthrough |
suppressor tRNA |
ACE-tRNA · sup-tRNA · anticodon-engineered tRNA · nonsense suppressor tRNA · PTC readthrough (distinct from small-molecule inducers) |
|
IX |
Innate immune receptor modulation |
CpG-ODN / TLR agonist or antagonist oligonucleotide |
CpG oligodeoxynucleotide · immunostimulatory sequence (ISS) · IMO · TLR9 agonist · TLR7/8 agonist · inhibitory oligonucleotide (INH-ODN, suppressive ODN) |
|
X |
Targeted delivery conjugate |
ApDC / AOC (aptamer conjugate) |
aptamer–drug conjugate · aptamer–oligonucleotide conjugate · aptamer–siRNA chimera · aptamer-guided delivery |
|
XI |
Structure-based ligand candidate generation |
de novo aptamer scaffold |
aptamer library design · SELEX starting pool · structure-based candidate pool |
Adjacent technologies that fall outside this document's scope are distinguished as well. shRNA and miRNA sponges are expressed from vectors, so their manufacturing and regulatory paths differ from synthetic oligonucleotides; the CRISPR family (Cas9, Cas13, base editors, prime editors) carries a protein component and is a separate technology class. mRNA, self-amplifying RNA and circular RNA vaccines are different molecules but share the lipid nanoparticle delivery layer, and are therefore covered in Chapters 7 and 23.
|
Category |
Terms |
Relationship to this document |
|---|---|---|
|
Expressed analogues |
shRNA · artificial miRNA · miRNA sponge · expressed antisense |
Same mechanism but vector-expressed, so manufacturing and regulatory paths differ (out of scope) |
|
Protein-accompanied technologies |
CRISPR-Cas9 · Cas13 · base editors · prime editors |
The guide is a nucleic acid, but a protein or its mRNA must be co-delivered — a separate technology class (out of scope) |
|
Antigen-expressing platforms |
mRNA vaccines · self-amplifying RNA (saRNA*) · circular RNA vaccines |
Different molecules but a shared delivery layer — covered in Chapter 7 (delivery) and Chapter 23 (vaccines) |
|
Chemical elements |
2'-OMe · 2'-F · 2'-MOE · LNA · cEt · GNA · PS · PMO · 5'-VP · GalNAc |
Not modalities but the chemistry repertoire — covered in Chapter 6 |
|
Delivery elements |
LNP · SORT · exosome · polymer · VLP · GalNAc conjugation · gymnosis |
Not modalities but the delivery repertoire — covered in Chapter 7 |
A collision of abbreviations to watch: saRNA denotes two different things in the literature. In this document saRNA means the small activating RNA of Modality V; the self-amplifying RNA used in vaccine contexts is a different molecule and is written out in full as 'self-amplifying RNA', only in section 23.3.2.
1.6 Categories of deliverable
The output of design activity is not only a sequence. What actually feeds a development decision is the sequence, the reason it was chosen, and the quantitative evidence behind that reason. Deliverables fall into four categories.
|
Category |
Contents |
What it is used for |
|---|---|---|
|
Molecular specification |
Per-strand sequences, length and overhang convention, per-position chemistry map (sugar string and backbone string), terminal handling, conjugate attachment point and linker |
Synthesis orders, quality specifications, definition of the substance |
|
Selection evidence |
Efficacy predictions and their decomposition, off-target tables (per gene and per site), variant avoidance results, accessibility profiles, cross-species tables, cleavage coverage |
Candidate selection meetings, experimental prioritisation, review responses |
|
Development-readiness material |
Preclinical species suitability verdicts, synthesis impurity projections, formulation and stability risk items, route-dependent absorption summaries, patent-landscape advisories |
IND preparation, CDMO handover, start of formulation development |
|
Reproducibility material |
The run ledger — tool versions, database releases, parameters, random seed, UTC timestamps |
Journal Methods sections, regulatory filings, later audits |
The fourth category is routinely treated as a by-product but in fact determines the value of the other three. If it is impossible to reconstruct after the fact why a candidate was chosen, then the experimental results obtained with that candidate cannot be interpreted reproducibly either. A determinism contract guaranteeing byte-identical output for the same input and seed, together with a ledger written automatically for every run, is what this category consists of.
2. Shared Technical Foundation
The eleven modalities differ in operating principle but are designed on one shared foundation. How a target gene is resolved, what counts as an off-target, which coordinate system positions are expressed in, and how a result is made reproducible — if these four differ per modality, comparison between modalities does not hold. When two modalities are designed in parallel against the same gene, '3 off-targets' in one report must mean the same three things it means in the other. That is why this chapter exists.
2.1 The determinism contract
The design pipeline produces byte-identical output for the same input and the same random seed. This sounds unremarkable, yet a substantial fraction of bioinformatics pipelines do not have the property. The causes are usually three: results accumulate into a data structure in an order that depends on parallel execution; iteration depends on hash ordering; or an external tool breaks ties non-deterministically across threads.
Each of the three is blocked structurally. Parallel results are merged only in order-independent ways, ties carry explicit secondary sort keys, and caching is restricted to computations that are pure functions. That last restriction matters most in practice. Cross-species ortholog search is a pure function of (query sequence, species database), so caching it cannot change the value; off-target screening, by contrast, has two aligners writing into one structure — one by assignment and one by accumulation — so a deterministic cache would double-count. A performance optimisation that changes the result is not an optimisation but a defect, so the latter is deliberately not cached.
Determinism buys more than reproducibility. In a deterministic system, any difference between two runs must originate in a difference of input, so the effect of changing one parameter can be attributed exactly. In a non-deterministic system that attribution is impossible, and as a consequence the question 'does this parameter matter' cannot be answered at all.
2.2 Provenance — the run ledger
Every run leaves a self-contained ledger. It records start and end times as UTC ISO-8601 timestamps carrying explicit timezone information, the invoking principal, a fingerprint identifying the state of the code, the version of every external tool used, the release of every reference database consulted, the scoring weights applied, and the random seed.
Fixing timestamps to UTC removes ordering ambiguity caused by local timezone differences in multi-region collaboration, which is a basic requirement of the audit trails that electronic records regulations ask for. Recording database releases is even more practical: when two designs of the same gene give different results, it is the only basis on which to distinguish an algorithm change from a reference data update.
The ledger is also the record of pre-execution verification. Before any computation begins, the pipeline confirms that the expected reference data releases are actually present, and on failure it stops rather than silently computing against something else. Silent failure is far more expensive than explicit failure, because a wrong result then propagates downstream indistinguishable from a correct one.
2.3 Reference data
Design quality depends on the breadth and freshness of reference data as much as on algorithms. The principal resources and their roles are below.
|
Resource |
Release |
Role |
What breaks without it |
|---|---|---|---|
|
Population variation database |
v4.1 (constraint metrics v4.1.1) |
Target-site polymorphism avoidance, common-SNP checks across the seed, gene-level loss constraint |
Candidates that work only in part of a population rise to the top |
|
Transcript annotation |
GENCODE v49 |
Transcript coordinates, biotypes, canonical isoform designation, cleavage-coverage arithmetic |
It becomes impossible to know which isoform set 'knocking down a gene' actually knocks down |
|
Clinical variant database |
rolling |
Pathogenic variant avoidance, decomposition of the nonsense variant spectrum, variant profiles |
Designs land on disease-causing variants, and readthrough or editing suitability cannot be judged |
|
Circular RNA atlas |
circAtlas 3.0 |
Back-splice junction definitions |
Circle-specific targets cannot be separated from the linear parent gene |
|
Noncoding RNA family database |
Rfam 15 |
Noncoding RNA family classification |
Misclassifying the target class leads to the wrong mechanism |
|
Integrated noncoding RNA resources |
RNAcentral · LncATLAS · REDIportal |
Identifier integration, subcellular localisation, A-to-I editing sites |
Cytoplasmic mechanisms get applied to nuclear-retained targets |
|
Chromatin and protein–RNA resources |
ENCODE (H3K4me3 · ATAC · H3K27ac · eCLIP) |
Promoter accessibility, RNA-binding protein occupancy masks |
Activation preconditions cannot be checked, and protein-occupied sites get targeted |
|
RNA modification atlas |
m6A-Atlas |
Modification-site interference masks |
Sites where modification blocks binding are not avoided |
|
miRNA resources |
miRBase mature set · TargetScan 8.0 · ENCORI |
Mature miRNA anchors, context-based target prediction, AGO-CLIP and degradome cross-checks |
Prediction alone is dominated by false positives |
|
Tissue expression resource |
GTEx v11 |
Expression breadth, tissue-specificity index, target tissue calls |
Delivery target tissue and systemic exposure risk are assumed without basis |
|
Cell dependency resource |
DepMap (1,208 cell lines) |
Essentiality calls, residual fitness dependency |
A well-designed molecule silences an essential gene |
|
Tractability and liability resource |
Open Targets |
Tractability buckets, documented adverse-event precedent |
Targets with existing failure precedent are chosen again |
|
tRNA resource |
gtRNAdb (hg38 · mm39) |
Scaffold body sequences |
Risk of using invented sequence |
|
RNA modification resource |
MODOMICS |
tRNA modification-impact assessment |
Perturbation of modification sites goes undetected |
|
Protein interaction graph |
13 integrated sources |
1.73M consensus pairs, 6.70M pairs across 13 evidence channels, 22,824 complexes |
No basis for finding a detour target when the direct target is not viable |
|
Knowledge graph |
integrated build |
4,228 pathways · 48,329 GO terms / 858,529 annotations · 28.09M gene–disease · 140,134 paralog pairs |
No basis for family selectivity or disease-association verdicts |
2.4 Computational tool layer and licensing posture
The platform combines openly licensed software with components developed in house. Public tools are named directly so a reader can verify the method independently; components developed in house are identified as proprietary in-house algorithms and carry GC-prefixed names. This is a naming policy, not concealment — what each component computes is stated below.
|
Component |
Licence class |
What it computes |
|---|---|---|
|
BLAST+ (blastn · blastn-short · tblastx) |
public |
Off-target search, paralog and ortholog matching, cross-species transcript comparison |
|
Bowtie |
public |
Transcriptome-wide mismatch-tolerant alignment (first-pass off-target scan) |
|
Infernal (cmsearch) |
public |
Covariance-model validation of structured RNA families |
|
RIblast · RIsearch |
public |
RNA–RNA interaction free energy, target-site accessibility |
|
CPC2 |
public |
Coding-potential call (the gate on a 'noncoding' premise) |
|
miranda |
public |
Strict-seed miRNA target search |
|
patman |
public |
High-throughput short-sequence matching |
|
GC-ThermoFold |
proprietary in-house algorithm |
Secondary-structure minimum free energy and partition function, duplex binding free energy, local accessibility profiles, G-quadruplex propensity |
|
GC-SpliceNet |
proprietary in-house algorithm |
Deep-learning splice-site probability; in-silico masking differential |
|
GC-SpliceMotif |
proprietary in-house algorithm |
Maximum-entropy splice-site scoring, latent cryptic-site scans |
|
GC-ESRScan |
proprietary in-house algorithm |
Exonic splicing regulatory element motif scanning |
Stating the licensing posture in a whitepaper has a practical purpose. Academic publication cares about reproducibility of method; open distribution cares about downstream freedom; commercial development cares about whether a component may be incorporated into a product. These are three different conditions. The public tools used here generally carry permissive licences, so method disclosure and reproduction are unconstrained, and the proprietary in-house algorithms are distinguished by name so that which computations fall in that category is documented.
GC-ThermoFold is the collective name for the thermodynamic layer: minimum-free-energy folding, duplex binding energy (ΔG_bind = ΔG_duplex − ΔG_A − ΔG_B), local accessibility profiling, and G-quadruplex propensity. Wherever this document says 'secondary structure', 'accessibility' or 'ΔG', this is the component doing the work. The physical model underneath is the published nearest-neighbour thermodynamic parameter set, so the parameters themselves are verifiable against the literature.
2.5 Defining an off-target — four buckets
A count of alignment hits cannot be used as an index of off-target risk. A hit against another isoform of the target gene is on-target, not off-target; a hit against a same-family paralog may be desirable or harmful depending on the therapeutic hypothesis; and a hit against another species' ortholog is a property preclinical species selection actively requires. Summing all three into one number causes excellent candidates to be discarded on a raw count.
|
Bucket |
Definition |
Interpretation |
Treatment in design |
|---|---|---|---|
|
On-target extension |
Another isoform of the target gene |
Coverage, not off-target — often the broader the better |
Quantified separately as cleavage coverage |
|
Paralog |
Another member of the same gene family |
Allow or exclude depending on the therapeutic hypothesis — not automatically decidable |
Handled as a hard filter under explicit instruction |
|
Ortholog |
The corresponding gene in another species |
A property required for preclinical species selection |
Presented as per-species complementarity with mismatch positions |
|
Unrelated hit |
Anything not in the three above |
A true off-target — risk graded by seed context and position |
Ranked by seed-based scoring and positional weighting |
Risk is not uniform even within the unrelated bucket. Full complementarity across the seed (guide positions 2–8) produces miRNA-like repression that lowers expression without cleavage, and the phenomenon is predicted from the distribution of seed complementarity in target 3' untranslated regions. Conversely, hits carrying many mismatches outside the seed often do not translate into repression despite a high alignment score. Off-target assessment is therefore a question of mechanism rather than of alignment; alignment is only the input to that assessment.
2.6 The double-stranded synthesis specification
The point at which design output and synthesis orders most often diverge is the overhang convention. This platform's double-stranded standard is a 21 bp duplex core carrying a 2 nt native overhang at each end, so each strand is 23 nt.
sense 5'- UAGGUCAUCGAUGCUAGCUAG GC -3' (21 + 2 = 23 nt)
antisense 5'- CUAGCUAGCAUCGAUGACCUA AG -3' (21 + 2 = 23 nt)
└────── 21 bp core ──────┘ └┘ native 2 nt overhang
antisense overhang AG = revcomp(upstream CU)
Keeping the overhang bases native to the target context has two reasons. First, an artificial thymidine dimer is unrelated to the target, so there is no basis on which to predict what those two nucleotides interact with in the real molecule during off-target assessment. Second, a native overhang extends complementarity to the target RNA two nucleotides beyond the core, which keeps 3'-end recognition and loading geometry interpretable.
The acceptance test is not length but full-length reverse complementarity. Two strands that are each other's reverse complement across a 21 bp core meet the specification; a sequence that is merely 23 nt long does not. This check runs automatically at export, and keeping sequences produced by paths that bypass that check out of the synthesis order is the essence of specification control.
2.7 Progressive refinement — where computational cost is placed
The pipeline places the most expensive evaluations last. Enumeration and first-pass filters run over thousands of candidates in seconds; transcriptome-wide off-target scanning and deep-model inference run over a few dozen. This is not merely a performance optimisation but an allocation of information value: spending precise evaluation on candidates that will obviously fail wastes compute and, more importantly, forfeits precision that could have been applied to the finalists.
|
Stage |
Scale |
Nature of the evaluation |
Kind of rejection |
|---|---|---|---|
|
Enumeration and composition filters |
thousands |
Deterministic rules |
GC composition out of range, homopolymer runs, immune-stimulatory motifs |
|
Rule-based efficacy scoring |
hundreds |
Ensemble of literature-derived positional rules |
Insufficient thermodynamic asymmetry, positional rule violations |
|
Population variant masking |
~150 |
Database lookup |
Common polymorphism at the target site, overlap with pathogenic variants |
|
Structure and thermodynamic filters |
~30 |
Free-energy computation |
Self-hairpin, seed binding too strong or too weak, poor local accessibility |
|
Off-target screening |
~30 |
Transcriptome-wide alignment plus mechanistic interpretation |
Seed complementarity in unrelated genes, junction-proximal hits |
|
Accessibility scoring |
~30 |
Ensemble accessibility plus interaction energy |
Target site buried in structure |
|
Deep-model inference |
~30 |
Ensemble of pretrained models |
Low predicted potency |
|
Chemistry and regulatory checks |
top few |
Rules plus databases |
Chemistry constraint violations, preclinical species requirements unmet |
3. Modality I — siRNA / RNAi (Double-Stranded Catalytic Silencing, RISC Cleavage)
3.1 Form and definition
The common name is siRNA (small interfering RNA), and the operating principle is called RNAi (RNA interference). A synthetic duplex with a 21 bp double-stranded RNA core carrying a 2 nt overhang at each end, 23 nt per strand. The two strands are not symmetric: one (the guide, antisense) is loaded into an Argonaute protein and recognises the target, while the other (the passenger, sense) is discarded during loading. The molecule is therefore dosed as a duplex but acts as a single strand, and controlling which of the two is loaded is the first task of design.
The property that most fundamentally separates this modality from every other is catalysis. The loaded guide–protein complex cleaves a target and then moves on to the next one rather than being consumed. One molecule removes many, and that turnover produces strong repression at low molar concentration. This is why the required dose differs fundamentally from a stoichiometrically acting steric blocker.
3.2 Molecular mechanism
3.2.1 Loading and strand selection
In the cytoplasm the synthetic duplex associates with an Argonaute protein to form the RNA-induced silencing complex. During assembly, whichever end of the duplex frays more easily opens first, and the strand whose 5' terminus sits at that loose end is retained as the guide. This asymmetry rule was reported independently in the same year by Khvorova and colleagues (Cell 2003) and Schwarz and colleagues (Cell 2003), and has underpinned every rational siRNA design approach since.
In design the rule is implemented as a free-energy differential. Nearest-neighbour free energies are computed for the four or five terminal base pairs at each end of a candidate duplex, and designs in which the intended antisense 5' end is the destabilised one are rewarded. Candidates with symmetric ends are penalised: both strands compete for loading, halving the effective on-target dose and doubling the passenger-derived off-target burden.
Loading also depends on the identity of the 5' nucleotide itself. The MID domain of Argonaute recognises the guide's 5' nucleotide in a base-specific manner, and crystallographic work reported a preference for uridine or adenosine (Frank, Sonenberg and Nagar, Nature 2010). The practically important point is that this 5' nucleotide is buried in the MID pocket and does not base-pair with the target. A strong candidate whose 5' end is guanosine or cytidine therefore need not be discarded; the standard remedy is to substitute that one position with uridine. Where such a substitution is made, the claim of 'full complementarity' must be restated as 'fully complementary from position 2 onward, with position 1 as the MID residue'.
3.2.2 Cleavage — what determines the rate
Once loaded, the guide searches for its target, and on encountering a fully complementary one the PIWI domain of Argonaute cleaves the target phosphodiester bond. The cleavage rate constant varies substantially with sequence; a recent quantitative study using 22-nt guides reported over 250-fold variation even among fully complementary targets (Wang and Bartel, Mol Cell 2024). The positional determinants identified in that work are as follows.
|
Guide position |
Preferred identity |
Reported rate difference |
Interpretation |
|---|---|---|---|
|
10 |
purine (A or G) |
8.2 – 9.4× |
The single strongest determinant. The evidence supports a purine/pyrimidine dichotomy; there is no basis for distinguishing A from G |
|
17 |
W (A or U) |
2.1 – 6.3× |
G and C are the slow identities; there is no basis for distinguishing C from G |
|
7 |
W (A or U) |
1.1 – 4.4× |
Both G and C are slow. Treating position 7 = C as optimal would equate an identity the literature calls slow with the fastest one |
|
6–7 junction |
weak pairing or a mismatch |
partially redundant with the above determinants |
A backbone kink driven by a sugar pucker change at position 6 promotes cleavage — the same mechanism as the position 7 determinant |
|
All three aligned |
— |
4.7 – 51× |
Cumulative effect observed in triple variants |
These determinants were independently corroborated by a 2026 cryo-electron microscopy study of catalytic activation in human Argonaute 2. That work showed that guide–target base pairing alone is insufficient for slicing and that duplex distortion is required, and reported that a pyrimidine at target position 10 optimally aligns a catalytic residue. A pyrimidine on the target corresponds to a purine on the guide, exactly matching the position-10 determinant above. The same study observed that a kink after guide nucleotide 6 releases the seed-only pairing conformation and promotes the extended pairing catalysis requires, corroborating the 6–7 axis structurally.
3.2.3 A common misreading of central pairing strength
Design guidance frequently advises raising GC content in the central region to strengthen pairing. This is a misreading that comes from citing the literature without its conditional clause. Wang and Bartel (2024) state explicitly that formation of a continuous helix does not limit the cleavage rate of fully complementary targets, and report that hydroxyl-radical footprinting profiles for four different guides were similar to one another despite a 250-fold spread in rate. Thermodynamic conformational occupancy is not the rate-limiting step for fully complementary cleavage.
The widely quoted 600-fold figure comes from a different condition. It is the range of tolerance to 3' mismatches, measured on 16 bp targets where complementarity ends at target position 16. Guides with strong central pairing barely slow on such partially complementary targets; guides with weak central pairing slow by more than a hundredfold, and in the extreme case by 600-fold. The original text is conditional: pairing beyond position 16 is dispensable for efficient slicing only when the central region has high predicted pairing stability.
Two design implications follow, and both run counter to conventional advice. First, in a fully complementary design central GC content is not a determinant of on-target potency. Second, the direction inverts: strong central pairing means high tolerance to 3' mismatches, which means the guide also cleaves partially complementary off-targets efficiently. From a specificity standpoint, rewarding central GC promotes candidates that cleave off-targets well. The correct form of the axis is not a main-effect bonus but a conditional risk term — flagging combinations of weak central pairing with weak 3' pairing — and it belongs on the off-target risk axis, not the on-target score.
The structural evidence points the same way. The catalytic activation study reports that expansion of the central major groove positions the scissile phosphate, and GC-rich duplexes resist distortion, so structure does not support a central GC bonus either. A separate structural study reports that on fully complementary binding an N-domain rotation licenses rapid slicing — but that licensing is a function of complementarity, not of GC content.
3.2.4 The mechanism of off-targets — seed-mediated repression
Binding that is not fully complementary also represses. When guide positions 2–8, the seed, are complementary to a target's 3' untranslated region, the result is miRNA-like translational repression and transcript destabilisation. Jackson and colleagues (Nat Biotechnol 2003) first showed the scale of this by microarray, and Birmingham and colleagues (Nat Methods 2006) established that 3' UTR seed matches, rather than overall identity, explain the off-target signature.
This mechanism imposes two design requirements. First, off-target assessment must be performed at seed granularity rather than by full-length alignment: a single seed 7-mer occurs in hundreds of transcripts, so a candidate with no full-length hits can still carry a broad burden. Second, it becomes possible to introduce a modification that weakens binding within the seed and thereby selectively lowers seed-mediated repression. That strategy exploits an asymmetry — the effect on a fully complementary on-target is small, while the effect on seed-only binding is large.
3.2.5 Innate immune stimulation
Synthetic double-stranded RNA can stimulate innate immunity in a sequence-dependent way. Judge and colleagues (Nat Biotechnol 2005) showed that particular U- and GU-rich motifs trigger Toll-like-receptor-mediated cytokine responses, and reported that 2'-O-methyl substitution suppresses them. In design, known stimulatory motifs are filtered at enumeration and residual risk is lowered by 2'-O-methyl placement during the chemistry stage. There is a constraint: 2'-O-methyl substitution early in the seed (positions 2–5) has been reported to reduce Argonaute loading, so placement must balance immune suppression against loading efficiency.
3.3 Clinical precedent and patent landscape
This modality has the deepest approval record among RNA therapeutics. Since the first approval in 2018 of a lipid-nanoparticle formulation against transthyretin, a succession of N-acetylgalactosamine-conjugated products targeting hepatocytes has established 'hepatic silencing' as a mature development path.
|
Year |
Target |
Delivery |
Indication area |
What the approval established |
|---|---|---|---|---|
|
2018 |
TTR |
Lipid nanoparticle, intravenous |
Hereditary transthyretin amyloidosis |
First demonstration that siRNA produces therapeutic effect in humans |
|
2019 |
ALAS1 |
GalNAc conjugate, subcutaneous |
Acute hepatic porphyria |
Practicality of the subcutaneous GalNAc route |
|
2020 |
HAO1 |
GalNAc conjugate, subcutaneous |
Primary hyperoxaluria type 1 |
Repeat application in a rare metabolic disease |
|
2020–2021 |
PCSK9 |
GalNAc conjugate, subcutaneous |
Hypercholesterolaemia |
Twice-yearly dosing — durability as a clinical differentiator |
|
2022 |
TTR |
GalNAc conjugate, subcutaneous |
Hereditary transthyretin amyloidosis |
Generational replacement of delivery on the same target |
|
2023 |
LDHA |
GalNAc conjugate, subcutaneous |
Primary hyperoxaluria type 1 |
Reproducibility of the route |
The patent landscape has three central axes. The first is chemical placement patterns — claim families specifying alternating positional arrangements of 2'-O-methyl and 2'-fluoro together with terminal phosphorothioate placement, the so-called enhanced stabilisation chemistry family. The second is the triantennary N-acetylgalactosamine ligand and its linker architecture. The third is off-target mitigation modifications such as acyclic sugar analogues placed in the seed. Because these claims are generally constructed as 'a specific modification at a specific position', emitting an explicit per-position chemistry map at design time is itself the input to a freedom-to-operate review.
Patent statements here are informational and are not legal opinion. Freedom-to-operate depends on jurisdiction and claim construction and is the domain of qualified counsel.
3.4 Chemistry requirement — mandatory, with positional rules
Chemical modification is not optional in this modality. An unmodified RNA duplex degrades within minutes in plasma, stimulates innate immunity, and is not taken up by cells. Modification serves four purposes: nuclease resistance, immune suppression, seed off-target mitigation, and providing an attachment point for conjugation.
|
Modification |
Role |
Typical placement |
Caution |
|---|---|---|---|
|
2'-O-methyl |
Nuclease resistance, immune suppression |
Alternating across both strands |
Heavy placement early in the seed (2–5) reduces loading |
|
2'-fluoro |
Maintains binding affinity, local stabilisation |
Alternating with 2'-O-methyl |
Excessive total content carries cellular and cost burden |
|
Phosphorothioate |
Terminal exonuclease resistance, protein binding |
1–2 linkages at each terminus |
Beyond the termini it buys nothing under encapsulation and costs potency |
|
Acyclic sugar analogue |
Weakens seed binding to mitigate off-targets |
One specific position within the guide seed |
Small effect on-target, but position choice is decisive |
|
5'-terminal stabilisation |
Loading efficiency and terminal protection |
Guide 5' terminus |
An enhancement for duplexes; an absolute requirement for single strands (see Chapter 6) |
|
N-acetylgalactosamine conjugation |
Hepatocyte receptor-mediated uptake |
Passenger strand terminus |
Linker chemistry (O/S/C) changes in vivo stability |
The most important positional rule is protection of the cleavage site. Argonaute cleaves the target phosphate corresponding to guide positions 10–11, so a modification there that impedes cleavage damages catalysis, while leaving the region entirely bare lowers stability. Design resolves the conflict with the constraint 'only modifications that do not block cleavage are permitted'.
3.5 Delivery requirement — required
A duplex cannot carry a fully phosphorothioate backbone: duplex stability and Argonaute loading would both be compromised. The carrier-free uptake strategy that works for single strands therefore barely applies here, and a delivery vehicle is effectively mandatory. There are two branches.
- Hepatocyte targets — a triantennary N-acetylgalactosamine ligand conjugated to the passenger terminus. The asialoglycoprotein receptor is expressed at very high density on the hepatocyte surface and recycles rapidly, so subcutaneous dosing concentrates efficiently in the liver. This is why the route accounts for the majority of approvals in this modality.
- Non-hepatic targets — ionisable-lipid nanoparticles, formulations with an added organ-tropic lipid, extracellular vesicles, polymeric carriers, or virus-like particles. On systemic administration, biodistribution shifts markedly with plasma protein adsorption, so surface composition and the protein corona are the practical determinants.
The delivery choice feeds back into chemistry. Under encapsulation the particle performs both protection and uptake, so phosphorothioate beyond the termini costs potency for no gain; on a conjugate route the receptor performs part of the uptake, so the required backbone burden is intermediate. Chapter 6 treats this relationship as a quantitative window.
3.6 Design deliverables
- Per-strand sequences — 23 nt notation including the 21 bp core and native 2 nt overhangs
- Per-position chemistry map — sugar string, backbone string, 5' terminal handling, conjugation point
- Efficacy predictions — rule-based ensemble and deep-model ensemble reported separately
- Strand-selection asymmetry — terminal free-energy differential and the resulting verdict
- Cleavage determinant profile — identities at guide positions 7, 10 and 17, and flexibility at 6–7
- Off-target table — four-bucket classification with seed-context scoring
- Cleavage coverage — how many of the target gene's transcripts a cut at this site destroys, with coding-only and all-transcript counts separated
- Variant avoidance — population allele frequencies and pathogenic variant overlap
- Accessibility profile — local structure at the target site and binding free energy
- Cross-species table — per-species ortholog complementarity and mismatch positions
- Delivery specification — conjugate design or formulation candidates with route-dependent absorption summary
- Regulatory pre-flight — automatic verdict on the two-species preclinical expectation
- Synthesis impurity projection
3.7 Advanced capabilities — the deep ensemble and special design modes
3.7.1 Why five lineages are run at once
The efficacy prediction layer runs five models from different architectural lineages simultaneously in the ACTIVE state. The reason for not simply picking the single best-performing model is not accuracy but independence of bias. Models trained on the same data tend to be wrong in the same places, and that shared error is not removed by ensembling. Agreement among models with different representations and inductive biases, by contrast, generalises better than the confidence of any one of them.
|
Model lineage |
What it sees |
Its distinctive contribution |
|---|---|---|
|
Transformer with RNA language-model embeddings |
Long-range context across the whole sequence, transferring a pretrained RNA representation |
Off-target-aware training — a representation optimised jointly for efficacy and specificity |
|
Transformer combined with convolution |
Local motifs (convolution) combined with global context (attention) |
Retains short positional patterns without losing global context |
|
Chemistry-aware multi-view with cross-attention |
Sequence and chemical modification map encoded as separate views, then cross-referenced |
The only lineage that addresses modified molecules directly — models trained on unmodified data extrapolate when chemistry changes |
|
In-house spline-based model |
B-spline activations expressing positional non-linearity locally |
Smooth positional dependence that connects interpretively to the rule-based axes |
|
Seed-mediated off-target model |
A learned representation of seed–target binding |
Reinforces the paralog / ortholog / unrelated three-bucket classification (section 2.5) |
The five outputs are combined by calibrated weighting rather than plain averaging. Because model predictions are correlated, a plain average leaves the correlated error intact, while weighting that reflects the correlation structure reduces variance further. The deliverables present the individual model scores alongside the combined score so that candidates with large inter-model disagreement can be identified — large disagreement is itself the signal that a candidate lies outside the training distribution.
3.7.2 Cross-checking the rule axis against the learned axis
Presenting the rule-based ensemble and the deep ensemble separately follows the same logic. When both families rank the same candidate highly, confidence is high; where they diverge, that point deserves review. In particular, a candidate the rules like and the learned models score low often has no similar case in the training data, while the converse may indicate a contextual effect the rules do not capture. Collapsing the two into one number destroys that information.
3.7.3 Allele-selective design
In dominant-negative disease only the mutant allele should be lowered while the normal allele is preserved. The two alleles are identical apart from one position, so the guide must be made to discriminate that single base. Design places the variant at the position within the guide where discrimination is greatest and, where necessary, introduces a second artificial mismatch to weaken binding to the normal allele further.
The practical difficulty is a conflict between discrimination ratio and absolute potency. Placing the variant at a highly discriminating position raises selectivity but can also lower absolute potency against the mutant allele. The deliverables present predicted binding and potency for both alleles side by side so the selectivity ratio can be read directly.
3.7.4 Virus mode — designing on conserved regions
For a rapidly varying pathogen, designing against a single reference sequence selects resistance variants quickly. Virus mode takes a user-supplied multi-strain sequence set as input, identifies regions of low variation from the cross-strain alignment, and restricts design to them. In parallel it screens off-targets against host transcriptomes (human, mouse, non-human primate) so that only candidates leaving host genes untouched remain.
The deliverables show, per candidate, in how many strains it is fully conserved and where mismatches arise in which strains. That figure is the direct index of breadth of efficacy and the basis for ranking a candidate that is robust across many strains above one that is perfect against a single reference.
3.7.5 Cross-species reactivity — 27 species
The design layer holds transcriptome references for 27 species including human and evaluates cross-reactivity through predefined species combination modes. The list covers rodents, several non-human primates, companion animals, livestock, poultry and fish, so that preclinical species selection and veterinary indication extension are addressed together. For each species the ortholog transcript is aligned and per-candidate mismatch counts and positions are computed, with cross-reactivity interpreted differently depending on whether the mismatch falls in the seed.
3.7.6 The chemistry blueprint as checkable rules
The chemistry blueprint is emitted not as free description but as a set of checkable rules. Each is judged pass, warn or fail with the supporting figure shown.
- Is the duplex formation free energy within the specified window — too stable lowers turnover, too unstable leaves binding insufficient
- Is strand-selection asymmetry secured — the terminal free-energy differential
- Does the seed-weakening modification sit where the rules require — the recommended position shifts with seed GC content
- Is there no cleavage-blocking modification at the position corresponding to the cleavage site
- Does the phosphorothioate placement fall inside the productive window of the chosen delivery route
- Is conjugation on the passenger terminus rather than the active strand
- Do immune-stimulatory motifs remain, and if so are they masked by 2'-modification
3.7.7 Regulatory pre-flight and synthesis impurity projection
Two development-readiness artifacts accompany the top candidates. One is an automatic verdict on the preclinical species requirement, confirming that the chosen species combination includes one rodent and one non-rodent and that the candidate has sufficient complementarity in those species. The other is a synthesis impurity projection, giving the relative prominence of deletion sequences, addition sequences, incomplete sulfurisation and diastereomers derived from the sequence and chemistry (see Chapter 21). Both are drafts and do not replace final specification setting.
3.8 What this modality cannot solve
- Nuclear-retained transcripts — the Argonaute complex acts principally in the cytoplasm, so an RNA that functions in the nucleus and never exits is not efficiently removed by this mechanism
- Targets that must be increased — silencing runs in the opposite direction in principle
- Correction of point mutations — quantity can be lowered but sequence cannot be repaired, although allele-selective silencing of a mutant allele is a viable strategy
- Premature stop codons — lowering a transcript does not address an already-truncated protein
- Systemic delivery beyond the liver — not a limitation of this platform but the bottleneck of the industry as a whole, and the area where effort should be shifted to vehicle exploration
4. Modality II — ASO Gapmer (Single-Stranded Catalytic Silencing, RNase H1 Cleavage)
4.1 Form and definition
The common name is antisense oligonucleotide (ASO), and this particular architecture is called a gapmer. A single-stranded oligonucleotide of 16–20 nt divided internally into three compartments. The two wings are filled with 2'-modified nucleotides and carry target affinity and nuclease resistance; the central gap remains 2'-deoxy. Typical configurations are 5-10-5 or 5-8-5, and the backbone is fully or largely phosphorothioate. This three-compartment architecture is why the molecule is called a gapmer.
The reason the gap exists is singular. The host enzyme RNase H1 recognises DNA:RNA heteroduplexes and cleaves the RNA strand, and that recognition requires a contiguous 2'-deoxy stretch above a minimum length. Introducing even one 2'-modification into the gap destroys the heteroduplex and abolishes cleavage, demoting the molecule from a catalytic silencer to a plain steric blocker. Conversely, without wings the molecule lacks the affinity and stability to function in vivo. The architecture is therefore not an aesthetic choice but the spatial separation of two opposing requirements.
4.2 Molecular mechanism
4.2.1 Why it is catalytic
Once RNase H1 has cleaved the target RNA, the heteroduplex dissociates and the oligonucleotide is released intact. It then forms a new heteroduplex with the next target molecule, so it has turnover. In this respect it shares the character of Modality I, but the catalyst differs: I is an Argonaute protein that cuts by itself, whereas II builds a structure that a host enzyme recognises and cuts. That difference determines where in the cell each modality operates.
4.2.2 Nuclear activity as the decisive advantage
RNase H1 is present in both nucleus and cytoplasm, with substantial nuclear activity. This modality can therefore address targets an Argonaute complex reaches poorly: unspliced pre-mRNA, intronic sequence, and long noncoding RNAs that remain in the nucleus. Practically this widens the target space considerably — a gene with no suitable site in its mature mRNA may have one in an intron or a 5' untranslated region, and a nuclear-retained transcript is difficult to address by any other mechanism.
Site selection is also freer. Modality I is effectively fixed in length and geometry by Argonaute's requirements, whereas a gapmer's length, gap width and wing composition can all be tuned to the local structure and binding strength of the target site. When a target site is buried in strong secondary structure, moving the wing chemistry toward higher affinity to improve invasion is a legitimate response.
4.2.3 Cleavage site and the meaning of coverage
Argonaute cuts one phosphodiester bond precisely; RNase H1 cuts at several positions within the window the gap covers. The 'cleavage site' is therefore an interval rather than a single coordinate, and is reported as a window. The arithmetic of how many transcripts a cut destroys is mechanism-independent — it only asks how many transcripts carry that sequence — but the notation of where the cut occurred is mechanism-specific.
4.2.4 Carrier-free uptake — the backbone is the uptake mechanism
A phosphorothioate backbone binds broadly to plasma and cell-surface proteins. That binding retains the oligonucleotide in circulation and pushes it into endocytic routes. Stein and colleagues (Nucleic Acids Res 2010) showed that simply exposing cultured cells to oligonucleotides without transfection reagent produces gene silencing, and named the phenomenon gymnosis. The observation fundamentally changed this modality's development path, because it allowed entry into preclinical work without a separate delivery-vehicle programme.
Carrier-free uptake is not free, however. The phosphorothioate fraction that creates uptake also creates class toxicity arising from protein binding — complement activation, platelet effects, hepatotoxicity. Design in this modality is therefore not the question 'how much PS is needed' but 'which range is productive', and that range depends on the delivery route. Chapter 6 gives the quantitative model.
4.2.5 Hepatotoxicity — this modality's development bottleneck
Depending on sequence and chemistry, gapmers can cause hepatocyte toxicity, and this is the single most common cause of failure to enter the clinic in this modality. Mechanistic work suggested the toxicity is associated not with chemical burden alone but with RNase H1-mediated activity itself — that is, with cleavage of unintended transcripts (Burel and colleagues, Nucleic Acids Res 2016). A tendency for high-affinity bicyclic modifications to raise the risk has also been reported (Kasuya and colleagues, Sci Rep 2016).
Omitting this axis at the design stage produces a predictable, repeating failure. Ranking on potency alone brings candidates with high-affinity chemistry and strong binding to the top, and those are precisely the properties correlated with toxicity risk. Evaluating potency and toxicity in parallel is therefore not optional in this modality, and design emits a multi-tier risk classification based on sequence motif and chemistry combination alongside potency.
4.3 Clinical precedent and patent landscape
This modality has the longest approval history. The first product, dosed locally into the eye, was approved in 1998 and later withdrawn from the market; a systemically dosed gapmer approved in 2013 brought class-toxicity management into focus. From 2018 onward, generational changes in chemistry and conjugation drove approvals in central nervous system and metabolic disease.
|
Year |
Target |
Route |
Indication area |
What the approval established |
|---|---|---|---|---|
|
1998 |
A viral transcript |
Intraocular, local |
Cytomegalovirus retinitis |
First demonstration that an antisense oligonucleotide can be a medicine |
|
2013 |
APOB |
Subcutaneous, carrier-free |
Homozygous familial hypercholesterolaemia |
Feasibility of systemic carrier-free dosing, and the necessity of class-toxicity management |
|
2018 |
TTR |
Subcutaneous, carrier-free |
Hereditary transthyretin amyloidosis |
Direct competition with Modality I on the same target |
|
2023 |
SOD1 |
Intrathecal |
SOD1-associated amyotrophic lateral sclerosis |
Establishment of direct central nervous system dosing |
|
2023 |
TTR |
GalNAc conjugate, subcutaneous |
Hereditary transthyretin amyloidosis |
Receptor-targeted conjugation applied to a gapmer, sharply lowering dose |
|
2024 |
APOC3 |
GalNAc conjugate, subcutaneous |
Familial chylomicronaemia syndrome |
Extension of conjugated gapmers into metabolic disease |
The patent landscape centres on wing chemistry. A second-generation architecture using 2'-O-methoxyethyl wings and a 2.5-generation architecture using constrained-ethyl bicyclic sugars each carry substantial claim families. On top of these sit 5'/3' N-acetylgalactosamine conjugation and its linker architecture. A more recent layer specifies stereochemical control of the phosphorothioate backbone — that is, defined stereoisomers rather than a racemic mixture.
4.4 Chemistry requirement — mandatory and structurally enforced
In this modality chemistry is not an add-on but the definition of the molecule. An unmodified DNA oligonucleotide has a plasma half-life of minutes and insufficient target affinity to function systemically. The gapmer architecture is the resolution of that problem into 'where to modify and where not to'.
|
Compartment |
Chemistry |
What it carries |
Consequence of violation |
|---|---|---|---|
|
5' wing (typically 3–5 nt) |
2'-O-methoxyethyl, bicyclic sugars (LNA, cEt) and similar |
Target binding affinity, exonuclease resistance |
Too short or too low-affinity and binding is insufficient |
|
Central gap (typically 8–10 nt) |
2'-deoxy maintained, phosphorothioate backbone |
RNase H1 recognition and cleavage |
Introducing a 2'-modification abolishes catalysis — the molecule becomes a steric blocker |
|
3' wing (typically 3–5 nt) |
Same class as the 5' wing |
Affinity, 3'-exonuclease resistance |
Failure to protect the 3' terminus causes rapid degradation |
|
Backbone overall |
Phosphorothioate, full length or nearly so |
Plasma stability, protein-binding-driven carrier-free uptake |
Too little and uptake fails; too much and class toxicity appears |
|
Terminal conjugate |
N-acetylgalactosamine or other targeting ligand |
Receptor-mediated hepatocyte uptake, sharply reducing dose |
Insufficient linker stability causes premature loss of the ligand |
High-affinity bicyclic modifications are powerful but double-edged. They allow sufficient affinity from a short oligonucleotide, but they also strengthen unintended binding at partially complementary sites and correlate with hepatotoxicity risk. Wing chemistry should therefore be chosen as 'as much as the target site structure requires' rather than 'the strongest available', and design determines wing chemistry together with the local accessibility of the target site.
4.5 Delivery requirement — optional; unnecessary on many routes
The practical strength of this modality is that a development path exists without a delivery vehicle. Four routes coexist.
|
Route |
Delivery mode |
When it fits |
Constraint |
|---|---|---|---|
|
Systemic carrier-free |
Subcutaneous or intravenous, no carrier |
Tissues with efficient uptake such as liver and kidney; rapid preclinical entry |
Requires large doses, widening the scope for class toxicity |
|
Receptor conjugate |
GalNAc conjugate, subcutaneous |
Hepatocyte targets; dose falls sharply versus carrier-free |
Does not extend to non-hepatic tissue |
|
Local direct administration |
Intrathecal, intraocular, inhaled |
Central nervous system, eye, airway; minimises systemic exposure |
The administration itself is invasive or device-dependent |
|
Particle encapsulation |
Lipid nanoparticles and similar |
Tissues the three routes above do not reach |
Under encapsulation the phosphorothioate requirement falls, so the chemistry must be redesigned |
The fourth row matters. For the same sequence, whether it goes carrier-free or in a particle changes the optimal chemistry. On the carrier-free route most of the backbone must be phosphorothioate for uptake to occur; under encapsulation the particle performs protection and uptake, so the same level of phosphorothioate sacrifices potency and safety for nothing. Optimising chemistry without fixing the route yields a value optimal for neither.
4.6 Design deliverables
- A sequence at mechanism-appropriate length with a per-position chemistry map showing the wing–gap–wing architecture
- Rationale for gap width and wing composition, with an assessment of local structure at the target site
- DNA:RNA heteroduplex melting temperature modelling and target binding strength
- Multi-tier hepatotoxicity risk classification based on sequence motif and chemistry combination
- Exonic splicing regulatory element positional-weight-matrix scan — avoiding unintended splicing perturbation
- Risk score for creation of latent cryptic splice sites
- Junction-aware off-target filter — hits proximal to splice junctions flagged separately
- Four-bucket off-target classification with seed-context assessment
- Cleavage coverage measured on the gap window — how many transcripts a cut here destroys
- Carrier-free uptake diagnosis — phosphorothioate fraction, 2'-modification coverage, conjugation status, and the resulting verdict
- Comparison of chemistry alternatives per delivery route (carrier-free / conjugate / encapsulated)
- Synthesis impurity projection
4.7 Advanced capabilities — five mechanisms, toxicity prediction, splicing-interference avoidance
4.7.1 The five supported mechanisms
The design layer distinguishes five mechanisms explicitly, and once the mechanism is set the chemistry rules and evaluation axes change automatically. Specifying the wrong mechanism does not raise an error; it simply produces a different molecule, which makes mechanism specification the first decision of the design.
|
Mechanism |
What it does |
Mandatory structural requirement |
Representative clinical lineage |
|---|---|---|---|
|
Gapmer |
Recruits RNase H1 to cleave the target RNA |
A central 2'-deoxy gap is mandatory |
Systemic metabolic disease, hepatic targets, central nervous system |
|
Splice switching |
Changes the splicing decision |
No gap; uniform high-affinity modification |
Treated in depth as Modality III (Chapter 5) |
|
Steric block |
Physically obstructs translation initiation or protein binding |
No gap; uniform modification |
Manipulation of upstream open reading frames in 5' untranslated regions and similar |
|
Anti-miR |
Sequesters an endogenous miRNA |
No gap; fully modified backbone |
Treated in depth as Modality VI-b (Chapter 10) |
|
Dual-guide class |
Places two binding elements on one molecule |
Design-specific |
Exploratory |
These five are handled in one layer because the chemistry repertoire and off-target assessment are shared. The activity axis, however, diverges completely: the gapmer rewards catalytic recognisability while the other four reward uniform occupancy (section 6.3).
4.7.2 The hepatotoxicity risk index — a weighted motif sum
Hepatocyte toxicity is the most common single cause of failure to enter the clinic in this modality, so design quantifies toxicity risk in parallel with potency. The method is an index that multiplies the occurrence frequency of particular short motifs by weights and sums them, with the weights reflecting the strength of the relative risk signal reported in the literature. The index is for relative ranking and is not calibrated to an absolute risk probability.
|
Risk band |
Motifs |
Weight |
|---|---|---|
|
High |
TCCC · TGCC |
3.0 |
|
High |
TCCT · CCTCC |
2.5 |
|
High |
TCCA · TGCT |
2.0 |
|
Moderate |
AACC · GACC · GCCT |
1.2 |
|
Moderate |
CCTG · AGCC |
1.0 |
|
Low |
CCAG · GGCA · GGCT |
0.8 |
|
Low |
CAGG · GGAG · GCAG |
0.6 |
The deliverables present the index together with a list of which motifs occurred and how often. That decomposition matters because the response differs: if the index is high because of a single high-risk motif, shifting the site by one position may resolve it, whereas an accumulation of low-risk motifs requires changing the target site altogether.
This index is correlated with the potency axis, and that is stated explicitly. Candidates with high-affinity chemistry and strong binding tend also to carry higher toxicity risk, so ranking on potency alone systematically brings risky candidates to the top. Presenting the two axes in parallel is the only way to block that bias.
4.7.3 Avoiding splicing interference in advance
When a gapmer targets pre-mRNA and the target site overlaps a splicing regulatory element, splicing can be perturbed unintentionally. Design scans candidate sites with position weight matrices for exonic splicing enhancer binding proteins to assess that risk in advance. The matrices used are four well-validated SR protein recognition motif families, each with its own threshold.
|
Matrix |
Protein family |
Role |
|---|---|---|
|
SF2/ASF |
SR protein |
A principal enhancer for exon recognition and splice site choice |
|
SC35 |
SR protein |
Involved in exon definition |
|
SRp40 |
SR protein |
Promotes use of weak splice sites |
|
SRp55 |
SR protein |
Exonic enhancer recognition |
In addition, the chemistry and sequence profiles of representative approved products are built in as reference presets, so a new design can be compared with clinical precedent on the same axes. The presets are chosen to represent distinct development contexts — central nervous system gapmers, metabolic disease gapmers, splice-switching lineages.
4.7.4 Junction-aware off-target classification
An ordinary off-target scan treats hit positions only as transcript coordinates. For a gapmer, however, a hit near a splice junction carries different risk, because heteroduplex formation there can perturb the splicing of that gene rather than merely cleaving one transcript. Design consults known junction coordinates in the host transcriptome and classifies junction-proximal hits at a separate grade.
4.7.5 Generations and the choice of wing chemistry
|
Generation |
Wing chemistry |
Character |
Selection criterion |
|---|---|---|---|
|
Second |
2'-O-methoxyethyl |
The deepest clinical precedent and the richest safety data |
The default choice where the target site is not especially difficult |
|
2.5 |
Constrained ethyl and related bicyclics |
High affinity permitting a shorter molecule |
Where the target site is buried in structure or a short molecule is needed |
|
High-affinity bicyclic |
Locked nucleic acid family |
The highest affinity |
Requires care — correlation with hepatotoxicity risk has been reported |
|
Mixed placement |
Bicyclics at selected positions only |
A compromise between affinity and risk |
Placement must avoid creating a contiguous 2'-deoxy stretch |
4.7.6 Retrovirus mode
For retroviral targets a dedicated path forces the gapmer mechanism, identifies conserved regions from a user-supplied multi-sequence set and restricts design to that range. Off-target screening against the host transcriptome runs in parallel, and because proviral sequence integrates into the host genome, checking against host-genome-derived similar sequences is treated as especially important.
4.8 What this modality cannot solve
- Upregulation — it is a silencing mechanism and runs the other way
- Splice manipulation — a gapmer cleaves its target and cannot be used to change isoform choice. Although also single-stranded, it is structurally exclusive with Modality III, which forbids the gap
- Sequence correction — it repairs neither point mutations nor premature stop codons
- Sites where catalysis is not established — where the target structure prevents gap recognition, no cleavage occurs and the molecule acts only as a steric blocker
- Complete elimination of hepatotoxicity risk — risk can be quantified and ranked but not removed; final confirmation belongs to experimental toxicology
5. Modality III — SSO / Splice-Switching Oligonucleotide (Non-Cleaving Steric Block)
5.1 Form and definition
The common name is SSO (splice-switching oligonucleotide), also called splice-switching antisense or an exon-skipping oligonucleotide. A single-stranded oligonucleotide of 18–25 nt in which every position carries a uniform high-affinity modification. The form matches Modality II but the internal composition is its inverse: where a gapmer must place a 2'-deoxy window at the centre, this modality leaves no 2'-deoxy residue anywhere, because the objective is not to recruit a nuclease.
Three chemical branches are available: uniform 2'-O-methoxyethyl with a phosphorothioate backbone; a mixmer incorporating bicyclic sugars; and phosphorodiamidate morpholino oligomers, in which the sugar–phosphate backbone itself is replaced by morpholine rings and phosphorodiamidate linkages, yielding a charge-neutral molecule. Charge neutrality sharply reduces protein binding and class toxicity, at the cost of lower carrier-free uptake efficiency.
5.2 Molecular mechanism
5.2.1 Splicing is a decision, not a fixed procedure
The conversion of pre-mRNA into mature mRNA is a competitive decision rather than a fixed sequence of steps. The spliceosome recognises donor and acceptor signals, but that recognition is not determined by the signal sequence alone. Proteins binding to regulatory elements scattered through the exon and adjacent introns — exonic and intronic splicing enhancers and silencers — push and pull on whether each site is used. Which exon is included is therefore the outcome of a probabilistic competition, and tipping that balance changes the outcome.
This modality tips exactly that balance. When the oligonucleotide binds a regulatory element and physically masks it, the protein that should bind there cannot; if the element is inhibitory, inhibition is relieved and the exon is included, and if it is enhancing or is a splice signal itself, that site goes unused and the exon is excluded. Changing the decision without destroying the target is the essence of the modality.
5.2.2 Five forms of intervention
|
Form |
Element masked |
What happens |
Therapeutic logic |
|---|---|---|---|
|
Exon inclusion |
Intronic splicing silencer |
Inhibition is relieved and a normally skipped exon is included |
Increases the amount of functional full-length protein |
|
Exon skipping |
Donor/acceptor signals or an exonic enhancer |
The exon is excluded from the mature transcript |
Restores a reading frame broken by a deletion, yielding partially functional protein |
|
Pseudoexon suppression |
A latent splice site created by a deep-intronic variant |
The artificially created exon is not used |
Restores the proportion of normal transcript |
|
Intron retention resolution |
Elements that promote retention |
A retained intron is spliced normally |
Matures a nucleus-trapped transcript and restarts protein production |
|
PolyA switching |
An alternative polyadenylation signal |
A different 3' end is selected |
Adjusts isoform ratio or transcript stability |
5.2.3 Why precision of the target window is decisive
Regulatory elements are typically 6–20 nt long, and shifting the target by a few nucleotides abolishes the effect. The intronic splicing silencer targeted in spinal muscular atrophy is the canonical example: therapeutic design only became possible once the exact coordinates and inhibitory function of that element had been established (Singh and colleagues, Mol Cell Biol 2006; Hua and colleagues in subsequent work). This modality therefore demands not 'target this gene' but 'target this coordinate in this gene', and selection of the target window determines most of the design quality.
For that reason a generic 'risk that a cryptic splice site could arise' score is of limited help in candidate selection: it belongs to a region rather than to a candidate. What is actually needed is how far this particular oligonucleotide moves splicing at this target, and design computes that as an in-silico masking differential — the window a candidate covers is masked, splice-site probabilities are recomputed, and the change in donor and acceptor probability is attributed to that candidate. Two candidates with identical generic risk scores frequently have very different masking differentials, and that difference reorders the ranking.
5.2.4 The reading-frame constraint
Exon-skipping design carries one further genetic constraint. Excluding an exon shortens the sequence by its length, and unless that length is a multiple of three the reading frame shifts, rendering everything downstream meaningless and introducing a premature stop. The clinical manifestation of this rule is the split between severe and milder phenotypes of muscular dystrophy according to the position and size of the deletion. Therapeutic design therefore computes, per patient subgroup, which exon must be skipped to restore the frame, and this is why a reading-frame validator is part of the deliverable set.
5.2.5 Interpretive consequences of not cleaving
Because this modality does not destroy its target, the 'cleavage coverage' metric used elsewhere cannot be read the same way. Here coverage means the number of transcripts carrying the site — binding range, not destruction range. Two consequences cause repeated confusion in practice and are stated explicitly. First, a design targeting an intronic element is structurally reported at zero coverage because that sequence is absent from every mature transcript; that is correct behaviour, not a defect. Second, high coverage in an exon-skipping design means 'widely bound', not 'widely knocked down'.
5.3 Clinical precedent and patent landscape
This modality established the therapeutic logic of 'repair rather than reduce'. In 2016 a treatment for spinal muscular atrophy and an exon-51 skipping treatment for muscular dystrophy were approved in the same year, validating two different forms of intervention simultaneously; further approvals followed in the muscular dystrophy family targeting different exons.
|
Year |
Form |
Target |
Chemistry and route |
What the approval established |
|---|---|---|---|---|
|
2016 |
Exon inclusion |
SMN2 intron 7 silencer |
Uniform 2'-MOE, intrathecal |
That splice manipulation can change the phenotype of a severe genetic disease |
|
2016 |
Exon skipping |
DMD exon 51 |
PMO, intravenous |
Clinical entry of charge-neutral chemistry and the frame-restoration strategy |
|
2019 |
Exon skipping |
DMD exon 53 |
PMO, intravenous |
Extension to further patient subgroups |
|
2020 |
Exon skipping |
DMD exon 53 |
PMO, intravenous |
Multiple products against the same exon |
|
2021 |
Exon skipping |
DMD exon 45 |
PMO, intravenous |
Direction toward a complete target-exon portfolio |
A pseudoexon-suppression strategy targeting a deep-intronic variant in retinal disease has entered clinical development, and is also an example of minimising systemic exposure through local administration. The patent landscape has three layers: claims over the target element and its coordinates, claims over uniform high-affinity and morpholino chemistry, and sequence claims over specific exons. The third layer bears directly on design freedom, so patent-family advisories are included in the deliverables.
5.4 Chemistry requirement — mandatory and uniform
Chemistry here serves three interlocking requirements: affinity sufficient to bind the regulatory element and displace its protein; nuclease resistance sufficient to survive until the target is reached; and, most importantly, not recruiting RNase H1. The third creates the uniformity constraint — any stretch left 2'-deoxy could be recognised as a heteroduplex and the target cleaved, at which moment the molecule ceases to be a splice switcher and becomes a gapmer.
|
Chemistry |
Character |
Advantages |
Considerations |
|---|---|---|---|
|
2'-O-methoxyethyl + phosphorothioate |
Anionic, high affinity |
The deepest clinical precedent; systemic and intrathecal behaviour well characterised |
Phosphorothioate class toxicity must be managed |
|
Phosphorodiamidate morpholino |
Charge-neutral, non-natural backbone |
Low protein binding, so little class toxicity and little immune stimulation |
Low carrier-free uptake efficiency; requires high doses or a cell-penetrating conjugate |
|
Bicyclic mixmer |
Bicyclic sugars placed at selected positions |
High affinity from a short oligonucleotide |
Placement must avoid creating a contiguous 2'-deoxy stretch |
|
2'-O-methyl + phosphorothioate |
Classical uniform modification |
Simple synthesis and lower cost |
Relatively lower affinity, compensated by length |
5.5 Delivery requirement — usually unnecessary
Most approvals in this modality are dosed without a carrier. Central nervous system targets are reached by intrathecal administration bypassing the blood–brain barrier; muscle targets rely on tissue distribution after systemic dosing; ocular targets achieve local concentration by intravitreal injection. These are cases in which the delivery problem was solved by route rather than by vehicle.
That does not mean delivery is easy. For targets where sufficient tissue concentration is hard to reach after systemic dosing — muscle, heart — the required dose becomes very large, and this operates as the practical ceiling on the modality. Efforts to improve it through cell-penetrating peptide conjugation or receptor-targeted conjugation are ongoing, and from a design standpoint the effect of chemistry and attachment point on folding and binding must be assessed together.
5.6 Design deliverables
- Candidate sequences aligned to the target window with uniform chemistry notation (2'-MOE / PMO / bicyclic mixmer / 2'-O-methyl)
- In-silico masking differential — donor and acceptor probability change after masking the covered window. The single most important metric in this modality
- Binding free energy including the cost of unfolding local pre-mRNA structure
- Cryptic splice-site sweep ±500 nt with ranking
- Exonic splicing regulatory element motif scan results
- Junction-aware off-target filter — hits proximal to known junctions in the host transcriptome flagged separately
- Reading-frame validation — confirmation that an exon-skipping design preserves the frame
- Toxicity motif scan — G-quadruplex, immune-stimulatory motifs, hepatotoxicity-associated motifs
- Binding coverage, with an explicit note that it reads as 'mature transcripts carrying the site'
- Patent-family advisories
5.7 Advanced capabilities — quantitative splicing assessment and clinical presets
5.7.1 The five mechanisms and their clinical anchors
|
Mechanism |
What it masks |
Clinical anchor |
The key design decision |
|---|---|---|---|
|
Exon skipping |
An exonic enhancer or the donor/acceptor signal |
Muscular dystrophy exons 45 / 51 / 53 |
Which exon must be skipped for the frame to be restored |
|
Exon inclusion |
An intronic splicing silencer |
The spinal muscular atrophy intron 7 element |
The exact coordinates of the silencer — a few nucleotides off and the effect is lost |
|
Pseudoexon suppression |
A cryptic splice site created by a deep-intronic variant |
Deep-intronic variants in retinal disease |
Whether to mask the cryptic site or the branch point |
|
Intron retention resolution |
Elements promoting retention |
Exploratory |
Whether retention is caused by a weak splice site or by a repressive element |
|
PolyA switching |
An alternative polyadenylation signal |
Exploratory |
Whether the switch changes isoform ratio or transcript stability |
The spinal muscular atrophy case is the reference point for coordinate precision. The targeted intronic silencer sits in intron 7 at approximately +10 to +27, and therapeutic design became possible only after those coordinates and that repressive function had been established. This modality demands not 'target this gene' but 'target this 18-nucleotide window of this gene'.
5.7.2 Quantifying splice site strength — two layers
Design assesses splice site strength in two layers. The first is a maximum-entropy site score, computed over a 9-mer window for donors and a 23-mer window for acceptors. Because the model captures dependencies between positions, it separates weak from strong sites better than a simple position weight matrix. The second is a deep-learning site probability that takes far broader sequence context as input and predicts whether that position is actually used as a splice site.
The two are used together because they measure different things. The maximum-entropy score answers 'does this sequence look like a splice site'; the deep model answers 'will it be used in this context'. Latent sites exist that carry a strong signal yet go unused, and those are precisely the cryptic sites at risk of activation once an SSO is introduced.
5.7.3 The in-silico masking differential — the central metric
A generic cryptic risk score belongs to a region and therefore discriminates poorly between candidates. The masking differential masks the window a candidate covers, recomputes splice site probabilities and attributes the change in donor and acceptor probability to that candidate. The procedure is as follows.
1. Compute baseline donor and acceptor probabilities in the target pre-mRNA context
2. Generate a sequence with the window covered by the candidate SSO masked
3. Recompute probabilities on the masked sequence
4. delta(donor) = masked - baseline, delta(acceptor) = masked - baseline
5. Judge whether the sign matches the intended direction (up for inclusion, down for skipping)
-> Two candidates with identical generic risk scores frequently have very different deltas,
and that difference reorders the ranking
Checking the sign as well as the magnitude matters. Judged on magnitude alone, a candidate that moves splicing strongly rises to the top — but if the direction is wrong, that candidate produces the opposite of the intended result. The deliverables present magnitude and sign together with a verdict on agreement with the intended direction.
5.7.4 Advance scanning for cryptic sites
A sweep of ±500 nt around the target window enumerates latent donors and acceptors and ranks them by strength. The subjects are sites currently unused but with signal strength in the middle range — weak signals are unlikely to be activated, and already-strong ones would be in use. When an SSO masks the canonical site the spliceosome looks for the next best, so these middle-strength sites are the real risk list.
5.7.5 The toxicity motif scanner
Uniform high-affinity chemistry combined with particular sequences creates three classes of risk, each scanned separately.
- G-quadruplex formation — contiguous guanosine patterns, which complicate synthesis and purification and cause non-specific protein binding
- Innate immune stimulation — GU-rich sequence (the TLR7/8 route) and unmethylated CpG (the TLR9 route)
- Hepatotoxicity-associated motifs — the weighted index of section 4.7.2 applied here as well
5.7.6 The reading-frame validator
In an exon-skipping design, if the length of the target exon is not a multiple of three the reading frame shifts and everything downstream becomes meaningless. The validator reads the exon structure of the target transcript, confirms the length of the exon to be skipped and returns a pass or fail on frame preservation. On failure it also searches for combinations in which skipping an adjacent exon as well restores the frame.
Automating this check is a practical necessity: the deletion range differs by patient subgroup, so which exon to skip changes per subgroup, and computing the combinations by hand invites omissions.
5.7.7 Clinical presets and patent landscape notation
For well-known targets, the target coordinates and mechanism of clinical precedent are built in as presets so that a new design can be compared with that precedent on the same axes. A preset carries the target intron or exon number, the relative coordinates of the target window, the mechanism and the reference literature.
This is also an area in which sequence claims exist, so designs near a preset are accompanied by advisories on the relevant patent families. The notation is informational and is not a freedom-to-operate opinion.
5.8 What this modality cannot solve
- Reducing transcript quantity — it does not cleave, so total abundance is unchanged
- Disease mechanisms that do not run through splicing — there is no point of intervention
- Deletions whose frame cannot be restored — deletion types for which no skippable exon yields a multiple of three
- Targets whose regulatory elements are uncharacterised — without coordinates, precise targeting is impossible and element mapping becomes prerequisite research
- Achieving sufficient concentration in systemic muscle and cardiac tissue — the current practical bottleneck
6. The Chemical Modification Layer
The preceding three chapters described the chemistry each modality demands individually. This chapter integrates those demands into a single decision structure. Chemical modification is not a matter of adding more of a good thing; it is an optimisation in which four objective functions conflict, and the shape of that conflict differs by modality and again by delivery route. The purpose of this chapter is to make that structure explicit.
6.1 The repertoire
Modifications applicable to a synthetic oligonucleotide fall on three axes: substitution at the ribose 2' position, substitution of the phosphodiester linkage, and terminal additions.
|
Axis |
Modification |
Physical effect |
Typical use |
|---|---|---|---|
|
Sugar 2' |
2'-O-methyl (m) |
Raises nuclease resistance, slightly raises affinity, suppresses immune stimulation |
Baseline modification across all modalities |
|
Sugar 2' |
2'-fluoro (f) |
Raises binding affinity, locks the C3'-endo conformation |
Alternated with 2'-O-methyl in duplexes |
|
Sugar 2' |
2'-O-methoxyethyl (e) |
High affinity and high stability, bulky |
Gapmer wings, uniformly modified steric blockers |
|
Sugar 2' |
Locked nucleic acid and related bicyclics (l) |
Very high affinity, strongly locked conformation |
Where high affinity is needed from a short oligonucleotide |
|
Sugar 2' |
Constrained ethyl (c) |
Bicyclic class, high affinity and relatively tolerant |
Next-generation option for gapmer wings |
|
Sugar 2' |
Acyclic sugar analogue (g) |
Locally weakens binding |
Placed in the seed to mitigate off-targets |
|
Sugar 2' |
2'-deoxy (d) |
Natural DNA; forms a heteroduplex |
The gapmer central gap — the only place it is used deliberately |
|
Backbone |
Phosphorothioate (s) |
Nuclease resistance, plasma protein binding, uptake mediation |
All modalities; its fraction determines the delivery route |
|
Backbone |
Phosphodiester (o) |
Natural; low stability but no potency cost |
Internal linkages on encapsulated routes |
|
Backbone |
Phosphorodiamidate morpholino |
Charge-neutral, almost no protein binding |
Splice-switching family only |
|
Terminal |
5'-vinylphosphonate and related phosphate mimics |
Metabolically stable 5'-phosphate mimic |
Stabilises RISC loading; an absolute requirement for single-strand RISC |
|
Terminal |
Inverted deoxythymidine |
Blocks 3'-exonucleases |
Duplex 3' terminal protection |
|
Terminal |
N-acetylgalactosamine conjugate |
Hepatocyte receptor ligand |
Hepatic targets, subcutaneous |
|
Terminal |
Lipid anchor · polyethylene glycol |
Confers amphiphilicity, delays renal clearance |
Self-assembling conjugates, half-life extension |
In the deliverables, chemistry is expressed as two strings: a sugar string of length L and a backbone string of length L−1. What matters is that this notation is complete at per-position granularity. A summary such as '60% 2'-O-methyl' is usable neither for a synthesis order nor for a patent review.
Example notation (21-mer guide)
sugar : mfmfmfmmmfmfmfmfmfmfm (m = 2'-O-methyl, f = 2'-fluoro)
ps : ssoooooooooooooooooss (s = phosphorothioate, o = phosphodiester)
5' end: (vp) (metabolically stable 5'-phosphate mimic)
Gapmer example (5-10-5, 20-mer)
sugar : eeeee dddddddddd eeeee (e = 2'-MOE, d = 2'-deoxy gap)
ps : every linkage s
6.2 Four objective functions and their conflicts
Chemistry optimisation is formulated as a weighted sum over the axes below, with a fifth axis added on carrier-free delivery routes. Each axis is normalised to the interval 0 to 1 and computed from literature-derived rules.
|
Objective |
What it rewards |
What it penalises |
Principal conflict |
|---|---|---|---|
|
Stability |
Terminal phosphorothioate, 2'-ribose protection, terminal protection |
Exposed unmodified pyrimidine dinucleotides (ribonuclease hotspots) |
Activity — the catalytic site and the gap cannot be protected |
|
Immune suppression |
2'-O-methyl near guide position 2, overall 2'-O-methyl coverage |
GU-rich motifs, UGUGU repeats, CpG dinucleotides |
Loading efficiency — 2'-O-methyl early in the seed lowers loading |
|
Activity |
Mechanism-dependent (see 6.3) |
Destruction of the structure the mechanism requires |
Stability — protection cannot cover the catalytic site |
|
Seed off-target mitigation |
Binding-weakening modification in the seed (acyclic analogue > 2'-O-methyl) |
Excessive strengthening of binding across the seed |
On-target potency — weakening the seed also costs some on-target binding |
|
Self-assembly (conjugate platforms only) |
Amphiphilic monomer geometry — balance of hydrophilic and hydrophobic termini |
Arrangements that prevent micelle formation |
Synthesis complexity |
|
Delivery fitness (carrier-free routes only) |
Route-specific productive phosphorothioate window, full 2'-modification, lipophilic or targeting ligand, metabolically stable 5'-phosphate |
Exceeding the protein-binding ceiling, excessive bicyclic content |
Activity — a high phosphorothioate load costs potency |
One design decision within this formulation deserves emphasis: real thresholds must be implemented as constraints, not as weighted terms. The carrier-free phosphorothioate floor, for example, is not 'a value it is good to exceed' but 'a value below which the molecule does not enter the cell'. Encoded as one term of a weighted sum it is outvoted by the potency term and never binds, and the search then optimises a molecule that never arrives. Designs below the floor are therefore judged infeasible rather than merely scoring low, and the final design carries an explicit verdict of pass, blocked, or relaxed.
6.3 The activity axis follows mechanism
The most frequent error in chemistry optimisation is selecting a template by the molecule's topology — single- or double-stranded. The correct criterion is the effector mechanism. Two single strands, one that recruits RNase H1 and one that must not, require opposite chemistry; and two RISC-mechanism molecules, one double-stranded and one single-stranded, have different requirements.
|
Mechanism |
Modalities |
What the activity axis rewards |
Error from judging by topology alone |
|---|---|---|---|
|
RISC cleavage (duplex) |
I, part of IV, VI-a |
Protection of the cleavage site and 5' loading asymmetry |
— |
|
RISC without cleavage |
V transcriptional activation |
Loading optimisation; the cleavage rule does not apply |
Applying the silencing cleavage-site rule constrains the design for no reason |
|
RNase H1 |
II, part of IV |
Preservation of catalytic recognition in the central gap |
— |
|
Steric block |
III, VI-b |
Uniform high-affinity occupancy; no gap permitted |
Applying a gapmer template causes unintended cleavage |
|
RISC cleavage (single strand) |
the single-strand variant of I |
5'-phosphate presentation acts as a multiplier; phosphorothioate beyond two per terminus costs potency; over-modification is penalised; there is no strand-selection term |
Copying a duplex template produces, without error, a molecule with no activity |
6.3.1 The particular case of single-strand RISC
Single-stranded siRNA is single-stranded in topology but RISC in mechanism — a combination the conventional four-effector classification cannot express. Its existence was reported independently in the same issue by Lima and colleagues (Cell 2012) and Yu and colleagues (Cell 2012), showing that Argonaute-dependent activity is achieved without a passenger strand but that a 5'-phosphate is essential for activity in vivo.
A natural 5'-monophosphate, however, is a substrate for cellular phosphatases and is removed before RISC loading when appended to a synthetic molecule. Metabolically stable analogues were therefore sought on the basis of the crystal structure of the 5'-phosphate pocket, and (E)-vinylphosphonate was reported to adopt a conformation similar to the natural phosphate while being metabolically stable (Prakash and colleagues, Nucleic Acids Res 2015). In single-strand RISC design the 5'-phosphate mimic is therefore not a beneficial enhancement but a precondition for activity, and its absence must be applied as a multiplicative penalty no other axis can compensate for.
Phosphorothioate is handled differently too. Under encapsulation, two phosphorothioates per terminus have been reported optimal for a single-strand RISC molecule, with more reducing potency. Carrier-free, the same phosphorothioate is the only means of uptake. One parameter thus operates in opposite directions depending on the delivery route, and this is the clearest case for treating the route as a premise of chemistry.
6.3.2 Duplex chemistry does not transfer to single strands
Because this repeatedly generates cost in practice, it is stated separately. Holen and colleagues (Nucleic Acids Res 2003) reported that positional and accessibility rankings transfer well between single- and double-stranded formats (reported correlation r = 0.967 over seven sites). The same work showed that chemical tolerance does not transfer: a modification equivalent to wild type in a duplex impaired activity in a single strand.
Subsequent work points the same way; three of five active duplexes lost activity on conversion to a single strand (Pendergraff and colleagues, Nucleic Acid Ther 2016). The practical rule is clear: when moving a duplex-validated candidate to a single strand, carry the sequence but not the chemistry. Central unmodified-window conventions, phosphorothioate counts and 5'-terminal handling all encode duplex assumptions.
6.4 Phosphorothioate — one parameter, three consequences
Phosphorothioate fraction is the single most consequential parameter in this layer. The same value acts simultaneously on stability, uptake and toxicity, in different directions.
|
Fraction range |
Stability |
Uptake |
Toxicity burden |
Matching delivery route |
|---|---|---|---|---|
|
0.10 – 0.30 |
Terminal protection is sufficient |
Performed by the particle |
Low |
Particle encapsulation |
|
0.25 – 0.55 |
Sufficient |
Receptor binding carries part of it |
Moderate |
Receptor or lipophilic conjugate |
|
0.75 – 1.00 |
Very high |
Phosphorothioate itself is the uptake mechanism |
High |
Carrier-free (gymnosis) |
|
Above 0.90 |
Saturated |
Little further gain |
Protein-binding burden rises steeply |
Not recommended — above the ceiling |
The ranges rest on different evidence. The lower bound under encapsulation is the minimum needed for terminal protection, and the upper bound comes from the observation that more costs potency — two per terminus on a 20-mer corresponds to 4/19 ≈ 0.21. The carrier-free floor comes from the minimum protein binding required for uptake, based on the observation that roughly 75% or more of linkages must be phosphorothioate. The ceiling at 0.90 is where class effects arising from protein binding — complement activation, platelet effects, hepatotoxicity — rise sharply.
Phosphorothioate carries one further dimension. The phosphorus atom becomes a stereocentre, so each linkage exists as two stereoisomers and ordinary synthesis produces a racemic mixture. An oligonucleotide with n phosphorothioate linkages is a mixture of 2ⁿ diastereomers, which differ in nuclease resistance and target binding. Synthetic approaches that control stereochemistry exist and carry their own patent families. From a design standpoint, whether stereocontrol is assumed changes both the chemistry specification and the quality-control item list, so the assumption is stated explicitly in the deliverables.
6.5 Safety axes created by chemistry
|
Axis |
Mechanism |
Design response |
|---|---|---|
|
Innate immune stimulation |
GU-rich and CpG motifs activate endosomal receptors |
Motif filtering at enumeration, 2'-O-methyl placement at the chemistry stage |
|
Protein-binding class effects |
Phosphorothioate binds complement and coagulation factors |
Fraction constrained within the route-specific productive window, penalised above the ceiling |
|
Hepatocyte toxicity |
High-affinity chemistry combined with sequence drives unintended cleavage and accumulation |
Multi-tier risk classification from sequence and chemistry, emitted in parallel with potency |
|
Seed-mediated off-targets |
A seed 7-mer is complementary to hundreds of 3' untranslated regions |
Binding-weakening modification in the seed; seed-level burden quantification |
|
Fluorination burden |
Cellular burden when total 2'-fluoro content is excessive |
Total-content cap and a penalty term |
|
Bicyclic burden |
Higher bicyclic content correlates with toxicity risk |
Reflected as a penalty in the carrier-free delivery fitness axis |
6.6 Chemistry deliverables
- Per-strand, per-position sugar and backbone strings with 5' and 3' terminal notation
- The selected template name and its literature basis (enhanced-stabilisation family, gapmer family, uniform-modification family, single-strand RISC family)
- Score decomposition per objective — stability, immune stimulation, activity, seed off-target (plus self-assembly and delivery fitness where applicable)
- Delivery fitness diagnosis — pass / blocked / relaxed verdict and the name of any gate that failed
- Phosphorothioate fraction and its position relative to the route-specific productive window
- Conjugate design — lipid, polyethylene glycol and ligand placement per terminus, with rendered notation
- Durability ranking and an expected dosing-interval band
- Whether stereocontrol is assumed and the quality-control items that follow
Practical warning — a sequence without a fixed chemistry specification is not an orderable specification. The same sequence becomes an entirely different molecule in activity, stability and toxicity depending on chemistry, and in particular copying a duplex template onto a single strand yields, without any error being raised, a molecule with no activity.
7. The Delivery Layer
Nucleic acids are large and strongly anionic and do not cross a lipid bilayer by free diffusion. Every oligonucleotide therapeutic must therefore secure an uptake mechanism, and in principle there are only three: encapsulate in a particle, attach a receptor ligand, or make the backbone chemistry itself bind proteins and drive endocytosis. This chapter covers the biology and physical chemistry of the three routes, how that choice propagates back through the whole design, and what is evaluated on the path from formulation to product.
One easily overlooked fact should be stated first. Getting into the cell and reaching the place where the molecule can act are different problems. Most of what enters by endocytosis remains trapped in endosomes and proceeds to lysosomal degradation; the fraction escaping to the cytosol is very small. Endosomal escape efficiencies reported for lipid nanoparticle work are only a few per cent. The bottleneck in delivery is therefore not getting in but getting out, and much of vehicle design aims at that escape efficiency.
7.1 Decomposing the delivery problem into levels
|
Level |
Question |
What determines it |
What is observed on failure |
|---|---|---|---|
|
Circulation residence |
How long does it survive in blood |
Particle size and surface hydrophilicity, backbone chemistry, the renal excretion threshold |
Disappearance shortly after dosing — no tissue exposure at all |
|
Tissue arrival |
Which organ does it reach |
Protein corona composition, surface charge, receptor ligand, vascular permeability |
Accumulation only in liver and spleen; insufficient concentration in the target tissue |
|
Cell entry |
Does it enter the target cell |
Receptor expression density and recycling rate, backbone protein binding |
Arrival in the tissue but retention in the extracellular space |
|
Endosomal escape |
Does it reach the cytosol |
The behaviour of the ionisable lipid at endosomal pH, its membrane-destabilising capacity |
Intracellular fluorescence is visible but there is no pharmacology — the most common failure point |
|
Arrival at the site of action |
Must it reach the nucleus |
Depends on modality (transcriptional activation requires the nucleus) |
Present in the cytosol but never meeting the target |
The fourth level is the bottleneck for a quantitative reason. That only a few per cent of endocytosed molecules escape to the cytosol means most of the dose is lost before reaching its target. Doubling that fraction halves the dose required for the same effect, and halving the dose halves class toxicity and manufacturing cost together. This is why ionisable lipid design attracts the largest investment in the field.
7.2 Route 1 — receptor-targeted conjugation
7.2.1 Why this works exceptionally well in hepatocytes
Conjugating a triantennary N-acetylgalactosamine ligand to an oligonucleotide terminus causes the asialoglycoprotein receptor on the hepatocyte surface to recognise and internalise it. This works exceptionally well because three biological conditions are met at once, and checking those three one by one is the checklist for extending to any other tissue.
|
Condition |
Value in hepatocytes |
Why it is needed |
What to check in another tissue |
|---|---|---|---|
|
Surface receptor density |
Very high — on the order of hundreds of thousands per cell |
A high probability of encounter means capture even at low concentration |
Receptor expression in the target cell — checkable in advance from tissue expression data |
|
Recycling rate |
Releases its ligand and returns to the surface within minutes |
Repeated use means a small receptor pool absorbs a large amount |
Does it recycle after internalisation, or route to degradation |
|
Natural distribution of the organ |
Subcutaneous dosing concentrates naturally in the liver |
The ligand's job is not to redirect but to capture what is already passing |
Is the tissue naturally exposed on that route of administration |
The third condition is the one most often overlooked. Hepatic conjugates succeed not because the ligand drags the molecule to the liver but because it captures a molecule that would pass through the liver anyway. In a tissue with low natural exposure, the same receptor density gives a far lower capture probability.
7.2.2 The geometry of multivalent binding
A single galactosamine binds the receptor weakly. Three arranged on a branched scaffold with appropriate spacing engage three binding sites of the receptor trimer simultaneously and raise affinity by orders of magnitude. What matters is not the number but the spacing and flexibility: if the branch arm lengths do not match the spacing of the receptor's binding sites, multivalent binding does not occur and the result is merely the sum of monovalent interactions. This multivalent geometry is the core of the conjugation technology and forms one axis of the patent landscape.
7.2.3 What the linker chemistry determines
|
Linker type |
Bond |
In vivo behaviour |
When it fits |
|---|---|---|---|
|
Oxygen (O) |
Phosphodiester-like linkage |
Cleaved relatively quickly by plasma nucleases and esterases |
When the ligand should be shed after hepatic arrival |
|
Sulfur (S) |
Phosphorothioate-like linkage |
High cleavage resistance, so the ligand persists |
When premature loss in circulation must be prevented |
|
Carbon (C) |
Non-cleavable covalent bond |
Not cleaved in vivo |
When retaining the ligand is harmless or required for activity |
Linker choice looks minor and is not. Premature cleavage loses the ligand before the target tissue is reached and nullifies the benefit of conjugation; conversely a non-cleavable linker can leave the ligand acting as steric bulk inside the cell and lower activity. The deliverables present attachment point and linker type bound into one specification with sequence and chemistry.
7.2.4 The ligand catalogue — extending beyond the liver
Efforts to extend the same principle to other tissues are active. The design layer holds a ligand catalogue indexed by receptor, each entry carrying the target receptor, the ligand format (sugar, peptide, antibody fragment, small molecule) and the reference binding constant as reported in the literature. Twenty-nine entries are currently registered; binding constants without literature support are not generated and are marked unconfirmed.
|
Target family |
Receptor |
Ligand format |
Tissue or cell targeted |
|---|---|---|---|
|
Hepatocyte |
Asialoglycoprotein receptor |
Triantennary galactosamine |
Hepatocytes — the established route |
|
Tumour vasculature |
Integrin αvβ3 |
Cyclic RGD peptide |
Tumour neovascular endothelium |
|
Tumour metabolism |
Folate receptor alpha |
Folate |
Folate-receptor-overexpressing tumours |
|
Immune cells |
Mannose receptor |
Mannose |
Macrophages and dendritic cells |
|
Immune cells |
Dendritic cell surface lectin, endocytic receptor, Fc receptor |
Antibody fragments |
Cross-presenting dendritic cells, monocyte lineage |
|
Prostate |
Prostate-specific membrane antigen |
Small-molecule ligand |
Prostate cancer |
|
Breast and gastric |
HER2 |
Antibody fragment |
HER2-positive tumours |
|
Epithelial tumours |
TROP2 · EGFR |
Antibody fragment · peptide |
Epithelial-origin tumours |
|
Neuroblastoma |
GD2 ganglioside |
Antibody |
Neuroectodermal tumours |
|
Gastric and pancreatic |
Claudin 18.2 |
Antibody |
Gastric and pancreatic tumours |
|
Blood–brain barrier |
Transferrin receptor |
Antibody fragment and three peptides |
Receptor-mediated transcytosis |
|
Blood–brain barrier |
Insulin-like growth factor 1 receptor |
Single-domain antibody |
Receptor-mediated transcytosis |
|
Blood–brain barrier |
LDL receptor-related protein |
Angiopep-class peptide |
Receptor-mediated transcytosis |
|
Blood–brain barrier |
Glucose transporter · glutathione transporter |
Sugar · glutathione |
Transporter-mediated crossing |
|
Brain homing |
Cyclic brain-homing peptide family |
Cyclic peptides |
Brain tissue accumulation |
|
Nervous system |
Nicotinic acetylcholine receptor |
Cyclic peptides |
Neurons |
|
Metabolic |
Leptin receptor |
Peptides |
Hypothalamus and related |
The blood–brain barrier shuttle family is a particularly active area. Receptor-mediated transcytosis binds a specific receptor on brain capillary endothelium and traverses the cell, attempting to reach the central nervous system while avoiding the invasiveness of intrathecal dosing. Crossing efficiency is low and the receptors are expressed in other tissues too, so selectivity remains the challenge.
Every binding constant in the ligand catalogue is a literature reference value. The affinity of an arbitrary ligand–receptor pair is not computed, and uncertain entries are left empty and marked unconfirmed.
7.3 Route 2 — lipid nanoparticle encapsulation
7.3.1 Four components and their roles
|
Component |
Role |
Design variables |
Failure mode |
|---|---|---|---|
|
Ionisable lipid |
Neutral at physiological pH and cationic at endosomal pH — performs both encapsulation and endosomal membrane disruption |
Apparent pKa, tail architecture and branching, position of degradable linkages |
Too low a pKa and encapsulation fails; too high and circulating toxicity and non-specific binding follow |
|
Helper phospholipid |
Bilayer structure formation, fusion support |
Saturation, phase transition temperature, whether the geometry is conical |
Insufficient fusogenicity lowers endosomal escape |
|
Cholesterol |
Membrane fluidity and stability, particle rigidity |
Content, use of analogues |
Too little and the particle leaks; too much and fusion is impeded |
|
PEG-lipid |
Prevents aggregation, extends circulation time |
Mole fraction, chain length, lipid anchor length (which sets the shedding rate) |
Too much impedes uptake and escape; too little causes aggregation and rapid clearance |
Composition is specified as mole fractions, and a widely referenced baseline places ionisable lipid, helper, cholesterol and PEG-lipid at approximately 50 : 10 : 38.5 : 1.5 mol%. That ratio is itself the subject of patent claims, so the design layer reports the proximity of a searched composition to known composition claims — so that the centre of a claim scope is not stepped on for no reason.
7.3.2 The apparent pKa of the ionisable lipid
This is the single most consequential parameter on this route. Two requirements pull in opposite directions: at blood pH (7.4) the lipid must be neutral so that toxicity and non-specific binding stay low, and in the endosome (around pH 5.5) it must be cationic so that it can interact with and destabilise the anionic endosomal membrane.
The band satisfying both is narrow. Computing the protonated fraction at endosomal pH through the Henderson–Hasselbalch relation places the optimum near an apparent pKa of about 6.2 to 6.5. Lower, and the lipid does not become sufficiently cationic even in the endosome, so escape does not occur; higher, and it is already cationic in blood and binds plasma proteins and erythrocytes non-specifically.
One distinction matters. The apparent pKa is not the pKa of an isolated lipid molecule but the value at the surface of the assembled particle. When lipids pack densely at the surface, neighbouring positive charges repel one another and suppress protonation, so the same lipid has a different apparent pKa depending on mole fraction and surface density. The design layer computes an apparent pKa that accounts for this surface charge effect; a calculation using the single-molecule pKa directly misses the effect of composition changes.
7.3.3 The cargo determines the particle — N/P ratio and target size
With the same lipid composition, what is loaded changes the particle. A short duplex oligonucleotide and a transcript of several thousand bases differ in charge density, rigidity and volume, so the amount of cationic lipid required (the N/P ratio — ionisable lipid amines to nucleic acid phosphates) and the size and morphology of the resulting particle differ.
|
Cargo |
Typical length |
N/P ratio |
Target hydrodynamic diameter |
Expected morphology |
|---|---|---|---|---|
|
siRNA · ncRNA-directed · activating duplexes |
21 nt |
1 |
about 25 nm |
core–shell |
|
ASO · SSO single strands |
20 nt |
1 |
about 25 nm |
core–shell |
|
Generic reference |
— |
— |
about 60 nm |
core–shell |
|
Messenger RNA |
about 1,500 nt |
6 |
about 80 nm |
bleb |
|
Circular RNA |
about 1,200 nt |
6 |
about 60 nm |
bleb |
|
Self-amplifying RNA |
about 9,500 nt |
8 |
about 110 nm |
bleb |
|
Circular self-amplifying RNA |
about 11,500 nt |
9 |
about 125 nm |
bleb |
The morphological distinction matters in practice. Short oligonucleotides form a dense core with the ionisable lipid, wrapped by phospholipid and polyethylene glycol — a core–shell structure — whereas long transcripts form a bleb structure containing internal aqueous compartments. The two differ in encapsulation efficiency, release behaviour and stability, and cannot be managed under one quality specification. The design layer assigns an expected morphology per cargo type and confirms at the structural evaluation stage that it is actually obtained.
7.3.4 PEG-lipid — resolving a conflict in time rather than space
The PEG-lipid mole fraction creates a clear conflict. More gives a stable particle and long circulation but impedes cellular uptake and endosomal escape; less improves uptake but the particle aggregates and is cleared rapidly. The value near 1.5 mol% commonly used is the compromise.
A widely used design resolves the conflict in time rather than in space. Shortening the lipid anchor chain makes the polyethylene glycol shed rapidly from the lipid membrane in vivo, so stability and circulation are secured immediately after dosing while uptake occurs from a shed particle after tissue arrival. Anchor length is in effect the control variable for when shedding happens.
Polyethylene glycol carries a separate immunological consideration. Repeat dosing can raise anti-PEG antibodies, producing accelerated blood clearance from the second dose onward, and this is an independent evaluation item in the design layer.
7.3.5 The protein corona — the surface the cell actually sees
A systemically administered particle is coated with plasma proteins almost immediately, and that adsorbed layer becomes the surface the cell recognises. Biodistribution is governed by the composition of the adsorbed proteins rather than by the designed surface composition. When particular apolipoproteins adsorb, uptake proceeds through the corresponding hepatocyte receptor, and this is the principal reason lipid nanoparticles concentrate in the liver without any targeting ligand.
Sending them elsewhere is therefore a problem of changing the corona before it is a problem of attaching a ligand. Attaching a ligand achieves nothing if the corona buries it. The design layer treats corona formation and multivalent binding (avidity) as separate evaluation stages, assessing whether a surface ligand remains accessible beneath the corona.
7.3.6 Selective organ targeting
Selective organ targeting strategies follow from the corona observation. Adding a fifth lipid to the base four-component formulation adjusts the particle's apparent charge, which changes corona composition and shifts biodistribution.
|
Added lipid type |
Representative lipid |
Organ it shifts toward |
Mechanistic reading |
|---|---|---|---|
|
Permanently cationic |
Quaternary ammonium lipid family |
Lung |
Electrostatic capture in pulmonary capillary endothelium and a changed corona |
|
Anionic |
Phosphate-class anionic lipids |
Spleen |
A shift toward opsonin composition recognised by splenic macrophage lineages |
|
No fifth lipid |
The base four components |
Liver |
Apolipoprotein-mediated hepatocyte uptake (the default) |
|
Permanently cationic at high ratio |
Specialised cationic combinations |
Central nervous system access |
Used with the intrathecal route; systemic dosing does not cross the blood–brain barrier |
The design layer treats this redirection with a trained predictive model rather than a rule table. A model validated on a 1,593-row training set including 49 curated selective-organ-targeting formulations predicts organ distribution probability from composition, with an in-house validated balanced accuracy of about 0.717. Predictions are emitted as a probability distribution over liver, spleen, lung, kidney and lymph node, and are used for relative comparison among candidate compositions.
7.3.7 Endosomal escape — the real bottleneck
As the endosome acidifies, the ionisable lipid protonates and the now-cationic lipid forms ion pairs with anionic endosomal membrane phospholipids. Those ion pairs adopt a conical geometry that creates pressure to convert the bilayer (lamellar) organisation into an inverted hexagonal phase, and that phase transition destabilises the membrane so the cargo escapes to the cytosol. The conical geometry of the helper phospholipid assists the same transition.
A proton sponge effect is superimposed. As amine groups absorb incoming protons, chloride ions and water enter to compensate, osmotic pressure rises and the endosome can swell and rupture. Which of the two mechanisms dominates depends on the lipid, and the design layer computes a free energy profile under endosomal conditions as a separate evaluation stage.
7.3.8 What encapsulation changes in the chemistry
Because the particle performs both protection and uptake, the chemical burden on the oligonucleotide falls substantially. Phosphorothioate beyond the termini costs potency for no gain, and full 2'-modification is not mandatory. Conversely, properties affecting behaviour inside the particle — charge density, duplex stability, length — become new considerations. Carrying a chemistry optimised for the carrier-free route unchanged onto an encapsulated route loses potency and takes on class toxicity for nothing.
7.4 Route 3 — carrier-free uptake
A phosphorothioate backbone binds plasma and cell-surface proteins, and that binding pushes the molecule into endocytic routes. The observation that silencing occurs simply on exposing cultured cells to oligonucleotides without transfection reagent is the origin of this route, and it was reported together with the finding that a fairly high phosphorothioate density is required.
The practical value of this route is a shortened development path: preclinical entry without a separate delivery-vehicle programme, and a sharply reduced set of formulation-related chemistry-manufacturing-and-control items. The cost is a large required dose, correspondingly more scope for class toxicity, and no active means of steering tissue distribution.
Design emits carrier-free suitability as a diagnosis rather than a score. A rejected design that leaves only a low number does not say what to change, so the failing gate is returned by name.
- Phosphorothioate fraction below the carrier-free floor (about 75% of linkages)
- 2'-modification does not cover every position (unprotected positions remain)
- A conjugate route declared with no lipophilic or targeting ligand on any terminus — a contradiction of the uptake premise
- A single-strand RISC molecule without a metabolically stable 5'-phosphate mimic
- Phosphorothioate fraction above the protein-binding ceiling (0.90)
7.5 The self-assembling conjugate platform
Attaching a hydrophilic polyethylene glycol at the sense strand's 5' terminus and a hydrophobic lipid anchor at its 3' terminus makes the molecule itself an amphiphile. The monomer self-assembles into a micelle on the order of 100 nm whose shell protects the duplex. Rather than building a separate lipid particle, the molecule becomes the particle.
The practical consequence is a lighter chemical burden: because the shell performs the protection, the internal 2'-modification and phosphorothioate load is lower than for a naked conjugate. When this platform is selected the design layer adds a fifth objective, self-assembly, and switches to conjugate-aware templates. Conjugation is encoded per terminus — guide 5'/3' and passenger 5'/3', each none, lipid, polyethylene glycol or ligand.
Single-strand modalities receive a lipid (3') and polyethylene glycol (5') analogue on the drug strand itself, scored against a lower micelle ceiling. A duplex can use the passenger as a sacrificial strand, whereas a single strand must carry the conjugation on the active strand and therefore risks more activity loss.
Applying the self-assembling platform can put the overhang convention and the chemistry convention in tension. Conjugation is at the 3' terminus and this platform's standard is a native overhang, so the chemical treatment of the overhang bases must be stated in the specification. A 21-mer fully complementary guide with a 19-mer sense strand is the recommended configuration.
7.6 Alternative vehicles
|
Vehicle class |
Principle |
Strengths |
Current constraints |
Registered variants |
|---|---|---|---|---|
|
Extracellular vesicles (exosomes) |
Payload loaded into cell-derived vesicles; surface proteins confer targeting |
Low immunogenicity; tissue tropism engineerable through surface proteins; exploits natural intercellular transfer |
Manufacturing scale and homogeneity, loading efficiency, absence of characterisation standards |
6 |
|
Polymeric carriers |
Polyesters, polyamines and dendrimers condense nucleic acid electrostatically |
Wide compositional freedom, straightforward sustained-release design, degradation rate tunable chemically |
Cytotoxicity of cationic polymers, management of degradation products, polydispersity |
8 |
|
Virus-like particles |
Viral capsid proteins self-assemble around the nucleic acid |
Homogeneous particle size, natural cell-entry mechanism, high loading density |
Immunogenicity and repeat dosing, payload capacity, manufacturing complexity |
6 |
|
Aptamer conjugation |
A folded nucleic acid binds a surface receptor and internalises the payload |
The molecule is its own vehicle; non-immunogenic; synthetically homogeneous |
Limited to targets with a validated aptamer; the endosomal escape bottleneck remains |
see Modality X |
Extracellular vesicles and virus-like particles form by routes different from lipid particles and therefore need their own structural evaluation stage. At present they are registered and assessed at the level of composition and surface properties, with structural simulation a staged extension. The deliverables state that status explicitly — the principle of marking an unevaluated item 'not evaluated' rather than scoring it zero applies here too.
7.7 The formulation variant catalogue — 59 variants
The design layer registers combinations of cargo type and vehicle class as named product variants and evaluates them in parallel during a campaign. Each variant carries a predefined lipid composition, target size, expected morphology, default route of administration and available ligand set, so that comparison among candidates is made on identical terms.
|
Vehicle / cargo class |
Variants |
Character |
|---|---|---|
|
Self-assembling conjugate (specificity and immobilisation families) |
9 |
Lipid and polyethylene glycol amphiphilic micelle |
|
Single-strand oligonucleotide encapsulation |
2 |
Gapmer and steric-block cargo, core–shell |
|
Messenger RNA encapsulation |
5 |
Bleb route, with cap, polyA and modified nucleoside specifications |
|
Self-amplifying RNA encapsulation |
2 |
Large transcript, high N/P, large bleb |
|
Circular RNA encapsulation |
3 |
Internal ribosome entry site driven, no ends |
|
Generic reference formulation |
2 |
Cargo-agnostic control formulation |
|
Selective organ targeting |
8 |
Lung 3 · spleen 3 · central nervous system 2 |
|
Multiplexed cargo |
2 |
Multiple cargoes co-encapsulated |
|
Noncoding-target oligonucleotide |
2 |
ncRNA-directed duplexes |
|
Transcriptional activation oligonucleotide |
2 |
Promoter-targeted duplexes |
|
Steric-block oligonucleotide |
2 |
Uniformly modified single strands |
|
Extracellular vesicles |
6 |
Including surface targeting-protein engineered variants |
|
Polymers |
8 |
Polyester · polyamine · dendrimer · polysaccharide · block copolymer |
|
Virus-like particles |
6 |
Multiple capsid families |
|
Single-strand RISC structural path |
0 by default (2 when enabled) |
3'-only amphiphile, no 5' additions |
The design principle in the last row matters in practice. Registering a new modality must not, as a side effect, push that modality into a running campaign. New families requiring structural search are included only when explicitly enabled; otherwise they receive chemistry optimisation and skip the structural and dynamics stages. Registering and wiring are different acts.
7.8 Route of administration and pharmacokinetics
The route changes absorption rate and tissue distribution together. The design layer holds per-route pharmacokinetic parameters and emits an exposure profile and organ distribution bias per route for a candidate formulation. The values below are literature-based priors used for relative comparison among candidates — they are not absolute predictions.
|
Route |
Time to peak |
Half-life |
Normalised peak |
Organ bias (liver / lung / spleen / CNS) |
|---|---|---|---|---|
|
Intravenous |
0.25 h |
4.5 h |
1.00 |
0.50 / 0.10 / 0.30 / 0.02 |
|
Intramuscular |
6 h |
24 h |
0.40 |
0.20 / 0.05 / 0.10 / 0.01 |
|
Subcutaneous |
12 h |
36 h |
0.30 |
0.15 / 0.05 / 0.08 / 0.01 |
|
Inhalation |
0.5 h |
6 h |
0.80 |
0.05 / 0.75 / 0.05 / 0.01 |
|
Intrathecal |
2 h |
30 h |
0.70 |
0.02 / 0.02 / 0.02 / 0.85 |
|
Intraocular |
4 h |
48 h |
0.50 |
0.02 / 0.02 / 0.02 / 0.01 |
|
Intratumoral |
1 h |
12 h |
0.90 |
0.10 / 0.05 / 0.05 / 0.01 |
What to read from the table is the contrast between routes rather than the absolute values. Intravenous dosing gives the highest peak but a short half-life and strong hepatic and splenic bias; subcutaneous gives a lower peak with gentler, longer exposure. Intrathecal dosing has an overwhelming central nervous system bias of 0.85 with almost no systemic exposure, so class toxicity constraints relax substantially; inhalation achieves local concentration with a lung bias of 0.75 but a short half-life, making residence time the deciding factor.
The route feeds back into the whole formulation. On the inhaled route, shear during aerosolisation can destroy particles, so formulation rigidity matters; on the intrathecal route, cerebrospinal fluid distribution and osmolality and pH compatibility become new constraints; and intratumorally, interstitial pressure and diffusion distance govern efficacy.
7.8.1 Dose translation and in vitro–in vivo correlation
Translating a dose obtained in a preclinical species to humans, and connecting in vitro activity to in vivo activity, exist as separate evaluation stages. What makes that translation awkward for oligonucleotide therapeutics is the large divergence between plasma and tissue exposure: plasma clears within hours while the tissue depot persists for weeks to months, so a plasma-based translation grossly underestimates the real duration of effect. The design layer handles this by treating the tissue depot as a separate compartment, and the output is for candidate comparison rather than a proposed clinical regimen.
7.9 The stage map of delivery design
Delivery design runs as a sequence of stages from chemistry optimisation to formulation and process assessment. The overall skeleton is below, marked for what falls within standard design scope and what is arranged separately.
|
Block |
What it decides |
In standard scope |
|---|---|---|
|
Design front end — cargo verification, chemical-modification optimisation, conjugate structure search, multi-objective Pareto search, joint chemistry × vehicle search |
Sequence and chemistry, and the joint optimum that holds chemistry and vehicle together |
Yes |
|
Lipid and formulation — lipid screening, composition search, quality-by-design space, formulation fixing, vehicle structure assessment, sterol scoring |
Ionisable lipid, helper, cholesterol and PEG-lipid ratios and the permissible design space |
Yes |
|
Physicochemistry and safety — physicochemical and toxicity, stability, innate immunity, anti-PEG, lipid degradants, lipid phase behaviour, metal-catalysed oxidation |
Apparent pKa, particle properties, immunogenicity, degradation routes |
Yes |
|
Structure and dynamics — assembly structure generation, backmapping, molecular dynamics production, convergence, trajectory analysis, endosomal free energy profile |
Particle morphology (core–shell or bleb) and the energetic barrier to endosomal escape |
Requires large simulation assets — arranged separately |
|
Biology and pharmacokinetics — protein corona and avidity, physiologically based pharmacokinetics, dose translation, route absorption, lipid clearance, in vitro–in vivo correlation, PEG design, serum stability |
Biodistribution, dose, route-dependent absorption and correlation |
Yes |
|
Intellectual property, quality and ranking — patent filter, good-manufacturing assessment with roughly 40 quality half-stages, ranking, design-result emission, platemap |
Proximity to composition claims, formulation, process and container risk, final candidate ranking |
Yes (at the level of risk exposure) |
The whole span consists of 82 independent evaluation stages, each of which can be disabled or re-run individually. The purpose of that granularity is twofold: running expensive stages only after the candidate pool has narrowed, and making it possible to trace by name at which stage a candidate was rejected.
7.10 From formulation to product — quality items surfaced at design time
Some quality items are hard to reverse once the formulation is locked. The design layer treats them as independent evaluation stages so that they surface at candidate-selection time. This is not process design but risk exposure; actual process development and validation belong to later stages.
|
Category |
Items |
Why it must be seen at design time |
|---|---|---|
|
Physical properties and stability |
Apparent pKa, particle size and distribution, encapsulation efficiency, colloidal stability (electrostatic versus van der Waals balance), lipid phase behaviour, lipid degradants, metal-catalysed oxidation, serum stability, RNA integrity |
Composition can still be changed here; after the process is fixed, changing it is expensive |
|
Immune and toxicity |
Innate immune response, anti-PEG response, complement activation tendency |
Governs repeat-dosing design and whether premedication is needed |
|
Freezing and drying |
Lyophilisation conditions, reconstitution, repeated freeze–thaw, cryogenic pH shift, crystallisation, long-term storage |
Cold-chain requirements determine commercial viability |
|
Process |
Sterile filtration, tangential-flow filtration, microfluidic scale-up, aseptic fill, process capability |
Whether particle size crosses the filtration limit, and whether size holds on scale-up |
|
Container and transport |
Container-closure suitability, container-closure integrity, headspace oxygen, agitation and shear, transport vibration |
Logistics conditions set the formulation rigidity requirement |
|
Clinical use |
Infusion compatibility, deliverable volume, in-use stability, temperature excursion, photostability |
Handling conditions at the point of care constrain the formulation |
|
Regulatory and documentation |
Potency assay, comparability, impurities including nitrosamines, documentation package |
Components of the filing package are secured at design time |
The list is long because it is genuinely long. A substantial share of failures in oligonucleotide development arises not from sequence or chemistry but from items on this list, and many of them require reversing the composition if discovered after the formulation is locked. The value of surfacing them at design time lies not in the accuracy of the prediction but in moving the moment of discovery earlier.
7.11 Delivery deliverables
- The selected delivery route and an explicit statement of the constraints it imposes on chemistry
- Conjugate design — ligand class and target receptor, attachment terminus, linker type (O/S/C), reference binding constant from the literature and its source
- Formulation candidates — four-component mole fractions, presence of a fifth lipid, expected apparent pKa, N/P ratio, target hydrodynamic diameter, expected morphology
- Organ-targeting prediction — a probability distribution over liver, spleen, lung, kidney and lymph node, with the model basis
- Proximity to known composition claims
- Per-route exposure profile — time to peak, half-life, organ distribution bias, dose translation
- Carrier-free suitability diagnosis where that route is selected — per-gate pass or block verdict with reasons
- Self-assembling conjugate design — lipid, polyethylene glycol and ligand placement per terminus, with rendered notation
- Quality risk item list with the outcome of each assessment, and explicit marking of unevaluated items
- Durability ranking and an expected dosing-interval band
7.12 Limits of this layer
- Structure- and molecular-dynamics-based evaluation stages require large simulation assets and are outside the standard design scope; they are treated as a separate scope where needed
- Trained activity, property and organ-tropism models are predictors and do not replace measurement. The in-house validated balanced accuracy of the organ-tropism prediction is about 0.717, which suits relative comparison among candidates and is not a guarantee of absolute distribution
- Pharmacokinetic parameters are literature-based priors and do not replace measurement on a specific formulation. Tissue depot dynamics in particular vary greatly by tissue and species and rest on limited data
- Protein corona composition varies with plasma composition, disease state and species, so a predicted distribution is not guaranteed to reproduce in a patient population
- Ligand binding constants are literature reference values and are not computed. Unconfirmed entries are left empty and marked
- It does not design the manufacturing process. Quality items surface risk and propose draft specifications; actual process development and validation belong to the manufacturing stage
- Composition claim proximity is informational and is not a freedom-to-operate opinion
8. Modality IV — ncRNA-Directed Silencing (lncRNA · circRNA · snoRNA siRNA/ASO)
8.1 Form and definition
The common name is ncRNA-directed siRNA or ncRNA-directed ASO. Apart from the fact that the target makes no protein, the molecule itself is the same as in Modality I or II. Either the double-stranded RISC mechanism or the single-stranded gapmer mechanism may be used, and the choice follows from the target's class and subcellular localisation. This modality is treated separately not because the molecule differs but because the target's properties differ fundamentally.
There are four differences. Noncoding RNA is markedly less conserved across species than protein-coding genes. Much of it is nuclear-retained, out of reach of cytoplasmic mechanisms. Some targets, such as circular RNAs, share most of their sequence with a linear transcript, so specificity must be achieved differently. And function frequently resides in structure or binding partners rather than sequence, so 'where to cut for the function to disappear' is not self-evident.
8.2 Biology by target class
8.2.1 Long noncoding RNA
Noncoding transcripts longer than 200 nt, functioning as scaffolds for chromatin-remodelling complexes, decoys for transcription factors, or structural elements of nuclear bodies. Many are nuclear-retained, so the RNase H1 mechanism with its nuclear activity is often favoured over the cytoplasmic Argonaute mechanism. An important design fact is that these transcripts have no protein-coding isoforms, so the metric 'coding transcripts cut / total' is structurally 0/0. That is a property of the target rather than an error, and the all-transcript metric is the meaningful one.
8.2.2 Circular RNA
Transcripts whose 5' and 3' ends are covalently joined by back-splicing. Having no ends, they resist exonucleases and are long-lived, and functions as miRNA sponges and as translation templates have been reported. The difficulty as a therapeutic target is specificity: a circular RNA's sequence overlaps almost entirely with the exons of the linear transcript that produced it, so targeting an arbitrary site silences the linear parent gene as well.
The only sequence unique to the circle is the back-splice junction, where the 3' end of a downstream exon abuts the 5' start of an upstream exon. That boundary does not exist in the linear transcript, so only a guide spanning the junction is circle-specific. Design builds a junction-aware alignment index to secure that specificity, and computes accessibility on the circular topology, because a linear-assumption calculation mispredicts structure around the junction.
One interpretive note: a guide spanning the junction does not match the linear transcript annotation, so 'cut on representative transcript' is reported as negative. That is evidence of circle specificity, not a defect.
8.2.3 Small nucleolar and other small noncoding RNAs
Short structured RNAs that guide chemical modification of ribosomal RNA. They have strong secondary structure and exist in complex with proteins, so accessible sites are limited and single-stranded mechanisms are frequently favoured over double-stranded ones. Design classifies this class automatically and routes it to the corresponding mechanism.
8.2.4 Translation-enhancing elements
Antisense transcripts carrying particular repeat sequences have been reported to increase translation of a target mRNA, and engineered constructs based on this are under study. This class is both a silencing target and a design object, so the design layer applies different handling according to the class call.
8.3 Class calling and verification of the premise
The first step in this modality is confirming that the target really is noncoding. Annotation database classifications are not always current or correct, and some transcripts carry short open reading frames and do produce peptides. Design places a coding-potential tool as a gate on that premise and states the call in the report, because proceeding silently on a false premise is the most expensive kind of failure.
Once the class is fixed, the mechanism follows: nuclear-retained targets route to RNase H1, cytoplasmic targets to RISC, and strongly structured small RNAs to a single-stranded mechanism. That this branching is automatic matters in practice — selecting a different tool by hand for each class invites a missed branch, and a missed branch produces output on which no filter was applied.
8.4 Single-species design as a deliberate constraint
This modality designs against one species at a time. That is a deliberate constraint rather than a missing feature. Cross-species conservation of noncoding RNA is far weaker than for protein-coding genes, so demanding cross-species reactivity automatically causes one of two problems: the design space narrows without justification, or sequences that are not functional counterparts are matched as orthologs. The second is the more dangerous, because preclinical species selection based on a false ortholog match produces meaningless toxicology.
Where preclinical species coverage is needed, designing per species independently and comparing the resulting candidates at the sequence level rests on firmer evidence.
8.5 Interfering factors — RNA-binding proteins and chemical modification
Noncoding RNAs frequently function in complex with proteins, so targeting a protein-occupied site impedes binding. Design uses binding peaks derived from crosslinking and immunoprecipitation data as masks to avoid such sites. Likewise, regions around particular chemical modifications may differ in binding and structure, so a modification atlas is used as a mask as well.
These masks operate as weighted penalties rather than absolute exclusions. Binding profiles depend on cell type and condition, so a peak observed in one dataset cannot be assumed present in a customer's experimental system. Applying a penalty and recording the basis in the report is more accurate than absolute exclusion.
8.6 Therapeutic significance
The strategic value of targeting noncoding RNA is that it opens a route to targets considered undruggable. Where a regulatory axis cannot be addressed at the protein level by a small molecule or an antibody but runs through a noncoding RNA, lowering that RNA is the only point of intervention. Chromatin-level regulation, nuclear body formation and transcription-factor decoying in particular are difficult to reach by protein targeting and direct to reach by RNA targeting.
Clinical maturity in this area is lower than in the three preceding modalities. No noncoding-RNA-targeted therapeutic has been approved, and most programmes are preclinical or in early clinical development. That means development risk is high, and equally that much of the target space remains unexplored. Strategically the modality is best understood as low-competition and high-validation-burden.
8.7 Chemistry and delivery requirements
Chemistry requirements are inherited from the chosen mechanism — the Modality I rules for RISC and the Modality II rules for RNase H1. Delivery requirements follow similarly, with subcellular localisation as an additional variable. Nuclear-retained targets are not addressed by cytoplasmic delivery alone, so nuclear entry must be assumed; fortunately single-stranded oligonucleotides are known to distribute well to the nucleus, which makes that combination practically favourable.
8.8 Design deliverables
- Automatic target class call with its basis, and the outcome of the coding-potential gate
- The selected mechanism and the reason for that selection
- A sequence matching the mechanism with a per-position chemistry map
- For circular RNA targets — junction-aware alignment results and the circle-specificity verdict
- An accessibility profile computed on the correct topology
- Results of applying RNA-binding protein occupancy and chemical modification masks
- Off-target table with the effect on the linear parent gene separated out
- Cleavage coverage with the interpretive notes specific to noncoding targets
- Explicit statement that the design is single-species, with the species named
8.9 Advanced capabilities — class routing, circle specificity, interference masks
8.9.1 Class calling and branch routing
The first step fixes the target's class, and the call routes it to an entirely different design branch. The call combines annotation priority with a computational gate: the reference annotation's accession prefix is read first, and where that is ambiguous a coding-potential tool confirms it.
|
Class called |
Branch routed to |
Special handling in that branch |
|---|---|---|
|
Long noncoding RNA |
RISC or RNase H1, depending on nuclear retention |
Zero protein-coding isoforms, so the coverage metric is read differently |
|
Circular RNA |
Back-splice-junction-aware path |
Junction-specific alignment index, circular-topology accessibility, separate reporting of the effect on the linear parent gene |
|
Small nucleolar and other small structured RNAs |
Handed to a single-stranded mechanism |
Accessibility assessed first, given strong secondary structure and protein complexes |
|
Translation-enhancing elements |
Currently unsupported, returned as an explicit error |
Not quietly routed to an adjacent branch — silent wrong answers are prevented |
|
Actually protein-coding |
Terminated as an explicit failure with a pointer to the coding-target path |
That the 'noncoding' premise was wrong is signalled as a failure, not delivered as a result |
The last two rows express this layer's design philosophy. Quietly routing an unsupported class to a neighbouring branch produces output whose meaning nobody can establish. Explicit failure is far cheaper than a wrong success.
8.9.2 Circle specificity — junction design and a validation plan
A circular RNA's sequence overlaps almost entirely with the parent gene's exons, so the only sequence unique to the circle is the back-splice junction. Design accepts only guides spanning that junction as circle-specific and builds a dedicated junction-aware alignment index to verify the specificity.
It does not stop there: a wet-lab validation plan is emitted as well. Demonstrating circle-specific knockdown experimentally requires quantifying circular and linear species separately and reading their ratio, which needs particular primer designs — a divergent primer pair spanning the junction amplifies only the circle, while a convergent pair amplifies only the linear form. The deliverables present both pairs per candidate.
Candidate selection is also constrained to secure several guides sufficiently separated around the junction, because a circle-specific phenotype is only credible when three or more independent guides give the same direction of result.
8.9.3 Alternative mechanism advisory — when to recommend a different tool
Sometimes the circular RNA is short, or the sequence around the junction is unfavourable, and too few viable junction-spanning candidates can be found. Rather than padding the list, the design layer advises an alternative mechanism: a programmable RNA-targeting nuclease approach.
The basis for the advisory is a difference in complementarity thresholds. RISC cleavage requires roughly 17 to 19 nucleotides of contiguous complementarity, whereas that nuclease family has been reported to activate at longer complementarity, on the order of 23 nucleotides. Because of that difference, the window in which a guide cuts the circle without touching the linear parent opens differently for the two mechanisms. Design computes the contiguous complementarity length against the linear parent for each candidate and adjudicates which mechanism affords specificity.
8.9.4 Interference masks — modification and protein occupancy
Noncoding RNAs are complexed with proteins and chemically modified to a greater degree than mRNA, so both factors can badly distort accessibility prediction. Two masks are applied and are active by default.
|
Mask |
Data loaded |
What it avoids |
How it is applied |
|---|---|---|---|
|
RNA modification mask |
172,706 high-confidence modification sites, integrating several modification types (m6A, m5C, m7G, m1A, A-to-I, pseudouridine and others) |
Sites where modification alters base pairing and protein binding |
A weighted penalty, not an absolute exclusion |
|
Protein occupancy mask |
Aggregated crosslinking-immunoprecipitation peaks across many cell lines |
Sites physically occupied by protein and therefore inaccessible |
A weighted penalty, because occupancy is cell-type dependent |
|
Editing site handling |
An A-to-I editing site database |
An edited transcript has a different sequence, so the guide no longer matches |
A surrogate edited sequence (A→G substituted) is generated and complementarity is assessed against both versions |
The third row is specific to this class. Targeting a frequently edited site means a guide that is fully complementary against the reference carries a mismatch in the real cell. Design generates the edited surrogate and prefers candidates valid against both versions.
8.9.5 Accessibility on a circular topology
Accessibility computed on a linear assumption mispredicts structure around the back-splice junction, because a linear calculation treats the two ends as free termini while a circle has none, and sequences that were originally far apart become neighbours across the junction. Design computes topology-aware accessibility for circular targets; without that handling, candidates near the junction are systematically misassessed.
8.9.6 Single-species design and per-species references
This layer holds independent references for six species — human, mouse, rat, macaque, dog and pig — and designs against one species per run. Automatic cross-species extension is not offered for the reason given in 8.4; where preclinical species coverage is needed, designing per species and comparing at the sequence level rests on firmer evidence.
8.10 What this modality cannot solve
- Cross-species extension — conservation is weak, so automatic cross-reactive design is not offered
- Partial inactivation of a structural function — cleavage is close to all-or-nothing and is not suited to selectively disabling one domain
- Targets whose functional region is unmapped — without knowing where to cut, candidate priorities cannot be set without experiment
- Compensation at the protein level — removing the RNA has no phenotype if another pathway substitutes for its function
9. Modality V — saRNA / RNAa (Promoter-Targeted Transcriptional Activation)
9.1 Form and definition
The common name is saRNA (small activating RNA), and the operating principle is called RNAa (RNA activation). The form is the same double-stranded small RNA as Modality I, but the target is a gene's promoter region rather than its mRNA, and the outcome is activation rather than repression. The same material in the same form acts in opposite directions depending on where it is pointed. That fact produces this modality's most important design implication: which coordinate is targeted dominates the outcome more than the sequence design itself.
9.2 Molecular mechanism
9.2.1 An Argonaute-mediated event in the nucleus
That a small double-stranded RNA targeting a promoter can increase gene expression has been reported reproducibly across multiple genes. The leading mechanistic account is that an Argonaute protein moves into the nucleus, interacts with low-abundance noncoding transcripts produced at the promoter, and induces activation-associated chromatin marks at that locus. The target is therefore not genomic DNA itself but the RNA transcribed there, using the same recognition principle as silencing with the opposite outcome.
An important difference is that cleavage is not required. In silencing the catalytic activity of Argonaute is central, whereas in activation the presence of the complex at the locus is what produces the effect. Cleavage-site protection rules therefore do not apply in chemistry optimisation, and applying a silencing chemistry template unchanged imposes an unnecessary constraint. Which member of the Argonaute family is routed through may also differ from silencing, so the design includes a loading-route audit.
9.2.2 Chromatin state is a precondition
For activation to occur, the locus must be in a state capable of being activated. A fully condensed, inaccessible promoter gives the complex nowhere to settle, and an already maximally active promoter has no headroom. The modality therefore works best at loci in an intermediate state.
Design accordingly assesses the chromatin context of the target locus, consulting active-promoter marks, chromatin accessibility and active-enhancer marks to present evidence on whether the locus presents favourable conditions for an activation attempt. Public chromatin data, however, are measured in particular cell lines and tissues and may not represent a customer's experimental system. That limitation is stated in the report, and the output is best read as an evidence summary rather than a prediction of magnitude.
9.2.3 Natural antisense transcripts as a second axis
At many loci a transcript is produced in the opposite orientation alongside the sense transcript, and in some cases the antisense transcript represses the sense one. Where such a repressive axis exists, lowering the antisense transcript is itself a way of raising the target gene — achieving activation by means of silencing. Design maps antisense transcripts around the target locus and presents that possibility.
The practical advantage of this route is predictability. Silencing mechanisms have higher prediction accuracy and far deeper clinical precedent than activation, so achieving the same therapeutic objective by silencing lowers development risk. One of the first things the design layer checks in response to an activation request is therefore whether a repressive axis exists that can be silenced instead.
9.3 Therapeutic significance — why activation is needed
There are disease classes that silencing cannot address in principle: loss or silencing of a tumour suppressor, haploinsufficiency, enzyme deficiency in a metabolic pathway, and shortage of a protective factor. In these cases the therapeutic direction is upward, and the existing options were protein replacement, gene therapy, or a small molecule that raises expression. Each has constraints: replacement requires repeat dosing and carries immunogenicity risk, gene therapy carries vector-associated risk and difficulty in controlling expression, and small molecules struggle with target selectivity.
Promoter-targeted activation aims at that gap. Because the endogenous gene is raised within its own regulatory context, the increase is expected to stay near physiological range; sequence-based specificity provides target selectivity; and being a synthetic oligonucleotide, it uses the existing manufacturing and administration infrastructure of the field.
Clinical maturity is lower than for silencing. Clinical trials have been conducted in indications including hepatocellular carcinoma with proof-of-concept results reported, but there are no approvals. The patent landscape centres on target loci and sequences and on activation-specific chemistry profiles, and the design layer offers several chemistry profile families as options so that alternatives are available for freedom-to-operate review.
9.4 Chemistry requirement — mandatory, minus the cleavage rule
Stability and immune-suppression requirements match the silencing modalities. Two things differ. Cleavage is not required, so cleavage-site protection does not apply and chemical placement has correspondingly more freedom. And nuclear localisation matters, so bulky additions that might impede it warrant caution.
The design layer holds several activation-specific chemistry profiles and presents them for comparison — an enhanced-stabilisation family, placement families derived from the activation literature, and self-assembling conjugate families — each balancing stability, potency and manufacturing complexity differently.
9.5 Delivery requirement — required
This modality must reach the nucleus, so cytoplasmic delivery is insufficient, and being double-stranded it cannot use carrier-free uptake. A vehicle is therefore mandatory, and the options are those of Modality I: receptor conjugation for hepatic targets, particle encapsulation or an alternative vehicle otherwise.
The additional requirement of nuclear entry affects vehicle choice. Unlike a silencing molecule acting in the cytoplasm, an activating molecule must cross the nuclear envelope after endosomal escape. The efficiency of that step is hard to quantify and differs by vehicle, so in activation programmes it is advisable to obtain measured nuclear distribution data when narrowing vehicle candidates.
9.6 Design deliverables
- Transcription start site resolution — which of automatic resolution, an upstream/downstream window, or explicit coordinates was used, and why
- Double-stranded candidate sequences aligned to the target window with a per-position chemistry map
- Chromatin context — active-promoter marks, accessibility and active-enhancer marks, with an explicit statement of their limitations
- Mapping of natural antisense and adjacent noncoding transcripts — the possibility of a silencing detour route
- Activation prediction score from the transfer-learning prediction head, with its uncertainty
- Argonaute loading-route audit — which family member the design is predicted to route through
- Comparison of several chemistry profiles, each with its balance of stability, potency and manufacturing complexity
- A positive control set — literature-validated activating sequences with which the experimental system itself can be validated
- Off-target table with hits against promoter-like sequences separated out
Positive controls are included in the deliverables because of this modality's predictive uncertainty. When an experiment shows no effect, it must be possible to distinguish a design problem from an experimental-system problem, and a literature-validated control makes that distinction possible.
9.7 Advanced capabilities — promoter resolution and activation prediction
9.7.1 The transcription start site is not a single point
The first difficulty in activation design is that 'where is the transcription start site' has no single answer. Most genes have multiple transcripts with different start points, and annotation databases may designate different representative starts. A shift of 100 nucleotides moves the whole target window, so this resolution governs design quality.
Design offers three resolution modes and states in the deliverables which was used: automatic derivation by merging the 5' ends of multiple annotated transcripts, a user-specified upstream/downstream window, and direct genomic coordinates.
9.7.2 Promoter confidence — a weighted sum over five elements
In automatic mode the confidence of a candidate start site is a weighted sum over five elements. Each is a different class of evidence, so one being weak can be offset by others; all five being weak means the locus is unfavourable for an activation attempt.
|
Element |
Weight |
What it reads |
Interpretation |
|---|---|---|---|
|
Annotation-based 5' ends |
35% |
Merged 5' ends of multiple annotated transcripts |
The largest weight. Confidence is high when several transcripts support the same point |
|
Core promoter motifs |
15% |
TATA box (about −30 ± 5), initiator element (−3 to +5) |
Presence of classical core elements. Absence does not mean it is not a promoter |
|
CpG island |
15% |
Sliding-window CpG island call (observed/expected ratio and GC content) |
The majority of mammalian promoters overlap a CpG island |
|
Chromatin state |
30% |
Active promoter marks and chromatin accessibility peak summits |
The second largest weight — direct evidence of whether the locus can be activated |
|
Cross-species conservation |
5% |
Multi-species alignment conservation score (optional) |
Functional elements tend to be conserved, but exceptions are common, hence the low weight |
Chromatin carries 30% because it is the enabling condition for this modality: at a fully closed locus, no strength in the other four elements lets the complex settle. Public chromatin data are however measured in particular cell lines and tissues, so this element is 'the state in that cell' and may not represent a customer's system — a limitation stated alongside the value.
9.7.3 Predicted regulatory tracks — filling gaps where data are absent
Many tissues and cell types have no public chromatin data. To fill that gap the design layer can optionally use a deep model that predicts regulatory tracks from sequence alone. The model predicts transcription, initiation, accessibility and binding signals at tens-of-bases resolution over input windows hundreds of kilobases wide, and the initiation output is used for start site estimation and promoter activity estimation.
Predicted tracks do not replace measurement, and where measurement exists it takes precedence. The deliverables state at which loci measurement was used and at which prediction was used.
9.7.4 Argonaute loading-route audit
Which Argonaute family member activation routes through may differ from silencing, and this depends in part on the 5'-terminal nucleotide. Design emits a loading-route prediction based on guide 5'-end identity as an audit item.
One common overstatement is corrected explicitly here. It is conventional to mark a 5'-adenosine as 'compatible with both Ago2 and Ago1', but the reported distribution is spread fairly evenly across the four family members. A 5'-adenosine is therefore not suited to a particular route but at risk of being distributed across several, and the deliverables label it a mixing risk. Marking it as compatible would promote candidates that in fact disperse.
9.7.5 Mapping the repressive axis — the activation detour
Transcripts in the opposite orientation around the target locus are mapped and assessed for whether they act as a repressive axis. Where such an axis exists, silencing it is a way of raising the target, and because it uses a mature silencing modality the development risk falls substantially.
Mapping uses opposite-strand transcripts from the reference annotation together with a separate noncoding transcript annotation. The deliverables present, for each candidate repressive axis, its positional relationship to the target, expression correlation, and whether it can be converted into a silencing design.
9.7.6 The positive control set — 13 duplexes
Because predictive uncertainty is high in this modality, a means of distinguishing a design problem from an experimental-system problem is needed when an experiment shows no effect. The design layer bundles 13 control duplexes for which activation has been reported in the literature, each carrying its target gene and its coordinate relative to the transcription start site.
The control targets include cell-cycle inhibitors, an adhesion factor, a growth factor, transcription factor families, tumour suppressor families and a nuclear receptor, giving enough diversity that at least one is expressed in a customer's system. A control that fails to work is the signal to check delivery or the cell system before the design.
9.7.7 Activation-specific chemistry profiles
Because cleavage is not required, chemical placement has more freedom, and nuclear localisation and loading efficiency matter instead. The design layer presents several chemistry profiles built on different placement philosophies, each balancing stability, potency and manufacturing complexity differently. Applying a silencing profile unchanged imposes a cleavage-site protection rule that constrains the design for no reason.
Positional immune-avoidance rules are applied as well: modifications are placed at the balance point between immune suppression and loading efficiency, on the basis of reports that 2'-modification at particular positions suppresses innate immune recognition.
9.8 What this modality cannot solve
- A completely deleted gene — with no gene there is nothing to activate; that is the domain of gene replacement
- A locus already at maximum expression — there is no headroom for an effect
- Fully condensed, closed chromatin — the complex has nowhere to settle
- Fine quantitative control — specifying a numerical target level and raising expression to it is beyond current capability
- Cell-type-specific prediction — chromatin state differs by cell, limiting the transferability of predictions based on public data
10. Modality VI — miRNA Mimic and anti-miR / Antagomir (miRNA Axis Modulation)
10.1 Two opposite molecules
MicroRNAs are endogenous small RNAs of about 22 nt that load into Argonaute and repress tens to hundreds of target mRNAs simultaneously. Repression per individual target is modest, but the number of targets produces effects at network scale. When the abundance of a particular miRNA is abnormal in disease, there are two directions of intervention, and they yield entirely different molecules.
|
Mimic |
Antagonist (anti-miR) |
|
|---|---|---|
|
Role |
Agonist — replaces the function of a lost miRNA |
Antagonist — sequesters an overactive miRNA |
|
Form |
Double-stranded; the guide is the mature miRNA sequence |
Single-stranded; the reverse complement of the mature miRNA |
|
Mechanism |
Loads into RISC and represses that miRNA's natural target set |
Binds the endogenous miRNA and sequesters it sterically |
|
Nuclease involvement |
Argonaute-mediated repression (largely translational repression and destabilisation, without cleavage) |
None — RNase H must not be recruited |
|
Design freedom |
Sequence fixed in advance; freedom exists only in chemistry and passenger |
Freedom exists in length and chemistry |
|
Principal off-target |
Seed-mediated repression by the passenger strand |
Cross-inhibition of other miRNAs sharing the seed |
|
Chemical modification |
Required |
Required — uniform across every position, with a fully modified backbone |
|
Delivery vehicle |
Required (double-stranded) |
Optional — carrier-free uptake works |
10.2 The biology and design logic of a mimic
10.2.1 The deliverable is the target list, not the sequence
In mimic design the sequence is already determined: the mature miRNA entry is the guide strand. The question that actually has to be answered is therefore not 'what should be made' but 'what gets repressed if this is introduced'. The answer — the predicted target list — is this modality's central deliverable and simultaneously the basis of its safety assessment.
Target prediction by any single method is dominated by false positives. Selecting candidates by seed complementarity alone makes a large fraction of the transcriptome a candidate; selecting by thermodynamics alone fails to reflect accessibility in vivo. Design therefore combines evidence from five independent routes with weighting.
- Context-based standard prediction — a score combining seed type, 3' supplementary pairing and local context features of the target site
- Machine-learning binding prediction — a model learning binding directly from sequence features
- A seed classifier — strict classification of seed type (8mer, 7mer-m8, 7mer-A1, 6mer)
- Strict-seed alignment — rule-based search for reproducibility
- Hybridisation free energy and target-site accessibility — thermodynamic feasibility
The predicted list is then refined against experimental evidence and expression context. Argonaute crosslinking-immunoprecipitation and degradome data check measured binding, conservation in the mouse ortholog adds evolutionary support, and tissue expression data filter for whether a target is actually expressed in the tissue of interest. Consensus voting raises precision at some cost in recall, and that balance is the correct direction for therapeutic design: a wrongly included target is a more expensive error than a missed one.
10.2.2 Seed heterogeneity as a trap
A mature miRNA is not a single sequence. Processing generates variants whose 5' end is shifted by one nucleotide, and a single-nucleotide shift at the 5' end moves the entire seed region. A shifted seed means a shifted target set, so for miRNAs with substantial 5' heterogeneity, predicting targets from the single dominant mature sequence does not represent the real biology. Design flags this heterogeneity and reports the major variants.
10.2.3 Safety axes
|
Axis |
What is at stake |
How it is assessed |
|---|---|---|
|
Strand-loading asymmetry |
If the passenger loads, an entirely different target set is repressed |
Terminal free-energy differential; quantification of the passenger seed's off-target burden |
|
Seed toxicity |
Particular seeds repress essential genes broadly |
Essential-gene and genome-wide seed burden indices |
|
Innate immunity |
GU-rich and CpG motifs stimulate receptors |
Motif scanning and 2'-O-methyl placement |
|
Population polymorphism |
Common variants at seed-binding sites create inter-individual differences in efficacy |
Allele frequency scan across the seed site |
|
Pathway convergence |
If repressed targets crowd into one pathway, unexpected phenotypes follow |
Functional annotation convergence analysis of the target set |
10.3 The biology and design logic of an antagonist
10.3.1 Sequestration is stoichiometric
An anti-miR does not cleave its target; it binds and holds it. It therefore has no turnover and acts stoichiometrically, requiring an amount commensurate with the number of target miRNA molecules in the cell. That produces a dose logic fundamentally different from the catalytic silencing modalities — although if binding is very strong it becomes effectively irreversible, extending duration.
Because the means of raising binding strength is chemistry, chemistry is efficacy itself in this modality. Uniform high-affinity modification with a fully phosphorothioate backbone is standard, and that same combination enables carrier-free uptake.
10.3.2 Length strategy — the whole family or just one member
|
Strategy |
Structure |
Target scope |
When to choose it |
|---|---|---|---|
|
Full-length steric block |
Full reverse complement with uniform high-affinity modification and a fully phosphorothioate backbone |
High selectivity for that one miRNA |
When one specific miRNA must be inhibited precisely |
|
Short seed-directed oligonucleotide |
A high-affinity oligonucleotide of about 8 nt complementary to the seed |
The whole family sharing that seed |
When family members are functionally redundant and inhibiting one is compensated |
|
Mixmer |
Alternating high-affinity and natural residues |
Intermediate — a compromise between selectivity and affinity |
When following a structure with clinical precedent |
|
Receptor-conjugated form |
Any of the above plus a hepatocyte targeting ligand |
Maximised local concentration in liver |
Hepatic targets where systemic exposure should be reduced |
The logic of the short seed-directed strategy comes from miRNA biology itself. MiRNAs sharing a seed have largely overlapping target sets and are therefore functionally redundant; inhibiting one is compensated by the others. Suppressing the whole family is often required for a phenotype, and a short oligonucleotide targeting only the seed fits that purpose. Where family members differ in expressing tissue and have diverged functionally, full-length targeting is appropriate.
10.3.3 The character of off-targets
Anti-miR off-targets have two layers. The first is co-sequestration of miRNAs other than the intended one, predicted from seed-sharing relationships. The second is incidental complementarity to mRNAs; because uniform high-affinity chemistry does not recruit RNase H, no cleavage occurs, but translational interference is possible. Design reports the two layers separately, and where paralog instructions are supplied, applies adjudication to mRNA complementarity hits.
10.4 Therapeutic significance
The strategic character of targeting a miRNA axis is that one molecule moves a network. Where single-gene silencing removes one node, miRNA modulation moves the regulatory layer over a set of nodes. Where a disease is not the failure of a single gene but a failure of regulation — fibrosis, metabolic reprogramming, immune microenvironment — this approach fits conceptually.
That is also the risk. Moving many targets at once leaves broad scope for unanticipated phenotypes, and the safety assessment burden exceeds that of single-target silencing. This is the principal reason the modality is less clinically mature than the silencing modalities. A hepatic anti-miR showing antiviral effect in clinical trials is widely cited as proof of concept, but there are no approved products.
The recommended development approach is to establish the convergence of the target set first. If predicted targets converge on a single pathway consistent with the therapeutic hypothesis, the network effect is likely to act in the intended direction; if they are scattered across unrelated pathways, the risk of unpredictable effects is high. This is why functional convergence analysis of the target set is part of the deliverables.
10.5 Design deliverables
- The chosen direction (mimicry or antagonism) and the resulting molecular form
- Mimic — guide and passenger sequences with a per-position chemistry map, and a strand-loading asymmetry assessment
- Antagonist — length strategy, uniform chemistry notation, and conjugation status
- Predicted target list — evidence from five routes with a consensus score and the basis shown per target
- Experimental evidence cross-check — agreement with crosslinking-immunoprecipitation and degradome data
- Results of conservation and tissue-expression context filtering
- Safety axes — seed toxicity indices, innate-immune motifs, population polymorphism, pathway convergence
- 5' heterogeneity awareness — major variants and their seed shift
- For antagonists — the list of seed-sharing miRNAs and predicted cross-inhibition
- For antagonists — mRNA complementarity hits and any paralog adjudication results
10.6 Advanced capabilities — weighted evidence combination and safety axes
10.6.1 The five evidence weights — the actual coefficients
The central deliverable of mimic design is the target list, produced as a weighted sum over five lines of evidence. The weights are documented deterministic constants, so the same input reproduces the same list.
|
Evidence |
Weight |
What it measures |
Why this weight |
|---|---|---|---|
|
Context-based efficacy score |
0.34 |
A real-valued score combining seed type, 3' supplementary pairing and site context features |
The largest weight — the only real-valued evidence tied directly to efficacy |
|
Machine-learning binding probability |
0.24 |
Per-site probability from a model learning binding directly from sequence features |
Complements patterns the rules do not capture |
|
Target site accessibility |
0.16 |
The probability that the site is open in the target 3' untranslated region |
A physical prerequisite for binding — necessary but not sufficient, hence a middle weight |
|
Cross-species conservation |
0.14 |
Conservation of the site in the mouse ortholog |
Evolutionary support; sites can function without being conserved, hence a low weight |
|
Rule-based alignment consensus |
0.12 |
Agreement among strict-seed search tools |
A baseline for reproducibility; highly correlated with the other evidence, hence a low weight |
The weights are published because they make the output interpretable. Whether a target ranks highly because of its context score or its accessibility changes how it should be validated, and the deliverables present the individual values of all five for each target.
10.6.2 Cross-checking against measured binding evidence
The predicted list is checked against experimentally observed binding data. Argonaute crosslinking immunoprecipitation shows where the complex actually bound, and degradome data show traces of actual cleavage. Both are independent of the prediction score, so targets that are both highly predicted and experimentally supported form the most trustworthy subset.
Measured data are queried on demand and cached locally; where a query fails, that axis is marked 'not queried' rather than scored zero, so that absence of experimental evidence is not read as absence of binding — in tissues and conditions where the experiment was never performed, absent data are normal.
10.6.3 Seed toxicity — quantifying the essential-gene burden
A miRNA mimic represses hundreds of transcripts through one seed, so if that seed broadly represses essential genes, cytotoxicity follows. Design computes a per-seed off-target frequency index to quantify that burden, presented at two levels — a genome-wide burden and an essential-gene-restricted burden, the latter being more directly connected to phenotypic risk.
In a mimic the sequence is fixed, so the seed cannot be changed. The index therefore serves risk awareness rather than candidate selection: a high-burden miRNA signals that dose design and safety assessment should be more conservative. The passenger strand's seed, by contrast, can be changed by design, so passenger seed burden is a genuine selection axis.
10.6.4 5' heterogeneity — when the seed moves wholesale
A mature miRNA is not a single sequence. Processing generates variants shifted by one nucleotide at the 5' end, and a one-nucleotide shift moves the entire seed region (positions 2–8). A shifted seed means a shifted target set, so predicting targets from the single dominant mature sequence stops representing the real biology.
Design computes the seeds of the canonical form and of the +1 and −1 shifted variants and reports the target set of each separately. For a miRNA with substantial heterogeneity, the union of those three lists is closer to the real scope of effect.
10.6.5 Population polymorphism at target sites
Where a common polymorphism exists at a seed binding site, repression does not occur in individuals carrying that variant. Design queries population allele frequencies at the binding sites of predicted targets and flags sites where polymorphism is common. This is not a reason to change the mimic, but it identifies in advance where response heterogeneity would appear clinically.
10.6.6 Functional convergence of the target set
Whether predicted targets converge on one pathway or scatter across unrelated ones is analysed. Convergence makes it likely the network effect acts in the intended direction; scattering carries a high risk of unpredictable effects. The analysis uses the same function-axis resources as Chapter 11, with the same pathway size cap and per-source separation rules.
The practical use is modality suitability. A miRNA for which convergence cannot be established is risky to address with this modality, and silencing the individual targets of that axis directly becomes the more predictable approach.
10.6.7 Choosing among the four anti-miR strategies
In the antagonist direction there is design freedom in the sequence, so strategy selection is a real decision. The four strategies of 10.3.2 separate on the following criteria.
- Does the target miRNA have a function unique within its family — if so, full-length targeting secures selectivity
- Are family members functionally redundant so that inhibiting one is compensated — if so, a short seed-directed oligonucleotide
- Is there a constraint to follow clinical precedent — if so, a mixmer architecture
- Is the target tissue the liver — if so, conjugate a receptor ligand to any of the above to lower the dose
Design derives the family member list from seed-sharing relationships and presents each member's tissue expression, providing the basis for judging whether the whole family may be suppressed. If one member carries a protective function in the target tissue, family-wide suppression is inappropriate.
10.7 What this modality cannot solve
- Discovery of novel miRNAs — design anchors on known mature entries
- Sequence optimisation of a mimic — the sequence is fixed in advance, so efficacy cannot be improved by sequence; only chemistry and delivery remain
- Guaranteeing the predicted targets — the list is a prioritisation, and experimental confirmation is required
- Fine control of the network effect — which targets are repressed and by how much cannot be selectively tuned. This is the modality's strength and its limit
- Complete family selectivity — distinguishing seed-sharing miRNAs is difficult in principle, and in the short-oligonucleotide strategy it is not even the goal
11. Target Adjudication and Family Selectivity — Advanced Pathway and Network Analysis
Where the preceding chapters addressed what molecule to build, this chapter addresses which target, through which modality, and how far across its family to intervene. The deliverable is a verdict with evidence rather than a sequence, which is why it is described separately. In execution order it precedes design, and programmes that skip it fail not on design quality but on target definition: silence an essential gene with a well-designed molecule and the molecule is excellent while the experiment fails.
What fundamentally separates this layer from the other design stages is that its input is a graph rather than a sequence. Target suitability is not settled by examining one gene. What it interacts with, which pathways it belongs to, whether that pathway holds redundant members, what moves downstream when it is inhibited, and — if it cannot be touched at all — where else the same outcome can be produced: these are all network questions. This chapter describes how that network is constituted and what verdicts it produces.
11.1 The five questions this layer answers
|
# |
Question |
Which axis answers it |
Consequence of no answer |
|---|---|---|---|
|
Q1 |
May this gene be touched at all |
Safety axis — population loss constraint, cell dependency, adverse-event precedent |
An essential gene is silenced and the target cells die |
|
Q2 |
In which direction should it be moved |
Disease axis plus function axis plus the sign of the causal network |
The wrong direction can worsen the disease |
|
Q3 |
Which modality suits this gene |
Variant spectrum, subcellular localisation, surface exposure of the protein |
Resource is committed to a modality that cannot apply |
|
Q4 |
If it cannot be touched, what should be touched instead |
Two-step neighbourhood, signed causal layer, pathway ladder |
The programme stops at 'this target does not work' |
|
Q5 |
How far across the family should the intervention reach |
Paralog segregation — functional sharing, disease association, redundancy |
Compensation suppresses the phenotype, or a member that had to be spared is lost |
Q4 shows most clearly why this layer exists. A target assessment that ends in 'unsuitable' stops a programme without advancing it. By contrast, a conclusion that 'this gene is pan-essential and cannot be silenced directly, but two downstream effectors connected to it by inhibitory edges are tissue-restricted and weakly constrained' creates the next step. Reaching that conclusion requires a directional network.
11.2 The evidence graphs — what is loaded
Verdicts rest on two large graphs: a protein interaction graph and a knowledge graph holding pathways, functions, diseases and phenotypes. The two are normalised to share one gene identifier space and therefore join directly — a sampled verification found zero mismatches. If the identifier spaces diverge, using the two graphs together produces silent empty results, so this normalisation is the precondition for any advanced analysis.
11.2.1 The protein interaction graph
|
Component |
Scale |
What it is used for |
|---|---|---|
|
Integrated sources |
13 |
Removes single-source bias; degree is also reported per source |
|
Consensus interaction pairs |
about 1.73 million |
Interactions supported by multiple sources — the base set for first-shell partner search |
|
Pairs from the large integrated source |
about 6.70 million across 13 evidence channels |
Evidence separated by channel (experimental, co-expression, literature, genetic interaction and so on). Channels are reported separately, not summed |
|
Protein complexes |
22,824 |
Complex members act as one functional unit and are considered together in target selection |
|
Signed and directed causal edges |
From curated signalling sources |
Activation or inhibition sign with upstream and downstream direction — the core of substitute target search |
Evidence channels are reported separately because they differ in reliability and in meaning. Summing an edge supported by a physical binding experiment together with one inferred from literature co-mention into a single score lets the latter, being far more numerous, dilute the former. When searching for substitute targets in particular, only edges with physical or causal support form a meaningful path.
11.2.2 The knowledge graph
|
Layer |
Scale |
Role |
|---|---|---|
|
Pathways |
4,228 |
Loaded separately per source (Reactome, KEGG, WikiPathways and others), with pathway size retained |
|
Functional ontology |
48,329 terms / 858,529 annotations |
Separated into the three aspects: molecular function, biological process, cellular component |
|
Gene–disease associations |
about 28.09 million |
Curated and text-mined sources distinguished, with a separate normalised score column |
|
Paralog pairs |
140,134 |
Held with sequence identity — the input to family selectivity adjudication |
|
Population loss constraint |
18,140 genes (including 814 on sex chromosomes) |
pLI, LOEUF and LOEUF percentile. Holding the percentile alongside is the precondition for threshold setting |
|
Tissue expression |
Normalised per-tissue expression |
Expression breadth, tissue-specificity index, top tissues |
|
Cell dependency |
1,208 cell lines |
Gene effect values and the fraction of dependent cell lines |
|
Clinical variant profiles |
Decomposed by variant type |
The composition of pathogenic variants — a direct input to modality suitability |
|
Tractability assessment |
Bucket classification plus adverse-event precedent |
So that targets with existing failure precedent are not chosen again |
Scale is not the same as quality. Most of the 28.09 million gene–disease associations are text-mined, and curated associations are a small fraction of them. This is exactly why the layer reports the two classes separately and keeps a separate normalised column — see 11.6.
11.3 The safety axis — may this gene be touched
Safety adjudication reads four independent lines of evidence together, because any one of them alone misses a substantial fraction of risky genes.
|
Evidence |
What it tells you |
Blind spot when used alone |
|---|---|---|
|
Population loss constraint (pLI, LOEUF, percentile) |
How strongly loss-of-function variation in this gene has been suppressed in human populations — a proxy for evolutionary essentiality |
A single threshold misses many risky genes. Short genes appear weakly constrained for lack of statistical power, and sex chromosomes have coverage limits |
|
Cell dependency (gene effect, dependent cell line fraction) |
Whether removing the gene experimentally kills cells, and in what fraction of cell lines |
Measured in cultured cells, so tissue-restricted essentiality and adult-tissue requirement are not reflected |
|
Tractability bucket and adverse-event precedent |
Whether development against this target has been attempted, and what went wrong |
A novel target has no information — absence does not mean safety |
|
Validated synthetic-lethal partners |
The partner whose co-loss is lethal — the basis for tumour-selectivity strategies |
Context-dependent and may not reproduce across cell types |
Loss constraint is held with its percentile rather than only its raw value because of thresholding. The absolute value is affected by gene length and the number of observation opportunities, so a quantile basis is more stable for comparison across genes. The verdict presents raw value and percentile together so that a reader can reinterpret against their own criterion.
Notation of absent data matters most on this axis. Where constraint information is missing, the reason is stated — sex-chromosome coverage limits, a noncoding locus, insufficient statistical power — and a blank is never rendered as 'unconstrained'. A coverage limitation in the reference data turning directly into a false safety verdict is the most common and most expensive failure mode in this area.
11.4 The expression axis — where and how much is made
Expression information feeds two judgements: the choice of delivery target tissue, and the prediction of which tissues are affected under systemic exposure.
- Expression breadth — the number of tissues above a threshold. The broader it is, the wider the impact of systemic exposure
- Tissue-specificity index — how concentrated expression is in one tissue, summarised between 0 and 1. A high value means targeting one tissue suffices, simplifying delivery strategy and lowering systemic burden
- Top expressing tissues — the tissues with the highest absolute expression. A direct input to vehicle selection and organ-tropic formulation
This axis changes verdicts when combined with the safety axis. Even a gene with high cell dependency may be addressable if its expression is confined to one tissue and that tissue is the target, because local administration can avoid systemic essentiality. Conversely, a broadly expressed essential gene cannot be avoided by any delivery strategy.
11.5 The function axis — pathway analysis
11.5.1 Why pathways are read per source
Pathway databases embody different curation philosophies. Some subdivide at the level of individual reactions; others group by disease or biological theme. It is common for the same gene to belong to twelve pathways in one source and three in another. Summing them lets the finer-grained source dominate, so the verdict presents pathways separately per source and reports the size of each.
Pathway size matters because it bears directly on functional sharing verdicts. That two genes share a top-level pathway such as 'cellular metabolism', which contains thousands of genes, is not evidence that they are functionally related. The verdict applies a size cap that excludes overly large pathways from counting as evidence of functional sharing, defaulting to 300 genes.
11.5.2 The three aspects of the functional ontology
|
Aspect |
What it describes |
Use in adjudication |
Caution |
|---|---|---|---|
|
Molecular function |
What the protein does at the molecular level (enzyme activity, binding activity) |
The primary input to family selectivity — does it do the same job |
Sharing a molecular function term does not imply the same substrate. This is the main source of over-broad co-silencing verdicts |
|
Biological process |
Which process it participates in (cell cycle, immune response) |
Consistency with the therapeutic hypothesis; target-set convergence analysis |
Sharing at the process level does not imply functional redundancy |
|
Cellular component |
Where it resides (membrane, nucleus, secreted) |
A direct input to modality suitability — surface exposure is the precondition for aptamer conjugates (Modality X) |
Whether an annotation is experimental or predicted needs to be distinguished |
That the cellular component aspect connects directly to modality adjudication matters in practice. An aptamer conjugate binds surface proteins from outside the cell and cannot in principle reach cytosolic or nuclear proteins. A request for an aptamer against a target with no surface or secreted annotation does not receive 'not applicable' but 'redirected to an interaction partner carrying a surface annotation', and that redirection uses the network axis.
11.5.3 The pathway ladder
In a hierarchically organised pathway database, looking only at the most specific pathway a gene belongs to loses context. The verdict presents that pathway's parent and grandparent as a ladder. This serves two purposes: grasping the target's biological context at a glance, and defining the search scope for substitute targets as 'other specific pathways under the same parent'.
11.6 The disease axis — association and the scale problem
Gene–disease associations differ in character by source. Curated sources often assert an association verified by an expert without grading it; statistical sources assign a score between 0 and 1; and text-mining sources use their own scale based on literature co-mention frequency.
|
Source type |
Example |
Score character |
Treatment in adjudication |
|---|---|---|---|
|
Curated (variant-based) |
Clinical variant databases |
Ungraded — asserts association only |
Both score columns left empty. No arbitrary maximum is assigned |
|
Curated (phenotype) |
Phenotype ontologies, rare disease databases |
Ungraded |
Same |
|
Integrated statistical |
Drug target integration platforms |
0–1 scale |
Used directly in the normalised column |
|
Text-mined |
Literature co-mention channels |
Own scale (for example 0–5) |
Used in comparison only after mapping to the normalised column |
The error that recurs in practice is comparing across sources on the raw scale column. A value of 3.9 from a source whose maximum is 5 always looks larger than 0.9 from a source whose maximum is 1, so a single literature co-mention outranks a curated causal association and clears every threshold set against a 0–1 scale. Only the column mapped to a common interval is safe to compare, threshold or rank across sources, and the verdict presents both columns while stating which one is comparable.
The disease axis is also a direct input to modality suitability. Decomposing pathogenic variants by type — premature stop, splice site, deep intronic, particular transitions, frameshift deletions — yields, as data, what percentage of patients each modality covers. That decomposition is the basis for sequencing development in rare genetic disease, and it is treated further in section 22.8.
11.7 The network axis — interactions and the signed causal layer
11.7.1 Why degree is reported per source
Total degree alone makes well-studied genes look like hubs, because more literature means more reported interactions. That is research bias rather than biological centrality. The verdict presents total degree together with per-source degree, so that connections concentrated in one source can be distinguished from those supported by several independent ones.
- Total degree and per-source degree — separating research bias from real centrality
- Top partners with the list of sources supporting each edge — a direct display of evidence strength
- Complex members — sets that operate physically as one unit. Silencing one member can disable the whole complex, so they are considered together in target selection
11.7.2 The signed causal layer — the core asset of this layer
A plain interaction edge says only that two proteins meet; it has neither direction nor sign. Finding a substitute target requires knowing what activates or inhibits what. The verdict loads signed, directed edges from curated signalling sources as a separate layer and presents upstream and downstream relative to the target.
|
Layer |
What it holds |
Which verdict it serves |
|---|---|---|
|
Upstream activators |
Regulators that raise the target's expression or activity |
When the target must be lowered but cannot be silenced directly, silencing these is the detour |
|
Upstream inhibitors |
Regulators that lower the target's expression or activity |
When the target must be raised (Modality V), silencing these is the activation detour — achieving an upward objective with a mature silencing modality lowers development risk |
|
Downstream effectors (activated) |
What the target switches on |
When the target is pan-essential and untouchable, the downstream point that mediates the therapeutic effect |
|
Downstream effectors (inhibited) |
What the target switches off |
The expected consequence of silencing the target — what is released |
Sources sometimes report opposite signs for the same edge. The verdict marks these as 'sign disputed across sources', and that notation is itself a finding: a regulator whose direction depends on context is not a reliable point of manipulation. Picking one sign arbitrarily would destroy that information.
11.8 The two-step neighbourhood — searching for detours
Where direct partners do not yield a substitute target, the search extends to genes two hops away. Listing two hops raw is not informative, however, because most genes connect to most genes. Three treatments make it readable.
11.8.1 Sign composition
Signs along a path are multiplied to give the net effect. An inhibitor of an inhibitor is an activator, so silencing an inhibitor two hops away is a viable strategy for raising a target. The verdict shows, for each path, the signs of its constituent edges and the composed net sign.
Notation example
HSPA8 --| CFTR direct: HSPA8 inhibits CFTR
STUB1 --| HSPA8 --| CFTR net: STUB1 activates CFTR *also a direct partner
CFTR --| MCC --| APC net: CFTR activates APC
--| inhibits --> activates the net sign is the product of signs along the path
11.8.2 Demoting hub intermediates
Paths through a hub gene with very many input edges exist between almost any pair of genes. The statement 'X acts on the target through some growth factor' is true of most of the genome and therefore carries no information. The verdict ranks paths through intermediates above a threshold (default 30 inputs) last and labels them as hub-mediated.
11.8.3 Marking overlap with direct partners
When a gene two hops away is also directly connected to the target, that fact is marked separately. Such a gene is more likely to belong to a coherent functional module than to lie on a chance path, so its credibility as a substitute target candidate is higher.
11.8.4 The second shell
Genes two hops out from the direct partner set are ranked by how many distinct first-shell partners bridge to them. A gene on which several first-shell partners converge is more likely to be central to that functional module, and has stronger support as a point of manipulation than a gene connected by a single path.
11.9 Modality suitability adjudication
Synthesising the six axes above, the layer adjudicates for each modality whether the target meets its requirements. Those requirements are independent of oligonucleotide quality, and where they are unmet no therapy results regardless of design quality.
|
Modality |
Requirement imposed on the target |
Which axis adjudicates |
When unmet |
|---|---|---|---|
|
siRNA · ASO · ncRNA silencing |
Lowering it must be therapeutically correct, and it must not be pan-essential |
Disease axis direction plus safety axis |
Target cells die or broad toxicity appears |
|
saRNA (transcriptional activation) |
The therapeutic direction must be upward and the locus must be activatable |
Disease axis direction plus chromatin context |
The wrong direction can worsen the disease |
|
SSO (splice switching) |
The disease mechanism must run through splicing and the regulatory element must be characterised |
Variant type decomposition |
There is no point of intervention |
|
ADAR editing |
The pathogenic variant must be a transition that editing can revert |
Variant type decomposition |
Not applicable |
|
Suppressor tRNA |
Premature stop codons must account for a meaningful share |
Variant type decomposition |
Almost no patients are addressable |
|
miRNA modulation |
A disturbance of that axis must be established in the disease |
Disease axis plus target-set convergence |
The direction of the network effect cannot be predicted |
|
Aptamer conjugate |
The target protein must be surface-exposed or secreted |
Cellular component ontology |
An intracellular protein cannot be reached physically |
|
Immunomodulatory oligonucleotide |
The therapeutic purpose must be immune stimulation or suppression |
Therapeutic hypothesis |
The modality is unrelated to the target gene |
11.9.1 Verdict values and redirection
Verdicts take four values — recommended, conditional, not applicable, no data — each presented with the target gene, the modality, the evidence and the caveats. 'Not applicable' is a result rather than a blank and keeps its evidence: where a gene is excluded for direction reasons and is also pan-essential, both reasons are shown. Keeping only one would mean a later run in the opposite direction loses the safety warning.
When a requirement is unmet, the correct answer is frequently a substitute target rather than 'not applicable', and that redirection uses the network axes of 11.7 and 11.8.
Three canonical redirections
Aptamer requested against an intracellular protein
-> redirected to a first-shell interaction partner carrying a surface or secreted annotation
Silencing requested against a pan-essential gene
-> redirected to a downstream effector in the signed causal layer that is weakly constrained
and tissue-restricted
Activation requested for a gene that must go down
-> redirected to silencing an upstream inhibitor (an upward objective achieved with a mature
silencing modality)
11.9.2 A worked verdict
Adjudication for a cystic fibrosis gene (12 recommended / 2 conditional / 2 not applicable)
suppressor tRNA that gene 341 pathogenic stop_gained variants -- 23% of its pathogenic alleles
SSO that gene 216 pathogenic splice-site variants -- 15%
ADAR editing that gene 139 pathogenic G>A -- exactly the transition A-to-I editing reverts
aptamer conjugate that gene surface annotation present in the cellular component ontology
saRNA NFE2L2, YY1, GOPC
this gene should go down, so activating it is backwards ->
activate a curated upstream negative regulator instead
ncRNA silencing -- this gene is protein-coding; that modality targets ncRNA
11.10 Family selectivity adjudication — paralog analysis
11.10.1 The shape of the problem
Most genes have sequence-similar paralogs. Whether to co-silence or spare them cannot be decided automatically, because the correct answer inverts with the therapeutic hypothesis: where family members are functionally redundant and inhibiting one is compensated, all must be silenced; where each carries a distinct normal function, only the target must be.
This is why the adjudication depends on pathway and network analysis. Sequence identity alone cannot establish functional redundancy. Two genes at 60% identity may sit in entirely different pathways, while two at 30% identity may be interchangeable members of the same complex. The verdict uses sequence identity only as one input and takes functional sharing and disease association as its primary evidence.
11.10.2 Decision rules
|
Mode |
Rule |
Verdict |
Evidence axis |
|---|---|---|---|
|
Gene mode |
The paralog shares at least one pathway or molecular function with the target |
Co-silence |
Function axis (pathways plus molecular function ontology) |
|
Gene mode |
The paralog sits only in a non-overlapping pathway |
Spare |
Function axis |
|
Disease mode |
Select only disease-associated paralogs — functional agreement alone is insufficient |
All disease-related paralogs are co-silenced |
Disease axis plus function axis |
|
Disease mode |
Present only in the normal pathway and not disease-linked |
Spare |
Disease axis |
|
Safety override |
The paralog is pan-essential |
Converted to spare regardless of functional agreement |
Safety axis (cell dependency plus constraint) |
Defaults are sequence identity at or above 20%, a function term set size at or below 300, use of pathways and the molecular function ontology, and a minimum association of 0.1. The term set size cap exists for the reason given in 11.5.1 — sharing an overly large pathway is no evidence of functional relatedness.
11.10.3 The two auxiliary scores
redundancy score = 0.35 · sequence identity
+ 0.25 · min(shared function terms / 8, 1)
+ 0.20 · co-expression
+ 0.20 · joint dispensability
priority score = redundancy score
+ 0.10 (validated synthetic-lethal partner)
+ 0.15 · disease association
- 0.15 (documented adverse-event precedent)
- 0.10 (strong loss constraint)
- 0.10 (residual fitness dependency)
The redundancy score summarises whether this paralog can substitute for the target's function; the priority score summarises whether it may nonetheless be touched. They are kept separate because the two judgements diverge: a paralog that is functionally fully redundant but very strongly constrained scores high on redundancy and low on priority, and that combination is precisely the warning that co-silencing would be effective but dangerous.
11.10.4 The shape of the verdict output
paralog identity terms dep% LOEUF co-exp redundancy priority verdict
A 62% 17 7% 0.41 +0.80 0.79 0.79 co-silence
shares 17 terms, 11 of them specific
shared: oligosaccharyl transferase activity ... | tractability: structure with ligand
top tissue: ovary, tissue specificity 0.71
B 31% 0 100% 0.40 +0.61 0.23 -- spare
no shared pathway or molecular function term -- a different role, spare it
Instructions to carry into design:
--paralog-allow A (keep only candidates that also silence A)
--paralog-exclude B (drop any candidate that touches B)
11.10.5 Interpretation limits that must be read alongside
- Sharing a molecular function term does not imply the same substrate. Two proteins with the same enzyme-activity term acting on different substrates provide no basis for co-silencing, and this is the main source of over-broad verdicts. This is why the count of specific (lower-level) shared terms is reported alongside
- Paralogs with tumour-suppressor character are not detected by functional sharing alone. The disease axis must be read with it, and genes whose loss causes disease must be excluded from co-silencing
- The functional ontology is loaded principally on the molecular function aspect, so sharing at the biological process level is not reflected. Suspected process-level redundancy needs separate review
- Symbol mismatches produce not errors but silent empty results. Genes with aliases are resolved to the approved symbol before lookup, and the resolution is stated in the report
- A loss constraint threshold alone misses a substantial fraction of risky genes; cell dependency and adverse-event precedent are read with it
- Sex-chromosome genes may have empty constraint values from reference coverage limits. These are marked 'no data', not 'unconstrained', and a blank must never be read as a safety signal
- Paralog relationships are themselves annotation-dependent. Recently duplicated pairs and segmental duplication regions may be reported differently by different sources
11.10.6 How the verdict enters design
Verdicts act as two hard filters at the design stage. A co-silencing instruction keeps only candidates that silence all named paralogs; a sparing instruction removes any candidate that touches even one. Used together they express a compound requirement: these three must be hit, those two never.
Application differs by modality. Where there are many candidates the instruction is a filter over a list; where the sequence is fixed in advance and there is exactly one design object, as with a miRNA mimic, it is an adjudication on that one object. In that case, silently deleting an offending paralog from the predicted target list would be the opposite of useful, because that list is the off-target evidence the adjudication is made from. The list is preserved and only the verdict recorded.
The choice of application point also requires care. In a modality with several design branches by target class, such as ncRNA silencing, the filter must be applied at the upper point where final candidates converge; applying it to one branch leaves the others unfiltered. An unfiltered branch raises no error and simply produces output in which the instruction was not honoured, which makes it hard to detect.
11.11 Cleavage coverage — what 'knocking down' actually means
'Knocking down a gene' actually means knocking down some set of its isoforms. For genes with many isoforms two candidates commonly differ substantially in coverage, and that difference governs the interpretation of experimental results. Coverage is recomputed directly from the same transcript annotation the design engines read.
|
Item |
Rule |
Why |
|---|---|---|
|
Representative transcript |
Use the annotation's canonical isoform designation, falling back to the longest isoform |
The value only agrees if the same criterion as the design engine is used |
|
Definition of 'cut' |
Only when the transcript contains the site verbatim, with no mismatch tolerance |
The same exact-substring test the design engine uses |
|
Coding-only counts |
Separated by transcript biotype |
So that nonsense-mediated-decay and retained-intron isoforms do not dilute the denominator |
|
Alphabet normalisation |
Unified to the RNA alphabet |
Normalising to DNA instead makes every row report zero silently |
|
Cleavage-site notation |
Mechanism-dependent — a single phosphate (RISC), a gap window (RNase H1), or none (steric block) |
The arithmetic is mechanism-independent but the notation belongs to the mechanism |
One frequent confusion is worth stating. A candidate may hold the maximum coding-transcript coverage and still rank second overall. That is not a defect: the maximum is over coding transcripts alone, while the ranking sorts on coding and all-transcript counts together.
11.12 Cross-species correspondence — the basis for preclinical species selection
Before human clinical trials, safety assessment in one rodent and one non-rodent species is customarily expected. For oligonucleotide therapeutics that carries a particular meaning: meaningful pharmacology and toxicology are only possible if the molecule also binds the corresponding transcript in that species. Dosing a human-specific molecule into another species reveals chemistry class toxicity but nothing about target-related toxicity.
Species selection is therefore a design constraint rather than a matter of experimental convenience. This layer evaluates several human-primary species combination modes in parallel and presents which candidates survive in each. Because each combination demands a different set of cross-reactive species, the surviving candidate sets differ, and the question 'does a candidate exist that works in both rodent and non-human primate' can be answered with data.
That parallel evaluation is computationally heavy, so shared computations are reused under strict limits. Only pure functions may be cached; computations in which parallel threads write into one data structure in different ways are deliberately not cached, because a deterministic cache would double-count. A performance optimisation that changes the result is not an optimisation. Because of that distinction, each combination's result is equivalent to a standalone run.
11.13 Deliverables of this layer
- Identity — approved symbol and database identifiers, cytogenetic location, locus type, alias resolution result
- Safety axis — loss constraint (raw value and percentile), cell dependency and dependent cell line fraction, tractability bucket, adverse-event precedent, validated synthetic-lethal partners, and the reason where data are absent
- Expression axis — expression breadth, tissue-specificity index, top expressing tissues
- Function axis — pathways per source with sizes, the three ontology aspects separated, the pathway ladder
- Disease axis — per-source counts, curated disorders and mechanisms, standard identifiers, raw and normalised score columns, and the pathogenic variant spectrum decomposed by type
- Network axis — total and per-source degree, top partners with evidence sources, complex membership, the signed causal layer of upstream activators and inhibitors and downstream effectors, and sign-dispute markings
- Two-step neighbourhood — bidirectional two hops with composed signs, hub-mediated markings, direct-partner overlap markings, second-shell ranking, pathway ladder
- Family segregation — per-paralog verdicts with both auxiliary scores, shared term lists, and the instructions to carry into design
- Per-modality suitability verdicts — the four values with evidence and caveats, and where redirection occurred, the substitute target and its basis
- Cleavage coverage columns — coding-only, all-transcript and representative-transcript values separated
- Per-species-combination comparison of surviving candidates and the two-species preclinical verdict
11.14 What this layer cannot answer
- Relationships absent from the graph — an interaction or pathway not yet reported is treated as non-existent. Under-studied genes appear to have sparse networks, and that reflects literature volume rather than biology
- Context dependence — interactions and signs vary with cell type and stimulus, while graphs largely store context-merged assertions. Sign-dispute markings are the signal for this
- Quantitative prediction — this layer adjudicates where to intervene, not how far a target must be lowered for a phenotype to appear
- Establishing causation — association and causation differ. Even a curated causal edge is an observation in a particular experimental context and does not hold under all conditions
- Replacing experimental validation — substitute target proposals are prioritisations, and whether that point actually mediates the therapeutic effect must be confirmed experimentally
12. Modality VII — ADAR-Recruiting Editing Oligonucleotide (AIMer · EON · arRNA)
12.1 Form and definition
The common name is an ADAR-recruiting editing oligonucleotide, known by lineage as an AIMer, an EON (editing oligonucleotide), an Axiomer or an arRNA. A single-stranded, chemically modified oligonucleotide, in two length classes. The short class is roughly 20–40 nt and modified at every position; the long class is a linear guide RNA of roughly 50–100 nt with a comparatively lighter modification burden. Both form a duplex with the target mRNA while deliberately placing a cytidine opposite the adenosine to be edited, leaving one mismatch.
What separates this modality from every other is that the molecule does nothing itself. It does not cleave, block, or bind as a ligand. Its only action is to create the geometry in which an enzyme the cell already possesses does its work, and the chemistry is performed by that endogenous enzyme. Introducing no foreign protein is the entire strategic argument for the modality.
12.2 Molecular mechanism
12.2.1 Adenosine deamination and how it is read
Human cells endogenously express adenosine deaminases acting on RNA. These enzymes recognise double-stranded structure and remove the amino group at position 6 of adenosine, converting it to inosine. Inosine is read as guanosine by the translation machinery and by reverse transcriptases, so the net effect is an A-to-G change. The phenomenon occurs naturally across the human transcriptome, concentrated in double-stranded regions derived from repeat sequences.
The therapeutic idea is simple. If the enzyme recognises double-stranded structure, then creating double-stranded structure artificially at a chosen position should cause editing there. It does, and the class of pathogenic variant to which this applies is a G-to-A change, since reverting the A to G restores wild type. That variant type accounts for a substantial share of human pathogenic variants, so the applicable scope is not narrow.
12.2.2 The A:C mismatch as a geometric requirement
For the enzyme to flip the target adenosine into its catalytic pocket, the duplex must be locally destabilised at that position. Placing a cytidine opposite the target adenosine to create an A:C mismatch supplies that destabilisation, and it is an absolute design requirement. Without the mismatch the duplex is fully paired and editing efficiency falls sharply, and other mismatch types have been reported to be less effective than A:C.
12.2.3 Nearest-neighbour preferences — which targets edit well
The enzyme shows clear preferences for the bases flanking the target adenosine. In the triplet context with the target A at the centre, a uridine on the 5' side and a guanosine on the 3' side — 5'-UAG-3' — is most favourable, and a guanosine on the 5' side is the least favourable. These preferences change editing efficiency severalfold and are therefore a primary ranking axis.
The design-relevant point is that this context is already fixed by the target: the position of the pathogenic variant is the editing site, so the context cannot be chosen. The axis therefore functions not as a ranking among candidates but as an upfront assessment of how favourable this target is for the modality, and for targets in unfavourable contexts it is reasonable to consider an alternative modality.
12.2.4 Bystander editing — this modality's characteristic off-target
Within the duplex the guide creates there are adenosines other than the intended one. If those are edited too, unintended amino acid changes result; this is bystander editing. The means of suppressing it is clear: placing a uridine opposite each bystander adenosine forms a normal Watson–Crick pair, which is stable and therefore not edited.
Design enumerates every adenosine within the duplex footprint and reports, for each, the opposing base placement and the residual editing risk. Beyond the footprint there is a second class of off-target: the guide may incidentally form a duplex with another transcript and edit an adenosine there. That class is predicted from sequence complementarity and is therefore subject to transcriptome-wide scanning.
12.2.5 A direct conflict between chemistry and activity
This modality carries a structural tension no other modality has. Chemical modification stabilises the oligonucleotide, but modification within the central window around the orphan cytidine lowers editing efficiency, understood to be because phosphorothioate and 2'-modifications there impede local flexibility and enzyme access. The measure required for stability therefore directly costs activity.
Design accordingly sets a symmetric unmodified window at the centre and protects everything outside it. A narrow window gives good stability and lower activity; a wide window the reverse. The optimum depends on target context and delivery route, so there is no single answer, and the deliverables compare several window widths.
That chemical constraint enlarges the delivery requirement. Because the centre of the molecule is chemically unprotected, plasma stability cannot be secured by chemistry alone and the carrier's protective role is correspondingly larger.
12.3 Clinical precedent and patent landscape
The modality has entered clinical development, with two distinct chemical lineages each running trials. One is a short, fully modified oligonucleotide class for which human editing was reported in a hepatic alpha-1 antitrypsin deficiency programme; the other pursued hepatic and ocular programmes with a separate chemical design. There are no approved products, and classifying the modality as development-stage is accurate.
The patent landscape has three layers. The first is conceptual — recruiting the endogenous enzyme via an A:C mismatch — with early claim families from both academic groups and companies. The second is chemical, specifying particular nucleotide analogues and backbone stereocontrol. The third is structural, specifying guide designs with particular secondary structures such as hairpins.
The design layer uses only generic chemistry within a research-use scope and does not implement any company's proprietary chemistry or structure. It instead surfaces the existence of those claim families as advisories, providing a starting point for a customer's freedom-to-operate review.
12.4 Therapeutic significance
The strategic position of this modality is the gap between gene therapy and silencing. Silencing reduces quantity but cannot repair sequence; gene therapy repairs sequence but carries the burdens of a vector and of permanence. RNA-level editing repairs sequence while the change remains transient and confined to the transcript.
That transience is both a limitation and an advantage. The limitation is the need for repeat dosing; the advantage is reversibility. If an unexpected adverse effect appears, stopping treatment returns the system to baseline — a safety property that approaches permanently altering the genome do not have. The value of that property is greatest in paediatric treatment and for novel targets lacking long-term safety data.
A second advantage is the naturalness of expression control. An edited transcript remains under endogenous regulation, so expression does not leave the physiological range. The problems of overexpression or lost tissue specificity that arise when a gene is introduced under an external promoter do not occur in principle.
12.5 Chemistry and delivery requirements
|
Region |
Chemistry |
Reason |
|---|---|---|
|
Termini and flanks |
2'-O-methyl and 2'-fluoro modification with a phosphorothioate backbone |
Nuclease resistance and plasma stability |
|
The central window around the orphan cytidine |
Unmodified 2'-OH and phosphodiester linkages retained |
Modification here directly reduces editing efficiency |
|
Across the guide |
Complementarity to the target maintained; uridine fixed opposite each bystander adenosine |
Suppression of unintended editing |
|
Delivery |
Lipid nanoparticle or receptor conjugate |
Because of the unmodified central window, stability cannot be secured chemically and the carrier's role is large |
12.6 Design deliverables
- Guide sequence, the position of the A:C orphan, and the exact coordinate of the target adenosine
- Triplet context assessment — the 5' and 3' neighbouring bases and their preference grade
- A complete list of bystander adenosines with their opposing base placement and residual editing risk
- Off-target scan beyond the footprint — potential duplex formation with other transcripts
- Comparison of central unmodified window widths — the balance between stability and activity
- Per-position chemistry map with the central window marked explicitly
- Editing propensity index and candidate grade
- Comparison of the short and long length classes
- Patent-family advisories
12.7 Advanced capabilities — the bystander map and window-width search
12.7.1 A complete map of bystander adenosines
Every adenosine within the duplex footprint the guide creates is a potential bystander. Design enumerates them all and computes four things for each: distance from the target adenosine, the base placed opposite, that adenosine's triplet context, and a residual editing risk grade.
Risk depends jointly on distance and context for a reason. Editing probability falls with distance from the catalytic pocket, but a context close to the optimum can still be edited at distance. 'Far from the target' is therefore not sufficient reassurance, and a bystander in a favourable context must be protected by placing a uridine opposite it.
The consequence of bystander editing is not always the same, and that is assessed too. A bystander at the third codon position may be a synonymous substitution and harmless, while the first or second position changes the amino acid. Design reports the expected amino acid change for each bystander so that protection can be prioritised.
12.7.2 Searching the central unmodified window width
In a modality where stability and activity conflict directly there is no single optimum. Design evaluates several window widths and presents both axes for each as a comparison table.
|
Window width |
Stability axis |
Activity axis |
When it fits |
|---|---|---|---|
|
Narrow (orphan C ± 1–2) |
High — most positions protected |
Low — insufficient flexibility around the catalytic site |
Systemic dosing with long circulating exposure and no strong carrier protection |
|
Medium (orphan C ± 3–4) |
Medium |
Medium |
The standard starting point |
|
Wide (orphan C ± 5 or more) |
Low — a long unprotected stretch |
High — free enzyme access |
Particle encapsulation, where the carrier supplies protection |
That conflict connects directly to the delivery choice. Strong carrier protection allows a wide window and the activity that comes with it; conditions closer to carrier-free demand a narrow window for stability. In this modality chemistry and delivery must be decided together rather than in sequence.
12.7.3 Off-targets beyond the footprint
If the guide incidentally forms a duplex with another transcript, an adenosine there may be edited. This risk is predicted from sequence complementarity and is therefore subject to transcriptome-wide scanning. What differs from an ordinary silencing off-target scan is an additional condition: binding alone is not a problem unless an editable geometry is formed.
Design evaluates, for each complementarity hit, whether an arrangement corresponding to an A:C mismatch arises, and reports plain binding hits separately from editable hits. Only the latter is a substantive risk, and counting them together overstates it.
12.7.4 Choosing between the two length classes
|
Class |
Length |
Chemical burden |
How the enzyme is recruited |
Selection criterion |
|---|---|---|---|---|
|
Short |
about 20–40 nt |
Modified at every position except the central window |
The duplex itself is the recognition structure |
Where chemical stability matters and carrier burden should be reduced |
|
Long |
about 50–100 nt |
Comparatively light |
A longer duplex presents a wider enzyme binding surface |
Where editing efficiency is the priority and delivery can supply stability |
The two classes also differ in patent landscape. The deliverables present separate advisories on the known claim families for each, and state that only generic chemistry within a research-use scope is used.
12.8 What this modality cannot solve
- Variants other than A-to-G — the editing direction is fixed, so other variant types cannot be reverted. Editing in the C-to-U direction requires a different enzyme family and is a separate technology
- Cells with low enzyme expression — editing efficiency depends directly on endogenous enzyme levels, so efficiency falls in tissues where expression is low
- Targets in unfavourable context — a target with a 5' guanosine neighbour is intrinsically hard to edit and chemistry does not compensate
- Permanent correction — editing is at the transcript level and requires repeat dosing
- Complete elimination of bystander editing — it can be suppressed but not reduced to zero, and residual risk must be confirmed experimentally
13. Modality VIII — Suppressor tRNA / ACE-tRNA (Premature Stop Codon Readthrough)
13.1 Form and definition
The common name is a suppressor tRNA, and a cognate design that restores the wild-type amino acid is specifically called an ACE-tRNA. A single-chain RNA of roughly 72–76 nt that folds through a cloverleaf secondary structure into an L-shaped tertiary structure — a transfer RNA. Where every preceding modality was a ligand that recognises a target, this is a functional molecule that works in the ribosome: its anticodon pairs with the premature stop codon, and the amino acid carried at its 3' end is added to the growing polypeptide.
The central design concept is cognate suppression. Reading a stop codon with an arbitrary amino acid inserts the wrong residue, and the protein may not fold correctly. A cognate suppressor tRNA is designed to restore the original wild-type amino acid at that position, which is achieved by using the natural tRNA body cognate to that amino acid as a scaffold and editing only the anticodon.
13.2 Molecular mechanism
13.2.1 Nonsense mutation as a disease class
When a single base substitution converts an amino acid codon into a stop codon, translation terminates early, producing a truncated protein or triggering degradation of the transcript by the nonsense-mediated decay machinery. Either way no functional protein is made. This variant type has been reported to account for roughly 11% of variants causing inherited disease, and it represents a substantial patient fraction in cystic fibrosis, muscular dystrophy, epidermolysis bullosa and other conditions.
The previous approach was small-molecule readthrough inducers, which lower ribosomal fidelity so that stop codons are skipped. That approach has fundamental limitations: the amino acid inserted is not controlled, and normal stop codons across the transcriptome are affected too. A suppressor tRNA addresses both — the inserted amino acid is determined by design, and selectivity arises from the identity and context of the stop codon.
13.2.2 A six-stage design logic
|
Stage |
What it determines |
Constraint |
|---|---|---|
|
Scaffold selection |
Which natural tRNA body cognate to the wild-type amino acid to use |
The synthetase for that amino acid must still recognise and charge this tRNA. Scaffolds are ranked using gene copy number as an abundance proxy |
|
Anticodon engineering |
Replacing the anticodon triplet to be complementary to the premature stop codon |
Wobble rules apply. The aminoacylation identity elements — the acceptor stem and the discriminator base — must be invariant, and invariance is explicitly asserted |
|
Structure validation |
Whether the edited sequence still folds into a cloverleaf |
Covariance-model validation as primary, free-energy folding as secondary, plus checks on T-stem and D-stem topology |
|
Readthrough efficiency prediction |
The probability of winning the competition against termination factors |
The base immediately following the stop codon, its surrounding context, and the P-site codon strongly affect efficiency |
|
Safety assessment |
Quantifying the burden on normal stop codons across the transcriptome |
Every gene ending in the same codon and a similar context is at risk of C-terminal extension. Also checked: mischarging risk, clinical variant cross-reference, nonsense-mediated decay prediction |
|
Output |
Full tRNA sequence, anticodon, restored amino acid, efficiency and safety scores, optional DNA cassette |
Modification-impact assessment and a composite score |
13.2.3 Global burden is the decisive safety axis
A suppressor tRNA does not in principle read only the target gene. Other genes ending in the same stop codon in a similar context may also have their C-termini extended, and the human transcriptome may contain hundreds of such genes. If an extended protein loses function, aggregates, or mislocalises, broad toxicity follows.
A design that has not quantified this burden across the transcriptome is a design that looks good only at its target. The design layer scans the whole transcript annotation, enumerates genes sharing the stop codon and context, and separately flags those for which C-terminal extension is particularly dangerous — essential genes, transmembrane proteins, aggregation-prone proteins.
Two means of achieving selectivity exist. First, designing the anticodon to recognise only the stop codon the target actually carries leaves genes ending in the other two unaffected. Second, context preferences including the base immediately after the stop codon lower efficiency at genes with a different context. The resulting selectivity is partial but real, and it is quantifiable.
13.2.4 Nonsense-mediated decay as a second barrier
Many transcripts carrying a premature stop codon are degraded by the nonsense-mediated decay machinery. However efficient the suppressor tRNA, there is no effect if there is no transcript left to read. Susceptibility depends on the position of the stop codon — upstream of the last exon–exon junction it is likely to be a decay substrate, downstream it escapes — and the design emits that prediction.
Two practical implications follow. For highly decay-susceptible targets, a suppressor tRNA alone has limited effect and combination with decay inhibition is discussed. And within one gene, suitability for this modality varies with variant position, so stratifying the patient population by variant position is strategically advantageous.
13.3 Development status
The modality is moving from proof of concept into preclinical development. Reports have accumulated that cognate suppressor tRNAs read premature stop codons and restore functional full-length protein in cell and animal models (Lueck and colleagues, Nat Commun 2019, among others), and several companies are running programmes differing in delivery and indication. There are no approvals.
The principal development issues are three. Delivery: a folded functional RNA does not undergo carrier-free uptake, so either an expression cassette must be delivered by vector or particle, or a synthetic tRNA must be delivered in a particle. Global readthrough burden, as described above. And control of expression level, since an overexpressed suppressor tRNA can broadly affect normal translation termination.
13.4 Chemistry requirement — not required in principle
This is the one modality in which chemical modification is not required in principle, and the reason lies in the nature of the molecule. It is not a synthetic ligand but a functional RNA that must fold inside the cell, be recognised by a synthetase and charged with an amino acid, and bind an elongation factor to enter the ribosome. All three processes depend on the precise structure and modification pattern of a natural tRNA, and an artificial modification perturbing any one of them disables the molecule.
The standard route is therefore delivery as a DNA cassette so that transcription, modification and folding occur inside the cell. In that case the deliverable is not an RNA sequence but an expression cassette — promoter, tRNA gene, terminator. Where a synthetic tRNA is dosed directly, minimal stabilisation is discussed, restricted to a scope that does not perturb native modification sites or tertiary structure.
Design separately assesses the impact of modification sites. tRNAs carry numerous chemical modifications that contribute to folding stability and codon recognition fidelity, and anticodon editing that alters a modification site can produce unexpected results.
13.5 Delivery requirement — required
|
Delivery mode |
Form |
Advantages |
Constraints |
|---|---|---|---|
|
DNA expression cassette in a lipid nanoparticle |
A tRNA gene under a polymerase III promoter |
Native transcription, modification and folding occur in the cell |
Expression level is hard to control and duration is short |
|
DNA expression cassette in a viral vector |
The same cassette carried by a vector |
Sustained expression, and tissue tropism can be conferred |
Vector-associated immunogenicity and the burden of permanence |
|
Synthetic tRNA in a lipid nanoparticle |
Transcribed and purified, or chemically synthesised, tRNA |
Expression level is controlled by dose; transient |
Possible loss of activity from absent modifications; stability |
In every case a vehicle is required, which raises the development burden. Where the target tissue is well defined, however — airway epithelium in cystic fibrosis, muscle in muscular dystrophy — local administration to reduce systemic exposure is a viable strategy.
13.6 Design deliverables
- Full tRNA sequence, the edited anticodon, and the wild-type amino acid restored
- Scaffold selection rationale — which natural tRNA body was chosen and why, with the abundance-proxy ranking
- Aminoacylation identity element invariance check — explicit confirmation of the acceptor stem and discriminator base
- Structure validation — covariance model results, free-energy folding, T-stem and D-stem topology
- Readthrough efficiency prediction with the contributions of stop codon context and P-site codon decomposed
- Transcriptome-wide normal stop codon burden — enumeration of genes sharing codon and context, with high-risk genes flagged
- Mischarging risk assessment
- Clinical variant database cross-reference — whether the nonsense variant is reported, with disease annotations
- Nonsense-mediated decay susceptibility prediction from variant position relative to the last junction
- Modification impact assessment and a composite efficacy score
- Optional DNA expression cassette sequence
Scaffold bodies are taken from actual gene sequences in a public tRNA database; no sequence is invented. Design in this modality uses no machine learning and is a fully rule-based deterministic pipeline, because the structural and identity-element requirements are precisely specified and what matters is detecting rule violations definitively rather than estimating them statistically.
13.7 Advanced capabilities — identity element preservation and global burden quantification
13.7.1 Invariance checking of the aminoacylation identity elements
The most important constraint in this design is: change the anticodon, do not change the identity. The synthetase reads particular positions of the tRNA to decide which amino acid to load, and disturbing those positions loads the wrong amino acid or prevents charging altogether. Design compares the sequence before and after editing, asserts explicitly that the identity elements are invariant, and terminates as a failure on violation.
- The acceptor stem — base pairs near the 3' terminus where the amino acid attaches; the principal recognition site for many synthetases
- The discriminator base — the unpaired base at the end of the acceptor stem, a key determinant of amino acid specificity
- Particular positions in the D-stem and T-stem — involved in recognition depending on the amino acid family
- Amino acid families where the anticodon is itself an identity element — in that case editing the anticodon destroys the identity, so that amino acid is judged outside the scope of this approach
The last item matters in practice. For some amino acids the synthetase reads the anticodon directly, so changing it means the tRNA is no longer charged with the original amino acid. Design first checks whether the wild-type amino acid falls into that category and, if so, searches for an alternative scaffold or returns a not-applicable verdict.
13.7.2 Scaffold selection and the abundance proxy
Several tRNA genes exist for the same amino acid, and which is used as the scaffold affects the outcome. Building on an abundant family makes it more likely that the expression, processing and modification pathways are already optimised for that sequence. Design ranks scaffold candidates using gene copy number as an abundance proxy, and states as a limitation that copy number is not perfectly proportional to actual abundance.
13.7.3 Structure validation in three layers
|
Layer |
Method |
What it confirms |
Meaning of failure |
|---|---|---|---|
|
Primary — covariance model |
Alignment score against a tRNA-family covariance model |
Whether the edited sequence still has a structure recognised as a tRNA |
The cloverleaf itself has collapsed — discard the design |
|
Secondary — free-energy folding |
Minimum free energy secondary structure prediction |
Whether the predicted structure matches the cloverleaf and whether alternatives compete |
Competing folds lower the fraction of functional molecules |
|
Tertiary — domain checks |
T-stem elongation factor binding proxy, D-stem topology check |
Whether the tertiary elements required for ribosome entry are retained |
It folds but does not enter the ribosome |
All three layers must pass because the failure points differ. A molecule that does not fold, one that folds but is not charged, and one that is charged but cannot enter the ribosome all produce the same observation — it does not work — while demanding completely different responses.
13.7.4 Readthrough efficiency — the context model
The probability of winning the competition against termination factors depends far more on the context around the stop codon than on the codon itself. The base immediately following the stop is particularly decisive, so the tetranucleotide formed by the stop codon plus that base is reported as a principal determinant of termination efficiency. Design predicts readthrough efficiency with an extended model that adds the P-site codon and roughly six nucleotides of context on each side to that tetranucleotide table.
The practical value of the model is patient stratification. Even for the same variant type in the same gene, a different context gives a substantially different predicted efficiency, so which variant-position subgroup is likely to respond can be distinguished in advance.
13.7.5 Global stop codon burden — quantified transcriptome-wide
This is the decisive safety axis for a suppressor tRNA, and it is quantified by scanning the whole transcript annotation. The procedure is as follows.
1. Fix which stop codon the target carries (one of three)
2. Enumerate every gene in the annotation ending in the same stop codon
3. Compute context similarity (the tetranucleotide plus flanks) for each
4. Classify high-similarity genes as the readthrough risk set
5. Within that set, flag genes for which C-terminal extension is especially harmful:
- essential genes (loss of function on extension would be lethal)
- transmembrane proteins (extension changes localisation)
- aggregation-prone proteins (extension induces aggregation)
Two means of securing selectivity exist and both are partial. Designing the anticodon to recognise only the stop codon the target carries leaves genes ending in the other two unaffected, and context preferences lower efficiency at genes with a different context. The resulting selectivity is quantifiable but cannot be reduced to zero.
13.7.6 Nonsense-mediated decay susceptibility
Many transcripts carrying a premature stop are degraded by the decay machinery, and susceptibility depends on the position of the stop codon: upstream of the last exon–exon junction (by roughly 50 nucleotides or more) it is likely a decay substrate, while downstream or in the last exon it escapes. Design compares variant position with junction coordinates to emit that prediction.
The output has two uses. For highly susceptible targets a suppressor tRNA alone has limited effect and combination strategies must be considered; and because suitability varies with variant position within the same gene, it forms the basis for patient stratification.
13.7.7 Modification impact and the expression cassette
tRNAs carry numerous chemical modifications that contribute to folding stability and codon recognition fidelity. If anticodon editing alters a modification site or the recognition sequence of a modifying enzyme, unexpected results can follow, so design checks for overlap between known modification positions and the edited positions.
The standard deliverable is an expression cassette rather than an RNA sequence — a polymerase III promoter, the tRNA gene body and a terminator — because delivery in that form allows native transcription, modification and folding to occur inside the cell. A direct synthetic tRNA route is also supported, with the possible loss of activity from absent modifications stated in the deliverables.
13.8 What this modality cannot solve
- Variants other than nonsense — it does not apply to missense, deletion or splice variants
- Targets with strong nonsense-mediated decay — with no transcript left to read, the effect is limited
- Complete elimination of global readthrough burden — codon identity and context provide selectivity but cannot reduce it to zero
- Precise control of expression level — particularly difficult with cassette delivery
- Guaranteeing the restored protein's function — even with the wild-type amino acid inserted, whether the protein recovers normal folding and localisation requires separate experimental confirmation
14. Modality IX — CpG-ODN / Immunomodulatory Oligonucleotide (TLR7 · TLR8 · TLR9)
14.1 Form and definition
The common name is an immunomodulatory oligonucleotide; the TLR9 agonist class is specifically called a CpG-ODN (CpG oligodeoxynucleotide) or an immunostimulatory sequence (ISS), and the inhibitory direction an INH-ODN. A single-stranded DNA or RNA of 18–30 nt with a largely phosphorothioate backbone. What decisively separates it from the preceding modalities is that the target is a protein rather than a nucleic acid. This molecule does not recognise anything by base pairing. It presents the features that endosomal Toll-like receptors read as 'foreign nucleic acid' — unmethylated CpG dinucleotides, uridine-rich sequence — and thereby activates or blocks the receptor.
The meaning of 'off-target' therefore differs here. In other modalities an off-target is unintended binding to a transcript; here it is both unintended immune activation and incidental gene repression through sequence complementarity. Design assesses both layers.
14.2 Molecular mechanism
14.2.1 Three receptors and their ligands
|
Receptor |
What it recognises |
Response on activation |
Therapeutic direction |
|---|---|---|---|
|
TLR9 |
DNA carrying unmethylated CpG motifs |
Type I interferon from plasmacytoid dendritic cells; B cell activation |
Vaccine adjuvant, immuno-oncology, infection prophylaxis |
|
TLR7 |
Uridine- and guanosine-rich single-stranded RNA |
Predominantly type I interferon from plasmacytoid dendritic cells |
Antiviral stimulation, adjuvant |
|
TLR8 |
Uridine-rich single-stranded RNA (pronounced in humans) |
Predominantly inflammatory cytokines from monocytes and macrophages |
Reprogramming the tumour microenvironment |
|
All three (antagonism) |
Inhibitory sequences — telomeric repeats, guanosine runs, particular triplets |
Suppression of receptor signalling |
Autoimmune and chronic inflammatory disease |
That these receptors reside in endosomes has two design implications. First, the molecule must be endocytosed to reach its target, so endocytosis is itself delivery; the phosphorothioate backbone mediates that uptake, so no separate vehicle is needed. Second, the endosomal interior is acidic and contains nucleases, so the molecule needs stability under those conditions.
14.2.2 Structural classes of CpG oligonucleotide
TLR9 agonists produce markedly different immune profiles according to structure. This classification is not academic tidiness but a set of practical options, since the class is chosen according to the kind of immune response desired.
|
Class |
Structural features |
Immune profile |
Suited to |
|---|---|---|---|
|
Class A (D-type) |
Palindromic CpG core on a natural backbone with phosphorothioate guanosine runs at both termini |
Strong type I interferon from plasmacytoid dendritic cells; weak B cell activation |
Antiviral stimulation, natural killer cell activation |
|
Class B (K-type) |
Fully phosphorothioate linear with multiple CpGs |
Predominantly B cell activation and inflammatory cytokines; weak interferon |
Vaccine adjuvant — enhancing antibody responses |
|
Class C |
Fully phosphorothioate and palindromic |
Intermediate — induces both interferon and B cell activation |
Where a balanced response is needed |
|
Class P |
Two palindromes forming higher-order concatamers |
Very strong interferon induction |
Where a powerful interferon response is required |
Class A carries guanosine runs at both termini because those segments form G-quadruplex structures that assemble the molecules into higher-order aggregates. That aggregation directs trafficking to a particular endosomal compartment, and signalling from that compartment leads to the interferon pathway. In this modality, therefore, secondary and higher-order structure is part of the function rather than a side effect, and design treats G-quadruplex propensity as something to be tuned rather than avoided.
14.2.3 Species specificity — the central trap in preclinical interpretation
The optimal recognition motif for TLR9 differs between species. The optimal hexamer in humans differs from that in mice, and a strong agonist in one may be weak in the other. This bears directly on the interpretation of preclinical data: failing to distinguish a molecule problem from a species difference leads to wrong development decisions.
TLR8 is more extreme. It is functional in humans but responds very differently in mice, which limits the value of evaluating a human TLR8 agonist in murine models at all. Design evaluates the optimal motif per species and emits cross-species concordance as a separate item, providing the basis for preclinical species selection and data interpretation.
14.2.4 The mechanism of antagonists
Inhibitory oligonucleotides interfere with the receptor itself or its signalling pathway. Telomere-derived repeat sequences, guanosine runs and particular triplet repeats have been reported to carry inhibitory activity, and 2'-O-methyl RNA is known to block TLR7 and TLR8. Where chronic innate immune activation by nucleic acid autoantigens forms part of the pathology of an autoimmune disease, intervention in this direction fits logically.
14.3 Clinical precedent and therapeutic significance
This is one of the few non-silencing oligonucleotide modalities with approval precedent. A hepatitis B vaccine using a class B CpG oligonucleotide as adjuvant was approved in 2017, establishing that a synthetic oligonucleotide can obtain regulatory approval as a vaccine adjuvant. What makes that case significant is the scale of the safety dataset: adjuvants are given in large numbers to healthy adults, so the approval created broad safety evidence for that chemistry and dose range.
From a vaccine development standpoint the modality's value is that the direction of the response can be steered. Traditional aluminium salt adjuvants induce mainly antibody responses and weak cell-mediated immunity. A CpG oligonucleotide can selectively strengthen the interferon pathway or the B cell pathway depending on class, so the type of immune response can be designed to suit the pathogen. This is particularly useful for pathogens where cell-mediated immunity is protective and for populations such as the elderly whose responses are weak.
In immuno-oncology, intratumoral administration to reprogramme the microenvironment is under study. The tumour microenvironment is generally immunosuppressive, and the logic is to stimulate innate immunity locally to relieve that suppression and improve the effect of checkpoint inhibitors. In this approach minimising systemic exposure is the key to safety, and local administration with short residence becomes an advantage rather than a limitation.
14.4 Chemistry requirement — the modification is the function
Here phosphorothioate is both a stabiliser and part of the function. The backbone mediates endocytosis, provides stability inside the endosome, and in some classes participates in higher-order structure formation. The strategy used in other modalities — reduce phosphorothioate to lower class toxicity — therefore does not transfer directly.
|
Class |
Backbone chemistry |
Reason |
|---|---|---|
|
Class A agonist |
Natural backbone in the core, phosphorothioate only in the terminal guanosine runs |
A natural core favours optimal recognition while terminal phosphorothioate carries higher-order structure and stability |
|
Class B and C agonists |
Fully phosphorothioate |
Maximises stability and uptake; being linear, higher-order assembly is not required |
|
TLR7/8 agonists |
Phosphorothioate RNA preserving uridine-rich sequence |
RNA is less stable than DNA, so backbone protection matters more |
|
Antagonists |
Phosphorothioate combined with 2'-O-methyl RNA |
2'-O-methyl is an active component of TLR7/8 blockade |
A caution: in agonist design, heavy 2'-modification around the CpG motif abolishes recognition. The receptor reads specific chemical features, so piling on modifications for stability produces a molecule that is not recognised. Design treats the region around the motif as protected and secures stability outside it.
14.5 Delivery requirement — unnecessary
Because the target is a receptor inside the endosome, endocytosis is itself arrival, and the phosphorothioate backbone mediates that endocytosis, so no separate vehicle is required. In the approved case the oligonucleotide is formulated together with the vaccine antigen without a carrier. DNA CpG oligonucleotides are sometimes formulated with aluminium salts, not for delivery but for co-localisation with antigen and release control.
Attempts to increase targeting to particular cells do exist, since acting only on selected immune cell subsets would give the desired response without systemic cytokine release. Antibody or ligand conjugation and particle formulation are under study, and in those cases a vehicle is introduced.
14.6 Design deliverables
- Receptor target, direction (agonism or antagonism) and CpG class call
- Sequence and motif layout — CpG positions, spacing, 5' initiating triplet, palindromic structure
- Per-species optimal motif assessment — predicted activity in human, mouse and non-human primate
- Cross-species concordance — the basis for interpreting preclinical data
- Per-position chemistry map reflecting the class-specific backbone rules
- Structural assessment — nearest-neighbour melting temperature with backbone correction, secondary-structure free energy, G-quadruplex motif detection, motif accessibility
- Composite efficacy score combining potency, structure, safety and motif spacing
- Three-layer safety — melting-temperature-weighted sequence off-targets, cross-species concordance, immunotoxicity (G-quadruplex, self-duplex, backbone burden, CpG density)
- TLR7/TLR8 selectivity indices for RNA agonists
Prediction is rule-primary: literature-anchored motif scoring makes the first-pass call, and a light machine-learning rerank is optional and degrades to the identity function when no model is present. Machine learning never overrides a hard rule call by design, because immune response is an area with comparatively well-codified rules and detecting rule violations definitively matters more than statistical estimation.
14.7 Advanced capabilities — species specificity, structural scoring, three-layer safety
14.7.1 Species-specific recognition motifs — the premise for preclinical interpretation
The optimal TLR9 recognition hexamer differs by species. Humans and primates are reported to prefer 5'-GTCGTT-3' while mice prefer 5'-GACGTT-3', and that difference makes the same molecule behave quite differently in the two species. Design computes the occurrence and position of the species-optimal motif separately for each species and emits cross-species concordance as a distinct item.
Without that item, preclinical data cannot be interpreted: a weak effect in mice cannot be attributed to the molecule rather than to a species difference. Conversely, a molecule with high concordance allows comparatively trustworthy extrapolation of preclinical results — though TLR8 differs so much between species that this extrapolation is limited in principle.
14.7.2 Structural class calling and its consequences
Design calls the CpG oligonucleotide's class automatically from sequence structure. The call rests on the presence of palindromes, the number and spacing of CpGs, the presence of terminal guanosine runs, and the backbone pattern, and the expected immune profile follows from it. Automation matters because it catches cases where the class contradicts the design intent — a molecule intended as class B may form a palindrome and behave as class C.
- The number of CpG motifs and their spacing — too close and receptor binding interferes
- Presence of a 5'-terminal initiating triplet — a structure beginning TCG is reported to favour activity
- Presence and length of palindromes — the determinant of class C and class P
- Terminal guanosine runs and the higher-order structure they form — the defining feature of class A
- Whether CpG density reaches island level — excessive density raises immunotoxicity risk
14.7.3 Structural and thermodynamic scoring
In this modality secondary and higher-order structure is something to be tuned rather than avoided. Four structural indices are computed.
|
Index |
How it is computed |
How it reads in this modality |
|---|---|---|
|
Melting temperature |
Nearest-neighbour thermodynamics with a backbone chemistry correction |
Phosphorothioate lowers duplex stability, so the value is overestimated without the correction |
|
Secondary structure free energy |
Minimum free energy folding |
Self-duplex formation obstructs receptor recognition and is to be avoided |
|
G-quadruplex motif |
Detection of contiguous guanosine patterns |
Part of the function in class A and a risk in the other classes — the sign inverts with class |
|
Motif accessibility |
Local opening probability at the CpG motif position |
A motif buried in a hairpin is not recognised |
The third row shows this modality's peculiarity. In other modalities a G-quadruplex is something to avoid because it complicates synthesis and purification and causes non-specific binding, whereas in a class A CpG oligonucleotide it is a functional element that drives trafficking to a particular endosomal compartment through higher-order assembly. That the same index carries opposite sign depending on class is what makes automatic classification necessary.
14.7.4 Three-layer safety assessment
|
Layer |
What it examines |
Method |
Why it is needed |
|---|---|---|---|
|
Sequence off-targets |
Unintended gene repression through incidental complementarity |
Melting-temperature-weighted sequence similarity search |
Even when immune modulation is the purpose, the molecule is still a nucleic acid and binds complementarily |
|
Cross-species concordance |
Difference in activity between preclinical species and human |
Comparison of the per-species optimal motif assessment |
Quantifies how far animal data can be extrapolated to humans |
|
Immunotoxicity |
Excessive or unintended immune activation |
Composite assessment of G-quadruplex, self-duplex, backbone burden and CpG density |
Immune stimulation is the objective here, so judging the appropriate level is especially important |
The third layer is particularly subtle in this modality. Assessing immunotoxicity in a molecule whose purpose is immune stimulation means examining the magnitude and type of stimulation rather than its presence, and the boundary between wanted and unwanted response shifts with indication. A vaccine adjuvant aims at a local, transient response; intratumoral dosing tolerates a stronger one; and in the autoimmune antagonist direction any stimulation is a failure. The deliverables present component values rather than an absolute grade so that they can be read in the context of the indication.
14.7.5 Rules primary, learning auxiliary
Prediction is rule-primary. Literature-anchored motif scoring makes the first-pass call, and a light machine-learning rerank is optional and degrades to the identity function when no model is present. Machine learning never overrides a hard rule call by design.
That arrangement was chosen because the rules in this area are comparatively well codified and detecting rule violations definitively matters more than estimating statistically. A TLR9 agonist with no CpG motif, or a TLR7/8 agonist short of uridine, is not a statistically low-scoring molecule but a structurally non-functional one, and that verdict must be definite.
14.8 What this modality cannot solve
- Regulation of a specific gene's expression — the target is a protein receptor, so it is not a sequence-specific gene regulation tool
- Cross-species extrapolation — species specificity is pronounced, imposing a fundamental limit on translating animal results to humans, especially for TLR8
- Avoiding systemic cytokine responses — local administration reduces them but does not eliminate them in principle
- Fine composition of the immune response — class selection steers direction, but inducing only a chosen cytokine is beyond current control
- Chronic dosing — repeated stimulation risks tolerance or chronic inflammation, making dosing schedule design a separate problem
15. Modality X — Aptamer Conjugates ApDC / AOC (Targeted Delivery)
15.1 Form and definition
A folded single-stranded nucleic acid — an aptamer — joined through a linker to a payload. Where the payload is a small-molecule cytotoxin the construct is an aptamer–drug conjugate; where it is a therapeutic oligonucleotide it is an aptamer–oligonucleotide conjugate. Aptamers are typically 25–90 nt and fold into a defined tertiary structure that binds a specific site on a protein surface.
This modality sits at a different level from the others. It does not produce a therapeutic effect itself but carries something else to a destination. Where the preceding modalities addressed what to do, this one addresses where to go. It occupies, using nucleic acid, the role antibody–drug conjugates have already established.
15.2 The biology and physical chemistry of aptamers
15.2.1 Why a nucleic acid binds a protein
A single-stranded nucleic acid pairs partially with itself to form a tertiary structure of stems and loops. The surface that structure presents has shape and charge distribution much as an antibody's complementarity-determining regions do, and it can bind a complementary protein surface. Binding constants in the sub-nanomolar range have been reported, comparable to antibodies.
There are three advantages over antibodies. Chemical synthesis gives high batch-to-batch homogeneity and no biological contamination risk. Low molecular weight favours tissue penetration. And immunogenicity is low, since nucleic acids are not presented through the major histocompatibility complex as proteins are. The disadvantages are equally clear: nucleases degrade them, low molecular weight means rapid renal excretion, and there is no reliable computational method for predicting binding affinity.
15.2.2 Affinity is not computationally predictable
This fact determines the entire design policy of the modality and is therefore stated explicitly. No computational method currently exists that reliably predicts the binding affinity of an arbitrary nucleic acid sequence to an arbitrary target protein. Secondary structure prediction is well established, but with what affinity that structure binds a particular protein surface is a different order of question.
The design policy therefore is: binding constants are not generated. Aptamers used in conjugate design are restricted to experimentally validated ones, each carrying a citable sequence, binding constant and literature source. Uncertain entries are left empty and marked unconfirmed rather than filled in. The policy is inconvenient, but far cheaper than allowing an unsupported binding constant into a development decision.
15.2.3 Internalisation — binding alone is not enough
To bring a payload inside a cell, the receptor the aptamer binds must be endocytosed. A receptor that binds at the surface without internalising is unsuitable as a conjugate target. Target selection must therefore consider not only whether the protein is surface-expressed but whether it internalises and how quickly it recycles, and the design includes known internalisation characteristics among its assessment items.
Beyond internalisation lies a second barrier, endosomal escape. For cytotoxin payloads the common design has the linker cleaved in the lysosome so the freed drug crosses the membrane; for oligonucleotide payloads, which cannot cross membranes by themselves, escape efficiency becomes the limiting factor.
15.3 The three components of a conjugate
|
Component |
Options |
Design variables |
Failure mode |
|---|---|---|---|
|
Aptamer |
Indexed by target receptor from a validated library |
Sequence, stabilising chemistry, attachment point |
Payload attachment that disrupts the fold abolishes binding |
|
Linker (cytotoxin) |
Cleavable — protease-sensitive peptides, acid-sensitive, disulfide; or non-cleavable stable covalent |
Cleavage conditions, length, hydrophilicity |
Premature cleavage in circulation causes systemic toxicity; failure to cleave means no effect |
|
Linker (oligonucleotide) |
Stable linker or disulfide |
Length and flexibility |
Too short and the two moieties interfere; too long and renal excretion increases |
|
Payload (cytotoxin) |
Microtubule inhibitors, topoisomerase inhibitors, DNA-binding agents, RNA polymerase inhibitors and others |
Potency, whether freed drug kills neighbouring cells |
Insufficient potency for the number of molecules internalised |
|
Payload (oligonucleotide) |
A silencing guide or duplex, or an antisense oligonucleotide |
Inherits the payload modality's own requirements |
The payload's chemistry may interfere with aptamer folding |
|
Drug-to-aptamer ratio |
Typically 1–4 |
Number and position of attachment points |
High ratios increase aggregation and clearance; low ratios lack potency |
The oligonucleotide-payload conjugate is conceptually interesting because it fuses two modalities. A silencing molecule is potent but delivery-limited; an aptamer is a delivery device with little therapeutic effect of its own. Combining them makes cell-type-specific silencing possible without a lipid nanoparticle or a receptor conjugate on the oligonucleotide itself, which is attractive particularly outside the liver. The endosomal escape problem described above remains, however, so practical use of this combination depends on improving escape efficiency.
15.4 Clinical precedent and therapeutic significance
Aptamers themselves have approval precedent. An aptamer binding a vascular growth factor was approved in ocular disease in 2004, and an aptamer binding a complement factor in the same field in 2023. Both are administered locally in the eye, a strategy that circumvents renal excretion and nuclease degradation through local administration.
There is no approval precedent for the conjugate form. The principal development issue is the limited number of validated aptamers. Antibodies can be raised against essentially any target through immunisation, whereas aptamers require an in-vitro selection process that does not succeed for every target. The applicable scope of this modality is therefore currently limited to targets for which a validated aptamer exists.
Areas of strategic value nonetheless exist: targets where antibody conjugates are already crowded and differentiated properties are wanted; repeat-dosing situations where antibody immunogenicity is problematic; and tissues that require low molecular weight for penetration.
15.5 Chemistry and delivery requirements
Chemical modification is required. An unmodified aptamer is degraded within minutes by serum nucleases, so the standard is stabilisation with a combination of 2'-fluoro and 2'-O-methyl plus an inverted deoxythymidine at the 3' terminus to block exonucleases. Polyethylene glycol is sometimes added to slow renal excretion.
Modification is constrained, however: changing the chemistry of residues involved in binding can destroy affinity, so the binding interface must retain its original chemistry. Which residues are involved can only be known from structural information or experimental scanning, so design incorporates such information where it is reported and is conservative where it is not.
No delivery vehicle is required. The aptamer is itself the targeting device, and adding a carrier would dilute that targeting while increasing molecular weight. Administration is intravenous or local.
15.6 Design deliverables
- Full conjugate structure notation — the connection order of aptamer, linker and payload, and the attachment point
- Aptamer information — sequence, target receptor, the binding constant as reported in the literature, and its source
- Linker specification — type, cleavage conditions, length
- Drug-to-aptamer ratio and its basis
- Serum stability score and nuclease resistance profile
- Internalisation assessment — known internalisation and recycling characteristics of the target receptor
- Secondary structure check — the effect of payload attachment on the aptamer fold
- G-quadruplex detection and its interpretation
- PEGylation advisory
- Target tissue selectivity review from receptor expression distribution
- Explicit marking of unconfirmed entries
15.8 What this modality cannot solve
- Intracellular protein targets — an aptamer binds surface proteins from outside the cell and does not reach cytosolic or nuclear proteins
- Targets with no validated aptamer — discovering a new aptamer belongs to in-vitro selection and is not replaced by computation
- Endosomal escape — for oligonucleotide payloads this bottleneck remains
- Non-internalising receptors — binding alone does not bring the payload inside
15.7 Advanced capabilities — combinatorial search and the unconfirmed policy
15.7.1 Searching the three-component combination
Conjugate design is a combinatorial problem across aptamer, linker and payload. The axes are not independent, so combination-level evaluation is required rather than per-axis optimisation: a bulky payload demands a longer linker, a longer linker accelerates renal excretion, faster excretion requires a higher drug-to-aptamer ratio, and a higher ratio raises aggregation risk.
|
Axis |
Scale of options |
Principal constraint |
Coupling to other axes |
|---|---|---|---|
|
Aptamer |
A validated library of 20 or more, indexed by target receptor |
Only validated entries are used — novel discovery belongs to experiment |
Attachment point affects folding; payload size affects fold stability |
|
Linker |
Eight or more cleavable and non-cleavable types (protease-sensitive peptides, acid-sensitive, disulfide, glucuronide, carbonate, dual-cleavage) |
The window between premature cleavage in circulation and failure to cleave intracellularly |
Payload type sets the cleavage condition; length affects excretion |
|
Payload (cytotoxin) |
Ten or more — microtubule inhibitors, topoisomerase inhibitors, DNA-binding agents, RNA polymerase inhibitors |
Very high potency is required because few molecules internalise |
Lower potency demands a higher drug-to-aptamer ratio |
|
Payload (oligonucleotide) |
A silencing guide or duplex, or an antisense oligonucleotide |
Inherits the chemistry requirements of its own modality |
The payload's chemistry may interfere with aptamer folding |
|
Drug-to-aptamer ratio |
1–4 |
High raises aggregation and clearance; low lacks potency |
The number and position of attachment points set the ceiling |
15.7.2 Quantitative items emitted
- Serum stability score — an estimate of nuclease resistance derived from modification placement and terminal handling
- Nuclease resistance profile — vulnerable positions shown separately for exonucleases and endonucleases
- Internalisation characteristics — the known internalisation and recycling behaviour of the target receptor, from the literature
- Secondary structure check — folding compared before and after payload attachment, confirming that attachment does not break a stem
- G-quadruplex detection — read in both directions, as a binding motif for some aptamers and a cause of non-specific binding
- PEGylation advisory — whether delayed renal excretion is needed and whether it would obstruct binding
- Target tissue selectivity — predicted normal tissue exposure from receptor expression distribution
- Composite efficacy index — a weighted summary of the above, for relative comparison among candidates
15.7.3 The unconfirmed-marking policy in practice
The most important design policy in this modality is not to invent values. Because no reliable computational prediction of aptamer binding affinity to an arbitrary target exists, every binding constant in the library is a literature reference value with its source shown. Entries for which the literature has no value, or where the conditions are unclear, are left empty and marked unconfirmed.
The policy is as inconvenient as it is valuable. An unsupported binding constant entering a development decision contaminates the entire candidate selection and is only discovered at the experimental stage. A blank honestly exposes the absence of information, and that absence is itself actionable — it says this item must be established experimentally.
15.7.4 Special considerations for oligonucleotide payloads
Where the payload is a silencing molecule, two modalities' requirements must be satisfied at once. The payload must carry the chemistry its mechanism demands (section 6.3), the aptamer's fold must be preserved, and the linker must be long enough that the two moieties do not interfere yet not so large as to accelerate excretion.
The largest unresolved item is endosomal escape. A cytotoxin payload can cross the membrane after release in the lysosome, but an oligonucleotide payload cannot cross by itself, so escape efficiency remains the limiting factor. Design states that bottleneck in the deliverables and marks designs that co-deploy escape-assisting elements as exploratory.
16. Modality XI — de novo Aptamer Scaffolds (Structure-Based Ligand Candidate Generation)
16.1 What is produced
A diverse pool of well-folding nucleic acid candidates ranked by computable structural fitness, or the evaluation and improvement of sequences already in hand. The output is not a therapeutic molecule but an input to an experiment — specifically, the starting library for an in-vitro selection campaign or the candidate list for a docking campaign.
This modality exists as a direct consequence of the policy in the previous chapter. If conjugate design depends only on validated aptamers, then nothing can be done for a target that has none. Filling that gap is this modality's role, and what it does is not 'find sequences that look like they will bind' but 'generate a diverse pool of sequences that fold well'.
16.2 Why structure only, and no affinity
As stated in 15.2.2, no reliable computational prediction of binding affinity to an arbitrary target exists. What, then, can computation do — the answer is the necessary condition for binding. A sequence that does not fold well cannot present a binding surface and therefore cannot bind. Folding well does not imply binding, but a pool composed of well-folding candidates has a higher success rate in selection experiments than one that is not.
That is this modality's honest position: computation does not replace experiment, it improves the experiment's starting point. Relative to beginning a selection campaign with a random library, beginning with a structurally pre-selected library reduces the number of selection rounds and the sequencing depth required.
16.3 Evaluation axes
|
Axis |
What it measures |
Why it matters |
|---|---|---|
|
Minimum free energy |
The free energy of the most stable secondary structure |
Stability of the fold |
|
Length-normalised minimum free energy |
The value divided by length |
Makes candidates of different lengths comparable |
|
Paired-base fraction |
The proportion of bases paired in the secondary structure |
Degree of structuring |
|
Stem and loop content |
The number and size of defined stem and loop elements |
Binding surfaces are usually loops, so their presence and size matter |
|
G-quadruplex propensity |
Guanosine run patterns and their arrangement |
A binding motif for some aptamers, and conversely a cause of non-specific binding |
|
GC balance |
Base composition |
Too stable and the structure cannot switch; too unstable and it does not fold |
|
Structural ensemble diversity |
Whether several metastable structures coexist |
Candidates locked into a single structure fare better in selection |
A weighted sum of these axes forms a structural fitness index used to rank candidates. The purpose, however, is not to pick the single highest-ranked sequence but to obtain a diverse set of high-ranking candidates, because selection experiments must begin from diversity. The output is therefore a pool distributed across structural classes rather than a single optimum.
16.4 Three operating modes
|
Mode |
Input |
Processing |
Output |
|---|---|---|---|
|
De novo generation |
Length range, GC constraints, target structural class, optional fixed primer flanks |
Generate candidates satisfying the constraints, fold and score each |
A ranked candidate pool distributed across structural classes |
|
Evaluation |
Candidate sequences already in hand (one or many) |
Fold each and score on the same axes |
Per-candidate scores, structure summaries and relative ranking |
|
Optimisation |
One seed sequence |
Deterministic search over point mutations, preserving length |
Variants with improved structural fitness and the improvement path |
The fixed primer flank option is the interface to experimental design. An in-vitro selection library carries fixed amplification sequences at both ends, and those sequences affect the folding of the variable region. Folding evaluated with the flanks included reflects the structure of the actual library molecule.
16.5 Chemistry and delivery requirements
At this stage chemical modification is optional. Fixing chemistry during candidate exploration shrinks the search space, and selection experiments are in any case usually run with unmodified or restricted-modification libraries. The customary order is to apply stabilising chemistry after the sequence has been fixed by selection, at which point the Modality X chemistry rules apply.
Delivery is not applicable. The output of this stage is a laboratory input, not something administered in vivo.
16.6 Design deliverables
- A candidate pool distributed across structural classes, in a format ready for synthesis ordering
- Per-candidate structural fitness index with a decomposition by axis
- Predicted secondary structure and its stability
- G-quadruplex propensity assessment
- Explicit statement of the generation constraints — the length, composition and structural constraints used
- Determinism guarantee — the same seed regenerates the same pool
- Handoff format to the next stage — selection experiment or conjugate design
16.7 Advanced capabilities — pool diversity and the experimental interface
16.7.1 A distributed pool, not a single optimum
The output objective of this modality differs from every other chapter. Elsewhere the goal is a small number of top candidates; here it is a structurally diverse candidate set. Selection experiments must start from diversity, and a pool of structurally similar candidates narrows the selection space so much that nothing may emerge.
After ranking, therefore, a distributed draw across structural classes is performed. Candidates are clustered by stem-loop count, loop size distribution, presence of branching and G-quadruplex propensity, and the top candidates from each cluster form the pool. A plain top-N draw lets one structural class dominate.
16.7.2 Handling fixed primer flanks
An in-vitro selection library carries fixed amplification sequences at both ends of the variable region. Those flanks affect the folding of the variable region, so a structure evaluated without them is not the structure of the actual library molecule. Where the user supplies the flank sequences, design evaluates folding on the complete molecule including them.
There is a useful side effect: candidates in which the flanks pair with the variable region to form unwanted structures are filtered out. Such candidates amplify inefficiently or fold inconsistently and obstruct the experiment.
16.7.3 Practical use of the three modes
|
Mode |
When to use it |
Where the output goes next |
|---|---|---|
|
De novo generation |
A target with no validated aptamer; starting a new selection campaign |
Synthesis order → in-vitro selection library |
|
Evaluation |
Checking the structural plausibility of sequences from selection, or reviewing literature sequences |
Candidate narrowing → binding experiments |
|
Optimisation |
Improving the fold stability of a promising sequence from selection |
Synthesis of a few variants → comparative binding experiments |
Optimisation mode searches point mutations while preserving length, so the user can fix positions that may participate in binding. Where structural information or experimental scanning results exist, fixing those positions and searching only the rest is the safe approach.
16.7.4 Determinism and regeneration
Randomness enters the generation process but the seed is fixed, so the same constraints and the same seed regenerate the same pool. This carries practical value for reproducing selection experiments and for regulatory response: which library was used is fully specified by the seed and the constraints, so the library can be regenerated without being archived.
16.8 What this modality cannot solve
- Prediction of binding affinity — this is both why the modality exists and what it does not do
- Target-specific design — design informed by the target protein's structure belongs to docking and related methods
- Replacing the selection experiment — computation improves the starting library and does not substitute for selection
- Interpretation of high-throughput selection data — analysis of post-selection sequencing data is outside the current scope
17. The Biology of Target Site Accessibility
Accessibility is the axis most often underweighted in design. Complementarity does not guarantee binding. RNA inside a cell is not a naked linear chain but a folded, protein-coated, dynamically traversed structure with ribosomes and helicases moving through it. This chapter covers what that real environment demands of design.
The importance is quantitative. Two fully complementary candidates commonly differ in measured potency by an order of magnitude, and much of that difference is explained by whether the target site is buried in secondary structure. Ranking on sequence rules alone misses it, and the result is that candidates which are 'perfect on paper but do not work' rise to the top.
17.1 RNA is folded — and not into one structure
An mRNA pairs with itself to form secondary structure of hairpins, internal loops, bulges and multibranch junctions. Crucially that structure is not unique. A given sequence interconverts among several structures of similar free energy, and inside the cell it exists as a Boltzmann distribution over them. The precise form of the question 'is this site open' is therefore 'with what probability is this site open'.
Design handles that probability two ways. The first is local-window accessibility: the probability that the target site is unpaired within a bounded window containing it. The window is bounded because in a real cell there is limited time and opportunity for long-range base pairing to form, and full-length folding tends to predict unrealistic long-range structures. The second is interaction free energy: the total energy change when the oligonucleotide binds its target.
ΔG_bind = ΔG_duplex − ΔG_target_unfold − ΔG_oligo_unfold
ΔG_duplex : energy of forming the oligo–target duplex (negative, favourable)
ΔG_target_unfold : cost of opening the target's existing structure (positive, unfavourable)
ΔG_oligo_unfold : cost of opening the oligonucleotide's own structure (positive, unfavourable)
→ With identical complementarity, strong target structure makes ΔG_bind worse.
→ A hairpin in the oligonucleotide itself is also a cost.
The design implication of this decomposition is clear: there are two ways to strengthen binding — make the duplex itself stronger (raise affinity chemically), or choose a site where less target structure has to be opened. The latter is far cheaper and carries no side effects, so accessibility scoring should precede chemistry decisions.
17.2 Sites covered by protein
An mRNA in a cell is never naked. From the moment of transcription, numerous RNA-binding proteins associate to form a ribonucleoprotein complex, and that complex governs the transcript's stability, localisation and translation. A site occupied by protein is physically closed to an oligonucleotide.
Crosslinking and immunoprecipitation experiments reveal where proteins actually bind at nucleotide resolution. Design uses that data as a mask to avoid occupied sites. The mask operates as a weighted penalty rather than an absolute exclusion, because binding profiles depend on cell type and condition: a peak observed in one dataset cannot be assumed present in the customer's experimental system. Penalising and recording the basis is more accurate than absolute exclusion.
Chemical modification is a factor at the same level. A given modification changes the base-pairing capacity and protein binding of its position, so regions dense in modification are more likely to have predicted and actual structure diverge. That is why a modification atlas is used as a mask.
17.3 The dynamic environment created by translation and transcription
Ribosomes traverse coding regions, and they carry potent helicase activity that continuously unwinds secondary structure along their path. This cuts both ways: coding-region structure may be more open than prediction suggests, and a bound oligonucleotide may be displaced by a passing ribosome.
The practical consequence is that regions differ in character. The 3' untranslated region is not traversed by ribosomes, so its structure is relatively stable and predictions hold better; it is also the principal arena for seed-mediated repression. The 5' untranslated region is occupied by the translation initiation complex, limiting access. The coding region varies in openness with ribosome density, and structure is unwound more often in highly expressed genes.
Different factors operate in the nucleus. Splicing occurs co-transcriptionally, so the window during which an intron exists is limited, and the exon junction complex occupies a region upstream of each junction. The constraint that intronic designs must act within the short window before removal arises here.
17.4 How structure acts differently in each modality
|
Modality |
How accessibility operates |
Design response |
|---|---|---|
|
Catalytic silencing (RISC) |
Structure slows the loaded complex's search, although the complex itself can open some structure |
Assess local accessibility probability together with interaction energy |
|
Catalytic silencing (RNase H1) |
A heteroduplex must form, so strong structure prevents binding outright; the catalytic complex's structure-opening capacity is limited |
Accessibility requirements are stricter; wing chemistry compensates on affinity |
|
Splice switching |
The target is fixed as a regulatory element, so there is almost no freedom in site choice; chemistry must open that site's structure |
Instead of changing site, secure binding strength chemically. ΔG_bind is the key metric |
|
Base editing |
A guide–target duplex must form, so target structure interferes; the duplex formed must also present the geometry the enzyme recognises |
Search a narrow window satisfying both binding and editing geometry |
|
Noncoding RNA silencing |
Targets frequently function through structure, so structure is especially strong; circular RNAs have a different topology and linear assumptions fail |
Compute accessibility on the correct topology; single-stranded mechanisms favour strongly structured targets |
|
miRNA mimic |
The target site is one already used by a miRNA, so accessibility is in effect pre-validated |
Use accessibility as one evidence axis in target prediction |
17.5 Limits of accessibility assessment
- Thermodynamic models do not fully capture intracellular ionic conditions, molecular crowding or local temperature variation. Predicted structures are closer to in-vitro conditions.
- Protein binding data are measured in particular cell lines and conditions and are not guaranteed to reproduce in a customer's system. The mask is a penalty, not an exclusion.
- Co-transcriptional folding can produce intermediate structures that differ from the final one, and equilibrium folding calculations do not capture this.
- Good accessibility does not prevent rejection on other axes — efficacy rules, off-targets, chemistry constraints. Accessibility is a necessary rather than a sufficient condition.
- Where experimental structure-probing data exist they take precedence over computational prediction, but such data exist for only a small number of transcripts.
18. The Biology of Off-Targets — What the Real Risk Is
'Off-target' is not one phenomenon but a bundle of phenomena with different mechanisms. Different mechanisms require different prediction methods, different mitigation strategies and different experimental confirmation. Collapsing them into a single off-target score produces a number that reflects none of the risks accurately. This chapter separates them.
18.1 Classification by mechanism
|
Mechanism |
What happens |
In which modalities |
Predictability |
|---|---|---|---|
|
Seed-mediated repression |
Guide positions 2–8 are complementary to another transcript's 3' untranslated region, causing miRNA-like repression |
Double- and single-stranded RISC mechanisms, miRNA mimics |
High — candidates enumerable from the seed sequence |
|
Partially complementary cleavage |
Binding short of full complementarity but sufficient for cleavage |
RISC mechanisms |
Moderate — central pairing strength governs 3' mismatch tolerance |
|
Heteroduplex formation |
A heteroduplex forms at a partially complementary site and the host nuclease cleaves |
Gapmers |
Moderate — depends on melting temperature and gap position |
|
Passenger strand loading |
The unintended strand loads and represses an entirely different target set |
All double-stranded modalities |
High — predicted from terminal free-energy asymmetry |
|
Protein-binding class effects |
The backbone binds plasma and cellular proteins, causing complement activation, coagulation effects and hepatic accumulation |
Every modality using phosphorothioate |
Moderate — depends on fraction and total load, with little sequence specificity |
|
Innate immune activation |
Sequence motifs stimulate endosomal receptors |
All modalities (an intended effect in the immune-modulation modality) |
Moderate — known motifs are predictable |
|
Bystander editing |
Adenosines other than the intended one are edited |
Base editing |
High — enumerable within the duplex footprint |
|
Global readthrough |
Normal stop codons are read through, extending C-termini |
Readthrough |
High — enumerable transcriptome-wide from codon and context |
|
Cross-miRNA inhibition |
Other miRNAs sharing the seed are co-sequestered |
Anti-miRs |
High — enumerable from seed-sharing relationships |
Mechanisms marked as highly predictable are those for which candidates can be enumerated; that is not the same as predicting the degree of repression accurately. Which of the enumerated candidates is actually repressed depends on expression level, site accessibility and competition, and requires experimental confirmation. Enumeration nonetheless has value, because the enumerated list defines what the experiment should look for.
18.2 Seed-mediated repression — the largest off-target class
In double-stranded silencing modalities most off-target effect is seed-mediated. A single seed 7-mer occurs in the 3' untranslated regions of hundreds of human transcripts, so a candidate with no full-length alignment hits can still carry a broad burden. This is what makes off-target assessment a question of mechanism rather than of alignment.
Repression strength is not uniform. It varies with seed type — 8mer, 7mer with 3' supplementary pairing, 7mer with an adenosine at position 1, 6mer — and with local context and position within the untranslated region. Design emits a burden index weighted by seed type, and quantifies the seed burden against essential genes separately, since repression of essential genes is the most likely to produce a phenotype.
18.2.1 Mitigation strategies
- Seed-weakening modification — place a sugar analogue at a specific seed position that locally weakens binding. The strategy exploits the asymmetry that the effect on a fully complementary on-target is small while the effect on seed-only binding is large.
- Seed selection — prefer candidates whose seed carries a low essential-gene burden at the enumeration stage. This mitigation costs nothing chemically.
- Passenger suppression — securing terminal asymmetry so the passenger does not load removes its seed burden at source.
- Dose minimisation — seed-mediated repression tends to require higher concentrations than on-target cleavage, so choosing a more potent candidate and lowering the dose is itself a mitigation.
18.3 Central pairing strength and the specificity paradox
The material from Chapter 3 is restated here from a specificity standpoint. The cleavage rate on a fully complementary target is not determined by central pairing strength; the source states explicitly that formation of a continuous helix does not limit the cleavage rate of fully complementary targets. What central pairing strength determines is tolerance to 3' mismatches.
Here the direction inverts. A guide with strong central pairing also cleaves partially complementary targets whose 3' end is mismatched. Central pairing strength is therefore an index of off-target cleavage capability rather than of on-target potency. The conventional advice to raise central GC content runs precisely backwards from a specificity standpoint.
The correct form of the axis is a conditional risk term: candidates with strong central pairing are flagged toward higher off-target cleavage risk, and the term is placed on the specificity axis rather than the on-target score. Structural work points the same way — expansion of the central major groove is required to position the scissile phosphate, and GC-rich duplexes resist distortion, so structure does not support a central GC bonus either.
18.4 Counting off-targets per gene or per site
When the same site exists in several isoforms of one gene, alignment reports several hits, but biologically there is one effect on one gene. Counting hits directly over-counts genes with many isoforms, and candidates overlapping such genes are unfairly disadvantaged.
Off-target tallies must therefore be presented at two levels: a per-site hit list (where does it bind) and a per-gene summary (how many genes are affected). Risk assessment and candidate comparison need the latter; selecting experimental confirmation targets needs the former.
The same logic applies on the on-target side. Hits against other isoforms of the target gene are coverage rather than off-targets, and counting them as off-targets silently penalises a perfectly good candidate. That distinction is why the first of the four off-target buckets exists.
18.5 How reference data affect off-target counts
Off-target counts are a function of the reference transcript set as much as of the algorithm. The more isoforms an annotation includes the higher the hit count, and including noncoding transcripts raises it further. Comparing off-target counts produced against different reference sets therefore does not hold.
Two practical rules follow. Compare candidates only among values produced with the same reference set and the same settings. And trust relative ranking and mechanistic classification more than absolute counts: 'seed burden against essential genes is in the top decile' is far more stable information than 'three off-targets'.
The design layer handles this two ways — recording the reference release in the run ledger so later comparison remains possible, and presenting per-gene tallies by default to reduce distortion from isoform counts.
18.6 Connection to experimental confirmation
The purpose of computational prediction is not to replace experiment but to define it. In off-target assessment that connection is especially direct.
|
Predicted deliverable |
Corresponding experiment |
What it confirms |
|---|---|---|
|
List of top seed-burden genes |
Expression measurement across a target gene panel |
Whether predicted seed-mediated repression actually occurs |
|
Essential-gene seed burden index |
Cell viability assessment |
Whether toxicity appears at the phenotype level |
|
List of partially complementary cleavage candidates |
Detection of cleavage products at those sites |
Whether predicted cleavage actually occurs |
|
Passenger loading prediction |
Strand-specific loading measurement |
Whether the intended strand is the one loaded |
|
Bystander editing list |
Sequence analysis of the target region |
The real rate of unintended editing |
|
Global readthrough burden list |
Protein size confirmation for high-risk genes |
Whether C-terminal extension actually occurs |
|
Immune-stimulatory motifs |
Cytokine measurement |
The real level of innate immune response |
When these correspondences are explicit, computational deliverables convert directly into an experimental plan. A single number such as 'off-target score 0.3', by contrast, does not tell anyone which experiment to run. That is the substantive reason for presenting deliverables separated by mechanism.
18.7 Limits of off-target assessment
- Alignment hit counts are not an index of risk. Counting hits without mechanistic classification causes good candidates to be rejected unfairly.
- What fraction of predicted off-targets is actually repressed depends on expression and accessibility and is not determined computationally.
- Protein-binding class effects have little sequence specificity and are not amenable to sequence-based prediction; they must be managed through chemical fraction and total exposure.
- Off-target profiles differ between species. A candidate safe in humans may carry a different burden in a test species, and vice versa.
- Transcripts absent from the reference set — unannotated transcripts, individual-specific variants — are omitted from assessment altogether.
19. Durability — What Determines the Dosing Interval
In the clinical competitiveness of an oligonucleotide therapeutic, the dosing interval matters as much as potency. A drug given twice a year and one given every two weeks are different products even at equal efficacy, differing in adherence, administration infrastructure and cost structure. This chapter decomposes where durability comes from.
One fact should be stated first: the duration of action of an oligonucleotide therapeutic is largely unrelated to its plasma half-life. It disappears from plasma within hours but persists in tissue for weeks to months, and it is the latter that determines the pharmacodynamic effect. Applying conventional pharmacokinetic metrics directly therefore grossly underestimates durability.
19.1 Four contributors to durability
|
Contributor |
Mechanism |
Controllable by design? |
|---|---|---|
|
Tissue depot formation |
Endocytosed molecules accumulate in endosomal and lysosomal compartments and are slowly released to the cytosol, acting as a depot |
Partly — delivery mode and chemistry affect the amount accumulated and the release rate |
|
Intracellular chemical stability |
A modified oligonucleotide resists intracellular nucleases and remains intact longer |
Yes — a direct object of chemical design |
|
Catalysis |
One molecule processes many targets, so effect persists at low residual concentration |
Partly — a consequence of mechanism choice |
|
Complex stability |
The loaded protein–nucleic acid complex is stable and functions for longer |
Partly — loading efficiency and terminal chemistry contribute |
The first is understood to contribute most. Most of what is endocytosed does not act immediately but resides in intracellular compartments, releasing slowly and acting continuously. The size of that depot is proportional to the administered dose and release is slow, so the effect of a single administration extends over months.
The design implication is counterintuitive. Extending durability requires not only a more stable molecule but one that accumulates better, and accumulation depends heavily on delivery mode. Receptor-mediated uptake concentrates accumulation in a specific cell type and is therefore favourable for durability — part of the reason receptor-conjugated products achieve long dosing intervals.
19.2 Durability characteristics by modality
|
Modality |
Principal source of durability |
Typical character |
Constraint |
|---|---|---|---|
|
Catalytic silencing (RISC) |
Depot plus catalysis plus complex stability |
The longest durability in this technology family; multi-month intervals reported for receptor-conjugated products |
Depot formation is tissue-dependent and may differ outside the liver |
|
Catalytic silencing (RNase H1) |
Depot plus catalysis |
Intermediate — intervals of weeks are typical |
Accumulation is comparatively dispersed on carrier-free dosing |
|
Splice switching |
Depot plus chemical stability |
Varies greatly with target tissue; long intervals achieved with local central nervous system dosing |
No catalysis, so it is stoichiometric and highly depot-dependent |
|
miRNA antagonism |
Depot plus binding stability |
Very stable binding that is effectively irreversible extends duration |
Stoichiometric |
|
Base editing |
Depot |
Edited transcripts disappear by natural turnover, giving a double decay |
The unmodified central window limits intracellular stability |
|
Readthrough |
Persistence of the expression cassette, or the stability of a synthetic molecule |
With cassette delivery, duration follows expression persistence |
Comparatively short when synthetic tRNA is dosed directly |
|
Immune modulation |
Not applicable — a transient stimulus is the objective |
Short action is preferable |
Repeated stimulation risks tolerance or chronic inflammation |
The 'double decay' of base editing deserves explanation. Editing occurs at the transcript level, so edited transcripts disappear with natural mRNA turnover while newly transcribed molecules are unedited. Duration is therefore the product of two time constants — persistence of the guide and transcript turnover rate — and targets with fast turnover require more frequent dosing. That same property is also the substance of the modality's reversibility as a safety feature.
19.3 How chemistry contributes to durability
|
Chemical element |
Contribution to durability |
Conflicting axis |
|---|---|---|
|
2'-modification coverage |
Raises intracellular nuclease resistance so intact molecules persist longer |
Excess lowers activity, depending on mechanism |
|
Phosphorothioate fraction |
Raises both stability and tissue accumulation |
Class toxicity burden; potency loss on encapsulated routes |
|
Terminal protection |
Blocks exonucleolytic erosion |
Almost none — a high-value, low-cost measure |
|
Receptor ligand conjugation |
Concentrates accumulation in a specific cell type, enlarging the depot |
Restricted to particular target tissues |
|
Lipophilic conjugation |
Increases membrane interaction, prolonging tissue residence |
May increase non-specific distribution |
|
5'-terminal stabilisation |
Contributes to the stability of the loaded complex |
Almost none |
The design layer combines these into a durability ranking and an expected dosing-interval band. That output is a prediction, and the real interval varies with species, tissue and dose and must be confirmed by pharmacokinetic and pharmacodynamic study. Using it for relative comparison among candidates is its correct use.
19.4 The relationship between durability and safety
Long durability is not purely an advantage. If an adverse effect appears, stopping administration does not stop the effect, so it is hard to reverse. That matters particularly where the target has an essential function, or where long-term safety data are lacking for a novel target.
Durability is therefore not a property to be maximised but one to be matched to target and indication. In maintenance therapy of chronic disease a long interval carries great value; in early clinical development of an unproven target, shorter action provides safety margin. This is why the design layer presents several chemistry alternatives together with a durability ranking — different choices are rational at different development stages.
Where reversibility is paramount the modality choice itself changes. Unlike approaches that permanently alter the genome, RNA-level intervention is reversible in principle, and among those, stoichiometrically acting modalities recover faster than catalytically acting ones.
19.5 Limits of durability prediction
- Depot formation and release kinetics differ by tissue and cell type, and the data underpinning computational prediction are limited.
- Cross-species extrapolation carries large uncertainty; durations observed in rodents frequently do not reproduce in humans.
- The expected dosing-interval band is for relative comparison among candidates and does not propose a clinical regimen.
- Where the target protein has a long half-life, lowering the transcript takes additional time to manifest at the protein level, so real pharmacodynamic duration may exceed the prediction.
- Accumulation on repeat dosing does not follow directly from a single-dose prediction and requires a separate model.
20. Combination Strategies Across Modalities
Real development programmes frequently do not end with one modality. Two modalities may be designed in parallel against the same target and compared experimentally; different targets may be addressed with different modalities simultaneously; and the output of one modality may become the payload of another. This chapter sets out the types of combination and the conditions under which each holds.
20.1 Type 1 — same target, two competing modalities
It is common for both catalytic silencing mechanisms to be viable for lowering the same gene. Which is better is not determined in advance: target site accessibility, subcellular localisation, delivery route and tissue all shift the balance. Parallel design followed by experimental comparison is the rational approach, and there are cases in which products using both mechanisms against the same target have each been approved.
For that comparison to be fair, both designs must share one coordinate system and one off-target definition. Designed with different tools, '3 off-targets' becomes three things under two definitions and the comparison is meaningless. Designed on a common foundation, the metrics in the two reports carry the same meaning and the comparison becomes a basis for decision.
|
Comparison axis |
Catalytic silencing (RISC) |
Catalytic silencing (RNase H1) |
Deciding criterion |
|---|---|---|---|
|
Site of action |
Principally cytoplasmic |
Both nucleus and cytoplasm |
Nuclear-retained targets favour the latter |
|
Delivery burden |
Vehicle mandatory |
Carrier-free route available |
Without capacity for vehicle development, the latter |
|
Required dose |
Low (catalysis plus high potency) |
Comparatively high |
Considering dose-related toxicity margin, the former |
|
Principal toxicity risk |
Seed-mediated off-targets |
Hepatocyte toxicity |
Depends on target tissue and concomitant medication |
|
Freedom in site choice |
Form largely fixed by geometry |
Length, gap and wings adjustable |
Difficult target structure favours the latter |
|
Durability |
The longest in this technology family |
Intermediate |
Chronic maintenance therapy favours the former |
20.2 Type 2 — simultaneous silencing of multiple targets
Lowering a single gene sometimes produces no phenotype, because functionally redundant paralogs compensate or because the disease arises from the sum of several pathways. Two options then exist: find one molecule that silences all the targets, or combine molecules each targeting one.
The former means finding a conserved region within the family so that one molecule silences several members; this is what a co-silencing instruction expresses in family selectivity adjudication. It requires a sufficiently conserved site to exist, and where none does, the latter route follows. The latter combines several molecules, so doses and toxicities accumulate and the safety burden grows.
An important design point is that the off-target burden of a combination is not a simple sum. Two molecules with different seeds double the seed-mediated burden, and sharing chemistry means backbone-related class effects rise with total exposure. In combination design, therefore, the objective is not optimising each candidate but minimising the burden of the combination as a whole.
20.3 Type 3 — combining modalities that act in opposite directions
Where a disease mechanism is 'too much of one thing and too little of another', combining silencing with activation is logically coherent — lowering a pathogenic factor while raising a protective one.
In practice the combination is difficult. Activation carries greater predictive uncertainty, and for two molecules to reach the same cell they must travel in the same vehicle, which is awkward when their form and chemistry differ. The first option considered is therefore the detour route: silencing the repressive axis of the gene to be raised — a natural antisense transcript or an inhibitory regulator — achieves both objectives with silencing modalities, at which point both molecules share form and chemistry and co-formulation becomes straightforward.
20.4 Type 4 — combining sequence correction with expression control
In premature-stop-codon disease, two different interventions can be complementary. Even if a readthrough molecule allows the stop to be passed, there is nothing to read if the transcript has already been degraded by nonsense-mediated decay. Combining decay inhibition with readthrough induction is therefore under discussion.
Similarly, splice switching and silencing can be combined: silencing a pathogenic isoform while increasing production of the normal isoform by splice manipulation. Here both molecules are single-stranded with similar chemistry, so the technical burden of combination is comparatively low.
The relationship between base editing and other modalities is both competitive and complementary. For the same pathogenic variant, readthrough and editing may each be applicable, and which covers more patients can only be established by decomposing that gene's variant spectrum by type. This is why the target adjudication layer emits variant-type proportions.
20.5 Type 5 — one modality as the payload of another
A targeted conjugate produces no therapeutic effect itself but carries something else. When the payload is a silencing molecule, two modalities are fused into one construct, and because the conjugate substitutes for a delivery vehicle, silencing outside the liver becomes possible in principle.
Designing this combination must satisfy both modalities' requirements simultaneously. The silencing payload must carry the chemistry its mechanism demands; the conjugate must preserve the aptamer's fold; and the linker must be long enough that the two moieties do not interfere yet not so large as to accelerate renal excretion. Because the requirements conflict, the construct must be evaluated in its combined state rather than optimised independently and then joined.
The remaining bottleneck is endosomal escape. Even where the conjugate enters the cell, the silencing payload does nothing if it does not reach the cytosol. This is the main reason the combination remains of limited practical use, and from a design standpoint it means elements that assist escape must be considered alongside.
20.6 Common principles in combination design
- The burden of a combination can exceed the sum of its parts. Shared chemistry accumulates backbone-related class effects with total exposure, and a shared vehicle accumulates vehicle-related responses.
- Comparison requires a common foundation. Metrics produced under different definitions are not a basis for comparison.
- Consider detour routes first. Achieving the same objective with a mature modality substantially lowers development risk.
- Co-formulability governs practicality. Combining molecules of similar form and chemistry is technically far easier.
- Differing durability among components changes the ratio over time, and that mismatch can create unintended effects.
20.7 Deliverables of combination design
- The complete specification of each component molecule — sequence, chemistry, delivery
- Combination-level off-target burden — tallied for overlap and accumulation rather than summed naively
- Total load of shared chemical elements — combination-level assessment of backbone-related burden
- Co-formulability review — compatibility of form, charge and stability requirements
- Comparison of durability profiles — mismatch in duration of action among components
- Alternative route proposals — the design for a detour that achieves the same objective, where one exists
21. The Design Specification Seen from Synthesis and Quality
The ultimate use of a design deliverable is a synthesis order. It therefore matters whether the specification is complete from a synthesis standpoint, and what quality characteristics the material it produces will have. This chapter covers how sequence and chemistry affect synthesis and quality.
21.1 What it means for a specification to be complete
A synthesis order requires more than a sequence string. All the items below must be fixed for the specification to hold; where any is missing, the synthesis provider either interprets it or asks.
|
Item |
What it specifies |
What happens when it is blank |
|---|---|---|
|
Strand composition |
Single- or double-stranded, and for a duplex each strand individually |
A duplex is mistaken for a single strand or vice versa |
|
Base sequence |
The complete sequence in 5' to 3' orientation |
A misread orientation produces an entirely different molecule |
|
Sugar modification map |
The sugar identity at every position (natural ribose, 2'-deoxy, 2'-O-methyl, 2'-fluoro and so on) |
A summary such as '60% 2'-O-methyl' cannot be synthesised |
|
Backbone map |
The identity of every linkage (phosphodiester or phosphorothioate) |
Arbitrarily interpreted as fully phosphorothioate, changing the properties |
|
Terminal handling |
The chemistry at the 5' and 3' termini (phosphate, phosphate mimic, inverted residue, none) |
The presence of a terminal phosphate governs activity itself in some modalities |
|
Conjugate |
Attachment position, ligand class, linker chemistry |
A different attachment point changes targeting and folding |
|
Overhang convention |
The terminal form of a duplex and the overhang bases |
Artificial and target-derived overhangs are confused |
|
Stereocontrol |
Whether phosphorothioate stereochemistry is controlled |
The definition of the substance and its quality items both change |
|
Purification grade |
The required purity level |
Research-use and preclinical requirements differ |
This is why deliverables express chemistry as per-position strings. A summary statistic is not a specification, and it cannot be used for patent review either, because most chemistry claims are constructed as 'a specific modification at a specific position'.
21.2 How sequence and chemistry affect synthesis difficulty
|
Factor |
Effect on synthesis |
Response at the design stage |
|---|---|---|
|
Length |
Cumulative step yield lowers the proportion of full-length product |
Prefer shorter designs within what the mechanism allows |
|
Guanosine runs |
Higher-order structure formation lowers coupling efficiency and complicates purification |
Apply a run-length cap at enumeration |
|
Strong self-structure |
Aggregation during synthesis and purification |
Use self-secondary-structure as an evaluation axis |
|
Phosphorothioate fraction |
Adds a sulfurisation step and produces a diastereomer mixture |
Minimise within the route-specific productive window |
|
Bicyclic modifications |
Higher monomer cost and demanding coupling conditions |
Optimise placement so they are used only where needed |
|
Conjugates |
Add a conjugation reaction and subsequent purification steps |
Restrict attachment to the termini |
|
Double-stranded format |
Two strands must each be synthesised and purified, then annealed and verified |
Recognise in advance that quality items double |
These factors affect not only cost and timeline but the quality of the material. Length and phosphorothioate fraction in particular determine the impurity profile directly.
21.3 Expected impurity classes
Impurities in solid-phase synthesis fall into defined classes and are to some extent predictable from sequence and chemistry. Surfacing them at design time brings forward purification strategy and analytical method development.
|
Impurity class |
Mechanism |
Relationship to sequence and chemistry |
|---|---|---|
|
Deletion sequences (n−1, n−2 …) |
Chains missing a unit through coupling failure |
Increase with length and with the number of difficult coupling positions |
|
Addition sequences (n+1) |
Double coupling |
Increase with particular monomers and conditions |
|
Incomplete deprotection |
Chains retaining a protecting group |
Increase with bulky modifications |
|
Phosphodiester incorporation |
Incomplete sulfurisation leaves an intended phosphorothioate as a phosphodiester |
Proportional to the number of phosphorothioate linkages |
|
Diastereomers |
Stereochemistry at the phosphorothioate phosphorus |
2ⁿ for n linkages — intrinsically a mixture unless stereochemistry is controlled |
|
Depurination products |
Loss of a purine base under acidic conditions |
Depends on purine content and processing conditions |
|
Oxidation and dimers |
Oxidation during storage, thiol-related dimerisation |
Observed with particular conjugates and terminal chemistries |
The diastereomer entry warrants separate explanation. Without stereocontrol, one 'substance' is in reality a mixture of thousands to millions of isomers that behave differently in nuclease resistance and target binding. This is less an impurity question than a question of what the substance is, and from a regulatory standpoint the issue is demonstrating that the mixture is reproduced consistently. Adopting a stereocontrolled synthesis clarifies the definition of the substance at a substantial increase in cost and complexity. This is why the deliverables state explicitly whether stereocontrol is assumed.
21.4 Quality items specific to duplexes
- Per-strand purity — purity must be secured for each strand individually
- Annealing efficiency — the proportion of strands that actually formed the duplex
- Stoichiometry — whether the molar ratio of the two strands is close to 1:1
- Residual single strands — the amount of unannealed strand remaining
- Duplex stability — whether the duplex is maintained under storage conditions
- Confirmation of the overhang convention — whether the designed overhang is present in the actual material, and specifically whether a target-derived overhang has been replaced by an artificial one
The last item is a recurring practical problem. If the overhang convention is not stated in the specification, a conventional notation may be applied, and the molecule evaluated at design time then differs from the molecule synthesised. This is why the acceptance test is full-length reverse complementarity rather than length, and why it is verified automatically at export.
21.5 Quality-related information emitted at design time
- Per-position chemistry map in a form usable directly in a synthesis order
- Expected impurity profile — impurity classes and their relative prominence derived from sequence and chemistry
- Synthesis difficulty indicators — length, guanosine runs, self-structure, use of special monomers
- Whether stereocontrol is assumed and the resulting definition of the substance
- For duplexes, annealing-related items and the overhang convention
- For conjugates, attachment position and linker, and the heterogeneity expected after conjugation
- Storage and handling notes — stability issues expected for particular chemistries
The impurity information emitted here is a design-time prediction; setting actual specifications and validating analytical methods belong to the manufacturing stage. Its purpose is to prepare purification strategy and analytical development in advance, and to avoid at design time chemical choices that would become quality problems.
21.6 The handover between design and manufacturing
Information is lost between design and manufacturing at predictable points. First, summarisation of chemistry notation — per-position information compressed into statistics cannot be restored. Second, loss of design rationale — if why a modification was placed at a position is not conveyed, then when manufacturing changes it for synthetic convenience the effect on activity is unknown. Third, loss of alternatives — if the chemistry alternatives considered at design time are not conveyed, difficulty in synthesis forces a return to design.
Deliverables carry the per-position map, the selection rationale and the alternative comparison together to prevent all three. Because the run ledger records the parameters and reference data of the moment, a later return to design from manufacturing can be re-run under identical conditions.
22. A Disease-Area Perspective
Modality choice is not determined by the target's molecular properties alone. Which tissue the target sits in, what route reaches that tissue, and what development precedent has accumulated in that disease area all bear on it. This chapter sets out where this technology family stands in each major disease area.
The reason for organising by area is practical. 'I want to lower this gene' becomes an entirely different programme depending on whether the gene sits in the liver or the brain. The former has a mature path; the latter makes the route of administration itself a development item.
22.1 Liver — the most mature area
The liver is where approvals in this technology family are overwhelmingly concentrated, for biological reasons. Hepatocytes express the asialoglycoprotein receptor at very high surface density and it recycles rapidly, so a molecule bearing a triantennary galactosamine ligand concentrates efficiently in the liver after subcutaneous dosing. At the same time the liver is where particles naturally accumulate from systemic circulation, so the lipid particle route also works without separate targeting.
|
Disease type |
Suitable modalities |
Why |
|---|---|---|
|
Protein excess or misfolded accumulation |
Catalytic silencing (both mechanisms) |
Reducing production is the direct remedy; both mechanisms have approval precedent |
|
Metabolic enzyme deficiency |
Transcriptional activation, or silencing of the repressive axis |
The direction is upward; the silencing detour is the first option to examine |
|
Lipid metabolism disorders |
Catalytic silencing |
Lowering a regulator is an established approach |
|
Hepatic miRNA axes |
Anti-miR |
Both carrier-free uptake and receptor conjugation work |
|
Nonsense-variant metabolic disease |
Readthrough or base editing |
Delivery is comparatively easy, favouring validation of newer modalities |
The strategic implication is clear. Validating a new modality or new chemistry against a hepatic target controls for the delivery variable and allows the modality itself to be assessed. Validating a new modality outside the liver leaves failure indistinguishable between a modality problem and a delivery problem.
22.2 Central nervous system — the route is the solution
The brain and spinal cord are separated from the systemic circulation by the blood–brain barrier, so intravenously dosed oligonucleotides effectively do not reach them. What solved this was not a vehicle but a route: intrathecal administration distributes through the cerebrospinal fluid, and several products have been approved on that route.
Single-stranded modalities have an advantage here for specific reasons. Nuclease activity in cerebrospinal fluid is comparatively low, and phosphorothioate single strands distribute well through neural tissue and persist. Because dosing is local, systemic exposure is low and constraints from backbone class toxicity relax. Double strands, by contrast, require a vehicle, and intrathecal particle dosing brings additional considerations.
|
Disease type |
Suitable modalities |
Considerations |
|---|---|---|
|
Toxic gain-of-function variants in motor neuron disease |
Catalytic silencing (RNase H1), allele-selective design |
Selectively lowering only the mutant allele provides safety margin |
|
Splicing-defective neuromuscular disease |
Splice switching |
The flagship success of this area; local dosing manages the burden of repeat administration |
|
Repeat expansion disease |
Catalytic silencing or splice manipulation |
The repeat structure creates difficulty in both accessibility and specificity |
|
Neurodegenerative protein accumulation |
Catalytic silencing |
Where the target has a normal function, partial reduction becomes the objective |
|
Nonsense-variant neurological disease |
Readthrough or base editing |
The delivery burden is large and vehicle development runs in parallel |
The practical constraint is the invasiveness of the procedure and the resulting limit on dosing frequency. Durability therefore carries particular value in this area, and chemical design prioritises intracellular stability and tissue residence.
22.3 Eye — the advantage of local administration
The eye is a small, isolated compartment, so local administration achieves high local concentration while minimising systemic exposure. That several early approvals in this technology family are ocular is not a coincidence, and both aptamer approvals fall in this area.
Three characteristics define the area for design. Systemic toxicity constraints relax, widening the range of chemistry available. Administration is invasive, so durability matters. And diffusion behaviour differs by target tissue — retina, cornea — so vitreal diffusion and residence time become design considerations.
- Pseudoexons arising from deep-intronic variants — a flagship application of splice switching, with a comparatively low development burden when combined with local dosing
- Angiogenesis-related factors — aptamers or silencing modalities
- Complement pathway factors — an aptamer approval precedent exists
- Inherited retinal disease with diverse variant types — readthrough, editing or splice manipulation, depending on variant type
22.4 Muscle and heart — where delivery is the bottleneck
Muscle constitutes a large fraction of body mass, so obtaining adequate tissue concentration by systemic dosing requires very large amounts. That is why approved products in this area use high doses, and equally why unmet need remains large.
Improvement is being attempted in two directions: conjugating cell-penetrating peptides, and attaching ligands that target muscle cell surface receptors. From a design standpoint both require assessing the effect of conjugation on folding, binding and distribution, with attachment point and linker as new design variables.
The heart has delivery problems similar to muscle with a narrower safety margin. Unexpected effects on cardiac muscle are immediate and hard to reverse, so off-target and essentiality assessment are demanded particularly rigorously in this area.
22.5 Lung and airway — the possibility of the inhaled route
Inhalation reaches the target tissue directly while lowering systemic exposure. It fits logically for diseases targeting airway epithelium — cystic fibrosis, chronic airway disease, respiratory infection.
Three design challenges follow. Shear during aerosolisation can destroy particles, so formulation rigidity matters. Airway mucus acts as a physical barrier, making mucus penetration a design variable. And much of what is inhaled is removed by mucociliary clearance, so residence time governs efficacy.
In infectious disease response the route has strategic value. Targeting host factors required by respiratory viruses works independently of viral sequence variation, so resistance is hard to develop and effect against several viruses is possible. The counterweight is the safety burden of suppressing a host gene, which makes essentiality assessment of the target decisive.
22.6 Tumours — targeted delivery is the crux
Tumours are among the hardest areas for this technology family. Tumour tissue has irregular vasculature and high interstitial pressure, limiting molecular penetration, and tumour cells are heterogeneous, so a single target rarely addresses the whole. At the same time selectivity against normal tissue is difficult to secure.
|
Approach |
Modality |
Rationale and constraint |
|---|---|---|
|
Silencing tumour-dependency genes |
Catalytic silencing |
Essentiality must be tumour-specific; the cell-dependency axis of target adjudication is central |
|
Targeted delivery |
Aptamer conjugation (cytotoxin or oligonucleotide payload) |
Works where a tumour surface target exists; constrained by the availability of a validated aptamer |
|
Local immune stimulation |
Innate immune receptor agonists |
Intratumoral dosing to reprogramme the microenvironment; minimising systemic exposure is the crux |
|
Restoring tumour suppressors |
Transcriptional activation or silencing of the repressive axis |
The direction is upward; candidate genes are numerous, so combination strategies are discussed |
|
Rebalancing a miRNA axis |
Mimicry or antagonism |
Network-level intervention with a heavy safety assessment burden |
This is where the target adjudication layer shows its greatest value, because tumour dependency and normal-tissue essentiality must be assessed together, synthetic-lethal relationships and adverse-event precedent must be read alongside, and redirection to a substitute target is needed when the target proves unsuitable.
22.7 Infectious disease and vaccines — conservation and host factors
In targeting a pathogen directly, the central problem is variation. Targeting a single site in a rapidly varying pathogen selects resistance variants quickly, so regions conserved across strains and subtypes must be targeted. Identifying conserved regions from a multi-strain sequence set and restricting design to them is the corresponding approach.
The alternative is targeting host factors. Lowering a host gene the pathogen requires for replication works independently of the pathogen's sequence variation. The attraction is breadth and resistance robustness; the burden is the safety cost of suppressing a host gene. Essentiality assessment is therefore decisive in target selection, and targets that are tissue-restricted in expression, or that need suppressing only during the acute phase, are preferred.
In vaccines the most mature application is as an adjuvant. Innate immune receptor agonists have approval precedent, and the type of immune response induced can be selected by structural class. This has particular value for pathogens where cell-mediated immunity is protective and for populations with weak responses.
Veterinary and zoonotic applications belong in this context. Infectious disease management in livestock and poultry is both a market in itself and connected to human pandemic risk management, with comparatively shorter regulatory paths. What matters is designing against that species' own transcriptome and corresponding genes rather than repurposing a human product.
22.8 Rare genetic disease — variant type determines the modality
In rare genetic disease, modality selection begins from the variant type rather than the gene, because patients with the same gene disorder carry different variant types and different interventions apply to each.
|
Variant type |
Applicable modalities |
What determines the patient fraction covered |
|---|---|---|
|
Premature stop codon |
Readthrough |
The share of stop-gained variants in that gene |
|
Splice site variant |
Splice switching |
The share of splicing-related pathogenic variants |
|
Deep-intronic variant creating a pseudoexon |
Splice switching (pseudoexon suppression) |
How frequently that variant is reported |
|
G-to-A missense |
Base editing |
The share of that transition type |
|
Frameshift deletion |
Splice switching (exon skipping) |
The share of deletions for which skipping restores the frame |
|
Dominant-negative variant |
Allele-selective silencing |
Whether a sequence difference exists to distinguish the mutant allele |
|
Loss of function (haploinsufficiency) |
Transcriptional activation or silencing of the repressive axis |
Headroom for raising the normal allele |
The development strategy this table implies is patient stratification. No single molecule covers every patient, so quantifying the proportion of each variant type first and developing the intervention that covers the most patients is the rational order. This is why the target adjudication layer decomposes the pathogenic variant spectrum by type — that decomposition is the basis for development sequencing.
22.9 Difficulty by area, summarised
|
Area |
Delivery difficulty |
Safety margin |
Accumulated precedent |
Overall |
|---|---|---|---|---|
|
Liver |
Low — two mature routes |
Moderate |
Very deep |
Optimal for validating new modalities |
|
Central nervous system |
Moderate — solved by route |
Relatively wide (low systemic exposure) |
Deep |
Durability matters because of procedural invasiveness |
|
Eye |
Low — local administration |
Wide |
Moderate |
Favourable for early validation |
|
Muscle and heart |
High — the current bottleneck |
Narrow for the heart |
Moderate for muscle, thin for heart |
Delivery improvement is the crux |
|
Lung and airway |
Moderate — the inhaled route |
Moderate |
Thin |
Formulation and residence time are the challenges |
|
Tumours |
High — penetration and selectivity |
Narrow |
Thin |
Where target adjudication is most valuable |
|
Infectious disease |
Varies with target |
Narrow for host targets |
Moderate (deep for adjuvants) |
Conservation or host essentiality assessment is central |
23. A Drug and Vaccine Development Perspective
The structural characteristic of oligonucleotide therapeutics in drug development is that the design space is finite and computationally compressible. Once the target gene is fixed, the design space is bounded by the subsequences of that transcript, and efficacy, selectivity, accessibility, chemistry and cross-species conservation are all properties that can be evaluated in advance. This differs fundamentally from small molecules, which must search a vast chemical space, and from antibodies, which require immunisation and selection.
The practical difference that compressibility makes is in the time and cost of lead discovery. From a fixed target, the path to a set of synthesisable candidate sequences is short and mostly computational. The burdens of the later stages — delivery, chemistry-manufacturing-and-control, toxicology — are, by contrast, no lighter than for other modalities and in places add items of their own.
23.1 Deliverables mapped to development stages
|
Stage |
Central question |
What this layer provides |
What carries into the next stage |
|---|---|---|---|
|
Target validation |
Is intervening on this gene correct, and in which direction |
Modality suitability verdicts; safety, expression, disease and network evidence; substitute target proposals |
A fixed target and modality, and family selectivity instructions |
|
Lead discovery |
Which candidates should be synthesised |
Candidate sequences and ranking, off-target classification, cleavage coverage, accessibility |
Synthesis specifications for 10–30 candidates |
|
Lead optimisation |
How should chemistry and delivery be fixed |
Per-position chemistry maps, alternatives per delivery route, durability ranking, class toxicity risk classification |
Complete specifications for a small number of leads |
|
Preclinical |
In which species should it be evaluated, and what is risky |
Per-species-combination comparison of surviving candidates, two-species verdict, basis for off-target experiment design |
Documented rationale for toxicology species selection |
|
Regulatory preparation |
How should the filing package be assembled |
Run ledger, synthesis impurity projection, reference version verification records |
Machine-readable supporting material |
|
Chemistry, manufacturing and controls |
What will become a problem in formulation and process |
Formulation candidates and property predictions, stability, process and container risk item assessments |
The starting point for formulation development and a risk list |
Species selection at the preclinical stage is regulatorily significant. Safety assessment in one rodent and one non-rodent species is customarily expected before human trials, and for oligonucleotide therapeutics target-related toxicity can only be assessed if the molecule also binds the corresponding transcript in that species. Data from a species in which it does not bind show chemistry class toxicity and say nothing about target-related risk. Presenting the surviving candidate set per species combination is how that requirement is evidenced with data.
23.2 The decision flow for modality choice
What must be decided first in a real programme is not the sequence but the modality. Below is the flow from a target's properties to a modality.
What is the therapeutic direction?
├ Down
│ ├ target is mature cytoplasmic mRNA -> catalytic silencing (RISC). vehicle required
│ ├ target is nuclear-retained / pre-mRNA -> catalytic silencing (RNase H1). carrier-free possible
│ ├ target is a noncoding RNA -> class call, then one of the above
│ └ target is a miRNA itself -> miRNA antagonism. carrier-free possible
├ Up
│ ├ the locus is activatable -> transcriptional activation. vehicle required
│ ├ a repressive axis exists -> silence that axis (detour; better prediction accuracy)
│ └ restoring a lost miRNA function -> miRNA mimicry. vehicle required
├ Repair
│ ├ pathogenic variant is G>A -> ADAR-recruiting editing. vehicle required
│ ├ pathogenic variant is a premature stop-> suppressor tRNA. vehicle required
│ └ mechanism runs through splicing -> splice switching. vehicle usually unnecessary
└ Move the immune system / carry something else
├ innate immune stimulation or suppression -> TLR modulation. no vehicle needed
└ cell-type-specific delivery -> aptamer conjugation. it is the vehicle
The branch most often overlooked is the second under 'Up'. Activation carries greater predictive uncertainty and thinner clinical precedent, whereas silencing a repressive axis uses a mature silencing modality and therefore carries lower development risk. If the same therapeutic objective can be reached by a better-validated means, that route is preferable — and reaching that judgement requires mapping antisense transcripts around the target locus.
23.3 Applications in vaccine development
23.3.1 Oligonucleotides as adjuvants
A synthetic oligonucleotide has obtained regulatory approval as a vaccine adjuvant, and this is the application with the broadest human safety dataset in the technology family. Adjuvants are dosed in large numbers to healthy adults, so the scale of safety evidence required for approval differs from that for a therapeutic.
From a design standpoint the essential feature of an adjuvant oligonucleotide is selecting the type of immune response. Traditional aluminium salt adjuvants induce mainly antibody responses and weak cell-mediated immunity. An innate immune receptor agonist can selectively strengthen either the interferon pathway or the B cell pathway according to structural class, so the response type can be designed to suit the pathogen. The value is greatest for pathogens where cell-mediated immunity is protective and in populations, such as the elderly, whose responses are weak.
23.3.2 Antigen-expressing platforms share the delivery layer
Messenger RNA vaccines differ from the silencing and correction modalities in the molecule itself but share the delivery layer. Ionisable lipid particles, biodistribution driven by the protein corona, endosomal escape efficiency, cold-chain distribution and lyophilisation are the same formulation problems. The delivery design layer therefore applies to vaccine payloads unchanged, which is why messenger RNA, self-amplifying RNA and circular RNA cargo families appear in the formulation variant catalogue.
Self-amplifying RNA co-encodes a replicase and amplifies itself inside the cell, lowering the required dose; circular RNA has no ends and resists exonucleases, extending expression. Both change the payload's properties and therefore the particle design requirements, and the design layer applies different structural branches per cargo type.
23.3.3 Designing on conserved regions for infectious targets
Targeting a rapidly varying pathogen requires targeting regions conserved across strains and subtypes in order to achieve both breadth of efficacy and resistance suppression. The corresponding approach identifies low-variation regions from a multi-strain sequence set and restricts design to them, and the design layer performs that analysis on a user-supplied multi-strain sequence set.
A host-factor approach also exists. Lowering a host gene the virus requires for replication works independently of viral sequence variation, so resistance is difficult to develop and effect across several viruses is possible. The counterweight is the safety burden of suppressing a host gene, so essentiality assessment of the target is decisive — this is the area in which the safety axis of the target adjudication layer matters most.
23.3.4 Veterinary and zoonotic applications
Infectious disease management in livestock and poultry is both a market in its own right and connected to management of human pandemic risk: for pathogens with zoonotic transmission potential, control at the animal stage lowers human risk. Regulatory paths are shorter than for human medicines and areas of substantial unmet need exist, so species-specific design carries real value. What matters is designing against that species' transcriptome and orthologs rather than repurposing a human product, and per-species design modes support this.
23.4 Where development risk concentrates, by modality
Risk concentrates at a different stage in each modality. Knowing that distribution changes resource allocation and milestone design.
|
Modality |
Where risk concentrates |
Why |
Mitigation |
|---|---|---|---|
|
Catalytic silencing (RISC) |
Delivery (non-hepatic targets) |
The liver has a mature answer; elsewhere is the industry-wide bottleneck |
Start with hepatic targets, or concentrate resource on vehicle exploration |
|
Catalytic silencing (RNase H1) |
Toxicity, especially hepatocyte toxicity |
Potency and toxicity risk arise from the same properties |
Evaluate the toxicity axis in parallel with potency at design time; confirm experimentally early |
|
Splice switching |
Target tissue concentration |
Dosed without a carrier, so tissue delivery depends on dose |
Prioritise local routes; consider cell-penetrating conjugation |
|
Noncoding RNA silencing |
Target validation |
Functional regions are frequently uncharacterised |
State functional mapping as prerequisite work |
|
Transcriptional activation |
Prediction accuracy |
Strong chromatin-state dependence and poor transfer between cell types |
Run positive controls in parallel; examine the silencing detour route |
|
miRNA modulation |
Safety assessment |
Many targets move at once, leaving scope for unexpected phenotypes |
Confirm target-set convergence in advance; escalate dose in steps |
|
Base editing |
Conflict between efficiency and stability |
The unmodified central window constrains chemical protection |
Compare window widths; strengthen carrier protection |
|
Readthrough |
Global burden |
Effects on normal stop codons exist in principle |
Exploit codon and context selectivity; quantify global burden and monitor high-risk genes |
|
Immune modulation |
Cross-species extrapolation |
Species specificity limits translation of animal data |
Assess per-species optimal motifs; run human cell-based assessment in parallel |
|
Aptamer conjugation |
Applicable scope |
Restricted to targets with a validated aptamer |
Generate candidate pools and run selection in parallel |
23.5 What the computational stage actually saves
Stated without exaggeration: computation does not replace wet-lab work. What it does is reduce the number of candidates that must be validated, and that reduction becomes a reduction in time and cost. Validating a few dozen candidates instead of synthesising and screening several hundred cuts synthesis cost, reagents, personnel and elapsed time in proportion.
The second value is the elimination of handover losses between stages. When target adjudication, design, chemistry, delivery and regulatory preparation are dispersed across different tools and manual analysis, information is lost at each handover and reproducibility degrades. Performed continuously under one coordinate system and one provenance record, those losses do not occur, and the effect is cumulatively larger than the speed gain of any individual stage.
The third value is traceability of decisions. When why a candidate was selected, and which database versions and parameters were in force at that moment, are recorded automatically, the interpretation of later experimental results and the regulatory response both rest on that record. This is risk reduction rather than cost saving, and its value becomes visible when a problem arises at a late stage.
What computation does not do is stated equally plainly. Predictive scores rank; they do not guarantee absolute efficacy. Toxicity risk classification does not replace experimental toxicology. Delivery prediction does not replace measured biodistribution. The role of computation is to make the experiment smaller and better aimed.
24. Benchmarking Against Industry Standards
This chapter compares the capabilities and performance of this design layer against public tools and industry practice. Comparison is made at two levels: a qualitative comparison of whether a capability axis exists, and a quantitative comparison of predictive performance measured on public benchmark datasets. The scope in which quantitative comparison is possible is limited, so that boundary is stated first.
The boundary of comparison — quantitative comparison of predictive performance is confined to the area where a public benchmark exists, namely efficacy prediction for double-stranded silencing. For other modalities there is no widely adopted public benchmark dataset, so this document presents no quantitative figures there and offers only capability-axis comparison. Constructing a benchmark that does not exist in order to present one's own performance is not a comparison.
24.1 Double-stranded silencing efficacy prediction — quantitative comparison
The standard benchmark for efficacy prediction models is a large gene-silencing library dataset. Measured on the same dataset (n = 2,431), the coefficient of determination, correlation coefficient, and the area under the curve that expresses practical discrimination compare as follows.
|
Generation |
Model / method |
R² |
Correlation |
AUC@0.7 |
|---|---|---|---|---|
|
First-generation rules |
Positional rule-based (2004 family) |
~0.10 |
0.32 |
0.63 |
|
Second-generation statistics |
Linear kernel-based (2006 family) |
~0.18 |
0.42 |
0.69 |
|
Third-generation deep |
Transformer with RNA language-model embeddings (reported 2024) |
0.34 |
0.58 |
0.78 |
|
Third-generation deep |
Transformer combined with convolution (reported 2025) |
0.31 |
0.55 |
0.77 |
|
Third-generation deep |
Chemistry-aware multi-view with cross-attention (reported 2024) |
0.40 |
0.60 |
0.80 |
|
This layer |
Chemistry-aware model — best-fold checkpoint |
0.44 |
0.64 |
0.82 |
|
This layer |
Transformer plus convolution, retrained |
0.31 |
0.53 |
0.74 |
|
This layer |
In-house spline-based model |
0.28 |
0.52 |
0.74 |
|
This layer |
Cross-model ensemble (architectural diversity retained) |
0.36 |
0.58 |
0.79 |
What the table shows is the direction of generational transition. Every metric improves stepwise from rule-based to deep models, and in particular the area under the curve rises from 0.63 to around 0.80. A value of 0.63 is better than chance but insufficient for practical screening, while the 0.78–0.82 range provides adequate discrimination for narrowing candidates.
That the coefficient of determination remains in the 0.3–0.44 range reflects the intrinsic nature of the task. Silencing efficiency is a biological phenotype with substantial measurement noise, and values shift with cell line, delivery method and measurement time point. Repeat measurements of the same sequence are not perfectly reproducible, which caps what any predictive model can achieve. At this dataset size 0.44 is competitive, and reports claiming substantially more warrant a check on data leakage or evaluation design.
That the production checkpoint uses a single best fold rather than a fold average appears counterintuitive but is reasonable on a small dataset. Fold averaging averages away each fold model's weaknesses but also dilutes the strongest fold's signal, and at this dataset size the latter loss dominates. That decision is at the fold level only; ensembling across different architectures is retained — fold averaging is given up while the benefit of architectural diversity is kept.
24.2 Capability-axis comparison — double-stranded silencing design
Public design tools are generally based on rules formulated in the early to mid 2000s and were designed around a single species, a single database, and manual downstream analysis. They are adequate for proposing candidates, but do not provide the precise off-target classification, cross-species verification, chemistry design and regulatory traceability that clinical entry requires.
|
Capability axis |
Public rule-based tools |
Public web tools (improved) |
Pharmaceutical internal tools (inferred) |
This layer |
|---|---|---|---|---|
|
Efficacy prediction |
Positional rules only |
Rules plus some statistics |
Undisclosed (inferred to be trained on internal data) |
Rule ensemble plus deep-model ensemble across five distinct architectures |
|
Off-target classification |
Count of alignment hits |
Alignment plus seed search |
Full pipeline (inferred) |
Four-bucket classification (on-target extension / paralog / ortholog / unrelated) plus seed context |
|
Paralog awareness |
None |
None |
Present (inferred) |
Explicit co-silence and spare verdicts with hard filters |
|
Cross-species support |
1–2 species |
3–5 species |
Many (inferred) |
Parallel evaluation of multiple species combination modes with a two-species preclinical verdict |
|
Variant avoidance |
None or single database |
Partial |
Population variant database (inferred) |
Dual avoidance across population and clinical variants with multi-population allele frequencies |
|
Cleavage coverage |
None |
None |
Unknown |
Quantified per isoform, with coding-only and all-transcript counts separated |
|
Chemistry design |
None |
None |
Present (internal intellectual property) |
Per-position chemistry map, mechanism-specific templates, route-specific windows |
|
Delivery design |
None |
None |
Present (internal intellectual property) |
Conjugate and formulation candidates, route-specific absorption, formulation variant catalogue |
|
Regulatory support |
None |
None |
Present (internal) |
Run ledger, reference version verification, preclinical requirement verdict, impurity projection |
|
Reproducibility guarantee |
Not stated |
Not stated |
Unknown |
Byte-identical output contract for the same input and seed |
|
Automation interface |
Web only |
Web only |
Internal systems |
Web and command line, with batch processing |
The essential message of this table is not a claim of performance superiority but a difference in scope. Public tools have a clearly bounded purpose — proposing candidates — and are useful within it. The divergence begins after that: interpreting off-targets, verifying across species, fixing chemistry and delivery, generating regulatory evidence. Those stages have traditionally been dispersed across manual work and separate tools.
24.3 Capability-axis comparison — single-stranded modalities
For single-stranded modalities — gapmers, splice switching, anti-miRs — there is no public design tool as widely used as those for duplexes. Public tools mainly provide thermodynamic computation (melting temperature, secondary structure, accessibility) and do not address mechanism-specific chemistry rules or toxicity prediction. Comparison is therefore about which axes are automated.
|
Axis |
Public thermodynamic tools |
This layer |
|---|---|---|
|
Melting temperature and free energy |
Provided (core function) |
Provided, on published parameter sets |
|
Target accessibility |
Partially provided |
Provided — ensemble accessibility and interaction energy |
|
Mechanism-specific chemistry enforcement |
None |
Gap-mandatory, gap-forbidden and uniform-modification rules enforced automatically by mechanism |
|
Hepatotoxicity risk classification |
None |
Multi-tier classification from sequence and chemistry combination |
|
Quantifying the effect on splicing |
None |
Target-window masking differential — a quantity attributed to the candidate |
|
Reading-frame validation |
None |
Automatic validation of exon-skipping designs |
|
Junction-aware off-targets |
None |
Separate classification of junction-proximal hits |
|
Carrier-free suitability diagnosis |
None |
Pass/fail verdict per named gate with reasons |
|
Route-specific chemistry windows |
None |
Per-route backbone fraction windows implemented as constraints |
24.4 Target adjudication and family selectivity — the absence of a comparator
Modality suitability adjudication and paralog segregation are usually performed manually or not at all. The individual databases — population constraint, cell dependency, tissue expression, interaction networks — are all public, but an automated path that integrates them into one verdict and converts that verdict into design instructions is not standardised.
In this area the distinguishing feature of this layer is not algorithmic sophistication but explicitness of the verdict. The rules and coefficients are documented, the evidence used is presented item by item, and the limits of the verdict are stated alongside. Distinguishing 'no data' from 'no risk' in the notation is particularly important in practice, because reading a blank as a safety signal is the most common failure mode in this area.
24.5 Processing performance
The figures below are measured wall-clock times for the double-stranded silencing pipeline in an in-house benchmark environment. Turnaround for other modalities depends on scope and species count, and this document does not present estimates for values that were not measured.
|
Task |
Measured time |
Note |
|---|---|---|
|
Single gene, human only |
3–5 minutes |
Full eight-stage execution |
|
Five-species preclinical combination |
8–15 minutes |
Cross-species ortholog matching and evaluation run automatically in parallel |
|
Multi-target batch (10 genes) |
20–40 minutes |
Parallel execution |
|
Deep-ensemble re-ranking (200 candidates) |
under 5 seconds |
Forward passes over pretrained checkpoints only |
|
Cleavage coverage (per gene) |
tens of seconds |
Entails a transcript annotation rescan |
That a rule-based public tool is faster on a single gene reflects a difference in depth of evaluation. Computing positional rule scores and performing variant masking, transcriptome-wide off-target scanning, accessibility profiling, deep-model inference and chemistry design are not the same task. Conversely, the gap widens in cross-species evaluation because ortholog matching and per-species assessment run automatically in parallel, cutting time substantially against manual alignment analysis.
24.6 Cautions about comparison
- Benchmark performance and real-world performance are different. Public benchmarks come from particular cell systems and measurement protocols, and there is no guarantee that the same ranking reproduces on a customer's target and cell system.
- Area under the curve depends on where the threshold is set. The 0.7 threshold used in the table is a convention, and model ranking can change at other thresholds.
- Entries marked 'inferred' in the capability table concern undisclosed systems and are inferences from public information, not established fact.
- Processing times depend on hardware and reference data access speed and may not reproduce in other environments.
- No comparison in this chapter is evidence of a clinical outcome. That evidence comes only from wet-lab and clinical data.
25. Detailed Comparison Against Requirements
Design output is placed in three different review environments: academic peer review, use as public or open-source software, and validation as commercial or clinical software. The three demand different things, and satisfying one does not automatically satisfy the others. This chapter decomposes each environment's requirements item by item, states which technical artifact satisfies each, and where it does not, where the boundary lies.
Symbols are defined as follows. ✓ means the requirement is met by the tool itself without external supplementation; ◐ means the core infrastructure exists but site-specific configuration or an additional procedure is required; ✗ means it falls outside the design's scope. For ◐ and ✗ items the responsible party is named, because that clarity is what prevents disputes about accountability during review or audit. And marking ✗ items honestly is what makes the ✓ items credible.
25.1 Academic publication requirements
25.1.1 Three levels of reproducibility
What peer review means by reproducibility is not one concept but three levels. Each demands different evidence, and meeting one does not meet the others.
|
Level |
The question |
Evidence required |
How it is met |
Verdict |
|---|---|---|---|---|
|
Computational reproducibility |
Does the same input give the same result |
Output agreement for the same input and seed |
The determinism contract — order-independent parallel merging, explicit tie-break keys, caching restricted to pure functions |
✓ |
|
Methods reproducibility |
Can a third party reconstruct the same procedure |
Complete record of tool versions, database releases, parameters and seed |
Every item captured automatically in the run ledger — no need to reconstruct the Methods section afterwards |
✓ |
|
Results reproducibility |
Do other researchers independently reach the same conclusion |
Validation on independent data |
The design layer supplies the evidence; independent validation is the researcher's experimental domain |
◐ — a level no tool can guarantee |
25.1.2 Data availability and interoperability
|
Requirement |
What it asks for |
How it is met |
Verdict |
|---|---|---|---|
|
Machine readability |
Results in a format read by machines rather than only by eye |
Every output provided simultaneously in tabular and structured formats |
✓ |
|
Identifier standards |
Genes, transcripts and variants expressed with standard identifiers |
Approved symbols alongside each database identifier, with alias resolution stated |
✓ |
|
Provenance |
It must be possible to know where each value came from |
Per-item provenance attached, with reference data releases recorded |
✓ |
|
Supplementary format |
Reviewers must be able to read it without special tools |
Report, tables and structured data provided together |
✓ |
|
Long-term accessibility |
Data must remain accessible after publication |
Outputs are delivered as files; repository deposition is the author's choice |
◐ — author responsibility |
25.1.3 Completeness of methods description
A frequent reviewer criticism is incomplete description of what was run with which parameters — thresholds, filter conditions and database versions are commonly omitted. Because the run ledger captures these automatically, the Methods section can reference it directly, which prevents not only omissions but inaccuracies arising from recollection.
Care is also needed when citing predictive performance. The figures in Chapter 24 were measured on a specific public benchmark and are not a guarantee that the same performance reproduces on a customer's data. A paper citing them must state the dataset, and the corresponding conditions are recorded in the relevant deliverable items as well.
25.2 Public and open-source software requirements
25.2.1 Versioning and history
|
Requirement |
What it asks for |
How it is met |
Verdict |
|---|---|---|---|
|
Version identification |
The state of the executed code must be uniquely identifiable |
An immutable snapshot model — snapshot files do not change, so byte-level comparison confirms identity |
✓ |
|
Change history |
What changed, when and why must be traceable |
Per-snapshot history with reasons for change |
✓ |
|
Dependency disclosure |
External tools and libraries depended upon must be stated |
The computational tool layer table in 2.4 and version capture in the run ledger |
✓ |
|
Licence clarity |
Licence conditions of each component must be clear |
Public tools named directly; proprietary in-house algorithms distinguished by name |
✓ — the distinction is documented |
|
Downstream freedom |
Results must be usable without restriction |
Computation based on public tools is unencumbered, and items involving proprietary in-house algorithms are identifiable by name |
◐ — verification depends on the use context |
25.2.2 Robustness and graceful degradation
The most common way a public tool loses trust in practice is silent failure under exceptional conditions. If a missing reference file, an uninstalled external tool or an unreachable network resource causes quiet fallback to a default, a result is produced but nobody can say what it means.
|
Situation |
Poor handling |
How this layer handles it |
Verdict |
|---|---|---|---|
|
The expected reference data release is absent |
Quietly proceed with another version |
Stop at pre-execution verification |
✓ |
|
An optional external tool is absent |
Score that axis as zero |
Mark the axis 'not evaluated', exclude it from the composite score, and state so in the report |
✓ |
|
A model checkpoint is absent |
Return a random value |
Exclude that model from the ensemble and record which models were used |
✓ |
|
Target resolution is ambiguous |
Pick one arbitrarily |
State the resolution result and note where an alias was involved |
✓ |
|
An item has no data |
Blank or zero |
State 'no data' and the reason — distinguished from 'no risk' |
✓ |
The last row matters most. Failing to distinguish a blank from a safety signal turns a coverage limitation of the reference data directly into a false safety verdict. If, for example, constraint metrics for sex-chromosome genes are absent from the reference and left blank, they read as 'unconstrained', and a risky target passes a safety gate silently.
25.3 Commercial and clinical software requirements
25.3.1 Electronic records and audit trail
|
Requirement |
Standard |
How it is met |
Verdict |
|---|---|---|---|
|
Audit trail / electronic records |
The audit trail provisions of electronic records regulation |
A self-contained structured ledger persisted for every run |
✓ |
|
Time records |
Timestamps with explicit timezone |
Start and end times in Coordinated Universal Time in a standard format, removing ordering ambiguity across regions |
✓ |
|
Authority checks |
It must be possible to establish who executed the run |
The invoking principal is recorded in the ledger; strong authentication is the responsibility of the access control layer |
◐ — site infrastructure responsibility |
|
Record integrity |
It must be possible to confirm a record was not altered afterwards |
The immutable snapshot model — byte-level comparison confirms the producing code is unaltered |
✓ |
|
Record retention |
Records must remain readable for the retention period |
Stored in a standard structured format readable without proprietary tools; retention policy is a site responsibility |
◐ |
25.3.2 Validation and control of reference data
|
Requirement |
Standard |
How it is met |
Verdict |
|---|---|---|---|
|
Software validation |
Risk-based validation frameworks |
Regression and smoke testing run continuously; formal installation, operational and performance qualification are site deliverables tied to the deployed environment |
◐ |
|
Reference data version control |
Laboratory practice provisions on data management |
Expected releases verified at start-up, with the run stopped on failure |
✓ |
|
Reproducible methods |
The methods reproducibility expectation of nonclinical safety guidance |
Byte-identical output for the same input and seed |
✓ |
|
Change management |
Changes must follow a controlled procedure |
Per-snapshot history with reasons; approval procedure belongs to the site quality system |
◐ |
|
Data integrity principles |
Attributable, legible, contemporaneous, original, accurate |
The ledger secures attribution, timing and originality; the standard format secures legibility |
✓ |
25.3.3 Linkage to filing material
|
Requirement |
Standard |
How it is met |
Verdict |
|---|---|---|---|
|
Chemistry, manufacturing and controls material |
Impurity-related expectations of the quality module |
Synthesis impurity projection produced at design time |
✓ — at draft level; specification setting belongs to manufacturing |
|
Mutagenic impurity control |
The control logic of the relevant guideline |
Related impurity items assessed at the formulation and process stages |
◐ — at the level of risk exposure |
|
Nonclinical species selection |
The species expectations of nonclinical safety guidance |
Per-combination comparison of surviving candidates and an automatic two-species verdict |
✓ |
|
Pharmacokinetic material |
Biodistribution and dose translation |
Route-specific absorption and distribution summaries with dose translation |
◐ — predictions that do not replace measurement |
25.3.4 Items outside scope and why
|
Requirement |
Standard |
Verdict |
Reason and responsible party |
|---|---|---|---|
|
Patient information protection |
Technical safeguards of health information protection regulation |
✗ |
Processing is limited to sequences and public reference data and does not touch patient information. Where integrated into a workflow that includes patient data, those procedures are the integrating site's responsibility |
|
Medical device quality management |
Medical device quality management system standards |
✗ |
Supplied research-use-only; in-vitro diagnostic certification is a separate package undertaken by the integrating manufacturer |
|
Clinical decision support |
Regulation of clinical decision support software |
✗ |
It does not support individual patient care decisions and is a research and development tool |
|
Manufacturing process design |
Manufacturing regulation |
✗ |
It surfaces risk items and proposes draft specifications; actual process development and validation belong to the manufacturing stage |
Declaring out-of-scope items explicitly is itself an element of regulatory compliance. Leaving an obligation the tool was never designed for to be discharged by the tool means, in practice, that nobody discharges it. Naming the responsible party exposes that gap and lets the deploying site place the procedure inside its own quality system.
25.4 The three environments in one table
|
Nature of the requirement |
Academic publication |
Public software |
Commercial and clinical |
|---|---|---|---|
|
Reproducibility |
Essential — methods reproducibility is under review |
Essential — users must be able to verify |
Essential — required by regulation |
|
Provenance recording |
Recommended — the basis of the Methods section |
Recommended — dependency disclosure |
Essential — it is the substance of the audit trail |
|
Version immutability |
Recommended |
Essential — a citable version |
Essential — record integrity |
|
Licence clarity |
Recommended |
Essential — the precondition for downstream use |
Essential — product incorporation |
|
Explicit stop on failure |
Recommended |
Recommended |
Essential — silent failure violates data integrity |
|
Distinguishing absent data from safety |
Recommended |
Recommended |
Essential — prevents false safety verdicts |
|
Patient information handling |
Not applicable |
Not applicable |
Separate procedure where applicable (out of scope here) |
|
Formal qualification |
Not applicable |
Not applicable |
A site deliverable |
What the table shows is that the three environments do not conflict. What academic reproducibility requires and what a regulatory audit trail requires are largely the same technical artifacts, differing in stringency and format. Designing to satisfy the most stringent environment therefore satisfies the other two — except that organisational procedures such as formal qualification and certification cannot be discharged by a tool.
26. Limits and Disclaimers
This chapter states what this technical layer does not do and does not guarantee. It is the most important chapter in a technical whitepaper, because the boundaries recorded here define the credibility of everything in the preceding chapters.
26.1 Limits on the nature of predictions
- All efficacy, toxicity and delivery predictions are ranking tools. They do not guarantee absolute values and do not replace wet-lab validation. Their purpose is to compress the set requiring validation from tens or hundreds to a few so that experimental resource is concentrated.
- Reported predictive performance was measured on the stated public benchmark datasets. There is no guarantee of equivalent performance on other targets, other cell systems or other measurement protocols.
- Predictions from trained models are constrained by the distribution of their training data. Predictions for sequence and chemistry combinations unlike anything in training are less reliable, and in such cases disagreement with the rule-based axes acts as a warning signal.
- Structure prediction rests on thermodynamic models and does not fully reflect real intracellular factors such as protein binding, modification and molecular crowding.
26.2 Limits arising from data
- Reference data carry coverage limits. Notably, constraint metrics for sex-chromosome genes may be absent, in which case they are marked 'no data'. A blank must never be read as a safety signal.
- Binding, expression and dependency data are measured in particular cell lines and conditions and may not represent a customer's experimental system. This limitation is most pronounced in transcriptional activation, which depends heavily on chromatin state.
- Gene–disease association scores use different scales per source. Only the column mapped to a common interval can be compared, thresholded or ranked across sources; comparing raw scales produces inversions in which the source with the higher scale maximum always leads.
- Sharing a functional ontology term does not imply biological equivalence. Two proteins may share a molecular function term while acting on different substrates, and this is the main source of over-broad co-silencing verdicts in family adjudication.
- Predicted target lists are prioritisations and do not guarantee measured binding or repression.
26.3 Limits of scope
- This layer is supplied research-use-only. It carries no certification for direct diagnostic or therapeutic use, and in-vitro diagnostic certification is a separate process undertaken by the integrating manufacturer.
- No patient information is processed. Inputs are limited to sequences and public reference data.
- It does not design manufacturing processes. Chemistry-manufacturing-and-control items surface risk and propose draft specifications; actual process development and validation belong to the manufacturing stage.
- Structure- and molecular-dynamics-based evaluation stages require large simulation assets and are outside the standard design scope.
- Binding affinities (dissociation constants) are not computed. Only literature reference values are used, and uncertain entries are marked unconfirmed.
- Formal installation, operational and performance qualification, identity management, and medical device certification are the responsibility of the deploying site and the integrating manufacturer.
26.4 Intellectual property notice
Patent-related statements in this document and in the deliverables are informational and are not legal opinion. Freedom-to-operate depends on jurisdiction, claim construction and prosecution history and is the domain of qualified counsel. The design layer uses generic chemistry within a research-use scope and does not implement any particular company's proprietary chemistry or structures.
27. Selected References
The following are primary sources for the mechanisms, determinants and chemistry rules described in this document. Each judgement in the deliverables carries its corresponding source at item level.
27.1 Mechanism and design rules of RNA interference
- Fire A, et al. Potent and specific genetic interference by double-stranded RNA in Caenorhabditis elegans. Nature 1998;391:806-811.
- Elbashir SM, et al. Duplexes of 21-nucleotide RNAs mediate RNA interference in cultured mammalian cells. Nature 2001;411:494-498.
- Khvorova A, et al. Functional siRNAs and miRNAs exhibit strand bias. Cell 2003;115:209-216.
- Schwarz DS, et al. Asymmetry in the assembly of the RNAi enzyme complex. Cell 2003;115:199-208.
- Reynolds A, et al. Rational siRNA design for RNA interference. Nat Biotechnol 2004;22:326-330.
- Huesken D, et al. Design of a genome-wide siRNA library using an artificial neural network. Nat Biotechnol 2005;23:995-1001.
- Frank F, Sonenberg N, Nagar B. Structural basis for 5'-nucleotide base-specific recognition of guide RNA by human AGO2. Nature 2010;465:818-822.
- Wang PY, Bartel DP. The guide-RNA sequence dictates the slicing kinetics and conformational dynamics of the Argonaute silencing complex. Mol Cell 2024. (bioRxiv 2023.10.15.562437)
27.2 Off-targets and immunity
- Jackson AL, et al. Expression profiling reveals off-target gene regulation by RNAi. Nat Biotechnol 2003;21:635-637.
- Birmingham A, et al. 3' UTR seed matches, but not overall identity, are associated with RNAi off-targets. Nat Methods 2006;3:199-204.
- Judge AD, et al. Sequence-dependent stimulation of the mammalian innate immune response by synthetic siRNA. Nat Biotechnol 2005;23:457-462.
27.3 Single-stranded modalities and chemistry
- Holen T, et al. Similar behaviour of single-strand and double-strand siRNAs suggests they act through a common RNAi pathway. Nucleic Acids Res 2003;31:2401-2407.
- Stein CA, et al. Efficient gene silencing by delivery of locked nucleic acid antisense oligonucleotides, unassisted by transfection reagents. Nucleic Acids Res 2010;38:e3.
- Lima WF, et al. Single-stranded siRNAs activate RNAi in animals. Cell 2012;150:883-894.
- Yu D, et al. Single-stranded RNAs use RNAi to potently and allele-selectively inhibit mutant huntingtin expression. Cell 2012;150:895-908.
- Prakash TP, et al. Identification of metabolically stable 5'-phosphate analogs that support single-stranded siRNA activity. Nucleic Acids Res 2015;43:2993-3011. (erratum Nucleic Acids Res 2017;45:6994)
- Burel SA, et al. Hepatotoxicity of high affinity gapmer antisense oligonucleotides is mediated by RNase H1 dependent promiscuous reduction of very long pre-mRNA transcripts. Nucleic Acids Res 2016;44:2093-2109.
- Kasuya T, et al. Ribonuclease H1-dependent hepatotoxicity caused by locked nucleic acid-modified gapmer antisense oligonucleotides. Sci Rep 2016;6:30377.
- Pendergraff HM, et al. Single-stranded silencing RNAs: hit rate and chemical modification. Nucleic Acid Ther 2016;26:216-222.
27.4 Splicing regulation
- Singh NK, et al. Splicing of a critical exon of human survival motor neuron is regulated by a unique silencer element located in the last intron. Mol Cell Biol 2006;26:1333-1346.
- Hua Y, et al. Antisense masking of an hnRNP A1/A2 intronic splicing silencer corrects SMN2 splicing in transgenic mice. Am J Hum Genet 2008;82:834-848.
- Yeo G, Burge CB. Maximum entropy modeling of short sequence motifs with applications to RNA splicing signals. J Comput Biol 2004;11:377-394.
- Jaganathan K, et al. Predicting splicing from primary sequence with deep learning. Cell 2019;176:535-548.
27.5 Delivery
- Akinc A, et al. Targeted delivery of RNAi therapeutics with endogenous and exogenous ligand-based mechanisms. Mol Ther 2010;18:1357-1364.
- Jayaraman M, et al. Maximizing the potency of siRNA lipid nanoparticles for hepatic gene silencing in vivo. Angew Chem Int Ed 2012;51:8529-8533.
- Gilleron J, et al. Image-based analysis of lipid nanoparticle-mediated siRNA delivery, intracellular trafficking and endosomal escape. Nat Biotechnol 2013;31:638-646.
- Nair JK, et al. Multivalent N-acetylgalactosamine-conjugated siRNA localizes in hepatocytes and elicits robust RNAi-mediated gene silencing. J Am Chem Soc 2014;136:16958-16961.
- Cheng Q, et al. Selective organ targeting (SORT) nanoparticles for tissue-specific mRNA delivery and CRISPR-Cas gene editing. Nat Nanotechnol 2020;15:313-320.
27.6 Editing, readthrough and immune modulation
- Lueck JD, et al. Engineered transfer RNAs for suppression of premature termination codons. Nat Commun 2019;10:822.
- Cridge AG, et al. Eukaryotic translational termination efficiency is influenced by the 3' nucleotides within the ribosomal mRNA channel. Nucleic Acids Res 2018;46:1927-1944.
- Krieg AM. Therapeutic potential of Toll-like receptor 9 activation. Nat Rev Drug Discov 2006;5:471-484.
27.7 Predictive models and thermodynamics
- SantaLucia J Jr. A unified view of polymer, dumbbell, and oligonucleotide DNA nearest-neighbor thermodynamics. Proc Natl Acad Sci USA 1998;95:1460-1465.
- Turner DH, Mathews DH. NNDB: the nearest neighbor parameter database for predicting stability of nucleic acid secondary structure. Nucleic Acids Res 2010;38:D280-D282.
- Bai Y, et al. OligoFormer: an off-target-aware deep learning approach for siRNA design. Nat Commun 2024;15:9123.
- Liu B, et al. Cm-siRPred: a multi-view chemistry-aware predictor for modified siRNA efficacy. Brief Bioinform 2024;25:bbae123.
- Liao W, et al. DeepSilencer: a deep learning model for predicting siRNA knockdown efficacy. arXiv:2503.04200, 2025.
27.8 Reproducibility and regulation
- Wilkinson MD, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data 2016;3:160018.
- ICH Harmonised Tripartite Guideline M3(R2). Guidance on nonclinical safety studies for the conduct of human clinical trials and marketing authorization for pharmaceuticals. 2009.
- FDA. 21 CFR Part 11 — Electronic Records; Electronic Signatures. Final Rule, 1997.
28. Document Information
|
Item |
Detail |
|---|---|
|
Published by |
Bioneer BioFoundryCenter (바이오니아 바이오파운드리센터) |
|
Document |
Therapeutic Oligonucleotide Design Platform — Technical Whitepaper v2.0 (2026) |
|
Contact |
geneorder@bioneer.co.kr |
|
Nature of the document |
A public technical whitepaper — the biology of each modality, its design requirements, deliverables and limits |
Reference data releases and performance figures stated in this document are current as of writing; the values actually used by any individual run are recorded in that run's ledger. Contact: geneorder@bioneer.co.kr