Therapeutic Oligonucleotide Design Platform — Technical Whitepaper

Modality by modality: molecular form, chemistry, delivery, deliverables, and industry benchmarking

Bioneer BioFoundryCenter · Published 2026 · Document version 2.0 · Contact geneorder@bioneer.co.kr

0. What This Document Is

This is a public technical whitepaper. It is not a product manual. It describes, modality by modality, the biological mechanism by which a therapeutic oligonucleotide acts, the demands that mechanism places on design, chemistry and delivery, and the computation and data that meet those demands. The unit of description is the molecule, not the program. What a customer ultimately holds is not software but a sequence, a chemistry specification and a delivery specification — and understanding what determines those three is the substance of a development decision.

Every modality is therefore described under the same eight headings: (1) the form and definition of the molecule, (2) the molecular biology of its action, (3) the primary literature and structural evidence that established that mechanism, (4) clinical and regulatory precedent and the surrounding patent landscape, (5) whether chemical modification is required and why, (6) whether a delivery vehicle is required and why, (7) what the design deliverables consist of, and (8) what this modality cannot solve in principle. Never omitting the eighth heading is a principle of this document. A technical document that does not state a modality's boundaries leaves the reader to discover them as a failed experiment, at a cost many times that of the design itself.

Figures cited here fall into three classes. Values reported in the public literature are given with author, journal and year. Values measured on in-house benchmarks are given with the dataset they were measured on. Values whose basis is not established are not given at all — the fact that they are undetermined is stated instead. Refusing to fill that third class with estimates is the only thing that makes the first two credible.

1. What Kinds of Oligonucleotide Can Be Designed

Therapeutic oligonucleotides are not one technology but eleven distinct operating principles sharing a common chemical foundation. Two nucleic acids of similar length can behave entirely differently: one cleaves its target RNA, another occupies a site without cleaving it, another raises rather than lowers gene expression, another stimulates an immune receptor, and another folds into a ligand that binds a protein. This chapter lays out that whole landscape first; each subsequent chapter expands one row of the table.

1.1 Eleven roles by operating principle

#

Role

Common name

What the molecule does to its target

Type of target

I

Catalytic silencing (RISC cleavage)

siRNA · RNAi (RNA interference)

Enzymatically cleaves the target RNA, then is reused on the next target

Mature cytoplasmic mRNA

II

Catalytic silencing (RNase H1 cleavage)

ASO gapmer (antisense oligonucleotide)

Forms a DNA:RNA heteroduplex that recruits a host nuclease to cleave the RNA strand

Nuclear pre-mRNA and cytoplasmic mRNA

III

Splice switching (non-cleaving steric block)

SSO (splice-switching oligonucleotide)

Physically masks a splicing regulatory element, changing isoform choice

Intron/exon boundary elements of pre-mRNA

IV

Noncoding RNA silencing

ncRNA-directed siRNA / ASO

Removes an RNA that makes no protein, by mechanism I or II

lncRNA · circRNA · snoRNA · translation-enhancing elements

V

Transcriptional activation

saRNA · RNAa (small activating RNA)

Binds the promoter region and increases transcription of that gene

The genomic region around the transcription start site and the RNA made there

VI

miRNA axis modulation

miRNA mimic · anti-miR (antagomir)

Replaces the function of an endogenous miRNA (mimic) or sequesters it (antagonist)

A mature miRNA and its target mRNA set

VII

Site-directed base editing

ADAR-recruiting editing oligo (AIMer · EON · arRNA)

Recruits an endogenous editing enzyme to convert a specific adenosine to inosine

A single adenosine in a mature mRNA

VIII

Premature stop codon readthrough

suppressor tRNA (ACE-tRNA · sup-tRNA)

Reads a stop codon as an amino acid, restoring full-length protein synthesis

The premature termination codon of a nonsense-mutant mRNA

IX

Innate immune receptor modulation

CpG-ODN · immunomodulatory oligo (TLR agonist/antagonist)

Stimulates or suppresses an endosomal Toll-like receptor

TLR7 · TLR8 · TLR9 proteins

X

Targeted delivery (conjugate)

aptamer conjugate (ApDC · AOC)

A folded nucleic acid binds a surface receptor and carries a payload into that cell

A cell-surface protein

XI

Structure-based ligand candidate generation

de novo aptamer scaffold

Generates a candidate pool by folding fitness; affinity is determined experimentally

Any protein target (decided at the experimental stage)

Roles I and II both cleave the target, but the catalyst differs. In I the Argonaute protein is itself the slicer; in II the oligonucleotide builds a DNA:RNA heteroduplex that the host's RNase H1 recognises and cuts. That difference propagates straight into design: I must be double-stranded (or, if single-stranded, must satisfy the RISC loading requirements), while II must carry a central DNA window. Role III must not be cleaved at all, so the presence of that same DNA window makes the design fail. Within one class of single-stranded molecule, 'a central DNA gap is mandatory' and 'a central DNA gap is forbidden' stand opposed — and that single fact is the entire reason II and III are different modalities.

Role V runs in the opposite direction. Where I through IV reduce the target, V increases it. When the therapeutic hypothesis is not 'there is too much of this protein' but 'there is too little' — a lost tumour suppressor, a haploinsufficient gene, an insufficient compensatory pathway — silencing mechanisms cannot be the answer in principle. Roles VII and VIII change neither quantity nor occupancy but sequence: they apply where the problem is the quality of the transcript rather than its amount, that is, a point mutation or a premature stop.

Roles IX through XI do not target RNA at all. IX acts on a protein receptor as a ligand, X uses a folded nucleic acid as a targeting device in place of an antibody, and XI generates candidates for that device. They belong to the same technology family because the material and the manufacturing route are identical — the same solid-phase phosphoramidite synthesis and the same repertoire of chemical modifications.

1.2 The form of the molecule — single-stranded or double-stranded

Form is not a technical detail. A double-stranded molecule requires two strands to be synthesised and annealed, doubling synthesis cost and quality-control items; it requires thermodynamic control over which strand is loaded as the active guide; and the strand that is not loaded becomes an independent off-target burden. A single-stranded molecule has none of those problems, but it has no partner strand to protect it, so its chemical modification burden is higher and its target affinity has to be secured chemically.

Role

Form

Typical length

Strand composition

What the form imposes

I siRNA / RNAi — catalytic silencing (RISC)

double-stranded

21 bp core + 2 nt overhang ×2 = 23 nt per strand

guide (antisense) + passenger (sense)

Strand-selection asymmetry must be secured; passenger off-targets must be managed

I variant: ss-siRNA (single-strand RISC)

single-stranded

~20 nt

guide only

A 5'-phosphate mimic is an absolute requirement; duplex chemistry does not transfer

II ASO gapmer — catalytic silencing (RNase H1)

single-stranded

16–20 nt

wing–DNA gap–wing (typically 5-10-5)

Central DNA gap mandatory; wings carry the affinity

III SSO — splice switching

single-stranded

18–25 nt

uniform modification, no gap

DNA residues forbidden; uniform high affinity mandatory

IV ncRNA-directed silencing

double- or single-stranded

as I or II

follows the chosen mechanism

Target class dictates which mechanism applies

V saRNA / RNAa — transcriptional activation

double-stranded

21 bp core + overhangs

guide + passenger

No cleavage requirement; nuclear localisation and chromatin access are what matter

VI-a miRNA mimic

double-stranded

mature miRNA length (typically 21–23 nt)

guide = the mature miRNA sequence

Sequence is fixed in advance; design freedom exists only in chemistry and passenger

VI-b anti-miR / antagomir

single-stranded

8 nt (seed-directed) to full length

single strand

Must not recruit RNase H, so modification must be uniform

VII ADAR-recruiting editing (AIMer / EON)

single-stranded

~20–40 nt (short) / 50–100 nt (long)

single strand

A C mismatch opposite the target A is fixed; the central window must stay lightly modified

VIII suppressor tRNA (ACE-tRNA)

folded tRNA (~72–76 nt)

mature tRNA length

single chain, cloverleaf fold

Aminoacylation identity elements invariant; tertiary structure must be preserved

IX CpG-ODN / TLR ligand

single-stranded DNA or RNA

18–30 nt (class dependent)

single strand or palindromic dimer

The phosphorothioate backbone is part of the function; motif placement is decisive

X aptamer conjugate (ApDC / AOC)

folded single strand + payload

aptamer 25–90 nt + linker + payload

aptamer–linker–payload

The fold is the prerequisite for binding; payload attachment must not break it

XI aptamer scaffold pool

folded single strand

20–100 nt (per constraints)

single strand

Structural diversity is the goal, not a single optimum

The '21 bp core + 2 nt overhang' entry is this platform's double-stranded standard. The overhang bases are not an arbitrary thymidine dimer but the real bases from the target sequence context, and the antisense overhang is the reverse complement of the two upstream nucleotides. The rationale is given in 2.6.

1.3 Which modalities require chemical modification and which do not

Chemical modification is not a uniform procedure applied to every oligonucleotide. What it is for differs by modality, and in some cases modification destroys function. Its purposes are four: (a) nuclease resistance, (b) target binding affinity, (c) suppression of innate immune stimulation, and (d) mediating cellular uptake itself. The fourth applies only where there is no delivery vehicle; where there is one, it becomes a liability instead.

Role

Chemical modification

For what

Where modification is forbidden or limited

I siRNA / RNAi (RISC)

required

Plasma stability, immune suppression, seed off-target mitigation, hepatic conjugation

Around the position opposing the AGO2 cleavage site, only modifications that do not block cleavage

I variant: ss-siRNA

required, under different rules

A 5'-phosphate mimic is the precondition for activity; systemic stability

Excessive 2'-O-alkyl or bicyclic substitution lowers activity; duplex templates cannot be reused

II ASO gapmer (RNase H1)

required, and structurally enforced

Wing affinity plus preservation of catalytic recognition in the gap

The central gap must remain 2'-deoxy — a 2'-modification there abolishes cleavage

III SSO (splice switching)

required, uniform across every position

Maximum affinity, nuclease resistance, and avoidance of RNase H

No DNA residues — even one risks unintended cleavage

IV ncRNA-directed silencing

required

as I or II

Inherits the constraints of the chosen mechanism

V saRNA / RNAa

required, minus the cleavage rule

Stability and immune suppression. No cleavage is needed, so catalytic-site protection does not apply

Bulky additions that impede nuclear localisation warrant caution

VI-a miRNA mimic

required

Preserving the natural function of the mature miRNA while gaining stability

Over-modification of the seed region (positions 2–8) distorts target recognition

VI-b anti-miR / antagomir

required, uniform and fully backbone-modified

Maximum sequestration affinity while avoiding RNase H

Uniformity is mandatory, so partial-modification designs do not exist

VII ADAR-recruiting editing (AIMer)

required, but the centre must stay light

Stabilisation of the termini and flanks

The central window around the orphan cytidine must retain unmodified 2'-OH and phosphodiester linkages — modification there directly lowers editing efficiency

VIII suppressor tRNA (ACE-tRNA)

not required in principle

Stabilisation is discussed only where a synthetic RNA is dosed directly; the standard route is an expression cassette

Artificial modifications that perturb native modification sites or tertiary structure are avoided

IX CpG-ODN / TLR ligand

required — the modification is the function

The phosphorothioate backbone provides both stability and uptake; 2'-O-methyl is an active component of antagonist design

In agonist design, heavy 2'-modification around the CpG motif abolishes recognition

X aptamer conjugate (ApDC / AOC)

required

Serum nuclease resistance (2'-fluoro, 2'-O-methyl), delayed renal clearance (PEGylation)

Residues involved in binding must retain their original chemistry

XI aptamer scaffold pool

optional

Candidates are evaluated unmodified for folding; stabilisation is applied after selection

Fixing chemistry during exploration shrinks the search space

The two rows that deserve most attention are VII and VIII. VII is the only modality in which modification is mandatory and yet one specific window must remain unmodified — stability and activity collide directly on the same axis. VIII is the only modality in which modification is not required at all, because the molecule is not a synthetic ligand but a functional RNA that must fold inside the cell, be charged with an amino acid, and enter the ribosome. An artificial modification that perturbs any one of those three processes disables it.

1.4 Which modalities require a delivery vehicle and which do not

Delivery splits into two questions: does the molecule get into the cell, and does it reach the right tissue. Nucleic acids are large and strongly anionic, so they do not cross membranes by free diffusion. Some uptake mechanism is therefore mandatory, and there are only three options: encapsulate in a lipid particle, attach a receptor-binding ligand, or make the backbone chemistry itself bind proteins and drive endocytosis. The third is carrier-free (gymnotic) uptake, and the protein-binding character of the phosphorothioate backbone is what makes it work.

Role

Delivery vehicle

Route in practice

Why a vehicle is unnecessary or unsuitable

I siRNA / RNAi (RISC)

required

GalNAc conjugation for hepatic targets; lipid nanoparticles or alternative vehicles elsewhere

A duplex cannot carry a fully phosphorothioate backbone, so carrier-free uptake barely applies

I variant: ss-siRNA

optional

Lipid conjugation or carrier-free dosing are both reported

A single strand can carry a high PS fraction, but potency is traded away in doing so

II ASO gapmer (RNase H1)

optional — often unnecessary

Systemically via GalNAc conjugation or carrier-free; intrathecally for the central nervous system

A phosphorothioate single strand is the canonical case where carrier-free uptake works

III SSO (splice switching)

usually unnecessary

Intrathecal for CNS; systemic dosing with tissue distribution for muscle

Several approved products are dosed as naked oligonucleotides

IV ncRNA-directed silencing

follows the mechanism

as I or II

Subcellular localisation of the target drives the choice

V saRNA / RNAa

required

Lipid nanoparticles or conjugates

The molecule must reach the nucleus, so cytoplasmic delivery alone is insufficient

VI-a miRNA mimic

required

Lipid nanoparticles or conjugates

Duplex, so the same constraint as I

VI-b anti-miR / antagomir

optional

A fully phosphorothioate single strand supports carrier-free uptake; GalNAc conjugation for liver

Same logic as II

VII ADAR-recruiting editing (AIMer)

required

Lipid nanoparticles or conjugates — the unmodified central window lowers stability, so protection matters

Chemical protection cannot be placed at the centre, so the carrier's role is comparatively larger

VIII suppressor tRNA (ACE-tRNA)

required

An expression cassette in a lipid particle or viral vector, or synthetic tRNA in a particle

A folded functional RNA does not undergo carrier-free uptake

IX CpG-ODN / TLR ligand

unnecessary

Local or subcutaneous administration; DNA CpG is co-formulated with adjuvants such as aluminium salts

The target is an endosomal receptor, so endocytosis is itself arrival, and the PS backbone drives it

X aptamer conjugate (ApDC / AOC)

unnecessary — the molecule is the vehicle

Intravenous or local administration followed by receptor-mediated internalisation

The aptamer is the targeting device; adding a carrier dilutes that targeting

XI aptamer scaffold pool

not applicable

Laboratory-stage output

Not dosed in vivo

1.5 The chemistry × delivery map

Overlaying the two preceding tables yields four quadrants. This map is used directly to allocate effort early in a programme: in the quadrants that need a vehicle, no amount of sequence quality relieves the delivery bottleneck, and in the quadrants where chemistry is the function, a sequence without a fixed chemistry specification means nothing.

Delivery vehicle required

Vehicle unnecessary or optional

Chemical modification required

I siRNA/RNAi · V saRNA/RNAa · VI-a miRNA mimic · VII ADAR editing (AIMer)

II ASO gapmer · III SSO · VI-b anti-miR · IX CpG-ODN/TLR ligand · X aptamer conjugate (ApDC/AOC)

Chemical modification unnecessary or optional

VIII suppressor tRNA (ACE-tRNA)

XI aptamer scaffold pool (pre-in vivo stage)

The upper-left quadrant is where both problems must be solved, and is therefore the hardest — except for hepatic targets, where GalNAc conjugation is a mature answer that collapses the practical difficulty. The upper-right quadrant is where chemistry is delivery: a single parameter, phosphorothioate fraction, moves uptake, stability and toxicity simultaneously, so locating its productive window is the core of design. The lower-left is where the molecule is a functional RNA rather than a synthetic oligonucleotide, and the lower-right is a stage before in vivo administration.

1.7 Terminology map — the many names for the same molecule

This field calls the same molecule by different names in the literature, in regulatory documents and in industry. The table below maps the role numbers used in this document to their common names and to the synonyms and adjacent terms most often encountered, so that a reader knows which term to search on in literature or patent work.

Role

This document

Most widely used name

Synonyms and adjacent terms

I

Double-stranded catalytic silencing

siRNA (small interfering RNA), RNAi

double-stranded siRNA · duplex siRNA · GalNAc-siRNA · ESC / ESC+ siRNA · RNA interference · RISC-mediated silencing · shRNA (expressed analogue)

I variant

Single-stranded catalytic silencing

ss-siRNA (single-stranded siRNA)

ss-RNAi · single-stranded RNAi · 5'-VP guide

II

Single-stranded catalytic silencing

ASO gapmer (gapmer antisense oligonucleotide)

antisense oligonucleotide · AON · RNase H gapmer · 2'-MOE gapmer · cEt gapmer · LNA gapmer · GalNAc-ASO · third-generation antisense

III

Splice switching

SSO (splice-switching oligonucleotide)

SSA · splice-switching antisense · exon-skipping oligonucleotide · PMO (morpholino) · PPMO (peptide-conjugated PMO) · steric-block antisense · TOSS / TOES

IV

Noncoding RNA silencing

ncRNA-directed siRNA / ASO

lncRNA-targeting antisense · circRNA-targeting siRNA · anti-lncRNA · snoRNA-directed oligo · SINEUP

V

Transcriptional activation

saRNA (small activating RNA), RNAa

RNA activation · promoter-targeted RNA · pshRNA · NAT-targeting antisense (AntagoNAT)

VI-a

miRNA mimicry

miRNA mimic

miRNA replacement therapy · synthetic miRNA duplex

VI-b

miRNA antagonism

anti-miR / antagomir

antimiR · miRNA inhibitor · anti-miRNA oligonucleotide (AMO) · tiny LNA · miRNA sponge (expressed analogue) · mixmer

VII

Site-directed base editing

ADAR-recruiting editing oligonucleotide

AIMer · EON (editing oligonucleotide) · Axiomer · arRNA · RESTORE · LEAPER · ADAR-recruiting guide · A-to-I RNA editing

VIII

Premature stop codon readthrough

suppressor tRNA

ACE-tRNA · sup-tRNA · anticodon-engineered tRNA · nonsense suppressor tRNA · PTC readthrough (distinct from small-molecule inducers)

IX

Innate immune receptor modulation

CpG-ODN / TLR agonist or antagonist oligonucleotide

CpG oligodeoxynucleotide · immunostimulatory sequence (ISS) · IMO · TLR9 agonist · TLR7/8 agonist · inhibitory oligonucleotide (INH-ODN, suppressive ODN)

X

Targeted delivery conjugate

ApDC / AOC (aptamer conjugate)

aptamer–drug conjugate · aptamer–oligonucleotide conjugate · aptamer–siRNA chimera · aptamer-guided delivery

XI

Structure-based ligand candidate generation

de novo aptamer scaffold

aptamer library design · SELEX starting pool · structure-based candidate pool

Adjacent technologies that fall outside this document's scope are distinguished as well. shRNA and miRNA sponges are expressed from vectors, so their manufacturing and regulatory paths differ from synthetic oligonucleotides; the CRISPR family (Cas9, Cas13, base editors, prime editors) carries a protein component and is a separate technology class. mRNA, self-amplifying RNA and circular RNA vaccines are different molecules but share the lipid nanoparticle delivery layer, and are therefore covered in Chapters 7 and 23.

Category

Terms

Relationship to this document

Expressed analogues

shRNA · artificial miRNA · miRNA sponge · expressed antisense

Same mechanism but vector-expressed, so manufacturing and regulatory paths differ (out of scope)

Protein-accompanied technologies

CRISPR-Cas9 · Cas13 · base editors · prime editors

The guide is a nucleic acid, but a protein or its mRNA must be co-delivered — a separate technology class (out of scope)

Antigen-expressing platforms

mRNA vaccines · self-amplifying RNA (saRNA*) · circular RNA vaccines

Different molecules but a shared delivery layer — covered in Chapter 7 (delivery) and Chapter 23 (vaccines)

Chemical elements

2'-OMe · 2'-F · 2'-MOE · LNA · cEt · GNA · PS · PMO · 5'-VP · GalNAc

Not modalities but the chemistry repertoire — covered in Chapter 6

Delivery elements

LNP · SORT · exosome · polymer · VLP · GalNAc conjugation · gymnosis

Not modalities but the delivery repertoire — covered in Chapter 7

A collision of abbreviations to watch: saRNA denotes two different things in the literature. In this document saRNA means the small activating RNA of Modality V; the self-amplifying RNA used in vaccine contexts is a different molecule and is written out in full as 'self-amplifying RNA', only in section 23.3.2.

1.6 Categories of deliverable

The output of design activity is not only a sequence. What actually feeds a development decision is the sequence, the reason it was chosen, and the quantitative evidence behind that reason. Deliverables fall into four categories.

Category

Contents

What it is used for

Molecular specification

Per-strand sequences, length and overhang convention, per-position chemistry map (sugar string and backbone string), terminal handling, conjugate attachment point and linker

Synthesis orders, quality specifications, definition of the substance

Selection evidence

Efficacy predictions and their decomposition, off-target tables (per gene and per site), variant avoidance results, accessibility profiles, cross-species tables, cleavage coverage

Candidate selection meetings, experimental prioritisation, review responses

Development-readiness material

Preclinical species suitability verdicts, synthesis impurity projections, formulation and stability risk items, route-dependent absorption summaries, patent-landscape advisories

IND preparation, CDMO handover, start of formulation development

Reproducibility material

The run ledger — tool versions, database releases, parameters, random seed, UTC timestamps

Journal Methods sections, regulatory filings, later audits

The fourth category is routinely treated as a by-product but in fact determines the value of the other three. If it is impossible to reconstruct after the fact why a candidate was chosen, then the experimental results obtained with that candidate cannot be interpreted reproducibly either. A determinism contract guaranteeing byte-identical output for the same input and seed, together with a ledger written automatically for every run, is what this category consists of.

2. Shared Technical Foundation

The eleven modalities differ in operating principle but are designed on one shared foundation. How a target gene is resolved, what counts as an off-target, which coordinate system positions are expressed in, and how a result is made reproducible — if these four differ per modality, comparison between modalities does not hold. When two modalities are designed in parallel against the same gene, '3 off-targets' in one report must mean the same three things it means in the other. That is why this chapter exists.

2.1 The determinism contract

The design pipeline produces byte-identical output for the same input and the same random seed. This sounds unremarkable, yet a substantial fraction of bioinformatics pipelines do not have the property. The causes are usually three: results accumulate into a data structure in an order that depends on parallel execution; iteration depends on hash ordering; or an external tool breaks ties non-deterministically across threads.

Each of the three is blocked structurally. Parallel results are merged only in order-independent ways, ties carry explicit secondary sort keys, and caching is restricted to computations that are pure functions. That last restriction matters most in practice. Cross-species ortholog search is a pure function of (query sequence, species database), so caching it cannot change the value; off-target screening, by contrast, has two aligners writing into one structure — one by assignment and one by accumulation — so a deterministic cache would double-count. A performance optimisation that changes the result is not an optimisation but a defect, so the latter is deliberately not cached.

Determinism buys more than reproducibility. In a deterministic system, any difference between two runs must originate in a difference of input, so the effect of changing one parameter can be attributed exactly. In a non-deterministic system that attribution is impossible, and as a consequence the question 'does this parameter matter' cannot be answered at all.

2.2 Provenance — the run ledger

Every run leaves a self-contained ledger. It records start and end times as UTC ISO-8601 timestamps carrying explicit timezone information, the invoking principal, a fingerprint identifying the state of the code, the version of every external tool used, the release of every reference database consulted, the scoring weights applied, and the random seed.

Fixing timestamps to UTC removes ordering ambiguity caused by local timezone differences in multi-region collaboration, which is a basic requirement of the audit trails that electronic records regulations ask for. Recording database releases is even more practical: when two designs of the same gene give different results, it is the only basis on which to distinguish an algorithm change from a reference data update.

The ledger is also the record of pre-execution verification. Before any computation begins, the pipeline confirms that the expected reference data releases are actually present, and on failure it stops rather than silently computing against something else. Silent failure is far more expensive than explicit failure, because a wrong result then propagates downstream indistinguishable from a correct one.

2.3 Reference data

Design quality depends on the breadth and freshness of reference data as much as on algorithms. The principal resources and their roles are below.

Resource

Release

Role

What breaks without it

Population variation database

v4.1 (constraint metrics v4.1.1)

Target-site polymorphism avoidance, common-SNP checks across the seed, gene-level loss constraint

Candidates that work only in part of a population rise to the top

Transcript annotation

GENCODE v49

Transcript coordinates, biotypes, canonical isoform designation, cleavage-coverage arithmetic

It becomes impossible to know which isoform set 'knocking down a gene' actually knocks down

Clinical variant database

rolling

Pathogenic variant avoidance, decomposition of the nonsense variant spectrum, variant profiles

Designs land on disease-causing variants, and readthrough or editing suitability cannot be judged

Circular RNA atlas

circAtlas 3.0

Back-splice junction definitions

Circle-specific targets cannot be separated from the linear parent gene

Noncoding RNA family database

Rfam 15

Noncoding RNA family classification

Misclassifying the target class leads to the wrong mechanism

Integrated noncoding RNA resources

RNAcentral · LncATLAS · REDIportal

Identifier integration, subcellular localisation, A-to-I editing sites

Cytoplasmic mechanisms get applied to nuclear-retained targets

Chromatin and protein–RNA resources

ENCODE (H3K4me3 · ATAC · H3K27ac · eCLIP)

Promoter accessibility, RNA-binding protein occupancy masks

Activation preconditions cannot be checked, and protein-occupied sites get targeted

RNA modification atlas

m6A-Atlas

Modification-site interference masks

Sites where modification blocks binding are not avoided

miRNA resources

miRBase mature set · TargetScan 8.0 · ENCORI

Mature miRNA anchors, context-based target prediction, AGO-CLIP and degradome cross-checks

Prediction alone is dominated by false positives

Tissue expression resource

GTEx v11

Expression breadth, tissue-specificity index, target tissue calls

Delivery target tissue and systemic exposure risk are assumed without basis

Cell dependency resource

DepMap (1,208 cell lines)

Essentiality calls, residual fitness dependency

A well-designed molecule silences an essential gene

Tractability and liability resource

Open Targets

Tractability buckets, documented adverse-event precedent

Targets with existing failure precedent are chosen again

tRNA resource

gtRNAdb (hg38 · mm39)

Scaffold body sequences

Risk of using invented sequence

RNA modification resource

MODOMICS

tRNA modification-impact assessment

Perturbation of modification sites goes undetected

Protein interaction graph

13 integrated sources

1.73M consensus pairs, 6.70M pairs across 13 evidence channels, 22,824 complexes

No basis for finding a detour target when the direct target is not viable

Knowledge graph

integrated build

4,228 pathways · 48,329 GO terms / 858,529 annotations · 28.09M gene–disease · 140,134 paralog pairs

No basis for family selectivity or disease-association verdicts

2.4 Computational tool layer and licensing posture

The platform combines openly licensed software with components developed in house. Public tools are named directly so a reader can verify the method independently; components developed in house are identified as proprietary in-house algorithms and carry GC-prefixed names. This is a naming policy, not concealment — what each component computes is stated below.

Component

Licence class

What it computes

BLAST+ (blastn · blastn-short · tblastx)

public

Off-target search, paralog and ortholog matching, cross-species transcript comparison

Bowtie

public

Transcriptome-wide mismatch-tolerant alignment (first-pass off-target scan)

Infernal (cmsearch)

public

Covariance-model validation of structured RNA families

RIblast · RIsearch

public

RNA–RNA interaction free energy, target-site accessibility

CPC2

public

Coding-potential call (the gate on a 'noncoding' premise)

miranda

public

Strict-seed miRNA target search

patman

public

High-throughput short-sequence matching

GC-ThermoFold

proprietary in-house algorithm

Secondary-structure minimum free energy and partition function, duplex binding free energy, local accessibility profiles, G-quadruplex propensity

GC-SpliceNet

proprietary in-house algorithm

Deep-learning splice-site probability; in-silico masking differential

GC-SpliceMotif

proprietary in-house algorithm

Maximum-entropy splice-site scoring, latent cryptic-site scans

GC-ESRScan

proprietary in-house algorithm

Exonic splicing regulatory element motif scanning

Stating the licensing posture in a whitepaper has a practical purpose. Academic publication cares about reproducibility of method; open distribution cares about downstream freedom; commercial development cares about whether a component may be incorporated into a product. These are three different conditions. The public tools used here generally carry permissive licences, so method disclosure and reproduction are unconstrained, and the proprietary in-house algorithms are distinguished by name so that which computations fall in that category is documented.

GC-ThermoFold is the collective name for the thermodynamic layer: minimum-free-energy folding, duplex binding energy (ΔG_bind = ΔG_duplex − ΔG_A − ΔG_B), local accessibility profiling, and G-quadruplex propensity. Wherever this document says 'secondary structure', 'accessibility' or 'ΔG', this is the component doing the work. The physical model underneath is the published nearest-neighbour thermodynamic parameter set, so the parameters themselves are verifiable against the literature.

2.5 Defining an off-target — four buckets

A count of alignment hits cannot be used as an index of off-target risk. A hit against another isoform of the target gene is on-target, not off-target; a hit against a same-family paralog may be desirable or harmful depending on the therapeutic hypothesis; and a hit against another species' ortholog is a property preclinical species selection actively requires. Summing all three into one number causes excellent candidates to be discarded on a raw count.

Bucket

Definition

Interpretation

Treatment in design

On-target extension

Another isoform of the target gene

Coverage, not off-target — often the broader the better

Quantified separately as cleavage coverage

Paralog

Another member of the same gene family

Allow or exclude depending on the therapeutic hypothesis — not automatically decidable

Handled as a hard filter under explicit instruction

Ortholog

The corresponding gene in another species

A property required for preclinical species selection

Presented as per-species complementarity with mismatch positions

Unrelated hit

Anything not in the three above

A true off-target — risk graded by seed context and position

Ranked by seed-based scoring and positional weighting

Risk is not uniform even within the unrelated bucket. Full complementarity across the seed (guide positions 2–8) produces miRNA-like repression that lowers expression without cleavage, and the phenomenon is predicted from the distribution of seed complementarity in target 3' untranslated regions. Conversely, hits carrying many mismatches outside the seed often do not translate into repression despite a high alignment score. Off-target assessment is therefore a question of mechanism rather than of alignment; alignment is only the input to that assessment.

2.6 The double-stranded synthesis specification

The point at which design output and synthesis orders most often diverge is the overhang convention. This platform's double-stranded standard is a 21 bp duplex core carrying a 2 nt native overhang at each end, so each strand is 23 nt.

sense        5'- UAGGUCAUCGAUGCUAGCUAG GC -3'   (21 + 2 = 23 nt)
antisense    5'- CUAGCUAGCAUCGAUGACCUA AG -3'   (21 + 2 = 23 nt)
                 └──────  21 bp core  ──────┘ └┘ native 2 nt overhang
             antisense overhang AG = revcomp(upstream CU)

Keeping the overhang bases native to the target context has two reasons. First, an artificial thymidine dimer is unrelated to the target, so there is no basis on which to predict what those two nucleotides interact with in the real molecule during off-target assessment. Second, a native overhang extends complementarity to the target RNA two nucleotides beyond the core, which keeps 3'-end recognition and loading geometry interpretable.

The acceptance test is not length but full-length reverse complementarity. Two strands that are each other's reverse complement across a 21 bp core meet the specification; a sequence that is merely 23 nt long does not. This check runs automatically at export, and keeping sequences produced by paths that bypass that check out of the synthesis order is the essence of specification control.

2.7 Progressive refinement — where computational cost is placed

The pipeline places the most expensive evaluations last. Enumeration and first-pass filters run over thousands of candidates in seconds; transcriptome-wide off-target scanning and deep-model inference run over a few dozen. This is not merely a performance optimisation but an allocation of information value: spending precise evaluation on candidates that will obviously fail wastes compute and, more importantly, forfeits precision that could have been applied to the finalists.

Stage

Scale

Nature of the evaluation

Kind of rejection

Enumeration and composition filters

thousands

Deterministic rules

GC composition out of range, homopolymer runs, immune-stimulatory motifs

Rule-based efficacy scoring

hundreds

Ensemble of literature-derived positional rules

Insufficient thermodynamic asymmetry, positional rule violations

Population variant masking

~150

Database lookup

Common polymorphism at the target site, overlap with pathogenic variants

Structure and thermodynamic filters

~30

Free-energy computation

Self-hairpin, seed binding too strong or too weak, poor local accessibility

Off-target screening

~30

Transcriptome-wide alignment plus mechanistic interpretation

Seed complementarity in unrelated genes, junction-proximal hits

Accessibility scoring

~30

Ensemble accessibility plus interaction energy

Target site buried in structure

Deep-model inference

~30

Ensemble of pretrained models

Low predicted potency

Chemistry and regulatory checks

top few

Rules plus databases

Chemistry constraint violations, preclinical species requirements unmet

3. Modality I — siRNA / RNAi (Double-Stranded Catalytic Silencing, RISC Cleavage)

3.1 Form and definition

The common name is siRNA (small interfering RNA), and the operating principle is called RNAi (RNA interference). A synthetic duplex with a 21 bp double-stranded RNA core carrying a 2 nt overhang at each end, 23 nt per strand. The two strands are not symmetric: one (the guide, antisense) is loaded into an Argonaute protein and recognises the target, while the other (the passenger, sense) is discarded during loading. The molecule is therefore dosed as a duplex but acts as a single strand, and controlling which of the two is loaded is the first task of design.

The property that most fundamentally separates this modality from every other is catalysis. The loaded guide–protein complex cleaves a target and then moves on to the next one rather than being consumed. One molecule removes many, and that turnover produces strong repression at low molar concentration. This is why the required dose differs fundamentally from a stoichiometrically acting steric blocker.

3.2 Molecular mechanism

3.2.1 Loading and strand selection

In the cytoplasm the synthetic duplex associates with an Argonaute protein to form the RNA-induced silencing complex. During assembly, whichever end of the duplex frays more easily opens first, and the strand whose 5' terminus sits at that loose end is retained as the guide. This asymmetry rule was reported independently in the same year by Khvorova and colleagues (Cell 2003) and Schwarz and colleagues (Cell 2003), and has underpinned every rational siRNA design approach since.

In design the rule is implemented as a free-energy differential. Nearest-neighbour free energies are computed for the four or five terminal base pairs at each end of a candidate duplex, and designs in which the intended antisense 5' end is the destabilised one are rewarded. Candidates with symmetric ends are penalised: both strands compete for loading, halving the effective on-target dose and doubling the passenger-derived off-target burden.

Loading also depends on the identity of the 5' nucleotide itself. The MID domain of Argonaute recognises the guide's 5' nucleotide in a base-specific manner, and crystallographic work reported a preference for uridine or adenosine (Frank, Sonenberg and Nagar, Nature 2010). The practically important point is that this 5' nucleotide is buried in the MID pocket and does not base-pair with the target. A strong candidate whose 5' end is guanosine or cytidine therefore need not be discarded; the standard remedy is to substitute that one position with uridine. Where such a substitution is made, the claim of 'full complementarity' must be restated as 'fully complementary from position 2 onward, with position 1 as the MID residue'.

3.2.2 Cleavage — what determines the rate

Once loaded, the guide searches for its target, and on encountering a fully complementary one the PIWI domain of Argonaute cleaves the target phosphodiester bond. The cleavage rate constant varies substantially with sequence; a recent quantitative study using 22-nt guides reported over 250-fold variation even among fully complementary targets (Wang and Bartel, Mol Cell 2024). The positional determinants identified in that work are as follows.

Guide position

Preferred identity

Reported rate difference

Interpretation

10

purine (A or G)

8.2 – 9.4×

The single strongest determinant. The evidence supports a purine/pyrimidine dichotomy; there is no basis for distinguishing A from G

17

W (A or U)

2.1 – 6.3×

G and C are the slow identities; there is no basis for distinguishing C from G

7

W (A or U)

1.1 – 4.4×

Both G and C are slow. Treating position 7 = C as optimal would equate an identity the literature calls slow with the fastest one

6–7 junction

weak pairing or a mismatch

partially redundant with the above determinants

A backbone kink driven by a sugar pucker change at position 6 promotes cleavage — the same mechanism as the position 7 determinant

All three aligned

4.7 – 51×

Cumulative effect observed in triple variants

These determinants were independently corroborated by a 2026 cryo-electron microscopy study of catalytic activation in human Argonaute 2. That work showed that guide–target base pairing alone is insufficient for slicing and that duplex distortion is required, and reported that a pyrimidine at target position 10 optimally aligns a catalytic residue. A pyrimidine on the target corresponds to a purine on the guide, exactly matching the position-10 determinant above. The same study observed that a kink after guide nucleotide 6 releases the seed-only pairing conformation and promotes the extended pairing catalysis requires, corroborating the 6–7 axis structurally.

3.2.3 A common misreading of central pairing strength

Design guidance frequently advises raising GC content in the central region to strengthen pairing. This is a misreading that comes from citing the literature without its conditional clause. Wang and Bartel (2024) state explicitly that formation of a continuous helix does not limit the cleavage rate of fully complementary targets, and report that hydroxyl-radical footprinting profiles for four different guides were similar to one another despite a 250-fold spread in rate. Thermodynamic conformational occupancy is not the rate-limiting step for fully complementary cleavage.

The widely quoted 600-fold figure comes from a different condition. It is the range of tolerance to 3' mismatches, measured on 16 bp targets where complementarity ends at target position 16. Guides with strong central pairing barely slow on such partially complementary targets; guides with weak central pairing slow by more than a hundredfold, and in the extreme case by 600-fold. The original text is conditional: pairing beyond position 16 is dispensable for efficient slicing only when the central region has high predicted pairing stability.

Two design implications follow, and both run counter to conventional advice. First, in a fully complementary design central GC content is not a determinant of on-target potency. Second, the direction inverts: strong central pairing means high tolerance to 3' mismatches, which means the guide also cleaves partially complementary off-targets efficiently. From a specificity standpoint, rewarding central GC promotes candidates that cleave off-targets well. The correct form of the axis is not a main-effect bonus but a conditional risk term — flagging combinations of weak central pairing with weak 3' pairing — and it belongs on the off-target risk axis, not the on-target score.

The structural evidence points the same way. The catalytic activation study reports that expansion of the central major groove positions the scissile phosphate, and GC-rich duplexes resist distortion, so structure does not support a central GC bonus either. A separate structural study reports that on fully complementary binding an N-domain rotation licenses rapid slicing — but that licensing is a function of complementarity, not of GC content.

3.2.4 The mechanism of off-targets — seed-mediated repression

Binding that is not fully complementary also represses. When guide positions 2–8, the seed, are complementary to a target's 3' untranslated region, the result is miRNA-like translational repression and transcript destabilisation. Jackson and colleagues (Nat Biotechnol 2003) first showed the scale of this by microarray, and Birmingham and colleagues (Nat Methods 2006) established that 3' UTR seed matches, rather than overall identity, explain the off-target signature.

This mechanism imposes two design requirements. First, off-target assessment must be performed at seed granularity rather than by full-length alignment: a single seed 7-mer occurs in hundreds of transcripts, so a candidate with no full-length hits can still carry a broad burden. Second, it becomes possible to introduce a modification that weakens binding within the seed and thereby selectively lowers seed-mediated repression. That strategy exploits an asymmetry — the effect on a fully complementary on-target is small, while the effect on seed-only binding is large.

3.2.5 Innate immune stimulation

Synthetic double-stranded RNA can stimulate innate immunity in a sequence-dependent way. Judge and colleagues (Nat Biotechnol 2005) showed that particular U- and GU-rich motifs trigger Toll-like-receptor-mediated cytokine responses, and reported that 2'-O-methyl substitution suppresses them. In design, known stimulatory motifs are filtered at enumeration and residual risk is lowered by 2'-O-methyl placement during the chemistry stage. There is a constraint: 2'-O-methyl substitution early in the seed (positions 2–5) has been reported to reduce Argonaute loading, so placement must balance immune suppression against loading efficiency.

3.3 Clinical precedent and patent landscape

This modality has the deepest approval record among RNA therapeutics. Since the first approval in 2018 of a lipid-nanoparticle formulation against transthyretin, a succession of N-acetylgalactosamine-conjugated products targeting hepatocytes has established 'hepatic silencing' as a mature development path.

Year

Target

Delivery

Indication area

What the approval established

2018

TTR

Lipid nanoparticle, intravenous

Hereditary transthyretin amyloidosis

First demonstration that siRNA produces therapeutic effect in humans

2019

ALAS1

GalNAc conjugate, subcutaneous

Acute hepatic porphyria

Practicality of the subcutaneous GalNAc route

2020

HAO1

GalNAc conjugate, subcutaneous

Primary hyperoxaluria type 1

Repeat application in a rare metabolic disease

2020–2021

PCSK9

GalNAc conjugate, subcutaneous

Hypercholesterolaemia

Twice-yearly dosing — durability as a clinical differentiator

2022

TTR

GalNAc conjugate, subcutaneous

Hereditary transthyretin amyloidosis

Generational replacement of delivery on the same target

2023

LDHA

GalNAc conjugate, subcutaneous

Primary hyperoxaluria type 1

Reproducibility of the route

The patent landscape has three central axes. The first is chemical placement patterns — claim families specifying alternating positional arrangements of 2'-O-methyl and 2'-fluoro together with terminal phosphorothioate placement, the so-called enhanced stabilisation chemistry family. The second is the triantennary N-acetylgalactosamine ligand and its linker architecture. The third is off-target mitigation modifications such as acyclic sugar analogues placed in the seed. Because these claims are generally constructed as 'a specific modification at a specific position', emitting an explicit per-position chemistry map at design time is itself the input to a freedom-to-operate review.

Patent statements here are informational and are not legal opinion. Freedom-to-operate depends on jurisdiction and claim construction and is the domain of qualified counsel.

3.4 Chemistry requirement — mandatory, with positional rules

Chemical modification is not optional in this modality. An unmodified RNA duplex degrades within minutes in plasma, stimulates innate immunity, and is not taken up by cells. Modification serves four purposes: nuclease resistance, immune suppression, seed off-target mitigation, and providing an attachment point for conjugation.

Modification

Role

Typical placement

Caution

2'-O-methyl

Nuclease resistance, immune suppression

Alternating across both strands

Heavy placement early in the seed (2–5) reduces loading

2'-fluoro

Maintains binding affinity, local stabilisation

Alternating with 2'-O-methyl

Excessive total content carries cellular and cost burden

Phosphorothioate

Terminal exonuclease resistance, protein binding

1–2 linkages at each terminus

Beyond the termini it buys nothing under encapsulation and costs potency

Acyclic sugar analogue

Weakens seed binding to mitigate off-targets

One specific position within the guide seed

Small effect on-target, but position choice is decisive

5'-terminal stabilisation

Loading efficiency and terminal protection

Guide 5' terminus

An enhancement for duplexes; an absolute requirement for single strands (see Chapter 6)

N-acetylgalactosamine conjugation

Hepatocyte receptor-mediated uptake

Passenger strand terminus

Linker chemistry (O/S/C) changes in vivo stability

The most important positional rule is protection of the cleavage site. Argonaute cleaves the target phosphate corresponding to guide positions 10–11, so a modification there that impedes cleavage damages catalysis, while leaving the region entirely bare lowers stability. Design resolves the conflict with the constraint 'only modifications that do not block cleavage are permitted'.

3.5 Delivery requirement — required

A duplex cannot carry a fully phosphorothioate backbone: duplex stability and Argonaute loading would both be compromised. The carrier-free uptake strategy that works for single strands therefore barely applies here, and a delivery vehicle is effectively mandatory. There are two branches.

The delivery choice feeds back into chemistry. Under encapsulation the particle performs both protection and uptake, so phosphorothioate beyond the termini costs potency for no gain; on a conjugate route the receptor performs part of the uptake, so the required backbone burden is intermediate. Chapter 6 treats this relationship as a quantitative window.

3.6 Design deliverables

3.7 Advanced capabilities — the deep ensemble and special design modes

3.7.1 Why five lineages are run at once

The efficacy prediction layer runs five models from different architectural lineages simultaneously in the ACTIVE state. The reason for not simply picking the single best-performing model is not accuracy but independence of bias. Models trained on the same data tend to be wrong in the same places, and that shared error is not removed by ensembling. Agreement among models with different representations and inductive biases, by contrast, generalises better than the confidence of any one of them.

Model lineage

What it sees

Its distinctive contribution

Transformer with RNA language-model embeddings

Long-range context across the whole sequence, transferring a pretrained RNA representation

Off-target-aware training — a representation optimised jointly for efficacy and specificity

Transformer combined with convolution

Local motifs (convolution) combined with global context (attention)

Retains short positional patterns without losing global context

Chemistry-aware multi-view with cross-attention

Sequence and chemical modification map encoded as separate views, then cross-referenced

The only lineage that addresses modified molecules directly — models trained on unmodified data extrapolate when chemistry changes

In-house spline-based model

B-spline activations expressing positional non-linearity locally

Smooth positional dependence that connects interpretively to the rule-based axes

Seed-mediated off-target model

A learned representation of seed–target binding

Reinforces the paralog / ortholog / unrelated three-bucket classification (section 2.5)

The five outputs are combined by calibrated weighting rather than plain averaging. Because model predictions are correlated, a plain average leaves the correlated error intact, while weighting that reflects the correlation structure reduces variance further. The deliverables present the individual model scores alongside the combined score so that candidates with large inter-model disagreement can be identified — large disagreement is itself the signal that a candidate lies outside the training distribution.

3.7.2 Cross-checking the rule axis against the learned axis

Presenting the rule-based ensemble and the deep ensemble separately follows the same logic. When both families rank the same candidate highly, confidence is high; where they diverge, that point deserves review. In particular, a candidate the rules like and the learned models score low often has no similar case in the training data, while the converse may indicate a contextual effect the rules do not capture. Collapsing the two into one number destroys that information.

3.7.3 Allele-selective design

In dominant-negative disease only the mutant allele should be lowered while the normal allele is preserved. The two alleles are identical apart from one position, so the guide must be made to discriminate that single base. Design places the variant at the position within the guide where discrimination is greatest and, where necessary, introduces a second artificial mismatch to weaken binding to the normal allele further.

The practical difficulty is a conflict between discrimination ratio and absolute potency. Placing the variant at a highly discriminating position raises selectivity but can also lower absolute potency against the mutant allele. The deliverables present predicted binding and potency for both alleles side by side so the selectivity ratio can be read directly.

3.7.4 Virus mode — designing on conserved regions

For a rapidly varying pathogen, designing against a single reference sequence selects resistance variants quickly. Virus mode takes a user-supplied multi-strain sequence set as input, identifies regions of low variation from the cross-strain alignment, and restricts design to them. In parallel it screens off-targets against host transcriptomes (human, mouse, non-human primate) so that only candidates leaving host genes untouched remain.

The deliverables show, per candidate, in how many strains it is fully conserved and where mismatches arise in which strains. That figure is the direct index of breadth of efficacy and the basis for ranking a candidate that is robust across many strains above one that is perfect against a single reference.

3.7.5 Cross-species reactivity — 27 species

The design layer holds transcriptome references for 27 species including human and evaluates cross-reactivity through predefined species combination modes. The list covers rodents, several non-human primates, companion animals, livestock, poultry and fish, so that preclinical species selection and veterinary indication extension are addressed together. For each species the ortholog transcript is aligned and per-candidate mismatch counts and positions are computed, with cross-reactivity interpreted differently depending on whether the mismatch falls in the seed.

3.7.6 The chemistry blueprint as checkable rules

The chemistry blueprint is emitted not as free description but as a set of checkable rules. Each is judged pass, warn or fail with the supporting figure shown.

3.7.7 Regulatory pre-flight and synthesis impurity projection

Two development-readiness artifacts accompany the top candidates. One is an automatic verdict on the preclinical species requirement, confirming that the chosen species combination includes one rodent and one non-rodent and that the candidate has sufficient complementarity in those species. The other is a synthesis impurity projection, giving the relative prominence of deletion sequences, addition sequences, incomplete sulfurisation and diastereomers derived from the sequence and chemistry (see Chapter 21). Both are drafts and do not replace final specification setting.

3.8 What this modality cannot solve

4. Modality II — ASO Gapmer (Single-Stranded Catalytic Silencing, RNase H1 Cleavage)

4.1 Form and definition

The common name is antisense oligonucleotide (ASO), and this particular architecture is called a gapmer. A single-stranded oligonucleotide of 16–20 nt divided internally into three compartments. The two wings are filled with 2'-modified nucleotides and carry target affinity and nuclease resistance; the central gap remains 2'-deoxy. Typical configurations are 5-10-5 or 5-8-5, and the backbone is fully or largely phosphorothioate. This three-compartment architecture is why the molecule is called a gapmer.

The reason the gap exists is singular. The host enzyme RNase H1 recognises DNA:RNA heteroduplexes and cleaves the RNA strand, and that recognition requires a contiguous 2'-deoxy stretch above a minimum length. Introducing even one 2'-modification into the gap destroys the heteroduplex and abolishes cleavage, demoting the molecule from a catalytic silencer to a plain steric blocker. Conversely, without wings the molecule lacks the affinity and stability to function in vivo. The architecture is therefore not an aesthetic choice but the spatial separation of two opposing requirements.

4.2 Molecular mechanism

4.2.1 Why it is catalytic

Once RNase H1 has cleaved the target RNA, the heteroduplex dissociates and the oligonucleotide is released intact. It then forms a new heteroduplex with the next target molecule, so it has turnover. In this respect it shares the character of Modality I, but the catalyst differs: I is an Argonaute protein that cuts by itself, whereas II builds a structure that a host enzyme recognises and cuts. That difference determines where in the cell each modality operates.

4.2.2 Nuclear activity as the decisive advantage

RNase H1 is present in both nucleus and cytoplasm, with substantial nuclear activity. This modality can therefore address targets an Argonaute complex reaches poorly: unspliced pre-mRNA, intronic sequence, and long noncoding RNAs that remain in the nucleus. Practically this widens the target space considerably — a gene with no suitable site in its mature mRNA may have one in an intron or a 5' untranslated region, and a nuclear-retained transcript is difficult to address by any other mechanism.

Site selection is also freer. Modality I is effectively fixed in length and geometry by Argonaute's requirements, whereas a gapmer's length, gap width and wing composition can all be tuned to the local structure and binding strength of the target site. When a target site is buried in strong secondary structure, moving the wing chemistry toward higher affinity to improve invasion is a legitimate response.

4.2.3 Cleavage site and the meaning of coverage

Argonaute cuts one phosphodiester bond precisely; RNase H1 cuts at several positions within the window the gap covers. The 'cleavage site' is therefore an interval rather than a single coordinate, and is reported as a window. The arithmetic of how many transcripts a cut destroys is mechanism-independent — it only asks how many transcripts carry that sequence — but the notation of where the cut occurred is mechanism-specific.

4.2.4 Carrier-free uptake — the backbone is the uptake mechanism

A phosphorothioate backbone binds broadly to plasma and cell-surface proteins. That binding retains the oligonucleotide in circulation and pushes it into endocytic routes. Stein and colleagues (Nucleic Acids Res 2010) showed that simply exposing cultured cells to oligonucleotides without transfection reagent produces gene silencing, and named the phenomenon gymnosis. The observation fundamentally changed this modality's development path, because it allowed entry into preclinical work without a separate delivery-vehicle programme.

Carrier-free uptake is not free, however. The phosphorothioate fraction that creates uptake also creates class toxicity arising from protein binding — complement activation, platelet effects, hepatotoxicity. Design in this modality is therefore not the question 'how much PS is needed' but 'which range is productive', and that range depends on the delivery route. Chapter 6 gives the quantitative model.

4.2.5 Hepatotoxicity — this modality's development bottleneck

Depending on sequence and chemistry, gapmers can cause hepatocyte toxicity, and this is the single most common cause of failure to enter the clinic in this modality. Mechanistic work suggested the toxicity is associated not with chemical burden alone but with RNase H1-mediated activity itself — that is, with cleavage of unintended transcripts (Burel and colleagues, Nucleic Acids Res 2016). A tendency for high-affinity bicyclic modifications to raise the risk has also been reported (Kasuya and colleagues, Sci Rep 2016).

Omitting this axis at the design stage produces a predictable, repeating failure. Ranking on potency alone brings candidates with high-affinity chemistry and strong binding to the top, and those are precisely the properties correlated with toxicity risk. Evaluating potency and toxicity in parallel is therefore not optional in this modality, and design emits a multi-tier risk classification based on sequence motif and chemistry combination alongside potency.

4.3 Clinical precedent and patent landscape

This modality has the longest approval history. The first product, dosed locally into the eye, was approved in 1998 and later withdrawn from the market; a systemically dosed gapmer approved in 2013 brought class-toxicity management into focus. From 2018 onward, generational changes in chemistry and conjugation drove approvals in central nervous system and metabolic disease.

Year

Target

Route

Indication area

What the approval established

1998

A viral transcript

Intraocular, local

Cytomegalovirus retinitis

First demonstration that an antisense oligonucleotide can be a medicine

2013

APOB

Subcutaneous, carrier-free

Homozygous familial hypercholesterolaemia

Feasibility of systemic carrier-free dosing, and the necessity of class-toxicity management

2018

TTR

Subcutaneous, carrier-free

Hereditary transthyretin amyloidosis

Direct competition with Modality I on the same target

2023

SOD1

Intrathecal

SOD1-associated amyotrophic lateral sclerosis

Establishment of direct central nervous system dosing

2023

TTR

GalNAc conjugate, subcutaneous

Hereditary transthyretin amyloidosis

Receptor-targeted conjugation applied to a gapmer, sharply lowering dose

2024

APOC3

GalNAc conjugate, subcutaneous

Familial chylomicronaemia syndrome

Extension of conjugated gapmers into metabolic disease

The patent landscape centres on wing chemistry. A second-generation architecture using 2'-O-methoxyethyl wings and a 2.5-generation architecture using constrained-ethyl bicyclic sugars each carry substantial claim families. On top of these sit 5'/3' N-acetylgalactosamine conjugation and its linker architecture. A more recent layer specifies stereochemical control of the phosphorothioate backbone — that is, defined stereoisomers rather than a racemic mixture.

4.4 Chemistry requirement — mandatory and structurally enforced

In this modality chemistry is not an add-on but the definition of the molecule. An unmodified DNA oligonucleotide has a plasma half-life of minutes and insufficient target affinity to function systemically. The gapmer architecture is the resolution of that problem into 'where to modify and where not to'.

Compartment

Chemistry

What it carries

Consequence of violation

5' wing (typically 3–5 nt)

2'-O-methoxyethyl, bicyclic sugars (LNA, cEt) and similar

Target binding affinity, exonuclease resistance

Too short or too low-affinity and binding is insufficient

Central gap (typically 8–10 nt)

2'-deoxy maintained, phosphorothioate backbone

RNase H1 recognition and cleavage

Introducing a 2'-modification abolishes catalysis — the molecule becomes a steric blocker

3' wing (typically 3–5 nt)

Same class as the 5' wing

Affinity, 3'-exonuclease resistance

Failure to protect the 3' terminus causes rapid degradation

Backbone overall

Phosphorothioate, full length or nearly so

Plasma stability, protein-binding-driven carrier-free uptake

Too little and uptake fails; too much and class toxicity appears

Terminal conjugate

N-acetylgalactosamine or other targeting ligand

Receptor-mediated hepatocyte uptake, sharply reducing dose

Insufficient linker stability causes premature loss of the ligand

High-affinity bicyclic modifications are powerful but double-edged. They allow sufficient affinity from a short oligonucleotide, but they also strengthen unintended binding at partially complementary sites and correlate with hepatotoxicity risk. Wing chemistry should therefore be chosen as 'as much as the target site structure requires' rather than 'the strongest available', and design determines wing chemistry together with the local accessibility of the target site.

4.5 Delivery requirement — optional; unnecessary on many routes

The practical strength of this modality is that a development path exists without a delivery vehicle. Four routes coexist.

Route

Delivery mode

When it fits

Constraint

Systemic carrier-free

Subcutaneous or intravenous, no carrier

Tissues with efficient uptake such as liver and kidney; rapid preclinical entry

Requires large doses, widening the scope for class toxicity

Receptor conjugate

GalNAc conjugate, subcutaneous

Hepatocyte targets; dose falls sharply versus carrier-free

Does not extend to non-hepatic tissue

Local direct administration

Intrathecal, intraocular, inhaled

Central nervous system, eye, airway; minimises systemic exposure

The administration itself is invasive or device-dependent

Particle encapsulation

Lipid nanoparticles and similar

Tissues the three routes above do not reach

Under encapsulation the phosphorothioate requirement falls, so the chemistry must be redesigned

The fourth row matters. For the same sequence, whether it goes carrier-free or in a particle changes the optimal chemistry. On the carrier-free route most of the backbone must be phosphorothioate for uptake to occur; under encapsulation the particle performs protection and uptake, so the same level of phosphorothioate sacrifices potency and safety for nothing. Optimising chemistry without fixing the route yields a value optimal for neither.

4.6 Design deliverables

4.7 Advanced capabilities — five mechanisms, toxicity prediction, splicing-interference avoidance

4.7.1 The five supported mechanisms

The design layer distinguishes five mechanisms explicitly, and once the mechanism is set the chemistry rules and evaluation axes change automatically. Specifying the wrong mechanism does not raise an error; it simply produces a different molecule, which makes mechanism specification the first decision of the design.

Mechanism

What it does

Mandatory structural requirement

Representative clinical lineage

Gapmer

Recruits RNase H1 to cleave the target RNA

A central 2'-deoxy gap is mandatory

Systemic metabolic disease, hepatic targets, central nervous system

Splice switching

Changes the splicing decision

No gap; uniform high-affinity modification

Treated in depth as Modality III (Chapter 5)

Steric block

Physically obstructs translation initiation or protein binding

No gap; uniform modification

Manipulation of upstream open reading frames in 5' untranslated regions and similar

Anti-miR

Sequesters an endogenous miRNA

No gap; fully modified backbone

Treated in depth as Modality VI-b (Chapter 10)

Dual-guide class

Places two binding elements on one molecule

Design-specific

Exploratory

These five are handled in one layer because the chemistry repertoire and off-target assessment are shared. The activity axis, however, diverges completely: the gapmer rewards catalytic recognisability while the other four reward uniform occupancy (section 6.3).

4.7.2 The hepatotoxicity risk index — a weighted motif sum

Hepatocyte toxicity is the most common single cause of failure to enter the clinic in this modality, so design quantifies toxicity risk in parallel with potency. The method is an index that multiplies the occurrence frequency of particular short motifs by weights and sums them, with the weights reflecting the strength of the relative risk signal reported in the literature. The index is for relative ranking and is not calibrated to an absolute risk probability.

Risk band

Motifs

Weight

High

TCCC · TGCC

3.0

High

TCCT · CCTCC

2.5

High

TCCA · TGCT

2.0

Moderate

AACC · GACC · GCCT

1.2

Moderate

CCTG · AGCC

1.0

Low

CCAG · GGCA · GGCT

0.8

Low

CAGG · GGAG · GCAG

0.6

The deliverables present the index together with a list of which motifs occurred and how often. That decomposition matters because the response differs: if the index is high because of a single high-risk motif, shifting the site by one position may resolve it, whereas an accumulation of low-risk motifs requires changing the target site altogether.

This index is correlated with the potency axis, and that is stated explicitly. Candidates with high-affinity chemistry and strong binding tend also to carry higher toxicity risk, so ranking on potency alone systematically brings risky candidates to the top. Presenting the two axes in parallel is the only way to block that bias.

4.7.3 Avoiding splicing interference in advance

When a gapmer targets pre-mRNA and the target site overlaps a splicing regulatory element, splicing can be perturbed unintentionally. Design scans candidate sites with position weight matrices for exonic splicing enhancer binding proteins to assess that risk in advance. The matrices used are four well-validated SR protein recognition motif families, each with its own threshold.

Matrix

Protein family

Role

SF2/ASF

SR protein

A principal enhancer for exon recognition and splice site choice

SC35

SR protein

Involved in exon definition

SRp40

SR protein

Promotes use of weak splice sites

SRp55

SR protein

Exonic enhancer recognition

In addition, the chemistry and sequence profiles of representative approved products are built in as reference presets, so a new design can be compared with clinical precedent on the same axes. The presets are chosen to represent distinct development contexts — central nervous system gapmers, metabolic disease gapmers, splice-switching lineages.

4.7.4 Junction-aware off-target classification

An ordinary off-target scan treats hit positions only as transcript coordinates. For a gapmer, however, a hit near a splice junction carries different risk, because heteroduplex formation there can perturb the splicing of that gene rather than merely cleaving one transcript. Design consults known junction coordinates in the host transcriptome and classifies junction-proximal hits at a separate grade.

4.7.5 Generations and the choice of wing chemistry

Generation

Wing chemistry

Character

Selection criterion

Second

2'-O-methoxyethyl

The deepest clinical precedent and the richest safety data

The default choice where the target site is not especially difficult

2.5

Constrained ethyl and related bicyclics

High affinity permitting a shorter molecule

Where the target site is buried in structure or a short molecule is needed

High-affinity bicyclic

Locked nucleic acid family

The highest affinity

Requires care — correlation with hepatotoxicity risk has been reported

Mixed placement

Bicyclics at selected positions only

A compromise between affinity and risk

Placement must avoid creating a contiguous 2'-deoxy stretch

4.7.6 Retrovirus mode

For retroviral targets a dedicated path forces the gapmer mechanism, identifies conserved regions from a user-supplied multi-sequence set and restricts design to that range. Off-target screening against the host transcriptome runs in parallel, and because proviral sequence integrates into the host genome, checking against host-genome-derived similar sequences is treated as especially important.

4.8 What this modality cannot solve

5. Modality III — SSO / Splice-Switching Oligonucleotide (Non-Cleaving Steric Block)

5.1 Form and definition

The common name is SSO (splice-switching oligonucleotide), also called splice-switching antisense or an exon-skipping oligonucleotide. A single-stranded oligonucleotide of 18–25 nt in which every position carries a uniform high-affinity modification. The form matches Modality II but the internal composition is its inverse: where a gapmer must place a 2'-deoxy window at the centre, this modality leaves no 2'-deoxy residue anywhere, because the objective is not to recruit a nuclease.

Three chemical branches are available: uniform 2'-O-methoxyethyl with a phosphorothioate backbone; a mixmer incorporating bicyclic sugars; and phosphorodiamidate morpholino oligomers, in which the sugar–phosphate backbone itself is replaced by morpholine rings and phosphorodiamidate linkages, yielding a charge-neutral molecule. Charge neutrality sharply reduces protein binding and class toxicity, at the cost of lower carrier-free uptake efficiency.

5.2 Molecular mechanism

5.2.1 Splicing is a decision, not a fixed procedure

The conversion of pre-mRNA into mature mRNA is a competitive decision rather than a fixed sequence of steps. The spliceosome recognises donor and acceptor signals, but that recognition is not determined by the signal sequence alone. Proteins binding to regulatory elements scattered through the exon and adjacent introns — exonic and intronic splicing enhancers and silencers — push and pull on whether each site is used. Which exon is included is therefore the outcome of a probabilistic competition, and tipping that balance changes the outcome.

This modality tips exactly that balance. When the oligonucleotide binds a regulatory element and physically masks it, the protein that should bind there cannot; if the element is inhibitory, inhibition is relieved and the exon is included, and if it is enhancing or is a splice signal itself, that site goes unused and the exon is excluded. Changing the decision without destroying the target is the essence of the modality.

5.2.2 Five forms of intervention

Form

Element masked

What happens

Therapeutic logic

Exon inclusion

Intronic splicing silencer

Inhibition is relieved and a normally skipped exon is included

Increases the amount of functional full-length protein

Exon skipping

Donor/acceptor signals or an exonic enhancer

The exon is excluded from the mature transcript

Restores a reading frame broken by a deletion, yielding partially functional protein

Pseudoexon suppression

A latent splice site created by a deep-intronic variant

The artificially created exon is not used

Restores the proportion of normal transcript

Intron retention resolution

Elements that promote retention

A retained intron is spliced normally

Matures a nucleus-trapped transcript and restarts protein production

PolyA switching

An alternative polyadenylation signal

A different 3' end is selected

Adjusts isoform ratio or transcript stability

5.2.3 Why precision of the target window is decisive

Regulatory elements are typically 6–20 nt long, and shifting the target by a few nucleotides abolishes the effect. The intronic splicing silencer targeted in spinal muscular atrophy is the canonical example: therapeutic design only became possible once the exact coordinates and inhibitory function of that element had been established (Singh and colleagues, Mol Cell Biol 2006; Hua and colleagues in subsequent work). This modality therefore demands not 'target this gene' but 'target this coordinate in this gene', and selection of the target window determines most of the design quality.

For that reason a generic 'risk that a cryptic splice site could arise' score is of limited help in candidate selection: it belongs to a region rather than to a candidate. What is actually needed is how far this particular oligonucleotide moves splicing at this target, and design computes that as an in-silico masking differential — the window a candidate covers is masked, splice-site probabilities are recomputed, and the change in donor and acceptor probability is attributed to that candidate. Two candidates with identical generic risk scores frequently have very different masking differentials, and that difference reorders the ranking.

5.2.4 The reading-frame constraint

Exon-skipping design carries one further genetic constraint. Excluding an exon shortens the sequence by its length, and unless that length is a multiple of three the reading frame shifts, rendering everything downstream meaningless and introducing a premature stop. The clinical manifestation of this rule is the split between severe and milder phenotypes of muscular dystrophy according to the position and size of the deletion. Therapeutic design therefore computes, per patient subgroup, which exon must be skipped to restore the frame, and this is why a reading-frame validator is part of the deliverable set.

5.2.5 Interpretive consequences of not cleaving

Because this modality does not destroy its target, the 'cleavage coverage' metric used elsewhere cannot be read the same way. Here coverage means the number of transcripts carrying the site — binding range, not destruction range. Two consequences cause repeated confusion in practice and are stated explicitly. First, a design targeting an intronic element is structurally reported at zero coverage because that sequence is absent from every mature transcript; that is correct behaviour, not a defect. Second, high coverage in an exon-skipping design means 'widely bound', not 'widely knocked down'.

5.3 Clinical precedent and patent landscape

This modality established the therapeutic logic of 'repair rather than reduce'. In 2016 a treatment for spinal muscular atrophy and an exon-51 skipping treatment for muscular dystrophy were approved in the same year, validating two different forms of intervention simultaneously; further approvals followed in the muscular dystrophy family targeting different exons.

Year

Form

Target

Chemistry and route

What the approval established

2016

Exon inclusion

SMN2 intron 7 silencer

Uniform 2'-MOE, intrathecal

That splice manipulation can change the phenotype of a severe genetic disease

2016

Exon skipping

DMD exon 51

PMO, intravenous

Clinical entry of charge-neutral chemistry and the frame-restoration strategy

2019

Exon skipping

DMD exon 53

PMO, intravenous

Extension to further patient subgroups

2020

Exon skipping

DMD exon 53

PMO, intravenous

Multiple products against the same exon

2021

Exon skipping

DMD exon 45

PMO, intravenous

Direction toward a complete target-exon portfolio

A pseudoexon-suppression strategy targeting a deep-intronic variant in retinal disease has entered clinical development, and is also an example of minimising systemic exposure through local administration. The patent landscape has three layers: claims over the target element and its coordinates, claims over uniform high-affinity and morpholino chemistry, and sequence claims over specific exons. The third layer bears directly on design freedom, so patent-family advisories are included in the deliverables.

5.4 Chemistry requirement — mandatory and uniform

Chemistry here serves three interlocking requirements: affinity sufficient to bind the regulatory element and displace its protein; nuclease resistance sufficient to survive until the target is reached; and, most importantly, not recruiting RNase H1. The third creates the uniformity constraint — any stretch left 2'-deoxy could be recognised as a heteroduplex and the target cleaved, at which moment the molecule ceases to be a splice switcher and becomes a gapmer.

Chemistry

Character

Advantages

Considerations

2'-O-methoxyethyl + phosphorothioate

Anionic, high affinity

The deepest clinical precedent; systemic and intrathecal behaviour well characterised

Phosphorothioate class toxicity must be managed

Phosphorodiamidate morpholino

Charge-neutral, non-natural backbone

Low protein binding, so little class toxicity and little immune stimulation

Low carrier-free uptake efficiency; requires high doses or a cell-penetrating conjugate

Bicyclic mixmer

Bicyclic sugars placed at selected positions

High affinity from a short oligonucleotide

Placement must avoid creating a contiguous 2'-deoxy stretch

2'-O-methyl + phosphorothioate

Classical uniform modification

Simple synthesis and lower cost

Relatively lower affinity, compensated by length

5.5 Delivery requirement — usually unnecessary

Most approvals in this modality are dosed without a carrier. Central nervous system targets are reached by intrathecal administration bypassing the blood–brain barrier; muscle targets rely on tissue distribution after systemic dosing; ocular targets achieve local concentration by intravitreal injection. These are cases in which the delivery problem was solved by route rather than by vehicle.

That does not mean delivery is easy. For targets where sufficient tissue concentration is hard to reach after systemic dosing — muscle, heart — the required dose becomes very large, and this operates as the practical ceiling on the modality. Efforts to improve it through cell-penetrating peptide conjugation or receptor-targeted conjugation are ongoing, and from a design standpoint the effect of chemistry and attachment point on folding and binding must be assessed together.

5.6 Design deliverables

5.7 Advanced capabilities — quantitative splicing assessment and clinical presets

5.7.1 The five mechanisms and their clinical anchors

Mechanism

What it masks

Clinical anchor

The key design decision

Exon skipping

An exonic enhancer or the donor/acceptor signal

Muscular dystrophy exons 45 / 51 / 53

Which exon must be skipped for the frame to be restored

Exon inclusion

An intronic splicing silencer

The spinal muscular atrophy intron 7 element

The exact coordinates of the silencer — a few nucleotides off and the effect is lost

Pseudoexon suppression

A cryptic splice site created by a deep-intronic variant

Deep-intronic variants in retinal disease

Whether to mask the cryptic site or the branch point

Intron retention resolution

Elements promoting retention

Exploratory

Whether retention is caused by a weak splice site or by a repressive element

PolyA switching

An alternative polyadenylation signal

Exploratory

Whether the switch changes isoform ratio or transcript stability

The spinal muscular atrophy case is the reference point for coordinate precision. The targeted intronic silencer sits in intron 7 at approximately +10 to +27, and therapeutic design became possible only after those coordinates and that repressive function had been established. This modality demands not 'target this gene' but 'target this 18-nucleotide window of this gene'.

5.7.2 Quantifying splice site strength — two layers

Design assesses splice site strength in two layers. The first is a maximum-entropy site score, computed over a 9-mer window for donors and a 23-mer window for acceptors. Because the model captures dependencies between positions, it separates weak from strong sites better than a simple position weight matrix. The second is a deep-learning site probability that takes far broader sequence context as input and predicts whether that position is actually used as a splice site.

The two are used together because they measure different things. The maximum-entropy score answers 'does this sequence look like a splice site'; the deep model answers 'will it be used in this context'. Latent sites exist that carry a strong signal yet go unused, and those are precisely the cryptic sites at risk of activation once an SSO is introduced.

5.7.3 The in-silico masking differential — the central metric

A generic cryptic risk score belongs to a region and therefore discriminates poorly between candidates. The masking differential masks the window a candidate covers, recomputes splice site probabilities and attributes the change in donor and acceptor probability to that candidate. The procedure is as follows.

1. Compute baseline donor and acceptor probabilities in the target pre-mRNA context
2. Generate a sequence with the window covered by the candidate SSO masked
3. Recompute probabilities on the masked sequence
4. delta(donor) = masked - baseline,  delta(acceptor) = masked - baseline
5. Judge whether the sign matches the intended direction (up for inclusion, down for skipping)

   -> Two candidates with identical generic risk scores frequently have very different deltas,
      and that difference reorders the ranking

Checking the sign as well as the magnitude matters. Judged on magnitude alone, a candidate that moves splicing strongly rises to the top — but if the direction is wrong, that candidate produces the opposite of the intended result. The deliverables present magnitude and sign together with a verdict on agreement with the intended direction.

5.7.4 Advance scanning for cryptic sites

A sweep of ±500 nt around the target window enumerates latent donors and acceptors and ranks them by strength. The subjects are sites currently unused but with signal strength in the middle range — weak signals are unlikely to be activated, and already-strong ones would be in use. When an SSO masks the canonical site the spliceosome looks for the next best, so these middle-strength sites are the real risk list.

5.7.5 The toxicity motif scanner

Uniform high-affinity chemistry combined with particular sequences creates three classes of risk, each scanned separately.

5.7.6 The reading-frame validator

In an exon-skipping design, if the length of the target exon is not a multiple of three the reading frame shifts and everything downstream becomes meaningless. The validator reads the exon structure of the target transcript, confirms the length of the exon to be skipped and returns a pass or fail on frame preservation. On failure it also searches for combinations in which skipping an adjacent exon as well restores the frame.

Automating this check is a practical necessity: the deletion range differs by patient subgroup, so which exon to skip changes per subgroup, and computing the combinations by hand invites omissions.

5.7.7 Clinical presets and patent landscape notation

For well-known targets, the target coordinates and mechanism of clinical precedent are built in as presets so that a new design can be compared with that precedent on the same axes. A preset carries the target intron or exon number, the relative coordinates of the target window, the mechanism and the reference literature.

This is also an area in which sequence claims exist, so designs near a preset are accompanied by advisories on the relevant patent families. The notation is informational and is not a freedom-to-operate opinion.

5.8 What this modality cannot solve

6. The Chemical Modification Layer

The preceding three chapters described the chemistry each modality demands individually. This chapter integrates those demands into a single decision structure. Chemical modification is not a matter of adding more of a good thing; it is an optimisation in which four objective functions conflict, and the shape of that conflict differs by modality and again by delivery route. The purpose of this chapter is to make that structure explicit.

6.1 The repertoire

Modifications applicable to a synthetic oligonucleotide fall on three axes: substitution at the ribose 2' position, substitution of the phosphodiester linkage, and terminal additions.

Axis

Modification

Physical effect

Typical use

Sugar 2'

2'-O-methyl (m)

Raises nuclease resistance, slightly raises affinity, suppresses immune stimulation

Baseline modification across all modalities

Sugar 2'

2'-fluoro (f)

Raises binding affinity, locks the C3'-endo conformation

Alternated with 2'-O-methyl in duplexes

Sugar 2'

2'-O-methoxyethyl (e)

High affinity and high stability, bulky

Gapmer wings, uniformly modified steric blockers

Sugar 2'

Locked nucleic acid and related bicyclics (l)

Very high affinity, strongly locked conformation

Where high affinity is needed from a short oligonucleotide

Sugar 2'

Constrained ethyl (c)

Bicyclic class, high affinity and relatively tolerant

Next-generation option for gapmer wings

Sugar 2'

Acyclic sugar analogue (g)

Locally weakens binding

Placed in the seed to mitigate off-targets

Sugar 2'

2'-deoxy (d)

Natural DNA; forms a heteroduplex

The gapmer central gap — the only place it is used deliberately

Backbone

Phosphorothioate (s)

Nuclease resistance, plasma protein binding, uptake mediation

All modalities; its fraction determines the delivery route

Backbone

Phosphodiester (o)

Natural; low stability but no potency cost

Internal linkages on encapsulated routes

Backbone

Phosphorodiamidate morpholino

Charge-neutral, almost no protein binding

Splice-switching family only

Terminal

5'-vinylphosphonate and related phosphate mimics

Metabolically stable 5'-phosphate mimic

Stabilises RISC loading; an absolute requirement for single-strand RISC

Terminal

Inverted deoxythymidine

Blocks 3'-exonucleases

Duplex 3' terminal protection

Terminal

N-acetylgalactosamine conjugate

Hepatocyte receptor ligand

Hepatic targets, subcutaneous

Terminal

Lipid anchor · polyethylene glycol

Confers amphiphilicity, delays renal clearance

Self-assembling conjugates, half-life extension

In the deliverables, chemistry is expressed as two strings: a sugar string of length L and a backbone string of length L−1. What matters is that this notation is complete at per-position granularity. A summary such as '60% 2'-O-methyl' is usable neither for a synthesis order nor for a patent review.

Example notation (21-mer guide)
  sugar :  mfmfmfmmmfmfmfmfmfmfm      (m = 2'-O-methyl, f = 2'-fluoro)
  ps    :  ssoooooooooooooooooss      (s = phosphorothioate, o = phosphodiester)
  5' end:  (vp)                       (metabolically stable 5'-phosphate mimic)

Gapmer example (5-10-5, 20-mer)
  sugar :  eeeee dddddddddd eeeee     (e = 2'-MOE, d = 2'-deoxy gap)
  ps    :  every linkage s

6.2 Four objective functions and their conflicts

Chemistry optimisation is formulated as a weighted sum over the axes below, with a fifth axis added on carrier-free delivery routes. Each axis is normalised to the interval 0 to 1 and computed from literature-derived rules.

Objective

What it rewards

What it penalises

Principal conflict

Stability

Terminal phosphorothioate, 2'-ribose protection, terminal protection

Exposed unmodified pyrimidine dinucleotides (ribonuclease hotspots)

Activity — the catalytic site and the gap cannot be protected

Immune suppression

2'-O-methyl near guide position 2, overall 2'-O-methyl coverage

GU-rich motifs, UGUGU repeats, CpG dinucleotides

Loading efficiency — 2'-O-methyl early in the seed lowers loading

Activity

Mechanism-dependent (see 6.3)

Destruction of the structure the mechanism requires

Stability — protection cannot cover the catalytic site

Seed off-target mitigation

Binding-weakening modification in the seed (acyclic analogue > 2'-O-methyl)

Excessive strengthening of binding across the seed

On-target potency — weakening the seed also costs some on-target binding

Self-assembly (conjugate platforms only)

Amphiphilic monomer geometry — balance of hydrophilic and hydrophobic termini

Arrangements that prevent micelle formation

Synthesis complexity

Delivery fitness (carrier-free routes only)

Route-specific productive phosphorothioate window, full 2'-modification, lipophilic or targeting ligand, metabolically stable 5'-phosphate

Exceeding the protein-binding ceiling, excessive bicyclic content

Activity — a high phosphorothioate load costs potency

One design decision within this formulation deserves emphasis: real thresholds must be implemented as constraints, not as weighted terms. The carrier-free phosphorothioate floor, for example, is not 'a value it is good to exceed' but 'a value below which the molecule does not enter the cell'. Encoded as one term of a weighted sum it is outvoted by the potency term and never binds, and the search then optimises a molecule that never arrives. Designs below the floor are therefore judged infeasible rather than merely scoring low, and the final design carries an explicit verdict of pass, blocked, or relaxed.

6.3 The activity axis follows mechanism

The most frequent error in chemistry optimisation is selecting a template by the molecule's topology — single- or double-stranded. The correct criterion is the effector mechanism. Two single strands, one that recruits RNase H1 and one that must not, require opposite chemistry; and two RISC-mechanism molecules, one double-stranded and one single-stranded, have different requirements.

Mechanism

Modalities

What the activity axis rewards

Error from judging by topology alone

RISC cleavage (duplex)

I, part of IV, VI-a

Protection of the cleavage site and 5' loading asymmetry

RISC without cleavage

V transcriptional activation

Loading optimisation; the cleavage rule does not apply

Applying the silencing cleavage-site rule constrains the design for no reason

RNase H1

II, part of IV

Preservation of catalytic recognition in the central gap

Steric block

III, VI-b

Uniform high-affinity occupancy; no gap permitted

Applying a gapmer template causes unintended cleavage

RISC cleavage (single strand)

the single-strand variant of I

5'-phosphate presentation acts as a multiplier; phosphorothioate beyond two per terminus costs potency; over-modification is penalised; there is no strand-selection term

Copying a duplex template produces, without error, a molecule with no activity

6.3.1 The particular case of single-strand RISC

Single-stranded siRNA is single-stranded in topology but RISC in mechanism — a combination the conventional four-effector classification cannot express. Its existence was reported independently in the same issue by Lima and colleagues (Cell 2012) and Yu and colleagues (Cell 2012), showing that Argonaute-dependent activity is achieved without a passenger strand but that a 5'-phosphate is essential for activity in vivo.

A natural 5'-monophosphate, however, is a substrate for cellular phosphatases and is removed before RISC loading when appended to a synthetic molecule. Metabolically stable analogues were therefore sought on the basis of the crystal structure of the 5'-phosphate pocket, and (E)-vinylphosphonate was reported to adopt a conformation similar to the natural phosphate while being metabolically stable (Prakash and colleagues, Nucleic Acids Res 2015). In single-strand RISC design the 5'-phosphate mimic is therefore not a beneficial enhancement but a precondition for activity, and its absence must be applied as a multiplicative penalty no other axis can compensate for.

Phosphorothioate is handled differently too. Under encapsulation, two phosphorothioates per terminus have been reported optimal for a single-strand RISC molecule, with more reducing potency. Carrier-free, the same phosphorothioate is the only means of uptake. One parameter thus operates in opposite directions depending on the delivery route, and this is the clearest case for treating the route as a premise of chemistry.

6.3.2 Duplex chemistry does not transfer to single strands

Because this repeatedly generates cost in practice, it is stated separately. Holen and colleagues (Nucleic Acids Res 2003) reported that positional and accessibility rankings transfer well between single- and double-stranded formats (reported correlation r = 0.967 over seven sites). The same work showed that chemical tolerance does not transfer: a modification equivalent to wild type in a duplex impaired activity in a single strand.

Subsequent work points the same way; three of five active duplexes lost activity on conversion to a single strand (Pendergraff and colleagues, Nucleic Acid Ther 2016). The practical rule is clear: when moving a duplex-validated candidate to a single strand, carry the sequence but not the chemistry. Central unmodified-window conventions, phosphorothioate counts and 5'-terminal handling all encode duplex assumptions.

6.4 Phosphorothioate — one parameter, three consequences

Phosphorothioate fraction is the single most consequential parameter in this layer. The same value acts simultaneously on stability, uptake and toxicity, in different directions.

Fraction range

Stability

Uptake

Toxicity burden

Matching delivery route

0.10 – 0.30

Terminal protection is sufficient

Performed by the particle

Low

Particle encapsulation

0.25 – 0.55

Sufficient

Receptor binding carries part of it

Moderate

Receptor or lipophilic conjugate

0.75 – 1.00

Very high

Phosphorothioate itself is the uptake mechanism

High

Carrier-free (gymnosis)

Above 0.90

Saturated

Little further gain

Protein-binding burden rises steeply

Not recommended — above the ceiling

The ranges rest on different evidence. The lower bound under encapsulation is the minimum needed for terminal protection, and the upper bound comes from the observation that more costs potency — two per terminus on a 20-mer corresponds to 4/19 ≈ 0.21. The carrier-free floor comes from the minimum protein binding required for uptake, based on the observation that roughly 75% or more of linkages must be phosphorothioate. The ceiling at 0.90 is where class effects arising from protein binding — complement activation, platelet effects, hepatotoxicity — rise sharply.

Phosphorothioate carries one further dimension. The phosphorus atom becomes a stereocentre, so each linkage exists as two stereoisomers and ordinary synthesis produces a racemic mixture. An oligonucleotide with n phosphorothioate linkages is a mixture of 2ⁿ diastereomers, which differ in nuclease resistance and target binding. Synthetic approaches that control stereochemistry exist and carry their own patent families. From a design standpoint, whether stereocontrol is assumed changes both the chemistry specification and the quality-control item list, so the assumption is stated explicitly in the deliverables.

6.5 Safety axes created by chemistry

Axis

Mechanism

Design response

Innate immune stimulation

GU-rich and CpG motifs activate endosomal receptors

Motif filtering at enumeration, 2'-O-methyl placement at the chemistry stage

Protein-binding class effects

Phosphorothioate binds complement and coagulation factors

Fraction constrained within the route-specific productive window, penalised above the ceiling

Hepatocyte toxicity

High-affinity chemistry combined with sequence drives unintended cleavage and accumulation

Multi-tier risk classification from sequence and chemistry, emitted in parallel with potency

Seed-mediated off-targets

A seed 7-mer is complementary to hundreds of 3' untranslated regions

Binding-weakening modification in the seed; seed-level burden quantification

Fluorination burden

Cellular burden when total 2'-fluoro content is excessive

Total-content cap and a penalty term

Bicyclic burden

Higher bicyclic content correlates with toxicity risk

Reflected as a penalty in the carrier-free delivery fitness axis

6.6 Chemistry deliverables

Practical warning — a sequence without a fixed chemistry specification is not an orderable specification. The same sequence becomes an entirely different molecule in activity, stability and toxicity depending on chemistry, and in particular copying a duplex template onto a single strand yields, without any error being raised, a molecule with no activity.

7. The Delivery Layer

Nucleic acids are large and strongly anionic and do not cross a lipid bilayer by free diffusion. Every oligonucleotide therapeutic must therefore secure an uptake mechanism, and in principle there are only three: encapsulate in a particle, attach a receptor ligand, or make the backbone chemistry itself bind proteins and drive endocytosis. This chapter covers the biology and physical chemistry of the three routes, how that choice propagates back through the whole design, and what is evaluated on the path from formulation to product.

One easily overlooked fact should be stated first. Getting into the cell and reaching the place where the molecule can act are different problems. Most of what enters by endocytosis remains trapped in endosomes and proceeds to lysosomal degradation; the fraction escaping to the cytosol is very small. Endosomal escape efficiencies reported for lipid nanoparticle work are only a few per cent. The bottleneck in delivery is therefore not getting in but getting out, and much of vehicle design aims at that escape efficiency.

7.1 Decomposing the delivery problem into levels

Level

Question

What determines it

What is observed on failure

Circulation residence

How long does it survive in blood

Particle size and surface hydrophilicity, backbone chemistry, the renal excretion threshold

Disappearance shortly after dosing — no tissue exposure at all

Tissue arrival

Which organ does it reach

Protein corona composition, surface charge, receptor ligand, vascular permeability

Accumulation only in liver and spleen; insufficient concentration in the target tissue

Cell entry

Does it enter the target cell

Receptor expression density and recycling rate, backbone protein binding

Arrival in the tissue but retention in the extracellular space

Endosomal escape

Does it reach the cytosol

The behaviour of the ionisable lipid at endosomal pH, its membrane-destabilising capacity

Intracellular fluorescence is visible but there is no pharmacology — the most common failure point

Arrival at the site of action

Must it reach the nucleus

Depends on modality (transcriptional activation requires the nucleus)

Present in the cytosol but never meeting the target

The fourth level is the bottleneck for a quantitative reason. That only a few per cent of endocytosed molecules escape to the cytosol means most of the dose is lost before reaching its target. Doubling that fraction halves the dose required for the same effect, and halving the dose halves class toxicity and manufacturing cost together. This is why ionisable lipid design attracts the largest investment in the field.

7.2 Route 1 — receptor-targeted conjugation

7.2.1 Why this works exceptionally well in hepatocytes

Conjugating a triantennary N-acetylgalactosamine ligand to an oligonucleotide terminus causes the asialoglycoprotein receptor on the hepatocyte surface to recognise and internalise it. This works exceptionally well because three biological conditions are met at once, and checking those three one by one is the checklist for extending to any other tissue.

Condition

Value in hepatocytes

Why it is needed

What to check in another tissue

Surface receptor density

Very high — on the order of hundreds of thousands per cell

A high probability of encounter means capture even at low concentration

Receptor expression in the target cell — checkable in advance from tissue expression data

Recycling rate

Releases its ligand and returns to the surface within minutes

Repeated use means a small receptor pool absorbs a large amount

Does it recycle after internalisation, or route to degradation

Natural distribution of the organ

Subcutaneous dosing concentrates naturally in the liver

The ligand's job is not to redirect but to capture what is already passing

Is the tissue naturally exposed on that route of administration

The third condition is the one most often overlooked. Hepatic conjugates succeed not because the ligand drags the molecule to the liver but because it captures a molecule that would pass through the liver anyway. In a tissue with low natural exposure, the same receptor density gives a far lower capture probability.

7.2.2 The geometry of multivalent binding

A single galactosamine binds the receptor weakly. Three arranged on a branched scaffold with appropriate spacing engage three binding sites of the receptor trimer simultaneously and raise affinity by orders of magnitude. What matters is not the number but the spacing and flexibility: if the branch arm lengths do not match the spacing of the receptor's binding sites, multivalent binding does not occur and the result is merely the sum of monovalent interactions. This multivalent geometry is the core of the conjugation technology and forms one axis of the patent landscape.

7.2.3 What the linker chemistry determines

Linker type

Bond

In vivo behaviour

When it fits

Oxygen (O)

Phosphodiester-like linkage

Cleaved relatively quickly by plasma nucleases and esterases

When the ligand should be shed after hepatic arrival

Sulfur (S)

Phosphorothioate-like linkage

High cleavage resistance, so the ligand persists

When premature loss in circulation must be prevented

Carbon (C)

Non-cleavable covalent bond

Not cleaved in vivo

When retaining the ligand is harmless or required for activity

Linker choice looks minor and is not. Premature cleavage loses the ligand before the target tissue is reached and nullifies the benefit of conjugation; conversely a non-cleavable linker can leave the ligand acting as steric bulk inside the cell and lower activity. The deliverables present attachment point and linker type bound into one specification with sequence and chemistry.

7.2.4 The ligand catalogue — extending beyond the liver

Efforts to extend the same principle to other tissues are active. The design layer holds a ligand catalogue indexed by receptor, each entry carrying the target receptor, the ligand format (sugar, peptide, antibody fragment, small molecule) and the reference binding constant as reported in the literature. Twenty-nine entries are currently registered; binding constants without literature support are not generated and are marked unconfirmed.

Target family

Receptor

Ligand format

Tissue or cell targeted

Hepatocyte

Asialoglycoprotein receptor

Triantennary galactosamine

Hepatocytes — the established route

Tumour vasculature

Integrin αvβ3

Cyclic RGD peptide

Tumour neovascular endothelium

Tumour metabolism

Folate receptor alpha

Folate

Folate-receptor-overexpressing tumours

Immune cells

Mannose receptor

Mannose

Macrophages and dendritic cells

Immune cells

Dendritic cell surface lectin, endocytic receptor, Fc receptor

Antibody fragments

Cross-presenting dendritic cells, monocyte lineage

Prostate

Prostate-specific membrane antigen

Small-molecule ligand

Prostate cancer

Breast and gastric

HER2

Antibody fragment

HER2-positive tumours

Epithelial tumours

TROP2 · EGFR

Antibody fragment · peptide

Epithelial-origin tumours

Neuroblastoma

GD2 ganglioside

Antibody

Neuroectodermal tumours

Gastric and pancreatic

Claudin 18.2

Antibody

Gastric and pancreatic tumours

Blood–brain barrier

Transferrin receptor

Antibody fragment and three peptides

Receptor-mediated transcytosis

Blood–brain barrier

Insulin-like growth factor 1 receptor

Single-domain antibody

Receptor-mediated transcytosis

Blood–brain barrier

LDL receptor-related protein

Angiopep-class peptide

Receptor-mediated transcytosis

Blood–brain barrier

Glucose transporter · glutathione transporter

Sugar · glutathione

Transporter-mediated crossing

Brain homing

Cyclic brain-homing peptide family

Cyclic peptides

Brain tissue accumulation

Nervous system

Nicotinic acetylcholine receptor

Cyclic peptides

Neurons

Metabolic

Leptin receptor

Peptides

Hypothalamus and related

The blood–brain barrier shuttle family is a particularly active area. Receptor-mediated transcytosis binds a specific receptor on brain capillary endothelium and traverses the cell, attempting to reach the central nervous system while avoiding the invasiveness of intrathecal dosing. Crossing efficiency is low and the receptors are expressed in other tissues too, so selectivity remains the challenge.

Every binding constant in the ligand catalogue is a literature reference value. The affinity of an arbitrary ligand–receptor pair is not computed, and uncertain entries are left empty and marked unconfirmed.

7.3 Route 2 — lipid nanoparticle encapsulation

7.3.1 Four components and their roles

Component

Role

Design variables

Failure mode

Ionisable lipid

Neutral at physiological pH and cationic at endosomal pH — performs both encapsulation and endosomal membrane disruption

Apparent pKa, tail architecture and branching, position of degradable linkages

Too low a pKa and encapsulation fails; too high and circulating toxicity and non-specific binding follow

Helper phospholipid

Bilayer structure formation, fusion support

Saturation, phase transition temperature, whether the geometry is conical

Insufficient fusogenicity lowers endosomal escape

Cholesterol

Membrane fluidity and stability, particle rigidity

Content, use of analogues

Too little and the particle leaks; too much and fusion is impeded

PEG-lipid

Prevents aggregation, extends circulation time

Mole fraction, chain length, lipid anchor length (which sets the shedding rate)

Too much impedes uptake and escape; too little causes aggregation and rapid clearance

Composition is specified as mole fractions, and a widely referenced baseline places ionisable lipid, helper, cholesterol and PEG-lipid at approximately 50 : 10 : 38.5 : 1.5 mol%. That ratio is itself the subject of patent claims, so the design layer reports the proximity of a searched composition to known composition claims — so that the centre of a claim scope is not stepped on for no reason.

7.3.2 The apparent pKa of the ionisable lipid

This is the single most consequential parameter on this route. Two requirements pull in opposite directions: at blood pH (7.4) the lipid must be neutral so that toxicity and non-specific binding stay low, and in the endosome (around pH 5.5) it must be cationic so that it can interact with and destabilise the anionic endosomal membrane.

The band satisfying both is narrow. Computing the protonated fraction at endosomal pH through the Henderson–Hasselbalch relation places the optimum near an apparent pKa of about 6.2 to 6.5. Lower, and the lipid does not become sufficiently cationic even in the endosome, so escape does not occur; higher, and it is already cationic in blood and binds plasma proteins and erythrocytes non-specifically.

One distinction matters. The apparent pKa is not the pKa of an isolated lipid molecule but the value at the surface of the assembled particle. When lipids pack densely at the surface, neighbouring positive charges repel one another and suppress protonation, so the same lipid has a different apparent pKa depending on mole fraction and surface density. The design layer computes an apparent pKa that accounts for this surface charge effect; a calculation using the single-molecule pKa directly misses the effect of composition changes.

7.3.3 The cargo determines the particle — N/P ratio and target size

With the same lipid composition, what is loaded changes the particle. A short duplex oligonucleotide and a transcript of several thousand bases differ in charge density, rigidity and volume, so the amount of cationic lipid required (the N/P ratio — ionisable lipid amines to nucleic acid phosphates) and the size and morphology of the resulting particle differ.

Cargo

Typical length

N/P ratio

Target hydrodynamic diameter

Expected morphology

siRNA · ncRNA-directed · activating duplexes

21 nt

1

about 25 nm

core–shell

ASO · SSO single strands

20 nt

1

about 25 nm

core–shell

Generic reference

about 60 nm

core–shell

Messenger RNA

about 1,500 nt

6

about 80 nm

bleb

Circular RNA

about 1,200 nt

6

about 60 nm

bleb

Self-amplifying RNA

about 9,500 nt

8

about 110 nm

bleb

Circular self-amplifying RNA

about 11,500 nt

9

about 125 nm

bleb

The morphological distinction matters in practice. Short oligonucleotides form a dense core with the ionisable lipid, wrapped by phospholipid and polyethylene glycol — a core–shell structure — whereas long transcripts form a bleb structure containing internal aqueous compartments. The two differ in encapsulation efficiency, release behaviour and stability, and cannot be managed under one quality specification. The design layer assigns an expected morphology per cargo type and confirms at the structural evaluation stage that it is actually obtained.

7.3.4 PEG-lipid — resolving a conflict in time rather than space

The PEG-lipid mole fraction creates a clear conflict. More gives a stable particle and long circulation but impedes cellular uptake and endosomal escape; less improves uptake but the particle aggregates and is cleared rapidly. The value near 1.5 mol% commonly used is the compromise.

A widely used design resolves the conflict in time rather than in space. Shortening the lipid anchor chain makes the polyethylene glycol shed rapidly from the lipid membrane in vivo, so stability and circulation are secured immediately after dosing while uptake occurs from a shed particle after tissue arrival. Anchor length is in effect the control variable for when shedding happens.

Polyethylene glycol carries a separate immunological consideration. Repeat dosing can raise anti-PEG antibodies, producing accelerated blood clearance from the second dose onward, and this is an independent evaluation item in the design layer.

7.3.5 The protein corona — the surface the cell actually sees

A systemically administered particle is coated with plasma proteins almost immediately, and that adsorbed layer becomes the surface the cell recognises. Biodistribution is governed by the composition of the adsorbed proteins rather than by the designed surface composition. When particular apolipoproteins adsorb, uptake proceeds through the corresponding hepatocyte receptor, and this is the principal reason lipid nanoparticles concentrate in the liver without any targeting ligand.

Sending them elsewhere is therefore a problem of changing the corona before it is a problem of attaching a ligand. Attaching a ligand achieves nothing if the corona buries it. The design layer treats corona formation and multivalent binding (avidity) as separate evaluation stages, assessing whether a surface ligand remains accessible beneath the corona.

7.3.6 Selective organ targeting

Selective organ targeting strategies follow from the corona observation. Adding a fifth lipid to the base four-component formulation adjusts the particle's apparent charge, which changes corona composition and shifts biodistribution.

Added lipid type

Representative lipid

Organ it shifts toward

Mechanistic reading

Permanently cationic

Quaternary ammonium lipid family

Lung

Electrostatic capture in pulmonary capillary endothelium and a changed corona

Anionic

Phosphate-class anionic lipids

Spleen

A shift toward opsonin composition recognised by splenic macrophage lineages

No fifth lipid

The base four components

Liver

Apolipoprotein-mediated hepatocyte uptake (the default)

Permanently cationic at high ratio

Specialised cationic combinations

Central nervous system access

Used with the intrathecal route; systemic dosing does not cross the blood–brain barrier

The design layer treats this redirection with a trained predictive model rather than a rule table. A model validated on a 1,593-row training set including 49 curated selective-organ-targeting formulations predicts organ distribution probability from composition, with an in-house validated balanced accuracy of about 0.717. Predictions are emitted as a probability distribution over liver, spleen, lung, kidney and lymph node, and are used for relative comparison among candidate compositions.

7.3.7 Endosomal escape — the real bottleneck

As the endosome acidifies, the ionisable lipid protonates and the now-cationic lipid forms ion pairs with anionic endosomal membrane phospholipids. Those ion pairs adopt a conical geometry that creates pressure to convert the bilayer (lamellar) organisation into an inverted hexagonal phase, and that phase transition destabilises the membrane so the cargo escapes to the cytosol. The conical geometry of the helper phospholipid assists the same transition.

A proton sponge effect is superimposed. As amine groups absorb incoming protons, chloride ions and water enter to compensate, osmotic pressure rises and the endosome can swell and rupture. Which of the two mechanisms dominates depends on the lipid, and the design layer computes a free energy profile under endosomal conditions as a separate evaluation stage.

7.3.8 What encapsulation changes in the chemistry

Because the particle performs both protection and uptake, the chemical burden on the oligonucleotide falls substantially. Phosphorothioate beyond the termini costs potency for no gain, and full 2'-modification is not mandatory. Conversely, properties affecting behaviour inside the particle — charge density, duplex stability, length — become new considerations. Carrying a chemistry optimised for the carrier-free route unchanged onto an encapsulated route loses potency and takes on class toxicity for nothing.

7.4 Route 3 — carrier-free uptake

A phosphorothioate backbone binds plasma and cell-surface proteins, and that binding pushes the molecule into endocytic routes. The observation that silencing occurs simply on exposing cultured cells to oligonucleotides without transfection reagent is the origin of this route, and it was reported together with the finding that a fairly high phosphorothioate density is required.

The practical value of this route is a shortened development path: preclinical entry without a separate delivery-vehicle programme, and a sharply reduced set of formulation-related chemistry-manufacturing-and-control items. The cost is a large required dose, correspondingly more scope for class toxicity, and no active means of steering tissue distribution.

Design emits carrier-free suitability as a diagnosis rather than a score. A rejected design that leaves only a low number does not say what to change, so the failing gate is returned by name.

7.5 The self-assembling conjugate platform

Attaching a hydrophilic polyethylene glycol at the sense strand's 5' terminus and a hydrophobic lipid anchor at its 3' terminus makes the molecule itself an amphiphile. The monomer self-assembles into a micelle on the order of 100 nm whose shell protects the duplex. Rather than building a separate lipid particle, the molecule becomes the particle.

The practical consequence is a lighter chemical burden: because the shell performs the protection, the internal 2'-modification and phosphorothioate load is lower than for a naked conjugate. When this platform is selected the design layer adds a fifth objective, self-assembly, and switches to conjugate-aware templates. Conjugation is encoded per terminus — guide 5'/3' and passenger 5'/3', each none, lipid, polyethylene glycol or ligand.

Single-strand modalities receive a lipid (3') and polyethylene glycol (5') analogue on the drug strand itself, scored against a lower micelle ceiling. A duplex can use the passenger as a sacrificial strand, whereas a single strand must carry the conjugation on the active strand and therefore risks more activity loss.

Applying the self-assembling platform can put the overhang convention and the chemistry convention in tension. Conjugation is at the 3' terminus and this platform's standard is a native overhang, so the chemical treatment of the overhang bases must be stated in the specification. A 21-mer fully complementary guide with a 19-mer sense strand is the recommended configuration.

7.6 Alternative vehicles

Vehicle class

Principle

Strengths

Current constraints

Registered variants

Extracellular vesicles (exosomes)

Payload loaded into cell-derived vesicles; surface proteins confer targeting

Low immunogenicity; tissue tropism engineerable through surface proteins; exploits natural intercellular transfer

Manufacturing scale and homogeneity, loading efficiency, absence of characterisation standards

6

Polymeric carriers

Polyesters, polyamines and dendrimers condense nucleic acid electrostatically

Wide compositional freedom, straightforward sustained-release design, degradation rate tunable chemically

Cytotoxicity of cationic polymers, management of degradation products, polydispersity

8

Virus-like particles

Viral capsid proteins self-assemble around the nucleic acid

Homogeneous particle size, natural cell-entry mechanism, high loading density

Immunogenicity and repeat dosing, payload capacity, manufacturing complexity

6

Aptamer conjugation

A folded nucleic acid binds a surface receptor and internalises the payload

The molecule is its own vehicle; non-immunogenic; synthetically homogeneous

Limited to targets with a validated aptamer; the endosomal escape bottleneck remains

see Modality X

Extracellular vesicles and virus-like particles form by routes different from lipid particles and therefore need their own structural evaluation stage. At present they are registered and assessed at the level of composition and surface properties, with structural simulation a staged extension. The deliverables state that status explicitly — the principle of marking an unevaluated item 'not evaluated' rather than scoring it zero applies here too.

7.7 The formulation variant catalogue — 59 variants

The design layer registers combinations of cargo type and vehicle class as named product variants and evaluates them in parallel during a campaign. Each variant carries a predefined lipid composition, target size, expected morphology, default route of administration and available ligand set, so that comparison among candidates is made on identical terms.

Vehicle / cargo class

Variants

Character

Self-assembling conjugate (specificity and immobilisation families)

9

Lipid and polyethylene glycol amphiphilic micelle

Single-strand oligonucleotide encapsulation

2

Gapmer and steric-block cargo, core–shell

Messenger RNA encapsulation

5

Bleb route, with cap, polyA and modified nucleoside specifications

Self-amplifying RNA encapsulation

2

Large transcript, high N/P, large bleb

Circular RNA encapsulation

3

Internal ribosome entry site driven, no ends

Generic reference formulation

2

Cargo-agnostic control formulation

Selective organ targeting

8

Lung 3 · spleen 3 · central nervous system 2

Multiplexed cargo

2

Multiple cargoes co-encapsulated

Noncoding-target oligonucleotide

2

ncRNA-directed duplexes

Transcriptional activation oligonucleotide

2

Promoter-targeted duplexes

Steric-block oligonucleotide

2

Uniformly modified single strands

Extracellular vesicles

6

Including surface targeting-protein engineered variants

Polymers

8

Polyester · polyamine · dendrimer · polysaccharide · block copolymer

Virus-like particles

6

Multiple capsid families

Single-strand RISC structural path

0 by default (2 when enabled)

3'-only amphiphile, no 5' additions

The design principle in the last row matters in practice. Registering a new modality must not, as a side effect, push that modality into a running campaign. New families requiring structural search are included only when explicitly enabled; otherwise they receive chemistry optimisation and skip the structural and dynamics stages. Registering and wiring are different acts.

7.8 Route of administration and pharmacokinetics

The route changes absorption rate and tissue distribution together. The design layer holds per-route pharmacokinetic parameters and emits an exposure profile and organ distribution bias per route for a candidate formulation. The values below are literature-based priors used for relative comparison among candidates — they are not absolute predictions.

Route

Time to peak

Half-life

Normalised peak

Organ bias (liver / lung / spleen / CNS)

Intravenous

0.25 h

4.5 h

1.00

0.50 / 0.10 / 0.30 / 0.02

Intramuscular

6 h

24 h

0.40

0.20 / 0.05 / 0.10 / 0.01

Subcutaneous

12 h

36 h

0.30

0.15 / 0.05 / 0.08 / 0.01

Inhalation

0.5 h

6 h

0.80

0.05 / 0.75 / 0.05 / 0.01

Intrathecal

2 h

30 h

0.70

0.02 / 0.02 / 0.02 / 0.85

Intraocular

4 h

48 h

0.50

0.02 / 0.02 / 0.02 / 0.01

Intratumoral

1 h

12 h

0.90

0.10 / 0.05 / 0.05 / 0.01

What to read from the table is the contrast between routes rather than the absolute values. Intravenous dosing gives the highest peak but a short half-life and strong hepatic and splenic bias; subcutaneous gives a lower peak with gentler, longer exposure. Intrathecal dosing has an overwhelming central nervous system bias of 0.85 with almost no systemic exposure, so class toxicity constraints relax substantially; inhalation achieves local concentration with a lung bias of 0.75 but a short half-life, making residence time the deciding factor.

The route feeds back into the whole formulation. On the inhaled route, shear during aerosolisation can destroy particles, so formulation rigidity matters; on the intrathecal route, cerebrospinal fluid distribution and osmolality and pH compatibility become new constraints; and intratumorally, interstitial pressure and diffusion distance govern efficacy.

7.8.1 Dose translation and in vitro–in vivo correlation

Translating a dose obtained in a preclinical species to humans, and connecting in vitro activity to in vivo activity, exist as separate evaluation stages. What makes that translation awkward for oligonucleotide therapeutics is the large divergence between plasma and tissue exposure: plasma clears within hours while the tissue depot persists for weeks to months, so a plasma-based translation grossly underestimates the real duration of effect. The design layer handles this by treating the tissue depot as a separate compartment, and the output is for candidate comparison rather than a proposed clinical regimen.

7.9 The stage map of delivery design

Delivery design runs as a sequence of stages from chemistry optimisation to formulation and process assessment. The overall skeleton is below, marked for what falls within standard design scope and what is arranged separately.

Block

What it decides

In standard scope

Design front end — cargo verification, chemical-modification optimisation, conjugate structure search, multi-objective Pareto search, joint chemistry × vehicle search

Sequence and chemistry, and the joint optimum that holds chemistry and vehicle together

Yes

Lipid and formulation — lipid screening, composition search, quality-by-design space, formulation fixing, vehicle structure assessment, sterol scoring

Ionisable lipid, helper, cholesterol and PEG-lipid ratios and the permissible design space

Yes

Physicochemistry and safety — physicochemical and toxicity, stability, innate immunity, anti-PEG, lipid degradants, lipid phase behaviour, metal-catalysed oxidation

Apparent pKa, particle properties, immunogenicity, degradation routes

Yes

Structure and dynamics — assembly structure generation, backmapping, molecular dynamics production, convergence, trajectory analysis, endosomal free energy profile

Particle morphology (core–shell or bleb) and the energetic barrier to endosomal escape

Requires large simulation assets — arranged separately

Biology and pharmacokinetics — protein corona and avidity, physiologically based pharmacokinetics, dose translation, route absorption, lipid clearance, in vitro–in vivo correlation, PEG design, serum stability

Biodistribution, dose, route-dependent absorption and correlation

Yes

Intellectual property, quality and ranking — patent filter, good-manufacturing assessment with roughly 40 quality half-stages, ranking, design-result emission, platemap

Proximity to composition claims, formulation, process and container risk, final candidate ranking

Yes (at the level of risk exposure)

The whole span consists of 82 independent evaluation stages, each of which can be disabled or re-run individually. The purpose of that granularity is twofold: running expensive stages only after the candidate pool has narrowed, and making it possible to trace by name at which stage a candidate was rejected.

7.10 From formulation to product — quality items surfaced at design time

Some quality items are hard to reverse once the formulation is locked. The design layer treats them as independent evaluation stages so that they surface at candidate-selection time. This is not process design but risk exposure; actual process development and validation belong to later stages.

Category

Items

Why it must be seen at design time

Physical properties and stability

Apparent pKa, particle size and distribution, encapsulation efficiency, colloidal stability (electrostatic versus van der Waals balance), lipid phase behaviour, lipid degradants, metal-catalysed oxidation, serum stability, RNA integrity

Composition can still be changed here; after the process is fixed, changing it is expensive

Immune and toxicity

Innate immune response, anti-PEG response, complement activation tendency

Governs repeat-dosing design and whether premedication is needed

Freezing and drying

Lyophilisation conditions, reconstitution, repeated freeze–thaw, cryogenic pH shift, crystallisation, long-term storage

Cold-chain requirements determine commercial viability

Process

Sterile filtration, tangential-flow filtration, microfluidic scale-up, aseptic fill, process capability

Whether particle size crosses the filtration limit, and whether size holds on scale-up

Container and transport

Container-closure suitability, container-closure integrity, headspace oxygen, agitation and shear, transport vibration

Logistics conditions set the formulation rigidity requirement

Clinical use

Infusion compatibility, deliverable volume, in-use stability, temperature excursion, photostability

Handling conditions at the point of care constrain the formulation

Regulatory and documentation

Potency assay, comparability, impurities including nitrosamines, documentation package

Components of the filing package are secured at design time

The list is long because it is genuinely long. A substantial share of failures in oligonucleotide development arises not from sequence or chemistry but from items on this list, and many of them require reversing the composition if discovered after the formulation is locked. The value of surfacing them at design time lies not in the accuracy of the prediction but in moving the moment of discovery earlier.

7.11 Delivery deliverables

7.12 Limits of this layer

8. Modality IV — ncRNA-Directed Silencing (lncRNA · circRNA · snoRNA siRNA/ASO)

8.1 Form and definition

The common name is ncRNA-directed siRNA or ncRNA-directed ASO. Apart from the fact that the target makes no protein, the molecule itself is the same as in Modality I or II. Either the double-stranded RISC mechanism or the single-stranded gapmer mechanism may be used, and the choice follows from the target's class and subcellular localisation. This modality is treated separately not because the molecule differs but because the target's properties differ fundamentally.

There are four differences. Noncoding RNA is markedly less conserved across species than protein-coding genes. Much of it is nuclear-retained, out of reach of cytoplasmic mechanisms. Some targets, such as circular RNAs, share most of their sequence with a linear transcript, so specificity must be achieved differently. And function frequently resides in structure or binding partners rather than sequence, so 'where to cut for the function to disappear' is not self-evident.

8.2 Biology by target class

8.2.1 Long noncoding RNA

Noncoding transcripts longer than 200 nt, functioning as scaffolds for chromatin-remodelling complexes, decoys for transcription factors, or structural elements of nuclear bodies. Many are nuclear-retained, so the RNase H1 mechanism with its nuclear activity is often favoured over the cytoplasmic Argonaute mechanism. An important design fact is that these transcripts have no protein-coding isoforms, so the metric 'coding transcripts cut / total' is structurally 0/0. That is a property of the target rather than an error, and the all-transcript metric is the meaningful one.

8.2.2 Circular RNA

Transcripts whose 5' and 3' ends are covalently joined by back-splicing. Having no ends, they resist exonucleases and are long-lived, and functions as miRNA sponges and as translation templates have been reported. The difficulty as a therapeutic target is specificity: a circular RNA's sequence overlaps almost entirely with the exons of the linear transcript that produced it, so targeting an arbitrary site silences the linear parent gene as well.

The only sequence unique to the circle is the back-splice junction, where the 3' end of a downstream exon abuts the 5' start of an upstream exon. That boundary does not exist in the linear transcript, so only a guide spanning the junction is circle-specific. Design builds a junction-aware alignment index to secure that specificity, and computes accessibility on the circular topology, because a linear-assumption calculation mispredicts structure around the junction.

One interpretive note: a guide spanning the junction does not match the linear transcript annotation, so 'cut on representative transcript' is reported as negative. That is evidence of circle specificity, not a defect.

8.2.3 Small nucleolar and other small noncoding RNAs

Short structured RNAs that guide chemical modification of ribosomal RNA. They have strong secondary structure and exist in complex with proteins, so accessible sites are limited and single-stranded mechanisms are frequently favoured over double-stranded ones. Design classifies this class automatically and routes it to the corresponding mechanism.

8.2.4 Translation-enhancing elements

Antisense transcripts carrying particular repeat sequences have been reported to increase translation of a target mRNA, and engineered constructs based on this are under study. This class is both a silencing target and a design object, so the design layer applies different handling according to the class call.

8.3 Class calling and verification of the premise

The first step in this modality is confirming that the target really is noncoding. Annotation database classifications are not always current or correct, and some transcripts carry short open reading frames and do produce peptides. Design places a coding-potential tool as a gate on that premise and states the call in the report, because proceeding silently on a false premise is the most expensive kind of failure.

Once the class is fixed, the mechanism follows: nuclear-retained targets route to RNase H1, cytoplasmic targets to RISC, and strongly structured small RNAs to a single-stranded mechanism. That this branching is automatic matters in practice — selecting a different tool by hand for each class invites a missed branch, and a missed branch produces output on which no filter was applied.

8.4 Single-species design as a deliberate constraint

This modality designs against one species at a time. That is a deliberate constraint rather than a missing feature. Cross-species conservation of noncoding RNA is far weaker than for protein-coding genes, so demanding cross-species reactivity automatically causes one of two problems: the design space narrows without justification, or sequences that are not functional counterparts are matched as orthologs. The second is the more dangerous, because preclinical species selection based on a false ortholog match produces meaningless toxicology.

Where preclinical species coverage is needed, designing per species independently and comparing the resulting candidates at the sequence level rests on firmer evidence.

8.5 Interfering factors — RNA-binding proteins and chemical modification

Noncoding RNAs frequently function in complex with proteins, so targeting a protein-occupied site impedes binding. Design uses binding peaks derived from crosslinking and immunoprecipitation data as masks to avoid such sites. Likewise, regions around particular chemical modifications may differ in binding and structure, so a modification atlas is used as a mask as well.

These masks operate as weighted penalties rather than absolute exclusions. Binding profiles depend on cell type and condition, so a peak observed in one dataset cannot be assumed present in a customer's experimental system. Applying a penalty and recording the basis in the report is more accurate than absolute exclusion.

8.6 Therapeutic significance

The strategic value of targeting noncoding RNA is that it opens a route to targets considered undruggable. Where a regulatory axis cannot be addressed at the protein level by a small molecule or an antibody but runs through a noncoding RNA, lowering that RNA is the only point of intervention. Chromatin-level regulation, nuclear body formation and transcription-factor decoying in particular are difficult to reach by protein targeting and direct to reach by RNA targeting.

Clinical maturity in this area is lower than in the three preceding modalities. No noncoding-RNA-targeted therapeutic has been approved, and most programmes are preclinical or in early clinical development. That means development risk is high, and equally that much of the target space remains unexplored. Strategically the modality is best understood as low-competition and high-validation-burden.

8.7 Chemistry and delivery requirements

Chemistry requirements are inherited from the chosen mechanism — the Modality I rules for RISC and the Modality II rules for RNase H1. Delivery requirements follow similarly, with subcellular localisation as an additional variable. Nuclear-retained targets are not addressed by cytoplasmic delivery alone, so nuclear entry must be assumed; fortunately single-stranded oligonucleotides are known to distribute well to the nucleus, which makes that combination practically favourable.

8.8 Design deliverables

8.9 Advanced capabilities — class routing, circle specificity, interference masks

8.9.1 Class calling and branch routing

The first step fixes the target's class, and the call routes it to an entirely different design branch. The call combines annotation priority with a computational gate: the reference annotation's accession prefix is read first, and where that is ambiguous a coding-potential tool confirms it.

Class called

Branch routed to

Special handling in that branch

Long noncoding RNA

RISC or RNase H1, depending on nuclear retention

Zero protein-coding isoforms, so the coverage metric is read differently

Circular RNA

Back-splice-junction-aware path

Junction-specific alignment index, circular-topology accessibility, separate reporting of the effect on the linear parent gene

Small nucleolar and other small structured RNAs

Handed to a single-stranded mechanism

Accessibility assessed first, given strong secondary structure and protein complexes

Translation-enhancing elements

Currently unsupported, returned as an explicit error

Not quietly routed to an adjacent branch — silent wrong answers are prevented

Actually protein-coding

Terminated as an explicit failure with a pointer to the coding-target path

That the 'noncoding' premise was wrong is signalled as a failure, not delivered as a result

The last two rows express this layer's design philosophy. Quietly routing an unsupported class to a neighbouring branch produces output whose meaning nobody can establish. Explicit failure is far cheaper than a wrong success.

8.9.2 Circle specificity — junction design and a validation plan

A circular RNA's sequence overlaps almost entirely with the parent gene's exons, so the only sequence unique to the circle is the back-splice junction. Design accepts only guides spanning that junction as circle-specific and builds a dedicated junction-aware alignment index to verify the specificity.

It does not stop there: a wet-lab validation plan is emitted as well. Demonstrating circle-specific knockdown experimentally requires quantifying circular and linear species separately and reading their ratio, which needs particular primer designs — a divergent primer pair spanning the junction amplifies only the circle, while a convergent pair amplifies only the linear form. The deliverables present both pairs per candidate.

Candidate selection is also constrained to secure several guides sufficiently separated around the junction, because a circle-specific phenotype is only credible when three or more independent guides give the same direction of result.

8.9.3 Alternative mechanism advisory — when to recommend a different tool

Sometimes the circular RNA is short, or the sequence around the junction is unfavourable, and too few viable junction-spanning candidates can be found. Rather than padding the list, the design layer advises an alternative mechanism: a programmable RNA-targeting nuclease approach.

The basis for the advisory is a difference in complementarity thresholds. RISC cleavage requires roughly 17 to 19 nucleotides of contiguous complementarity, whereas that nuclease family has been reported to activate at longer complementarity, on the order of 23 nucleotides. Because of that difference, the window in which a guide cuts the circle without touching the linear parent opens differently for the two mechanisms. Design computes the contiguous complementarity length against the linear parent for each candidate and adjudicates which mechanism affords specificity.

8.9.4 Interference masks — modification and protein occupancy

Noncoding RNAs are complexed with proteins and chemically modified to a greater degree than mRNA, so both factors can badly distort accessibility prediction. Two masks are applied and are active by default.

Mask

Data loaded

What it avoids

How it is applied

RNA modification mask

172,706 high-confidence modification sites, integrating several modification types (m6A, m5C, m7G, m1A, A-to-I, pseudouridine and others)

Sites where modification alters base pairing and protein binding

A weighted penalty, not an absolute exclusion

Protein occupancy mask

Aggregated crosslinking-immunoprecipitation peaks across many cell lines

Sites physically occupied by protein and therefore inaccessible

A weighted penalty, because occupancy is cell-type dependent

Editing site handling

An A-to-I editing site database

An edited transcript has a different sequence, so the guide no longer matches

A surrogate edited sequence (A→G substituted) is generated and complementarity is assessed against both versions

The third row is specific to this class. Targeting a frequently edited site means a guide that is fully complementary against the reference carries a mismatch in the real cell. Design generates the edited surrogate and prefers candidates valid against both versions.

8.9.5 Accessibility on a circular topology

Accessibility computed on a linear assumption mispredicts structure around the back-splice junction, because a linear calculation treats the two ends as free termini while a circle has none, and sequences that were originally far apart become neighbours across the junction. Design computes topology-aware accessibility for circular targets; without that handling, candidates near the junction are systematically misassessed.

8.9.6 Single-species design and per-species references

This layer holds independent references for six species — human, mouse, rat, macaque, dog and pig — and designs against one species per run. Automatic cross-species extension is not offered for the reason given in 8.4; where preclinical species coverage is needed, designing per species and comparing at the sequence level rests on firmer evidence.

8.10 What this modality cannot solve

9. Modality V — saRNA / RNAa (Promoter-Targeted Transcriptional Activation)

9.1 Form and definition

The common name is saRNA (small activating RNA), and the operating principle is called RNAa (RNA activation). The form is the same double-stranded small RNA as Modality I, but the target is a gene's promoter region rather than its mRNA, and the outcome is activation rather than repression. The same material in the same form acts in opposite directions depending on where it is pointed. That fact produces this modality's most important design implication: which coordinate is targeted dominates the outcome more than the sequence design itself.

9.2 Molecular mechanism

9.2.1 An Argonaute-mediated event in the nucleus

That a small double-stranded RNA targeting a promoter can increase gene expression has been reported reproducibly across multiple genes. The leading mechanistic account is that an Argonaute protein moves into the nucleus, interacts with low-abundance noncoding transcripts produced at the promoter, and induces activation-associated chromatin marks at that locus. The target is therefore not genomic DNA itself but the RNA transcribed there, using the same recognition principle as silencing with the opposite outcome.

An important difference is that cleavage is not required. In silencing the catalytic activity of Argonaute is central, whereas in activation the presence of the complex at the locus is what produces the effect. Cleavage-site protection rules therefore do not apply in chemistry optimisation, and applying a silencing chemistry template unchanged imposes an unnecessary constraint. Which member of the Argonaute family is routed through may also differ from silencing, so the design includes a loading-route audit.

9.2.2 Chromatin state is a precondition

For activation to occur, the locus must be in a state capable of being activated. A fully condensed, inaccessible promoter gives the complex nowhere to settle, and an already maximally active promoter has no headroom. The modality therefore works best at loci in an intermediate state.

Design accordingly assesses the chromatin context of the target locus, consulting active-promoter marks, chromatin accessibility and active-enhancer marks to present evidence on whether the locus presents favourable conditions for an activation attempt. Public chromatin data, however, are measured in particular cell lines and tissues and may not represent a customer's experimental system. That limitation is stated in the report, and the output is best read as an evidence summary rather than a prediction of magnitude.

9.2.3 Natural antisense transcripts as a second axis

At many loci a transcript is produced in the opposite orientation alongside the sense transcript, and in some cases the antisense transcript represses the sense one. Where such a repressive axis exists, lowering the antisense transcript is itself a way of raising the target gene — achieving activation by means of silencing. Design maps antisense transcripts around the target locus and presents that possibility.

The practical advantage of this route is predictability. Silencing mechanisms have higher prediction accuracy and far deeper clinical precedent than activation, so achieving the same therapeutic objective by silencing lowers development risk. One of the first things the design layer checks in response to an activation request is therefore whether a repressive axis exists that can be silenced instead.

9.3 Therapeutic significance — why activation is needed

There are disease classes that silencing cannot address in principle: loss or silencing of a tumour suppressor, haploinsufficiency, enzyme deficiency in a metabolic pathway, and shortage of a protective factor. In these cases the therapeutic direction is upward, and the existing options were protein replacement, gene therapy, or a small molecule that raises expression. Each has constraints: replacement requires repeat dosing and carries immunogenicity risk, gene therapy carries vector-associated risk and difficulty in controlling expression, and small molecules struggle with target selectivity.

Promoter-targeted activation aims at that gap. Because the endogenous gene is raised within its own regulatory context, the increase is expected to stay near physiological range; sequence-based specificity provides target selectivity; and being a synthetic oligonucleotide, it uses the existing manufacturing and administration infrastructure of the field.

Clinical maturity is lower than for silencing. Clinical trials have been conducted in indications including hepatocellular carcinoma with proof-of-concept results reported, but there are no approvals. The patent landscape centres on target loci and sequences and on activation-specific chemistry profiles, and the design layer offers several chemistry profile families as options so that alternatives are available for freedom-to-operate review.

9.4 Chemistry requirement — mandatory, minus the cleavage rule

Stability and immune-suppression requirements match the silencing modalities. Two things differ. Cleavage is not required, so cleavage-site protection does not apply and chemical placement has correspondingly more freedom. And nuclear localisation matters, so bulky additions that might impede it warrant caution.

The design layer holds several activation-specific chemistry profiles and presents them for comparison — an enhanced-stabilisation family, placement families derived from the activation literature, and self-assembling conjugate families — each balancing stability, potency and manufacturing complexity differently.

9.5 Delivery requirement — required

This modality must reach the nucleus, so cytoplasmic delivery is insufficient, and being double-stranded it cannot use carrier-free uptake. A vehicle is therefore mandatory, and the options are those of Modality I: receptor conjugation for hepatic targets, particle encapsulation or an alternative vehicle otherwise.

The additional requirement of nuclear entry affects vehicle choice. Unlike a silencing molecule acting in the cytoplasm, an activating molecule must cross the nuclear envelope after endosomal escape. The efficiency of that step is hard to quantify and differs by vehicle, so in activation programmes it is advisable to obtain measured nuclear distribution data when narrowing vehicle candidates.

9.6 Design deliverables

Positive controls are included in the deliverables because of this modality's predictive uncertainty. When an experiment shows no effect, it must be possible to distinguish a design problem from an experimental-system problem, and a literature-validated control makes that distinction possible.

9.7 Advanced capabilities — promoter resolution and activation prediction

9.7.1 The transcription start site is not a single point

The first difficulty in activation design is that 'where is the transcription start site' has no single answer. Most genes have multiple transcripts with different start points, and annotation databases may designate different representative starts. A shift of 100 nucleotides moves the whole target window, so this resolution governs design quality.

Design offers three resolution modes and states in the deliverables which was used: automatic derivation by merging the 5' ends of multiple annotated transcripts, a user-specified upstream/downstream window, and direct genomic coordinates.

9.7.2 Promoter confidence — a weighted sum over five elements

In automatic mode the confidence of a candidate start site is a weighted sum over five elements. Each is a different class of evidence, so one being weak can be offset by others; all five being weak means the locus is unfavourable for an activation attempt.

Element

Weight

What it reads

Interpretation

Annotation-based 5' ends

35%

Merged 5' ends of multiple annotated transcripts

The largest weight. Confidence is high when several transcripts support the same point

Core promoter motifs

15%

TATA box (about −30 ± 5), initiator element (−3 to +5)

Presence of classical core elements. Absence does not mean it is not a promoter

CpG island

15%

Sliding-window CpG island call (observed/expected ratio and GC content)

The majority of mammalian promoters overlap a CpG island

Chromatin state

30%

Active promoter marks and chromatin accessibility peak summits

The second largest weight — direct evidence of whether the locus can be activated

Cross-species conservation

5%

Multi-species alignment conservation score (optional)

Functional elements tend to be conserved, but exceptions are common, hence the low weight

Chromatin carries 30% because it is the enabling condition for this modality: at a fully closed locus, no strength in the other four elements lets the complex settle. Public chromatin data are however measured in particular cell lines and tissues, so this element is 'the state in that cell' and may not represent a customer's system — a limitation stated alongside the value.

9.7.3 Predicted regulatory tracks — filling gaps where data are absent

Many tissues and cell types have no public chromatin data. To fill that gap the design layer can optionally use a deep model that predicts regulatory tracks from sequence alone. The model predicts transcription, initiation, accessibility and binding signals at tens-of-bases resolution over input windows hundreds of kilobases wide, and the initiation output is used for start site estimation and promoter activity estimation.

Predicted tracks do not replace measurement, and where measurement exists it takes precedence. The deliverables state at which loci measurement was used and at which prediction was used.

9.7.4 Argonaute loading-route audit

Which Argonaute family member activation routes through may differ from silencing, and this depends in part on the 5'-terminal nucleotide. Design emits a loading-route prediction based on guide 5'-end identity as an audit item.

One common overstatement is corrected explicitly here. It is conventional to mark a 5'-adenosine as 'compatible with both Ago2 and Ago1', but the reported distribution is spread fairly evenly across the four family members. A 5'-adenosine is therefore not suited to a particular route but at risk of being distributed across several, and the deliverables label it a mixing risk. Marking it as compatible would promote candidates that in fact disperse.

9.7.5 Mapping the repressive axis — the activation detour

Transcripts in the opposite orientation around the target locus are mapped and assessed for whether they act as a repressive axis. Where such an axis exists, silencing it is a way of raising the target, and because it uses a mature silencing modality the development risk falls substantially.

Mapping uses opposite-strand transcripts from the reference annotation together with a separate noncoding transcript annotation. The deliverables present, for each candidate repressive axis, its positional relationship to the target, expression correlation, and whether it can be converted into a silencing design.

9.7.6 The positive control set — 13 duplexes

Because predictive uncertainty is high in this modality, a means of distinguishing a design problem from an experimental-system problem is needed when an experiment shows no effect. The design layer bundles 13 control duplexes for which activation has been reported in the literature, each carrying its target gene and its coordinate relative to the transcription start site.

The control targets include cell-cycle inhibitors, an adhesion factor, a growth factor, transcription factor families, tumour suppressor families and a nuclear receptor, giving enough diversity that at least one is expressed in a customer's system. A control that fails to work is the signal to check delivery or the cell system before the design.

9.7.7 Activation-specific chemistry profiles

Because cleavage is not required, chemical placement has more freedom, and nuclear localisation and loading efficiency matter instead. The design layer presents several chemistry profiles built on different placement philosophies, each balancing stability, potency and manufacturing complexity differently. Applying a silencing profile unchanged imposes a cleavage-site protection rule that constrains the design for no reason.

Positional immune-avoidance rules are applied as well: modifications are placed at the balance point between immune suppression and loading efficiency, on the basis of reports that 2'-modification at particular positions suppresses innate immune recognition.

9.8 What this modality cannot solve

10. Modality VI — miRNA Mimic and anti-miR / Antagomir (miRNA Axis Modulation)

10.1 Two opposite molecules

MicroRNAs are endogenous small RNAs of about 22 nt that load into Argonaute and repress tens to hundreds of target mRNAs simultaneously. Repression per individual target is modest, but the number of targets produces effects at network scale. When the abundance of a particular miRNA is abnormal in disease, there are two directions of intervention, and they yield entirely different molecules.

Mimic

Antagonist (anti-miR)

Role

Agonist — replaces the function of a lost miRNA

Antagonist — sequesters an overactive miRNA

Form

Double-stranded; the guide is the mature miRNA sequence

Single-stranded; the reverse complement of the mature miRNA

Mechanism

Loads into RISC and represses that miRNA's natural target set

Binds the endogenous miRNA and sequesters it sterically

Nuclease involvement

Argonaute-mediated repression (largely translational repression and destabilisation, without cleavage)

None — RNase H must not be recruited

Design freedom

Sequence fixed in advance; freedom exists only in chemistry and passenger

Freedom exists in length and chemistry

Principal off-target

Seed-mediated repression by the passenger strand

Cross-inhibition of other miRNAs sharing the seed

Chemical modification

Required

Required — uniform across every position, with a fully modified backbone

Delivery vehicle

Required (double-stranded)

Optional — carrier-free uptake works

10.2 The biology and design logic of a mimic

10.2.1 The deliverable is the target list, not the sequence

In mimic design the sequence is already determined: the mature miRNA entry is the guide strand. The question that actually has to be answered is therefore not 'what should be made' but 'what gets repressed if this is introduced'. The answer — the predicted target list — is this modality's central deliverable and simultaneously the basis of its safety assessment.

Target prediction by any single method is dominated by false positives. Selecting candidates by seed complementarity alone makes a large fraction of the transcriptome a candidate; selecting by thermodynamics alone fails to reflect accessibility in vivo. Design therefore combines evidence from five independent routes with weighting.

The predicted list is then refined against experimental evidence and expression context. Argonaute crosslinking-immunoprecipitation and degradome data check measured binding, conservation in the mouse ortholog adds evolutionary support, and tissue expression data filter for whether a target is actually expressed in the tissue of interest. Consensus voting raises precision at some cost in recall, and that balance is the correct direction for therapeutic design: a wrongly included target is a more expensive error than a missed one.

10.2.2 Seed heterogeneity as a trap

A mature miRNA is not a single sequence. Processing generates variants whose 5' end is shifted by one nucleotide, and a single-nucleotide shift at the 5' end moves the entire seed region. A shifted seed means a shifted target set, so for miRNAs with substantial 5' heterogeneity, predicting targets from the single dominant mature sequence does not represent the real biology. Design flags this heterogeneity and reports the major variants.

10.2.3 Safety axes

Axis

What is at stake

How it is assessed

Strand-loading asymmetry

If the passenger loads, an entirely different target set is repressed

Terminal free-energy differential; quantification of the passenger seed's off-target burden

Seed toxicity

Particular seeds repress essential genes broadly

Essential-gene and genome-wide seed burden indices

Innate immunity

GU-rich and CpG motifs stimulate receptors

Motif scanning and 2'-O-methyl placement

Population polymorphism

Common variants at seed-binding sites create inter-individual differences in efficacy

Allele frequency scan across the seed site

Pathway convergence

If repressed targets crowd into one pathway, unexpected phenotypes follow

Functional annotation convergence analysis of the target set

10.3 The biology and design logic of an antagonist

10.3.1 Sequestration is stoichiometric

An anti-miR does not cleave its target; it binds and holds it. It therefore has no turnover and acts stoichiometrically, requiring an amount commensurate with the number of target miRNA molecules in the cell. That produces a dose logic fundamentally different from the catalytic silencing modalities — although if binding is very strong it becomes effectively irreversible, extending duration.

Because the means of raising binding strength is chemistry, chemistry is efficacy itself in this modality. Uniform high-affinity modification with a fully phosphorothioate backbone is standard, and that same combination enables carrier-free uptake.

10.3.2 Length strategy — the whole family or just one member

Strategy

Structure

Target scope

When to choose it

Full-length steric block

Full reverse complement with uniform high-affinity modification and a fully phosphorothioate backbone

High selectivity for that one miRNA

When one specific miRNA must be inhibited precisely

Short seed-directed oligonucleotide

A high-affinity oligonucleotide of about 8 nt complementary to the seed

The whole family sharing that seed

When family members are functionally redundant and inhibiting one is compensated

Mixmer

Alternating high-affinity and natural residues

Intermediate — a compromise between selectivity and affinity

When following a structure with clinical precedent

Receptor-conjugated form

Any of the above plus a hepatocyte targeting ligand

Maximised local concentration in liver

Hepatic targets where systemic exposure should be reduced

The logic of the short seed-directed strategy comes from miRNA biology itself. MiRNAs sharing a seed have largely overlapping target sets and are therefore functionally redundant; inhibiting one is compensated by the others. Suppressing the whole family is often required for a phenotype, and a short oligonucleotide targeting only the seed fits that purpose. Where family members differ in expressing tissue and have diverged functionally, full-length targeting is appropriate.

10.3.3 The character of off-targets

Anti-miR off-targets have two layers. The first is co-sequestration of miRNAs other than the intended one, predicted from seed-sharing relationships. The second is incidental complementarity to mRNAs; because uniform high-affinity chemistry does not recruit RNase H, no cleavage occurs, but translational interference is possible. Design reports the two layers separately, and where paralog instructions are supplied, applies adjudication to mRNA complementarity hits.

10.4 Therapeutic significance

The strategic character of targeting a miRNA axis is that one molecule moves a network. Where single-gene silencing removes one node, miRNA modulation moves the regulatory layer over a set of nodes. Where a disease is not the failure of a single gene but a failure of regulation — fibrosis, metabolic reprogramming, immune microenvironment — this approach fits conceptually.

That is also the risk. Moving many targets at once leaves broad scope for unanticipated phenotypes, and the safety assessment burden exceeds that of single-target silencing. This is the principal reason the modality is less clinically mature than the silencing modalities. A hepatic anti-miR showing antiviral effect in clinical trials is widely cited as proof of concept, but there are no approved products.

The recommended development approach is to establish the convergence of the target set first. If predicted targets converge on a single pathway consistent with the therapeutic hypothesis, the network effect is likely to act in the intended direction; if they are scattered across unrelated pathways, the risk of unpredictable effects is high. This is why functional convergence analysis of the target set is part of the deliverables.

10.5 Design deliverables

10.6 Advanced capabilities — weighted evidence combination and safety axes

10.6.1 The five evidence weights — the actual coefficients

The central deliverable of mimic design is the target list, produced as a weighted sum over five lines of evidence. The weights are documented deterministic constants, so the same input reproduces the same list.

Evidence

Weight

What it measures

Why this weight

Context-based efficacy score

0.34

A real-valued score combining seed type, 3' supplementary pairing and site context features

The largest weight — the only real-valued evidence tied directly to efficacy

Machine-learning binding probability

0.24

Per-site probability from a model learning binding directly from sequence features

Complements patterns the rules do not capture

Target site accessibility

0.16

The probability that the site is open in the target 3' untranslated region

A physical prerequisite for binding — necessary but not sufficient, hence a middle weight

Cross-species conservation

0.14

Conservation of the site in the mouse ortholog

Evolutionary support; sites can function without being conserved, hence a low weight

Rule-based alignment consensus

0.12

Agreement among strict-seed search tools

A baseline for reproducibility; highly correlated with the other evidence, hence a low weight

The weights are published because they make the output interpretable. Whether a target ranks highly because of its context score or its accessibility changes how it should be validated, and the deliverables present the individual values of all five for each target.

10.6.2 Cross-checking against measured binding evidence

The predicted list is checked against experimentally observed binding data. Argonaute crosslinking immunoprecipitation shows where the complex actually bound, and degradome data show traces of actual cleavage. Both are independent of the prediction score, so targets that are both highly predicted and experimentally supported form the most trustworthy subset.

Measured data are queried on demand and cached locally; where a query fails, that axis is marked 'not queried' rather than scored zero, so that absence of experimental evidence is not read as absence of binding — in tissues and conditions where the experiment was never performed, absent data are normal.

10.6.3 Seed toxicity — quantifying the essential-gene burden

A miRNA mimic represses hundreds of transcripts through one seed, so if that seed broadly represses essential genes, cytotoxicity follows. Design computes a per-seed off-target frequency index to quantify that burden, presented at two levels — a genome-wide burden and an essential-gene-restricted burden, the latter being more directly connected to phenotypic risk.

In a mimic the sequence is fixed, so the seed cannot be changed. The index therefore serves risk awareness rather than candidate selection: a high-burden miRNA signals that dose design and safety assessment should be more conservative. The passenger strand's seed, by contrast, can be changed by design, so passenger seed burden is a genuine selection axis.

10.6.4 5' heterogeneity — when the seed moves wholesale

A mature miRNA is not a single sequence. Processing generates variants shifted by one nucleotide at the 5' end, and a one-nucleotide shift moves the entire seed region (positions 2–8). A shifted seed means a shifted target set, so predicting targets from the single dominant mature sequence stops representing the real biology.

Design computes the seeds of the canonical form and of the +1 and −1 shifted variants and reports the target set of each separately. For a miRNA with substantial heterogeneity, the union of those three lists is closer to the real scope of effect.

10.6.5 Population polymorphism at target sites

Where a common polymorphism exists at a seed binding site, repression does not occur in individuals carrying that variant. Design queries population allele frequencies at the binding sites of predicted targets and flags sites where polymorphism is common. This is not a reason to change the mimic, but it identifies in advance where response heterogeneity would appear clinically.

10.6.6 Functional convergence of the target set

Whether predicted targets converge on one pathway or scatter across unrelated ones is analysed. Convergence makes it likely the network effect acts in the intended direction; scattering carries a high risk of unpredictable effects. The analysis uses the same function-axis resources as Chapter 11, with the same pathway size cap and per-source separation rules.

The practical use is modality suitability. A miRNA for which convergence cannot be established is risky to address with this modality, and silencing the individual targets of that axis directly becomes the more predictable approach.

10.6.7 Choosing among the four anti-miR strategies

In the antagonist direction there is design freedom in the sequence, so strategy selection is a real decision. The four strategies of 10.3.2 separate on the following criteria.

Design derives the family member list from seed-sharing relationships and presents each member's tissue expression, providing the basis for judging whether the whole family may be suppressed. If one member carries a protective function in the target tissue, family-wide suppression is inappropriate.

10.7 What this modality cannot solve

11. Target Adjudication and Family Selectivity — Advanced Pathway and Network Analysis

Where the preceding chapters addressed what molecule to build, this chapter addresses which target, through which modality, and how far across its family to intervene. The deliverable is a verdict with evidence rather than a sequence, which is why it is described separately. In execution order it precedes design, and programmes that skip it fail not on design quality but on target definition: silence an essential gene with a well-designed molecule and the molecule is excellent while the experiment fails.

What fundamentally separates this layer from the other design stages is that its input is a graph rather than a sequence. Target suitability is not settled by examining one gene. What it interacts with, which pathways it belongs to, whether that pathway holds redundant members, what moves downstream when it is inhibited, and — if it cannot be touched at all — where else the same outcome can be produced: these are all network questions. This chapter describes how that network is constituted and what verdicts it produces.

11.1 The five questions this layer answers

#

Question

Which axis answers it

Consequence of no answer

Q1

May this gene be touched at all

Safety axis — population loss constraint, cell dependency, adverse-event precedent

An essential gene is silenced and the target cells die

Q2

In which direction should it be moved

Disease axis plus function axis plus the sign of the causal network

The wrong direction can worsen the disease

Q3

Which modality suits this gene

Variant spectrum, subcellular localisation, surface exposure of the protein

Resource is committed to a modality that cannot apply

Q4

If it cannot be touched, what should be touched instead

Two-step neighbourhood, signed causal layer, pathway ladder

The programme stops at 'this target does not work'

Q5

How far across the family should the intervention reach

Paralog segregation — functional sharing, disease association, redundancy

Compensation suppresses the phenotype, or a member that had to be spared is lost

Q4 shows most clearly why this layer exists. A target assessment that ends in 'unsuitable' stops a programme without advancing it. By contrast, a conclusion that 'this gene is pan-essential and cannot be silenced directly, but two downstream effectors connected to it by inhibitory edges are tissue-restricted and weakly constrained' creates the next step. Reaching that conclusion requires a directional network.

11.2 The evidence graphs — what is loaded

Verdicts rest on two large graphs: a protein interaction graph and a knowledge graph holding pathways, functions, diseases and phenotypes. The two are normalised to share one gene identifier space and therefore join directly — a sampled verification found zero mismatches. If the identifier spaces diverge, using the two graphs together produces silent empty results, so this normalisation is the precondition for any advanced analysis.

11.2.1 The protein interaction graph

Component

Scale

What it is used for

Integrated sources

13

Removes single-source bias; degree is also reported per source

Consensus interaction pairs

about 1.73 million

Interactions supported by multiple sources — the base set for first-shell partner search

Pairs from the large integrated source

about 6.70 million across 13 evidence channels

Evidence separated by channel (experimental, co-expression, literature, genetic interaction and so on). Channels are reported separately, not summed

Protein complexes

22,824

Complex members act as one functional unit and are considered together in target selection

Signed and directed causal edges

From curated signalling sources

Activation or inhibition sign with upstream and downstream direction — the core of substitute target search

Evidence channels are reported separately because they differ in reliability and in meaning. Summing an edge supported by a physical binding experiment together with one inferred from literature co-mention into a single score lets the latter, being far more numerous, dilute the former. When searching for substitute targets in particular, only edges with physical or causal support form a meaningful path.

11.2.2 The knowledge graph

Layer

Scale

Role

Pathways

4,228

Loaded separately per source (Reactome, KEGG, WikiPathways and others), with pathway size retained

Functional ontology

48,329 terms / 858,529 annotations

Separated into the three aspects: molecular function, biological process, cellular component

Gene–disease associations

about 28.09 million

Curated and text-mined sources distinguished, with a separate normalised score column

Paralog pairs

140,134

Held with sequence identity — the input to family selectivity adjudication

Population loss constraint

18,140 genes (including 814 on sex chromosomes)

pLI, LOEUF and LOEUF percentile. Holding the percentile alongside is the precondition for threshold setting

Tissue expression

Normalised per-tissue expression

Expression breadth, tissue-specificity index, top tissues

Cell dependency

1,208 cell lines

Gene effect values and the fraction of dependent cell lines

Clinical variant profiles

Decomposed by variant type

The composition of pathogenic variants — a direct input to modality suitability

Tractability assessment

Bucket classification plus adverse-event precedent

So that targets with existing failure precedent are not chosen again

Scale is not the same as quality. Most of the 28.09 million gene–disease associations are text-mined, and curated associations are a small fraction of them. This is exactly why the layer reports the two classes separately and keeps a separate normalised column — see 11.6.

11.3 The safety axis — may this gene be touched

Safety adjudication reads four independent lines of evidence together, because any one of them alone misses a substantial fraction of risky genes.

Evidence

What it tells you

Blind spot when used alone

Population loss constraint (pLI, LOEUF, percentile)

How strongly loss-of-function variation in this gene has been suppressed in human populations — a proxy for evolutionary essentiality

A single threshold misses many risky genes. Short genes appear weakly constrained for lack of statistical power, and sex chromosomes have coverage limits

Cell dependency (gene effect, dependent cell line fraction)

Whether removing the gene experimentally kills cells, and in what fraction of cell lines

Measured in cultured cells, so tissue-restricted essentiality and adult-tissue requirement are not reflected

Tractability bucket and adverse-event precedent

Whether development against this target has been attempted, and what went wrong

A novel target has no information — absence does not mean safety

Validated synthetic-lethal partners

The partner whose co-loss is lethal — the basis for tumour-selectivity strategies

Context-dependent and may not reproduce across cell types

Loss constraint is held with its percentile rather than only its raw value because of thresholding. The absolute value is affected by gene length and the number of observation opportunities, so a quantile basis is more stable for comparison across genes. The verdict presents raw value and percentile together so that a reader can reinterpret against their own criterion.

Notation of absent data matters most on this axis. Where constraint information is missing, the reason is stated — sex-chromosome coverage limits, a noncoding locus, insufficient statistical power — and a blank is never rendered as 'unconstrained'. A coverage limitation in the reference data turning directly into a false safety verdict is the most common and most expensive failure mode in this area.

11.4 The expression axis — where and how much is made

Expression information feeds two judgements: the choice of delivery target tissue, and the prediction of which tissues are affected under systemic exposure.

This axis changes verdicts when combined with the safety axis. Even a gene with high cell dependency may be addressable if its expression is confined to one tissue and that tissue is the target, because local administration can avoid systemic essentiality. Conversely, a broadly expressed essential gene cannot be avoided by any delivery strategy.

11.5 The function axis — pathway analysis

11.5.1 Why pathways are read per source

Pathway databases embody different curation philosophies. Some subdivide at the level of individual reactions; others group by disease or biological theme. It is common for the same gene to belong to twelve pathways in one source and three in another. Summing them lets the finer-grained source dominate, so the verdict presents pathways separately per source and reports the size of each.

Pathway size matters because it bears directly on functional sharing verdicts. That two genes share a top-level pathway such as 'cellular metabolism', which contains thousands of genes, is not evidence that they are functionally related. The verdict applies a size cap that excludes overly large pathways from counting as evidence of functional sharing, defaulting to 300 genes.

11.5.2 The three aspects of the functional ontology

Aspect

What it describes

Use in adjudication

Caution

Molecular function

What the protein does at the molecular level (enzyme activity, binding activity)

The primary input to family selectivity — does it do the same job

Sharing a molecular function term does not imply the same substrate. This is the main source of over-broad co-silencing verdicts

Biological process

Which process it participates in (cell cycle, immune response)

Consistency with the therapeutic hypothesis; target-set convergence analysis

Sharing at the process level does not imply functional redundancy

Cellular component

Where it resides (membrane, nucleus, secreted)

A direct input to modality suitability — surface exposure is the precondition for aptamer conjugates (Modality X)

Whether an annotation is experimental or predicted needs to be distinguished

That the cellular component aspect connects directly to modality adjudication matters in practice. An aptamer conjugate binds surface proteins from outside the cell and cannot in principle reach cytosolic or nuclear proteins. A request for an aptamer against a target with no surface or secreted annotation does not receive 'not applicable' but 'redirected to an interaction partner carrying a surface annotation', and that redirection uses the network axis.

11.5.3 The pathway ladder

In a hierarchically organised pathway database, looking only at the most specific pathway a gene belongs to loses context. The verdict presents that pathway's parent and grandparent as a ladder. This serves two purposes: grasping the target's biological context at a glance, and defining the search scope for substitute targets as 'other specific pathways under the same parent'.

11.6 The disease axis — association and the scale problem

Gene–disease associations differ in character by source. Curated sources often assert an association verified by an expert without grading it; statistical sources assign a score between 0 and 1; and text-mining sources use their own scale based on literature co-mention frequency.

Source type

Example

Score character

Treatment in adjudication

Curated (variant-based)

Clinical variant databases

Ungraded — asserts association only

Both score columns left empty. No arbitrary maximum is assigned

Curated (phenotype)

Phenotype ontologies, rare disease databases

Ungraded

Same

Integrated statistical

Drug target integration platforms

0–1 scale

Used directly in the normalised column

Text-mined

Literature co-mention channels

Own scale (for example 0–5)

Used in comparison only after mapping to the normalised column

The error that recurs in practice is comparing across sources on the raw scale column. A value of 3.9 from a source whose maximum is 5 always looks larger than 0.9 from a source whose maximum is 1, so a single literature co-mention outranks a curated causal association and clears every threshold set against a 0–1 scale. Only the column mapped to a common interval is safe to compare, threshold or rank across sources, and the verdict presents both columns while stating which one is comparable.

The disease axis is also a direct input to modality suitability. Decomposing pathogenic variants by type — premature stop, splice site, deep intronic, particular transitions, frameshift deletions — yields, as data, what percentage of patients each modality covers. That decomposition is the basis for sequencing development in rare genetic disease, and it is treated further in section 22.8.

11.7 The network axis — interactions and the signed causal layer

11.7.1 Why degree is reported per source

Total degree alone makes well-studied genes look like hubs, because more literature means more reported interactions. That is research bias rather than biological centrality. The verdict presents total degree together with per-source degree, so that connections concentrated in one source can be distinguished from those supported by several independent ones.

11.7.2 The signed causal layer — the core asset of this layer

A plain interaction edge says only that two proteins meet; it has neither direction nor sign. Finding a substitute target requires knowing what activates or inhibits what. The verdict loads signed, directed edges from curated signalling sources as a separate layer and presents upstream and downstream relative to the target.

Layer

What it holds

Which verdict it serves

Upstream activators

Regulators that raise the target's expression or activity

When the target must be lowered but cannot be silenced directly, silencing these is the detour

Upstream inhibitors

Regulators that lower the target's expression or activity

When the target must be raised (Modality V), silencing these is the activation detour — achieving an upward objective with a mature silencing modality lowers development risk

Downstream effectors (activated)

What the target switches on

When the target is pan-essential and untouchable, the downstream point that mediates the therapeutic effect

Downstream effectors (inhibited)

What the target switches off

The expected consequence of silencing the target — what is released

Sources sometimes report opposite signs for the same edge. The verdict marks these as 'sign disputed across sources', and that notation is itself a finding: a regulator whose direction depends on context is not a reliable point of manipulation. Picking one sign arbitrarily would destroy that information.

11.8 The two-step neighbourhood — searching for detours

Where direct partners do not yield a substitute target, the search extends to genes two hops away. Listing two hops raw is not informative, however, because most genes connect to most genes. Three treatments make it readable.

11.8.1 Sign composition

Signs along a path are multiplied to give the net effect. An inhibitor of an inhibitor is an activator, so silencing an inhibitor two hops away is a viable strategy for raising a target. The verdict shows, for each path, the signs of its constituent edges and the composed net sign.

Notation example
  HSPA8 --| CFTR                     direct: HSPA8 inhibits CFTR
  STUB1 --| HSPA8 --| CFTR           net: STUB1 activates CFTR   *also a direct partner
  CFTR --| MCC --| APC               net: CFTR activates APC

  --| inhibits   --> activates   the net sign is the product of signs along the path

11.8.2 Demoting hub intermediates

Paths through a hub gene with very many input edges exist between almost any pair of genes. The statement 'X acts on the target through some growth factor' is true of most of the genome and therefore carries no information. The verdict ranks paths through intermediates above a threshold (default 30 inputs) last and labels them as hub-mediated.

11.8.3 Marking overlap with direct partners

When a gene two hops away is also directly connected to the target, that fact is marked separately. Such a gene is more likely to belong to a coherent functional module than to lie on a chance path, so its credibility as a substitute target candidate is higher.

11.8.4 The second shell

Genes two hops out from the direct partner set are ranked by how many distinct first-shell partners bridge to them. A gene on which several first-shell partners converge is more likely to be central to that functional module, and has stronger support as a point of manipulation than a gene connected by a single path.

11.9 Modality suitability adjudication

Synthesising the six axes above, the layer adjudicates for each modality whether the target meets its requirements. Those requirements are independent of oligonucleotide quality, and where they are unmet no therapy results regardless of design quality.

Modality

Requirement imposed on the target

Which axis adjudicates

When unmet

siRNA · ASO · ncRNA silencing

Lowering it must be therapeutically correct, and it must not be pan-essential

Disease axis direction plus safety axis

Target cells die or broad toxicity appears

saRNA (transcriptional activation)

The therapeutic direction must be upward and the locus must be activatable

Disease axis direction plus chromatin context

The wrong direction can worsen the disease

SSO (splice switching)

The disease mechanism must run through splicing and the regulatory element must be characterised

Variant type decomposition

There is no point of intervention

ADAR editing

The pathogenic variant must be a transition that editing can revert

Variant type decomposition

Not applicable

Suppressor tRNA

Premature stop codons must account for a meaningful share

Variant type decomposition

Almost no patients are addressable

miRNA modulation

A disturbance of that axis must be established in the disease

Disease axis plus target-set convergence

The direction of the network effect cannot be predicted

Aptamer conjugate

The target protein must be surface-exposed or secreted

Cellular component ontology

An intracellular protein cannot be reached physically

Immunomodulatory oligonucleotide

The therapeutic purpose must be immune stimulation or suppression

Therapeutic hypothesis

The modality is unrelated to the target gene

11.9.1 Verdict values and redirection

Verdicts take four values — recommended, conditional, not applicable, no data — each presented with the target gene, the modality, the evidence and the caveats. 'Not applicable' is a result rather than a blank and keeps its evidence: where a gene is excluded for direction reasons and is also pan-essential, both reasons are shown. Keeping only one would mean a later run in the opposite direction loses the safety warning.

When a requirement is unmet, the correct answer is frequently a substitute target rather than 'not applicable', and that redirection uses the network axes of 11.7 and 11.8.

Three canonical redirections

  Aptamer requested against an intracellular protein
    -> redirected to a first-shell interaction partner carrying a surface or secreted annotation

  Silencing requested against a pan-essential gene
    -> redirected to a downstream effector in the signed causal layer that is weakly constrained
       and tissue-restricted

  Activation requested for a gene that must go down
    -> redirected to silencing an upstream inhibitor (an upward objective achieved with a mature
       silencing modality)

11.9.2 A worked verdict

Adjudication for a cystic fibrosis gene (12 recommended / 2 conditional / 2 not applicable)

  suppressor tRNA    that gene        341 pathogenic stop_gained variants -- 23% of its pathogenic alleles
  SSO                that gene        216 pathogenic splice-site variants -- 15%
  ADAR editing       that gene        139 pathogenic G>A -- exactly the transition A-to-I editing reverts
  aptamer conjugate  that gene        surface annotation present in the cellular component ontology
  saRNA              NFE2L2, YY1, GOPC
                                      this gene should go down, so activating it is backwards ->
                                      activate a curated upstream negative regulator instead
  ncRNA silencing    --               this gene is protein-coding; that modality targets ncRNA

11.10 Family selectivity adjudication — paralog analysis

11.10.1 The shape of the problem

Most genes have sequence-similar paralogs. Whether to co-silence or spare them cannot be decided automatically, because the correct answer inverts with the therapeutic hypothesis: where family members are functionally redundant and inhibiting one is compensated, all must be silenced; where each carries a distinct normal function, only the target must be.

This is why the adjudication depends on pathway and network analysis. Sequence identity alone cannot establish functional redundancy. Two genes at 60% identity may sit in entirely different pathways, while two at 30% identity may be interchangeable members of the same complex. The verdict uses sequence identity only as one input and takes functional sharing and disease association as its primary evidence.

11.10.2 Decision rules

Mode

Rule

Verdict

Evidence axis

Gene mode

The paralog shares at least one pathway or molecular function with the target

Co-silence

Function axis (pathways plus molecular function ontology)

Gene mode

The paralog sits only in a non-overlapping pathway

Spare

Function axis

Disease mode

Select only disease-associated paralogs — functional agreement alone is insufficient

All disease-related paralogs are co-silenced

Disease axis plus function axis

Disease mode

Present only in the normal pathway and not disease-linked

Spare

Disease axis

Safety override

The paralog is pan-essential

Converted to spare regardless of functional agreement

Safety axis (cell dependency plus constraint)

Defaults are sequence identity at or above 20%, a function term set size at or below 300, use of pathways and the molecular function ontology, and a minimum association of 0.1. The term set size cap exists for the reason given in 11.5.1 — sharing an overly large pathway is no evidence of functional relatedness.

11.10.3 The two auxiliary scores

redundancy score = 0.35 · sequence identity
                 + 0.25 · min(shared function terms / 8, 1)
                 + 0.20 · co-expression
                 + 0.20 · joint dispensability

priority score   = redundancy score
                 + 0.10  (validated synthetic-lethal partner)
                 + 0.15 · disease association
                 - 0.15  (documented adverse-event precedent)
                 - 0.10  (strong loss constraint)
                 - 0.10  (residual fitness dependency)

The redundancy score summarises whether this paralog can substitute for the target's function; the priority score summarises whether it may nonetheless be touched. They are kept separate because the two judgements diverge: a paralog that is functionally fully redundant but very strongly constrained scores high on redundancy and low on priority, and that combination is precisely the warning that co-silencing would be effective but dangerous.

11.10.4 The shape of the verdict output

paralog   identity  terms  dep%  LOEUF  co-exp  redundancy  priority  verdict
  A          62%      17     7%   0.41   +0.80      0.79      0.79    co-silence
             shares 17 terms, 11 of them specific
             shared: oligosaccharyl transferase activity ...  |  tractability: structure with ligand
             top tissue: ovary, tissue specificity 0.71

  B          31%       0   100%   0.40   +0.61      0.23        --    spare
             no shared pathway or molecular function term -- a different role, spare it

  Instructions to carry into design:
    --paralog-allow A        (keep only candidates that also silence A)
    --paralog-exclude B      (drop any candidate that touches B)

11.10.5 Interpretation limits that must be read alongside

11.10.6 How the verdict enters design

Verdicts act as two hard filters at the design stage. A co-silencing instruction keeps only candidates that silence all named paralogs; a sparing instruction removes any candidate that touches even one. Used together they express a compound requirement: these three must be hit, those two never.

Application differs by modality. Where there are many candidates the instruction is a filter over a list; where the sequence is fixed in advance and there is exactly one design object, as with a miRNA mimic, it is an adjudication on that one object. In that case, silently deleting an offending paralog from the predicted target list would be the opposite of useful, because that list is the off-target evidence the adjudication is made from. The list is preserved and only the verdict recorded.

The choice of application point also requires care. In a modality with several design branches by target class, such as ncRNA silencing, the filter must be applied at the upper point where final candidates converge; applying it to one branch leaves the others unfiltered. An unfiltered branch raises no error and simply produces output in which the instruction was not honoured, which makes it hard to detect.

11.11 Cleavage coverage — what 'knocking down' actually means

'Knocking down a gene' actually means knocking down some set of its isoforms. For genes with many isoforms two candidates commonly differ substantially in coverage, and that difference governs the interpretation of experimental results. Coverage is recomputed directly from the same transcript annotation the design engines read.

Item

Rule

Why

Representative transcript

Use the annotation's canonical isoform designation, falling back to the longest isoform

The value only agrees if the same criterion as the design engine is used

Definition of 'cut'

Only when the transcript contains the site verbatim, with no mismatch tolerance

The same exact-substring test the design engine uses

Coding-only counts

Separated by transcript biotype

So that nonsense-mediated-decay and retained-intron isoforms do not dilute the denominator

Alphabet normalisation

Unified to the RNA alphabet

Normalising to DNA instead makes every row report zero silently

Cleavage-site notation

Mechanism-dependent — a single phosphate (RISC), a gap window (RNase H1), or none (steric block)

The arithmetic is mechanism-independent but the notation belongs to the mechanism

One frequent confusion is worth stating. A candidate may hold the maximum coding-transcript coverage and still rank second overall. That is not a defect: the maximum is over coding transcripts alone, while the ranking sorts on coding and all-transcript counts together.

11.12 Cross-species correspondence — the basis for preclinical species selection

Before human clinical trials, safety assessment in one rodent and one non-rodent species is customarily expected. For oligonucleotide therapeutics that carries a particular meaning: meaningful pharmacology and toxicology are only possible if the molecule also binds the corresponding transcript in that species. Dosing a human-specific molecule into another species reveals chemistry class toxicity but nothing about target-related toxicity.

Species selection is therefore a design constraint rather than a matter of experimental convenience. This layer evaluates several human-primary species combination modes in parallel and presents which candidates survive in each. Because each combination demands a different set of cross-reactive species, the surviving candidate sets differ, and the question 'does a candidate exist that works in both rodent and non-human primate' can be answered with data.

That parallel evaluation is computationally heavy, so shared computations are reused under strict limits. Only pure functions may be cached; computations in which parallel threads write into one data structure in different ways are deliberately not cached, because a deterministic cache would double-count. A performance optimisation that changes the result is not an optimisation. Because of that distinction, each combination's result is equivalent to a standalone run.

11.13 Deliverables of this layer

11.14 What this layer cannot answer

12. Modality VII — ADAR-Recruiting Editing Oligonucleotide (AIMer · EON · arRNA)

12.1 Form and definition

The common name is an ADAR-recruiting editing oligonucleotide, known by lineage as an AIMer, an EON (editing oligonucleotide), an Axiomer or an arRNA. A single-stranded, chemically modified oligonucleotide, in two length classes. The short class is roughly 20–40 nt and modified at every position; the long class is a linear guide RNA of roughly 50–100 nt with a comparatively lighter modification burden. Both form a duplex with the target mRNA while deliberately placing a cytidine opposite the adenosine to be edited, leaving one mismatch.

What separates this modality from every other is that the molecule does nothing itself. It does not cleave, block, or bind as a ligand. Its only action is to create the geometry in which an enzyme the cell already possesses does its work, and the chemistry is performed by that endogenous enzyme. Introducing no foreign protein is the entire strategic argument for the modality.

12.2 Molecular mechanism

12.2.1 Adenosine deamination and how it is read

Human cells endogenously express adenosine deaminases acting on RNA. These enzymes recognise double-stranded structure and remove the amino group at position 6 of adenosine, converting it to inosine. Inosine is read as guanosine by the translation machinery and by reverse transcriptases, so the net effect is an A-to-G change. The phenomenon occurs naturally across the human transcriptome, concentrated in double-stranded regions derived from repeat sequences.

The therapeutic idea is simple. If the enzyme recognises double-stranded structure, then creating double-stranded structure artificially at a chosen position should cause editing there. It does, and the class of pathogenic variant to which this applies is a G-to-A change, since reverting the A to G restores wild type. That variant type accounts for a substantial share of human pathogenic variants, so the applicable scope is not narrow.

12.2.2 The A:C mismatch as a geometric requirement

For the enzyme to flip the target adenosine into its catalytic pocket, the duplex must be locally destabilised at that position. Placing a cytidine opposite the target adenosine to create an A:C mismatch supplies that destabilisation, and it is an absolute design requirement. Without the mismatch the duplex is fully paired and editing efficiency falls sharply, and other mismatch types have been reported to be less effective than A:C.

12.2.3 Nearest-neighbour preferences — which targets edit well

The enzyme shows clear preferences for the bases flanking the target adenosine. In the triplet context with the target A at the centre, a uridine on the 5' side and a guanosine on the 3' side — 5'-UAG-3' — is most favourable, and a guanosine on the 5' side is the least favourable. These preferences change editing efficiency severalfold and are therefore a primary ranking axis.

The design-relevant point is that this context is already fixed by the target: the position of the pathogenic variant is the editing site, so the context cannot be chosen. The axis therefore functions not as a ranking among candidates but as an upfront assessment of how favourable this target is for the modality, and for targets in unfavourable contexts it is reasonable to consider an alternative modality.

12.2.4 Bystander editing — this modality's characteristic off-target

Within the duplex the guide creates there are adenosines other than the intended one. If those are edited too, unintended amino acid changes result; this is bystander editing. The means of suppressing it is clear: placing a uridine opposite each bystander adenosine forms a normal Watson–Crick pair, which is stable and therefore not edited.

Design enumerates every adenosine within the duplex footprint and reports, for each, the opposing base placement and the residual editing risk. Beyond the footprint there is a second class of off-target: the guide may incidentally form a duplex with another transcript and edit an adenosine there. That class is predicted from sequence complementarity and is therefore subject to transcriptome-wide scanning.

12.2.5 A direct conflict between chemistry and activity

This modality carries a structural tension no other modality has. Chemical modification stabilises the oligonucleotide, but modification within the central window around the orphan cytidine lowers editing efficiency, understood to be because phosphorothioate and 2'-modifications there impede local flexibility and enzyme access. The measure required for stability therefore directly costs activity.

Design accordingly sets a symmetric unmodified window at the centre and protects everything outside it. A narrow window gives good stability and lower activity; a wide window the reverse. The optimum depends on target context and delivery route, so there is no single answer, and the deliverables compare several window widths.

That chemical constraint enlarges the delivery requirement. Because the centre of the molecule is chemically unprotected, plasma stability cannot be secured by chemistry alone and the carrier's protective role is correspondingly larger.

12.3 Clinical precedent and patent landscape

The modality has entered clinical development, with two distinct chemical lineages each running trials. One is a short, fully modified oligonucleotide class for which human editing was reported in a hepatic alpha-1 antitrypsin deficiency programme; the other pursued hepatic and ocular programmes with a separate chemical design. There are no approved products, and classifying the modality as development-stage is accurate.

The patent landscape has three layers. The first is conceptual — recruiting the endogenous enzyme via an A:C mismatch — with early claim families from both academic groups and companies. The second is chemical, specifying particular nucleotide analogues and backbone stereocontrol. The third is structural, specifying guide designs with particular secondary structures such as hairpins.

The design layer uses only generic chemistry within a research-use scope and does not implement any company's proprietary chemistry or structure. It instead surfaces the existence of those claim families as advisories, providing a starting point for a customer's freedom-to-operate review.

12.4 Therapeutic significance

The strategic position of this modality is the gap between gene therapy and silencing. Silencing reduces quantity but cannot repair sequence; gene therapy repairs sequence but carries the burdens of a vector and of permanence. RNA-level editing repairs sequence while the change remains transient and confined to the transcript.

That transience is both a limitation and an advantage. The limitation is the need for repeat dosing; the advantage is reversibility. If an unexpected adverse effect appears, stopping treatment returns the system to baseline — a safety property that approaches permanently altering the genome do not have. The value of that property is greatest in paediatric treatment and for novel targets lacking long-term safety data.

A second advantage is the naturalness of expression control. An edited transcript remains under endogenous regulation, so expression does not leave the physiological range. The problems of overexpression or lost tissue specificity that arise when a gene is introduced under an external promoter do not occur in principle.

12.5 Chemistry and delivery requirements

Region

Chemistry

Reason

Termini and flanks

2'-O-methyl and 2'-fluoro modification with a phosphorothioate backbone

Nuclease resistance and plasma stability

The central window around the orphan cytidine

Unmodified 2'-OH and phosphodiester linkages retained

Modification here directly reduces editing efficiency

Across the guide

Complementarity to the target maintained; uridine fixed opposite each bystander adenosine

Suppression of unintended editing

Delivery

Lipid nanoparticle or receptor conjugate

Because of the unmodified central window, stability cannot be secured chemically and the carrier's role is large

12.6 Design deliverables

12.7 Advanced capabilities — the bystander map and window-width search

12.7.1 A complete map of bystander adenosines

Every adenosine within the duplex footprint the guide creates is a potential bystander. Design enumerates them all and computes four things for each: distance from the target adenosine, the base placed opposite, that adenosine's triplet context, and a residual editing risk grade.

Risk depends jointly on distance and context for a reason. Editing probability falls with distance from the catalytic pocket, but a context close to the optimum can still be edited at distance. 'Far from the target' is therefore not sufficient reassurance, and a bystander in a favourable context must be protected by placing a uridine opposite it.

The consequence of bystander editing is not always the same, and that is assessed too. A bystander at the third codon position may be a synonymous substitution and harmless, while the first or second position changes the amino acid. Design reports the expected amino acid change for each bystander so that protection can be prioritised.

12.7.2 Searching the central unmodified window width

In a modality where stability and activity conflict directly there is no single optimum. Design evaluates several window widths and presents both axes for each as a comparison table.

Window width

Stability axis

Activity axis

When it fits

Narrow (orphan C ± 1–2)

High — most positions protected

Low — insufficient flexibility around the catalytic site

Systemic dosing with long circulating exposure and no strong carrier protection

Medium (orphan C ± 3–4)

Medium

Medium

The standard starting point

Wide (orphan C ± 5 or more)

Low — a long unprotected stretch

High — free enzyme access

Particle encapsulation, where the carrier supplies protection

That conflict connects directly to the delivery choice. Strong carrier protection allows a wide window and the activity that comes with it; conditions closer to carrier-free demand a narrow window for stability. In this modality chemistry and delivery must be decided together rather than in sequence.

12.7.3 Off-targets beyond the footprint

If the guide incidentally forms a duplex with another transcript, an adenosine there may be edited. This risk is predicted from sequence complementarity and is therefore subject to transcriptome-wide scanning. What differs from an ordinary silencing off-target scan is an additional condition: binding alone is not a problem unless an editable geometry is formed.

Design evaluates, for each complementarity hit, whether an arrangement corresponding to an A:C mismatch arises, and reports plain binding hits separately from editable hits. Only the latter is a substantive risk, and counting them together overstates it.

12.7.4 Choosing between the two length classes

Class

Length

Chemical burden

How the enzyme is recruited

Selection criterion

Short

about 20–40 nt

Modified at every position except the central window

The duplex itself is the recognition structure

Where chemical stability matters and carrier burden should be reduced

Long

about 50–100 nt

Comparatively light

A longer duplex presents a wider enzyme binding surface

Where editing efficiency is the priority and delivery can supply stability

The two classes also differ in patent landscape. The deliverables present separate advisories on the known claim families for each, and state that only generic chemistry within a research-use scope is used.

12.8 What this modality cannot solve

13. Modality VIII — Suppressor tRNA / ACE-tRNA (Premature Stop Codon Readthrough)

13.1 Form and definition

The common name is a suppressor tRNA, and a cognate design that restores the wild-type amino acid is specifically called an ACE-tRNA. A single-chain RNA of roughly 72–76 nt that folds through a cloverleaf secondary structure into an L-shaped tertiary structure — a transfer RNA. Where every preceding modality was a ligand that recognises a target, this is a functional molecule that works in the ribosome: its anticodon pairs with the premature stop codon, and the amino acid carried at its 3' end is added to the growing polypeptide.

The central design concept is cognate suppression. Reading a stop codon with an arbitrary amino acid inserts the wrong residue, and the protein may not fold correctly. A cognate suppressor tRNA is designed to restore the original wild-type amino acid at that position, which is achieved by using the natural tRNA body cognate to that amino acid as a scaffold and editing only the anticodon.

13.2 Molecular mechanism

13.2.1 Nonsense mutation as a disease class

When a single base substitution converts an amino acid codon into a stop codon, translation terminates early, producing a truncated protein or triggering degradation of the transcript by the nonsense-mediated decay machinery. Either way no functional protein is made. This variant type has been reported to account for roughly 11% of variants causing inherited disease, and it represents a substantial patient fraction in cystic fibrosis, muscular dystrophy, epidermolysis bullosa and other conditions.

The previous approach was small-molecule readthrough inducers, which lower ribosomal fidelity so that stop codons are skipped. That approach has fundamental limitations: the amino acid inserted is not controlled, and normal stop codons across the transcriptome are affected too. A suppressor tRNA addresses both — the inserted amino acid is determined by design, and selectivity arises from the identity and context of the stop codon.

13.2.2 A six-stage design logic

Stage

What it determines

Constraint

Scaffold selection

Which natural tRNA body cognate to the wild-type amino acid to use

The synthetase for that amino acid must still recognise and charge this tRNA. Scaffolds are ranked using gene copy number as an abundance proxy

Anticodon engineering

Replacing the anticodon triplet to be complementary to the premature stop codon

Wobble rules apply. The aminoacylation identity elements — the acceptor stem and the discriminator base — must be invariant, and invariance is explicitly asserted

Structure validation

Whether the edited sequence still folds into a cloverleaf

Covariance-model validation as primary, free-energy folding as secondary, plus checks on T-stem and D-stem topology

Readthrough efficiency prediction

The probability of winning the competition against termination factors

The base immediately following the stop codon, its surrounding context, and the P-site codon strongly affect efficiency

Safety assessment

Quantifying the burden on normal stop codons across the transcriptome

Every gene ending in the same codon and a similar context is at risk of C-terminal extension. Also checked: mischarging risk, clinical variant cross-reference, nonsense-mediated decay prediction

Output

Full tRNA sequence, anticodon, restored amino acid, efficiency and safety scores, optional DNA cassette

Modification-impact assessment and a composite score

13.2.3 Global burden is the decisive safety axis

A suppressor tRNA does not in principle read only the target gene. Other genes ending in the same stop codon in a similar context may also have their C-termini extended, and the human transcriptome may contain hundreds of such genes. If an extended protein loses function, aggregates, or mislocalises, broad toxicity follows.

A design that has not quantified this burden across the transcriptome is a design that looks good only at its target. The design layer scans the whole transcript annotation, enumerates genes sharing the stop codon and context, and separately flags those for which C-terminal extension is particularly dangerous — essential genes, transmembrane proteins, aggregation-prone proteins.

Two means of achieving selectivity exist. First, designing the anticodon to recognise only the stop codon the target actually carries leaves genes ending in the other two unaffected. Second, context preferences including the base immediately after the stop codon lower efficiency at genes with a different context. The resulting selectivity is partial but real, and it is quantifiable.

13.2.4 Nonsense-mediated decay as a second barrier

Many transcripts carrying a premature stop codon are degraded by the nonsense-mediated decay machinery. However efficient the suppressor tRNA, there is no effect if there is no transcript left to read. Susceptibility depends on the position of the stop codon — upstream of the last exon–exon junction it is likely to be a decay substrate, downstream it escapes — and the design emits that prediction.

Two practical implications follow. For highly decay-susceptible targets, a suppressor tRNA alone has limited effect and combination with decay inhibition is discussed. And within one gene, suitability for this modality varies with variant position, so stratifying the patient population by variant position is strategically advantageous.

13.3 Development status

The modality is moving from proof of concept into preclinical development. Reports have accumulated that cognate suppressor tRNAs read premature stop codons and restore functional full-length protein in cell and animal models (Lueck and colleagues, Nat Commun 2019, among others), and several companies are running programmes differing in delivery and indication. There are no approvals.

The principal development issues are three. Delivery: a folded functional RNA does not undergo carrier-free uptake, so either an expression cassette must be delivered by vector or particle, or a synthetic tRNA must be delivered in a particle. Global readthrough burden, as described above. And control of expression level, since an overexpressed suppressor tRNA can broadly affect normal translation termination.

13.4 Chemistry requirement — not required in principle

This is the one modality in which chemical modification is not required in principle, and the reason lies in the nature of the molecule. It is not a synthetic ligand but a functional RNA that must fold inside the cell, be recognised by a synthetase and charged with an amino acid, and bind an elongation factor to enter the ribosome. All three processes depend on the precise structure and modification pattern of a natural tRNA, and an artificial modification perturbing any one of them disables the molecule.

The standard route is therefore delivery as a DNA cassette so that transcription, modification and folding occur inside the cell. In that case the deliverable is not an RNA sequence but an expression cassette — promoter, tRNA gene, terminator. Where a synthetic tRNA is dosed directly, minimal stabilisation is discussed, restricted to a scope that does not perturb native modification sites or tertiary structure.

Design separately assesses the impact of modification sites. tRNAs carry numerous chemical modifications that contribute to folding stability and codon recognition fidelity, and anticodon editing that alters a modification site can produce unexpected results.

13.5 Delivery requirement — required

Delivery mode

Form

Advantages

Constraints

DNA expression cassette in a lipid nanoparticle

A tRNA gene under a polymerase III promoter

Native transcription, modification and folding occur in the cell

Expression level is hard to control and duration is short

DNA expression cassette in a viral vector

The same cassette carried by a vector

Sustained expression, and tissue tropism can be conferred

Vector-associated immunogenicity and the burden of permanence

Synthetic tRNA in a lipid nanoparticle

Transcribed and purified, or chemically synthesised, tRNA

Expression level is controlled by dose; transient

Possible loss of activity from absent modifications; stability

In every case a vehicle is required, which raises the development burden. Where the target tissue is well defined, however — airway epithelium in cystic fibrosis, muscle in muscular dystrophy — local administration to reduce systemic exposure is a viable strategy.

13.6 Design deliverables

Scaffold bodies are taken from actual gene sequences in a public tRNA database; no sequence is invented. Design in this modality uses no machine learning and is a fully rule-based deterministic pipeline, because the structural and identity-element requirements are precisely specified and what matters is detecting rule violations definitively rather than estimating them statistically.

13.7 Advanced capabilities — identity element preservation and global burden quantification

13.7.1 Invariance checking of the aminoacylation identity elements

The most important constraint in this design is: change the anticodon, do not change the identity. The synthetase reads particular positions of the tRNA to decide which amino acid to load, and disturbing those positions loads the wrong amino acid or prevents charging altogether. Design compares the sequence before and after editing, asserts explicitly that the identity elements are invariant, and terminates as a failure on violation.

The last item matters in practice. For some amino acids the synthetase reads the anticodon directly, so changing it means the tRNA is no longer charged with the original amino acid. Design first checks whether the wild-type amino acid falls into that category and, if so, searches for an alternative scaffold or returns a not-applicable verdict.

13.7.2 Scaffold selection and the abundance proxy

Several tRNA genes exist for the same amino acid, and which is used as the scaffold affects the outcome. Building on an abundant family makes it more likely that the expression, processing and modification pathways are already optimised for that sequence. Design ranks scaffold candidates using gene copy number as an abundance proxy, and states as a limitation that copy number is not perfectly proportional to actual abundance.

13.7.3 Structure validation in three layers

Layer

Method

What it confirms

Meaning of failure

Primary — covariance model

Alignment score against a tRNA-family covariance model

Whether the edited sequence still has a structure recognised as a tRNA

The cloverleaf itself has collapsed — discard the design

Secondary — free-energy folding

Minimum free energy secondary structure prediction

Whether the predicted structure matches the cloverleaf and whether alternatives compete

Competing folds lower the fraction of functional molecules

Tertiary — domain checks

T-stem elongation factor binding proxy, D-stem topology check

Whether the tertiary elements required for ribosome entry are retained

It folds but does not enter the ribosome

All three layers must pass because the failure points differ. A molecule that does not fold, one that folds but is not charged, and one that is charged but cannot enter the ribosome all produce the same observation — it does not work — while demanding completely different responses.

13.7.4 Readthrough efficiency — the context model

The probability of winning the competition against termination factors depends far more on the context around the stop codon than on the codon itself. The base immediately following the stop is particularly decisive, so the tetranucleotide formed by the stop codon plus that base is reported as a principal determinant of termination efficiency. Design predicts readthrough efficiency with an extended model that adds the P-site codon and roughly six nucleotides of context on each side to that tetranucleotide table.

The practical value of the model is patient stratification. Even for the same variant type in the same gene, a different context gives a substantially different predicted efficiency, so which variant-position subgroup is likely to respond can be distinguished in advance.

13.7.5 Global stop codon burden — quantified transcriptome-wide

This is the decisive safety axis for a suppressor tRNA, and it is quantified by scanning the whole transcript annotation. The procedure is as follows.

1. Fix which stop codon the target carries (one of three)
2. Enumerate every gene in the annotation ending in the same stop codon
3. Compute context similarity (the tetranucleotide plus flanks) for each
4. Classify high-similarity genes as the readthrough risk set
5. Within that set, flag genes for which C-terminal extension is especially harmful:
     - essential genes (loss of function on extension would be lethal)
     - transmembrane proteins (extension changes localisation)
     - aggregation-prone proteins (extension induces aggregation)

Two means of securing selectivity exist and both are partial. Designing the anticodon to recognise only the stop codon the target carries leaves genes ending in the other two unaffected, and context preferences lower efficiency at genes with a different context. The resulting selectivity is quantifiable but cannot be reduced to zero.

13.7.6 Nonsense-mediated decay susceptibility

Many transcripts carrying a premature stop are degraded by the decay machinery, and susceptibility depends on the position of the stop codon: upstream of the last exon–exon junction (by roughly 50 nucleotides or more) it is likely a decay substrate, while downstream or in the last exon it escapes. Design compares variant position with junction coordinates to emit that prediction.

The output has two uses. For highly susceptible targets a suppressor tRNA alone has limited effect and combination strategies must be considered; and because suitability varies with variant position within the same gene, it forms the basis for patient stratification.

13.7.7 Modification impact and the expression cassette

tRNAs carry numerous chemical modifications that contribute to folding stability and codon recognition fidelity. If anticodon editing alters a modification site or the recognition sequence of a modifying enzyme, unexpected results can follow, so design checks for overlap between known modification positions and the edited positions.

The standard deliverable is an expression cassette rather than an RNA sequence — a polymerase III promoter, the tRNA gene body and a terminator — because delivery in that form allows native transcription, modification and folding to occur inside the cell. A direct synthetic tRNA route is also supported, with the possible loss of activity from absent modifications stated in the deliverables.

13.8 What this modality cannot solve

14. Modality IX — CpG-ODN / Immunomodulatory Oligonucleotide (TLR7 · TLR8 · TLR9)

14.1 Form and definition

The common name is an immunomodulatory oligonucleotide; the TLR9 agonist class is specifically called a CpG-ODN (CpG oligodeoxynucleotide) or an immunostimulatory sequence (ISS), and the inhibitory direction an INH-ODN. A single-stranded DNA or RNA of 18–30 nt with a largely phosphorothioate backbone. What decisively separates it from the preceding modalities is that the target is a protein rather than a nucleic acid. This molecule does not recognise anything by base pairing. It presents the features that endosomal Toll-like receptors read as 'foreign nucleic acid' — unmethylated CpG dinucleotides, uridine-rich sequence — and thereby activates or blocks the receptor.

The meaning of 'off-target' therefore differs here. In other modalities an off-target is unintended binding to a transcript; here it is both unintended immune activation and incidental gene repression through sequence complementarity. Design assesses both layers.

14.2 Molecular mechanism

14.2.1 Three receptors and their ligands

Receptor

What it recognises

Response on activation

Therapeutic direction

TLR9

DNA carrying unmethylated CpG motifs

Type I interferon from plasmacytoid dendritic cells; B cell activation

Vaccine adjuvant, immuno-oncology, infection prophylaxis

TLR7

Uridine- and guanosine-rich single-stranded RNA

Predominantly type I interferon from plasmacytoid dendritic cells

Antiviral stimulation, adjuvant

TLR8

Uridine-rich single-stranded RNA (pronounced in humans)

Predominantly inflammatory cytokines from monocytes and macrophages

Reprogramming the tumour microenvironment

All three (antagonism)

Inhibitory sequences — telomeric repeats, guanosine runs, particular triplets

Suppression of receptor signalling

Autoimmune and chronic inflammatory disease

That these receptors reside in endosomes has two design implications. First, the molecule must be endocytosed to reach its target, so endocytosis is itself delivery; the phosphorothioate backbone mediates that uptake, so no separate vehicle is needed. Second, the endosomal interior is acidic and contains nucleases, so the molecule needs stability under those conditions.

14.2.2 Structural classes of CpG oligonucleotide

TLR9 agonists produce markedly different immune profiles according to structure. This classification is not academic tidiness but a set of practical options, since the class is chosen according to the kind of immune response desired.

Class

Structural features

Immune profile

Suited to

Class A (D-type)

Palindromic CpG core on a natural backbone with phosphorothioate guanosine runs at both termini

Strong type I interferon from plasmacytoid dendritic cells; weak B cell activation

Antiviral stimulation, natural killer cell activation

Class B (K-type)

Fully phosphorothioate linear with multiple CpGs

Predominantly B cell activation and inflammatory cytokines; weak interferon

Vaccine adjuvant — enhancing antibody responses

Class C

Fully phosphorothioate and palindromic

Intermediate — induces both interferon and B cell activation

Where a balanced response is needed

Class P

Two palindromes forming higher-order concatamers

Very strong interferon induction

Where a powerful interferon response is required

Class A carries guanosine runs at both termini because those segments form G-quadruplex structures that assemble the molecules into higher-order aggregates. That aggregation directs trafficking to a particular endosomal compartment, and signalling from that compartment leads to the interferon pathway. In this modality, therefore, secondary and higher-order structure is part of the function rather than a side effect, and design treats G-quadruplex propensity as something to be tuned rather than avoided.

14.2.3 Species specificity — the central trap in preclinical interpretation

The optimal recognition motif for TLR9 differs between species. The optimal hexamer in humans differs from that in mice, and a strong agonist in one may be weak in the other. This bears directly on the interpretation of preclinical data: failing to distinguish a molecule problem from a species difference leads to wrong development decisions.

TLR8 is more extreme. It is functional in humans but responds very differently in mice, which limits the value of evaluating a human TLR8 agonist in murine models at all. Design evaluates the optimal motif per species and emits cross-species concordance as a separate item, providing the basis for preclinical species selection and data interpretation.

14.2.4 The mechanism of antagonists

Inhibitory oligonucleotides interfere with the receptor itself or its signalling pathway. Telomere-derived repeat sequences, guanosine runs and particular triplet repeats have been reported to carry inhibitory activity, and 2'-O-methyl RNA is known to block TLR7 and TLR8. Where chronic innate immune activation by nucleic acid autoantigens forms part of the pathology of an autoimmune disease, intervention in this direction fits logically.

14.3 Clinical precedent and therapeutic significance

This is one of the few non-silencing oligonucleotide modalities with approval precedent. A hepatitis B vaccine using a class B CpG oligonucleotide as adjuvant was approved in 2017, establishing that a synthetic oligonucleotide can obtain regulatory approval as a vaccine adjuvant. What makes that case significant is the scale of the safety dataset: adjuvants are given in large numbers to healthy adults, so the approval created broad safety evidence for that chemistry and dose range.

From a vaccine development standpoint the modality's value is that the direction of the response can be steered. Traditional aluminium salt adjuvants induce mainly antibody responses and weak cell-mediated immunity. A CpG oligonucleotide can selectively strengthen the interferon pathway or the B cell pathway depending on class, so the type of immune response can be designed to suit the pathogen. This is particularly useful for pathogens where cell-mediated immunity is protective and for populations such as the elderly whose responses are weak.

In immuno-oncology, intratumoral administration to reprogramme the microenvironment is under study. The tumour microenvironment is generally immunosuppressive, and the logic is to stimulate innate immunity locally to relieve that suppression and improve the effect of checkpoint inhibitors. In this approach minimising systemic exposure is the key to safety, and local administration with short residence becomes an advantage rather than a limitation.

14.4 Chemistry requirement — the modification is the function

Here phosphorothioate is both a stabiliser and part of the function. The backbone mediates endocytosis, provides stability inside the endosome, and in some classes participates in higher-order structure formation. The strategy used in other modalities — reduce phosphorothioate to lower class toxicity — therefore does not transfer directly.

Class

Backbone chemistry

Reason

Class A agonist

Natural backbone in the core, phosphorothioate only in the terminal guanosine runs

A natural core favours optimal recognition while terminal phosphorothioate carries higher-order structure and stability

Class B and C agonists

Fully phosphorothioate

Maximises stability and uptake; being linear, higher-order assembly is not required

TLR7/8 agonists

Phosphorothioate RNA preserving uridine-rich sequence

RNA is less stable than DNA, so backbone protection matters more

Antagonists

Phosphorothioate combined with 2'-O-methyl RNA

2'-O-methyl is an active component of TLR7/8 blockade

A caution: in agonist design, heavy 2'-modification around the CpG motif abolishes recognition. The receptor reads specific chemical features, so piling on modifications for stability produces a molecule that is not recognised. Design treats the region around the motif as protected and secures stability outside it.

14.5 Delivery requirement — unnecessary

Because the target is a receptor inside the endosome, endocytosis is itself arrival, and the phosphorothioate backbone mediates that endocytosis, so no separate vehicle is required. In the approved case the oligonucleotide is formulated together with the vaccine antigen without a carrier. DNA CpG oligonucleotides are sometimes formulated with aluminium salts, not for delivery but for co-localisation with antigen and release control.

Attempts to increase targeting to particular cells do exist, since acting only on selected immune cell subsets would give the desired response without systemic cytokine release. Antibody or ligand conjugation and particle formulation are under study, and in those cases a vehicle is introduced.

14.6 Design deliverables

Prediction is rule-primary: literature-anchored motif scoring makes the first-pass call, and a light machine-learning rerank is optional and degrades to the identity function when no model is present. Machine learning never overrides a hard rule call by design, because immune response is an area with comparatively well-codified rules and detecting rule violations definitively matters more than statistical estimation.

14.7 Advanced capabilities — species specificity, structural scoring, three-layer safety

14.7.1 Species-specific recognition motifs — the premise for preclinical interpretation

The optimal TLR9 recognition hexamer differs by species. Humans and primates are reported to prefer 5'-GTCGTT-3' while mice prefer 5'-GACGTT-3', and that difference makes the same molecule behave quite differently in the two species. Design computes the occurrence and position of the species-optimal motif separately for each species and emits cross-species concordance as a distinct item.

Without that item, preclinical data cannot be interpreted: a weak effect in mice cannot be attributed to the molecule rather than to a species difference. Conversely, a molecule with high concordance allows comparatively trustworthy extrapolation of preclinical results — though TLR8 differs so much between species that this extrapolation is limited in principle.

14.7.2 Structural class calling and its consequences

Design calls the CpG oligonucleotide's class automatically from sequence structure. The call rests on the presence of palindromes, the number and spacing of CpGs, the presence of terminal guanosine runs, and the backbone pattern, and the expected immune profile follows from it. Automation matters because it catches cases where the class contradicts the design intent — a molecule intended as class B may form a palindrome and behave as class C.

14.7.3 Structural and thermodynamic scoring

In this modality secondary and higher-order structure is something to be tuned rather than avoided. Four structural indices are computed.

Index

How it is computed

How it reads in this modality

Melting temperature

Nearest-neighbour thermodynamics with a backbone chemistry correction

Phosphorothioate lowers duplex stability, so the value is overestimated without the correction

Secondary structure free energy

Minimum free energy folding

Self-duplex formation obstructs receptor recognition and is to be avoided

G-quadruplex motif

Detection of contiguous guanosine patterns

Part of the function in class A and a risk in the other classes — the sign inverts with class

Motif accessibility

Local opening probability at the CpG motif position

A motif buried in a hairpin is not recognised

The third row shows this modality's peculiarity. In other modalities a G-quadruplex is something to avoid because it complicates synthesis and purification and causes non-specific binding, whereas in a class A CpG oligonucleotide it is a functional element that drives trafficking to a particular endosomal compartment through higher-order assembly. That the same index carries opposite sign depending on class is what makes automatic classification necessary.

14.7.4 Three-layer safety assessment

Layer

What it examines

Method

Why it is needed

Sequence off-targets

Unintended gene repression through incidental complementarity

Melting-temperature-weighted sequence similarity search

Even when immune modulation is the purpose, the molecule is still a nucleic acid and binds complementarily

Cross-species concordance

Difference in activity between preclinical species and human

Comparison of the per-species optimal motif assessment

Quantifies how far animal data can be extrapolated to humans

Immunotoxicity

Excessive or unintended immune activation

Composite assessment of G-quadruplex, self-duplex, backbone burden and CpG density

Immune stimulation is the objective here, so judging the appropriate level is especially important

The third layer is particularly subtle in this modality. Assessing immunotoxicity in a molecule whose purpose is immune stimulation means examining the magnitude and type of stimulation rather than its presence, and the boundary between wanted and unwanted response shifts with indication. A vaccine adjuvant aims at a local, transient response; intratumoral dosing tolerates a stronger one; and in the autoimmune antagonist direction any stimulation is a failure. The deliverables present component values rather than an absolute grade so that they can be read in the context of the indication.

14.7.5 Rules primary, learning auxiliary

Prediction is rule-primary. Literature-anchored motif scoring makes the first-pass call, and a light machine-learning rerank is optional and degrades to the identity function when no model is present. Machine learning never overrides a hard rule call by design.

That arrangement was chosen because the rules in this area are comparatively well codified and detecting rule violations definitively matters more than estimating statistically. A TLR9 agonist with no CpG motif, or a TLR7/8 agonist short of uridine, is not a statistically low-scoring molecule but a structurally non-functional one, and that verdict must be definite.

14.8 What this modality cannot solve

15. Modality X — Aptamer Conjugates ApDC / AOC (Targeted Delivery)

15.1 Form and definition

A folded single-stranded nucleic acid — an aptamer — joined through a linker to a payload. Where the payload is a small-molecule cytotoxin the construct is an aptamer–drug conjugate; where it is a therapeutic oligonucleotide it is an aptamer–oligonucleotide conjugate. Aptamers are typically 25–90 nt and fold into a defined tertiary structure that binds a specific site on a protein surface.

This modality sits at a different level from the others. It does not produce a therapeutic effect itself but carries something else to a destination. Where the preceding modalities addressed what to do, this one addresses where to go. It occupies, using nucleic acid, the role antibody–drug conjugates have already established.

15.2 The biology and physical chemistry of aptamers

15.2.1 Why a nucleic acid binds a protein

A single-stranded nucleic acid pairs partially with itself to form a tertiary structure of stems and loops. The surface that structure presents has shape and charge distribution much as an antibody's complementarity-determining regions do, and it can bind a complementary protein surface. Binding constants in the sub-nanomolar range have been reported, comparable to antibodies.

There are three advantages over antibodies. Chemical synthesis gives high batch-to-batch homogeneity and no biological contamination risk. Low molecular weight favours tissue penetration. And immunogenicity is low, since nucleic acids are not presented through the major histocompatibility complex as proteins are. The disadvantages are equally clear: nucleases degrade them, low molecular weight means rapid renal excretion, and there is no reliable computational method for predicting binding affinity.

15.2.2 Affinity is not computationally predictable

This fact determines the entire design policy of the modality and is therefore stated explicitly. No computational method currently exists that reliably predicts the binding affinity of an arbitrary nucleic acid sequence to an arbitrary target protein. Secondary structure prediction is well established, but with what affinity that structure binds a particular protein surface is a different order of question.

The design policy therefore is: binding constants are not generated. Aptamers used in conjugate design are restricted to experimentally validated ones, each carrying a citable sequence, binding constant and literature source. Uncertain entries are left empty and marked unconfirmed rather than filled in. The policy is inconvenient, but far cheaper than allowing an unsupported binding constant into a development decision.

15.2.3 Internalisation — binding alone is not enough

To bring a payload inside a cell, the receptor the aptamer binds must be endocytosed. A receptor that binds at the surface without internalising is unsuitable as a conjugate target. Target selection must therefore consider not only whether the protein is surface-expressed but whether it internalises and how quickly it recycles, and the design includes known internalisation characteristics among its assessment items.

Beyond internalisation lies a second barrier, endosomal escape. For cytotoxin payloads the common design has the linker cleaved in the lysosome so the freed drug crosses the membrane; for oligonucleotide payloads, which cannot cross membranes by themselves, escape efficiency becomes the limiting factor.

15.3 The three components of a conjugate

Component

Options

Design variables

Failure mode

Aptamer

Indexed by target receptor from a validated library

Sequence, stabilising chemistry, attachment point

Payload attachment that disrupts the fold abolishes binding

Linker (cytotoxin)

Cleavable — protease-sensitive peptides, acid-sensitive, disulfide; or non-cleavable stable covalent

Cleavage conditions, length, hydrophilicity

Premature cleavage in circulation causes systemic toxicity; failure to cleave means no effect

Linker (oligonucleotide)

Stable linker or disulfide

Length and flexibility

Too short and the two moieties interfere; too long and renal excretion increases

Payload (cytotoxin)

Microtubule inhibitors, topoisomerase inhibitors, DNA-binding agents, RNA polymerase inhibitors and others

Potency, whether freed drug kills neighbouring cells

Insufficient potency for the number of molecules internalised

Payload (oligonucleotide)

A silencing guide or duplex, or an antisense oligonucleotide

Inherits the payload modality's own requirements

The payload's chemistry may interfere with aptamer folding

Drug-to-aptamer ratio

Typically 1–4

Number and position of attachment points

High ratios increase aggregation and clearance; low ratios lack potency

The oligonucleotide-payload conjugate is conceptually interesting because it fuses two modalities. A silencing molecule is potent but delivery-limited; an aptamer is a delivery device with little therapeutic effect of its own. Combining them makes cell-type-specific silencing possible without a lipid nanoparticle or a receptor conjugate on the oligonucleotide itself, which is attractive particularly outside the liver. The endosomal escape problem described above remains, however, so practical use of this combination depends on improving escape efficiency.

15.4 Clinical precedent and therapeutic significance

Aptamers themselves have approval precedent. An aptamer binding a vascular growth factor was approved in ocular disease in 2004, and an aptamer binding a complement factor in the same field in 2023. Both are administered locally in the eye, a strategy that circumvents renal excretion and nuclease degradation through local administration.

There is no approval precedent for the conjugate form. The principal development issue is the limited number of validated aptamers. Antibodies can be raised against essentially any target through immunisation, whereas aptamers require an in-vitro selection process that does not succeed for every target. The applicable scope of this modality is therefore currently limited to targets for which a validated aptamer exists.

Areas of strategic value nonetheless exist: targets where antibody conjugates are already crowded and differentiated properties are wanted; repeat-dosing situations where antibody immunogenicity is problematic; and tissues that require low molecular weight for penetration.

15.5 Chemistry and delivery requirements

Chemical modification is required. An unmodified aptamer is degraded within minutes by serum nucleases, so the standard is stabilisation with a combination of 2'-fluoro and 2'-O-methyl plus an inverted deoxythymidine at the 3' terminus to block exonucleases. Polyethylene glycol is sometimes added to slow renal excretion.

Modification is constrained, however: changing the chemistry of residues involved in binding can destroy affinity, so the binding interface must retain its original chemistry. Which residues are involved can only be known from structural information or experimental scanning, so design incorporates such information where it is reported and is conservative where it is not.

No delivery vehicle is required. The aptamer is itself the targeting device, and adding a carrier would dilute that targeting while increasing molecular weight. Administration is intravenous or local.

15.6 Design deliverables

15.8 What this modality cannot solve

15.7 Advanced capabilities — combinatorial search and the unconfirmed policy

15.7.1 Searching the three-component combination

Conjugate design is a combinatorial problem across aptamer, linker and payload. The axes are not independent, so combination-level evaluation is required rather than per-axis optimisation: a bulky payload demands a longer linker, a longer linker accelerates renal excretion, faster excretion requires a higher drug-to-aptamer ratio, and a higher ratio raises aggregation risk.

Axis

Scale of options

Principal constraint

Coupling to other axes

Aptamer

A validated library of 20 or more, indexed by target receptor

Only validated entries are used — novel discovery belongs to experiment

Attachment point affects folding; payload size affects fold stability

Linker

Eight or more cleavable and non-cleavable types (protease-sensitive peptides, acid-sensitive, disulfide, glucuronide, carbonate, dual-cleavage)

The window between premature cleavage in circulation and failure to cleave intracellularly

Payload type sets the cleavage condition; length affects excretion

Payload (cytotoxin)

Ten or more — microtubule inhibitors, topoisomerase inhibitors, DNA-binding agents, RNA polymerase inhibitors

Very high potency is required because few molecules internalise

Lower potency demands a higher drug-to-aptamer ratio

Payload (oligonucleotide)

A silencing guide or duplex, or an antisense oligonucleotide

Inherits the chemistry requirements of its own modality

The payload's chemistry may interfere with aptamer folding

Drug-to-aptamer ratio

1–4

High raises aggregation and clearance; low lacks potency

The number and position of attachment points set the ceiling

15.7.2 Quantitative items emitted

15.7.3 The unconfirmed-marking policy in practice

The most important design policy in this modality is not to invent values. Because no reliable computational prediction of aptamer binding affinity to an arbitrary target exists, every binding constant in the library is a literature reference value with its source shown. Entries for which the literature has no value, or where the conditions are unclear, are left empty and marked unconfirmed.

The policy is as inconvenient as it is valuable. An unsupported binding constant entering a development decision contaminates the entire candidate selection and is only discovered at the experimental stage. A blank honestly exposes the absence of information, and that absence is itself actionable — it says this item must be established experimentally.

15.7.4 Special considerations for oligonucleotide payloads

Where the payload is a silencing molecule, two modalities' requirements must be satisfied at once. The payload must carry the chemistry its mechanism demands (section 6.3), the aptamer's fold must be preserved, and the linker must be long enough that the two moieties do not interfere yet not so large as to accelerate excretion.

The largest unresolved item is endosomal escape. A cytotoxin payload can cross the membrane after release in the lysosome, but an oligonucleotide payload cannot cross by itself, so escape efficiency remains the limiting factor. Design states that bottleneck in the deliverables and marks designs that co-deploy escape-assisting elements as exploratory.

16. Modality XI — de novo Aptamer Scaffolds (Structure-Based Ligand Candidate Generation)

16.1 What is produced

A diverse pool of well-folding nucleic acid candidates ranked by computable structural fitness, or the evaluation and improvement of sequences already in hand. The output is not a therapeutic molecule but an input to an experiment — specifically, the starting library for an in-vitro selection campaign or the candidate list for a docking campaign.

This modality exists as a direct consequence of the policy in the previous chapter. If conjugate design depends only on validated aptamers, then nothing can be done for a target that has none. Filling that gap is this modality's role, and what it does is not 'find sequences that look like they will bind' but 'generate a diverse pool of sequences that fold well'.

16.2 Why structure only, and no affinity

As stated in 15.2.2, no reliable computational prediction of binding affinity to an arbitrary target exists. What, then, can computation do — the answer is the necessary condition for binding. A sequence that does not fold well cannot present a binding surface and therefore cannot bind. Folding well does not imply binding, but a pool composed of well-folding candidates has a higher success rate in selection experiments than one that is not.

That is this modality's honest position: computation does not replace experiment, it improves the experiment's starting point. Relative to beginning a selection campaign with a random library, beginning with a structurally pre-selected library reduces the number of selection rounds and the sequencing depth required.

16.3 Evaluation axes

Axis

What it measures

Why it matters

Minimum free energy

The free energy of the most stable secondary structure

Stability of the fold

Length-normalised minimum free energy

The value divided by length

Makes candidates of different lengths comparable

Paired-base fraction

The proportion of bases paired in the secondary structure

Degree of structuring

Stem and loop content

The number and size of defined stem and loop elements

Binding surfaces are usually loops, so their presence and size matter

G-quadruplex propensity

Guanosine run patterns and their arrangement

A binding motif for some aptamers, and conversely a cause of non-specific binding

GC balance

Base composition

Too stable and the structure cannot switch; too unstable and it does not fold

Structural ensemble diversity

Whether several metastable structures coexist

Candidates locked into a single structure fare better in selection

A weighted sum of these axes forms a structural fitness index used to rank candidates. The purpose, however, is not to pick the single highest-ranked sequence but to obtain a diverse set of high-ranking candidates, because selection experiments must begin from diversity. The output is therefore a pool distributed across structural classes rather than a single optimum.

16.4 Three operating modes

Mode

Input

Processing

Output

De novo generation

Length range, GC constraints, target structural class, optional fixed primer flanks

Generate candidates satisfying the constraints, fold and score each

A ranked candidate pool distributed across structural classes

Evaluation

Candidate sequences already in hand (one or many)

Fold each and score on the same axes

Per-candidate scores, structure summaries and relative ranking

Optimisation

One seed sequence

Deterministic search over point mutations, preserving length

Variants with improved structural fitness and the improvement path

The fixed primer flank option is the interface to experimental design. An in-vitro selection library carries fixed amplification sequences at both ends, and those sequences affect the folding of the variable region. Folding evaluated with the flanks included reflects the structure of the actual library molecule.

16.5 Chemistry and delivery requirements

At this stage chemical modification is optional. Fixing chemistry during candidate exploration shrinks the search space, and selection experiments are in any case usually run with unmodified or restricted-modification libraries. The customary order is to apply stabilising chemistry after the sequence has been fixed by selection, at which point the Modality X chemistry rules apply.

Delivery is not applicable. The output of this stage is a laboratory input, not something administered in vivo.

16.6 Design deliverables

16.7 Advanced capabilities — pool diversity and the experimental interface

16.7.1 A distributed pool, not a single optimum

The output objective of this modality differs from every other chapter. Elsewhere the goal is a small number of top candidates; here it is a structurally diverse candidate set. Selection experiments must start from diversity, and a pool of structurally similar candidates narrows the selection space so much that nothing may emerge.

After ranking, therefore, a distributed draw across structural classes is performed. Candidates are clustered by stem-loop count, loop size distribution, presence of branching and G-quadruplex propensity, and the top candidates from each cluster form the pool. A plain top-N draw lets one structural class dominate.

16.7.2 Handling fixed primer flanks

An in-vitro selection library carries fixed amplification sequences at both ends of the variable region. Those flanks affect the folding of the variable region, so a structure evaluated without them is not the structure of the actual library molecule. Where the user supplies the flank sequences, design evaluates folding on the complete molecule including them.

There is a useful side effect: candidates in which the flanks pair with the variable region to form unwanted structures are filtered out. Such candidates amplify inefficiently or fold inconsistently and obstruct the experiment.

16.7.3 Practical use of the three modes

Mode

When to use it

Where the output goes next

De novo generation

A target with no validated aptamer; starting a new selection campaign

Synthesis order → in-vitro selection library

Evaluation

Checking the structural plausibility of sequences from selection, or reviewing literature sequences

Candidate narrowing → binding experiments

Optimisation

Improving the fold stability of a promising sequence from selection

Synthesis of a few variants → comparative binding experiments

Optimisation mode searches point mutations while preserving length, so the user can fix positions that may participate in binding. Where structural information or experimental scanning results exist, fixing those positions and searching only the rest is the safe approach.

16.7.4 Determinism and regeneration

Randomness enters the generation process but the seed is fixed, so the same constraints and the same seed regenerate the same pool. This carries practical value for reproducing selection experiments and for regulatory response: which library was used is fully specified by the seed and the constraints, so the library can be regenerated without being archived.

16.8 What this modality cannot solve

17. The Biology of Target Site Accessibility

Accessibility is the axis most often underweighted in design. Complementarity does not guarantee binding. RNA inside a cell is not a naked linear chain but a folded, protein-coated, dynamically traversed structure with ribosomes and helicases moving through it. This chapter covers what that real environment demands of design.

The importance is quantitative. Two fully complementary candidates commonly differ in measured potency by an order of magnitude, and much of that difference is explained by whether the target site is buried in secondary structure. Ranking on sequence rules alone misses it, and the result is that candidates which are 'perfect on paper but do not work' rise to the top.

17.1 RNA is folded — and not into one structure

An mRNA pairs with itself to form secondary structure of hairpins, internal loops, bulges and multibranch junctions. Crucially that structure is not unique. A given sequence interconverts among several structures of similar free energy, and inside the cell it exists as a Boltzmann distribution over them. The precise form of the question 'is this site open' is therefore 'with what probability is this site open'.

Design handles that probability two ways. The first is local-window accessibility: the probability that the target site is unpaired within a bounded window containing it. The window is bounded because in a real cell there is limited time and opportunity for long-range base pairing to form, and full-length folding tends to predict unrealistic long-range structures. The second is interaction free energy: the total energy change when the oligonucleotide binds its target.

ΔG_bind = ΔG_duplex − ΔG_target_unfold − ΔG_oligo_unfold

  ΔG_duplex        : energy of forming the oligo–target duplex (negative, favourable)
  ΔG_target_unfold : cost of opening the target's existing structure (positive, unfavourable)
  ΔG_oligo_unfold  : cost of opening the oligonucleotide's own structure (positive, unfavourable)

→ With identical complementarity, strong target structure makes ΔG_bind worse.
→ A hairpin in the oligonucleotide itself is also a cost.

The design implication of this decomposition is clear: there are two ways to strengthen binding — make the duplex itself stronger (raise affinity chemically), or choose a site where less target structure has to be opened. The latter is far cheaper and carries no side effects, so accessibility scoring should precede chemistry decisions.

17.2 Sites covered by protein

An mRNA in a cell is never naked. From the moment of transcription, numerous RNA-binding proteins associate to form a ribonucleoprotein complex, and that complex governs the transcript's stability, localisation and translation. A site occupied by protein is physically closed to an oligonucleotide.

Crosslinking and immunoprecipitation experiments reveal where proteins actually bind at nucleotide resolution. Design uses that data as a mask to avoid occupied sites. The mask operates as a weighted penalty rather than an absolute exclusion, because binding profiles depend on cell type and condition: a peak observed in one dataset cannot be assumed present in the customer's experimental system. Penalising and recording the basis is more accurate than absolute exclusion.

Chemical modification is a factor at the same level. A given modification changes the base-pairing capacity and protein binding of its position, so regions dense in modification are more likely to have predicted and actual structure diverge. That is why a modification atlas is used as a mask.

17.3 The dynamic environment created by translation and transcription

Ribosomes traverse coding regions, and they carry potent helicase activity that continuously unwinds secondary structure along their path. This cuts both ways: coding-region structure may be more open than prediction suggests, and a bound oligonucleotide may be displaced by a passing ribosome.

The practical consequence is that regions differ in character. The 3' untranslated region is not traversed by ribosomes, so its structure is relatively stable and predictions hold better; it is also the principal arena for seed-mediated repression. The 5' untranslated region is occupied by the translation initiation complex, limiting access. The coding region varies in openness with ribosome density, and structure is unwound more often in highly expressed genes.

Different factors operate in the nucleus. Splicing occurs co-transcriptionally, so the window during which an intron exists is limited, and the exon junction complex occupies a region upstream of each junction. The constraint that intronic designs must act within the short window before removal arises here.

17.4 How structure acts differently in each modality

Modality

How accessibility operates

Design response

Catalytic silencing (RISC)

Structure slows the loaded complex's search, although the complex itself can open some structure

Assess local accessibility probability together with interaction energy

Catalytic silencing (RNase H1)

A heteroduplex must form, so strong structure prevents binding outright; the catalytic complex's structure-opening capacity is limited

Accessibility requirements are stricter; wing chemistry compensates on affinity

Splice switching

The target is fixed as a regulatory element, so there is almost no freedom in site choice; chemistry must open that site's structure

Instead of changing site, secure binding strength chemically. ΔG_bind is the key metric

Base editing

A guide–target duplex must form, so target structure interferes; the duplex formed must also present the geometry the enzyme recognises

Search a narrow window satisfying both binding and editing geometry

Noncoding RNA silencing

Targets frequently function through structure, so structure is especially strong; circular RNAs have a different topology and linear assumptions fail

Compute accessibility on the correct topology; single-stranded mechanisms favour strongly structured targets

miRNA mimic

The target site is one already used by a miRNA, so accessibility is in effect pre-validated

Use accessibility as one evidence axis in target prediction

17.5 Limits of accessibility assessment

18. The Biology of Off-Targets — What the Real Risk Is

'Off-target' is not one phenomenon but a bundle of phenomena with different mechanisms. Different mechanisms require different prediction methods, different mitigation strategies and different experimental confirmation. Collapsing them into a single off-target score produces a number that reflects none of the risks accurately. This chapter separates them.

18.1 Classification by mechanism

Mechanism

What happens

In which modalities

Predictability

Seed-mediated repression

Guide positions 2–8 are complementary to another transcript's 3' untranslated region, causing miRNA-like repression

Double- and single-stranded RISC mechanisms, miRNA mimics

High — candidates enumerable from the seed sequence

Partially complementary cleavage

Binding short of full complementarity but sufficient for cleavage

RISC mechanisms

Moderate — central pairing strength governs 3' mismatch tolerance

Heteroduplex formation

A heteroduplex forms at a partially complementary site and the host nuclease cleaves

Gapmers

Moderate — depends on melting temperature and gap position

Passenger strand loading

The unintended strand loads and represses an entirely different target set

All double-stranded modalities

High — predicted from terminal free-energy asymmetry

Protein-binding class effects

The backbone binds plasma and cellular proteins, causing complement activation, coagulation effects and hepatic accumulation

Every modality using phosphorothioate

Moderate — depends on fraction and total load, with little sequence specificity

Innate immune activation

Sequence motifs stimulate endosomal receptors

All modalities (an intended effect in the immune-modulation modality)

Moderate — known motifs are predictable

Bystander editing

Adenosines other than the intended one are edited

Base editing

High — enumerable within the duplex footprint

Global readthrough

Normal stop codons are read through, extending C-termini

Readthrough

High — enumerable transcriptome-wide from codon and context

Cross-miRNA inhibition

Other miRNAs sharing the seed are co-sequestered

Anti-miRs

High — enumerable from seed-sharing relationships

Mechanisms marked as highly predictable are those for which candidates can be enumerated; that is not the same as predicting the degree of repression accurately. Which of the enumerated candidates is actually repressed depends on expression level, site accessibility and competition, and requires experimental confirmation. Enumeration nonetheless has value, because the enumerated list defines what the experiment should look for.

18.2 Seed-mediated repression — the largest off-target class

In double-stranded silencing modalities most off-target effect is seed-mediated. A single seed 7-mer occurs in the 3' untranslated regions of hundreds of human transcripts, so a candidate with no full-length alignment hits can still carry a broad burden. This is what makes off-target assessment a question of mechanism rather than of alignment.

Repression strength is not uniform. It varies with seed type — 8mer, 7mer with 3' supplementary pairing, 7mer with an adenosine at position 1, 6mer — and with local context and position within the untranslated region. Design emits a burden index weighted by seed type, and quantifies the seed burden against essential genes separately, since repression of essential genes is the most likely to produce a phenotype.

18.2.1 Mitigation strategies

18.3 Central pairing strength and the specificity paradox

The material from Chapter 3 is restated here from a specificity standpoint. The cleavage rate on a fully complementary target is not determined by central pairing strength; the source states explicitly that formation of a continuous helix does not limit the cleavage rate of fully complementary targets. What central pairing strength determines is tolerance to 3' mismatches.

Here the direction inverts. A guide with strong central pairing also cleaves partially complementary targets whose 3' end is mismatched. Central pairing strength is therefore an index of off-target cleavage capability rather than of on-target potency. The conventional advice to raise central GC content runs precisely backwards from a specificity standpoint.

The correct form of the axis is a conditional risk term: candidates with strong central pairing are flagged toward higher off-target cleavage risk, and the term is placed on the specificity axis rather than the on-target score. Structural work points the same way — expansion of the central major groove is required to position the scissile phosphate, and GC-rich duplexes resist distortion, so structure does not support a central GC bonus either.

18.4 Counting off-targets per gene or per site

When the same site exists in several isoforms of one gene, alignment reports several hits, but biologically there is one effect on one gene. Counting hits directly over-counts genes with many isoforms, and candidates overlapping such genes are unfairly disadvantaged.

Off-target tallies must therefore be presented at two levels: a per-site hit list (where does it bind) and a per-gene summary (how many genes are affected). Risk assessment and candidate comparison need the latter; selecting experimental confirmation targets needs the former.

The same logic applies on the on-target side. Hits against other isoforms of the target gene are coverage rather than off-targets, and counting them as off-targets silently penalises a perfectly good candidate. That distinction is why the first of the four off-target buckets exists.

18.5 How reference data affect off-target counts

Off-target counts are a function of the reference transcript set as much as of the algorithm. The more isoforms an annotation includes the higher the hit count, and including noncoding transcripts raises it further. Comparing off-target counts produced against different reference sets therefore does not hold.

Two practical rules follow. Compare candidates only among values produced with the same reference set and the same settings. And trust relative ranking and mechanistic classification more than absolute counts: 'seed burden against essential genes is in the top decile' is far more stable information than 'three off-targets'.

The design layer handles this two ways — recording the reference release in the run ledger so later comparison remains possible, and presenting per-gene tallies by default to reduce distortion from isoform counts.

18.6 Connection to experimental confirmation

The purpose of computational prediction is not to replace experiment but to define it. In off-target assessment that connection is especially direct.

Predicted deliverable

Corresponding experiment

What it confirms

List of top seed-burden genes

Expression measurement across a target gene panel

Whether predicted seed-mediated repression actually occurs

Essential-gene seed burden index

Cell viability assessment

Whether toxicity appears at the phenotype level

List of partially complementary cleavage candidates

Detection of cleavage products at those sites

Whether predicted cleavage actually occurs

Passenger loading prediction

Strand-specific loading measurement

Whether the intended strand is the one loaded

Bystander editing list

Sequence analysis of the target region

The real rate of unintended editing

Global readthrough burden list

Protein size confirmation for high-risk genes

Whether C-terminal extension actually occurs

Immune-stimulatory motifs

Cytokine measurement

The real level of innate immune response

When these correspondences are explicit, computational deliverables convert directly into an experimental plan. A single number such as 'off-target score 0.3', by contrast, does not tell anyone which experiment to run. That is the substantive reason for presenting deliverables separated by mechanism.

18.7 Limits of off-target assessment

19. Durability — What Determines the Dosing Interval

In the clinical competitiveness of an oligonucleotide therapeutic, the dosing interval matters as much as potency. A drug given twice a year and one given every two weeks are different products even at equal efficacy, differing in adherence, administration infrastructure and cost structure. This chapter decomposes where durability comes from.

One fact should be stated first: the duration of action of an oligonucleotide therapeutic is largely unrelated to its plasma half-life. It disappears from plasma within hours but persists in tissue for weeks to months, and it is the latter that determines the pharmacodynamic effect. Applying conventional pharmacokinetic metrics directly therefore grossly underestimates durability.

19.1 Four contributors to durability

Contributor

Mechanism

Controllable by design?

Tissue depot formation

Endocytosed molecules accumulate in endosomal and lysosomal compartments and are slowly released to the cytosol, acting as a depot

Partly — delivery mode and chemistry affect the amount accumulated and the release rate

Intracellular chemical stability

A modified oligonucleotide resists intracellular nucleases and remains intact longer

Yes — a direct object of chemical design

Catalysis

One molecule processes many targets, so effect persists at low residual concentration

Partly — a consequence of mechanism choice

Complex stability

The loaded protein–nucleic acid complex is stable and functions for longer

Partly — loading efficiency and terminal chemistry contribute

The first is understood to contribute most. Most of what is endocytosed does not act immediately but resides in intracellular compartments, releasing slowly and acting continuously. The size of that depot is proportional to the administered dose and release is slow, so the effect of a single administration extends over months.

The design implication is counterintuitive. Extending durability requires not only a more stable molecule but one that accumulates better, and accumulation depends heavily on delivery mode. Receptor-mediated uptake concentrates accumulation in a specific cell type and is therefore favourable for durability — part of the reason receptor-conjugated products achieve long dosing intervals.

19.2 Durability characteristics by modality

Modality

Principal source of durability

Typical character

Constraint

Catalytic silencing (RISC)

Depot plus catalysis plus complex stability

The longest durability in this technology family; multi-month intervals reported for receptor-conjugated products

Depot formation is tissue-dependent and may differ outside the liver

Catalytic silencing (RNase H1)

Depot plus catalysis

Intermediate — intervals of weeks are typical

Accumulation is comparatively dispersed on carrier-free dosing

Splice switching

Depot plus chemical stability

Varies greatly with target tissue; long intervals achieved with local central nervous system dosing

No catalysis, so it is stoichiometric and highly depot-dependent

miRNA antagonism

Depot plus binding stability

Very stable binding that is effectively irreversible extends duration

Stoichiometric

Base editing

Depot

Edited transcripts disappear by natural turnover, giving a double decay

The unmodified central window limits intracellular stability

Readthrough

Persistence of the expression cassette, or the stability of a synthetic molecule

With cassette delivery, duration follows expression persistence

Comparatively short when synthetic tRNA is dosed directly

Immune modulation

Not applicable — a transient stimulus is the objective

Short action is preferable

Repeated stimulation risks tolerance or chronic inflammation

The 'double decay' of base editing deserves explanation. Editing occurs at the transcript level, so edited transcripts disappear with natural mRNA turnover while newly transcribed molecules are unedited. Duration is therefore the product of two time constants — persistence of the guide and transcript turnover rate — and targets with fast turnover require more frequent dosing. That same property is also the substance of the modality's reversibility as a safety feature.

19.3 How chemistry contributes to durability

Chemical element

Contribution to durability

Conflicting axis

2'-modification coverage

Raises intracellular nuclease resistance so intact molecules persist longer

Excess lowers activity, depending on mechanism

Phosphorothioate fraction

Raises both stability and tissue accumulation

Class toxicity burden; potency loss on encapsulated routes

Terminal protection

Blocks exonucleolytic erosion

Almost none — a high-value, low-cost measure

Receptor ligand conjugation

Concentrates accumulation in a specific cell type, enlarging the depot

Restricted to particular target tissues

Lipophilic conjugation

Increases membrane interaction, prolonging tissue residence

May increase non-specific distribution

5'-terminal stabilisation

Contributes to the stability of the loaded complex

Almost none

The design layer combines these into a durability ranking and an expected dosing-interval band. That output is a prediction, and the real interval varies with species, tissue and dose and must be confirmed by pharmacokinetic and pharmacodynamic study. Using it for relative comparison among candidates is its correct use.

19.4 The relationship between durability and safety

Long durability is not purely an advantage. If an adverse effect appears, stopping administration does not stop the effect, so it is hard to reverse. That matters particularly where the target has an essential function, or where long-term safety data are lacking for a novel target.

Durability is therefore not a property to be maximised but one to be matched to target and indication. In maintenance therapy of chronic disease a long interval carries great value; in early clinical development of an unproven target, shorter action provides safety margin. This is why the design layer presents several chemistry alternatives together with a durability ranking — different choices are rational at different development stages.

Where reversibility is paramount the modality choice itself changes. Unlike approaches that permanently alter the genome, RNA-level intervention is reversible in principle, and among those, stoichiometrically acting modalities recover faster than catalytically acting ones.

19.5 Limits of durability prediction

20. Combination Strategies Across Modalities

Real development programmes frequently do not end with one modality. Two modalities may be designed in parallel against the same target and compared experimentally; different targets may be addressed with different modalities simultaneously; and the output of one modality may become the payload of another. This chapter sets out the types of combination and the conditions under which each holds.

20.1 Type 1 — same target, two competing modalities

It is common for both catalytic silencing mechanisms to be viable for lowering the same gene. Which is better is not determined in advance: target site accessibility, subcellular localisation, delivery route and tissue all shift the balance. Parallel design followed by experimental comparison is the rational approach, and there are cases in which products using both mechanisms against the same target have each been approved.

For that comparison to be fair, both designs must share one coordinate system and one off-target definition. Designed with different tools, '3 off-targets' becomes three things under two definitions and the comparison is meaningless. Designed on a common foundation, the metrics in the two reports carry the same meaning and the comparison becomes a basis for decision.

Comparison axis

Catalytic silencing (RISC)

Catalytic silencing (RNase H1)

Deciding criterion

Site of action

Principally cytoplasmic

Both nucleus and cytoplasm

Nuclear-retained targets favour the latter

Delivery burden

Vehicle mandatory

Carrier-free route available

Without capacity for vehicle development, the latter

Required dose

Low (catalysis plus high potency)

Comparatively high

Considering dose-related toxicity margin, the former

Principal toxicity risk

Seed-mediated off-targets

Hepatocyte toxicity

Depends on target tissue and concomitant medication

Freedom in site choice

Form largely fixed by geometry

Length, gap and wings adjustable

Difficult target structure favours the latter

Durability

The longest in this technology family

Intermediate

Chronic maintenance therapy favours the former

20.2 Type 2 — simultaneous silencing of multiple targets

Lowering a single gene sometimes produces no phenotype, because functionally redundant paralogs compensate or because the disease arises from the sum of several pathways. Two options then exist: find one molecule that silences all the targets, or combine molecules each targeting one.

The former means finding a conserved region within the family so that one molecule silences several members; this is what a co-silencing instruction expresses in family selectivity adjudication. It requires a sufficiently conserved site to exist, and where none does, the latter route follows. The latter combines several molecules, so doses and toxicities accumulate and the safety burden grows.

An important design point is that the off-target burden of a combination is not a simple sum. Two molecules with different seeds double the seed-mediated burden, and sharing chemistry means backbone-related class effects rise with total exposure. In combination design, therefore, the objective is not optimising each candidate but minimising the burden of the combination as a whole.

20.3 Type 3 — combining modalities that act in opposite directions

Where a disease mechanism is 'too much of one thing and too little of another', combining silencing with activation is logically coherent — lowering a pathogenic factor while raising a protective one.

In practice the combination is difficult. Activation carries greater predictive uncertainty, and for two molecules to reach the same cell they must travel in the same vehicle, which is awkward when their form and chemistry differ. The first option considered is therefore the detour route: silencing the repressive axis of the gene to be raised — a natural antisense transcript or an inhibitory regulator — achieves both objectives with silencing modalities, at which point both molecules share form and chemistry and co-formulation becomes straightforward.

20.4 Type 4 — combining sequence correction with expression control

In premature-stop-codon disease, two different interventions can be complementary. Even if a readthrough molecule allows the stop to be passed, there is nothing to read if the transcript has already been degraded by nonsense-mediated decay. Combining decay inhibition with readthrough induction is therefore under discussion.

Similarly, splice switching and silencing can be combined: silencing a pathogenic isoform while increasing production of the normal isoform by splice manipulation. Here both molecules are single-stranded with similar chemistry, so the technical burden of combination is comparatively low.

The relationship between base editing and other modalities is both competitive and complementary. For the same pathogenic variant, readthrough and editing may each be applicable, and which covers more patients can only be established by decomposing that gene's variant spectrum by type. This is why the target adjudication layer emits variant-type proportions.

20.5 Type 5 — one modality as the payload of another

A targeted conjugate produces no therapeutic effect itself but carries something else. When the payload is a silencing molecule, two modalities are fused into one construct, and because the conjugate substitutes for a delivery vehicle, silencing outside the liver becomes possible in principle.

Designing this combination must satisfy both modalities' requirements simultaneously. The silencing payload must carry the chemistry its mechanism demands; the conjugate must preserve the aptamer's fold; and the linker must be long enough that the two moieties do not interfere yet not so large as to accelerate renal excretion. Because the requirements conflict, the construct must be evaluated in its combined state rather than optimised independently and then joined.

The remaining bottleneck is endosomal escape. Even where the conjugate enters the cell, the silencing payload does nothing if it does not reach the cytosol. This is the main reason the combination remains of limited practical use, and from a design standpoint it means elements that assist escape must be considered alongside.

20.6 Common principles in combination design

20.7 Deliverables of combination design

21. The Design Specification Seen from Synthesis and Quality

The ultimate use of a design deliverable is a synthesis order. It therefore matters whether the specification is complete from a synthesis standpoint, and what quality characteristics the material it produces will have. This chapter covers how sequence and chemistry affect synthesis and quality.

21.1 What it means for a specification to be complete

A synthesis order requires more than a sequence string. All the items below must be fixed for the specification to hold; where any is missing, the synthesis provider either interprets it or asks.

Item

What it specifies

What happens when it is blank

Strand composition

Single- or double-stranded, and for a duplex each strand individually

A duplex is mistaken for a single strand or vice versa

Base sequence

The complete sequence in 5' to 3' orientation

A misread orientation produces an entirely different molecule

Sugar modification map

The sugar identity at every position (natural ribose, 2'-deoxy, 2'-O-methyl, 2'-fluoro and so on)

A summary such as '60% 2'-O-methyl' cannot be synthesised

Backbone map

The identity of every linkage (phosphodiester or phosphorothioate)

Arbitrarily interpreted as fully phosphorothioate, changing the properties

Terminal handling

The chemistry at the 5' and 3' termini (phosphate, phosphate mimic, inverted residue, none)

The presence of a terminal phosphate governs activity itself in some modalities

Conjugate

Attachment position, ligand class, linker chemistry

A different attachment point changes targeting and folding

Overhang convention

The terminal form of a duplex and the overhang bases

Artificial and target-derived overhangs are confused

Stereocontrol

Whether phosphorothioate stereochemistry is controlled

The definition of the substance and its quality items both change

Purification grade

The required purity level

Research-use and preclinical requirements differ

This is why deliverables express chemistry as per-position strings. A summary statistic is not a specification, and it cannot be used for patent review either, because most chemistry claims are constructed as 'a specific modification at a specific position'.

21.2 How sequence and chemistry affect synthesis difficulty

Factor

Effect on synthesis

Response at the design stage

Length

Cumulative step yield lowers the proportion of full-length product

Prefer shorter designs within what the mechanism allows

Guanosine runs

Higher-order structure formation lowers coupling efficiency and complicates purification

Apply a run-length cap at enumeration

Strong self-structure

Aggregation during synthesis and purification

Use self-secondary-structure as an evaluation axis

Phosphorothioate fraction

Adds a sulfurisation step and produces a diastereomer mixture

Minimise within the route-specific productive window

Bicyclic modifications

Higher monomer cost and demanding coupling conditions

Optimise placement so they are used only where needed

Conjugates

Add a conjugation reaction and subsequent purification steps

Restrict attachment to the termini

Double-stranded format

Two strands must each be synthesised and purified, then annealed and verified

Recognise in advance that quality items double

These factors affect not only cost and timeline but the quality of the material. Length and phosphorothioate fraction in particular determine the impurity profile directly.

21.3 Expected impurity classes

Impurities in solid-phase synthesis fall into defined classes and are to some extent predictable from sequence and chemistry. Surfacing them at design time brings forward purification strategy and analytical method development.

Impurity class

Mechanism

Relationship to sequence and chemistry

Deletion sequences (n−1, n−2 …)

Chains missing a unit through coupling failure

Increase with length and with the number of difficult coupling positions

Addition sequences (n+1)

Double coupling

Increase with particular monomers and conditions

Incomplete deprotection

Chains retaining a protecting group

Increase with bulky modifications

Phosphodiester incorporation

Incomplete sulfurisation leaves an intended phosphorothioate as a phosphodiester

Proportional to the number of phosphorothioate linkages

Diastereomers

Stereochemistry at the phosphorothioate phosphorus

2ⁿ for n linkages — intrinsically a mixture unless stereochemistry is controlled

Depurination products

Loss of a purine base under acidic conditions

Depends on purine content and processing conditions

Oxidation and dimers

Oxidation during storage, thiol-related dimerisation

Observed with particular conjugates and terminal chemistries

The diastereomer entry warrants separate explanation. Without stereocontrol, one 'substance' is in reality a mixture of thousands to millions of isomers that behave differently in nuclease resistance and target binding. This is less an impurity question than a question of what the substance is, and from a regulatory standpoint the issue is demonstrating that the mixture is reproduced consistently. Adopting a stereocontrolled synthesis clarifies the definition of the substance at a substantial increase in cost and complexity. This is why the deliverables state explicitly whether stereocontrol is assumed.

21.4 Quality items specific to duplexes

The last item is a recurring practical problem. If the overhang convention is not stated in the specification, a conventional notation may be applied, and the molecule evaluated at design time then differs from the molecule synthesised. This is why the acceptance test is full-length reverse complementarity rather than length, and why it is verified automatically at export.

The impurity information emitted here is a design-time prediction; setting actual specifications and validating analytical methods belong to the manufacturing stage. Its purpose is to prepare purification strategy and analytical development in advance, and to avoid at design time chemical choices that would become quality problems.

21.6 The handover between design and manufacturing

Information is lost between design and manufacturing at predictable points. First, summarisation of chemistry notation — per-position information compressed into statistics cannot be restored. Second, loss of design rationale — if why a modification was placed at a position is not conveyed, then when manufacturing changes it for synthetic convenience the effect on activity is unknown. Third, loss of alternatives — if the chemistry alternatives considered at design time are not conveyed, difficulty in synthesis forces a return to design.

Deliverables carry the per-position map, the selection rationale and the alternative comparison together to prevent all three. Because the run ledger records the parameters and reference data of the moment, a later return to design from manufacturing can be re-run under identical conditions.

22. A Disease-Area Perspective

Modality choice is not determined by the target's molecular properties alone. Which tissue the target sits in, what route reaches that tissue, and what development precedent has accumulated in that disease area all bear on it. This chapter sets out where this technology family stands in each major disease area.

The reason for organising by area is practical. 'I want to lower this gene' becomes an entirely different programme depending on whether the gene sits in the liver or the brain. The former has a mature path; the latter makes the route of administration itself a development item.

22.1 Liver — the most mature area

The liver is where approvals in this technology family are overwhelmingly concentrated, for biological reasons. Hepatocytes express the asialoglycoprotein receptor at very high surface density and it recycles rapidly, so a molecule bearing a triantennary galactosamine ligand concentrates efficiently in the liver after subcutaneous dosing. At the same time the liver is where particles naturally accumulate from systemic circulation, so the lipid particle route also works without separate targeting.

Disease type

Suitable modalities

Why

Protein excess or misfolded accumulation

Catalytic silencing (both mechanisms)

Reducing production is the direct remedy; both mechanisms have approval precedent

Metabolic enzyme deficiency

Transcriptional activation, or silencing of the repressive axis

The direction is upward; the silencing detour is the first option to examine

Lipid metabolism disorders

Catalytic silencing

Lowering a regulator is an established approach

Hepatic miRNA axes

Anti-miR

Both carrier-free uptake and receptor conjugation work

Nonsense-variant metabolic disease

Readthrough or base editing

Delivery is comparatively easy, favouring validation of newer modalities

The strategic implication is clear. Validating a new modality or new chemistry against a hepatic target controls for the delivery variable and allows the modality itself to be assessed. Validating a new modality outside the liver leaves failure indistinguishable between a modality problem and a delivery problem.

22.2 Central nervous system — the route is the solution

The brain and spinal cord are separated from the systemic circulation by the blood–brain barrier, so intravenously dosed oligonucleotides effectively do not reach them. What solved this was not a vehicle but a route: intrathecal administration distributes through the cerebrospinal fluid, and several products have been approved on that route.

Single-stranded modalities have an advantage here for specific reasons. Nuclease activity in cerebrospinal fluid is comparatively low, and phosphorothioate single strands distribute well through neural tissue and persist. Because dosing is local, systemic exposure is low and constraints from backbone class toxicity relax. Double strands, by contrast, require a vehicle, and intrathecal particle dosing brings additional considerations.

Disease type

Suitable modalities

Considerations

Toxic gain-of-function variants in motor neuron disease

Catalytic silencing (RNase H1), allele-selective design

Selectively lowering only the mutant allele provides safety margin

Splicing-defective neuromuscular disease

Splice switching

The flagship success of this area; local dosing manages the burden of repeat administration

Repeat expansion disease

Catalytic silencing or splice manipulation

The repeat structure creates difficulty in both accessibility and specificity

Neurodegenerative protein accumulation

Catalytic silencing

Where the target has a normal function, partial reduction becomes the objective

Nonsense-variant neurological disease

Readthrough or base editing

The delivery burden is large and vehicle development runs in parallel

The practical constraint is the invasiveness of the procedure and the resulting limit on dosing frequency. Durability therefore carries particular value in this area, and chemical design prioritises intracellular stability and tissue residence.

22.3 Eye — the advantage of local administration

The eye is a small, isolated compartment, so local administration achieves high local concentration while minimising systemic exposure. That several early approvals in this technology family are ocular is not a coincidence, and both aptamer approvals fall in this area.

Three characteristics define the area for design. Systemic toxicity constraints relax, widening the range of chemistry available. Administration is invasive, so durability matters. And diffusion behaviour differs by target tissue — retina, cornea — so vitreal diffusion and residence time become design considerations.

22.4 Muscle and heart — where delivery is the bottleneck

Muscle constitutes a large fraction of body mass, so obtaining adequate tissue concentration by systemic dosing requires very large amounts. That is why approved products in this area use high doses, and equally why unmet need remains large.

Improvement is being attempted in two directions: conjugating cell-penetrating peptides, and attaching ligands that target muscle cell surface receptors. From a design standpoint both require assessing the effect of conjugation on folding, binding and distribution, with attachment point and linker as new design variables.

The heart has delivery problems similar to muscle with a narrower safety margin. Unexpected effects on cardiac muscle are immediate and hard to reverse, so off-target and essentiality assessment are demanded particularly rigorously in this area.

22.5 Lung and airway — the possibility of the inhaled route

Inhalation reaches the target tissue directly while lowering systemic exposure. It fits logically for diseases targeting airway epithelium — cystic fibrosis, chronic airway disease, respiratory infection.

Three design challenges follow. Shear during aerosolisation can destroy particles, so formulation rigidity matters. Airway mucus acts as a physical barrier, making mucus penetration a design variable. And much of what is inhaled is removed by mucociliary clearance, so residence time governs efficacy.

In infectious disease response the route has strategic value. Targeting host factors required by respiratory viruses works independently of viral sequence variation, so resistance is hard to develop and effect against several viruses is possible. The counterweight is the safety burden of suppressing a host gene, which makes essentiality assessment of the target decisive.

22.6 Tumours — targeted delivery is the crux

Tumours are among the hardest areas for this technology family. Tumour tissue has irregular vasculature and high interstitial pressure, limiting molecular penetration, and tumour cells are heterogeneous, so a single target rarely addresses the whole. At the same time selectivity against normal tissue is difficult to secure.

Approach

Modality

Rationale and constraint

Silencing tumour-dependency genes

Catalytic silencing

Essentiality must be tumour-specific; the cell-dependency axis of target adjudication is central

Targeted delivery

Aptamer conjugation (cytotoxin or oligonucleotide payload)

Works where a tumour surface target exists; constrained by the availability of a validated aptamer

Local immune stimulation

Innate immune receptor agonists

Intratumoral dosing to reprogramme the microenvironment; minimising systemic exposure is the crux

Restoring tumour suppressors

Transcriptional activation or silencing of the repressive axis

The direction is upward; candidate genes are numerous, so combination strategies are discussed

Rebalancing a miRNA axis

Mimicry or antagonism

Network-level intervention with a heavy safety assessment burden

This is where the target adjudication layer shows its greatest value, because tumour dependency and normal-tissue essentiality must be assessed together, synthetic-lethal relationships and adverse-event precedent must be read alongside, and redirection to a substitute target is needed when the target proves unsuitable.

22.7 Infectious disease and vaccines — conservation and host factors

In targeting a pathogen directly, the central problem is variation. Targeting a single site in a rapidly varying pathogen selects resistance variants quickly, so regions conserved across strains and subtypes must be targeted. Identifying conserved regions from a multi-strain sequence set and restricting design to them is the corresponding approach.

The alternative is targeting host factors. Lowering a host gene the pathogen requires for replication works independently of the pathogen's sequence variation. The attraction is breadth and resistance robustness; the burden is the safety cost of suppressing a host gene. Essentiality assessment is therefore decisive in target selection, and targets that are tissue-restricted in expression, or that need suppressing only during the acute phase, are preferred.

In vaccines the most mature application is as an adjuvant. Innate immune receptor agonists have approval precedent, and the type of immune response induced can be selected by structural class. This has particular value for pathogens where cell-mediated immunity is protective and for populations with weak responses.

Veterinary and zoonotic applications belong in this context. Infectious disease management in livestock and poultry is both a market in itself and connected to human pandemic risk management, with comparatively shorter regulatory paths. What matters is designing against that species' own transcriptome and corresponding genes rather than repurposing a human product.

22.8 Rare genetic disease — variant type determines the modality

In rare genetic disease, modality selection begins from the variant type rather than the gene, because patients with the same gene disorder carry different variant types and different interventions apply to each.

Variant type

Applicable modalities

What determines the patient fraction covered

Premature stop codon

Readthrough

The share of stop-gained variants in that gene

Splice site variant

Splice switching

The share of splicing-related pathogenic variants

Deep-intronic variant creating a pseudoexon

Splice switching (pseudoexon suppression)

How frequently that variant is reported

G-to-A missense

Base editing

The share of that transition type

Frameshift deletion

Splice switching (exon skipping)

The share of deletions for which skipping restores the frame

Dominant-negative variant

Allele-selective silencing

Whether a sequence difference exists to distinguish the mutant allele

Loss of function (haploinsufficiency)

Transcriptional activation or silencing of the repressive axis

Headroom for raising the normal allele

The development strategy this table implies is patient stratification. No single molecule covers every patient, so quantifying the proportion of each variant type first and developing the intervention that covers the most patients is the rational order. This is why the target adjudication layer decomposes the pathogenic variant spectrum by type — that decomposition is the basis for development sequencing.

22.9 Difficulty by area, summarised

Area

Delivery difficulty

Safety margin

Accumulated precedent

Overall

Liver

Low — two mature routes

Moderate

Very deep

Optimal for validating new modalities

Central nervous system

Moderate — solved by route

Relatively wide (low systemic exposure)

Deep

Durability matters because of procedural invasiveness

Eye

Low — local administration

Wide

Moderate

Favourable for early validation

Muscle and heart

High — the current bottleneck

Narrow for the heart

Moderate for muscle, thin for heart

Delivery improvement is the crux

Lung and airway

Moderate — the inhaled route

Moderate

Thin

Formulation and residence time are the challenges

Tumours

High — penetration and selectivity

Narrow

Thin

Where target adjudication is most valuable

Infectious disease

Varies with target

Narrow for host targets

Moderate (deep for adjuvants)

Conservation or host essentiality assessment is central

23. A Drug and Vaccine Development Perspective

The structural characteristic of oligonucleotide therapeutics in drug development is that the design space is finite and computationally compressible. Once the target gene is fixed, the design space is bounded by the subsequences of that transcript, and efficacy, selectivity, accessibility, chemistry and cross-species conservation are all properties that can be evaluated in advance. This differs fundamentally from small molecules, which must search a vast chemical space, and from antibodies, which require immunisation and selection.

The practical difference that compressibility makes is in the time and cost of lead discovery. From a fixed target, the path to a set of synthesisable candidate sequences is short and mostly computational. The burdens of the later stages — delivery, chemistry-manufacturing-and-control, toxicology — are, by contrast, no lighter than for other modalities and in places add items of their own.

23.1 Deliverables mapped to development stages

Stage

Central question

What this layer provides

What carries into the next stage

Target validation

Is intervening on this gene correct, and in which direction

Modality suitability verdicts; safety, expression, disease and network evidence; substitute target proposals

A fixed target and modality, and family selectivity instructions

Lead discovery

Which candidates should be synthesised

Candidate sequences and ranking, off-target classification, cleavage coverage, accessibility

Synthesis specifications for 10–30 candidates

Lead optimisation

How should chemistry and delivery be fixed

Per-position chemistry maps, alternatives per delivery route, durability ranking, class toxicity risk classification

Complete specifications for a small number of leads

Preclinical

In which species should it be evaluated, and what is risky

Per-species-combination comparison of surviving candidates, two-species verdict, basis for off-target experiment design

Documented rationale for toxicology species selection

Regulatory preparation

How should the filing package be assembled

Run ledger, synthesis impurity projection, reference version verification records

Machine-readable supporting material

Chemistry, manufacturing and controls

What will become a problem in formulation and process

Formulation candidates and property predictions, stability, process and container risk item assessments

The starting point for formulation development and a risk list

Species selection at the preclinical stage is regulatorily significant. Safety assessment in one rodent and one non-rodent species is customarily expected before human trials, and for oligonucleotide therapeutics target-related toxicity can only be assessed if the molecule also binds the corresponding transcript in that species. Data from a species in which it does not bind show chemistry class toxicity and say nothing about target-related risk. Presenting the surviving candidate set per species combination is how that requirement is evidenced with data.

23.2 The decision flow for modality choice

What must be decided first in a real programme is not the sequence but the modality. Below is the flow from a target's properties to a modality.

What is the therapeutic direction?
├ Down
│  ├ target is mature cytoplasmic mRNA     -> catalytic silencing (RISC).   vehicle required
│  ├ target is nuclear-retained / pre-mRNA -> catalytic silencing (RNase H1). carrier-free possible
│  ├ target is a noncoding RNA             -> class call, then one of the above
│  └ target is a miRNA itself              -> miRNA antagonism.  carrier-free possible
├ Up
│  ├ the locus is activatable              -> transcriptional activation.  vehicle required
│  ├ a repressive axis exists              -> silence that axis (detour; better prediction accuracy)
│  └ restoring a lost miRNA function       -> miRNA mimicry.  vehicle required
├ Repair
│  ├ pathogenic variant is G>A             -> ADAR-recruiting editing.  vehicle required
│  ├ pathogenic variant is a premature stop-> suppressor tRNA.  vehicle required
│  └ mechanism runs through splicing       -> splice switching.  vehicle usually unnecessary
└ Move the immune system / carry something else
   ├ innate immune stimulation or suppression -> TLR modulation.  no vehicle needed
   └ cell-type-specific delivery              -> aptamer conjugation.  it is the vehicle

The branch most often overlooked is the second under 'Up'. Activation carries greater predictive uncertainty and thinner clinical precedent, whereas silencing a repressive axis uses a mature silencing modality and therefore carries lower development risk. If the same therapeutic objective can be reached by a better-validated means, that route is preferable — and reaching that judgement requires mapping antisense transcripts around the target locus.

23.3 Applications in vaccine development

23.3.1 Oligonucleotides as adjuvants

A synthetic oligonucleotide has obtained regulatory approval as a vaccine adjuvant, and this is the application with the broadest human safety dataset in the technology family. Adjuvants are dosed in large numbers to healthy adults, so the scale of safety evidence required for approval differs from that for a therapeutic.

From a design standpoint the essential feature of an adjuvant oligonucleotide is selecting the type of immune response. Traditional aluminium salt adjuvants induce mainly antibody responses and weak cell-mediated immunity. An innate immune receptor agonist can selectively strengthen either the interferon pathway or the B cell pathway according to structural class, so the response type can be designed to suit the pathogen. The value is greatest for pathogens where cell-mediated immunity is protective and in populations, such as the elderly, whose responses are weak.

23.3.2 Antigen-expressing platforms share the delivery layer

Messenger RNA vaccines differ from the silencing and correction modalities in the molecule itself but share the delivery layer. Ionisable lipid particles, biodistribution driven by the protein corona, endosomal escape efficiency, cold-chain distribution and lyophilisation are the same formulation problems. The delivery design layer therefore applies to vaccine payloads unchanged, which is why messenger RNA, self-amplifying RNA and circular RNA cargo families appear in the formulation variant catalogue.

Self-amplifying RNA co-encodes a replicase and amplifies itself inside the cell, lowering the required dose; circular RNA has no ends and resists exonucleases, extending expression. Both change the payload's properties and therefore the particle design requirements, and the design layer applies different structural branches per cargo type.

23.3.3 Designing on conserved regions for infectious targets

Targeting a rapidly varying pathogen requires targeting regions conserved across strains and subtypes in order to achieve both breadth of efficacy and resistance suppression. The corresponding approach identifies low-variation regions from a multi-strain sequence set and restricts design to them, and the design layer performs that analysis on a user-supplied multi-strain sequence set.

A host-factor approach also exists. Lowering a host gene the virus requires for replication works independently of viral sequence variation, so resistance is difficult to develop and effect across several viruses is possible. The counterweight is the safety burden of suppressing a host gene, so essentiality assessment of the target is decisive — this is the area in which the safety axis of the target adjudication layer matters most.

23.3.4 Veterinary and zoonotic applications

Infectious disease management in livestock and poultry is both a market in its own right and connected to management of human pandemic risk: for pathogens with zoonotic transmission potential, control at the animal stage lowers human risk. Regulatory paths are shorter than for human medicines and areas of substantial unmet need exist, so species-specific design carries real value. What matters is designing against that species' transcriptome and orthologs rather than repurposing a human product, and per-species design modes support this.

23.4 Where development risk concentrates, by modality

Risk concentrates at a different stage in each modality. Knowing that distribution changes resource allocation and milestone design.

Modality

Where risk concentrates

Why

Mitigation

Catalytic silencing (RISC)

Delivery (non-hepatic targets)

The liver has a mature answer; elsewhere is the industry-wide bottleneck

Start with hepatic targets, or concentrate resource on vehicle exploration

Catalytic silencing (RNase H1)

Toxicity, especially hepatocyte toxicity

Potency and toxicity risk arise from the same properties

Evaluate the toxicity axis in parallel with potency at design time; confirm experimentally early

Splice switching

Target tissue concentration

Dosed without a carrier, so tissue delivery depends on dose

Prioritise local routes; consider cell-penetrating conjugation

Noncoding RNA silencing

Target validation

Functional regions are frequently uncharacterised

State functional mapping as prerequisite work

Transcriptional activation

Prediction accuracy

Strong chromatin-state dependence and poor transfer between cell types

Run positive controls in parallel; examine the silencing detour route

miRNA modulation

Safety assessment

Many targets move at once, leaving scope for unexpected phenotypes

Confirm target-set convergence in advance; escalate dose in steps

Base editing

Conflict between efficiency and stability

The unmodified central window constrains chemical protection

Compare window widths; strengthen carrier protection

Readthrough

Global burden

Effects on normal stop codons exist in principle

Exploit codon and context selectivity; quantify global burden and monitor high-risk genes

Immune modulation

Cross-species extrapolation

Species specificity limits translation of animal data

Assess per-species optimal motifs; run human cell-based assessment in parallel

Aptamer conjugation

Applicable scope

Restricted to targets with a validated aptamer

Generate candidate pools and run selection in parallel

23.5 What the computational stage actually saves

Stated without exaggeration: computation does not replace wet-lab work. What it does is reduce the number of candidates that must be validated, and that reduction becomes a reduction in time and cost. Validating a few dozen candidates instead of synthesising and screening several hundred cuts synthesis cost, reagents, personnel and elapsed time in proportion.

The second value is the elimination of handover losses between stages. When target adjudication, design, chemistry, delivery and regulatory preparation are dispersed across different tools and manual analysis, information is lost at each handover and reproducibility degrades. Performed continuously under one coordinate system and one provenance record, those losses do not occur, and the effect is cumulatively larger than the speed gain of any individual stage.

The third value is traceability of decisions. When why a candidate was selected, and which database versions and parameters were in force at that moment, are recorded automatically, the interpretation of later experimental results and the regulatory response both rest on that record. This is risk reduction rather than cost saving, and its value becomes visible when a problem arises at a late stage.

What computation does not do is stated equally plainly. Predictive scores rank; they do not guarantee absolute efficacy. Toxicity risk classification does not replace experimental toxicology. Delivery prediction does not replace measured biodistribution. The role of computation is to make the experiment smaller and better aimed.

24. Benchmarking Against Industry Standards

This chapter compares the capabilities and performance of this design layer against public tools and industry practice. Comparison is made at two levels: a qualitative comparison of whether a capability axis exists, and a quantitative comparison of predictive performance measured on public benchmark datasets. The scope in which quantitative comparison is possible is limited, so that boundary is stated first.

The boundary of comparison — quantitative comparison of predictive performance is confined to the area where a public benchmark exists, namely efficacy prediction for double-stranded silencing. For other modalities there is no widely adopted public benchmark dataset, so this document presents no quantitative figures there and offers only capability-axis comparison. Constructing a benchmark that does not exist in order to present one's own performance is not a comparison.

24.1 Double-stranded silencing efficacy prediction — quantitative comparison

The standard benchmark for efficacy prediction models is a large gene-silencing library dataset. Measured on the same dataset (n = 2,431), the coefficient of determination, correlation coefficient, and the area under the curve that expresses practical discrimination compare as follows.

Generation

Model / method

Correlation

AUC@0.7

First-generation rules

Positional rule-based (2004 family)

~0.10

0.32

0.63

Second-generation statistics

Linear kernel-based (2006 family)

~0.18

0.42

0.69

Third-generation deep

Transformer with RNA language-model embeddings (reported 2024)

0.34

0.58

0.78

Third-generation deep

Transformer combined with convolution (reported 2025)

0.31

0.55

0.77

Third-generation deep

Chemistry-aware multi-view with cross-attention (reported 2024)

0.40

0.60

0.80

This layer

Chemistry-aware model — best-fold checkpoint

0.44

0.64

0.82

This layer

Transformer plus convolution, retrained

0.31

0.53

0.74

This layer

In-house spline-based model

0.28

0.52

0.74

This layer

Cross-model ensemble (architectural diversity retained)

0.36

0.58

0.79

What the table shows is the direction of generational transition. Every metric improves stepwise from rule-based to deep models, and in particular the area under the curve rises from 0.63 to around 0.80. A value of 0.63 is better than chance but insufficient for practical screening, while the 0.78–0.82 range provides adequate discrimination for narrowing candidates.

That the coefficient of determination remains in the 0.3–0.44 range reflects the intrinsic nature of the task. Silencing efficiency is a biological phenotype with substantial measurement noise, and values shift with cell line, delivery method and measurement time point. Repeat measurements of the same sequence are not perfectly reproducible, which caps what any predictive model can achieve. At this dataset size 0.44 is competitive, and reports claiming substantially more warrant a check on data leakage or evaluation design.

That the production checkpoint uses a single best fold rather than a fold average appears counterintuitive but is reasonable on a small dataset. Fold averaging averages away each fold model's weaknesses but also dilutes the strongest fold's signal, and at this dataset size the latter loss dominates. That decision is at the fold level only; ensembling across different architectures is retained — fold averaging is given up while the benefit of architectural diversity is kept.

24.2 Capability-axis comparison — double-stranded silencing design

Public design tools are generally based on rules formulated in the early to mid 2000s and were designed around a single species, a single database, and manual downstream analysis. They are adequate for proposing candidates, but do not provide the precise off-target classification, cross-species verification, chemistry design and regulatory traceability that clinical entry requires.

Capability axis

Public rule-based tools

Public web tools (improved)

Pharmaceutical internal tools (inferred)

This layer

Efficacy prediction

Positional rules only

Rules plus some statistics

Undisclosed (inferred to be trained on internal data)

Rule ensemble plus deep-model ensemble across five distinct architectures

Off-target classification

Count of alignment hits

Alignment plus seed search

Full pipeline (inferred)

Four-bucket classification (on-target extension / paralog / ortholog / unrelated) plus seed context

Paralog awareness

None

None

Present (inferred)

Explicit co-silence and spare verdicts with hard filters

Cross-species support

1–2 species

3–5 species

Many (inferred)

Parallel evaluation of multiple species combination modes with a two-species preclinical verdict

Variant avoidance

None or single database

Partial

Population variant database (inferred)

Dual avoidance across population and clinical variants with multi-population allele frequencies

Cleavage coverage

None

None

Unknown

Quantified per isoform, with coding-only and all-transcript counts separated

Chemistry design

None

None

Present (internal intellectual property)

Per-position chemistry map, mechanism-specific templates, route-specific windows

Delivery design

None

None

Present (internal intellectual property)

Conjugate and formulation candidates, route-specific absorption, formulation variant catalogue

Regulatory support

None

None

Present (internal)

Run ledger, reference version verification, preclinical requirement verdict, impurity projection

Reproducibility guarantee

Not stated

Not stated

Unknown

Byte-identical output contract for the same input and seed

Automation interface

Web only

Web only

Internal systems

Web and command line, with batch processing

The essential message of this table is not a claim of performance superiority but a difference in scope. Public tools have a clearly bounded purpose — proposing candidates — and are useful within it. The divergence begins after that: interpreting off-targets, verifying across species, fixing chemistry and delivery, generating regulatory evidence. Those stages have traditionally been dispersed across manual work and separate tools.

24.3 Capability-axis comparison — single-stranded modalities

For single-stranded modalities — gapmers, splice switching, anti-miRs — there is no public design tool as widely used as those for duplexes. Public tools mainly provide thermodynamic computation (melting temperature, secondary structure, accessibility) and do not address mechanism-specific chemistry rules or toxicity prediction. Comparison is therefore about which axes are automated.

Axis

Public thermodynamic tools

This layer

Melting temperature and free energy

Provided (core function)

Provided, on published parameter sets

Target accessibility

Partially provided

Provided — ensemble accessibility and interaction energy

Mechanism-specific chemistry enforcement

None

Gap-mandatory, gap-forbidden and uniform-modification rules enforced automatically by mechanism

Hepatotoxicity risk classification

None

Multi-tier classification from sequence and chemistry combination

Quantifying the effect on splicing

None

Target-window masking differential — a quantity attributed to the candidate

Reading-frame validation

None

Automatic validation of exon-skipping designs

Junction-aware off-targets

None

Separate classification of junction-proximal hits

Carrier-free suitability diagnosis

None

Pass/fail verdict per named gate with reasons

Route-specific chemistry windows

None

Per-route backbone fraction windows implemented as constraints

24.4 Target adjudication and family selectivity — the absence of a comparator

Modality suitability adjudication and paralog segregation are usually performed manually or not at all. The individual databases — population constraint, cell dependency, tissue expression, interaction networks — are all public, but an automated path that integrates them into one verdict and converts that verdict into design instructions is not standardised.

In this area the distinguishing feature of this layer is not algorithmic sophistication but explicitness of the verdict. The rules and coefficients are documented, the evidence used is presented item by item, and the limits of the verdict are stated alongside. Distinguishing 'no data' from 'no risk' in the notation is particularly important in practice, because reading a blank as a safety signal is the most common failure mode in this area.

24.5 Processing performance

The figures below are measured wall-clock times for the double-stranded silencing pipeline in an in-house benchmark environment. Turnaround for other modalities depends on scope and species count, and this document does not present estimates for values that were not measured.

Task

Measured time

Note

Single gene, human only

3–5 minutes

Full eight-stage execution

Five-species preclinical combination

8–15 minutes

Cross-species ortholog matching and evaluation run automatically in parallel

Multi-target batch (10 genes)

20–40 minutes

Parallel execution

Deep-ensemble re-ranking (200 candidates)

under 5 seconds

Forward passes over pretrained checkpoints only

Cleavage coverage (per gene)

tens of seconds

Entails a transcript annotation rescan

That a rule-based public tool is faster on a single gene reflects a difference in depth of evaluation. Computing positional rule scores and performing variant masking, transcriptome-wide off-target scanning, accessibility profiling, deep-model inference and chemistry design are not the same task. Conversely, the gap widens in cross-species evaluation because ortholog matching and per-species assessment run automatically in parallel, cutting time substantially against manual alignment analysis.

24.6 Cautions about comparison

25. Detailed Comparison Against Requirements

Design output is placed in three different review environments: academic peer review, use as public or open-source software, and validation as commercial or clinical software. The three demand different things, and satisfying one does not automatically satisfy the others. This chapter decomposes each environment's requirements item by item, states which technical artifact satisfies each, and where it does not, where the boundary lies.

Symbols are defined as follows. ✓ means the requirement is met by the tool itself without external supplementation; ◐ means the core infrastructure exists but site-specific configuration or an additional procedure is required; ✗ means it falls outside the design's scope. For ◐ and ✗ items the responsible party is named, because that clarity is what prevents disputes about accountability during review or audit. And marking ✗ items honestly is what makes the ✓ items credible.

25.1 Academic publication requirements

25.1.1 Three levels of reproducibility

What peer review means by reproducibility is not one concept but three levels. Each demands different evidence, and meeting one does not meet the others.

Level

The question

Evidence required

How it is met

Verdict

Computational reproducibility

Does the same input give the same result

Output agreement for the same input and seed

The determinism contract — order-independent parallel merging, explicit tie-break keys, caching restricted to pure functions

Methods reproducibility

Can a third party reconstruct the same procedure

Complete record of tool versions, database releases, parameters and seed

Every item captured automatically in the run ledger — no need to reconstruct the Methods section afterwards

Results reproducibility

Do other researchers independently reach the same conclusion

Validation on independent data

The design layer supplies the evidence; independent validation is the researcher's experimental domain

◐ — a level no tool can guarantee

25.1.2 Data availability and interoperability

Requirement

What it asks for

How it is met

Verdict

Machine readability

Results in a format read by machines rather than only by eye

Every output provided simultaneously in tabular and structured formats

Identifier standards

Genes, transcripts and variants expressed with standard identifiers

Approved symbols alongside each database identifier, with alias resolution stated

Provenance

It must be possible to know where each value came from

Per-item provenance attached, with reference data releases recorded

Supplementary format

Reviewers must be able to read it without special tools

Report, tables and structured data provided together

Long-term accessibility

Data must remain accessible after publication

Outputs are delivered as files; repository deposition is the author's choice

◐ — author responsibility

25.1.3 Completeness of methods description

A frequent reviewer criticism is incomplete description of what was run with which parameters — thresholds, filter conditions and database versions are commonly omitted. Because the run ledger captures these automatically, the Methods section can reference it directly, which prevents not only omissions but inaccuracies arising from recollection.

Care is also needed when citing predictive performance. The figures in Chapter 24 were measured on a specific public benchmark and are not a guarantee that the same performance reproduces on a customer's data. A paper citing them must state the dataset, and the corresponding conditions are recorded in the relevant deliverable items as well.

25.2 Public and open-source software requirements

25.2.1 Versioning and history

Requirement

What it asks for

How it is met

Verdict

Version identification

The state of the executed code must be uniquely identifiable

An immutable snapshot model — snapshot files do not change, so byte-level comparison confirms identity

Change history

What changed, when and why must be traceable

Per-snapshot history with reasons for change

Dependency disclosure

External tools and libraries depended upon must be stated

The computational tool layer table in 2.4 and version capture in the run ledger

Licence clarity

Licence conditions of each component must be clear

Public tools named directly; proprietary in-house algorithms distinguished by name

✓ — the distinction is documented

Downstream freedom

Results must be usable without restriction

Computation based on public tools is unencumbered, and items involving proprietary in-house algorithms are identifiable by name

◐ — verification depends on the use context

25.2.2 Robustness and graceful degradation

The most common way a public tool loses trust in practice is silent failure under exceptional conditions. If a missing reference file, an uninstalled external tool or an unreachable network resource causes quiet fallback to a default, a result is produced but nobody can say what it means.

Situation

Poor handling

How this layer handles it

Verdict

The expected reference data release is absent

Quietly proceed with another version

Stop at pre-execution verification

An optional external tool is absent

Score that axis as zero

Mark the axis 'not evaluated', exclude it from the composite score, and state so in the report

A model checkpoint is absent

Return a random value

Exclude that model from the ensemble and record which models were used

Target resolution is ambiguous

Pick one arbitrarily

State the resolution result and note where an alias was involved

An item has no data

Blank or zero

State 'no data' and the reason — distinguished from 'no risk'

The last row matters most. Failing to distinguish a blank from a safety signal turns a coverage limitation of the reference data directly into a false safety verdict. If, for example, constraint metrics for sex-chromosome genes are absent from the reference and left blank, they read as 'unconstrained', and a risky target passes a safety gate silently.

25.3 Commercial and clinical software requirements

25.3.1 Electronic records and audit trail

Requirement

Standard

How it is met

Verdict

Audit trail / electronic records

The audit trail provisions of electronic records regulation

A self-contained structured ledger persisted for every run

Time records

Timestamps with explicit timezone

Start and end times in Coordinated Universal Time in a standard format, removing ordering ambiguity across regions

Authority checks

It must be possible to establish who executed the run

The invoking principal is recorded in the ledger; strong authentication is the responsibility of the access control layer

◐ — site infrastructure responsibility

Record integrity

It must be possible to confirm a record was not altered afterwards

The immutable snapshot model — byte-level comparison confirms the producing code is unaltered

Record retention

Records must remain readable for the retention period

Stored in a standard structured format readable without proprietary tools; retention policy is a site responsibility

25.3.2 Validation and control of reference data

Requirement

Standard

How it is met

Verdict

Software validation

Risk-based validation frameworks

Regression and smoke testing run continuously; formal installation, operational and performance qualification are site deliverables tied to the deployed environment

Reference data version control

Laboratory practice provisions on data management

Expected releases verified at start-up, with the run stopped on failure

Reproducible methods

The methods reproducibility expectation of nonclinical safety guidance

Byte-identical output for the same input and seed

Change management

Changes must follow a controlled procedure

Per-snapshot history with reasons; approval procedure belongs to the site quality system

Data integrity principles

Attributable, legible, contemporaneous, original, accurate

The ledger secures attribution, timing and originality; the standard format secures legibility

25.3.3 Linkage to filing material

Requirement

Standard

How it is met

Verdict

Chemistry, manufacturing and controls material

Impurity-related expectations of the quality module

Synthesis impurity projection produced at design time

✓ — at draft level; specification setting belongs to manufacturing

Mutagenic impurity control

The control logic of the relevant guideline

Related impurity items assessed at the formulation and process stages

◐ — at the level of risk exposure

Nonclinical species selection

The species expectations of nonclinical safety guidance

Per-combination comparison of surviving candidates and an automatic two-species verdict

Pharmacokinetic material

Biodistribution and dose translation

Route-specific absorption and distribution summaries with dose translation

◐ — predictions that do not replace measurement

25.3.4 Items outside scope and why

Requirement

Standard

Verdict

Reason and responsible party

Patient information protection

Technical safeguards of health information protection regulation

Processing is limited to sequences and public reference data and does not touch patient information. Where integrated into a workflow that includes patient data, those procedures are the integrating site's responsibility

Medical device quality management

Medical device quality management system standards

Supplied research-use-only; in-vitro diagnostic certification is a separate package undertaken by the integrating manufacturer

Clinical decision support

Regulation of clinical decision support software

It does not support individual patient care decisions and is a research and development tool

Manufacturing process design

Manufacturing regulation

It surfaces risk items and proposes draft specifications; actual process development and validation belong to the manufacturing stage

Declaring out-of-scope items explicitly is itself an element of regulatory compliance. Leaving an obligation the tool was never designed for to be discharged by the tool means, in practice, that nobody discharges it. Naming the responsible party exposes that gap and lets the deploying site place the procedure inside its own quality system.

25.4 The three environments in one table

Nature of the requirement

Academic publication

Public software

Commercial and clinical

Reproducibility

Essential — methods reproducibility is under review

Essential — users must be able to verify

Essential — required by regulation

Provenance recording

Recommended — the basis of the Methods section

Recommended — dependency disclosure

Essential — it is the substance of the audit trail

Version immutability

Recommended

Essential — a citable version

Essential — record integrity

Licence clarity

Recommended

Essential — the precondition for downstream use

Essential — product incorporation

Explicit stop on failure

Recommended

Recommended

Essential — silent failure violates data integrity

Distinguishing absent data from safety

Recommended

Recommended

Essential — prevents false safety verdicts

Patient information handling

Not applicable

Not applicable

Separate procedure where applicable (out of scope here)

Formal qualification

Not applicable

Not applicable

A site deliverable

What the table shows is that the three environments do not conflict. What academic reproducibility requires and what a regulatory audit trail requires are largely the same technical artifacts, differing in stringency and format. Designing to satisfy the most stringent environment therefore satisfies the other two — except that organisational procedures such as formal qualification and certification cannot be discharged by a tool.

26. Limits and Disclaimers

This chapter states what this technical layer does not do and does not guarantee. It is the most important chapter in a technical whitepaper, because the boundaries recorded here define the credibility of everything in the preceding chapters.

26.1 Limits on the nature of predictions

26.2 Limits arising from data

26.3 Limits of scope

26.4 Intellectual property notice

Patent-related statements in this document and in the deliverables are informational and are not legal opinion. Freedom-to-operate depends on jurisdiction, claim construction and prosecution history and is the domain of qualified counsel. The design layer uses generic chemistry within a research-use scope and does not implement any particular company's proprietary chemistry or structures.

27. Selected References

The following are primary sources for the mechanisms, determinants and chemistry rules described in this document. Each judgement in the deliverables carries its corresponding source at item level.

27.1 Mechanism and design rules of RNA interference

27.2 Off-targets and immunity

27.3 Single-stranded modalities and chemistry

27.4 Splicing regulation

27.5 Delivery

27.6 Editing, readthrough and immune modulation

27.7 Predictive models and thermodynamics

27.8 Reproducibility and regulation

28. Document Information

Item

Detail

Published by

Bioneer BioFoundryCenter (바이오니아 바이오파운드리센터)

Document

Therapeutic Oligonucleotide Design Platform — Technical Whitepaper v2.0 (2026)

Contact

geneorder@bioneer.co.kr

Nature of the document

A public technical whitepaper — the biology of each modality, its design requirements, deliverables and limits

Reference data releases and performance figures stated in this document are current as of writing; the values actually used by any individual run are recorded in that run's ledger. Contact: geneorder@bioneer.co.kr