G Fun Facts Online explores advanced technological topics and their wide-ranging implications across various fields, from geopolitics and neuroscience to AI, digital ownership, and environmental conservation.

How a Natural Enzyme Just Accurately Transcribed an Eight-Letter DNA Alphabet

How a Natural Enzyme Just Accurately Transcribed an Eight-Letter DNA Alphabet

A research team at the University of California San Diego has demonstrated that unmodified cellular RNA polymerase—the standard molecular engine found in ordinary Escherichia coli bacteria—can accurately read and transcribe a synthetic eight-letter DNA alphabet into functional RNA.

Published in Nature Communications, the study reveals that nature’s native transcription machinery does not require extensive directed evolution or laboratory re-engineering to process unnatural genetic letters. Using high-resolution cryo-electron microscopy (cryo-EM), the researchers captured structural snapshots at near-atomic resolution, proving that the bacterial enzyme processes synthetic base pairs by deploying the exact same structural checks, catalytic movements, and spatial geometries that it uses for standard biological DNA.

The findings resolve a fundamental question in chemical biology: whether the four-letter genetic alphabet shared by all known life is a strict biochemical requirement for cellular machinery, or simply an evolutionary contingency. By confirming that wild-type multi-subunit RNA polymerases can natively transcribe expanded genetic systems, this discovery removes a significant bottleneck in synthetic biology, opening practical routes toward living organisms with expanded genetic codes, ultra-dense molecular data storage, and custom therapeutics.

NATURAL VS. EXPANDED GENETIC BASES
========================================================================
Standard DNA (4 Letters):
   Purines:  Adenine (A)     --pairs with-->  Thymine (T)     [2 H-bonds]
   Purines:  Guanine (G)     --pairs with-->  Cytosine (C)    [3 H-bonds]

Hachimoji DNA (8 Letters - Expanded System):
   Standard: A : T           and              G : C
   Synthetic Purine Analog:  B (Isoguanine)   --pairs with--> S (Isocytosine)  [3 H-bonds]
   Synthetic Purine Analog:  P (Imidazotriazinone) --pairs with--> Z (Nitropyridone) [3 H-bonds]
========================================================================

The Four-Letter Constraint and the Transcription Bottleneck

For roughly four billion years, every organism on Earth—from deep-sea archaea to humans—has relied on the same four chemical bases to store and execute genetic instructions: adenine (A), cytosine (C), guanine (G), and thymine (T). In RNA, uracil (U) replaces thymine. This four-letter alphabet forms two canonical base pairs: A pairs with T via two hydrogen bonds, and G pairs with C via three.

CANONICAL BASE PAIRING (WATSON-CRICK):
        H-Bond Donors (D) and Acceptors (A)

   A : T Pairing (2 Hydrogen Bonds)
   [Adenine]       N1 (A) <====== H-N3 (D)       [Thymine]
                   C6-NH2 (D) ====> O4 (A)

   G : C Pairing (3 Hydrogen Bonds)
   [Guanine]       O6 (A) <====== H-N4 (D)       [Cytosine]
                   N1-H (D) ======> N3 (A)
                   C2-NH2 (D) ====> O2 (A)

While this system has generated the entirety of terrestrial biodiversity, its structural and functional scope is inherently bounded. In information terms, a four-letter language provides only 64 possible three-letter codons ($4^3 = 64$). Three of these act as stop signals, leaving 61 codons to specify just 20 standard amino acids.

Synthetic biologists have long recognized that this limited lexicon imposes severe boundaries on modern biotechnology:

  • Constrained Chemical Repertoire in Aptamers: Functional RNA and DNA aptamers rely on only four distinct side-chain functionalities, limiting their binding affinities and catalytic profiles compared to proteins, which draw on 20 distinct amino acid side chains.
  • Density Limits in Molecular Data Storage: In synthetic DNA data storage, a four-letter system provides a theoretical information density ceiling of 2 bits per nucleotide position ($\log_2 4 = 2$).
  • Codon Exhaustion in Synthetic Translation: Expanding the genetic code to incorporate non-canonical amino acids (ncAAs) has historically required hijacking existing stop codons (such as amber UAG) or using complex four-base frame-shift codons, which frequently cause translational stalling, low yields, and cellular toxicity.

To overcome these boundaries, researchers have spent decades synthesizing non-standard nucleobases. In 2019, a consortium led by Steven A. Benner at the Foundation for Applied Molecular Evolution synthesized the "Hachimoji" (from the Japanese hachi for eight and moji for letter) genetic system.

The Hachimoji system adds four synthetic building blocks—designated B, S, P, and Z—that pair through mutually exclusive hydrogen-bonding configurations:

  1. B (6-amino-9-(1′-β-D-ribofuranosyl)-4-hydroxy-5-(hydroxymethyl)oxolan-2-yl-1H-purin-2-one, or isoguanine) pairs with S (2-amino-1-(1′-β-D-ribofuranosyl)-4(1H)-pyrimidinone, or isocytosine) through three hydrogen bonds.
  2. P (2-amino-8-(1′-β-D-ribofuranosyl)imidazo[1,2-a]-1,3,5-triazin-4(8H)-one) pairs with Z (6-amino-3-(1′-β-D-ribofuranosyl)-5-nitro-1H-pyridin-2-one) through three hydrogen bonds.

HACHIMOJI BASE PAIRING (HYDROGEN BOND REARRANGEMENT):

   B : S Pairing (3 Hydrogen Bonds - Inverted Donor/Acceptor Pattern)
   [Base B (isoG)]   O6 (A) <====== H-N4 (D)       [Base S (isoC)]
                     N1-H (D) ======> O2 (A)
                     C2-NH2 (D) ====> N3 (A)

   P : Z Pairing (3 Hydrogen Bonds - Pyridine/Imidazotriazine System)
   [Base P]          C4=O (A) <===== H-N6 (D)       [Base Z]
                     N3-H (D) ======> C2=O (A)
                     C2-NH2 (D) ====> C5-NO2 / Acceptor Ring

Although chemists could synthesize these unnatural bases and demonstrate that double-stranded Hachimoji DNA forms stable, predictable helices, a critical functional challenge emerged: transcription.

For an expanded genetic system to be functional, enzymes must read the DNA sequence and accurately synthesize the corresponding RNA transcript. Early efforts stalled because natural polymerases were thought incapable of processing synthetic bases. In initial trials, wild-type enzymes rejected the synthetic nucleotides, stalled at insertion sites, or introduced misincorporation errors.

To bypass this hurdle, researchers initially turned to protein engineering. A team at the University of Texas at Austin led by Andrew Ellington engineered mutant variants of single-subunit bacteriophage T7 RNA polymerase (such as the Y639F/H784A mutant) to force the transcription of Hachimoji templates.

However, relying on engineered, single-subunit viral enzymes created a deep biological divide. Phage enzymes operate via simplified mechanisms that do not reflect the complex, multi-subunit regulatory machinery of cellular life. To transition an eight-letter DNA alphabet from an in vitro curiosity into autonomous, living cellular platforms, researchers had to determine whether complex multi-subunit cellular RNA polymerases could natively read, proofread, and transcribe unnatural genetic codes.


The Mechanical Dilemma: Why Cellular Enzymes Were Expected to Reject Synthetic Bases

Multi-subunit cellular RNA polymerases—such as the 400-kilodalton, five-subunit ($\alpha_2\beta\beta'\omega$) enzyme complex in E. coli—are vastly more stringent than single-subunit viral polymerases. Cellular life depends on extreme transcriptional fidelity; an error rate higher than roughly one mistake in 10,000 to 100,000 base pairs can lead to proteotoxic stress and cellular death.

STRUCTURE OF THE BACTERIAL TRANSCRIPTION ELONGATION COMPLEX:

         +-------------------------------------------------------+
         |               E. coli RNA Polymerase Core             |
         |                                                       |
         |   Subunits:  alpha-I, alpha-II, beta, beta', omega    |
         |                                                       |
         |                    [Main Channel]                     |
         |                          |                            |
  Downstream DNA -------------------+---> Unwinding Cleft        |
                                    |                            |
                               [Active Site] <== Catalytic Mg2+  |
                                    |                            |
    Template Strand (dDNA) --------(+)-------- Incoming NTP      |
    RNA Transcript (rRNA)  ----------+                            |
                                    |                            |
                        [Trigger Loop Domain]                    |
                     (Open = Inactive / Closed = Active)         |
         +-------------------------------------------------------+

To maintain this fidelity, cellular RNA polymerases do not simply check for basic hydrogen bonding. Instead, they subject every incoming nucleoside triphosphate (NTP) to a multi-tiered structural and biochemical interrogation:

1. The Spatial Geometry Filter

The active site of RNA polymerase is shaped to accommodate the precise geometry of a standard Watson-Crick base pair. The distance between the C1′ carbon of the template deoxyribose sugar and the C1′ carbon of the substrate ribose sugar must span approximately 10.8 Å (1.08 nanometers), with a distinct pseudo-twofold symmetry axis. If an incoming base pair is too bulky (purine-purine), too narrow (pyrimidine-pyrimidine), or tilted out of plane, steric clashes with surrounding amino acid residues prevent the active site from adopting a catalytic state.

WATSON-CRICK GEOMETRIC CONSTRAINTS IN THE ACTIVE SITE:

                 ~10.8 Angstroms
      C1' (Template) <-------------------> C1' (Incoming NTP)
            \                                 /
             [Template Base] === [Substrate Base]
                     \               /
                 Pseudo-Twofold Symmetry Axis

2. Minor Groove Electrostatic Checkpoints

DNA and RNA polymerases interact with the minor groove of the newly forming base pair. Standard bases present a regular pattern of hydrogen-bond acceptors in the minor groove: the N3 atom of purines (A and G) and the O2 atom of pyrimidines (T and C) display localized negative electrostatic potentials. Polymerases project conserved amino acid residues into this minor groove pocket to verify the presence of these electron densities. In synthetic bases where chemical substituents shift or eliminate these minor-groove electron pairs, polymerases were historically observed to stall.

3. Trigger Loop Folding and Catalytic Gating

The catalytic core of the polymerase includes an essential, mobile protein element located on the $\beta'$ subunit known as the trigger loop (residues $\beta'$ 916–1146 in E. coli). When the active site is unoccupied or contains a mismatched nucleotide, the trigger loop remains in an unstructured, open conformation, rendering the catalytic center inactive.

Only when a precisely matched nucleotide binds at the insertion site ($i+1$) does the trigger loop undergo a conformational transition, folding into a paired $\alpha$-helical bundle (the trigger helices) that caps the active site. This closure positions critical basic residues—specifically Arg933, Lys939, and His940 in E. coli—to coordinate the triphosphate group of the substrate and align two catalytic divalent magnesium ions ($\text{Mg}^{2+}_A$ and $\text{Mg}^{2+}_B$). These ions facilitate the nucleophilic attack of the RNA primer's $3'\text{-OH}$ group on the $\alpha$-phosphate of the incoming NTP, forming the phosphodiester bond.

THE CATALYTIC GATEKEEPING MECHANISM (TRIGGER LOOP ACTION):

    INCORRECT BASE / MISMATCH:          CORRECT BASE MATCH (NATIVE OR EXPANDED):
    
       +-----------------------+           +-----------------------+
       | Active Site Open      |           | Active Site Closed    |
       | Trigger Loop: Flexible|           | Trigger Loop: Folded  |
       | Catalytic Ions: Misaligned        | Catalytic Ions: Aligned
       +-----------------------+           +-----------------------+
                  |                                    |
                  v                                    v
          [No Catalysis]                     [Phosphodiester Bond]
          Substrate Rejected                 Chain Extension (~15-50 nt/s)

Because unnatural base pairs modify the functional groups, electrostatics, and chemical surfaces of the nucleobase core, the prevailing assumption was that cellular RNA polymerases would fail this three-stage checkpoint. Scientists expected that either steric hindrance would block trigger loop closure, or the absence of canonical minor-groove contacts would trigger transcriptional arrest.


How the Bacterial Enzyme Transcribed the Eight-Letter Alphabet

Led by Dong Wang, PhD, professor at the UC San Diego Skaggs School of Pharmacy and Pharmaceutical Sciences, and first author Qingrong Li, the research team designed an integrated experimental platform combining quantitative in vitro transcription kinetics with single-particle cryo-electron microscopy.

The researchers reconstituted fully functional E. coli RNA polymerase elongation complexes containing synthetic DNA templates carrying Hachimoji non-standard nucleotides (B, S, P, and Z) alongside canonical bases. They then monitored single-nucleotide incorporation rates, extension kinetics across multi-base unnatural cassettes, and structural dynamics.

EXPERIMENTAL ELONGATION COMPLEX DESIGN:

   Non-Template DNA:  5'- ... G T A C C T G C C G C C A C C T ... -3'
   Template DNA:      3'- ... C A T G G A C G G X G G T G G A ... -5'  (where X = dZ, dP, dB, or dS)
   RNA Primer:        5'-       A U G G A G A G G -3' (3'-OH attack ready)
                                                |
                                        Incoming Substrate:
                                        (riboPTP, riboZTP, riboSTP, or riboBTP)

The kinetic results established that wild-type E. coli RNA polymerase does not merely tolerate the unnatural bases; it processes them with remarkable efficiency and fidelity:

  • P:Z Pair Kinetics: When reading a template deoxyribo-Z ($\text{dZ}$), the natural enzyme incorporated the complementary substrate ribo-P triphosphate ($\text{riboPTP}$) with catalytic efficiency ($k_{cat}/K_m$) approaching standard native purine-pyrimidine base pairs. The transcription velocity of the synthetic $\text{P:Z}$ pair occurred at a rate only approximately two times slower than a canonical $\text{G:C}$ pair under identical physiological conditions.
  • B:S Pair Fidelity: For the $\text{dB:STP}$ and $\text{dS:BTP}$ combinations, the natural polymerase exhibited strong selectivity, preferentially pairing the unnatural triphosphate substrate opposite its designated unnatural template partner over standard cellular NTPs ($\text{ATP}$, $\text{GTP}$, $\text{CTP}$, $\text{UTP}$) by several orders of magnitude.
  • Continuous Synthesis: The enzyme did not permanently stall after incorporating a synthetic nucleotide. It successfully translocated downstream, reloaded the next natural or unnatural triphosphate, and continued processive elongation along the template.

SINGLE-NUCLEOTIDE TRANSCRIPTION PERFORMANCE (RELATIVE RATES):

   Base Pair Context       Relative Transcription Efficiency (kcat / Km)
   ---------------------------------------------------------------------
   Canonical G : C         ====================================  100%
   Synthetic P : Z         ==================                     48%
   Canonical A : T         ================                42%
   Synthetic B : S         ============                           31%
   Mismatched Pair (dZ:A)  -                                      <0.01%
   ---------------------------------------------------------------------

To determine how the enzyme accomplished this, the UC San Diego team used high-end Titan Krios transmission electron microscopes equipped with Gatan direct electron detectors to solve structures of the transcription elongation complexes trapped at various catalytic stages.

The cryo-EM structures revealed why the wild-type enzyme succeeds:

  1. Perfect Watson-Crick Isostericity: Within the catalytic insertion site ($i+1$), both the synthetic $\text{P:Z}$ and $\text{B:S}$ base pairs adopt an edge-to-edge coplanar geometry that matches the canonical $\text{G:C}$ and $\text{A:T}$ base-pairing envelope. The $\text{C1′–C1′}$ inter-sugar distance remains locked at approximately 10.7 to 10.8 Å, preventing any distortion of the phosphodiester backbone.
  2. Native-State Trigger Loop Closure: Despite the non-canonical chemical structures of the synthetic bases, their accommodation in the active site allows the trigger loop to transition smoothly into its fully closed, active conformation. Residues Arg933, Lys939, and His940 form identical hydrogen-bonding networks with the triphosphate moiety of the incoming unnatural nucleotide, correctly positioning the catalytic $\text{Mg}^{2+}$ ions.
  3. Absence of Active-Site Repulsion: Surrounding amino acid side chains within the $\beta$ and $\beta'$ subunits—including conserved residues in the bridge helix and the F-loop—interact with the synthetic base pairs via neutral van der Waals packing, avoiding electrostatic clashes that would otherwise destabilize the active complex.

CRYO-EM RESOLUTION OF CATALYTIC CORE WITH EXPANDED BASES:

      Bridge Helix (beta' Subunit)
      ============================
             |             |
         [Base dZ] <===> [riboPTP]   <-- Watson-Crick Coplanar Alignment
             |             |             (C1'-C1' = 10.8 Angstroms)
      ----------------------------
      Catalytic Center:
         * Mg2+ (A) aligned with Primer 3'-OH
         * Mg2+ (B) stabilizing P-alpha, P-beta, P-gamma
         * Trigger Loop: FULLY FOLDED (Closed State)
           (Arg933, Lys939, His940 locked onto triphosphate)

Expanding the Model: Lessons from Hydrophobic Unnatural Base Pairs

In parallel work published in the Proceedings of the National Academy of Sciences (PNAS), Dong Wang’s laboratory, collaborating with synthetic biologist Ichiro Hirao and colleagues, investigated an entirely different class of unnatural genetic letters: hydrophobic unnatural base pairs (UBPs).

Unlike Hachimoji bases, which rely on rearranged hydrogen-bonding networks, hydrophobic base pairs—such as the 7-(2-thienyl)imidazo[4,5-b]pyridine ($\text{Ds}$) and pyrrole-2-carbaldehyde ($\text{Pa}$) system—lack canonical complementary hydrogen-bonding groups entirely. Instead, they pair through shape complementarity, hydrophobic packing, and $\pi\text{–}\pi$ aromatic stacking.

CHEMICAL FOUNDATIONS: HYDROGEN-BONDED VS. HYDROPHOBIC BASES

   Hachimoji System (P : Z / B : S)         Hydrophobic System (Ds : Pa)
   ---------------------------------         ----------------------------
   * Governed by H-bonding donor/acceptor     * Zero canonical H-bonds across pair
   * Isosteric with Watson-Crick geometry    * Governed by shape complementarity
   * Strict electrostatic alignment           * Governed by London dispersion & pi-stacking

The PNAS study demonstrated that wild-type E. coli RNA polymerase can also process these hydrophobic pairs, although it revealed an intriguing mechanistic asymmetry:

  • When the template contained the unnatural base $\text{dPa}$, incoming $\text{DsTP}$ was incorporated with high efficiency, driving full trigger loop closure and rapid phosphodiester bond formation.
  • Conversely, when the template contained $\text{dDs}$, the incorporation of incoming $\text{PaTP}$ was nearly 30-fold slower, because the active site struggled to stabilize the smaller $\text{Pa}$ moiety prior to chemistry.

Cryo-EM structures of the $\text{dPa:DsTP}$ complex resolved at 3.20 Å resolution captured the exact pre-catalytic state, demonstrating that the trigger loop can fold into its active conformation even when hydrogen bonds between the pairing bases are entirely absent.

CRYO-EM STRUCTURAL COMPARISONS:

   Metric / Feature          Canonical G:C       Hachimoji P:Z       Hydrophobic Ds:Pa
   -----------------------------------------------------------------------------------
   Inter-sugar C1'-C1'       10.8 Angstroms      10.8 Angstroms      10.9 Angstroms
   Hydrogen Bonds            3                   3                   0
   Trigger Loop State        Closed (Folded)     Closed (Folded)     Closed (Folded for dPa:DsTP)
   Cryo-EM Resolution        ~2.8 - 3.1 A        ~3.0 - 3.2 A        3.20 A
   Catalytic Rate (kcat)     Fast (~30-50 s^-1)  Moderate-Fast       Asymmetric (Ds > Pa)

Together, the Nature Communications and PNAS studies rewrite standard models of transcription. They prove that cellular RNA polymerases are not hardwired to recognize specific chemical elements unique to adenine, thymine, cytosine, and guanine.

Instead, the enzyme functions as a versatile geometric and spatial filter. As long as a synthetic base pair satisfies the strict boundary conditions of Watson-Crick dimensional volume and does not project steric bulk into forbidden zones, the natural multi-subunit enzyme will close its trigger loop, recruit magnesium ions, and transcribe the code.


Schrödinger's Aperiodic Crystal and the Limits of Life

The fact that natural enzymes natively transcribe an eight-letter DNA alphabet provides direct experimental evidence for a core concept in theoretical biology: Erwin Schrödinger’s 1944 proposal of the aperiodic crystal.

In his foundational text What is Life?, Schrödinger posited that the hereditary material of living systems must satisfy two competing thermodynamic requirements:

  1. Regularity (The Crystal): The molecule must possess a repetitive, uniform structural backbone so that physical enzymes can replicate, read, and manipulate it using a single set of standardized molecular machinery regardless of length.
  2. Aperiodicity (Information Storage): The molecule must allow virtually infinite variations in sequence order—swapping chemical letters at any position without disrupting the overall crystalline geometry of the macroscopic fiber.

THE APERIODIC CRYSTAL PARADOX:

   Periodic Crystal (e.g., Salt / Diamond):
   [ A ] --- [ A ] --- [ A ] --- [ A ] --- [ A ]
   * Maximum stability
   * Zero information storage capacity

   Standard DNA (4-Letter Aperiodic Crystal):
   [ G ] === [ C ] --- [ A ] === [ T ] --- [ G ]
   * Uniform helical parameters
   * Variable sequence = High information storage

   Hachimoji DNA (8-Letter Aperiodic Crystal):
   [ G ] === [ Z ] --- [ P ] === [ S ] --- [ B ] === [ T ] --- [ A ] === [ C ]
   * Same uniform helical parameters (10.8 A base-pair width)
   * Doubled combinatorial variety = Doubled information density per unit length

Natural four-letter DNA solves this problem because the $\text{A:T}$ and $\text{G:C}$ pairs have identical overall sizes and external shapes, forming an invariant double helix regardless of sequence.

The structural and biochemical data from the UC San Diego experiments confirm that Hachimoji DNA and its corresponding transcripts fully satisfy the Schrödinger requirement. By preserving identical inter-strand distances, base-stacking properties, and backbone orientations across eight letters, Hachimoji biopolymers maintain a stable physical architecture that natural enzymes can navigate.

This finding has major implications for astrobiology and the search for extraterrestrial life:

  • The Fallacy of Terrestrial Uniqueness: The universal use of A, T, C, and G across terrestrial life is not a chemical imperative. There is nothing uniquely privileged about the four bases that emerged in the prebiotic broth of early Earth.
  • Alternative Evolutionary Trajectories: Life emerging on other worlds, or synthesized in future exobiology laboratories, could readily operate on 6-, 8-, 10-, or 12-letter genetic alphabets using alternative hydrogen-bonding arrangements, provided those letters conform to the constraints of an aperiodic crystal.


Practical Applications: High-Density Data Storage, XenoAptamers, and Synthetic Therapeutics

The ability of natural enzymes to transcribe an eight-letter DNA alphabet without genetic modification removes a major barrier for several emerging technologies.

IMPACT SECTORS FOR EXPANDED ALPHABETS:

   +--------------------------+--------------------------+--------------------------+
   |   Molecular Data Storage |    Medical Diagnostics   |  Therapeutic Aptamers    |
   +--------------------------+--------------------------+--------------------------+
   | * 3 bits/base vs 2 bits  | * Zero background noise  | * Picomolar target affinity
   | * 50% increase in density| * Multiplexed detection  | * Hydrophobic binding    |
   | * Shorter oligo lengths  | * Viral variant typing   | * Extreme serum stability|
   +--------------------------+--------------------------+--------------------------+

1. Ultra-High-Density DNA Data Storage

Modern digital data storage faces an impending silicon and physical footprint crisis. DNA data storage offers a theoretical density millions of times greater than magnetic tape or flash memory, capable of storing hundreds of petabytes in a single gram of biomaterial.

However, physical DNA synthesis has historically been throttled by synthesis length limits and error rates. Expanding from a four-letter code to an eight-letter alphabet transforms the mathematics of molecular storage:

$$\text{Information per nucleotide (bits)} = \log_2(N)$$

  • For standard 4-letter DNA ($N = 4$): $\log_2(4) = 2.0\text{ bits per base}$.
  • For 8-letter Hachimoji DNA ($N = 8$): $\log_2(8) = 3.0\text{ bits per base}$.

This represents an immediate 50% increase in raw data storage density per nucleotide position. A digital file that requires a 300-nucleotide strand in standard DNA can be stored in just 200 nucleotides using an eight-letter alphabet.

Because the error rate of chemical oligonucleotide synthesis scales exponentially with strand length, compressing data into shorter strands increases synthesis yields, lowers chemical reagent costs, and speeds up sequencing readouts. The ability of wild-type enzymes to read and transcribe these strands into RNA opens direct enzymatic routes for copying, amplifying, and retrieving stored molecular data.

DATA DENSITY COMPARISON (THEORETICAL MAXIMUM):

   System       Letters   Bits / Base   Bytes / 100-mer Strand   Relative Capacity
   -------------------------------------------------------------------------------
   Binary       2 (0,1)   1.0           12.5 Bytes               Baseline (1x)
   Native DNA   4 (ATGC)  2.0           25.0 Bytes               2.0x
   Hachimoji    8 (+BSZP) 3.0           37.5 Bytes               3.0x
   Expanded-16  16        4.0           50.0 Bytes               4.0x
   -------------------------------------------------------------------------------

2. Next-Generation XenoAptamers for Oncology and Infectious Disease

Aptamers—often termed "chemical antibodies"—are short single-stranded DNA or RNA oligonucleotides that fold into intricate tertiary shapes to bind specific molecular targets with high affinity. While natural four-letter aptamers have achieved clinical success (such as pegaptanib for macular degeneration), their chemical diversity is limited.

Incorporating synthetic bases like P, Z, B, S, and hydrophobic residues like Ds expands the functional landscape of aptamers:

  • Nitro- and Imidazo-Functionalization: The nitro ($-\text{NO}_2$) group on base Z and the amino-heterocyclic surface of base P create electrostatic and hydrogen-bonding motifs that do not exist in standard nucleic acids.
  • Cancer Cell Targeting: Synthetic aptamers incorporating P and Z have already been selected to recognize human liver cancer cell lines (HepG2) with high specificity, distinguishing malignant tissue from healthy hepatocytes while delivering targeted chemotherapeutic payloads.
  • Dengue and Viral Diagnostics: Aptamers utilizing expanded genetic letters (XenoAptamers) targeting the Non-Structural Protein 1 (NS1) of the Dengue virus achieve binding affinities ($K_D$) in the low picomolar range, enabling rapid serotyping that distinguishes between four closely related viral variants.

Because natural E. coli RNA polymerase can transcribe these motifs, researchers can now deploy standard, high-throughput in vitro selection protocols (SELEX) using natural transcription machinery, eliminating the need to optimize custom mutant enzymes for every candidate library.

XENOAPTAMER EVOLUTION CYCLE WITH NATIVE POLYMERASE:

      [Synthetic 8-Letter DNA Library (10^15 variants)]
                            |
                            v
   [Transcription by WILD-TYPE E. coli RNA Polymerase]
                            |
                            v
          [8-Letter RNA Aptamer Folded Pool]
                            |
                            v  Binding Screen
                 [Target Protein / Cancer Marker]
                            |
                            v  Elution of Hits
            [High-Affinity Bound RNA Fractions]
                            |
                            v
        [Reverse Transcription & Deep Sequencing]

3. Molecular Diagnostics with Zero Background Noise

In clinical diagnostics, non-specific binding and cross-hybridization create background noise that limits the sensitivity of PCR, multiplexed panels, and hybridization assays. Synthetic bases like P, Z, B, and S do not pair with natural A, T, C, or G. Diagnostic probes engineered with non-standard bases can bind exclusively to synthetic target sequences with near-zero off-target background, enabling the ultra-sensitive detection of low-copy-number pathogens, including HIV, Hepatitis C, and early-stage circulating tumor DNA (ctDNA).


Remaining Roadblocks: The Path to Living Synthetic Organisms

While the demonstration that natural RNA polymerase can transcribe an eight-letter DNA alphabet in vitro marks a critical milestone, transitioning this chemistry into living, self-sustaining organisms presents several complex biochemical challenges.

TECHNICAL HURDLES: IN VITRO RECONSTITUTION TO IN VIVO LIFE

   +------------------------------------------------------------------------+
   | 1. NUCLEOTIDE BIOSYNTHESIS & UPTAKE                                    |
   |    Challenge: Cells do not possess metabolic pathways to make B,S,P,Z. |
   |    Solution: Express exogenous nucleoside triphosphate transporters   |
   |              (e.g., algal PtNTT2) to import precursors from media.     |
   +------------------------------------------------------------------------+
                                      |
                                      v
   +------------------------------------------------------------------------+
   | 2. DNA REPLICATION FIDELITY ACROSS GENERATIONS                         |
   |    Challenge: Replicative polymerases (Pol III) must copy 8 letters    |
   |               with >99.99% fidelity to avoid mutational loss.          |
   |    Solution: Directed evolution of sliding clamps and proofreaders.    |
   +------------------------------------------------------------------------+
                                      |
                                      v
   +------------------------------------------------------------------------+
   | 3. INTRACELLULAR REPAIR SYSTEM EVASION                                 |
   |    Challenge: MutS, Uracil-DNA glycosylases (UDG) may flag synthetic   |
   |               bases as damaged DNA and excise them.                    |
   |    Solution: Genetic knockout or tuning of non-essential repair paths. |
   +------------------------------------------------------------------------+
                                      |
                                      v
   +------------------------------------------------------------------------+
   | 4. EXPANDED TRANSLATION (8-LETTER TRANSLATION CODE)                    |
   |    Challenge: Ribosomes and tRNAs must read 8-letter mRNA codons.      |
   |    Solution: Engineering orthogonal synthetase/tRNA pairs to assign    |
   |              expanded codons to non-canonical amino acids.             |
   +------------------------------------------------------------------------+

1. Metabolic Synthesis and Nucleotide Transport

Currently, synthetic nucleotides must be chemically synthesized and supplied externally. A fully autonomous eight-letter organism must either:

  • Express engineered metabolic pathways consisting of novel synthetic enzymes capable of synthesizing dPTP, dZTP, dBTP, and dSTP from simple cellular metabolites, or
  • Express broad-spectrum nucleoside triphosphate transporters in its outer and inner membranes—such as the algal transporter PtNTT2 from Phaeodactylum tricornutum—to import phosphorylated synthetic precursors directly from growth media.

2. Genomic Replication Fidelity

Transcription is only one half of the central dogma. For an organism to proliferate, its DNA replication machinery (such as DNA Polymerase III holoenzyme in bacteria) must replicate the eight-letter genome across generations with error rates below the error catastrophe threshold.

While DNA polymerases have replicated synthetic six-letter plasmids in pioneering work by Floyd Romesberg’s group, replicating an eight-letter genome containing contiguous stretches of synthetic bases requires optimizing polymerase proofreading domains ($\epsilon$ subunit exonucleases) to ensure unnatural bases are not erroneously excised during synthesis.

3. Tautomerization and Mispairing Control

Synthetic nucleobases exist in a dynamic equilibrium between different chemical tautomers (such as keto-enol or amino-imino forms). If a synthetic base temporarily shifts into a minor tautomeric state, its hydrogen-bond donor and acceptor pattern alters, potentially leading to mispairing. For example, the rare imino tautomer of base B can mispair with standard thymine, leading to transition mutations during prolonged replication.

Synthetic chemists are addressing this by fine-tuning ring substituents—such as attaching electron-withdrawing groups to stabilize the desired amino-keto tautomeric state—to maintain high pairing fidelity over multiple cellular generations.

4. Genetic Containment and Biosafety

A natural safety feature of synthetic base systems is intrinsic biocontainment. Organisms engineered to rely on an eight-letter DNA alphabet are auxotrophic for unnatural, laboratory-synthesized chemical building blocks. If such an organism escapes into the wild, it cannot find synthetic nucleobases (P, Z, B, S) in the natural environment. Without these building blocks, the cell cannot replicate its genome or transcribe essential synthetic genes, causing replication arrest and cell death. This provides a robust genetic firewall preventing ecological contamination.


The Trajectory of Expanded Genetic Systems

The structural demonstration that natural bacterial RNA polymerase processes an eight-letter genetic alphabet with high catalytic efficiency shifts synthetic biology from an era of bespoke protein engineering to one of systemic integration.

TIMELINE OF EXPANDED GENETIC CODE MILESTONES:

   1944: Erwin Schrödinger publishes "What is Life?", predicting the aperiodic crystal.
     |
   1989: Steven Benner's lab synthesizes the first non-standard base pair (isoC:isoG).
     |
   2014: Floyd Romesberg's team creates the first semi-synthetic organism with a 6-letter DNA code.
     |
   2019: Benner and colleagues publish the 8-letter Hachimoji DNA system in Science.
     |
   2023: Early cryo-EM models reveal basic geometric tolerance in single-subunit polymerases.
     |
   2026: UC San Diego team proves WILD-TYPE cellular multi-subunit RNA polymerase accurately
         transcribes an 8-letter DNA alphabet via native trigger loop closure (Nature Comms/PNAS).
     |
   FUTURE: Fully autonomous 8-letter living cells, clinical XenoAptamers, and 3-bit DNA storage.

Over the coming years, researchers are focused on several defined milestones:

  • In Vivo Reconstitution: Reconstituting active Hachimoji transcription inside live, intact E. coli cells equipped with membrane nucleotide transporters.
  • Eight-Letter In Vitro Translation: Coupling eight-letter transcription directly to cell-free translation systems, using synthetic mRNA transcripts containing expanded codons to incorporate multiple distinct non-canonical amino acids into single designer enzymes.
  • Commercialization of High-Density Storage: Integrating enzymatic eight-letter transcription into automated, microfluidic DNA storage and retrieval platforms to lower the cost of archival data storage.

By demonstrating that life's core molecular machinery can read and transcribe an expanded genetic code using its native architecture, this discovery confirms that the language of life is not limited to four letters. The chemical toolkit of biology is fundamentally modular, open to expansion, and capable of operating across a wider molecular landscape than nature originally selected.

Reference:

Share this article

Enjoyed this article? Support G Fun Facts by shopping on Amazon.

Shop on Amazon
As an Amazon Associate, we earn from qualifying purchases.