Medical Biology · Year 1 · Medical University of Sofia
10
RNA replication. Reverse transcription
Free notes for topic 10 of the Medical Biology syllabus, open without an account. Written by a senior student against the syllabus question and checked line by line by a second student before publishing. How content is made
Updated
What this topic covers
note
RNA replication: RNA-dependent RNA polymerase, viral RNA sense, the replication of negative-sense RNA and its mechanism, the replication/transcription switch, and how RNA from different viruses can enter a single virion.
Reverse transcription: the different roles of reverse transcription, the reproductive cycle of retroviruses, reverse transcriptase as two enzymes in one, the replication of chromosome telomeres, and telomeres and ageing.
Both processes are exceptions to the central dogma as it was first written, and both were discovered because something failed to behave as the dogma predicted.
Palm-fingers-thumb fold shown as a hand cartoon beside a ribbon structure, the shared architecture of RNA replicase, RNA polymerase and reverse transcriptase
1. Viruses and the Baltimore classification
note
For organisms that consist of cells there is one presumed common ancestor, LUCA. For viruses there is probably nothing like a common ancestor.
We are nearly sure that evolution created different viruses at different times, using different materials and in different ways. The main evidence is that the genetic mechanisms used by viruses differ a great deal from one another.
The Baltimore classification, named after David Baltimore, clusters viruses into 7 groups according to their type of genome. Different viral genomes use different ways to produce mRNA, and different ways to multiply their genomes.
Two of those ways are new here: RNA replication and reverse transcription.
Baltimore classification of virus genomes into seven groups by genetic material, showing the route from each group to mRNA and, for groups VI and VII, reverse transcription to DNA(+/-)
2. RNA replication
note
How it was discovered. In the early 1960s, studies on mengovirus and poliovirus observed that these viruses were not sensitive to actinomycin D, a drug that inhibits cellular DNA-directed RNA synthesis. That insensitivity suggested there must be a virus-specific enzyme able to copy RNA without a DNA template.
The process: synthesis of RNA using an RNA template.
The official name: RNA-dependent RNA replication.
The main enzyme: RNA-dependent RNA polymerase (RdRP), also called RNA replicase. It is encoded by the viral genome.
Substrates: nucleoside triphosphates (NTPs).
Who uses it: RNA-dependent RNA replication is reserved exclusively for RNA viruses, not for cellular RNAs. Almost all RNA viruses except the retroviruses undergo it - in other words, all RNA-containing viruses with no DNA stage.
The two stages:
Initiation. RNA synthesis begins at or near the 3' end of the RNA template. It can be primer-independent (de novo) or primer-dependent. The primer-dependent mechanism requires a viral protein called the genome-linked primer (VPg). De novo initiation consists in the addition of an NTP to the 3'-OH of the first initiating NTP.
Elongation. The same nucleotidyl transfer reaction is repeated with subsequent NTPs, generating the complementary RNA product.
RNA-dependent RNA polymerase in action: single-stranded RNA template entry, NTP entry and Mg2+ ions at the catalytic site, and double-stranded RNA exit
Where the RNA replicase came from
note
Many eukaryotes have their own RNA-dependent RNA polymerase, used for the amplification of regulatory microRNAs.
It is possible that viral RNA replicases evolved from this enzyme.
Various lineages of eukaryotes have, however, lost the RNA replicase during their evolution. The vertebrates have lost it - which is why we have no cellular equivalent of the process.
Phylogenetic tree of animal lineages marking which groups keep a eukaryotic RNA-dependent RNA polymerase, filled circles for example nematoda, arachnida and myriapoda, and which have lost it, open circles including vertebrata
The structure of RdRP
note
RdRP catalyses the synthesis of the RNA strand complementary to a given RNA template, in the 5' to 3' direction. The substrates are the ribonucleotide 5' triphosphates ATP, GTP, UTP and CTP.
The structural arrangement of RdRP forms two channels that meet at the active site:
the main channel accommodates the RNA template;
the secondary channel allows the entry of the incoming nucleoside triphosphates.
RNA-dependent RNA polymerase active site: single-stranded RNA template entry, NTP entry, Mg2+ ions in the catalytic site, and double-stranded RNA exit
3. Sense in viruses
note
First, a reminder of what antisense RNA means in a cell: an RNA sequence complementary to an endogenous mRNA. When the mRNA forms a duplex with it, translation is blocked. This is related to RNA interference, and it is applied experimentally as antisense therapy.
For viruses the term "sense" has a slightly different meaning, and it is used as a basis for classifying them.
Positive-sense (plus-strand) viral RNA, read 5' to 3', can be directly translated into viral proteins. The viral RNA genome can be considered viral mRNA and is immediately translatable by the host cell. Because of this, these viruses do not need an RNA replicase packaged into the virion - the replicase will be one of the first proteins produced by the host cell.
Negative-sense (minus-strand) viral RNA, read 3' to 5', is complementary to the viral mRNA. Like DNA, it cannot be translated directly. It must first be transcribed into a positive-sense RNA that acts as an mRNA. Some viruses, for example the influenza viruses, have negative-sense genomes and must carry an RNA replicase inside the virion.
Double-stranded RNA viruses also carry a ready-to-use RNA replicase; viral mRNAs are synthesised using one of the viral genomic strands as a template.
A point worth holding on to: in the multiplication of RNA viruses, transcription and replication mean largely the same thing.
Baltimore classification chart contrasting RNA(+) genomes that lead straight to mRNA with RNA(-) genomes that must first be copied to an RNA(-) template, alongside the DNA and reverse-transcribing groups
4. Principles of negative-sense RNA replication
note
A (-)RNA genome must be transcribed into mRNAs before translation.
Transcription is performed by the RNA replicase (RdRP), which is packaged in the virion.
The (+) copies serve as template both for making viral proteins and for making new genomic (-)RNA.
Some (-)RNA viruses have one genomic RNA molecule (nonsegmented genomes); some have several RNA molecules in the virion (segmented genomes).
The important difference between them is the location of replication.
Nonsegmented genomes replicate in the cytoplasm. The single RNA is a template for multiple mRNAs.
Segmented genomes replicate in the nucleus, and the RdRP produces one mRNA strand from each genome segment.
Negative-sense RNA virus cycle: entry, RdRP-mediated mRNA transcription and replication to positive-sense RNA, translation, assembly and release of new virions
The mechanism of (-)RNA replication
note
Synthesis starts from the leader sequence.
(+) copies are produced.
The (+) copies become templates for new (-) viral genomes.
The synthesis of the new (-) viral RNA starts from a sequence called the trailer.
Mechanism of negative-sense RNA replication: polymerase complex starts at the leader on the (-) genome, produces the (+) antigenome, then restarts at the trailer to make new (-) genomes
The replication/transcription switch
note
In (-)RNA viruses the RNA replicase works as a transcriptase by default. It turns into a replicase when that becomes necessary, and the trigger is the concentration of capsid proteins: if there are enough capsid proteins to make new virions, then new viral genomes must be produced by replication.
This is an elegant piece of accounting - the virus reads how many empty shells it has and makes exactly as many genomes as it can package.
How transcription runs:
RdRP initiates transcription by binding to the leader.
The enzyme produces a short leader RNA, then stops and restarts on a transcription initiation signal.
The RNA initiated on this signal is capped.
At the end of each viral gene there is a transcription stop signal, on which the RdRP produces a poly-A tail by stuttering on a U stretch, before releasing the mRNA.
RdRP can then scan to the next transcription initiation signal and resume transcription on the next gene.
Replication-transcription switch: polymerase complex initiates at the transcription start signal, produces leader RNA, and finishes each capped mRNA with a poly-A tail at the transcription stop signal
Segmented genomes allow mixing between viruses
note
Influenza has a (-)RNA segmented genome consisting of 8 (-)RNA molecules.
If two separate strains simultaneously infect a host cell, the new virions can acquire a new combination of segments, producing a third strain.
Antigenic shift: two influenza A strains co-infect a host cell, their segmented genomes intermix, and a new reassortant virus C emerges
It is exactly such combined influenza viruses that cause pandemics.
Influenza virion cross-section labelled haemagglutinin (HA), matrix (M1), neuraminidase (NA), polymerase complex (PA, PB1, PB2), nucleoprotein (NP), M2 ion channel and nuclear export protein (NEP)
An exception: human deltavirus (HDV)
note
HDV is a satellite virus: infection requires the host cell to be co-infected with hepatitis B virus.
Its replication cycle:
Inside the nucleus, the (-)RNA genome is replicated by a rolling circle mechanism, producing a chain of multiple (+) genomes.
These are self-cleaved and processed by viral ribozymes.
The new (+)RNA then serves as template for the synthesis of genomic RNA.
As viral proteins are synthesised in the cytoplasm, they are recruited to the nucleus to form RNA-protein complexes (RNP).
These are subsequently exported to the cytoplasm.
There they interact with the surface antigen of hepatitis B virus (HBsAg) to form mature HDV virions - which is why HDV cannot manage without HBV.
HDV replication cycle: nuclear rolling-circle replication of the (-)RNA genome, ribozyme processing, RNP export to the cytoplasm and assembly with hepatitis B surface antigen, HBsAg, into HDV virions
5. Principles of positive-sense RNA replication
note
(+)RNA can function both as a genome and as messenger RNA. It can be directly translated into protein by host ribosomes.
The genes necessary for RNA replication are expressed first.
A single polyprotein is produced, which is then processed into individual proteins.
The newly made RNA replicase copies the (+)RNA into complementary (-)RNA, and the new (-) strands serve as template for new (+) strand synthesis.
The (+)RNA produced can then be translated, replicated, or packaged into new virions.
Example: the HCV genome. The hepatitis C virus genome is a single-stranded RNA encoding a single large open reading frame (ORF), flanked by 5' and 3' non-coding regions. Translation of that ORF generates a large polyprotein, which undergoes a complex co- and post-translational series of cleavage events catalysed by both host and viral proteases, to produce the 10 individual HCV proteins. These are integrated into the ER membrane.
Positive-sense RNA virus cycle: ribosome translation of the (+)RNA, RdRP-mediated synthesis of negative-sense RNA, replication back to (+)RNA, assembly and releaseHepatitis C virus genome translated as one open reading frame into a polyprotein, processed by host signal peptidase and viral NS2/3 and NS3/4A proteases into structural and nonstructural proteins
Virus factories
note
To avoid antiviral defence, (+)RNA virus replication occurs in protected compartments.
The places of viral replication and assembly in the cell are given various names: virus factories, virus inclusions, or viroplasm. In most cases they are wrapped in a membrane.
The viral genomes encode proteins that manipulate the cell membranes and cytoskeleton. Their task is to provoke membrane invaginations in a variety of organelles - the endoplasmic reticulum, mitochondria, vacuoles, Golgi apparatus, peroxisomes, plasma membrane, and the outer nuclear membrane. These membranous invaginations are used as the virus factories.
Membranous invagination forming a virus factory: (+)RNA genome translated, RdRP builds the double-stranded RNA replication form, and new (+)RNA genomes are released by strand-displacement synthesis
The steps of (+)RNA replication and translation
note
(+)RNA viruses enter cells by endocytosis.
Once inside, the (+)RNA genome is released into the cytosol, where it is translated by host ribosomes.
The resulting viral replication proteins recruit the (+)RNA to subcellular membrane compartments, where the virus factories are assembled.
A small amount of (-)RNA is synthesised, and it serves as template for the synthesis of a large number of (+)RNA copies.
The new (+)RNAs are released from the virus factory, whereas the (-)RNA is retained - keeping the dangerous double-stranded intermediate hidden.
The released (+)RNAs start a new cycle of translation and replication, become encapsidated, and then exit the cell.
Steps of positive-sense RNA replication: entry, uncoating and disassembly, translation, degradation of some (+)RNA, encapsidation, egress and cell-to-cell movement via plasmodesmata
6. Replication of double-stranded viral RNA
note
The governing constraint is a defensive one. Normally dsRNA is not produced in the cell, and host cells have various antiviral systems that detect and inactivate dsRNA. That is why naked genomic dsRNA never enters the cytoplasm.
All dsRNA viruses have segmented genomes. Each dsRNA segment codes for a single polypeptide.
Each virion carries its own RNA replicase, responsible for both transcription and replication. Both processes occur inside the capsid.
The incoming virion is only partially uncoated. This partial uncoating activates transcription of all the (-) strands, creating (+)mRNAs.
The mRNAs are extruded from the virion and translated by the cellular machinery.
Each mRNA is later encapsidated and copied once, inside the capsid, to form dsRNA. Several hours pass between the synthesis of the (+) and the (-) strands.
Viruses with dsRNA never enter the nucleus; transcription and translation occur in the host cytoplasm.
Example: rotaviruses. After penetration the virion does not disassemble completely, in order to block cellular defences against dsRNA. The (-) strands are transcribed to produce (+)mRNAs, which are translated into viral proteins in the cytoplasm and packed into new virions. Protected inside the virions, the (+) strands replicate to double-stranded viral genomes.
dsRNA viral genome inside the capsid: RNA-dependent RNA polymerase either transcribes the (-) strand into (+) viral mRNA or copies both strands to make a progeny dsRNA genomeRotavirus cycle: partially uncoated virion releases (+)mRNAs for translation, viral proteins and (+)RNAs are assorted and packaged in the virus factory, replicated into cores, then assembled and released via the ER membrane
Cap snatching
note
To be translated, viral mRNAs need caps.
Viruses with (+)RNA generally use their own capping enzymes.
Viruses with (-) or double-stranded RNA acquire caps from cellular mRNAs, by a process called cap snatching.
During cap snatching, cellular mRNAs are cut 10 to 15 nucleotides after the cap, and the resulting capped RNA fragment is used as a primer to initiate transcription of the viral genome.
Cap snatching: influenza vRNPs enter the nucleus, the polymerase complex steals a capped RNA fragment from host RNA polymerase II transcripts to prime viral mRNA synthesis, which is then exported and translated
7. Reverse transcription
note
How it was discovered. The central dogma held that DNA is transcribed to RNA, which is translated into protein. This was challenged in the 1970s, when new enzymes associated with the replication of retroviruses were identified. These enzymes synthesise, on the viral RNA, a complementary DNA (cDNA), which is then capable of integrating into the host genome.
The enzyme. RNA-dependent DNA polymerase is called reverse transcriptase, because in contrast to the DNA-to-RNA flow of the central dogma it transcribes RNA templates into cDNA.
It occurs in many genetic systems. Reverse transcriptases have been identified in viruses, bacteria, animals and plants. The general role of reverse transcription is to synthesise, on RNA sequences, cDNA sequences capable of inserting into different areas of the genome. In this way it contributes to:
propagation of retroviruses;
genetic diversity in eukaryotes, through mobile transposable elements called retrotransposons, which are reverse-transcribed for integration into the genome;
replication of the chromosomal ends, the telomeres, by telomerase reverse transcriptase (TERT);
synthesis of extrachromosomal single-stranded DNAs (msDNA) in bacteria. These are not discussed further, but it is worth knowing that they are related to cell ageing in bacteria.
Retrotransposition: donor DNA transcribed to RNA, reverse-transcribed to transposition DNA, then inserted into target DNAmsDNA biosynthesis: reverse transcriptase primes DNA synthesis from an internal guanosine on a folded RNA, extending a branched DNA-RNA hybrid to form msDNA
The reproductive cycle of retroviruses
note
Binding to the cell membrane;
entry into the cell;
reverse transcription;
transport of the proviral genome into the nucleus;
integration of the proviral genome into the cellular DNA;
transcription of the proviral genome;
translation of the viral mRNA into new viral proteins;
virion assembly inside the cell;
maturation of the immature virion into a completely infectious particle.
Retrovirus reproductive cycle: binding, fusion and entry, reverse transcription to a preintegration complex, nuclear entry and integration as a provirus, transcription, translation, assembly, budding and maturation
Reverse transcriptase: two enzymes in one
note
Reverse transcriptase performs several functions, using two separate active sites at opposite ends of the molecule.
It builds a DNA strand on an RNA template. This reaction is performed in the polymerase active site, formed by two sets of arms that surround the RNA and DNA.
It then removes the original RNA strand, cleaving it into pieces. This is performed by a nuclease active site, located at the opposite end of the enzyme.
Finally it builds a second DNA strand matched to the one just created, forming the final DNA double helix. This is performed again by the polymerase site.
Structure of HIV reverse transcriptase, a heterodimer of subunits p66 and p51, with the polymerase active site and the RNase H active site circled at opposite ends of the molecule
A sloppy enzyme, and why that matters clinically
note
The reverse transcriptase has two subunits, both encoded by the same gene.
The enzyme makes many mistakes, up to about one in every 2,000 nucleotides.
Is that a problem for the virus? In fact, the high level of medical care turns this high error rate into an advantage for the virus: the high mutation rate makes it resistant to drugs and vaccines.
The solution is to combine several drugs. If we do that, the virus has to create a combination of several mutations, each effective against a single drug - and the chance of all of them arising together is small.
Structure of HIV reverse transcriptase showing its two subunits, p66 and p51, both encoded by the same pol gene
The mechanism of reverse transcription in HIV
note
The initial step: a tRNA for lysine acts as a primer. It hybridises to a complementary part of the viral RNA, the primer binding site (PBS).
Reverse transcriptase adds DNA nucleotides onto the 3' end of the primer, synthesising DNA complementary to the R and U regions of the viral RNA. These are non-coding regions, and there are copies of R and U at both ends (U3 and U5).
The RNase H domain of the enzyme degrades the U5 and R regions at the 5' end of the RNA.
The tRNA primer then "jumps" to the 3' end of the viral genome, and the newly synthesised DNA strand hybridises to the complementary R region on the RNA.
The complementary cDNA added in step 2 is further extended.
The majority of the viral RNA is degraded by RNase H, leaving only the PP sequence.
Synthesis of the second DNA strand begins, using the remaining PP fragment of viral RNA as a primer.
The tRNA primer leaves and is ready for a new jump.
The PBS from the second strand hybridises with the complementary PBS on the first strand.
Both strands are extended to form a complete double-stranded DNA copy of the original viral RNA genome, which can then be integrated into the host genome by the enzyme integrase.
HIV genome regions PBS, gag, pol, env, PP, U3, R and U5, tracking early reverse transcription steps: minus-strand synthesis from the PBS primer, RNase H removal of U5 and R at the 5' end, and the tRNA jump to the 3' R regionLater steps of HIV reverse transcription: RNase H degrades viral RNA leaving the PP fragment, second-strand synthesis initiates from PP, and the two PBS regions hybridise to complete double-stranded proviral DNA
8. Telomere replication
note
The problem. Replication of the leading strand continues until the end of the chromosome is reached. On the lagging strand, however, DNA polymerase cannot copy the very end: when the last primer is removed there is no free OH group, which the polymerase requires.
End replication problem: the newly made lagging-strand copy, red, is shorter than the parental strand, orange, at the chromosome's 3' end after the last primer is removed
The consequence. The chromosomal ends remain unpaired, because the lagging strand stays shorter. Over time these ends get progressively shorter as cells continue to divide - the telomeres are shortened with each round of DNA replication, that is, with each cell division.
The solution. The function of the enzyme telomerase is to maintain and elongate the chromosome ends. Telomerase works as a reverse transcriptase: in its active centre it carries its own RNA molecule, used as the template for telomere elongation.
Once the lagging DNA strand has been sufficiently elongated, an RNA primer is made and DNA polymerase can add nucleotides to finish replication of that strand.
How it runs, step by step. Telomerase has an associated RNA that complements the 3' overhang at the end of the chromosome. That RNA template is used to synthesise the complementary strand. The telomerase then shifts along and the process is repeated, adding repeat after repeat. Finally primase and DNA polymerase synthesise the complementary strand.
Telomerase mechanism: its RNA template pairs with the chromosome's 3' overhang, synthesises a complementary repeat, shifts along and repeats, after which primase and DNA polymerase finish the lagging strand
Telomeres and ageing
note
The gene for telomerase is not active in adult somatic cells. In those cells the telomeres therefore become shorter and shorter with each division. When the telomeres become too short, the cell is not allowed to initiate replication and mitosis, and the cell is considered aged.
Normally the gene for telomerase is active only in embryonic cells and adult stem cells. Some types of cancer cell illegally reactivate that gene, which is one of the ways a tumour escapes the limit on how many times a cell may divide.
Telomerase enzyme with its own RNA template pairing to the chromosome's single-stranded DNA end and adding a nucleotide to extend it
Palm, fingers and thumb
note
A closing structural observation that ties the three enzymes together.
RNA replicase, RNA polymerase and reverse transcriptase all share the same architecture, conventionally described as palm, fingers and thumb. Three enzymes doing three different jobs on three different template-product combinations have converged on, or inherited, one hand-shaped fold.
Palm-fingers-thumb fold, shown as a hand cartoon beside the ribbon structure, the shared architecture of RNA replicase, RNA polymerase and reverse transcriptase
The most important things to know
note
RNA replication was found because polio and mengovirus ignored actinomycin D, which meant they were copying RNA without DNA.
RdRP is encoded by the viral genome and is used by all RNA viruses except the retroviruses. Vertebrates have lost their own version.
(+) sense = is mRNA = translated immediately = no replicase needed in the virion.(-) sense and dsRNA = cannot be translated = replicase must be carried in the virion.
Nonsegmented (-)RNA replicates in the cytoplasm; segmented (-)RNA replicates in the nucleus, one mRNA per segment.
The replication/transcription switch is set by the concentration of capsid proteins.
Influenza's 8 segments allow reassortment between strains, and that is what causes pandemics.
(+)RNA viruses hide their double-stranded intermediate in virus factories made from invaginated host membranes; dsRNA viruses hide it inside the capsid and never fully uncoat.
Cap snatching: (-) and dsRNA viruses steal 10 to 15 nucleotides of capped cellular mRNA to prime their own transcription.
Reverse transcriptase is two enzymes in one: a polymerase site at one end and a nuclease (RNase H) site at the other. It is error-prone at about 1 in 2,000, which is why HIV is treated with combinations of drugs.
Reverse transcription is not only viral. It also drives retrotransposons, telomerase, and bacterial msDNA.
Telomerase is a reverse transcriptase carrying its own RNA template. It is off in adult somatic cells, on in embryonic and stem cells, and illegally switched back on in some cancers.
Palm-fingers-thumb fold shown as a hand cartoon beside a ribbon structure, the shared architecture uniting RNA replicase, RNA polymerase and reverse transcriptase
All 105 Medical Biology topics · exam study support, not clinical guidance. SuperMed is not affiliated with the Medical University of Sofia.
Start free. No card needed.
Every account starts free, with free topics in Cytology and Medical Biology. Super opens the rest. Super is €15 a month.