Medical Biology · Year 1 · Medical University of Sofia
06
RNA processing
Free notes for topic 06 of the Medical Biology syllabus, open without an account. Written by a senior student against the syllabus question and checked line by line by a second student before publishing. How content is made
Updated
What this topic covers
note
The maturation of eukaryotic mRNAs, including the 5' cap, the 3' poly(A) tail and mixed tails, the connection between cap and tail, splicing and the action of the spliceosome, and alternative splicing; the maturation of eukaryotic rRNAs, tRNAs and microRNAs; and, by contrast, processing in prokaryotes.
Maturation of eukaryotic mRNA: hnRNA gains a 5' cap and poly(A) tail, then splicing removes the two introns (pink arcs) to give mature spliced mRNA with cap and poly(A) tail
1. Maturation of eukaryotic mRNA
note
The primary transcripts of RNA polymerase II are called heterogeneous nuclear RNA (hnRNA).
An hnRNA is not yet a usable message. Before it can leave the nucleus it acquires a cap at one end, a tail at the other, and loses its introns. Each of those three changes is also a licence: a transcript that has not had them done is not allowed out.
Primary mRNA transcript: 5' cap, exons 1 to 5 (introns between them not shown as boxes) and the 3' poly(A) tail, AAAA
The 5' cap
note
The 5' cap is a modified nucleotide that contains G.
The cap is added by the enzyme guanylyl transferase, which forms an unusual 5'-5' linkage with the terminal 5' nucleotide. The G in the cap is then methylated by a methyl transferase, and the two closest riboses can be methylated too.
The functions of the cap:
it ensures the stability of the mRNA during translation;
it is required as a signal allowing transfer through the nuclear pores into the cytoplasm;
it gives the signal for the splicing of the first exons;
it is recognised by the small ribosomal subunit and indicates the direction for translation.
Chemical structure of the 5' cap: 7-methylguanosine (m7G, with 7-CH3) joined by an unusual 5'-to-5' triphosphate bridge to the first transcribed nucleotide
The 3' poly(A) tail
note
At the end of protein genes there is a consensus sequence AATAAA in the DNA, which appears as AAUAAA in the transcript. It is the signal for the end of transcription.
A protein complex recognises AAUAAA in the growing mRNA.
The complex cuts the chain 10 to 30 bases after AAUAAA.
The enzyme poly(A) polymerase (PAP) joins the complex and adds 100 to 200 A nucleotides to the 3' end.
Cleavage of the pre-mRNA: CPSF binds the AAUAAA signal and CstF binds the GU or U-rich element downstream, CFI and CFII cut the chain as RNA polymerase continues transcribing the DNA
In eukaryotes all mRNAs are polyadenylated at the 3' end, with histone mRNA as the exception.
The functions of the tail mirror those of the cap:
it ensures the stability of the mRNA during translation;
it is a signal that allows transfer through the nuclear pores to the cytoplasm;
it gives the signal for the splicing of the last exons.
Addition of the poly(A) tail: after cleavage, CPSF, CstF, CFI and CFII remain bound while PAP (poly(A) polymerase) synthesises the poly(A) tail on the cut 3' end
Communication between the 5' cap and the 3' poly(A) tail
note
Communication between the two ends is necessary for the initiation of translation.
The poly(A)-binding protein (PABP) interacts with the translation initiation factor eIF4G, which in turn interacts with a cap-binding protein. The effect is to bring the 5' and 3' ends of the mRNA together, so a message being translated is effectively a circle.
This is also why the cap and the tail both report on stability: a molecule that has lost either end can no longer close the loop.
Circularisation of an mRNA being translated: PABP on the poly(A) tail binds eIF4G, which binds the cap-binding protein eIF4E at the 5' cap, bringing the cap and tail together near the AUG start codon
Mixed tails
note
Some human poly(A) tails do not contain only A. These are called mixed tails.
The non-A nucleotide, most frequently G, is found in the last or next-to-last nucleotide of the tail.
The role of mixed tails is to protect the mRNA tail from rapid shortening, that is, from rapid deadenylation - and since deadenylation is the usual first step of degradation, a single G at the end buys the message time.
Mixed poly(A) tail cartoon: mRNA ending in A, A, A, G nucleotides, with the terminal G nucleotide holding off a nuclease that would otherwise degrade the tail
2. Splicing
note
Most protein genes contain several introns, and in total the introns are longer than the exons.
Splicing is the process of removing the introns from the pre-mRNA and connecting the exons. It occurs in the nucleus, either during transcription or immediately after it.
According to the mechanism, introns are subdivided into two groups.
Self-splicing introns, removed by ribozymes formed within their own RNA sequences. These ribozymes have a specific conformation based on double-stranded RNA loops.
Introns removed by a complex of small nuclear ribonucleoproteins (snRNPs). That complex is the spliceosome.
Secondary structure (panels A, B) and 3D tertiary structure (panels C, D) of a self-splicing group II intron ribozyme, with its domains IA, IB, IC, ID1, ID2, I(i), I(ii), II to VI and the catalytic triad coloured separately
The mechanism of self-splicing
note
The chemistry is the same in outline for both groups, and it is worth following once.
A certain adenine nucleotide within the intron attacks the splicing site and cuts the sugar-phosphate backbone of the RNA.
The cut-off 5' end of the intron remains covalently bound to that adenine, which creates a loop in the RNA molecule: the lariat.
The released 3'-OH end of the exon then reacts with the next exon. The two exons connect, and at that point they are released from the intron, which still has the shape of a lariat.
The detached intron quickly degrades.
Mechanism of self-splicing: the branch-point adenine (A) in the intron attacks the 5' exon-intron junction, forming a lariat; the freed 3'-OH of the 5' exon then attacks the 3' splice site, joining the exons and releasing the lariat intron
The sequences the spliceosome recognises
note
Spliceosomes recognise specific nucleotide sequences at the borders between exons and introns.
For a vertebrate intron, the consensus border is: the intron begins with GUR and ends in a polypyrimidine tract followed by YAG. These are not to be learned.
Human splice sites are best represented as consensus logos, in which the size of each letter shows the probability of that variant at that position. The figures are usually drawn as DNA sequences; in the mRNA, T is replaced by U.
Consensus intron borders: exon 1 ends AG, the intron begins GURAGU (5' splice junction, donor site), contains the branch point A within YUNAY and a polypyrimidine tract, and ends YAG before exon 2 (3' splice junction, acceptor site)
The action of the spliceosome
note
The spliceosome members recognise the intron-exon borders on the pre-mRNA, bringing the two ends of the intron together.
The branch-point site, the reactive A, is first recognised by the branch-point binding protein (BBP) and a helper protein (U2AF). Their names are not to be learned.
The U2 snRNP displaces BBP and U2AF and forms base pairs with the branch-point consensus sequence, while the U1 snRNP forms base pairs with the 5' splice junction.
At this point the U4/U6/U5 triple snRNP is formed. Within this triple complex the snRNAs U4, U5 and U6 facilitate lariat formation and the splicing reaction.
Spliceosome assembly: BBP and U2AF bind the intron near the branch-point A, then U1 and U2 snRNP replace them, the U4/U6.U5 tri-snRNP joins to form a loop, and lariat formation with 5' splice site cleavage releases U1 and U4, leaving U6 snRNP holding the lariat
The specificity of the mRNA-snRNP connection is based on complementary interaction between the splice sites on the pre-mRNA and the snRNAs. In other words, the spliceosome finds its target by base pairing, exactly as everything else in this course does.
Only "export-ready" mRNA is allowed into the cytoplasm
note
Not all synthesised mRNAs are allowed to enter the cytosol. Aberrantly spliced pre-mRNAs, broken RNAs, and mRNAs without caps and tails - histone mRNAs excepted - are not merely useless but potentially dangerous, and they must be destroyed in the nucleus.
How does the cell distinguish healthy mature mRNA from the overwhelming amount of debris produced by RNA processing? The answer is that transport of mRNA from the nucleus to the cytoplasm is highly selective. It is controlled by the nuclear pore complex, which recognises certain marker proteins attached to the mRNA. Some of these proteins are deposited at exon-exon boundaries after successful splicing - so the proof that splicing worked is physically written onto the molecule.
Some of those proteins remain in the nucleus, whereas others travel with the mRNA into the cytoplasm; without them the mRNA would be very unstable there. Once in the cytoplasm, the mRNA sheds its previously bound proteins and acquires new ones, for example the initiation factors for translation.
Export-ready mRNA in the nucleus, coated with hnRNP proteins, poly-A-binding proteins, SR proteins and the cap-binding complex (CBC), passes through the nuclear pore complex while nuclear-restricted proteins are stripped off, then eIF-4G and eIF-4E bind in the cytosol before translation
The exome
note
The exome is the part of the genome formed by exons.
The exome of the human genome consists of roughly 180,000 exons, constituting about 1% of the total genome.
Though it comprises a very small fraction of the genome, the exome is thought to harbour 85% of disease-causing mutations. That ratio is the entire justification for exome sequencing as a diagnostic method.
DNA with alternating introns and exons above a diagram showing introns looped out and exons joined; below, a whole apple representing the whole genome next to a single slice labelled exome equals 1%
3. Alternative splicing
note
Very frequently, one exon encodes one protein domain with a particular function and specific contact sites.
During mRNA processing, some of the exons may be removed together with the introns. That is called alternative splicing.
The consequence is large: by changing the splicing, the cell can play with the protein's activities, creating different combinations of contact sites from a single gene.
Alternative splicing of a five-exon gene: the same DNA and RNA are spliced into three different mRNAs (combinations of exons 1,2,3,4,5 or 1,2,4,5 or 1,2,3,5) that translate into differently folded proteins A, B and C
The regulation of alternative splicing
note
In some cases alternative splicing occurs simply because the spliceosome is unable to distinguish cleanly between two or more alternative splice sites, so that different choices are made by chance. The result is that several versions of the protein are made in all cells in which the gene is expressed. This is constitutive alternative splicing.
In many cases, however, alternative splicing is strictly regulated. Regulated splicing is used, for example, to switch from the production of a non-functional protein to the production of a functional one. Regulation can also generate different versions of a protein in different cell types, according to the needs of the tissue and the differentiation stage of the cell.
There are two directions of control:
Negative control - a repressor protein binds to the pre-mRNA in one tissue, thereby preventing the splicing machinery from removing an intron sequence.
Positive control - the splicing machinery is unable to remove a particular intron without assistance from an activator protein.
Negative and positive control of alternative splicing: in tissue 1 a primary transcript splices normally, but in tissue 2 a repressor blocks splicing of a yellow exon (negative control) or an activator enables splicing that would not otherwise occur (positive control)
Example: alternative splicing in different cell types
note
Specific forms of tropomyosin are produced in different types of cell, all from the same gene.
Cell-type-specific forms of many other proteins are produced in the same way.
Tropomyosin gene alternative splicing: the primary pre-mRNA transcript with exons 1a to 9d is spliced into different tissue-specific mRNAs in striated muscle, smooth muscle, brain (TMBr-1, TMBr-2, TMBr-3) and fibroblast (TM-2, TM-3, TM-5a, TM-5b)
Example: alternative splicing at different stages of cell differentiation
note
IgM exists in two forms: secreted and membrane-bound, the latter with a receptor function.
They differ only in the C-terminus. The secreted protein has a hydrophilic terminus motif, while the membrane-bound IgM has a C-terminal hydrophobic transmembrane domain, the anchor region.
The transition from membrane to secreted IgM is achieved by alternative splicing that ignores the exons encoding the anchor region. This happens after contact between the antigen and the membrane IgM receptor, which triggers the next steps of B cell differentiation.
The same gene therefore makes the sensor and the weapon, and which one is made is decided by whether the sensor has fired.
IgM heavy chain gene with exons 3 to 6 and two poly(A) sites: splicing that includes exon 6 (with the transmembrane anchor) gives membrane-bound IgM on the cell surface, while splicing that stops after exon 4 gives secreted IgM
When splicing goes wrong: β thalassaemia
note
Abnormal processing of the β-globin pre-mRNA leads to β thalassaemia.
Normal: the pre-mRNA has three intact exons and is translated into normal β-globin.
A mutation in exon 1 creates an additional splice site. The spliceosome cannot distinguish between the original splice site and the new one, so the result can be either correctly spliced or aberrantly spliced mRNA, translated into normal or abnormal β-globin respectively.
In other cases a new splice site is created by a mutation in intron 1, or by a mutation in intron 2.
HBB gene with exon 1, IVS1, exon 2, IVS2, exon 3: normal splicing gives the correct mRNA (A), while a splicing mutation in exon 1 (B), intron 1 (C) or intron 2 (D) each produces an aberrantly spliced mRNA missing or retaining part of that region
The lesson is that a mutation does not have to be in a coding sequence to destroy a protein. It only has to be somewhere the spliceosome reads.
4. Maturation of eukaryotic rRNAs
note
The eukaryotic rRNAs are 28S, 18S, 5.8S and 5S.
RNA polymerase I transcribes the genes for 28S, 18S and 5.8S rRNA.
RNA polymerase III transcribes the genes for 5S rRNA.
The genes for 28S, 18S and 5.8S are present in multiple copies, several hundred in higher eukaryotes, with non-transcribed spacer regions between the transcription units for 45S RNA.
The steps of rRNA processing:
28S, 18S and 5.8S are synthesised together as a single 45S chain.
Some nucleotides are methylated, and some ribosomal proteins bind to the long 45S chain.
The long 45S transcript is cut into three separate rRNA molecules.
Non-coding parts are removed.
Other proteins and the 5S rRNA join the complex, and the ribosomal subunits are assembled.
Processing of the 45S pre-rRNA: the single transcript containing 18S, 5.8S and 28S segments undergoes nucleotide modification (methyl groups and pseudouridine), then cleavage into the three separate mature 18S, 5.8S and 28S rRNAs
The Christmas tree
note
In each cluster, multiple copies of the 45S pre-rRNA are synthesised simultaneously. Between and within the transcription units lie the external transcribed spacers (ETS) and the internal transcribed spacers (ITS1 and ITS2).
Under the electron microscope an actively transcribed rDNA locus has a characteristic Christmas tree appearance:
the chromatin is the trunk;
the closely packed rRNA transcripts are the branches;
at the tip of each transcript is a terminal ball, a complex of rRNA with the proteins involved in processing.
Christmas tree diagram: an rDNA repeat unit with 5'ETS, 18S, ITS1, 5.8S, ITS2, 25/28S and 3'ETS regions, and below it multiple pre-rRNA transcripts of increasing length branching from the chromatin axis, each ending in a terminal ballElectron micrograph of an actively transcribed rDNA locus showing the Christmas-tree pattern of RNA transcripts of increasing length branching from the DNA molecule
5. Maturation of eukaryotic tRNAs
note
Three things happen to a tRNA that do not happen to other transcripts.
The amino-acid-accepting end, CCA, is missing from the tRNA genes. An enzyme, tRNA nucleotidyl transferase, adds it after transcription.
Some tRNAs contain an intron located in the anticodon arm.
There is complex post-transcriptional modification: in mature tRNAs, up to 10% of the nucleosides are chemically modified.
tRNA maturation: RNase P removes the 5' leader and RNase Z the 3' trailer from the pre-tRNA, processing and modification give the mature cloverleaf tRNA with acceptor stem, D-stem and loop, T-stem and loop, V-loop and anticodon stem and loop, and aminoacylation attaches the amino acid
Processing of human tRNAs
note
The human genome contains more than 500 tRNA-encoding genes, although a significant number of them are either not expressed at all, or are expressed only in particular tissues or cells.
Transcription of the functional tRNA genes by RNA polymerase III generates precursor tRNAs, which then undergo:
trimming of the leader and trailer sequences, done by the endonucleases RNase P (leader) and RNase Z (trailer);
addition of CCA residues at the 3' end;
removal of the intron;
base modifications.
The mature tRNAs are then exported to the cytosol by specific transporters.
Human pre-tRNA processing detail: RNase P cuts the 5' leader and endo/exonucleases the 3' trailer, nucleotidyl transferase adds the CCA and discriminator nucleotide, and specific positions are modified (2'-OMe, m2 2G26, inosine at 34, i6A37, m1G, m7G46, T and pseudouridine at 54 and 55, m1A58)
6. Processing of microRNAs
note
miRNAs are synthesised as double-stranded molecules. Most of the molecule is removed - this is the passenger strand - and the rest is the functional guide strand.
Some microRNAs have their own genes. Some miRNA genes are located in the introns of protein genes; these are called mirtrons.
The pathway:
The primary miRNAs consist of one or more stem-loop structures, with a common 5' cap and poly(A) tail.
A microprocessor protein complex in the nucleus cuts the primary miRNAs into precursor miRNAs (pre-miRNAs), which are single hairpin structures.
The pre-miRNAs are transported into the cytosol, where they are further processed by the RNase Dicer into mature miRNAs, 18 to 22 nucleotides long.
One strand of the mature miRNA enters the RNA-induced silencing complex (RISC) and binds to the 3' untranslated region of the target mRNA.
miRNA pathway: RNA polymerase II transcribes the primary microRNA transcript with stem-loops and a cap and poly(A) tail, the Drosha-DGCR8 microprocessor complex cuts it into pre-miRNAs, Exportin-5 exports them to the cytosol, Dicer cuts the miRNA duplex, and one strand loads into the Ago-RISC complex to inhibit translation or destabilise the target mRNA
The outcomes are abbreviated RNAi for RNA interference, meaning repression, and RNAa for RNA activation.
7. Processing in prokaryotes
note
The contrast is instructive, because it shows how much of the above exists only because of the nuclear envelope.
mRNA. The prokaryotic mRNA is ready for translation without any processing at all.
rRNA. The three rRNAs, 16S, 23S and 5S, are copied as one primary transcript, which is cut into three chains. These chains need additional shortening of the ends and methylation.
tRNA. Some prokaryotic tRNAs are produced as a common primary transcript and need processing into individual molecules. The redundant parts are cut by endo- or exoribonucleases (RNases). These enzymes recognise either specific RNA sequences or specific conformations, and some of them use a short RNA as a cofactor.
Generic pre-tRNA processing pathway: RNase P and RNase Z trim the leader and trailer and processing modifies the molecule into a mature cloverleaf tRNA before aminoacylation, illustrating the same kind of endonuclease trimming used on prokaryotic tRNA precursors
The most important things to know
note
The primary transcript of RNA polymerase II is hnRNA. It needs a cap, a tail and splicing before it is allowed out of the nucleus.
The cap is a methylated G joined by an unusual 5'-5' linkage; the tail is 100 to 200 A residues added by poly(A) polymerase after cleavage 10 to 30 bases past AAUAAA. Histone mRNA is the exception to polyadenylation.
Cap and tail have the same three jobs: stability, export licence, and a splicing signal for the first and last exons respectively - and they are physically joined through PABP and eIF4G.
Splicing produces a lariat through attack by a branch-point adenine. Self-splicing introns are ribozymes; the rest need the spliceosome, which works by base pairing between snRNAs and the splice sites.
Export is selective: marker proteins deposited at exon-exon boundaries are the receipt for successful splicing.
The exome is 1% of the genome and holds 85% of disease-causing mutations.
Alternative splicing can be constitutive (by chance) or regulated (repressor or activator). Tropomyosin varies by cell type; IgM switches from membrane-bound to secreted by dropping the anchor exons.
β thalassaemia shows that a mutation in an intron can be as destructive as one in an exon.
rRNA: 45S is cut into 28S, 18S and 5.8S; 5S comes separately from polymerase III.
tRNA: CCA is not in the gene and is added afterwards; up to 10% of nucleosides are modified.
miRNA: primary, then pre-miRNA by the microprocessor, then mature 18 to 22 nt by Dicer, then one strand into RISC.
Prokaryotic mRNA needs no processing at all.
Maturation of eukaryotic mRNA: hnRNA gains a 5' cap and poly(A) tail, then splicing removes the introns to give the mature spliced mRNA with cap and poly(A) tail
All 105 Medical Biology topics · exam study support, not clinical guidance. SuperMed is not affiliated with the Medical University of Sofia.
Start free. No card needed.
Every account starts free, with free topics in Cytology and Medical Biology. Super opens the rest. Super is €15 a month.