RNA Synthesis, Processing & Modification
← Back πŸ“‹ Q-Bank 🏠 All Units
HIGH YIELD ⭐
Molecular Biology Β· Unit 24 of 26

RNA Synthesis, Processing & Modification

TMU Lecture 22 β€” RNA Harper's ch. 36 β€” RNA Synthesis, Processing & Modification RNA structure from Harper's ch. 34 β€” promoter was set in BOTH papers
01

RNA versus DNA β˜…β˜…β˜…

DNARNA
Sugar2β€²-deoxyriboseRibose
PyrimidinesCytosine, thymineCytosine, uracil
StrandednessDouble-strandedUsually single-stranded
Base ratiosA = T and G = C (Chargaff)Its guanine content does not necessarily equal its cytosine content, nor does its adenine content necessarily equal its uracil content
AlkaliStableCan be hydrolyzed by alkali to 2β€²,3β€² cyclic diesters of the mononucleotides β€” useful both diagnostically and analytically
CatalysisNoneSome RNA molecules have intrinsic catalytic activity β€” RIBOZYMES
Why one extra hydroxyl group changes everything

The single structural difference β€” a 2β€²-OH on the ribose β€” accounts for almost every entry in that table, and it is worth tracing.

Alkaline hydrolysis: the 2β€²-OH sits right beside the phosphodiester bond and, when deprotonated, attacks it β€” forming the 2β€²,3β€² cyclic diester Harper's names. DNA, lacking that hydroxyl, cannot do this. RNA is chemically self-destructive; DNA is not.

Which is why DNA is the archive and RNA the working copy. A molecule that must survive for the lifetime of an organism cannot carry a built-in cleavage mechanism; a message that must be made, used and disposed of benefits from one.

Catalysis: the same reactive hydroxyl, plus the single-stranded chain's freedom to fold, is what lets RNA act as an enzyme at all. Two ribozymes are the peptidyl transferase that catalyzes peptide bond formation on the ribosome, and ribozymes involved in RNA splicing β€” the two most fundamental reactions in this module are both catalysed by RNA, not protein.

Why thymine in DNA and uracil in RNA

Thymine is simply 5-methyluracil. The methyl group is a tag marking the base as belonging in DNA.

It matters because cytosine spontaneously deaminates to uracil. If uracil were a normal DNA base, that damage would be invisible; because it is not, a repair glycosylase can recognise every uracil in DNA as an error and excise it. The methyl group is the difference between a repairable lesion and a silent mutation.

Test yourself
  • Sugar difference? → 2β€²-deoxyribose in DNA, ribose in RNA
  • Base difference? → Thymine in DNA, uracil in RNA
  • Does Chargaff's rule apply to RNA? → No β€” G need not equal C, nor A equal U
  • What does alkali do to RNA? → Hydrolyses it to 2β€²,3β€² cyclic diesters of the mononucleotides
  • What is a ribozyme? → An RNA molecule with intrinsic catalytic activity β€” e.g. peptidyl transferase
02

The classes of RNA β˜…β˜…β˜…

ClassFunction
Messenger RNA (mRNA)Carries the coding sequence from gene to ribosome; the least stable and most heterogeneous class
Transfer RNA (tRNA)The adapter molecule β€” recognises a codon through its anticodon and carries the corresponding amino acid
Ribosomal RNA (rRNA)The structural and catalytic core of the ribosome; the most abundant class
Small nuclear RNAs (snRNAs)rRNA and mRNA processing and gene regulation β€” the components of the spliceosome
Micro-RNAs (miRNAs)Silence mRNAs by annealing to their 3β€² untranslated regions
Small interfering RNAs (siRNAs)RNA interference; protect the host from RNA viruses
How the four classes differ

Harper's compares them on four axes: abundance, size, function and general stability.

rRNA is the most abundant and the most stable; mRNA is the least stable, which is exactly what a regulatory message should be β€” a signal that persisted indefinitely could not be switched off.

The structure of tRNA

The cloverleaf, with four arms:

The acceptor arm β€” the 3β€²-CCA-OH terminus, the site of attachment of the specific amino acid.
The anticodon arm β€” consists of seven nucleotides and recognises the three-letter codon in mRNA.
The TψC arm β€” involved in binding of the aminoacyl-tRNA to the ribosomal surface at the site of protein synthesis.
The D arm β€” one of the sites important for the proper recognition of a given tRNA species by its proper aminoacyl-tRNA synthetase.

tRNA is rich in nontraditional nucleotides introduced by post-transcriptional modification β€” methylated guanosines, pseudouridine (ψ), inosine and others.

Test yourself
  • Name the four principal RNA classes → mRNA, tRNA, rRNA and the small RNAs
  • Which is most abundant? → rRNA. Least stable? → mRNA
  • Where does the amino acid attach to tRNA? → The 3β€²-CCA-OH of the acceptor arm
  • How many nucleotides in the anticodon arm? → Seven, of which three are the anticodon
  • What do snRNAs do? → rRNA and mRNA processing and gene regulation β€” the spliceosome
Four genes on one stretch of DNA β€” note that the template strand is NOT the same strand for every gene
Four genes on one stretch of DNA β€” note that the template strand is NOT the same strand for every gene
Harper's Illustrated Biochemistry, Figure 36–1, p.395
03

The central dogma β˜…β˜…

Transcription

RNA biosynthesis from a DNA template is called transcription. Its products are mRNA, tRNA and rRNA.

The synthesis of an RNA molecule from DNA is a complex process involving one of the group of RNA polymerase enzymes and a number of associated proteins. The general steps required to synthesize the primary transcript are initiation, elongation and termination.

ProkaryotesEukaryotes
The primary transcriptEquivalent to the mRNA moleculeA precursor (pre-mRNA) to the mRNA
ProcessingEssentially none for mRNAModified at both ends, and introns are removed; occurs primarily within the nucleus
CouplingTranslation can begin before transcription is completeSeparated in space β€” after processing, the mRNA is exported to the cytoplasm for translation
PolymerasesOne RNA polymeraseThree distinct nuclear polymerases
Why eukaryotes gained a nucleus and lost coupled translation

In a bacterium, ribosomes attach to the 5β€² end of an mRNA while its 3β€² end is still being transcribed. That is fast and economical β€” but it makes splicing impossible. You cannot cut an intron out of a message that is already being read.

The nuclear membrane is what buys the time. By separating transcription from translation in space, it creates a compartment in which the transcript can be capped, polyadenylated, spliced and inspected before any ribosome sees it. Everything in Β§8 and Β§9 depends on that separation.

The cost is speed; the return is alternative splicing, and hence many proteins from one gene. Errors or changes in synthesis, processing, splicing, stability or function of mRNA transcripts are a cause of disease β€” a whole class of pathology that prokaryotes simply cannot have.

Test yourself
  • Define transcription → RNA biosynthesis from a DNA template
  • In prokaryotes, what is the primary transcript equivalent to? → The mRNA itself
  • In eukaryotes? → Pre-mRNA, a precursor requiring processing
  • Where does eukaryotic processing occur? → Primarily within the nucleus
04

The promoter and the transcription unit ⭐

Set in BOTH past papers β€” Section I
Promoter
2019 paper, Section I
Promoter
2020/21 paper, Section I
Promoter

A promoter is the DNA sequence to which RNA polymerase binds to initiate transcription of a gene.

DNA-dependent RNA polymerase attaches at this specific site on the template strand. This is followed by initiation of RNA synthesis at the starting point, and the process continues until a termination sequence is reached.

The promoter determines two things: where transcription is to commence along the DNA, and how frequently this event is to occur.

The transcription unit

A transcription unit is the region of DNA that includes the signals for transcription initiation, elongation and termination.

Position +1 is the transcript initiation site.
Upstream sequences (negative numbers) β€” the promoter.
Downstream sequences β€” introns and exons.
The RNA product is the primary transcript.

Why a polymerase needs a promoter at all

Consider the search problem. E. coli has 4 Γ— 10Β³ transcription initiation sites in 4.2 Γ— 10⁢ base pairs of DNA, and humans have about 10⁡ promoters in 3 Γ— 10⁹ base pairs. The polymerase must find a few thousand specific addresses among millions of possible ones.

It does so by a strategy worth stating in an exam: RNA polymerase can bind, with low affinity, to many regions of DNA, but it scans the DNA sequence β€” at a rate of β‰₯10Β³ bp per second β€” until it recognizes certain specific regions to which it binds with higher affinity.

So a promoter is not a lock that only one key opens; it is a region of unusually high binding affinity in a sequence the polymerase is already sliding along. And because binding affinity is a continuous quantity, a stronger promoter is transcribed more often β€” which is how the same sequence element sets both where and how much.

Test yourself
  • Define a promoter → The DNA sequence to which RNA polymerase binds to initiate transcription of a gene
  • Define a transcription unit → The region of DNA that includes the signals for initiation, elongation and termination
  • What is position +1? → The transcript initiation site
  • The two things a promoter determines? → Where transcription commences, and how frequently
  • How does polymerase find a promoter? → It binds DNA weakly and scans at β‰₯10Β³ bp/s until it meets a higher-affinity region
05

RNA polymerase β˜…β˜…β˜…

DNA-dependent RNA polymerase

The enzyme responsible for the polymerization of ribonucleotides into a sequence complementary to the template strand of the gene.

Four features distinguish it from DNA polymerase, and all four are examinable:

1 Β· It adheres to Watson-Crick base-pairing rules, using ATP, GTP, CTP and UTP β€” U replacing T.
2 Β· It synthesises with 5β€²β†’3β€² polarity, reading the template strand in the 3β€²β†’5β€² direction.
3 Β· A primer is NOT involved in RNA synthesis, as RNA polymerases have the ability to initiate synthesis de novo.
4 Β· Initiation requires large, multicomponent initiation complexes.

The bacterial enzyme

Core enzyme: Ξ±β‚‚Ξ²Ξ²β€². Holoenzyme: Ξ±β‚‚Ξ²Ξ²β€²Οƒ.

Functions of the subunits:
Ξ± β€” assembly of the tetrameric core
Ξ² β€” ribonucleoside triphosphate binding site
Ξ²β€² β€” DNA template binding region
Οƒ β€” helps the core enzyme recognize and bind to the promoter region

The transcription β€œbubble” is 20 bp of DNA, and the entire complex covers 30–75 bp.

The three eukaryotic nuclear polymerases

Mammalian cells possess three distinct nuclear DNA-dependent RNA polymerases.

Pol I β€” most rRNA
Pol II β€” mRNA and most snRNAs and miRNAs
Pol III β€” tRNA and 5S rRNA

Ξ±-Amanitin is a specific differential inhibitor of the eukaryotic nuclear DNA-dependent RNA polymerases and as such has proved to be a powerful research tool β€” it blocks the translocation of RNA polymerase during phosphodiester bond formation. It distinguishes the three because they differ in sensitivity to it.

Why Οƒ is separable, and why that is the point

Notice the division of labour in the bacterial enzyme. The core can polymerise but cannot find a promoter; Οƒ helps the core enzyme recognize and bind to the promoter region and is then released.

Making promoter recognition a detachable function is what makes bacterial gene regulation possible. Swap one Οƒ factor for another and the same core polymerase transcribes an entirely different set of genes β€” heat-shock genes, sporulation genes, and so on. One catalytic machine, many programmes.

The eukaryotic solution to the same problem is different in form but identical in logic: instead of interchangeable Οƒ factors, all eukaryotic RNA polymerase forms require other proteins known as general transcription factors (GTFs), and it is these β€” not the polymerase β€” that recognise the promoter.

Test yourself
  • Does RNA synthesis need a primer? → No β€” RNA polymerases initiate de novo
  • In which direction is the template read? → 3β€²β†’5β€², while RNA is made 5β€²β†’3β€²
  • Bacterial core enzyme vs holoenzyme? → Ξ±β‚‚Ξ²Ξ²β€² vs Ξ±β‚‚Ξ²Ξ²β€²Οƒ
  • What does Οƒ do? → Helps the core recognise and bind the promoter
  • Size of the transcription bubble? → 20 bp; the whole complex covers 30–75 bp
  • Which polymerase makes mRNA? → Pol II
  • What is Ξ±-amanitin? → A differential inhibitor of the eukaryotic nuclear polymerases, blocking translocation
The bacterial RNA polymerase complex on DNA: Ξ±β‚‚Ξ²Ξ²β€² core plus the Οƒ subunit that finds the promoter, with the nascent 5β€²-PPP transcript emerging
The bacterial RNA polymerase complex on DNA: Ξ±β‚‚Ξ²Ξ²β€² core plus the Οƒ subunit that finds the promoter, with the nascent 5β€²-PPP transcript emerging
Harper's Illustrated Biochemistry, Figure 36–2, p.395
06

Initiation, elongation, termination β˜…β˜…β˜…

The three stages

Initiation: RNA polymerase binds to the promoter of DNA, and then a transcription β€œbubble” is formed. The holoenzyme must bind DNA and locate a promoter; then comes localized unwinding of the two strands by RNA polymerase to provide a single-stranded template, and formation of phosphodiester bonds between the first few ribonucleotides in the nascent RNA chain. The unwound complex is the preinitiation complex (PIC).

Elongation: the polymerase catalyzes formation of 3β€²,5β€²-phosphodiester bonds in the 5β€²β†’3β€² direction, using NTPs as building units. The nascent chain is attached to the polymerization site on the Ξ² subunit.

Termination: when the polymerase reaches a termination sequence on DNA, the reaction stops and the newly synthesized RNA is released.

Promoter clearance

RNA polymerase continues to incorporate nucleotides 3 to ~10, at which point the polymerase undergoes another conformational change and moves away from the promoter; this reaction is termed PROMOTER CLEARANCE.

Until it happens, the polymerase repeatedly makes and releases very short abortive transcripts. Promoter clearance is therefore the commitment step of transcription.

Two mechanisms of bacterial termination

Intrinsic terminators. The predominant bacterial transcription termination signal contains an inverted, hyphenated repeat followed by a stretch of AT base pairs. The inverted repeat, when transcribed into RNA, generates a secondary structure β€” an RNA hairpin β€” which causes RNA polymerase to pause. The weak rU:dA hybrid that follows then lets the transcript fall off. About 50% of genes use an inverted palindrome plus poly-A.

Rho-dependent termination. Rho is an ATP-dependent, RNA-stimulated helicase that disrupts the ternary transcription elongation complex composed of RNA polymerase, nascent RNA and DNA. It interacts with the paused polymerase and induces chain termination.

Why a hairpin can stop an enzyme

It looks improbable that a fold in the product could halt the machine making it, but the geometry makes it inevitable.

The nascent RNA emerges from an exit channel in the polymerase. An inverted repeat β€” a sequence that reads the same on both strands β€” transcribes into RNA that can base-pair with itself, forming a hairpin. That hairpin is too bulky for the channel, and forming it physically wrenches RNA out of the enzyme, causing RNA polymerase to pause.

Then the stretch of AT base pairs does the rest. The RNA:DNA hybrid holding transcript to template at that point is rU:dA β€” the weakest hybrid there is. Stall the enzyme over the weakest possible grip and the transcript simply lets go.

A sequence, transcribed, becomes a mechanical device. That is a genuinely elegant piece of design, and it is worth being able to explain rather than merely name.

Test yourself
  • The three stages of transcription? → Initiation, elongation, termination
  • What is promoter clearance? → After ~3–10 nucleotides the polymerase changes conformation and moves away from the promoter
  • The two bacterial termination mechanisms? → Intrinsic (hairpin + AT stretch) and rho-dependent
  • What is rho? → An ATP-dependent, RNA-stimulated helicase that disrupts the elongation complex
  • Which subunit carries the polymerisation site? → Ξ²
The six steps of the transcription cycle: template binding, open complex formation, chain initiation, promoter clearance, elongation, and termination with RNAP release
The six steps of the transcription cycle: template binding, open complex formation, chain initiation, promoter clearance, elongation, and termination with RNAP release
Harper's Illustrated Biochemistry, Figure 36–3, p.396
07

Eukaryotic promoters and transcription factors β˜…β˜…β˜…

Bacterial promoters are simple: approximately 40 nucleotides in length, with an eight-nucleotide-pair sequence about 35 bp upstream of the transcription start site and a six-nucleotide-pair A+T-rich sequence about 10 nucleotides upstream β€” the classical βˆ’35 and βˆ’10 (Pribnow) boxes. Eukaryotic promoters are more complex, and are built from three classes of element.

ClassPositionElements
The promoter properAt and around +1TATA box, initiator sequence (Inr), downstream promoter element (DPE). The TATA box has a particularly rigid requirement for both position and orientation.
Promoter-proximal elements50–200 bp upstreamSequence elements bound by specific transcription factors, setting the frequency of initiation
Distal elements1000–10⁡ bp awayEnhancers and repressors (silencers) β€” a third class that can either increase or decrease the rate of transcription initiation
How often each combination occurs

TATA box and Inr: 30% Β· Inr alone: 30% Β· Inr and DPE: 25% Β· all three elements: 15%.

Note what this means: the TATA box is present in only a minority of genes. The textbook picture of β€œevery eukaryotic promoter has a TATA box” is wrong.

cis-acting elements and trans-acting factors

cis-acting elements are DNA sequences on the same molecule as the gene they control β€” promoters, enhancers, silencers.

trans-acting factors are diffusible proteins that bind them β€” the transcription factors.

Transcription factors have two functional parts: DNA-binding domains (DBDs) and activation domains (ADs).

The preinitiation complex (PIC)

A complex consisting of 50 unique proteins provides accurate and regulatable transcription of eukaryotic genes.

RNA polymerase II requires TFIIA, B, D (or TBP), E, F and H to both facilitate promoter-specific binding of the enzyme and formation of the preinitiation complex (PIC). RNA polymerases I and III require their own polymerase-specific GTFs.

TFIID binds to the TATA box promoter element through its TATA-binding protein (TBP) subunit; TFIID consists of 15 subunits β€” TBP and 14 TBP-associated factors (TAFs).

Crucially: RNA polymerase II and the GTFs can only catalyze BASAL or UNREGULATED transcription in vitro. Regulated transcription needs more.

Coactivators and chromatin

The coactivators, or coregulators, work in conjunction with the DNA-binding transactivator proteins to communicate with Pol II and the GTFs to regulate the rate of transcription.

Promoter accessibility, and hence PIC formation, is often modulated by nucleosomes: nucleosome eviction by chromatin-active coregulators facilitates PIC formation and transcription. The machinery therefore includes Mediator, chromatin remodellers and chromatin modifying factors alongside the polymerase and GTFs.

Phosphorylation activates Pol II

Eukaryotic pol II consists of 12 subunits. The largest carries, at its carboxyl terminus, a carboxyl terminal repeat domain (CTD) with the consensus sequence Tyr-Ser-Pro-Thr-Ser-Pro-Ser, repeated many times.

The CTD is a substrate for several enzymes, and CTD phosphorylation/dephosphorylation is critical for promoter clearance, elongation, termination, and even appropriate mRNA processing.

Why the CTD is the best-designed part of the whole machine

Ask what the cell needs. Capping must happen immediately after initiation, splicing during elongation, and polyadenylation at termination. Each processing enzyme must arrive at exactly the right moment.

The CTD solves this by acting as a moving scaffold whose phosphorylation state encodes the stage of transcription. Different patterns of phosphorylation on the repeated heptapeptide recruit different sets of enzymes, and the pattern changes as the polymerase progresses.

That is why Harper's can say the CTD is critical for promoter clearance, elongation, termination AND mRNA processing β€” four apparently separate jobs. It is also why eukaryotic transcription and processing are described as cotranscriptionally coupled: the processing machinery rides on the polymerase. One tail, carrying a clock.

Test yourself
  • The three classes of eukaryotic promoter element? → The promoter proper, promoter-proximal elements (50–200 bp), distal elements (1000–10⁡ bp)
  • Which element has a rigid position and orientation requirement? → The TATA box
  • What fraction of genes have all three core elements? → 15%
  • cis vs trans? → cis = DNA sequences; trans = diffusible protein factors
  • Which GTF binds the TATA box, and through what? → TFIID, through its TBP subunit
  • How many subunits has TFIID? → 15 β€” TBP plus 14 TAFs
  • What can Pol II + GTFs achieve alone? → Only basal, unregulated transcription
  • What is the CTD consensus sequence? → Tyr-Ser-Pro-Thr-Ser-Pro-Ser
The transcription control regions of a eukaryotic mRNA gene β€” distal regulatory elements, enhancers and repressors, promoter-proximal elements, then TATA, INR and DPE at the start site
The transcription control regions of a eukaryotic mRNA gene β€” distal regulatory elements, enhancers and repressors, promoter-proximal elements, then TATA, INR and DPE at the start site
Harper's Illustrated Biochemistry, Figure 36–8, p.400
A real example β€” the herpes simplex thymidine kinase promoter, with Sp1 bound at two GC boxes, CTF at the CAAT box, and TFIID at the TATA box 25 bp upstream of +1
A real example β€” the herpes simplex thymidine kinase promoter, with Sp1 bound at two GC boxes, CTF at the CAAT box, and TFIID at the TATA box 25 bp upstream of +1
Harper's Illustrated Biochemistry, Figure 36–7, p.400
Two models for assembly of the RNA polymerase II preinitiation complex: stepwise recruitment of the GTFs, or arrival as a preassembled holoenzyme
Two models for assembly of the RNA polymerase II preinitiation complex: stepwise recruitment of the GTFs, or arrival as a preassembled holoenzyme
Harper's Illustrated Biochemistry, Figure 36–11, p.404
08

Capping and polyadenylation β˜…β˜…β˜…

The 5β€² cap

Mammalian mRNA molecules contain a 7-methylguanosine cap structure at their 5β€² terminal.

The 5β€² cap of the RNA transcript is required both for efficient translation initiation and protection of the 5β€² end of mRNA from attack by 5β€²β†’3β€² exonucleases.

The poly(A) tail

Most mRNAs have a poly(A) tail at the 3β€² terminal.

The mRNA is first cleaved about 20 nucleotides downstream from an AAUAAA sequence; poly(A) polymerase then adds a poly(A) tail, which is subsequently extended to about 200 A residues.

The poly(A) tail both protects the 3β€² end of mRNA from 3β€²β†’5β€² exonuclease attack and facilitates translation.

Note the exception worth knowing: histone mRNA lacks a poly(A) tail.

Both modifications do the same two jobs

Read the two definitions side by side and the symmetry is exact. The cap protects against 5β€²β†’3β€² exonucleases and promotes translation; the tail protects against 3β€²β†’5β€² exonucleases and promotes translation. Both ends are capped against attack from the direction they are exposed to.

And the two cooperate: the cap and poly(A) tail structures have a synergistic effect on protein synthesis, because initiation factors bridge them and effectively circularise the message. A ribosome finishing at the 3β€² end is handed straight back to the 5β€² end.

That circularisation is also a quality check. Only a message with both a cap and a tail can be circularised β€” that is, only a transcript that was completed and processed properly. A truncated or damaged mRNA fails the test and is not translated.

The full list of post-transcriptional modifications

Processing β€” cleavage of the 45S rRNA precursor, and base modifications of tRNAs and rRNAs
Capping β€” mRNAs and snRNAs
Polyadenylation β€” mRNAs
Splicing β€” mRNAs, some tRNAs

Both ribosomal RNAs and most transfer RNAs are processed from larger precursors; the 45S transcript is cleaved to yield the 18S, 5.8S and 28S rRNAs.

Test yourself
  • What is the 5β€² cap? → A 7-methylguanosine structure
  • Its two functions? → Efficient translation initiation, and protection from 5β€²β†’3β€² exonucleases
  • Where is the mRNA cleaved before polyadenylation? → About 20 nucleotides downstream of AAUAAA
  • Length of the poly(A) tail? → Extended to about 200 A residues
  • Which mRNA lacks a poly(A) tail? → Histone mRNA
  • What is the rRNA precursor? → The 45S transcript
Transcription and processing are cotranscriptionally coupled: the phosphorylated CTD of Pol II carries the capping, splicing and polyadenylation machinery, and the TREX factors that package the message for export
Transcription and processing are cotranscriptionally coupled: the phosphorylated CTD of Pol II carries the capping, splicing and polyadenylation machinery, and the TREX factors that package the message for export
Harper's Illustrated Biochemistry, Figure 36–12, p.406
09

Introns, exons and splicing β˜…β˜…β˜…

Exons and introns

The RNA sequences that appear in mature RNAs are termed EXONS.

In mRNA-encoding genes, exons are often interrupted by long sequences of DNA that neither appear in mature mRNA, nor contribute to the genetic information ultimately translated into the amino acid sequence of a protein molecule. These intervening sequences are termed INTRONS.

The intron RNA sequences are cleaved out of the transcript, and the exons are appropriately spliced together in the nucleus before the resulting mRNA molecule appears in the cytoplasm for translation.

The spliceosome

Splicing depends on consensus sequences at the splice junctions and on an internal branch site.

The spliceosome is assembled from snRNAs (small nuclear RNAs, the U series) and snRNPs (small nuclear ribonucleoprotein particles).

1 Β· Pre-mRNA combines with the snRNPs and other proteins to form a spliceosome.
2 Β· Within the spliceosome, snRNA base-pairs with nucleotides at the ends of the intron.
3 Β· The RNA transcript is cut to release the intron, and the exons are spliced together; the spliceosome then comes apart, releasing mRNA, which now contains only exons.

Why the exon-intron arrangement was worth the cost

Introns look wasteful. A human gene may be tens of kilobases long and encode a protein from a few kilobases of exon; the rest is transcribed at full metabolic cost and then thrown away. Why tolerate that?

Alternative splicing provides for different mRNAs. By joining the same exons in different combinations, one gene yields several proteins. That is how roughly 20,000 human genes specify a far larger proteome β€” and it is a capability a prokaryote, with no nucleus and no time to splice, simply cannot have.

A related device operates at the other end. Alternative promoter utilization provides a form of regulation: in the glucokinase gene, the Ξ²-cell promoter and exon 1B are located about 30 kbp upstream from the liver promoter and exon 1L; each promoter has a unique structure and is regulated differently, while exons 2–10 are identical and the proteins have identical kinetic properties.

So the same enzyme you met in Unit 22 β€” glucokinase, sensing glucose in the Ξ² cell and trapping it in the liver β€” is one protein transcribed from two independently regulated promoters. Two jobs, two control systems, one coding sequence.

Test yourself
  • Define exon and intron → Exons appear in the mature RNA; introns are intervening sequences removed from it
  • What removes introns? → The spliceosome, built from snRNAs and snRNPs
  • Where does splicing occur? → In the nucleus, before export
  • What does alternative splicing achieve? → Different mRNAs, and hence different proteins, from one gene
  • Give an example of alternative promoter use → The glucokinase gene β€” separate liver and Ξ²-cell promoters ~30 kbp apart
The consensus sequences at splice junctions and the internal branch site β€” the addresses the spliceosome reads
The consensus sequences at splice junctions and the internal branch site β€” the addresses the spliceosome reads
Harper's Illustrated Biochemistry, Figure 36–14, p.407
The chemistry of splicing: nucleophilic attack at the 5β€² end of the intron, lariat formation, cut at the 3β€² end, then ligation of the two exons
The chemistry of splicing: nucleophilic attack at the 5β€² end of the intron, lariat formation, cut at the 3β€² end, then ligation of the two exons
Harper's Illustrated Biochemistry, Figure 36–13, p.407
10

Small RNAs and RNA editing β˜…β˜…

microRNAs (miRNAs)

The majority of miRNAs are transcribed by RNA pol II into primary transcripts termed pri-miRNAs. These are cut by Drosha into hairpins, transported into the cytoplasm and cut by Dicer, and the product anneals to the 3β€² untranslated region of the target mRNA and interferes with protein translation.

Three mechanisms: (a) promoting mRNA degradation directly; (b) stimulating CCR4/NOT complex-mediated poly(A) tail degradation; (c) inhibition of translation by targeting the 5β€²-methyl cap binding translation factor eIF4.

Small interfering RNAs (siRNAs) and RNA interference

RNA interference is triggered by the Dicer ribonuclease, which generates short interfering RNAs (siRNAs) of 21–28 bp. These are used to degrade target RNA by the RNA-induced silencing complex (RISC), and protect the host from RNA viruses.

RNA editing

RNA editing refers to the reactions that can change the nucleotide sequence of an mRNA molecule by non-splicing mechanisms. The change may include nucleotide change, deletion or insertion.

The classic example: the mRNA for apolipoprotein B in the liver is translated to apolipoprotein B100, while in the small intestine the mRNA is changed to yield a new termination codon (UAA), resulting in a much shorter protein, apolipoprotein B48.

The apo B story closes a loop from Unit 18

In Unit 18 you learned that apo B-48 is 48% of the length of apo B-100, from the same gene, and that this matters because B-48 lacks the LDL-receptor-binding domain β€” which is why chylomicron remnants must be cleared through apo E instead.

Here is the mechanism behind that sentence. A single C→U change in the intestinal transcript converts a glutamine codon (CAA) into a stop codon (UAA). The ribosome stops halfway, and the receptor-binding domain — encoded downstream — is never made.

One base, edited in one tissue, redirects an entire lipoprotein pathway. It is worth carrying as an example because it demonstrates the whole point of post-transcriptional control: the gene is identical in liver and intestine, and the difference is imposed entirely after transcription.

Test yourself
  • Which polymerase transcribes miRNAs? → Pol II, as pri-miRNAs
  • Which two enzymes process them? → Drosha in the nucleus, Dicer in the cytoplasm
  • Where does an miRNA bind its target? → The 3β€² untranslated region
  • Size of siRNAs, and what degrades the target? → 21–28 bp; the RNA-induced silencing complex (RISC)
  • Define RNA editing → Change of an mRNA's nucleotide sequence by non-splicing mechanisms
  • The classic example? → apo B-100 in liver vs apo B-48 in intestine, via a new UAA stop codon
11

Revision layer

Transcription versus replication β€” the contrasts examiners set

Replication (Unit 23)Transcription (Unit 24)
ProductDNARNA
SubstratesdNTPsATP, GTP, CTP, UTP
PrimerRequired β€” RNA, made by primaseNOT required β€” initiation is de novo
ExtentThe whole genomeSelected genes only
Strands copiedBothOne β€” the template strand, and not necessarily the same strand for every gene
Direction5β€²β†’3β€²5β€²β†’3β€² (template read 3β€²β†’5β€²)
ProofreadingYes β€” 3β€²β†’5β€² exonucleaseMuch less accurate; errors matter less, since transcripts are disposable

Numbers to have ready

QuantityValue
Transcription bubble20 bp; whole complex 30–75 bp
Bacterial promoter length~40 nucleotides; boxes at βˆ’35 (8 bp) and βˆ’10 (6 bp, A+T-rich)
Promoter clearanceAfter nucleotides 3 to ~10
Polymerase scanning rateβ‰₯10Β³ bp/s
Transcription initiation sites4 Γ— 10Β³ in E. coli; ~10⁡ in humans
Eukaryotic promoter element combinationsTATA+Inr 30% Β· Inr 30% Β· Inr+DPE 25% Β· all three 15%
Promoter-proximal / distal elements50–200 bp / 1000–10⁡ bp
TFIID15 subunits β€” TBP + 14 TAFs
Pol II subunits12; CTD consensus Tyr-Ser-Pro-Thr-Ser-Pro-Ser
Poly(A) tail~200 A residues, cleaved ~20 nt downstream of AAUAAA
siRNAs21–28 bp
rRNA precursor45S

Prokaryotic versus eukaryotic transcription

ProkaryoteEukaryote
PolymerasesOne, core Ξ±β‚‚Ξ²Ξ²β€² + ΟƒThree β€” Pol I, II, III
Promoter recognitionσ factorGeneral transcription factors (GTFs)
Promoter~40 nt; βˆ’35 and βˆ’10 boxesTATA / Inr / DPE + proximal + distal elements
Primary transcriptIs the mRNAPre-mRNA, requiring processing
ProcessingNone for mRNACap, poly(A), splicing
CouplingTranslation before transcription endsSeparated by the nuclear membrane
InhibitorRifampicin (Ξ² subunit)Ξ±-Amanitin
Two hooks

β€œPromoter = where the polymerase parks.” The full mark-earning sentence is β€œthe DNA sequence to which RNA polymerase binds to initiate transcription of a gene” β€” write that verbatim, since it was set in both papers.

β€œCap guards the 5β€², tail guards the 3′” β€” each end is protected against the exonuclease that attacks from its own direction, and both help translation.

Final self-test β€” cover the answers
  • Define a promoter → The DNA sequence to which RNA polymerase binds to initiate transcription of a gene
  • Define a transcription unit → The region of DNA including the signals for initiation, elongation and termination
  • Does transcription need a primer? → No β€” RNA polymerase initiates de novo
  • Bacterial holoenzyme composition, and Οƒ's job? → Ξ±β‚‚Ξ²Ξ²β€²Οƒ; Οƒ helps the core recognise the promoter
  • Which eukaryotic polymerase makes mRNA, and what inhibits it? → Pol II; Ξ±-amanitin
  • Which GTF binds the TATA box? → TFIID, via its TBP subunit
  • What does the CTD do? → Its phosphorylation state controls promoter clearance, elongation, termination and mRNA processing
  • The two bacterial termination mechanisms? → Intrinsic hairpin plus AT stretch, and rho-dependent
  • The three eukaryotic mRNA processing events? → 5β€² capping, 3β€² polyadenylation, and splicing
  • What does alternative splicing achieve? → Several proteins from one gene