Proteins: Determination of Primary Structure
← Back 📋 Q-Bank 🏠 All Units
HIGH YIELD ★★★
Proteins & Enzymes · Unit 2 of 26

Proteins: Determination of Primary Structure

TMU Lecture 2 — Lin Yu, Dept of Biochemistry & Molecular Biology Harper's ch. 4 — Proteins: Determination of Primary Structure, pp. 25–34 Ties directly to Practicals 1 (gel filtration) and 2 (BCA protein assay)
01

Why primary structure matters

Unit 1 gave you twenty letters. This unit asks the obvious next question: given an unknown protein pulled out of a cell, how do you read the order of those letters? That order — the primary structure — is not a trivial piece of bookkeeping. Everything the protein does follows from it, because the sequence is what dictates how the chain folds, and the fold is what makes an enzyme an enzyme.

Harper's opens the chapter with the clinical reason. An important goal of molecular medicine is to identify proteins whose presence, absence or deficiency marks a specific disease state. The primary sequence gives you two things at once: a molecular fingerprint that identifies the protein, and enough information to find and clone the gene that encodes it. Read the protein, and you have a route back to the DNA.

The shape of this whole unit

There are only four steps, and every exam answer to “describe how to determine the sequence of a polypeptide” walks through them in order:

1. Purify the protein · 2. Check it is pure and dissociate it into single chains · 3. Cut the long chain into short peptides · 4. Sequence the peptides (Edman or mass spectrometry) and reassemble them using overlaps.

Hold those four steps in your head and the rest of this page is just detail hanging off them.

Test yourself
  • What is primary structure? → The linear order of amino acid residues in a polypeptide, read N-terminus to C-terminus
  • Why does molecular medicine care about it? → It fingerprints the protein and leads back to the gene that encodes it
  • Name the four steps of sequencing a protein → Purify → assess purity and dissociate chains → cleave into peptides → sequence and overlap
02

Step one: proteins must be purified first ★★

A cell extract contains thousands of proteins at once. Before you can say anything about one of them you have to get it away from all the others, and the classic tricks all exploit differences in relative solubility. Change the conditions until your protein comes out of solution while the rest stay in — or the reverse.

MethodThe property exploitedHow it works
Isoelectric precipitationpHAt its pI a protein has no net charge, so molecules stop repelling each other and aggregate out of solution. Straight from Unit 1.
Solvent precipitationPolarityEthanol or acetone lowers the polarity of the water, stripping proteins of their hydration shell
Salting outSalt concentrationAmmonium sulfate at high concentration competes for water; different proteins drop out at different salt concentrations
Notice what just happened

Isoelectric precipitation is Unit 1's isoelectric point doing real laboratory work. This is the pattern for the whole course: a definition you learned as an abstraction turns out to be a technique. If an examiner asks why a protein precipitates at its pI, the answer is that net charge is zero, so the electrostatic repulsion that kept the molecules apart is gone.

For amino acids and sugars — small molecules — you can get away with the simplest chromatography of all: a sheet of filter paper (paper chromatography) or a thin layer of cellulose, silica or alumina (thin-layer chromatography, TLC). Proteins need something better, and that means a column.

Test yourself
  • Name three classic solubility-based purification methods → Isoelectric precipitation, solvent (ethanol/acetone) precipitation, salting out with ammonium sulfate
  • Why does a protein precipitate at its pI? → Net charge is zero, so the molecules no longer repel one another
  • Which salt is used for salting out? → Ammonium sulfate
03

Column chromatography — the six types ★★★

All chromatography works on one idea. There are two phases — a stationary phase (beads packed in a column) and a mobile phase (liquid flowing through it) — and every protein in the mixture partitions between them. A protein that clings to the beads is held back; a protein that prefers the flowing liquid comes out early. Change what the beads are coated with, and you change which property does the separating.

Partition chromatography — the general principle

Separation depends on the relative affinity of each protein for the stationary phase versus the mobile phase. The association is weak and transient; proteins that interact more strongly with the stationary phase are retained longer. Optimal separation is achieved by manipulating the composition of both phases.

The six types below are the whole of this section, and the examiner's question is always the same: on what basis does it separate? Learn the property, not the plumbing.

TypeSeparates on the basis ofThe key detail
Size exclusion
(gel filtration)
Stokes radiusPorous beads. Proteins too big to enter the pores are excluded and travel with the flow; small ones enter the pores and lag behind. So proteins emerge in descending order of size — big first.
Ion exchangeNet chargeCation exchangers carry negative groups (carboxylate, sulfate) and bind positively charged proteins; anion exchangers carry positive groups (tertiary/quaternary amines, e.g. DEAE-cellulose) and bind negatively charged proteins. Elute by raising ionic strength.
Hydrophobic interactionExposed hydrophobic surfaceMatrix coated with phenyl- or octyl-Sepharose. Binding is enhanced by high salt; elute by lowering salt — the opposite of ion exchange.
AffinityLigand binding — biological specificityImmobilised substrate, product, coenzyme or inhibitor. The most selective method. A Ni²⁺ matrix binds His-tagged recombinant proteins; a glutathione matrix binds GST fusions.
AbsorptionStrength of adsorption to the matrixProtein binds so tightly the partition coefficient is essentially 1. Non-binders wash through; bound proteins released by a rising salt gradient, in descending order of affinity.
Reversed-phase HPLCHydrophobicity, at high pressureIncompressible silica/alumina microbeads at up to a few thousand psi. Stationary phase = aliphatic chains 3–18 carbons long; eluted with a gradient of acetonitrile or methanol. This is how peptides are purified.
⭐ The three that come up most
Define gel filtration (size-exclusion) chromatography.
A method that separates proteins according to their Stokes radius — the radius of the sphere they occupy as they tumble in solution — using porous beads. Proteins too large to enter the pores are excluded and elute first; smaller proteins enter the pores, are retarded, and elute later. Proteins therefore emerge in descending order of Stokes radius.
Harper's ch.4, p.27 · TMU Lecture 2 Slide 6 · Practical Exp 1
Why is Stokes radius, not molecular mass, the correct answer?
Because Stokes radius is a function of both mass and shape. A tumbling elongated protein sweeps out a larger effective volume than a spherical protein of the same mass, so it behaves as though it were bigger. Writing “separates by molecular weight” will lose you the mark that “Stokes radius” earns.
Harper's ch.4, p.27
Why is affinity chromatography the most powerful single step?
Because it separates on biological specificity rather than a bulk physical property. In theory only proteins that recognise the immobilised ligand adhere at all, so a single passage can achieve what several conventional steps cannot. Bound protein is eluted by competition with free ligand, or less selectively with urea, guanidine HCl, mildly acidic pH or high salt.
Harper's ch.4, p.28 · TMU Lecture 2 Slide 10
The salt rule — worth memorising as a pair

Ion exchange: bind at low salt, elute by raising salt (salt ions compete for the charged sites).
Hydrophobic interaction: bind at high salt, elute by lowering salt (salt strengthens hydrophobic association).

They are mirror images, and a favourite way to catch students who memorised “gradient of salt” without asking in which direction.

Test yourself
  • Gel filtration separates by what? → Stokes radius — and large proteins elute FIRST
  • DEAE-cellulose is which kind of exchanger? → Anion exchanger (positively charged tertiary amine), so it binds negatively charged proteins
  • How do you elute from a hydrophobic interaction column? → By lowering the salt concentration — opposite to ion exchange
  • What binds a polyhistidine-tagged protein? → A Ni²⁺ affinity matrix
  • Which method purifies peptides after cleavage? → Reversed-phase HPLC
A typical liquid chromatography apparatus: R1/R2 mobile-phase reservoirs, P the programmable pumps with mixing chamber M, C the column, F the fraction collector
A typical liquid chromatography apparatus: R1/R2 mobile-phase reservoirs, P the programmable pumps with mixing chamber M, C the column, F the fraction collector
Harper's Illustrated Biochemistry, Figure 4–2, p.27
Size-exclusion chromatography: small molecules (red) enter the pores and lag behind, large molecules (brown) are excluded and elute first
Size-exclusion chromatography: small molecules (red) enter the pores and lag behind, large molecules (brown) are excluded and elute first
Harper's Illustrated Biochemistry, Figure 4–3, p.28
04

Checking purity: SDS-PAGE, IEF and 2-D ★★★

You have a fraction off a column. Is it one protein or five? The standard answer is SDS-PAGE, and the reason it works is worth understanding rather than memorising, because the logic is elegant.

Electrophoresis separates charged molecules by how fast they migrate in an electric field — which normally depends on both charge and size, an awkward mixture of two variables. SDS removes one of them. The detergent binds at a ratio of roughly one SDS molecule per two peptide bonds, unfolding the protein and coating it in negative charge. Because each SDS carries a charge of −1 and they are attached in proportion to length, every polypeptide ends up with about the same charge-to-mass ratio. Charge has been neutralised as a variable. What is left is purely physical resistance through the acrylamide mesh — so migration now reports relative molecular mass (Mr) and nothing else.

SDS-PAGE

Polyacrylamide gel electrophoresis in the presence of the anionic detergent sodium dodecyl sulfate. SDS denatures the polypeptide and confers a uniform charge-to-mass ratio, so separation depends on Mr alone. Used with 2-mercaptoethanol or dithiothreitol to reduce disulfide bonds, it separates the individual subunits of a multimeric protein. Bands are visualised with a dye such as Coomassie Blue.

⭐ Why the reducing agent matters
Why add 2-mercaptoethanol or DTT to an SDS-PAGE sample?
SDS unfolds a protein but cannot break covalent bonds. Disulfide bridges would hold subunits — or distant parts of one chain — together, so the protein would run as a single large band and you would misread its subunit composition. Reducing the –S–S– bonds releases the separate polypeptides. This is exactly what Sanger had to do to insulin before he could sequence it.
Harper's ch.4, p.28, Figure 4–4
Name two ways of cleaving disulfide bonds.
Oxidative cleavage with performic acid, which gives cysteic acid residues; and reductive cleavage with β-mercaptoethanol (or DTT), which gives cysteinyl residues.
Harper's ch.4, p.28, Figure 4–4

The second technique separates on a completely different property. In isoelectric focusing, ionic buffers called ampholytes plus an applied field set up a pH gradient down the gel. A protein put into that gradient migrates — and keeps migrating only until it reaches the pH that equals its own pI. There its net charge becomes zero, the field can no longer pull it, and it stops. Every protein parks at its own pI. It is a beautifully self-correcting method: drift either way and the protein picks up charge again and is pushed back.

Isoelectric focusing (IEF)

Separation of proteins in a pH gradient generated within a polyacrylamide matrix using ampholytes and an electric field. Each protein migrates until it reaches the pH equal to its isoelectric point (pI), the pH at which its net charge is zero, and there it stops.

Put the two together and you get two-dimensional electrophoresis: IEF first, separating by pI along one axis; then the IEF gel is laid across the top of an SDS gel and run again, separating by Mr down the other. Two independent properties, two axes. A crude bacterial extract that gives a smear of overlapping bands on a one-dimensional gel resolves into hundreds of discrete spots — which is why 2-D electrophoresis became the workhorse of early proteomics.

Test yourself
  • What does SDS do? → Denatures the protein and gives every polypeptide the same charge-to-mass ratio, so separation depends on Mr alone
  • In what ratio does SDS bind? → About one SDS molecule per two peptide bonds
  • Why add DTT or 2-mercaptoethanol? → To reduce disulfide bonds so subunits separate
  • What stops a protein moving in IEF? → Reaching the pH equal to its pI, where net charge is zero
  • What are the two dimensions of 2-D electrophoresis? → pI by IEF, then Mr by SDS-PAGE
SDS-PAGE following successive purification steps, stained with Coomassie Blue — lane S carries the M<sub>r</sub> standards
SDS-PAGE following successive purification steps, stained with Coomassie Blue — lane S carries the Mr standards
Harper's Illustrated Biochemistry, Figure 4–5, p.29
Two-dimensional IEF-SDS-PAGE — pI horizontally, M<sub>r</sub> vertically; note the resolution gained over a one-dimensional gel
Two-dimensional IEF-SDS-PAGE — pI horizontally, Mr vertically; note the resolution gained over a one-dimensional gel
Harper's Illustrated Biochemistry, Figure 4–6, p.29
Cleavage of disulfide bonds: oxidative with performic acid (left), giving cysteic acid residues, or reductive with β-mercaptoethanol (right), giving cysteinyl residues
Cleavage of disulfide bonds: oxidative with performic acid (left), giving cysteic acid residues, or reductive with β-mercaptoethanol (right), giving cysteinyl residues
Harper's Illustrated Biochemistry, Figure 4–4, p.28
05

Sanger and the first sequence

The first protein anyone sequenced was insulin, and the story is worth knowing because it contains, in miniature, every step still used today. Insulin is two chains — a 21-residue A chain and a 30-residue B chain — held together by disulfide bonds. Frederick Sanger reduced those bonds, separated the two chains, and cleaved each into smaller peptides with trypsin, chymotrypsin and pepsin.

Then came the clever part. He treated each peptide with 1-fluoro-2,4-dinitrobenzene — Sanger reagent — which labels the exposed α-amino group of the N-terminal residue. Hydrolyse the peptide afterwards and the labelled amino acid tells you which residue was at the front. Do that to overlapping fragments of increasing size and the whole sequence can be reconstructed. He received the Nobel Prize for it in 1958 — and a second one later for DNA sequencing.

⭐ Sanger reagent vs Edman reagent — the distinction that carries the marks
Both label the N-terminal residue. So why did Edman's method replace Sanger's?
Because of what happens to the rest of the peptide. Sanger reagent labels the N-terminus but the peptide must then be completely hydrolysed to identify it — the peptide is destroyed, and you learn exactly one residue per sample. Edman reagent's phenylthiohydantoin derivative can be removed under mild conditions, leaving the rest of the peptide intact with a brand-new N-terminus. So the same sample can be cycled again and again. One residue per sample versus many — that is the whole difference.
Harper's ch.4, p.29
A trap: does lysine interfere with Sanger reagent?
It reacts, but it does not confuse the result. Lysine's ε-amino group also reacts with Sanger reagent — however an N-terminal lysine reacts with 2 mol of reagent (α- and ε-amino), so it is readily distinguished from an internal lysine, which reacts with only one.
Harper's ch.4, p.29
Test yourself
  • Which protein was sequenced first, and by whom? → Insulin, by Frederick Sanger (Nobel Prize 1958)
  • What is Sanger reagent? → 1-fluoro-2,4-dinitrobenzene, which labels the free α-amino group of the N-terminal residue
  • Why is it inferior to Edman? → The peptide must be hydrolysed to read the label, so only one residue is learned per sample
06

The Edman reaction ★★★

This is the definition most likely to appear in Section I, so learn it as a cycle of three moves rather than as a sentence. Pehr Edman's insight was to find a reagent that would grab the N-terminal residue, let go of it under conditions gentle enough to leave the remaining peptide bonds intact, and then be applied again to the residue newly exposed.

The Edman reaction

Phenylisothiocyanate (Edman reagent) derivatises the amino-terminal residue of a peptide as a phenylthiohydantoic acid. Treatment with acid in a non-hydroxylic solvent releases a phenylthiohydantoin (PTH) amino acid — identified by its chromatographic mobility — and a peptide one residue shorter. The process is then repeated.

StepWhat happensReagent / condition
1 · CoupleThe reagent attaches to the free α-amino group of the N-terminal residue, forming a phenylthiohydantoic acidPhenylisothiocyanate, mildly alkaline
2 · CleaveThat one residue is released as a phenylthiohydantoin; the rest of the chain is untouchedAcid in a non-hydroxylic solvent (e.g. nitromethane)
3 · Identify & repeatThe PTH–amino acid is identified by chromatographic mobility; the shortened peptide has a new N-terminus and the cycle begins againAutomated sequenator
⭐ How many residues can Edman actually read?
State the read length and explain the limit.
5 to 30 residues, depending on the quantity and purity of the peptide. The reason is cumulative inefficiency: the twenty amino acids are chemically heterogeneous, so every step is a compromise and none runs at 100% efficiency. Chains that fail to react in a given cycle fall out of phase with the rest, and the resulting mixture of N-termini eventually makes it impossible to tell the correct PTH–amino acid from the contaminants.

Note: your TMU slide gives “the first 20–30 residues”. Harper's gives 5–30. Both describe the same limit — quote a figure “of the order of 20–30 residues” and explain the out-of-phase reason, which is what actually earns the mark.
Harper's ch.4, p.29 · TMU Lecture 2 Slide 17
Why must the acid step use a non-hydroxylic solvent?
A hydroxylic solvent such as water would let the acid hydrolyse the internal peptide bonds as well, destroying the chain. The non-hydroxylic condition confines cleavage to the derivatised N-terminal residue — which is the entire basis of the method's repeatability.
Harper's ch.4, p.30, Figure 4–7
Test yourself
  • Name the Edman reagent → Phenylisothiocyanate
  • What is released at each cycle? → A phenylthiohydantoin (PTH) amino acid, plus a peptide one residue shorter
  • How is the released residue identified? → By its chromatographic mobility
  • How many residues can be read? → Of the order of 5–30, limited by cycles falling out of phase
The Edman reaction — phenylisothiocyanate derivatises the N-terminal residue as a phenylthiohydantoic acid; acid in a non-hydroxylic solvent then releases a phenylthiohydantoin plus a peptide one residue shorter
The Edman reaction — phenylisothiocyanate derivatises the N-terminal residue as a phenylthiohydantoic acid; acid in a non-hydroxylic solvent then releases a phenylthiohydantoin plus a peptide one residue shorter
Harper's Illustrated Biochemistry, Figure 4–7, p.30
07

Cleaving large polypeptides, and why overlaps matter ★★

Edman reads perhaps thirty residues. Most polypeptides are several hundred. So the long chain must first be cut into pieces short enough to sequence — and there is a second reason to cut, which the slides mention and students often miss: post-translational modification can leave the α-amino group blocked and unreactive with Edman reagent. Cleaving generates fresh N-termini that will react.

Now the crucial idea. Cut a chain into four pieces and sequence each one, and you know four sequences but not the order they came in — and there are 24 possible orders. The solution is to take a second sample of the intact protein and cut it with a different reagent, one that cuts in different places. The second set of peptides straddles the junctions of the first set. Those overlaps establish continuity and fix the order.

The torn-newspaper analogy

Tear a page of newsprint into four strips and shuffle them: you can read each strip, but you cannot tell which came first. Now take a second copy of the same page and tear it at different places. Each strip from the second copy contains the end of one first-copy strip and the beginning of another — so it tells you which two go together. That is the entire logic of overlapping peptides, and it is why you always need more than one method of cleavage.

ReagentCleaves the peptide bond on the C-side ofType
TrypsinArg and Lys — basic residuesEnzyme
ChymotrypsinAromatic residues — Phe, Trp, TyrEnzyme
Cyanogen bromide (CNBr)Met onlyChemical
S. aureus V8 proteaseAcidic residues — Glu (and Asp)Enzyme
⭐ Predicting fragment numbers — a favourite calculation
A 38-residue polypeptide contains 1 Arg, 2 Lys and 2 Met. How many fragments does trypsin give? How many does CNBr give?
Trypsin: four. It cleaves after each Arg and each Lys — that is 3 cleavage sites, and n cuts in a linear chain give n+1 pieces.
CNBr: three. Two Met residues, so 2 cuts, so 3 pieces.

The general rule to state: fragments = cleavage sites + 1. Watch for the trap where the residue is already at the C-terminus — cutting after it produces nothing new.
Lehninger (2004), pp. 100–101
How do you know which fragment is the C-terminal one after a tryptic digest?
The C-terminal fragment is the one that does not end in Arg or Lys. Every other tryptic peptide must, by definition, end at the residue trypsin cut after.
Lehninger (2004), p. 101

After cleavage the peptides are purified by reversed-phase HPLC — occasionally by SDS-PAGE — and then sequenced. Note the practical drawback Harper's flags: because you must run several different fragmentation and purification conditions, direct chemical sequencing needs large quantities of purified protein. That, together with the slowness, is why the field moved on.

Test yourself
  • Why cleave a large polypeptide? → Edman reads only ~30 residues, and cleavage also bypasses a blocked N-terminus
  • Why use more than one cleavage reagent? → To generate overlapping peptides that establish the order of the fragments
  • Trypsin cleaves after which residues? → Arg and Lys
  • CNBr cleaves after which residue? → Met
  • How are the peptides purified before sequencing? → Reversed-phase HPLC
08

Mass spectrometry — the method that took over ★★★

Harper's is blunt about it: the superior sensitivity, speed and versatility of mass spectrometry have replaced the Edman technique as the principal method for sequencing peptides and proteins. Understand why, and the section writes itself.

MS discriminates molecules on mass alone. That has one consequence that matters enormously in medicine: a post-translational modification — a phosphate group, a hydroxyl, a sugar — adds mass. So MS detects it directly. Edman sequencing struggles to identify which modification it has hit, and a DNA-derived sequence cannot see modifications at all, because they are added after translation. If a question asks why DNA sequencing has not made protein chemistry obsolete, this is the answer.

⭐ Why not just sequence the gene and translate it?
DNA sequencing is faster and cheaper. So why sequence proteins?
Because DNA tells you the order in which amino acids were added on the ribosome, and nothing more. It provides no information about post-translational modification — proteolytic processing, methylation, glycosylation, phosphorylation, hydroxylation of proline and lysine, or disulfide bond formation. The mature, functional protein may differ substantially from its gene-predicted sequence. In practice a hybrid approach is used: Edman or MS to read a short stretch of the real protein, then DNA cloning to obtain the rest.
Harper's ch.4, p.30 · TMU Lecture 2 Slides 19–20

How the instrument works

The sample is vaporised under vacuum in the presence of a proton donor, so the molecules pick up positive charge. An electric field accelerates the cations down a flight tube. From there the two common designs diverge:

DesignHow it measures massBest for
Quadrupole (magnetic sector)A magnetic field deflects the ions at right angles; the current needed to bend an ion's path onto the detector is proportional to its mass (for ions of equal charge)Molecules of 4000 Da or less
Time-of-flight (TOF)A straight flight tube. Time taken to reach the detector is inversely proportional to mass — heavy ions accelerate less and arrive laterWhole proteins, large masses
⭐ Another slide-vs-textbook discrepancy — read this before the exam
What is the mass limit of a conventional/quadrupole mass spectrometer?
Your TMU slide says 1000 Da. The Harper's edition in your reference folder says 4000 Da. This is an edition difference, not an error on either side — instruments improved. Answer with your lecturer's figure (1000 Da), since that is what the marking key will hold. What is not in dispute, and is what the question is really testing, is the contrast: quadrupole for small molecules, time-of-flight for whole proteins.
TMU Lecture 2 Slide 24 vs Harper's ch.4, p.31

Getting big molecules into the vapour phase

The obstacle that held MS back for years was simple: you can vaporise a small organic molecule by heating it in a vacuum, but a protein heated that way is destroyed. Three techniques solved it, and two of them are examinable by name.

MethodHow it avoids destroying the protein
Electrospray ionisationThe sample, dissolved in a volatile solvent, is sprayed through a capillary into the chamber. The solvent flashes away, leaving the macromolecule suspended in the gas phase. Convenient because peptides can be fed straight from an HPLC column into the spectrometer.
MALDI
(matrix-assisted laser desorption/ionisation)
The sample is mixed with a liquid matrix containing a light-absorbing dye and a proton source. A laser excites the matrix, which disperses into the vapour phase so fast that the embedded protein is carried along without being heated.
Fast atom bombardment (FAB)Macromolecules dispersed in glycerol or another protonic matrix are bombarded with a stream of neutral atoms

The pay-off is remarkable precision. MALDI and electrospray allow the masses of polypeptides above 100 000 Da to be determined to within about ±1 Da — accurate enough to see a single added phosphate.

Sequencing by fragmentation

Knowing a peptide's total mass is not a sequence. To get the order, the peptide is broken up inside the instrument by collision with neutral helium atoms (collision-induced dissociation) and the fragments weighed. Peptide bonds are much more labile than carbon–carbon bonds, so the chain preferentially breaks between residues — meaning the most abundant fragments differ from one another by exactly one amino acid. Since the molecular mass of each amino acid is unique, the difference in mass between two successive fragments names the residue that was lost, and the sequence can be reconstructed from the ladder of masses.

⭐ The one pair MS cannot tell apart
Which amino acids cannot be distinguished by mass spectrometry, and why?
Leucine and isoleucine. They are structural isomers — same formula, C₆H₁₃NO₂, therefore identical molecular mass. Every other amino acid has a unique mass, which is what makes the method work. This exception is the standard exam sting.
Harper's ch.4, p.33 · TMU Lecture 2 Slide 24
Tandem mass spectrometry (MS–MS or MS²)

Two mass spectrometers linked in series, allowing complex peptide mixtures to be analysed without prior purification. The first separates individual peptides by mass and directs a single chosen peptide into the second, where it is fragmented and the fragment masses determined.

Where you will meet this on the wards

Tandem MS is used to screen newborn blood samples for amino acids, fatty acids and other metabolites. Abnormal metabolite levels are diagnostic indicators for genetic disorders — Harper's names phenylketonuria, ethylmalonic encephalopathy and glutaric acidaemia type 1. This is the single most clinically important sentence in the chapter: the technique in your biochemistry lecture is the technique behind the heel-prick test.

Test yourself
  • Why has MS replaced Edman? → Greater sensitivity, speed and versatility, and it detects post-translational modifications by their added mass
  • Quadrupole vs TOF? → Quadrupole for small molecules; time-of-flight for whole proteins
  • Name two ways of volatilising a protein → Electrospray ionisation and MALDI (also fast atom bombardment)
  • How is sequence read from the fragments? → Successive fragments differ by one residue, and each amino acid has a unique mass
  • Which two amino acids cannot be distinguished? → Leucine and isoleucine — isomers of identical mass
  • What is tandem MS used for clinically? → Newborn screening for metabolic disorders such as phenylketonuria
Basic components of a simple mass spectrometer — the greater the mass of the ion, the higher the magnetic field needed to focus it onto the detector
Basic components of a simple mass spectrometer — the greater the mass of the ion, the higher the magnetic field needed to focus it onto the detector
Harper's Illustrated Biochemistry, Figure 4–8, p.32
Three ways of vaporising molecules in the sample chamber: heating, electrospray ionisation, and MALDI
Three ways of vaporising molecules in the sample chamber: heating, electrospray ionisation, and MALDI
Harper's Illustrated Biochemistry, Figure 4–9, p.32
09

Proteomics — the endpoint of the chapter

The chapter closes by scaling up from one protein to all of them. The genome is fixed and static; the set of proteins actually present is neither. Genes switch on and off, muscle cells express proteins neural cells never touch, the subunits of haemoglobin change between fetal and adult life, and proteins are modified after synthesis. Knowing the genome is therefore only the beginning.

The proteome

The set of all the proteins expressed by an individual cell at a particular time. Because the body contains thousands of cell types each containing thousands of proteins — and because expression changes with growth, differentiation and external stimuli — the proteome is a moving target, not a fixed list like the genome.

The goal of proteomics is to identify proteins whose level of expression correlates with medically significant events, on the presumption that a protein appearing or disappearing alongside a disease is linked to its cause or mechanism. The problem is scale. Antibody and enzyme assays are exquisitely specific but can only look at proteins you already suspect; total-protein assays such as the Lowry or Bradford method, and stains such as Coomassie Blue, are universal but tell you nothing about which protein you are looking at.

First-generation proteomics threaded between the two: resolve everything on a two-dimensional gel, extract individual spots, and identify each by Edman sequencing or mass spectrometry, matching Mr and pI against the databases. A single gel resolves only about a thousand proteins, but its advantage is that it examines the proteins themselves. The complementary approach — gene arrays, or DNA chips — detects the mRNAs instead. Arrays are more sensitive and cover more gene products, but carry a real caveat: a change in mRNA level does not necessarily mean a comparable change in the protein.

Finally, bioinformatics lets you guess a new protein's function from its sequence alone. Nature reuses structural themes, so algorithms look for conserved amino acids at key positions that mark a known domain — the Rossmann fold that binds NAD(P)H, nuclear targeting sequences, EF hands that bind Ca²⁺. Find the motif and you have a strong hypothesis about what the protein does before you have run a single assay.

Test yourself
  • Define the proteome → All the proteins expressed by an individual cell at a particular time
  • Why is it a “moving target”? → Expression varies with cell type, time, differentiation and stimuli, and proteins are modified after synthesis
  • Which two methods survey protein expression? → Two-dimensional electrophoresis (the proteins themselves) and gene array / DNA chips (the mRNAs)
  • What is the caveat with gene arrays? → mRNA level does not necessarily reflect protein level
  • Name two conserved domains bioinformatics looks for → The Rossmann fold (binds NAD(P)H) and EF hands (bind Ca²⁺)
10

Revision layer

The 2019 paper asked, in Section II, simply: “Describe the methods of determining the sequence of a polypeptide.” That is this whole page in one question, and it is worth 8 marks. The model answer is the four-step skeleton, each step named with its reagent.

The model answer — memorise this skeleton

StepWhat you doName the specifics
1 · PurifyIsolate the protein from the cell extractSalting out (ammonium sulfate), isoelectric precipitation; then column chromatography — ion exchange, gel filtration, affinity
2 · Assess purity & dissociateConfirm you have one protein, and break it into single chainsSDS-PAGE with Coomassie Blue; reduce disulfide bonds with 2-mercaptoethanol or DTT (or oxidise with performic acid)
3 · CleaveCut the chain into peptides short enough to sequence, using two different reagents so the peptides overlapTrypsin (after Arg/Lys), chymotrypsin (after aromatics), CNBr (after Met), V8 protease (after Glu); purify peptides by reversed-phase HPLC
4 · Sequence & assembleRead each peptide, then use the overlaps to order the fragmentsEdman degradation (phenylisothiocyanate → PTH amino acid, ~5–30 residues) or tandem mass spectrometry; or the hybrid approach — short protein sequence + DNA cloning

Definitions from this unit — Section I material

TermDefinition
Primary structureThe linear sequence of amino acid residues in a polypeptide chain, read from the N-terminus to the C-terminus
Gel filtration (size-exclusion) chromatographySeparation of proteins according to their Stokes radius using porous beads; excluded (large) proteins elute first, included (small) proteins are retarded and elute later
Stokes radiusThe radius of the sphere a protein occupies as it tumbles in solution; a function of both molecular mass and shape
Affinity chromatographyPurification exploiting a protein's specific binding to an immobilised ligand — substrate, product, coenzyme or inhibitor; only proteins that recognise the ligand adhere
SDS-PAGEPolyacrylamide gel electrophoresis in the presence of sodium dodecyl sulfate, which denatures the protein and confers a uniform charge-to-mass ratio so that separation depends on relative molecular mass alone
Isoelectric focusingSeparation of proteins in a pH gradient generated with ampholytes, each protein migrating until it reaches the pH equal to its isoelectric point, where its net charge is zero
Edman reactionPhenylisothiocyanate derivatises the N-terminal residue as a phenylthiohydantoic acid; acid in a non-hydroxylic solvent then releases a phenylthiohydantoin, identified chromatographically, plus a peptide one residue shorter — and the cycle repeats
Tandem mass spectrometryTwo mass spectrometers in series, allowing complex peptide mixtures to be analysed without prior purification: the first selects a peptide by mass, the second fragments it and weighs the fragments
ProteomeThe set of all the proteins expressed by an individual cell at a particular time

Separation methods at a glance — what separates on what

MethodSeparates by
Gel filtration / size exclusionStokes radius  (large elute first)
Ion exchangeNet charge  (elute by raising salt)
Hydrophobic interactionExposed hydrophobic surface  (elute by lowering salt)
AffinitySpecific ligand binding
Reversed-phase HPLCHydrophobicity, at high pressure  (purifies peptides)
SDS-PAGERelative molecular mass (Mr)
Isoelectric focusingIsoelectric point (pI)
2-D electrophoresispI in one dimension, Mr in the other

Cleavage specificities — learn all four

ReagentCleaves after
TrypsinArg, Lys
ChymotrypsinPhe, Trp, Tyr (aromatic)
Cyanogen bromideMet
S. aureus V8 proteaseGlu (acidic)

Numbers worth carrying in

FigureValue
Insulin A chain / B chain21 residues / 30 residues
Sanger's Nobel Prize1958
SDS binding ratio1 SDS per 2 peptide bonds
Edman read length~5–30 residues
Quadrupole MS mass limit1000 Da (slide) · 4000 Da (Harper's)
MALDI / electrospray accuracy>100 000 Da to within ±1 Da
Proteins resolved on one 2-D gel~1000
Your practicals sit inside this chapter

Practical 1 — Gel Filtration Chromatography is §3 of this page done with your own hands: you separate on Stokes radius and watch the large molecules come off first. Practical 2 — quantitative protein assay with a BCA kit is the “universal but non-identifying” category Harper's mentions alongside Lowry and Bradford — it tells you how much protein, never which. Expect the viva to ask why.

Final check — can you do these cold?
  • Write the four-step skeleton for sequencing a polypeptide, naming a reagent at each step
  • Define gel filtration, SDS-PAGE, isoelectric focusing, the Edman reaction and the proteome in exam wording
  • State what each of the six chromatographies separates on
  • Give the four cleavage reagents and their specificities, and calculate fragment numbers from a composition
  • Explain why DNA sequencing cannot replace protein sequencing
  • Explain why leucine and isoleucine defeat mass spectrometry