Proteins: Determination of Primary Structure
Why primary structure matters
Unit 1 gave you twenty letters. This unit asks the obvious next question: given an unknown protein pulled out of a cell, how do you read the order of those letters? That order — the primary structure — is not a trivial piece of bookkeeping. Everything the protein does follows from it, because the sequence is what dictates how the chain folds, and the fold is what makes an enzyme an enzyme.
Harper's opens the chapter with the clinical reason. An important goal of molecular medicine is to identify proteins whose presence, absence or deficiency marks a specific disease state. The primary sequence gives you two things at once: a molecular fingerprint that identifies the protein, and enough information to find and clone the gene that encodes it. Read the protein, and you have a route back to the DNA.
There are only four steps, and every exam answer to “describe how to determine the sequence of a polypeptide” walks through them in order:
1. Purify the protein · 2. Check it is pure and dissociate it into single chains · 3. Cut the long chain into short peptides · 4. Sequence the peptides (Edman or mass spectrometry) and reassemble them using overlaps.
Hold those four steps in your head and the rest of this page is just detail hanging off them.
- What is primary structure? → The linear order of amino acid residues in a polypeptide, read N-terminus to C-terminus
- Why does molecular medicine care about it? → It fingerprints the protein and leads back to the gene that encodes it
- Name the four steps of sequencing a protein → Purify → assess purity and dissociate chains → cleave into peptides → sequence and overlap
Step one: proteins must be purified first ★★
A cell extract contains thousands of proteins at once. Before you can say anything about one of them you have to get it away from all the others, and the classic tricks all exploit differences in relative solubility. Change the conditions until your protein comes out of solution while the rest stay in — or the reverse.
| Method | The property exploited | How it works |
|---|---|---|
| Isoelectric precipitation | pH | At its pI a protein has no net charge, so molecules stop repelling each other and aggregate out of solution. Straight from Unit 1. |
| Solvent precipitation | Polarity | Ethanol or acetone lowers the polarity of the water, stripping proteins of their hydration shell |
| Salting out | Salt concentration | Ammonium sulfate at high concentration competes for water; different proteins drop out at different salt concentrations |
Isoelectric precipitation is Unit 1's isoelectric point doing real laboratory work. This is the pattern for the whole course: a definition you learned as an abstraction turns out to be a technique. If an examiner asks why a protein precipitates at its pI, the answer is that net charge is zero, so the electrostatic repulsion that kept the molecules apart is gone.
For amino acids and sugars — small molecules — you can get away with the simplest chromatography of all: a sheet of filter paper (paper chromatography) or a thin layer of cellulose, silica or alumina (thin-layer chromatography, TLC). Proteins need something better, and that means a column.
- Name three classic solubility-based purification methods → Isoelectric precipitation, solvent (ethanol/acetone) precipitation, salting out with ammonium sulfate
- Why does a protein precipitate at its pI? → Net charge is zero, so the molecules no longer repel one another
- Which salt is used for salting out? → Ammonium sulfate
Column chromatography — the six types ★★★
All chromatography works on one idea. There are two phases — a stationary phase (beads packed in a column) and a mobile phase (liquid flowing through it) — and every protein in the mixture partitions between them. A protein that clings to the beads is held back; a protein that prefers the flowing liquid comes out early. Change what the beads are coated with, and you change which property does the separating.
Separation depends on the relative affinity of each protein for the stationary phase versus the mobile phase. The association is weak and transient; proteins that interact more strongly with the stationary phase are retained longer. Optimal separation is achieved by manipulating the composition of both phases.
The six types below are the whole of this section, and the examiner's question is always the same: on what basis does it separate? Learn the property, not the plumbing.
| Type | Separates on the basis of | The key detail |
|---|---|---|
| Size exclusion (gel filtration) | Stokes radius | Porous beads. Proteins too big to enter the pores are excluded and travel with the flow; small ones enter the pores and lag behind. So proteins emerge in descending order of size — big first. |
| Ion exchange | Net charge | Cation exchangers carry negative groups (carboxylate, sulfate) and bind positively charged proteins; anion exchangers carry positive groups (tertiary/quaternary amines, e.g. DEAE-cellulose) and bind negatively charged proteins. Elute by raising ionic strength. |
| Hydrophobic interaction | Exposed hydrophobic surface | Matrix coated with phenyl- or octyl-Sepharose. Binding is enhanced by high salt; elute by lowering salt — the opposite of ion exchange. |
| Affinity | Ligand binding — biological specificity | Immobilised substrate, product, coenzyme or inhibitor. The most selective method. A Ni²⁺ matrix binds His-tagged recombinant proteins; a glutathione matrix binds GST fusions. |
| Absorption | Strength of adsorption to the matrix | Protein binds so tightly the partition coefficient is essentially 1. Non-binders wash through; bound proteins released by a rising salt gradient, in descending order of affinity. |
| Reversed-phase HPLC | Hydrophobicity, at high pressure | Incompressible silica/alumina microbeads at up to a few thousand psi. Stationary phase = aliphatic chains 3–18 carbons long; eluted with a gradient of acetonitrile or methanol. This is how peptides are purified. |
Ion exchange: bind at low salt, elute by raising salt (salt ions compete for the charged sites).
Hydrophobic interaction: bind at high salt, elute by lowering salt (salt strengthens hydrophobic association).
They are mirror images, and a favourite way to catch students who memorised “gradient of salt” without asking in which direction.
- Gel filtration separates by what? → Stokes radius — and large proteins elute FIRST
- DEAE-cellulose is which kind of exchanger? → Anion exchanger (positively charged tertiary amine), so it binds negatively charged proteins
- How do you elute from a hydrophobic interaction column? → By lowering the salt concentration — opposite to ion exchange
- What binds a polyhistidine-tagged protein? → A Ni²⁺ affinity matrix
- Which method purifies peptides after cleavage? → Reversed-phase HPLC


Checking purity: SDS-PAGE, IEF and 2-D ★★★
You have a fraction off a column. Is it one protein or five? The standard answer is SDS-PAGE, and the reason it works is worth understanding rather than memorising, because the logic is elegant.
Electrophoresis separates charged molecules by how fast they migrate in an electric field — which normally depends on both charge and size, an awkward mixture of two variables. SDS removes one of them. The detergent binds at a ratio of roughly one SDS molecule per two peptide bonds, unfolding the protein and coating it in negative charge. Because each SDS carries a charge of −1 and they are attached in proportion to length, every polypeptide ends up with about the same charge-to-mass ratio. Charge has been neutralised as a variable. What is left is purely physical resistance through the acrylamide mesh — so migration now reports relative molecular mass (Mr) and nothing else.
Polyacrylamide gel electrophoresis in the presence of the anionic detergent sodium dodecyl sulfate. SDS denatures the polypeptide and confers a uniform charge-to-mass ratio, so separation depends on Mr alone. Used with 2-mercaptoethanol or dithiothreitol to reduce disulfide bonds, it separates the individual subunits of a multimeric protein. Bands are visualised with a dye such as Coomassie Blue.
The second technique separates on a completely different property. In isoelectric focusing, ionic buffers called ampholytes plus an applied field set up a pH gradient down the gel. A protein put into that gradient migrates — and keeps migrating only until it reaches the pH that equals its own pI. There its net charge becomes zero, the field can no longer pull it, and it stops. Every protein parks at its own pI. It is a beautifully self-correcting method: drift either way and the protein picks up charge again and is pushed back.
Separation of proteins in a pH gradient generated within a polyacrylamide matrix using ampholytes and an electric field. Each protein migrates until it reaches the pH equal to its isoelectric point (pI), the pH at which its net charge is zero, and there it stops.
Put the two together and you get two-dimensional electrophoresis: IEF first, separating by pI along one axis; then the IEF gel is laid across the top of an SDS gel and run again, separating by Mr down the other. Two independent properties, two axes. A crude bacterial extract that gives a smear of overlapping bands on a one-dimensional gel resolves into hundreds of discrete spots — which is why 2-D electrophoresis became the workhorse of early proteomics.
- What does SDS do? → Denatures the protein and gives every polypeptide the same charge-to-mass ratio, so separation depends on Mr alone
- In what ratio does SDS bind? → About one SDS molecule per two peptide bonds
- Why add DTT or 2-mercaptoethanol? → To reduce disulfide bonds so subunits separate
- What stops a protein moving in IEF? → Reaching the pH equal to its pI, where net charge is zero
- What are the two dimensions of 2-D electrophoresis? → pI by IEF, then Mr by SDS-PAGE



Sanger and the first sequence
The first protein anyone sequenced was insulin, and the story is worth knowing because it contains, in miniature, every step still used today. Insulin is two chains — a 21-residue A chain and a 30-residue B chain — held together by disulfide bonds. Frederick Sanger reduced those bonds, separated the two chains, and cleaved each into smaller peptides with trypsin, chymotrypsin and pepsin.
Then came the clever part. He treated each peptide with 1-fluoro-2,4-dinitrobenzene — Sanger reagent — which labels the exposed α-amino group of the N-terminal residue. Hydrolyse the peptide afterwards and the labelled amino acid tells you which residue was at the front. Do that to overlapping fragments of increasing size and the whole sequence can be reconstructed. He received the Nobel Prize for it in 1958 — and a second one later for DNA sequencing.
- Which protein was sequenced first, and by whom? → Insulin, by Frederick Sanger (Nobel Prize 1958)
- What is Sanger reagent? → 1-fluoro-2,4-dinitrobenzene, which labels the free α-amino group of the N-terminal residue
- Why is it inferior to Edman? → The peptide must be hydrolysed to read the label, so only one residue is learned per sample
The Edman reaction ★★★
This is the definition most likely to appear in Section I, so learn it as a cycle of three moves rather than as a sentence. Pehr Edman's insight was to find a reagent that would grab the N-terminal residue, let go of it under conditions gentle enough to leave the remaining peptide bonds intact, and then be applied again to the residue newly exposed.
Phenylisothiocyanate (Edman reagent) derivatises the amino-terminal residue of a peptide as a phenylthiohydantoic acid. Treatment with acid in a non-hydroxylic solvent releases a phenylthiohydantoin (PTH) amino acid — identified by its chromatographic mobility — and a peptide one residue shorter. The process is then repeated.
| Step | What happens | Reagent / condition |
|---|---|---|
| 1 · Couple | The reagent attaches to the free α-amino group of the N-terminal residue, forming a phenylthiohydantoic acid | Phenylisothiocyanate, mildly alkaline |
| 2 · Cleave | That one residue is released as a phenylthiohydantoin; the rest of the chain is untouched | Acid in a non-hydroxylic solvent (e.g. nitromethane) |
| 3 · Identify & repeat | The PTH–amino acid is identified by chromatographic mobility; the shortened peptide has a new N-terminus and the cycle begins again | Automated sequenator |
Note: your TMU slide gives “the first 20–30 residues”. Harper's gives 5–30. Both describe the same limit — quote a figure “of the order of 20–30 residues” and explain the out-of-phase reason, which is what actually earns the mark.
- Name the Edman reagent → Phenylisothiocyanate
- What is released at each cycle? → A phenylthiohydantoin (PTH) amino acid, plus a peptide one residue shorter
- How is the released residue identified? → By its chromatographic mobility
- How many residues can be read? → Of the order of 5–30, limited by cycles falling out of phase

Cleaving large polypeptides, and why overlaps matter ★★
Edman reads perhaps thirty residues. Most polypeptides are several hundred. So the long chain must first be cut into pieces short enough to sequence — and there is a second reason to cut, which the slides mention and students often miss: post-translational modification can leave the α-amino group blocked and unreactive with Edman reagent. Cleaving generates fresh N-termini that will react.
Now the crucial idea. Cut a chain into four pieces and sequence each one, and you know four sequences but not the order they came in — and there are 24 possible orders. The solution is to take a second sample of the intact protein and cut it with a different reagent, one that cuts in different places. The second set of peptides straddles the junctions of the first set. Those overlaps establish continuity and fix the order.
Tear a page of newsprint into four strips and shuffle them: you can read each strip, but you cannot tell which came first. Now take a second copy of the same page and tear it at different places. Each strip from the second copy contains the end of one first-copy strip and the beginning of another — so it tells you which two go together. That is the entire logic of overlapping peptides, and it is why you always need more than one method of cleavage.
| Reagent | Cleaves the peptide bond on the C-side of | Type |
|---|---|---|
| Trypsin | Arg and Lys — basic residues | Enzyme |
| Chymotrypsin | Aromatic residues — Phe, Trp, Tyr | Enzyme |
| Cyanogen bromide (CNBr) | Met only | Chemical |
| S. aureus V8 protease | Acidic residues — Glu (and Asp) | Enzyme |
CNBr: three. Two Met residues, so 2 cuts, so 3 pieces.
The general rule to state: fragments = cleavage sites + 1. Watch for the trap where the residue is already at the C-terminus — cutting after it produces nothing new.
After cleavage the peptides are purified by reversed-phase HPLC — occasionally by SDS-PAGE — and then sequenced. Note the practical drawback Harper's flags: because you must run several different fragmentation and purification conditions, direct chemical sequencing needs large quantities of purified protein. That, together with the slowness, is why the field moved on.
- Why cleave a large polypeptide? → Edman reads only ~30 residues, and cleavage also bypasses a blocked N-terminus
- Why use more than one cleavage reagent? → To generate overlapping peptides that establish the order of the fragments
- Trypsin cleaves after which residues? → Arg and Lys
- CNBr cleaves after which residue? → Met
- How are the peptides purified before sequencing? → Reversed-phase HPLC
Mass spectrometry — the method that took over ★★★
Harper's is blunt about it: the superior sensitivity, speed and versatility of mass spectrometry have replaced the Edman technique as the principal method for sequencing peptides and proteins. Understand why, and the section writes itself.
MS discriminates molecules on mass alone. That has one consequence that matters enormously in medicine: a post-translational modification — a phosphate group, a hydroxyl, a sugar — adds mass. So MS detects it directly. Edman sequencing struggles to identify which modification it has hit, and a DNA-derived sequence cannot see modifications at all, because they are added after translation. If a question asks why DNA sequencing has not made protein chemistry obsolete, this is the answer.
How the instrument works
The sample is vaporised under vacuum in the presence of a proton donor, so the molecules pick up positive charge. An electric field accelerates the cations down a flight tube. From there the two common designs diverge:
| Design | How it measures mass | Best for |
|---|---|---|
| Quadrupole (magnetic sector) | A magnetic field deflects the ions at right angles; the current needed to bend an ion's path onto the detector is proportional to its mass (for ions of equal charge) | Molecules of 4000 Da or less |
| Time-of-flight (TOF) | A straight flight tube. Time taken to reach the detector is inversely proportional to mass — heavy ions accelerate less and arrive later | Whole proteins, large masses |
Getting big molecules into the vapour phase
The obstacle that held MS back for years was simple: you can vaporise a small organic molecule by heating it in a vacuum, but a protein heated that way is destroyed. Three techniques solved it, and two of them are examinable by name.
| Method | How it avoids destroying the protein |
|---|---|
| Electrospray ionisation | The sample, dissolved in a volatile solvent, is sprayed through a capillary into the chamber. The solvent flashes away, leaving the macromolecule suspended in the gas phase. Convenient because peptides can be fed straight from an HPLC column into the spectrometer. |
| MALDI (matrix-assisted laser desorption/ionisation) | The sample is mixed with a liquid matrix containing a light-absorbing dye and a proton source. A laser excites the matrix, which disperses into the vapour phase so fast that the embedded protein is carried along without being heated. |
| Fast atom bombardment (FAB) | Macromolecules dispersed in glycerol or another protonic matrix are bombarded with a stream of neutral atoms |
The pay-off is remarkable precision. MALDI and electrospray allow the masses of polypeptides above 100 000 Da to be determined to within about ±1 Da — accurate enough to see a single added phosphate.
Sequencing by fragmentation
Knowing a peptide's total mass is not a sequence. To get the order, the peptide is broken up inside the instrument by collision with neutral helium atoms (collision-induced dissociation) and the fragments weighed. Peptide bonds are much more labile than carbon–carbon bonds, so the chain preferentially breaks between residues — meaning the most abundant fragments differ from one another by exactly one amino acid. Since the molecular mass of each amino acid is unique, the difference in mass between two successive fragments names the residue that was lost, and the sequence can be reconstructed from the ladder of masses.
Two mass spectrometers linked in series, allowing complex peptide mixtures to be analysed without prior purification. The first separates individual peptides by mass and directs a single chosen peptide into the second, where it is fragmented and the fragment masses determined.
Tandem MS is used to screen newborn blood samples for amino acids, fatty acids and other metabolites. Abnormal metabolite levels are diagnostic indicators for genetic disorders — Harper's names phenylketonuria, ethylmalonic encephalopathy and glutaric acidaemia type 1. This is the single most clinically important sentence in the chapter: the technique in your biochemistry lecture is the technique behind the heel-prick test.
- Why has MS replaced Edman? → Greater sensitivity, speed and versatility, and it detects post-translational modifications by their added mass
- Quadrupole vs TOF? → Quadrupole for small molecules; time-of-flight for whole proteins
- Name two ways of volatilising a protein → Electrospray ionisation and MALDI (also fast atom bombardment)
- How is sequence read from the fragments? → Successive fragments differ by one residue, and each amino acid has a unique mass
- Which two amino acids cannot be distinguished? → Leucine and isoleucine — isomers of identical mass
- What is tandem MS used for clinically? → Newborn screening for metabolic disorders such as phenylketonuria


Proteomics — the endpoint of the chapter
The chapter closes by scaling up from one protein to all of them. The genome is fixed and static; the set of proteins actually present is neither. Genes switch on and off, muscle cells express proteins neural cells never touch, the subunits of haemoglobin change between fetal and adult life, and proteins are modified after synthesis. Knowing the genome is therefore only the beginning.
The set of all the proteins expressed by an individual cell at a particular time. Because the body contains thousands of cell types each containing thousands of proteins — and because expression changes with growth, differentiation and external stimuli — the proteome is a moving target, not a fixed list like the genome.
The goal of proteomics is to identify proteins whose level of expression correlates with medically significant events, on the presumption that a protein appearing or disappearing alongside a disease is linked to its cause or mechanism. The problem is scale. Antibody and enzyme assays are exquisitely specific but can only look at proteins you already suspect; total-protein assays such as the Lowry or Bradford method, and stains such as Coomassie Blue, are universal but tell you nothing about which protein you are looking at.
First-generation proteomics threaded between the two: resolve everything on a two-dimensional gel, extract individual spots, and identify each by Edman sequencing or mass spectrometry, matching Mr and pI against the databases. A single gel resolves only about a thousand proteins, but its advantage is that it examines the proteins themselves. The complementary approach — gene arrays, or DNA chips — detects the mRNAs instead. Arrays are more sensitive and cover more gene products, but carry a real caveat: a change in mRNA level does not necessarily mean a comparable change in the protein.
Finally, bioinformatics lets you guess a new protein's function from its sequence alone. Nature reuses structural themes, so algorithms look for conserved amino acids at key positions that mark a known domain — the Rossmann fold that binds NAD(P)H, nuclear targeting sequences, EF hands that bind Ca²⁺. Find the motif and you have a strong hypothesis about what the protein does before you have run a single assay.
- Define the proteome → All the proteins expressed by an individual cell at a particular time
- Why is it a “moving target”? → Expression varies with cell type, time, differentiation and stimuli, and proteins are modified after synthesis
- Which two methods survey protein expression? → Two-dimensional electrophoresis (the proteins themselves) and gene array / DNA chips (the mRNAs)
- What is the caveat with gene arrays? → mRNA level does not necessarily reflect protein level
- Name two conserved domains bioinformatics looks for → The Rossmann fold (binds NAD(P)H) and EF hands (bind Ca²⁺)
Revision layer
The 2019 paper asked, in Section II, simply: “Describe the methods of determining the sequence of a polypeptide.” That is this whole page in one question, and it is worth 8 marks. The model answer is the four-step skeleton, each step named with its reagent.
The model answer — memorise this skeleton
| Step | What you do | Name the specifics |
|---|---|---|
| 1 · Purify | Isolate the protein from the cell extract | Salting out (ammonium sulfate), isoelectric precipitation; then column chromatography — ion exchange, gel filtration, affinity |
| 2 · Assess purity & dissociate | Confirm you have one protein, and break it into single chains | SDS-PAGE with Coomassie Blue; reduce disulfide bonds with 2-mercaptoethanol or DTT (or oxidise with performic acid) |
| 3 · Cleave | Cut the chain into peptides short enough to sequence, using two different reagents so the peptides overlap | Trypsin (after Arg/Lys), chymotrypsin (after aromatics), CNBr (after Met), V8 protease (after Glu); purify peptides by reversed-phase HPLC |
| 4 · Sequence & assemble | Read each peptide, then use the overlaps to order the fragments | Edman degradation (phenylisothiocyanate → PTH amino acid, ~5–30 residues) or tandem mass spectrometry; or the hybrid approach — short protein sequence + DNA cloning |
Definitions from this unit — Section I material
| Term | Definition |
|---|---|
| Primary structure | The linear sequence of amino acid residues in a polypeptide chain, read from the N-terminus to the C-terminus |
| Gel filtration (size-exclusion) chromatography | Separation of proteins according to their Stokes radius using porous beads; excluded (large) proteins elute first, included (small) proteins are retarded and elute later |
| Stokes radius | The radius of the sphere a protein occupies as it tumbles in solution; a function of both molecular mass and shape |
| Affinity chromatography | Purification exploiting a protein's specific binding to an immobilised ligand — substrate, product, coenzyme or inhibitor; only proteins that recognise the ligand adhere |
| SDS-PAGE | Polyacrylamide gel electrophoresis in the presence of sodium dodecyl sulfate, which denatures the protein and confers a uniform charge-to-mass ratio so that separation depends on relative molecular mass alone |
| Isoelectric focusing | Separation of proteins in a pH gradient generated with ampholytes, each protein migrating until it reaches the pH equal to its isoelectric point, where its net charge is zero |
| Edman reaction | Phenylisothiocyanate derivatises the N-terminal residue as a phenylthiohydantoic acid; acid in a non-hydroxylic solvent then releases a phenylthiohydantoin, identified chromatographically, plus a peptide one residue shorter — and the cycle repeats |
| Tandem mass spectrometry | Two mass spectrometers in series, allowing complex peptide mixtures to be analysed without prior purification: the first selects a peptide by mass, the second fragments it and weighs the fragments |
| Proteome | The set of all the proteins expressed by an individual cell at a particular time |
Separation methods at a glance — what separates on what
| Method | Separates by |
|---|---|
| Gel filtration / size exclusion | Stokes radius (large elute first) |
| Ion exchange | Net charge (elute by raising salt) |
| Hydrophobic interaction | Exposed hydrophobic surface (elute by lowering salt) |
| Affinity | Specific ligand binding |
| Reversed-phase HPLC | Hydrophobicity, at high pressure (purifies peptides) |
| SDS-PAGE | Relative molecular mass (Mr) |
| Isoelectric focusing | Isoelectric point (pI) |
| 2-D electrophoresis | pI in one dimension, Mr in the other |
Cleavage specificities — learn all four
| Reagent | Cleaves after |
|---|---|
| Trypsin | Arg, Lys |
| Chymotrypsin | Phe, Trp, Tyr (aromatic) |
| Cyanogen bromide | Met |
| S. aureus V8 protease | Glu (acidic) |
Numbers worth carrying in
| Figure | Value |
|---|---|
| Insulin A chain / B chain | 21 residues / 30 residues |
| Sanger's Nobel Prize | 1958 |
| SDS binding ratio | 1 SDS per 2 peptide bonds |
| Edman read length | ~5–30 residues |
| Quadrupole MS mass limit | 1000 Da (slide) · 4000 Da (Harper's) |
| MALDI / electrospray accuracy | >100 000 Da to within ±1 Da |
| Proteins resolved on one 2-D gel | ~1000 |
Practical 1 — Gel Filtration Chromatography is §3 of this page done with your own hands: you separate on Stokes radius and watch the large molecules come off first. Practical 2 — quantitative protein assay with a BCA kit is the “universal but non-identifying” category Harper's mentions alongside Lowry and Bradford — it tells you how much protein, never which. Expect the viva to ask why.
- Write the four-step skeleton for sequencing a polypeptide, naming a reagent at each step
- Define gel filtration, SDS-PAGE, isoelectric focusing, the Edman reaction and the proteome in exam wording
- State what each of the six chromatographies separates on
- Give the four cleavage reagents and their specificities, and calculate fragment numbers from a composition
- Explain why DNA sequencing cannot replace protein sequencing
- Explain why leucine and isoleucine defeat mass spectrometry