Proteins: Higher Orders of Structure
← Back 📋 Q-Bank 🏠 All Units
HIGH YIELD ★★★
Proteins & Enzymes · Unit 3 of 26

Proteins: Higher Orders of Structure

TMU Lecture 3 — Lin Yu, Dept of Biochemistry & Molecular Biology Harper's ch. 5 — Proteins: Higher Orders of Structure, pp. 35–48 Two study questions on the slide deck came straight from past papers
01

Form follows function

Unit 1 gave you the letters, Unit 2 the sentence. This unit is about what happens when the sentence folds up into a three-dimensional object — because a polypeptide chain lying flat does nothing at all. A protein works by positioning specific chemical groups in a precise three-dimensional arrangement, and everything from enzyme catalysis to oxygen transport to the tensile strength of a tendon follows from that single idea.

Harper's puts the difficulty starkly. A typical polypeptide can adopt ≥10⁵⁰ distinct conformations. If a newly made chain had to search them at random it would take billions of years to find the right one — yet real proteins fold in milliseconds. Something must be guiding the process. That tension between an astronomically large search space and an absurdly fast result is the intellectual heart of the chapter, and §10 resolves it.

Why the clinician cares

Failures of folding are diseases in their own right. Creutzfeldt-Jakob disease, scrapie and bovine spongiform encephalopathy are caused by a protein adopting the wrong shape; Alzheimer's disease features misfolded β-amyloid; and scurvy is a nutritional deficiency that stops collagen maturing properly. In each case the amino acid sequence may be perfectly normal. It is the conformation that is wrong — which is why these are called protein conformation diseases.

Test yourself
  • Why must a polypeptide fold? → Function depends on positioning specific chemical groups in a precise three-dimensional arrangement
  • How many conformations can a typical polypeptide adopt? → ≥10⁵⁰ — yet folding takes milliseconds
  • Name three protein conformation diseases → Creutzfeldt-Jakob disease, Alzheimer's disease, and (nutritionally) scurvy
02

Configuration versus conformation ★★

These two words are constantly confused, and the distinction is a clean two-mark question. The test is simple: do you have to break a covalent bond to get from one to the other?

ConfigurationConformation
Refers toThe geometric relationship between a given set of atomsThe spatial relationship of every atom in the molecule
Interconversion requiresBreaking covalent bondsNo bond rupture — rotation about single bonds
ExampleL- versus D-amino acidsA folded protein versus the same protein denatured
The one-line test

L- and D-alanine are different configurations: no amount of twisting turns one into the other, you would have to break bonds. A folded and an unfolded protein are different conformations of the same molecule: only rotations about single bonds separate them, and the configuration — L throughout — is retained.

Test yourself
  • What distinguishes configuration from conformation? → Changing configuration requires breaking covalent bonds; changing conformation does not
  • Which one distinguishes L- from D-amino acids? → Configuration
03

The four orders of protein structure ★★★

This is the single most predictable question in the whole of Module A — the TMU slide deck ends with it as a study question, and it has appeared on the papers. Learn all four with their defining feature and the force that holds each one together, because the second half of the question is almost always “what forces stabilise protein structure?”

OrderDefinitionStabilised mainly by
PrimaryThe sequence of amino acids in a polypeptide chainCovalent peptide bonds
SecondaryThe folding of short (3–30 residue), contiguous segments of polypeptide into geometrically ordered units — the α-helix, β-sheet, bends and loopsHydrogen bonds (backbone C=O to N–H)
TertiaryThe three-dimensional assembly of secondary structural units into larger functional units — the mature polypeptide and its domainsHydrophobic interactions chiefly, plus H-bonds, salt bridges, van der Waals, and sometimes disulfide bonds
QuaternaryThe number and types of polypeptide subunits of an oligomeric protein and their spatial arrangementThe same non-covalent forces; sometimes interchain disulfide bonds
⭐ The framing that gets full marks
Describe the primary, secondary, tertiary and quaternary structure of a protein.
State the definition and the stabilising force for each, then add the organising sentence the examiner is looking for: primary structure is the basic structure, held together by covalent bonds; secondary, tertiary and quaternary structure together constitute the spatial structure, or conformation, and are stabilised primarily by non-covalent forces. Finish with the consequence — conformation is dictated by the primary sequence, so a single amino acid substitution can abolish function. That last line converts a list into an answer.
TMU Lecture 3 Slides 6 and 35 · Harper's ch.5, p.36
Only one order of structure involves more than one polypeptide chain — which?
Quaternary. A monomeric protein has primary, secondary and tertiary structure but no quaternary structure. Myoglobin, one chain, has none; haemoglobin, four chains, does — which is exactly the comparison Unit 4 is built on.
Harper's ch.5, p.41
Test yourself
  • Define secondary structure → The folding of short, contiguous segments of polypeptide (3–30 residues) into geometrically ordered units
  • Are side chains part of secondary structure? → No — the backbone folds; side chains only influence its stability and type
  • Which force stabilises secondary structure? → Hydrogen bonds between backbone carbonyl and amide groups
  • Which force chiefly stabilises tertiary and quaternary structure? → Hydrophobic interactions
  • Which proteins lack quaternary structure? → Monomeric ones — a single polypeptide chain
04

What the backbone will and will not allow ★★

Unit 1 ended with the peptide bond's partial double-bond character. Now watch it pay off. Because the C–N bond cannot rotate, the carbonyl carbon, carbonyl oxygen and α-nitrogen are locked coplanar, and only two of the three backbone bonds are free to turn: the Cα–N bond, whose angle is phi (Φ), and the Cα–Co bond, whose angle is psi (Ψ).

Even those two are heavily constrained. For every residue except glycine, most combinations of φ and ψ are disallowed by steric hindrance — the side chains would simply collide. Proline is more restricted still, because its ring removes free rotation about the N–Cα bond altogether. Plot the permitted combinations and you get the Ramachandran plot.

The Ramachandran plot

A plot of the main-chain φ against ψ angles, on which dots mark allowable combinations and blank spaces mark prohibited ones.

Learn the two locations: the angles that define the α-helix fall in the lower left-hand quadrant; those of the β-sheet in the upper left-hand quadrant.

Why this matters more than it looks

Regions of ordered secondary structure arise simply when a run of consecutive residues adopts similar φ and ψ angles. That is the entire definition. The α-helix is not a special molecular machine — it is what you get when every residue in a stretch takes φ ≈ −57° and ψ ≈ −47°. Extended segments such as loops are the parts where the angles vary from residue to residue.

Test yourself
  • Which backbone bonds can rotate? → Cα–N (angle φ, phi) and Cα–Co (angle ψ, psi)
  • Why can the peptide C–N bond not rotate? → Its partial double-bond character keeps the carbonyl C, O and α-nitrogen coplanar
  • Where do the α-helix and β-sheet sit on a Ramachandran plot? → Lower left quadrant and upper left quadrant respectively
  • How does regular secondary structure arise? → When a series of consecutive residues adopts similar φ and ψ angles
Ramachandran plot of main-chain φ and ψ angles — dots mark allowable combinations, blank spaces prohibited ones. The α-helix falls in the lower left quadrant, the β-sheet in the upper left
Ramachandran plot of main-chain φ and ψ angles — dots mark allowable combinations, blank spaces prohibited ones. The α-helix falls in the lower left quadrant, the β-sheet in the upper left
Harper's Illustrated Biochemistry, Figure 5–1, p.37
05

The α-helix ★★★

Here the numbers matter, and they are small enough to memorise outright. The backbone is twisted by an equal amount about each α-carbon, with φ ≈ −57° and ψ ≈ −47°.

FeatureValue / fact
Residues per complete turn3.6 (average)
Pitch — the rise per turn0.54 nm
φ and ψ angles−57° and −47°
HandednessRight-handed only — proteins contain only L-amino acids, for which the right-handed helix is far more stable
Position of the R groupsFacing outward, on the outside of the helix
Stabilised byHydrogen bonds parallel to the helix axis, between the carbonyl oxygen of one residue and the amide hydrogen of the fourth residue down the chain
⭐ The hydrogen-bond question — get the “fourth residue” in
How is the α-helix stabilised?
By hydrogen bonds running parallel to the axis of the helix, formed between the oxygen of a peptide-bond carbonyl and the hydrogen on the peptide-bond nitrogen of the fourth residue further along the chain. Harper's adds the thermodynamic reason: the α-helix is the conformation that allows the maximum number of such hydrogen bonds, supplemented by van der Waals interactions in the tightly packed core — and it is that maximisation which drives its formation.

Note the answer is about the backbone. Side chains point outward and take no part.
Harper's ch.5, p.38 · TMU Lecture 3 Slide 14
Why do proline and glycine break α-helices?
Two different reasons, and the examiner wants both.
Proline: its peptide-bond nitrogen has no hydrogen, so it physically cannot donate the hydrogen bond the helix depends on. It can only be accommodated within the first turn; anywhere else it produces a bend.
Glycine: its R group is so small that the residue is too flexible — the restriction that normally holds a helix in register is lost — so glycine also frequently induces bends.
Harper's ch.5, p.38
Where you will see this again

Hold on to “α-helix breaker” for proline. In Unit 4 you will meet the haemoglobin subunits, which are almost entirely helical; in Unit 25 the same idea explains why certain mutations abolish a protein's structure. Schematic protein diagrams draw α-helices as coils or cylinders — worth knowing so you can read a figure in an exam.

Test yourself
  • How many residues per turn, and what is the pitch? → 3.6 residues; 0.54 nm
  • Which residue does the hydrogen bond reach? → The fourth residue down the chain
  • Which way do R groups point? → Outward
  • Why are only right-handed α-helices found? → Proteins contain only L-amino acids, for which the right-handed helix is much more stable
  • Name the two helix breakers and why → Proline (its N has no hydrogen to donate) and glycine (too small and flexible)
Main-chain atoms about the axis of an α-helix — pitch 0.54 nm, 3.6 residues per turn
Main-chain atoms about the axis of an α-helix — pitch 0.54 nm, 3.6 residues per turn
Harper's Illustrated Biochemistry, Figure 5–2, p.38
Viewed down the axis: the R groups lie on the OUTSIDE, and there is almost no free space inside the helix
Viewed down the axis: the R groups lie on the OUTSIDE, and there is almost no free space inside the helix
Harper's Illustrated Biochemistry, Figure 5–3, p.38
Hydrogen bonds (dotted) between the carbonyl oxygen and the amide hydrogen of the fourth residue down the chain
Hydrogen bonds (dotted) between the carbonyl oxygen and the amide hydrogen of the fourth residue down the chain
Harper's Illustrated Biochemistry, Figure 5–4, p.38
06

The β-pleated sheet ★★★

The second regular structure — hence “beta” — is built on the same hydrogen bond but arranged completely differently. Where the α-helix backbone is compact, the β-sheet backbone is highly extended, and instead of bonding to a residue four places along the same chain, it bonds to an adjacent segment of chain lying alongside.

Viewed edge-on, the residues form a zigzag or pleated pattern — hence the name — with the R groups of adjacent residues pointing in opposite directions, alternately above and below the sheet.

α-helixβ-sheet
BackboneCompact, coiledHighly extended, pleated
Hydrogen bonds formedWithin the same chain, to the 4th residue alongBetween adjacent segments of chain
Direction of H-bondsParallel to the helix axisRoughly perpendicular to the chains (antiparallel sheet)
R groupsAll point outwardAlternate — adjacent residues point in opposite directions
Parallel and antiparallel β-sheets

In a parallel β-sheet the adjacent segments of chain run in the same direction, amino to carboxyl. In an antiparallel sheet they run in opposite directions.

The hydrogen bonds differ accordingly: in the antiparallel sheet, pairs of hydrogen bonds alternate between close together and wide apart and lie approximately perpendicular to the backbone; in the parallel sheet they are evenly spaced but slant in alternate directions.

Test yourself
  • What shape does a β-sheet backbone take? → Highly extended, forming a zigzag or pleated pattern edge-on
  • Where do β-sheet hydrogen bonds form? → Between carbonyl oxygens and amide hydrogens of ADJACENT segments of chain
  • Which way do R groups point? → Adjacent residues point in opposite directions
  • Parallel vs antiparallel? → Adjacent strands run in the same direction, or in opposite directions, amino to carboxyl
Antiparallel (top) and parallel (bottom) β-sheets. In the antiparallel sheet, pairs of hydrogen bonds alternate close together and wide apart and lie roughly perpendicular to the backbone; in the parallel sheet they are evenly spaced but slanted
Antiparallel (top) and parallel (bottom) β-sheets. In the antiparallel sheet, pairs of hydrogen bonds alternate close together and wide apart and lie roughly perpendicular to the backbone; in the parallel sheet they are evenly spaced but slanted
Harper's Illustrated Biochemistry, Figure 5–5, p.39
07

Turns, loops, motifs and domains ★★

Helices and sheets are the rigid pieces. Something has to connect them, and those connectors are not junk — they carry a surprising amount of function.

TermDefinition
Turns / bendsSegments of amino acids that connect adjacent regions of secondary structure. Proline and glycine are often present — the two residues that break helices are exactly the two that make bends. A β-turn typically spans four residues, hydrogen bonded between the first and fourth.
LoopsRegions containing residues beyond the minimum number necessary to connect adjacent regions of secondary structure
Random coilA part of a protein that is loose — often at the N- or C-terminal — and does not contribute to tertiary structure
Supersecondary structureCombinations of secondary structure, typically 10–40 residues, found recurrently in many proteins — intermediate between secondary and tertiary structure. The commonest are combinations of α-helix and β-sheet, e.g. the helix-loop-helix
MotifA short, highly conserved region, frequently the most conserved part of a domain, and critical to its function — in an enzyme it may contain the active site. Examples: the zinc finger, nuclear localisation sequences
DomainA section of protein structure that folds independently into a stable conformation and is sufficient to perform a particular chemical or physical task, such as binding a substrate or a regulatory molecule
⭐ Domain — a near-certain Section I definition
Define a domain and give an example.
A domain is a section of protein structure that folds independently into a stable conformation, sufficient to perform a particular chemical or physical task; different domains work together to give the complete function of the protein.

Harper's example, which makes the point perfectly: protein kinases contain two domains. The amino-terminal domain, rich in β-sheet, binds ATP; the carboxyl-terminal domain, rich in α-helix, binds the peptide or protein substrate. Two jobs, two domains — and each with its own characteristic secondary structure. A small protein such as triose phosphate isomerase or myoglobin consists of a single domain.
Harper's ch.5, pp. 40–41 · TMU Lecture 3 Slides 25, 27
The zinc finger — the numbers to quote
About 30 amino acid residues form an elongated loop held together at the base by a single Zn²⁺ ion, coordinated to four of the residues — either four cysteines, or two cysteines and two histidines. You will meet it again in Unit 26 as a DNA-binding motif of transcription factors.
TMU Lecture 3 Slide 22
Why loops are the target of the immune system

Many loops and bends sit on the surface of a protein, exposed to solvent. That makes them readily accessible sites — epitopes — for recognition and binding by antibodies. When you meet antigen recognition in Immunology, this is the structural reason antibodies bind where they do.

One refinement worth carrying. Not every part of a protein is ordered at all. Proteins may contain “disordered” regions, often at the extreme amino or carboxyl terminal, with high conformational flexibility. In many cases these assume an ordered conformation only on binding a ligand — which lets them act as ligand-controlled switches that change the protein's structure and function.

Test yourself
  • Which two amino acids are often found in β-turns? → Proline and glycine
  • Define a domain → A section of protein structure that folds independently into a stable conformation, sufficient to perform a particular task
  • How many domains does a protein kinase have, and what does each bind? → Two — the β-rich N-terminal binds ATP, the α-rich C-terminal binds the substrate
  • What is supersecondary structure? → Recurring combinations of secondary structure, 10–40 residues, intermediate between secondary and tertiary
  • Why are loops important immunologically? → They lie on the surface and form epitopes for antibody binding
A β-turn linking two segments of antiparallel β-sheet — hydrogen bonded between the first and fourth residues of the four-residue segment
A β-turn linking two segments of antiparallel β-sheet — hydrogen bonded between the first and fourth residues of the four-residue segment
Harper's Illustrated Biochemistry, Figure 5–7, p.41
A polypeptide containing two domains — each folds independently into a stable conformation
A polypeptide containing two domains — each folds independently into a stable conformation
Harper's Illustrated Biochemistry, Figure 5–8, p.42
08

Tertiary and quaternary structure ★★★

Tertiary structure

The entire three-dimensional conformation of a polypeptide — how the secondary structural features (helices, sheets, bends, turns and loops) assemble into domains, and how those domains relate spatially to one another.

Quaternary structure

The number and types of polypeptide subunits (protomers) of an oligomeric protein and their spatial arrangement. Only proteins built from more than one chain possess it.

The nomenclature is straightforward and occasionally examined. A monomeric protein has one chain; a dimeric protein two. A homodimer contains two copies of the same chain; in a heterodimer the two differ. Greek letters distinguish the different subunits of a hetero-oligomer and subscripts give how many of each — so α₄ is a homotetramer, and α₂β₂γ is a protein of five subunits of three different types.

Read the subscript notation and you have already learned haemoglobin

Adult haemoglobin is α₂β₂ — a heterotetramer, two α chains and two β chains. Unit 4 is entirely about what that quaternary structure buys you that myoglobin, a single chain with no quaternary structure at all, cannot do. If the notation makes sense now, the next unit will be much easier.

Test yourself
  • Define tertiary structure → The entire three-dimensional conformation of one polypeptide — how secondary units assemble into domains and how domains relate
  • Define quaternary structure → The number and types of subunits of an oligomeric protein and their spatial arrangement
  • What does α₂β₂γ mean? → Five subunits of three different types
  • What is a homodimer? → A protein of two identical polypeptide chains
Tertiary structure: the elegant, symmetrical alternation of β-sheets and α-helices in triose phosphate isomerase
Tertiary structure: the elegant, symmetrical alternation of β-sheets and α-helices in triose phosphate isomerase
Harper's Illustrated Biochemistry, Figure 5–6, p.40
09

The forces that stabilise structure ★★★

This is the second half of the recurring exam question, and it is worth being precise. Higher orders of protein structure are stabilised primarily — and often exclusively — by non-covalent interactions. Individually each is weak; there are simply a great many of them.

ForceNatureNote
Hydrogen bondsNon-covalentDominate secondary structure; also stabilise loops and tertiary contacts
Hydrophobic interactionsNon-covalentThe main force in tertiary and quaternary structure — non-polar side chains are driven into the interior, away from solvent
Salt bridges (electrostatic / ionic bonds)Non-covalentBetween oppositely charged side chains, e.g. Lys⁺ with Asp⁻
van der Waals interactionsNon-covalentWeak and short-range, but numerous in a tightly packed core
Disulfide (S–S) bondsCovalentBetween the sulfhydryl groups of two cysteinyl residues — the exception to the rule
⭐ The trick question the slide deck itself ends on
“Non-covalent bonds that stabilise secondary, tertiary and quaternary structure do NOT include…”
Disulfide bonds. They stabilise structure, certainly — but they are covalent, not non-covalent. Every other option (hydrophobic forces, electrostatic bonds, van der Waals forces) is non-covalent. This exact question closes the TMU lecture slide deck, which makes it about as strong a signal as you will ever get.
TMU Lecture 3 Slide 49
Distinguish intrachain from interchain disulfide bonds.
Intrachain disulfide bonds form within a single polypeptide and further enhance the stability of its folded conformation — i.e. they reinforce tertiary structure. Interchain disulfide bonds link different polypeptides and so stabilise the quaternary structure of certain multimeric proteins. Insulin's A and B chains are the classic interchain example.
TMU Lecture 3 Slide 33
The three-line summary to write down first

Primary structure is stabilised by covalent peptide bonds.
Secondary structure is stabilised mainly by hydrogen bonds.
Tertiary and quaternary structure are stabilised mainly by hydrophobic interactions.

Three lines, straight from the lecture summary slide. Write them, then expand each with the supporting forces.

Test yourself
  • Which force dominates secondary structure? → Hydrogen bonds
  • Which force dominates tertiary and quaternary structure? → Hydrophobic interactions
  • Which stabilising bond is covalent? → The disulfide bond between two cysteinyl residues
  • Intrachain vs interchain disulfide bonds? → Intrachain stabilise tertiary structure; interchain stabilise quaternary structure
10

Folding, the helpers, and what happens when it fails ★★

Back to the 10⁵⁰ problem. The resolution is that folding is modular and proceeds in two stages, so the chain never has to search the whole space.

StageWhat happens
1 · Local orderAs the new polypeptide emerges from the ribosome, short segments fold into secondary structural units. The problem is now reduced to arranging a relatively small number of pre-formed elements
2 · The molten globuleForces driving hydrophobic regions into the interior, away from solvent, collapse the chain into a “molten globule” — a partially folded state in which the modules of secondary structure rearrange until the mature conformation is reached
The molten globule

A partially folded polypeptide formed when hydrophobic regions segregate into the interior away from solvent, and within which the modules of secondary structure rearrange until the native conformation is attained.

The process is orderly but not rigid — considerable flexibility exists in the order in which elements can rearrange. For oligomeric proteins, individual protomers tend to fold before they associate with other subunits.

Why does it reach the right answer? Because the native conformation is the thermodynamically favoured one — generally the most energetically favourable — so the information specifying it is already contained in the primary sequence. That is the sentence to write: conformation is dictated by primary structure.

Three proteins that assist folding

HelperWhat it does
ChaperonesFound from bacteria to humans, and involved in the folding of over half of mammalian proteins. They bind proteins before synthesis is complete, preventing premature folding into an incorrect conformation; and they rescue proteins thermodynamically trapped in a misfolded dead end by unfolding hydrophobic regions and giving them a second chance
Protein disulfide isomeraseDisulfide bond formation is non-specific — under oxidising conditions a cysteine can bond to any accessible cysteine. By catalysing disulfide exchange (rupture of an S–S bond and its reformation with a different partner), this enzyme drives the protein towards the correct, native pairings
Proline-cis,trans-isomeraseAll X-Pro peptide bonds are synthesised in the trans configuration, yet about 6% of the X-Pro bonds of mature proteins are cis — particularly common in β-turns. This enzyme catalyses the trans → cis isomerisation

Note the contrast between the cell and the test tube. In vitro, many denatured proteins will refold spontaneously — but the process is far slower than in vivo, and some proteins fail entirely, forming insoluble aggregates: disordered complexes of unfolded or partly folded chains held together by hydrophobic interactions. That word aggregate is the bridge to the next paragraph.

When folding goes wrong

Prion disease — the mechanism, stated exactly

Prions are protein particles that lack nucleic acid yet cause fatal transmissible diseases: Creutzfeldt-Jakob disease in humans, scrapie in sheep, bovine spongiform encephalopathy in cattle.

The normal human prion-related protein PrPc is monomeric and rich in α-helix, with a normal role in the nervous system. The pathological form PrPSc is rich in β-sheet, with many hydrophobic side chains exposed to solvent — so the molecules associate strongly with one another to form insoluble, protease-resistant aggregates.

The transmission mechanism is the examinable part: PrPSc serves as a template that converts normal PrPc into PrPSc. Nothing is copied; a shape is propagated. Hence “protein conformation disease”: the sequence is normal, the conformation is not.

Alzheimer's disease

Misfolding of another brain protein, β-amyloid — a 4.3-kDa polypeptide produced by proteolytic cleavage of a larger amyloid precursor protein — is a prominent feature of Alzheimer's disease. Same theme: a normal protein, wrongly folded, aggregating.

Finally, how these structures are known at all. X-ray crystallography requires precipitating the protein under conditions in which it forms ordered crystals that diffract X-rays. NMR spectroscopy, its powerful complement, measures the absorbance of radiofrequency energy by certain nuclei — the “NMR-active” isotopes being ¹H, ¹³C, ¹⁵N and ³¹P — and has the advantage of analysing proteins in solution. Molecular modelling adds a computational layer, including homology modelling, in which a known structure serves as a template for a related protein.

Test yourself
  • What are the two stages of protein folding? → Short segments fold into secondary units; then hydrophobic collapse into a molten globule, within which the units rearrange
  • Define the molten globule → A partially folded polypeptide in which secondary structural modules rearrange to reach the native conformation
  • Name three proteins that assist folding → Chaperones, protein disulfide isomerase, proline-cis,trans-isomerase
  • What proportion of X-Pro bonds in mature proteins are cis? → About 6%, commonest in β-turns
  • How do PrPᶜ and PrPˢᶜ differ? → PrPᶜ is monomeric and α-helix rich; PrPˢᶜ is β-sheet rich and forms insoluble protease-resistant aggregates
  • Name two methods for determining three-dimensional structure → X-ray crystallography and NMR spectroscopy
Isomerisation of an X-Pro peptide bond from cis to trans, catalysed by proline-cis,trans-isomerase
Isomerisation of an X-Pro peptide bond from cis to trans, catalysed by proline-cis,trans-isomerase
Harper's Illustrated Biochemistry, Figure 5–10, p.45
11

Collagen — post-translational processing made visible ★★★

Harper's devotes the end of the chapter to collagen for a reason: it demonstrates every theme of the unit at once — an unusual secondary structure, extensive post-translational modification, and two diseases that follow directly from getting it wrong. It is also the most clinically examinable material in Module A.

FactDetail
AbundanceThe most abundant fibrous protein — more than 25% of the protein mass of the human body
Repeating unitTropocollagen — three collagen polypeptides, each about 1000 amino acids, bundled into the collagen triple helix
The triple helixThree strands twisting to the left wrap around one another in a right-handed fashion. The opposing handedness makes the structure highly resistant to unwinding — the same principle as the steel cables of a suspension bridge
Geometry3.3 residues per turn, with a rise per residue nearly twice that of an α-helix
SequenceEvery third residue is glycine, giving the repeating Gly-X-Y pattern, in which Y is generally proline or hydroxyproline
Axial ratioA mature fibre is an elongated rod of axial ratio about 200
⭐ Why must every third residue be glycine?
Explain the necessity of glycine at every third position.
Because of packing. The R groups of the three strands of the triple helix pack so closely together that, in order to fit at the point where the three chains come into contact, one of the three must be a hydrogen atom — and glycine is the only amino acid whose R group is a hydrogen. Staggering the three strands positions the requisite glycines throughout the helix.

This is Unit 1 paying off again: glycine is small, so it fits where nothing else can. Here that property is not a curiosity but a structural requirement, and a mutation replacing a single glycine causes osteogenesis imperfecta.
Harper's ch.5, p.47

Maturation — where vitamin C and copper come in

Collagen is synthesised as a larger precursor, procollagen, and then modified in a defined order. Numerous prolyl and lysyl residues are hydroxylated by prolyl hydroxylase and lysyl hydroxylase — and both enzymes require ascorbic acid (vitamin C). The resulting hydroxyprolyl and hydroxylysyl residues provide extra hydrogen-bonding capacity that stabilises the mature protein. Glucosyl and galactosyl transferases then attach sugars to specific hydroxylysyl residues.

The central portion of the precursor then associates with others to form the triple helix, and the globular amino- and carboxyl-terminal extensions are removed by selective proteolysis. Finally lysyl oxidase, a copper-containing enzyme, converts certain lysyl ε-amino groups to aldehydes, which form the covalent cross-links that give the mature fibre its strength.

Three diseases straight out of that paragraph

Scurvy — dietary deficiency of vitamin C, which prolyl and lysyl hydroxylases require. Fewer hydroxyproline and hydroxylysine residues means fewer stabilising hydrogen bonds, so collagen fibres are conformationally unstable: bleeding gums, swollen joints, poor wound healing, and ultimately death.

Menkes syndromekinky hair and growth retardation, reflecting a dietary deficiency of the copper required by lysyl oxidase, so the covalent cross-links never form.

Ehlers-Danlos syndrome — a group of connective tissue disorders with hypermobile joints and skin abnormalities, from defects in the genes encoding α collagen-1, procollagen N-peptidase, or lysyl hydroxylase. And several forms of osteogenesis imperfecta, characterised by fragile bones.

The pattern to notice

Each of those three diseases knocks out one step of the maturation sequence — the hydroxylation (scurvy), the cross-linking (Menkes), or the gene/enzyme itself (Ehlers-Danlos). If you can recite the maturation steps in order, you can derive the diseases rather than memorising them, and vice versa. That is the way to hold this section.

Test yourself
  • What fraction of body protein is collagen? → More than 25% of protein mass — the most abundant fibrous protein
  • What is the repeating sequence? → Gly-X-Y, with Y usually proline or hydroxyproline
  • Why must every third residue be glycine? → The three strands pack so closely that one R group at the contact point must be a hydrogen
  • Which vitamin do prolyl and lysyl hydroxylase require? → Vitamin C (ascorbic acid) — its deficiency causes scurvy
  • Which metal does lysyl oxidase require? → Copper — deficiency causes Menkes syndrome
  • Why is the triple helix so resistant to unwinding? → The strands twist left while the superhelix winds right — opposing handedness, like a steel cable
Collagen — the repeating Gly-X-Y sequence, the extended secondary structure, and the three strands wound into the triple helix
Collagen — the repeating Gly-X-Y sequence, the extended secondary structure, and the three strands wound into the triple helix
Harper's Illustrated Biochemistry, Figure 5–11, p.46
12

Revision layer

Two questions close the TMU deck, and both are past-paper questions: “Describe the primary, secondary, tertiary and quaternary structure of protein” and “What are the major forces that stabilise protein structure?” The two tables below answer them. Everything else on this page is supporting detail.

The four orders — the core answer

OrderDefinitionStabilised by
PrimaryThe sequence of amino acids in the polypeptide chainCovalent peptide bonds
SecondaryFolding of short (3–30 residue) contiguous segments into geometrically ordered units — α-helix, β-sheet, bends, loopsHydrogen bonds
TertiaryThe whole three-dimensional conformation of one polypeptide; assembly of secondary units into domainsHydrophobic interactions (+ H-bonds, salt bridges, van der Waals, disulfides)
QuaternaryNumber, type and spatial arrangement of the subunits of an oligomeric proteinSame non-covalent forces; interchain disulfides

The stabilising forces

ForceCovalent?
Hydrogen bondsNo
Hydrophobic interactionsNo
Salt bridges (electrostatic / ionic)No
van der Waals interactionsNo
Disulfide bondsYES — the exception

α-helix versus β-sheet — the comparison table

α-helixβ-sheet
Residues per turn3.6
Pitch0.54 nm
φ / ψ≈ −57° / −47°
Ramachandran quadrantLower leftUpper left
BackboneCompact, coiled, right-handed onlyHighly extended, pleated
H-bond partnerThe 4th residue down the same chainAn adjacent segment of chain
H-bond directionParallel to the helix axisPerpendicular to the strands (antiparallel)
R groupsAll face outwardAdjacent residues point in opposite directions
Broken byProline (no N–H to donate), glycine (too flexible)

Definitions from this unit — Section I material

TermDefinition
ConformationThe spatial relationship of every atom in a molecule; interconversion between conformers occurs without breaking covalent bonds, by rotation about single bonds
ConfigurationThe geometric relationship between a given set of atoms; interconversion requires breaking covalent bonds — e.g. L- versus D-amino acids
Secondary structureThe folding of short (3–30 residue) contiguous segments of polypeptide backbone into geometrically ordered units
DomainA section of protein structure that folds independently into a stable conformation, sufficient to perform a particular chemical or physical task
MotifA short, highly conserved region within a domain, critical to its function; in an enzyme it may contain the active site
Supersecondary structureRecurring combinations of secondary structure, 10–40 residues long, intermediate between secondary and tertiary structure
Quaternary structureThe number and types of polypeptide subunits of an oligomeric protein and their spatial arrangement
Molten globuleA partially folded polypeptide, formed by segregation of hydrophobic regions away from solvent, within which the modules of secondary structure rearrange to reach the native conformation
PrionA protein particle that lacks nucleic acid and causes fatal transmissible disease by acting as a template that converts the normal host protein into the pathological conformation

Numbers worth carrying in

FigureValue
Conformations available to a typical polypeptide≥10⁵⁰ — yet folding takes milliseconds
Secondary structure segment length3–30 residues
α-helix3.6 residues per turn · pitch 0.54 nm · φ −57° ψ −47°
Globular protein axial rationot more than 3
Fibrous protein axial ratio10 or more (collagen fibre ≈ 200)
Supersecondary structure10–40 residues
Zinc finger~30 residues, one Zn²⁺, four coordinating residues
X-Pro bonds in cisabout 6%
Collagen triple helix3.3 residues per turn · every 3rd residue Gly · ~1000 aa per chain
Collagen as % of body proteinmore than 25%
β-amyloid4.3 kDa
Final check — can you do these cold?
  • Define all four orders of structure and name the force stabilising each
  • List the five stabilising forces and say which one is covalent
  • Give the α-helix numbers and explain the hydrogen bond precisely
  • Contrast α-helix and β-sheet on backbone, H-bonding and R groups
  • Define domain, motif, supersecondary structure and molten globule
  • Explain the prion mechanism without using the word “infection”
  • Trace collagen maturation and derive scurvy, Menkes and Ehlers-Danlos from it