Amino Acids, Proteins & Nucleic Acids
Sickle cell disease is caused by a single amino acid change (glutamate → valine) in haemoglobin. PKU results from inability to metabolise phenylalanine. The entire genetic code is written in just four nucleotides. This chapter bridges organic chemistry with molecular biology and clinical medicine.
Amino Acid Structure
An amino acid is a molecule that has both an amino group (–NH₂) and a carboxylic acid group (–COOH) on the same carbon — the α-carbon. This central carbon also bears a hydrogen and a variable side chain called the R group. It is the R group that distinguishes the 20 standard amino acids from each other and determines each amino acid's chemical character: size, charge, polarity, and reactivity.
Because the α-carbon has four different groups attached (–NH₂, –COOH, –H, and –R), it is a chiral centre in all amino acids except glycine (where R = H, making two groups the same). This chirality means amino acids can exist as two non-superimposable mirror-image forms. Almost all biologically active amino acids are in the L-configuration (more on this in §8.3).
Classification by R group:
• Non-polar/hydrophobic: Gly, Ala, Val, Leu, Ile, Pro, Phe, Met, Trp
• Polar, uncharged: Ser, Thr, Cys, Tyr, Asn, Gln
• Positively charged (basic): Lys, Arg, His
• Negatively charged (acidic): Asp, Glu
• Which amino acid has no chiral centre? → Glycine (R = H; two identical groups on α-C).
• What determines the chemical character of each amino acid? → The R group (side chain).
Zwitterion & Isoelectric Point (pI)
An amino acid has two ionisable groups: –COOH (pKa ~2) and –NH₂/–NH₃⁺ (pKa ~9–10). At the pH of a typical solution, neither group is in its fully protonated state. At physiological pH, most amino acids exist as a zwitterion (from German: "double ion"): the –COOH has donated its proton (now –COO⁻) and the –NH₂ has accepted a proton (now –NH₃⁺). The molecule carries both a positive and a negative charge simultaneously — it is internally neutralised but electrically polar.
The isoelectric point (pI) is the pH at which the amino acid carries zero net charge — where the number of positive charges exactly equals the number of negative charges. At the pI, the amino acid is in zwitterionic form with no net charge, and it does not migrate in an electric field. For simple amino acids with one amino and one carboxyl group, pI = (pKa₁ + pKa₂) / 2. In electrophoresis, amino acids migrate toward the anode (–) if above their pI (net negative) and toward the cathode (+) if below their pI (net positive).
pKa₁ < pH < pKa₂: zwitterion — +NH₃–CHR–COO⁻ (net 0) ← dominant at body pH
pH > pKa₂ (~9): fully deprotonated — NH₂–CHR–COO⁻ (net –1)
pI = (pKa₁ + pKa₂) / 2 for a simple amino acid (no ionisable R group)
• At what pH does an amino acid not move in an electric field? → At its pI (isoelectric point).
• Direction of migration at pH > pI? → Toward the anode (positive electrode) — the amino acid carries a net negative charge.
• pI formula for simple amino acid? → pI = (pKa₁ + pKa₂) / 2.
D/L Configuration
The D/L naming system for amino acids is based on comparison to glyceraldehyde, not on the Cahn-Ingold-Prelog (R/S) system. In a Fischer projection of an L-amino acid, the –NH₂ group is on the left side of the α-carbon. For D-amino acids, –NH₂ is on the right. All amino acids found in proteins are L-amino acids — this is a universal feature of life. D-amino acids do exist (in bacterial cell walls, some antibiotics like gramicidin and vancomycin, and a few neuropeptides), but they are exceptions.
L-amino acid: –NH₂ on the LEFT | D-amino acid: –NH₂ on the RIGHT
Memory: L = Left = Life (all protein amino acids are L)
• Are protein amino acids D or L? → All L (with rare exceptions in bacterial peptides).
• D-amino acids are found where? → Bacterial cell walls (D-Ala, D-Glu in peptidoglycan); vancomycin targets D-Ala–D-Ala.
Essential Amino Acids
Essential amino acids cannot be synthesised by the human body at all or in sufficient quantities; they must come from dietary protein. There are 9 essential amino acids. Deficiency causes protein malnutrition — even if total caloric intake is adequate, a diet lacking essential amino acids results in growth failure, muscle wasting, and impaired immunity (kwashiorkor is protein malnutrition; marasmus is combined calorie + protein deficiency).
Phenylalanine · Valine · Threonine · Tryptophan · Isoleucine · Methionine · Histidine · Leucine · Lysine
Conditionally essential (in infants or illness): Arginine, Cysteine, Tyrosine, Glutamine
• Name them using PVT TIM HaLL. → Phenylalanine, Valine, Threonine, Tryptophan, Isoleucine, Methionine, Histidine, Leucine, Lysine.
• PKU enzyme deficiency? → Phenylalanine hydroxylase (cannot convert Phe → Tyr).
Peptide Bond Formation
A peptide bond forms when the carboxyl group (–COOH) of one amino acid reacts with the amino group (–NH₂) of another, releasing water. This is a condensation reaction — the same amide bond formation we saw in Chapter 6, just in biological context. The resulting dipeptide still has a free –NH₂ at one end (N-terminus) and a free –COOH at the other (C-terminus). Chains of many amino acids joined by peptide bonds are polypeptides; folded polypeptides with biological function are proteins.
The peptide bond is planar due to resonance between the C–N bond and the adjacent carbonyl (just like amide bonds in Chapter 6). This planarity constrains protein backbone geometry and is fundamental to secondary structure formation. In biology, peptide bonds are formed by the ribosome and are hydrolysed by proteases (trypsin, chymotrypsin, pepsin).
• Trans configuration predominates (R groups on opposite sides)
• No rotation around the C–N bond
• Hydrolysis: acid (6M HCl, 110°C) or proteases at physiological conditions
• Ninhydrin test: detects free α-amino groups → purple colour (proline → yellow)
• Why is the peptide bond planar? → Resonance: N lone pair delocalises into C=O → partial C–N double bond → no rotation.
• What is the N-terminus vs C-terminus? → N-terminus = free –NH₂ end; C-terminus = free –COOH end. Convention: write N → C left to right.
Protein Structure: Four Levels
Protein function depends entirely on three-dimensional shape, which is determined by the amino acid sequence. The shape is described at four levels of organisation, each building on the previous. Understanding these levels helps explain how mutations (like in sickle cell disease) disrupt function and how denaturing agents (heat, urea, acids) disrupt protein structure.
| Level | Description | Stabilised by | Disrupted by |
|---|---|---|---|
| 1° (Primary) | Sequence of amino acids in the chain | Covalent peptide bonds | Acid/base hydrolysis, proteases |
| 2° (Secondary) | Local folding: α-helix, β-sheet | H-bonds between backbone C=O and N–H | Heat, pH extremes, urea |
| 3° (Tertiary) | Overall 3D fold of one chain | H-bonds, ionic bonds, hydrophobic interactions, disulfide bridges (Cys–Cys) | Heat, urea, SDS, reducing agents |
| 4° (Quaternary) | Assembly of multiple chains (subunits) | Same as 3° between subunits | Same denaturing conditions |
• What interactions stabilise secondary structure (α-helix/β-sheet)? → H-bonds between backbone C=O and N–H groups.
• Which covalent bond stabilises tertiary structure? → Disulfide bridge (–S–S–) between two Cys residues.
• Sickle cell disease is a defect at which structural level? → Primary (amino acid sequence — Val replaces Glu at position 6).
Nucleic Acids — Building Blocks
Nucleic acids (DNA and RNA) are the molecules that store, transmit, and express genetic information. They are polymers of nucleotides. Each nucleotide has three components: a nitrogenous base, a five-carbon sugar (ribose in RNA, deoxyribose in DNA), and one or more phosphate groups. The backbone of the nucleic acid chain alternates: sugar–phosphate–sugar–phosphate, with bases hanging off the side.
The nitrogenous bases are divided into two structural families. Purines (Adenine, Guanine) have a double-ring structure — a six-membered ring fused to a five-membered ring. Pyrimidines (Cytosine, Thymine, Uracil) have a single six-membered ring. The way to remember which is which: pyrimidine has a Y in the middle, and the three pyrimidines (C, T, U) are the smaller single-ring bases. Purines are bigger (double ring): A and G.
Nucleotide = base + sugar + phosphate (the monomer of nucleic acids)
| Base | Type | In DNA? | In RNA? |
|---|---|---|---|
| Adenine (A) | Purine | Yes | Yes |
| Guanine (G) | Purine | Yes | Yes |
| Cytosine (C) | Pyrimidine | Yes | Yes |
| Thymine (T) | Pyrimidine | Yes | No |
| Uracil (U) | Pyrimidine | No | Yes |
• What is a nucleoside? → Base + sugar only (no phosphate).
• Name the two purines. → Adenine (A) and Guanine (G) — double ring.
• Name the pyrimidines in DNA. → Cytosine (C) and Thymine (T).
DNA vs RNA & Watson-Crick Base Pairing
DNA (deoxyribonucleic acid) uses deoxyribose (no –OH at C-2) and is double-stranded, arranged in a famous antiparallel double helix stabilised by hydrogen bonds between complementary base pairs and hydrophobic stacking interactions between adjacent bases. RNA (ribonucleic acid) uses ribose (has –OH at C-2, making it more reactive and less stable), is generally single-stranded, and comes in three key forms: mRNA (messenger), tRNA (transfer), rRNA (ribosomal).
Watson-Crick base pairing is specific: A pairs with T (in DNA) or U (in RNA) via two hydrogen bonds; G pairs with C via three hydrogen bonds. The G≡C pair has one more H-bond, making it stronger — GC-rich DNA has a higher melting temperature (Tm) than AT-rich DNA. This is the basis of DNA hybridisation, PCR, and sequencing. Chargaff's rules follow from this: in any double-stranded DNA, [A] = [T] and [G] = [C].
| Feature | DNA | RNA |
|---|---|---|
| Sugar | Deoxyribose (no 2′-OH) | Ribose (has 2′-OH) |
| Strands | Double (antiparallel) | Single |
| Unique base | Thymine (T) | Uracil (U) |
| Location | Nucleus (+ mitochondria) | Nucleus, cytoplasm, ribosomes |
| Function | Genetic storage | Gene expression (mRNA/tRNA/rRNA) |
• Which base is in DNA but not RNA? → Thymine (T). Which is in RNA but not DNA? → Uracil (U).
• How many H-bonds in A=T? In G≡C? → A=T has 2; G≡C has 3.
• Why does GC-rich DNA have a higher melting temperature? → More H-bonds per base pair (3 vs 2) → requires more energy to separate strands.
Past-paper Drill
Essential (9): PVT TIM HaLL
Peptide bond: planar (resonance); 1° = covalent; 2° = H-bonds; 3° = H + ionic + hydrophobic + S–S; 4° = subunit assembly
DNA vs RNA: D = deoxyribose + T + double-stranded; R = ribose + U + single
Base pairs: A=T (2H) · G≡C (3H); higher GC → higher Tm