Mutation, Variation and the Human Genome Project
← Back 📝 The Exam 🏠 All Units
QUIZ TOPIC ★★★
Genetics · Unit 6 of 9

Mutation, Variation and the Human Genome Project

TMU genetics primer, documents F–J, and the HGP review-question sheet The review sheet is the most exam-shaped document in the folder — its questions come with model answers
01

Mutation ★★

Mutation

A change that occurs in our DNA sequence, either due to mistakes when the DNA is copied or as the result of environmental factors such as UV light and cigarette smoke.

PointDetail
ConsequenceChanges in the proteins that are made — this can be a bad or a good thing
When it happensDuring DNA replication if errors are made and not corrected in time; or from exposure to smoking, sunlight and radiation
RepairCells can often recognise potentially mutation-causing damage and repair it before it becomes a fixed mutation
Why it mattersMutations contribute to genetic variation within species, and can be inherited, particularly if they have a positive effect
The two words that matter in that definition

Notice “or a good thing” and “before it becomes a fixed mutation”.

Damage is not a mutation. A base altered by UV light is damage; it becomes a mutation only if replication passes over it before repair does. So most DNA damage never becomes mutation at all — which is why the repair machinery matters more than the exposure.

And a mutation is not by definition harmful. Without mutation there would be no variation, no alleles, and no evolution. The same process that causes cancer is the process that made you different from your parents.

Test yourself
  • Define mutation. → A change in the DNA sequence, from copying errors or environmental factors such as UV light and cigarette smoke
  • Name three environmental causes. → Smoking, sunlight, radiation
  • When does damage become a fixed mutation? → When it is not repaired before replication passes over it
  • Are mutations always harmful? → No — they can be good, and they contribute to genetic variation
02

Genetic variation and SNPs ★★★

Genetic variation

The variation in the DNA sequence in each of our genomes. It is what makes us all unique — hair colour, skin colour, even the shape of our faces. Individuals of a species have similar characteristics but are rarely identical; the difference between them is called variation.

Single nucleotide polymorphism (SNP)

The most common type of genetic variation among people. Each SNP represents a difference in a single DNA base — A, C, G or T. On average they occur once in every 300 bases, and are often found in the DNA between genes. Genetic variation results in different forms, or alleles, of genes.

ComparisonShared DNA
Any two human beings99%
Parent and child99.5%
Humans and apes (related genes, in general)more than 95–98%
Human and mouse genesabout 70–90%, averaging 85%
Everything about you is in the 1%

Read that first row again. Every difference between any two people on earth — appearance, disease susceptibility, drug response — is carried in 1% of the sequence. The other 99% is identical in all of us.

And even that 1% is mostly single-letter changes occurring about once every 300 bases, frequently between genes rather than inside them. That is why SNPs became the focus of the work in §5: if you want to find what makes people differ, you look where the differences actually are.

The mouse comparison is the same lesson from further away. Gene for gene we are 85% similar to a mouse — which is precisely why a mouse is a usable model organism.

Test yourself
  • Define genetic variation. → The variation in DNA sequence between genomes — what makes each of us unique
  • What is a SNP? → A difference in a single DNA base, the commonest type of human genetic variation
  • How often do SNPs occur? → About once in every 300 bases, often between genes
  • How much DNA do any two humans share? → 99% — and 99.5% between parent and child
  • How similar are human and mouse genes? → About 70–90%, averaging 85%
03

The six goals of the Human Genome Project ★★★

The review sheet asks first for the two primary goals and then for all six. Both versions are worth having ready.

The two primary goals

To determine the sequence of chemical base pairs which make up DNA, and to identify and map the 20,000 genes of the human genome from both a physical and functional standpoint.

  • Identify all the genes in human DNA
  • Determine the sequences of the 3.3 billion chemical base pairs that make up human DNA
  • Store this information in databases
  • Improve tools for data analysis
  • Transfer related technologies to the private sector
  • Address the ethical, legal and social issues (ELSI) that may arise from the project
Only two of the six goals are about DNA

Look at what goals 3 to 6 actually are: databases, software, technology transfer, and ethics. Two-thirds of the project's stated aims were about infrastructure and consequences rather than sequence.

That was deliberate and it was prescient. A sequence nobody can search is useless, so the databases and analysis tools were not administration — they were the deliverable. And ELSI anticipated, before any of it existed, that knowing a person's genome would raise questions about insurance, employment and privacy.

For an exam question asking for the six goals, that framing is the answer's structure: two goals about the DNA, four about what happens to the information.

Test yourself
  • State the two primary goals. → Determine the sequence of base pairs, and identify and map the 20,000 genes physically and functionally
  • How many base pairs? → 3.3 billion
  • How many genes? → About 20,000
  • What does ELSI stand for? → Ethical, legal and social issues
  • Name the six goals. → Identify genes · sequence the base pairs · store in databases · improve analysis tools · transfer technology to the private sector · address ELSI
04

How the genome was actually sequenced ★★

The BAC and shotgun method
  1. The genome was broken into smaller pieces, approximately 150,000 base pairs in length.
  2. These pieces were ligated into bacterial artificial chromosomes (BACs) — vectors derived from bacterial chromosomes that have been genetically engineered.
  3. The vectors were inserted into bacteria, where they are copied by the bacterial DNA replication machinery.
  4. Each piece was sequenced separately as a small “shotgun” project and then assembled.
  5. The 150,000-base-pair pieces go together to create chromosomes, which are then mapped before being selected for sequencing.

The bacterium is doing the copying — the project used living replication machinery as its photocopier.

The review sheet also asks which non-human organisms the HGP sequenced as model organisms: E. coli, fruit fly, mice, zebrafish, yeast, nematodes, plants, and many microbial organisms and parasites.

And it asks what the published data actually represents — a question worth getting right, because the answer is not 'the human genome'. It is the combined “reference genome” of a small number of anonymous donors. All humans have unique gene sequences; the reference is a composite, not any individual.

Test yourself
  • How large were the fragments? → About 150,000 base pairs
  • What vector was used? → Bacterial artificial chromosomes (BACs)
  • What copies the inserted DNA? → The bacterium's own DNA replication machinery
  • Name five model organisms. → E. coli, fruit fly, mice, zebrafish, yeast, nematodes, plants
  • What does the published data represent? → The combined reference genome of a small number of anonymous donors
05

After the sequence — SNPs and the HapMap ★★

The review sheet asks what the next step was after the sequence was known. The answer: identifying the genetic variants that increase the risks for common diseases like cancer and diabetes.

The “shortcut” — and the assumption inside it

The method was explicitly a shortcut, and it is worth understanding because the reasoning is checkable.

Rather than sequence everyone, researchers looked only at sites where many people have a variant DNA unit. The idea was that since the major diseases are common, the genetic variants causing them would be common too.

That is an assumption, not a fact — a common disease could equally be caused by many different rare variants. The shortcut was a bet on which was true, and it shaped a decade of genetics. Being able to state the assumption, not just the method, is what a good answer looks like.

The International HapMap Project

A multi-country effort to identify and catalogue genetic similarities and differences in human beings — where these variants occur in our DNA, and how they are distributed among people within populations and among populations in different parts of the world. Using the HapMap, researchers can find genes that affect health, disease, and individual responses to medications and environmental factors.

QuestionAnswer
Which two structures identify differences among individuals?Single nucleotide polymorphisms (SNPs) and the HapMap
Whose genomes does HapMap catalogue?European, East Asian and African genomes
Which common diseases were targeted?Cancer · diabetes · heart disease · hypertension · Alzheimer's disease · dementia · asthma · and autoimmune diseases such as rheumatoid arthritis, type 1 diabetes and Graves' disease

One application named in the sheet: companies such as Myriad Genetics began offering easy-to-administer genetic tests showing predisposition to a variety of illnesses, including breast cancer. That is the practical face of the whole project — and precisely the territory ELSI was created to worry about.

Test yourself
  • What was the next step after sequencing? → Identifying genetic variants that increase risk for common diseases
  • What was the shortcut, and its assumption? → Look only at common variants — assuming common diseases have common causes
  • What is the HapMap? → A multi-country project cataloguing human genetic similarities and differences and how they are distributed
  • Which three population groups? → European, East Asian, African
  • Which two structures identify individual differences? → SNPs and the HapMap
06

Revision

QuestionAnswer
Define mutationA change in the DNA sequence from copying errors or environmental factors
Define SNPA difference in a single DNA base — the commonest human genetic variation
SNP frequency?About once in every 300 bases
Two humans share?99% of DNA (parent–child 99.5%)
Human vs mouse genes?70–90%, averaging 85%
Two primary HGP goals?Sequence the base pairs · identify and map the 20,000 genes physically and functionally
How many base pairs?3.3 billion
ELSI?Ethical, legal and social issues
Fragment size and vector?~150,000 bp, in bacterial artificial chromosomes
Model organisms?E. coli · fruit fly · mice · zebrafish · yeast · nematodes · plants
What is the published genome?A combined reference genome of anonymous donors
HapMap populations?European · East Asian · African
Test yourself — the whole unit
  • Name the six goals of the HGP. → Identify genes · sequence 3.3 billion base pairs · store in databases · improve analysis tools · transfer technology · address ELSI
  • What is a SNP and how common is it? → A single-base difference, about once every 300 bases
  • How was the genome physically sequenced? → Broken into ~150 kb pieces, cloned into BACs, copied in bacteria, shotgun-sequenced and assembled
  • What does the HapMap do? → Catalogues where common variants occur and how they are distributed within and between populations
  • Why did they only look at common variants? → The assumption that common diseases have common genetic causes