Overview

This lecture introduces clinical genetics through three worked scenarios (BRCA1/2 breast cancer, a child with a structural heart defect, and polygenic breast cancer risk) that together illustrate the tools of modern genomic medicine: the family history and pedigree, karyotyping and chromosomal microarray, single-gene and genome-scale sequencing, and polygenic risk scores. It then situates these tools within the broader architecture of human genetic variation (rare high-effect variants versus common low-effect variants) and closes by naming genetic fallacies to avoid and why genetics is core clinical knowledge.

The clinical geneticist and the family history

A clinical geneticist is a specialty-trained physician who consults on individuals and families with presumptive genetic conditions, sitting at the intersection of molecular pathology and clinical medicine, and needs literacy in both the clinical practice and the scientific basis of genetics.

The family history is described as the most important genetic tool for a physician. A worked example shows why: an initially reported pedigree (mother died of stomach cancer, daughter had bone cancer, son had prostate cancer, another daughter is the proband) looked unremarkable, but revising it against actual pathology reports revealed a different picture (mother had ovarian cancer diagnosed at 43, died at 49; daughter had breast cancer diagnosed at 45, died at 59; son had benign prostatic hypertrophy). The revised, pathology-confirmed pedigree fits a BRCA-related cancer pattern that the family’s recollected history had obscured.

Example 1: BRCA1/2 and breast cancer

Key facts and process:

  • About 5% of people with breast cancer have a susceptibility due to mutations in one of two genes, BRCA1 (on chromosome 17) or BRCA2 (on chromosome 13), out of the 23 pairs of human chromosomes.
  • Current clinical process: risk is calculated from the pedigree; a woman above a predetermined risk threshold (the slide gives 10% as an example) has BRCA1/2 sequenced; variants are identified by a genetic pathologist and results explained by a genetic counsellor; actions that may follow include changes to surveillance or surgery.
  • BRCA1/2 shows reduced penetrance: not everyone who inherits a mutation develops breast and/or ovarian cancer.

Determining whether a detected sequence variant in BRCA1/2 is pathogenic requires answering four questions:

  1. Where in the gene is it, which requires understanding the gene’s architecture.
  2. What does the variant do to the sequence, i.e. is it synonymous or non-synonymous, does it truncate translation, does it cause a frameshift, and can computational tools help distinguish benign from pathogenic variants.
  3. Does it segregate with the disease trait within the family.
  4. Has it been observed before in this clinical context, using variant databases, checking that the database used is appropriate to the patient’s ancestry.

Risk management for confirmed BRCA1/2 mutation carriers:

  • Breast: annual clinical breast examination; mammography starting 10 years before the age of the youngest affected family member, with a baseline at age 25; consideration of prophylactic mastectomy.
  • Ovarian: prophylactic salpingo-oophorectomy (surgical removal of the oviducts and ovaries) at age 40.

Example 2: a child with a structural heart defect

The clinical approach starts with thinking clinically first: are there other co-morbidities suggesting a syndrome, developmental delay, or other malformations, and is there a family history, before taking and drawing a pedigree.

The worked family history: the proband is male, with a congenital heart defect, and a healthy sister; parents are healthy and unrelated. The father has two younger sisters: the youngest has a child with cystic fibrosis, and the older has a single healthy son from a whāngai adoption from her sister. The mother has two siblings: a younger sister with two healthy daughters, and a brother who died in the neonatal period with spina bifida. The paternal grandfather is deceased, the paternal grandmother is alive and well. The maternal grandparents were related to one another (degree of relatedness uncertain), and the maternal grandmother is deceased.

The resulting pedigree diagram marks the heart defect, spina bifida and cystic fibrosis cases and shows the maternal grandparents joined by a double line (consanguinity), the whāngai adoption, and the affected/deceased individuals across generations, but the exact connections and affected/unaffected status of each individual are difficult to resolve with full confidence from the diagram alone; the general structure matches the family history above.

Chromosomal and sequence-level investigation

Once family history and pedigree are taken, the question becomes whether the problem is chromosomal or at the sequence level.

  • Traditional karyotyping images and pairs chromosomes 1-22 plus X and Y from a metaphase spread.
  • The modern approach is a genome-wide chromosomal microarray, which screens for chromosomal imbalance across the genome. A worked example detects a chromosome 22q11 deletion: the microarray plot shows a cluster of data points dropping to a lower band (deletion) at 22q11, compared with balanced regions (two copies) and duplications (higher band). The 22q11 microdeletion has a prevalence of 1:4000 children.
  • If not chromosomal, the search moves to the sequence level: chromosomes are paired strands of double helical DNA, and around one third of paired sites can vary between individuals. The two questions become which variable sites are relevant, and how to pick a disease-causing variant from a benign one.

The human genome and human genetic variation

The search space for sequence-level analysis: the human genome comprises 46 chromosomes (paternal and maternal components), about 2 metres of DNA per cell, around 21,000 genes, and 5.7 x 10^9 base pairs of DNA. Each human has approximately 4 million sequence variants.

Genomic DNA consists of exons (coding) interspersed with introns (removed to give the mRNA/coding sequence). The exome, the protein-coding portion of the genome, makes up only about 1% of the whole genome.

At least one third of the genome is variable at the sequence level, and this variation exists at two levels:

  • Gene-level variation: the 1% of the genome that is best understood; coding regions vary, with both common and rare variation.
  • Non-coding variation: the remaining 99% (the “dark matter” of the genome), which houses most of the variation but has more nuanced functional effects that are difficult to predict and measure.

Across 2,504 sequenced genomes, a plot of allele frequency against number of variants shows a steep spike of very high variant counts at very low frequency (rare variants) and a much lower, broader band across mid-to-high frequencies (common variants): most genetic variants are rare, and comparatively few are common.

The genomics toolbox

Sequencing approaches range from targeted to genome-wide:

  • Iterative gene sequencing: sequencing single genes one by one (e.g. BRCA1).
  • Highly parallel genomic sequencing, which includes: gene panels/families (10-500 pre-selected genes of interest); genotype signatures (panels of disease-associated individual variants across the genome); sequencing the entire coding genome, i.e. the exome (about 21,000 protein-coding genes, 1% of the genome, 50 megabases, about 3 gigabytes of data); and sequencing the whole genome (3,000 megabases, 300-400 gigabytes of data).

Genetic architecture of disease

Different diseases have differing genetic architectures, plotted as allele frequency (very rare to common) against effect size/odds ratio (low to high):

  • Rare alleles causing Mendelian disease sit at very rare frequency with high effect size, are usually coding, and have major individual effects.
  • Common variants predisposing to common disease sit at common frequency with low effect size, are usually non-coding, and have minor individual effects.
  • Low-frequency variants with intermediate effects lie between these two extremes.

Worked example, obesity: many common variants each contribute only a tiny amount of liability. A plot of number of loci against cumulative variance explained rises steeply at first then levels off, reaching only about 2.7% cumulative variance explained at around 100 loci, illustrating that a useful genotype signature for a highly polygenic trait is hard to find from common variants alone.

Example 3: polygenic risk and the other 95% of breast cancer

Since only 5% of breast cancer is explained by BRCA1/2 mutations (a rare, high-effect Mendelian factor), and breast cancer is common with many families showing familial clustering without a BRCA1/2 mutation, the question is whether genetics can explain and help manage risk in the other 95%. Currently everybody over 45 is simply screened.

On the allele-frequency/effect-size plot, BRCA1/2 sits in the rare, high-effect Mendelian region, while “the other 95%” maps to the low-frequency/intermediate-effect and common/low-effect regions.

A polygenic risk score (PRS) built from 100 different genetic variants was used to test this:

  • The distribution of risk allele count is shifted slightly higher in breast cancer cases than in controls, though the two distributions overlap substantially.
  • Higher PRS lowers the age at which a woman reaches a 2.4% ten-year breast cancer risk (UK data): women with no family history and low PRS reach this risk around age 80, falling to about 37 at high PRS; women with a family history reach it earlier at every PRS level, from about 54 at low PRS down to about 32 at high PRS. Both family history and PRS category independently and jointly lower the age of reaching a given risk threshold.

Genetic fallacies to resist

  • “On average, we are all the same, one size fits all” is countered by precision medicine approaches.
  • Genetic determinism (genetics determines outcomes) is countered by gene-environment interaction, e.g. genes plus red meat intake elevating bowel cancer risk, or genes plus cigarette smoking causing early-onset emphysema.
  • Genetic exceptionalism (genetics is special and specialist, separate from the rest of medicine) is countered by recognising that although genetics is familial (which does make it different), it is also the blueprint for human biology generally.
  • “Genetics is hard to understand” is dismissed as old thinking for the easily intimidated; the lecture states it quite simply is not.

Conclusions

Genetics is altering medicine, and most diseases are genetic to a degree. The course aims to equip students with background knowledge and an appreciation of the clinical importance of genetics, while encouraging scepticism about deterministic or fatalistic framing of genetics. Since understanding susceptibility and resilience to disease is a core part of medicine, genetics is core medical knowledge.

Self-test

  1. Describe the role and position of a clinical geneticist within medicine.
  2. Explain, using the breast cancer family history example, why checking pathology reports revised the pedigree’s interpretation and why this matters clinically.
  3. What proportion of breast cancer is attributable to BRCA1/2 mutations, and on which chromosomes do BRCA1 and BRCA2 sit?
  4. Describe the current clinical process for identifying BRCA1/2 mutation carriers, from pedigree to result explanation.
  5. What is meant by reduced penetrance in the context of BRCA1/2, and why does it matter for counselling carriers?
  6. List the four questions used to determine whether a BRCA1/2 sequence variant is disease-causing rather than benign.
  7. Describe the risk management recommendations for a confirmed BRCA1/2 mutation carrier, for both breast and ovarian risk, including ages.
  8. When a child presents with a structural heart defect, what should the clinician think about before taking a pedigree?
  9. Distinguish a chromosomal microarray from traditional karyotyping in what it detects, and state the prevalence of the 22q11 microdeletion.
  10. Describe the size of the human genome (chromosomes, DNA length, gene number, base pairs) and roughly how many sequence variants an individual carries.
  11. Distinguish gene-level (coding) variation from non-coding variation in the human genome, including their relative proportions and how well each is understood.
  12. What does the rare-variant vs common-variant frequency plot from the 2,504 genomes study show about the distribution of human genetic variation?
  13. List the range of sequencing approaches in the genomic toolbox, from iterative single-gene sequencing to whole genome sequencing, with an example of what each targets.
  14. Distinguish the genetic architecture of a Mendelian disease from that of a common polygenic disease, in terms of allele frequency, effect size, and coding versus non-coding location.
  15. What did the obesity genetic architecture study find about the cumulative variance explained by common risk loci, and what does this imply about finding a useful genotype signature?
  16. Describe how a polygenic risk score built from 100 variants relates to breast cancer risk, and how family history and PRS category together affect the age at which a given risk threshold is reached.
  17. List the four genetic fallacies described in the lecture and, for each, the corrective concept given.
  18. Why does the lecture conclude that genetics is core knowledge for a career in medicine?

Answers