Overview
This lecture introduces clinical genetics through three worked scenarios (BRCA1/2 breast cancer, a child with a structural heart defect, and polygenic breast cancer risk) that together illustrate the tools of modern genomic medicine: the family history and pedigree, karyotyping and chromosomal microarray, single-gene and genome-scale sequencing, and polygenic risk scores. It then situates these tools within the broader architecture of human genetic variation (rare high-effect variants versus common low-effect variants) and closes by naming genetic fallacies to avoid and why genetics is core clinical knowledge.
The clinical geneticist and the family history
A clinical geneticist is a specialty-trained physician who consults on individuals and families with presumptive genetic conditions, sitting at the intersection of molecular pathology and clinical medicine, and needs literacy in both the clinical practice and the scientific basis of genetics.
The family history is described as the most important genetic tool for a physician. A worked example shows why: an initially reported pedigree (mother died of stomach cancer, daughter had bone cancer, son had prostate cancer, another daughter is the proband) looked unremarkable, but revising it against actual pathology reports revealed a different picture (mother had ovarian cancer diagnosed at 43, died at 49; daughter had breast cancer diagnosed at 45, died at 59; son had benign prostatic hypertrophy). The revised, pathology-confirmed pedigree fits a BRCA-related cancer pattern that the family’s recollected history had obscured.
Example 1: BRCA1/2 and breast cancer
Key facts and process:
- About 5% of people with breast cancer have a susceptibility due to mutations in one of two genes, BRCA1 (on chromosome 17) or BRCA2 (on chromosome 13), out of the 23 pairs of human chromosomes.
- Current clinical process: risk is calculated from the pedigree; a woman above a predetermined risk threshold (the slide gives 10% as an example) has BRCA1/2 sequenced; variants are identified by a genetic pathologist and results explained by a genetic counsellor; actions that may follow include changes to surveillance or surgery.
- BRCA1/2 shows reduced penetrance: not everyone who inherits a mutation develops breast and/or ovarian cancer.
Determining whether a detected sequence variant in BRCA1/2 is pathogenic requires answering four questions:
- Where in the gene is it, which requires understanding the gene’s architecture.
- What does the variant do to the sequence, i.e. is it synonymous or non-synonymous, does it truncate translation, does it cause a frameshift, and can computational tools help distinguish benign from pathogenic variants.
- Does it segregate with the disease trait within the family.
- Has it been observed before in this clinical context, using variant databases, checking that the database used is appropriate to the patient’s ancestry.
Risk management for confirmed BRCA1/2 mutation carriers:
- Breast: annual clinical breast examination; mammography starting 10 years before the age of the youngest affected family member, with a baseline at age 25; consideration of prophylactic mastectomy.
- Ovarian: prophylactic salpingo-oophorectomy (surgical removal of the oviducts and ovaries) at age 40.
Example 2: a child with a structural heart defect
The clinical approach starts with thinking clinically first: are there other co-morbidities suggesting a syndrome, developmental delay, or other malformations, and is there a family history, before taking and drawing a pedigree.
The worked family history: the proband is male, with a congenital heart defect, and a healthy sister; parents are healthy and unrelated. The father has two younger sisters: the youngest has a child with cystic fibrosis, and the older has a single healthy son from a whāngai adoption from her sister. The mother has two siblings: a younger sister with two healthy daughters, and a brother who died in the neonatal period with spina bifida. The paternal grandfather is deceased, the paternal grandmother is alive and well. The maternal grandparents were related to one another (degree of relatedness uncertain), and the maternal grandmother is deceased.
The resulting pedigree diagram marks the heart defect, spina bifida and cystic fibrosis cases and shows the maternal grandparents joined by a double line (consanguinity), the whāngai adoption, and the affected/deceased individuals across generations, but the exact connections and affected/unaffected status of each individual are difficult to resolve with full confidence from the diagram alone; the general structure matches the family history above.
Chromosomal and sequence-level investigation
Once family history and pedigree are taken, the question becomes whether the problem is chromosomal or at the sequence level.
- Traditional karyotyping images and pairs chromosomes 1-22 plus X and Y from a metaphase spread.
- The modern approach is a genome-wide chromosomal microarray, which screens for chromosomal imbalance across the genome. A worked example detects a chromosome 22q11 deletion: the microarray plot shows a cluster of data points dropping to a lower band (deletion) at 22q11, compared with balanced regions (two copies) and duplications (higher band). The 22q11 microdeletion has a prevalence of 1:4000 children.
- If not chromosomal, the search moves to the sequence level: chromosomes are paired strands of double helical DNA, and around one third of paired sites can vary between individuals. The two questions become which variable sites are relevant, and how to pick a disease-causing variant from a benign one.
The human genome and human genetic variation
The search space for sequence-level analysis: the human genome comprises 46 chromosomes (paternal and maternal components), about 2 metres of DNA per cell, around 21,000 genes, and 5.7 x 10^9 base pairs of DNA. Each human has approximately 4 million sequence variants.
Genomic DNA consists of exons (coding) interspersed with introns (removed to give the mRNA/coding sequence). The exome, the protein-coding portion of the genome, makes up only about 1% of the whole genome.
At least one third of the genome is variable at the sequence level, and this variation exists at two levels:
- Gene-level variation: the 1% of the genome that is best understood; coding regions vary, with both common and rare variation.
- Non-coding variation: the remaining 99% (the “dark matter” of the genome), which houses most of the variation but has more nuanced functional effects that are difficult to predict and measure.
Across 2,504 sequenced genomes, a plot of allele frequency against number of variants shows a steep spike of very high variant counts at very low frequency (rare variants) and a much lower, broader band across mid-to-high frequencies (common variants): most genetic variants are rare, and comparatively few are common.
The genomics toolbox
Sequencing approaches range from targeted to genome-wide:
- Iterative gene sequencing: sequencing single genes one by one (e.g. BRCA1).
- Highly parallel genomic sequencing, which includes: gene panels/families (10-500 pre-selected genes of interest); genotype signatures (panels of disease-associated individual variants across the genome); sequencing the entire coding genome, i.e. the exome (about 21,000 protein-coding genes, 1% of the genome, 50 megabases, about 3 gigabytes of data); and sequencing the whole genome (3,000 megabases, 300-400 gigabytes of data).
Genetic architecture of disease
Different diseases have differing genetic architectures, plotted as allele frequency (very rare to common) against effect size/odds ratio (low to high):
- Rare alleles causing Mendelian disease sit at very rare frequency with high effect size, are usually coding, and have major individual effects.
- Common variants predisposing to common disease sit at common frequency with low effect size, are usually non-coding, and have minor individual effects.
- Low-frequency variants with intermediate effects lie between these two extremes.
Worked example, obesity: many common variants each contribute only a tiny amount of liability. A plot of number of loci against cumulative variance explained rises steeply at first then levels off, reaching only about 2.7% cumulative variance explained at around 100 loci, illustrating that a useful genotype signature for a highly polygenic trait is hard to find from common variants alone.
Example 3: polygenic risk and the other 95% of breast cancer
Since only 5% of breast cancer is explained by BRCA1/2 mutations (a rare, high-effect Mendelian factor), and breast cancer is common with many families showing familial clustering without a BRCA1/2 mutation, the question is whether genetics can explain and help manage risk in the other 95%. Currently everybody over 45 is simply screened.
On the allele-frequency/effect-size plot, BRCA1/2 sits in the rare, high-effect Mendelian region, while “the other 95%” maps to the low-frequency/intermediate-effect and common/low-effect regions.
A polygenic risk score (PRS) built from 100 different genetic variants was used to test this:
- The distribution of risk allele count is shifted slightly higher in breast cancer cases than in controls, though the two distributions overlap substantially.
- Higher PRS lowers the age at which a woman reaches a 2.4% ten-year breast cancer risk (UK data): women with no family history and low PRS reach this risk around age 80, falling to about 37 at high PRS; women with a family history reach it earlier at every PRS level, from about 54 at low PRS down to about 32 at high PRS. Both family history and PRS category independently and jointly lower the age of reaching a given risk threshold.
Genetic fallacies to resist
- “On average, we are all the same, one size fits all” is countered by precision medicine approaches.
- Genetic determinism (genetics determines outcomes) is countered by gene-environment interaction, e.g. genes plus red meat intake elevating bowel cancer risk, or genes plus cigarette smoking causing early-onset emphysema.
- Genetic exceptionalism (genetics is special and specialist, separate from the rest of medicine) is countered by recognising that although genetics is familial (which does make it different), it is also the blueprint for human biology generally.
- “Genetics is hard to understand” is dismissed as old thinking for the easily intimidated; the lecture states it quite simply is not.
Conclusions
Genetics is altering medicine, and most diseases are genetic to a degree. The course aims to equip students with background knowledge and an appreciation of the clinical importance of genetics, while encouraging scepticism about deterministic or fatalistic framing of genetics. Since understanding susceptibility and resilience to disease is a core part of medicine, genetics is core medical knowledge.
Self-test
- Describe the role and position of a clinical geneticist within medicine.
- Explain, using the breast cancer family history example, why checking pathology reports revised the pedigree’s interpretation and why this matters clinically.
- What proportion of breast cancer is attributable to BRCA1/2 mutations, and on which chromosomes do BRCA1 and BRCA2 sit?
- Describe the current clinical process for identifying BRCA1/2 mutation carriers, from pedigree to result explanation.
- What is meant by reduced penetrance in the context of BRCA1/2, and why does it matter for counselling carriers?
- List the four questions used to determine whether a BRCA1/2 sequence variant is disease-causing rather than benign.
- Describe the risk management recommendations for a confirmed BRCA1/2 mutation carrier, for both breast and ovarian risk, including ages.
- When a child presents with a structural heart defect, what should the clinician think about before taking a pedigree?
- Distinguish a chromosomal microarray from traditional karyotyping in what it detects, and state the prevalence of the 22q11 microdeletion.
- Describe the size of the human genome (chromosomes, DNA length, gene number, base pairs) and roughly how many sequence variants an individual carries.
- Distinguish gene-level (coding) variation from non-coding variation in the human genome, including their relative proportions and how well each is understood.
- What does the rare-variant vs common-variant frequency plot from the 2,504 genomes study show about the distribution of human genetic variation?
- List the range of sequencing approaches in the genomic toolbox, from iterative single-gene sequencing to whole genome sequencing, with an example of what each targets.
- Distinguish the genetic architecture of a Mendelian disease from that of a common polygenic disease, in terms of allele frequency, effect size, and coding versus non-coding location.
- What did the obesity genetic architecture study find about the cumulative variance explained by common risk loci, and what does this imply about finding a useful genotype signature?
- Describe how a polygenic risk score built from 100 variants relates to breast cancer risk, and how family history and PRS category together affect the age at which a given risk threshold is reached.
- List the four genetic fallacies described in the lecture and, for each, the corrective concept given.
- Why does the lecture conclude that genetics is core knowledge for a career in medicine?
Answers
Reveal answers
- A clinical geneticist is a specialty-trained physician who consults on individuals and families with presumptive genetic conditions, located at the intersection of molecular pathology and clinical medicine, requiring literacy in both clinical practice and the scientific basis of genetics.
- The originally recalled family history (stomach cancer, bone cancer, prostate cancer) looked unremarkable, but pathology reports revealed the actual diagnoses were ovarian cancer, breast cancer, and benign prostatic hypertrophy, a pattern consistent with a BRCA mutation; this matters because relying on recalled history alone can miss a hereditary cancer syndrome that pathology-confirmed diagnoses reveal.
- About 5% of breast cancer is attributable to BRCA1/2 mutations; BRCA1 is on chromosome 17 and BRCA2 is on chromosome 13.
- Risk is first calculated from the pedigree; women above a risk threshold (e.g. 10%) have BRCA1/2 sequenced; a genetic pathologist identifies variants; a genetic counsellor explains results; this can lead to changes in surveillance or surgery.
- Reduced penetrance means not everyone who inherits a BRCA1/2 mutation goes on to develop breast and/or ovarian cancer, so carrier status indicates elevated risk rather than certainty of disease, which is important for how risk is communicated.
- (1) Where in the gene the variant lies, requiring knowledge of gene architecture; (2) what the variant does to the sequence, i.e. synonymous/non-synonymous change, truncation, frameshift, and whether computational tools can help classify it; (3) whether it segregates with the disease trait in the family; (4) whether it has been observed before in this clinical context, checked against variant databases appropriate to the patient’s ancestry.
- Breast: annual clinical breast examination, mammography starting 10 years before the age of the youngest affected family member with a baseline at 25, and consideration of prophylactic mastectomy. Ovarian: prophylactic salpingo-oophorectomy (removal of oviducts and ovaries) at age 40.
- Think about other co-morbidities that might indicate a syndrome, developmental delay, other malformations, and whether there is a family history, before taking and drawing a pedigree.
- Traditional karyotyping detects gross chromosomal changes from a metaphase spread; a chromosomal microarray is a generalised genome-wide survey that can detect smaller chromosomal imbalances (deletions/duplications) not visible by eye. The 22q11 microdeletion has a prevalence of 1:4000 children.
- The human genome has 46 chromosomes (paternal and maternal), about 2 metres of DNA per cell, around 21,000 genes, and 5.7 x 10^9 base pairs; each individual carries approximately 4 million sequence variants.
- Gene-level (coding) variation makes up about 1% of the genome and is the best understood, involving common and rare coding variants; non-coding variation makes up about 99% of the genome, houses most of the variation, but has more nuanced effects that are harder to predict and measure.
- It shows that variants are overwhelmingly rare (a steep spike of very high variant counts at very low allele frequency), with comparatively few variants being common (a broader, lower band across mid-to-high frequencies).
- Iterative gene sequencing sequences single genes one at a time (e.g. BRCA1). Highly parallel sequencing includes: gene panels of 10-500 pre-selected genes; genotype signatures (panels of known disease-associated variants); exome sequencing (the ~21,000 coding genes, ~1% of the genome, ~50 Mb, ~3 GB of data); and whole genome sequencing (~3,000 Mb, 300-400 GB of data).
- Mendelian disease is caused by rare alleles with high effect size (odds ratio), usually in coding regions. Common polygenic disease is driven by common variants with low effect size, usually in non-coding regions, with low-frequency/intermediate-effect variants occupying the middle ground.
- Cumulative variance explained rose steeply then plateaued at only about 2.7% at around 100 loci, implying that common variants each contribute only a tiny amount of liability, making a clinically useful genotype signature for obesity from common variants alone hard to achieve.
- A 100-variant PRS shows a modestly higher risk allele count distribution in breast cancer cases versus controls; higher PRS category lowers the age at which a woman reaches a 2.4% ten-year breast cancer risk (from about 80 down to about 37 with no family history, and from about 54 down to about 32 with a family history), so family history and PRS both independently and jointly bring forward the age of reaching a given risk threshold.
- “We are all the same, one size fits all” is countered by precision medicine; genetic determinism is countered by gene-environment interaction (e.g. genes plus red meat and bowel cancer, genes plus smoking and early emphysema); genetic exceptionalism is countered by recognising genetics as familial but also the general blueprint of human biology; “genetics is hard to understand” is dismissed as outdated thinking.
- Because most diseases are genetic to a degree, and understanding susceptibility and resilience to disease is a core part of medicine, so genetics is core knowledge for a medical career, approached with scepticism toward deterministic or fatalistic framing.