GWAS & Polygenic Risk Scores
Log in to starLast updated 3mo ago
A genome-wide association study (GWAS) scans the genome for common genetic variants (typically SNPs with minor allele frequency >1%) that are statistically associated with a trait or disease. The method has driven most of modern common-disease genetics over the past ~20 years and underlies polygenic risk scores (PRS), the clinical product layered on top of GWAS results.
A standard GWAS compares cases vs controls (or correlates SNP dosage with a quantitative trait) at hundreds of thousands to millions of SNPs simultaneously:
- Genotype thousands of individuals on a SNP array (~500K to 2M markers).
- Impute millions more SNPs using a reference panel (1000 Genomes, TOPMed, HRC).
- At each SNP, run a logistic regression of phenotype on genotype, with principal components as covariates to control for population stratification.
- Plot the −log₁₀(P) of every SNP against genomic position: the Manhattan plot. Peaks above the significance line are association hits.
- Check the QQ plot (observed vs expected −log₁₀(P)) for genomic inflation, which signals uncontrolled stratification or other systematic bias.
The conventional threshold is P < 5 × 10⁻⁸. This comes from a Bonferroni correction over ~1 million effectively independent common-variant tests across the genome. Hits below this threshold are "genome-wide significant."
Two facts to keep in mind:
- Effect sizes are small. Most genome-wide-significant common-variant hits have OR ~1.05 to 1.3. They identify biology, not clinical actionability individually. See odds ratios and relative risk for interpretation.
- Winner's curse: the discovery-cohort effect size is systematically inflated. Replication in an independent cohort gives a less biased estimate.
A GWAS hit is almost never the causal variant. It is typically a tag SNP in linkage disequilibrium with the underlying functional variant. The follow-up step is fine-mapping, which uses denser genotyping, ancestry diversity, and functional annotation to identify the true causal allele.
Canonical example: HLA-A3 and hereditary hemochromatosis. Pre-molecular-era serotyping showed a strong association between HLA-A3 and hemochromatosis. The causal variant is actually HFE C282Y on chromosome 6p22, ~5 Mb from HLA-A. The two sit on the same ancestral haplotype because of long-range LD across the MHC. A3 is the historical tag; C282Y is the driver. Modern hemochromatosis testing goes directly to HFE. The catalog of allele-disease associations lives on the HLA & MHC leaf.
The MHC region dominates GWAS for autoimmune diseases (T1D, celiac, MS, RA, lupus, ankylosing spondylitis, narcolepsy): it produces the largest peak in the Manhattan plot, often with effect sizes (OR 2 to 6, sometimes much larger) that dwarf everything else in the genome. Two structural reasons:
- Extreme LD. The MHC's recombination-suppressed extended haplotypes mean a single signal can span megabases, making fine-mapping uniquely hard.
- Real biology. HLA class II directly determines which self-peptides are presented to T cells. Functional impact on the immune repertoire is large and direct, unlike most non-coding GWAS hits in other regions.
That combination is why HLA dominates the catalog of autoimmune risk variants. See HLA & MHC for the specific allele-disease pairs.
A PRS aggregates the small effects of many GWAS-identified variants into a single per-individual score:
PRS = Σ (effect size of SNP i × allele dosage at SNP i)
For most common diseases, the top decile of PRS carries 2 to 4× the lifetime risk of the bottom decile. PRS is becoming clinically meaningful for:
- Coronary artery disease: risk stratification beyond traditional Framingham factors
- Breast cancer: adds discrimination on top of BRCA testing for general-population risk
- Type 2 diabetes: early identification of high-risk individuals
- Prostate cancer: informs screening intensity
Even the best PRSs explain less than half of trait heritability. PRS is risk stratification, not a diagnostic test.
Almost all early GWAS were done in European-ancestry cohorts. The resulting PRSs perform substantially worse in non-European populations because:
- LD patterns differ across ancestries, so the tag SNPs identified in Europeans no longer flag the causal variant in other groups.
- Allele frequencies and effect sizes can differ.
This is a major equity concern for clinical implementation. Diverse-ancestry GWAS efforts (PAGE, H3Africa, All of Us, TOPMed) are gradually closing the gap. When counseling a non-European-ancestry patient about a PRS-based risk estimate, the ancestry caveat belongs in the conversation.
GWAS detects common variants with small effect sizes. Rare variants with large effects (the classic Mendelian model) are invisible to GWAS; they require sequencing studies (whole exome, whole genome) or family-based linkage. Most common diseases sit on a continuum:
- Rare, large effect: Mendelian variants (e.g., BRCA1/2 in hereditary breast cancer)
- Low-frequency, moderate effect: intermediate-risk variants (e.g., CHEK2, ATM)
- Common, small effect: GWAS hits, aggregated into PRS
The "missing heritability" problem refers to the gap between twin-study estimates of heritability and what GWAS-identified variants explain. Some of that gap is in rare-variant contributions invisible to GWAS, some in gene-by-gene or gene-by-environment interactions, and some in poorly tagged regions of the genome.
A few non-HLA GWAS hits have either large effect sizes or major clinical influence:
| Variant | Trait | Notes |
|---|---|---|
| APOE ε4 | Alzheimer disease | ε4/ε4 carries OR ~12 for late-onset AD; pre-GWAS but confirmed in every modern AD GWAS |
| TCF7L2 | Type 2 diabetes | First major non-HLA T2D hit; OR ~1.4 per risk allele |
| PTPN22 R620W | Autoimmunity (T1D, RA, lupus, Graves) | Shared risk allele across multiple autoimmune diseases |
| 9p21 (CDKN2A/B) | Coronary artery disease, T2D | Highly replicated, mechanism still being worked out |
| FTO | Obesity, BMI | First non-syndromic obesity gene; modest per-allele effect, large at population scale |
| Application | Status |
|---|---|
| HLA pharmacogenomics (abacavir, carbamazepine, allopurinol) | Routine, required pretesting; see predictive testing |
| APOE genotyping for Alzheimer disease risk | Available; counseling implications are contested |
| CAD polygenic risk | Emerging clinical use; not standard of care |
| Breast cancer PRS on top of BRCA testing | Increasingly integrated in high-risk clinics |
| Disease-association research | Continuous discovery of new genes and pathways feeding bench-to-bedside translation |
- GWAS hits ≠ causal variants until fine-mapped. The hit is usually a tag SNP in LD with the real driver. HLA-A3 in hemochromatosis (tags HFE C282Y) is the textbook example.
- Genome-wide significance is P < 5 × 10⁻⁸. This is Bonferroni over ~1 million independent tests. Less-stringent thresholds for replication or candidate-region studies are acceptable.
- PRS is risk stratification, not diagnosis. A high PRS shifts pretest probability but does not confirm or rule out disease.
- Ancestry portability is the major clinical-translation barrier for PRS. European-derived PRSs underperform in African, East Asian, and Latino patients; flag this in counseling.
- The MHC dominates autoimmune GWAS because of extreme LD plus genuinely large immune-effect sizes; the MHC peak in a T1D or celiac Manhattan plot is a wall.
- Common variant vs rare variant: GWAS finds common-MAF, small-effect variants. Mendelian disease still requires sequencing.