StudyRareStudyRare

Gene Structure and Organization

Log in to star

Last updated 2mo ago

Log in to add personal notes on this page.

A eukaryotic gene is more than its coding sequence: it is a transcription unit (promoter, exons, introns, UTRs, polyadenylation signal) embedded in a regulatory landscape of distal enhancers, silencers, and insulators. Knowing how these elements are arranged is essential for predicting the consequences of variants (including those outside coding exons) and for interpreting gene panels, copy-number changes, and structural variants.

  • Anatomy of a transcription unit: A typical protein-coding gene contains, in 5' → 3' order: a promoter, a 5' UTR, alternating exons and introns, a 3' UTR, and a polyadenylation signal. The transcription start site (TSS) defines position +1; nucleotides upstream are numbered with a minus sign.
  • Core promoter elements: Bind the general transcription machinery and Pol II. Common elements include the TATA box (~25-30 bp upstream, recognized by TBP), the initiator (Inr) at the TSS, the BRE (TFIIB recognition element), and the DPE (downstream promoter element). Many housekeeping genes use TATA-less promoters embedded in CpG islands (~1 kb GC-rich, CpG-dense regions covering most ubiquitously expressed gene promoters).
  • Proximal promoter elements: Within ~200 bp of the TSS. Bind sequence-specific transcription factors, e.g., CAAT box (CCAAT, bound by NF-Y) and GC box (GGGCGG, bound by Sp1). Mutations in proximal promoter elements cause disease (e.g., HBB promoter variants in beta-thalassemia).
  • Distal regulatory elements: Enhancers activate transcription from up to ~1 Mb away in either orientation, looping to promoters via cohesin/mediator complexes. Silencers repress transcription. Insulators (e.g., CTCF-bound boundary elements) block enhancer-promoter contacts and define topologically associating domains (TADs), structural compartments that constrain regulatory interactions. Disrupting a TAD boundary can cause limb malformations by allowing inappropriate enhancer-gene contacts (e.g., EPHA4 locus rearrangements causing brachydactyly/F-syndrome).
  • Exons and introns: Exons are retained in mature mRNA; introns are spliced out. The human genome contains roughly 180,000 exons across ~20,000 genes, with an average of 8-9 exons per gene and average exon length ~150 bp. Introns are typically much larger than exons and contain regulatory elements, intronic enhancers, and host most of the genome's repeat content.
  • Splice site consensus: Introns begin with GT (GU in RNA) at the 5' donor and end with AG at the 3' acceptor; an internal branch point adenosine (typically 18-40 nt upstream of the 3' splice site) is followed by a polypyrimidine tract. Together these define the canonical (U2-type) splice signals; a minor (U12-type) spliceosome processes a small subset of introns with AT-AC boundaries.
  • Untranslated regions (UTRs): The 5' UTR modulates translation efficiency through secondary structure, upstream open reading frames (uORFs), and internal ribosome entry sites (IRES). The 3' UTR carries the polyadenylation signal (AAUAAA), AU-rich elements (mRNA stability), and microRNA binding sites. UTR variants cause disease, e.g., 5' UTR uORF-creating variants reducing protein output, or 3' UTR variants disrupting miRNA-mediated regulation.
  • Polyadenylation: Cleavage and addition of a ~200-nt poly(A) tail occurs ~10-30 nt downstream of the AAUAAA hexamer. Alternative polyadenylation generates transcript isoforms with different 3' UTRs, often shifting the regulatory landscape (cancer cells frequently shorten 3' UTRs, escaping miRNA repression).
  • Alternative splicing: ~95% of multi-exon human genes undergo alternative splicing, producing tissue- and developmental-stage-specific isoforms (cassette exons, mutually exclusive exons, alternative 5'/3' splice sites, retained introns, alternative promoters). Tissue-specific isoforms (e.g., DMD muscle vs brain isoforms) explain why the same gene can present with different organ involvement.
  • Gene families and clustered genes: Many human genes occur in clusters generated by tandem duplication: HBA/HBB globin clusters, HOX clusters, olfactory receptor genes, immunoglobulin/TCR loci. Clustering enables coordinated regulation (e.g., locus control region of the beta-globin cluster) but also predisposes to non-allelic homologous recombination and copy-number variation.
  • Pseudogenes: Non-functional gene-like sequences. Processed pseudogenes are reverse-transcribed mRNAs reinserted into the genome (intronless, often with a poly-A tail). Unprocessed pseudogenes arise by gene duplication followed by inactivating mutations. Pseudogenes complicate sequencing (e.g., PMS2/PMS2CL, SMN1/SMN2, CYP21A2/CYP21A1P) and require allele-specific assays for accurate clinical testing.
  • Non-coding genes: Beyond protein-coding genes, the genome encodes rRNAs (rDNA on acrocentric short arms), tRNAs (~500 dispersed loci), microRNAs (~2,000), long non-coding RNAs (lncRNAs) including XIST and H19, and small nuclear/nucleolar RNAs. ENCODE estimates ~80% of the genome is transcribed in some context, although the functional fraction is debated.
  • Gene density and chromosomal context: Gene density varies markedly across the genome. R-bands (light G-bands) are gene-rich, GC-rich, and early-replicating; G-bands (dark) are gene-poor, AT-rich, and late-replicating. Chromosome 19 has the highest gene density; chromosomes 13, 18, and Y are gene-sparse, which is partly why trisomies of these chromosomes are viable while most other autosomal trisomies are lethal.
  • Gene nomenclature: Approved by HGNC (HUGO Gene Nomenclature Committee). Gene symbols are written in uppercase italics (BRCA1); the corresponding protein in non-italic uppercase (BRCA1). Mouse orthologs use sentence case (Brca1). HGVS nomenclature (e.g., NM_000059.4:c.5266dupC) ties variants to specific transcript references.
  • Non-coding pathogenic variants: Promoter, enhancer, intronic, and UTR variants account for a growing share of solved cases. Examples: deep-intronic ABCA4 variants creating pseudoexons in Stargardt disease; ZRS enhancer mutations causing preaxial polydactyly; uORF-creating 5' UTR variants in IRF6 (van der Woude syndrome).
  • Splicing variants: Roughly 15-30% of disease-causing mutations affect splicing: canonical GT/AG variants, branch point variants, exonic splice enhancer disruption (synonymous variants that look benign but break splicing), and deep-intronic variants creating cryptic splice sites. RNA studies or splicing predictors (SpliceAI) are increasingly required for variant classification.
  • Structural variants and gene clusters: Tandemly duplicated genes (PMP22 in CMT1A/HNPP, NF1 in NF1 microdeletions) and segmental duplications mediate recurrent rearrangements via non-allelic homologous recombination. Whole-gene deletions/duplications are frequently missed by sequencing alone. MLPA or read-depth CNV calling is required.
  • Pseudogene interference: Variant calling in PMS2, SMN1, CYP21A2, and the NF1 region requires pseudogene-aware assays. Failure to account for pseudogenes leads to false negatives (variant filtered as pseudogene noise) or false positives (pseudogene reads mapped to the active gene).
  • Imprinted and X-linked loci: Gene structure interacts with epigenetic regulation. Imprinted clusters (15q11-q13 PWS/AS, 11p15 BWS/SRS) and the X-inactivation center (XIC, harboring XIST) illustrate how non-coding RNAs and differentially methylated regions are themselves part of "gene structure" in a broader sense.
  • Tissue-specific isoforms: When a gene shows different phenotypes in different tissues (e.g., DMD: muscle vs brain isoforms; LMNA: Emery-Dreifuss vs progeria; RYR1: malignant hyperthermia vs central core disease), the explanation often lies in tissue-specific promoters or alternative splicing.

"GT-AG Rule: GU Goes, AG Arrives": Introns begin with GT (GU in RNA) at the donor and end with AG at the acceptor.

"AAUAAA = Add A's Aaaaa": The polyadenylation signal AAUAAA marks the spot to add the poly-A tail.

"R-bands are Rich": R-bands (reverse / light G-bands) are gene-Rich, GC-Rich, and Replicate early.

"PMS2, SMN1, CYP21: Pseudogene Trio": Three classic loci where a nearly identical pseudogene complicates clinical sequencing and demands allele-specific assays.