Residency · Residency · Medical Genetics Genomics
Bioinformatics for the Clinical Geneticist
Introduction
Clinical geneticists increasingly interact with bioinformatics outputs in their daily practice. While deep computational expertise is not expected, a working knowledge of key concepts, databases, tools, and analytical frameworks is essential for critically evaluating genomic reports, understanding their limitations, and communicating effectively with laboratory teams.
Essential Databases
Variant and Disease Databases
ClinVar is an NCBI-hosted archive of variant-disease interpretations submitted by clinical labs and expert panels; it aggregates classifications but does not adjudicate conflicts. ClinGen provides expert-curated gene-disease validity assessments, variant classifications, and dosage sensitivity maps. OMIM (Online Mendelian Inheritance in Man) is a comprehensive catalog of human genes and genetic disorders with detailed phenotype-genotype descriptions. HGMD (Human Gene Mutation Database) is a curated collection of published disease-causing variants requiring subscription for full access. LOVD (Leiden Open Variation Database) hosts gene-specific variant databases maintained by expert curators.
Population Frequency Databases
gnomAD (Genome Aggregation Database) contains over 800,000 alleles from exomes and genomes across diverse populations and is essential for filtering common variants. ALFA (Allele Frequency Aggregator) is NCBI-hosted population frequency data. Population-specific databases for underrepresented groups are increasingly important to reduce misclassification bias.
Functional Annotation Resources
UniProt provides protein sequence, structure, and functional annotation. The Protein Data Bank (PDB) and AlphaFold Protein Structure Database offer 3D protein structures for assessing variant impact on protein function. GTEx (Genotype-Tissue Expression) catalogs gene expression across 54 human tissues, informing tissue-specific disease mechanisms. ENCODE provides a comprehensive catalog of functional elements in the human genome.
In-Silico Prediction Tools
Missense Variant Predictors
REVEL is an ensemble method combining 13 individual predictors into a single score from 0 to 1 and is widely used in ACMG classification. CADD (Combined Annotation Dependent Depletion) integrates diverse annotations into a single Phred-scaled deleteriousness score. AlphaMissense is DeepMind's deep learning tool trained on protein structure and evolutionary data, predicting pathogenicity of all possible missense variants. PolyPhen-2, SIFT, and MutationTaster are older tools still referenced but increasingly superseded by ensemble and deep learning methods.
| Tool/Resource | Type | Key Feature | Clinical Use |
|---|---|---|---|
| REVEL | Missense predictor (ensemble) | Combines 13 tools; score 0–1 | ACMG PP3/BP4 evidence |
| CADD | General variant predictor | Phred-scaled; integrates diverse annotations | Prioritization of coding + non-coding variants |
| AlphaMissense | Missense predictor (deep learning) | Protein structure-based; predicts all possible missense | Emerging use for novel missense VUS |
| SpliceAI | Splicing predictor (deep learning) | Delta scores for donor/acceptor gain/loss | Evaluating splice-region variants |
| gnomAD | Population database | >800K alleles; constraint metrics (LOEUF, pLI) | Variant filtering (BA1, PM2); gene constraint |
| ClinVar | Variant-disease archive | Aggregated lab classifications | Checking prior classifications; conflict identification |
| ClinGen | Expert curation | Gene-disease validity; variant expert panels | Authoritative gene and variant assessments |
Splicing Predictors
SpliceAI is a deep neural network predicting splice-altering effects of variants, providing delta scores for acceptor/donor gain/loss. MaxEntScan is an information theory-based scoring system of splice site strength. Human Splicing Finder is a web-based tool for splice site and regulatory motif analysis.
Conservation Metrics
PhyloP and PhastCons measure evolutionary conservation at individual nucleotides across vertebrate alignments. GERP++ identifies constrained elements where fewer substitutions have occurred than expected under a neutral model.
Gene-Level Constraint Metrics
pLI (probability of loss-of-function intolerance) from gnomAD identifies genes intolerant to heterozygous loss of function when the score exceeds 0.9. LOEUF (loss-of-function observed/expected upper bound fraction) is the preferred metric replacing pLI, with lower values indicating greater constraint. The missense constraint Z-score identifies genes under selective constraint against missense variation.
Variant Interpretation Frameworks
ACMG/AMP Guidelines
The framework uses 28 criteria organized into evidence categories: pathogenic (PVS1, PS1-4, PM1-6, PP1-5) and benign (BA1, BS1-4, BP1-7). Evidence is combined using a Bayesian-compatible point system (Tavtigian et al., 2018). The five-tier classification comprises Pathogenic, Likely Pathogenic, Variant of Uncertain Significance (VUS), Likely Benign, and Benign. The ClinGen Sequence Variant Interpretation (SVI) working group provides gene-specific and rule-specific refinements.
Key ACMG Criteria for the Clinician
PVS1 applies to null variants in genes where loss of function is a known mechanism, requiring careful application with ClinGen's PVS1 decision tree. PM2 indicates absence or extreme rarity in population databases (gnomAD). PP1/PS2 relate to cosegregation with disease in families and de novo occurrence (with confirmed maternity and paternity). PS3/BS3 represent functional studies supporting or refuting pathogenicity. PP3/BP4 provide computational evidence from in-silico predictions.
Practical Bioinformatics Skills
Reading a VCF File
The VCF format contains CHROM, POS, REF, and ALT fields (chromosome, position, reference allele, alternate allele), QUAL (variant call quality score), FILTER (whether the variant passed quality filters), INFO (annotations including allele frequency, consequence, and gene), and FORMAT with sample columns (genotype GT, allele depth AD, genotype quality GQ, and read depth DP).
Understanding Genomic Coordinates
Two reference genome builds are in clinical use: GRCh37 (hg19) and GRCh38 (hg38), with different coordinate systems. The LiftOver tool converts coordinates between builds. It is essential to always confirm which build is used when comparing variants across reports or databases.
HGVS Nomenclature
Genomic nomenclature uses the prefix g. (for example, NC_000017.11:g.43045684G>A). Coding DNA uses c. (for example, NM_007294.4:c.5266dupC). Protein uses p. (for example, NP_009225.1:p.Gln1756Profs*74). RNA uses r. (for example, r.5266dupc). The Mutalyzer tool validates HGVS nomenclature and performs coordinate conversions.
Genome Browsers
The UCSC Genome Browser offers extensive track options including conservation, regulation, variation, and clinical annotations. Ensembl is a European-based genome browser with strong support for comparative genomics and regulatory annotation. IGV (Integrative Genomics Viewer) is a desktop application for visualizing aligned sequencing reads (BAM files), useful for confirming variant calls.
Emerging Computational Tools
AI-assisted variant interpretation uses machine learning models trained on classified variants to prioritize novel VUS. Phenotype-driven gene prioritization tools like Exomiser, LIRICAL, and Phevor rank candidate genes based on HPO (Human Phenotype Ontology) terms. Automated ACMG classification tools such as InterVar, Franklin, and Varsome apply ACMG rules computationally but require human review.
Clinical Pearls
gnomAD population frequency is among the most powerful filters in variant interpretation; a variant present at greater than 1% in any population is unlikely to cause a rare Mendelian disorder (BA1 criterion). No single in-silico predictor is sufficient; ensemble tools like REVEL and CADD should be used, and computational evidence alone is supporting, not standalone. Gene-level constraint metrics (LOEUF) help assess whether haploinsufficiency is a plausible mechanism for a given gene. Understanding HGVS nomenclature and reference genome builds is essential for accurate communication about variants across clinical teams and databases.
References
- Richards S, Aziz N, Bale S, et al. Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the ACMG and AMP. Genetics in Medicine. 2015;17(5):405-424.
- Tavtigian SV, Greenblatt MS, Harrison SM, et al. Modeling the ACMG/AMP variant classification guidelines as a Bayesian classification framework. Genetics in Medicine. 2018;20(9):1054-1060.
- Karczewski KJ, Francioli LC, Tiao G, et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature. 2020;581(7809):434-443.
- Smedley D, Jacobsen JOB, Jager M, et al. Next-generation diagnostics and disease-gene discovery with the Exomiser. Nature Protocols. 2015;10(12):2004-2015.