# Rare Disease Diagnosis: The Undiagnosed Diseases Approach

## Introduction

Rare diseases collectively affect approximately 300-400 million people worldwide, with over 7,000 recognized rare diseases, the majority of which have a genetic basis. Despite advances in genomic technologies, a substantial proportion of individuals with suspected genetic conditions remain without a molecular diagnosis after standard clinical evaluation. Structured programs for undiagnosed diseases have emerged to address this diagnostic gap systematically.

## The Diagnostic Odyssey

### Scale of the Problem

The average time to diagnosis for a rare disease patient is 5-7 years, during which patients see an average of 7-8 physicians before receiving a correct diagnosis. Approximately 50% of patients with suspected Mendelian disorders remain undiagnosed after standard clinical genetic testing. An estimated 3,500-5,500 Mendelian conditions have been molecularly characterized, but thousands more gene-disease relationships likely remain to be discovered.

### Barriers to Diagnosis

Several factors conspire to make rare disease diagnosis so difficult. Phenotypic heterogeneity means the same genetic condition can present differently across individuals, while genetic heterogeneity means similar phenotypes can be caused by variants in different genes. Ultra-rare conditions may have insufficient published cases to establish gene-disease associations in the first place. Technical limitations present another challenge, as standard testing may miss structural variants, deep intronic variants, repeat expansions, somatic mosaicism, or epigenetic alterations. In some cases the causal gene may not yet be associated with human disease at all. Finally, non-Mendelian inheritance patterns such as digenic, oligogenic, or complex inheritance can evade detection by standard analytical approaches.

## Undiagnosed Diseases Programs

### NIH Undiagnosed Diseases Program (UDP) and Network (UDN)

The UDP was established at the NIH Clinical Center in 2008 as a flagship program for undiagnosed patients, and the UDN expanded in 2014 to a network of 12 clinical sites across the United States. These programs employ multidisciplinary evaluation teams including clinical geneticists, specialists, and researchers, alongside advanced testing modalities such as exome and genome sequencing, RNA-seq, metabolomics, and functional studies in model organisms. The diagnostic yield across the network has been approximately 25-35%, even among patients who had already been tested extensively, and the program has identified multiple novel gene-disease associations.

### International Programs

Similar efforts have been established worldwide. The Undiagnosed Diseases Network International (UDNI) coordinates programs across more than 20 countries. Solve-RD in Europe focuses on large-scale re-analysis and data sharing for unsolved rare disease cases. Japan's Initiative on Rare and Undiagnosed Diseases (IRUD) offers extensive genomic testing and functional validation at a national level, while Canada's Care4Rare program integrates functional genomics into research on unsolved rare diseases.

![World map showing locations of major undiagnosed diseases programs and networks with their key features](images/undiagnosed-diseases-programs.png)

## Systematic Diagnostic Approach

### Step 1: Deep Phenotyping

The diagnostic process begins with comprehensive clinical evaluation using Human Phenotype Ontology (HPO) terms for standardized phenotype description. This involves a systematic organ-by-organ review including neurological, ophthalmological, audiological, cardiac, skeletal, dermatological, and metabolic assessments, along with review of all prior imaging, laboratory, and genetic testing. A detailed three-generation pedigree is constructed with particular attention to consanguinity, multiple affected individuals, and early deaths. AI-based facial analysis tools such as Face2Gene and GestaltMatcher can also suggest syndromic diagnoses based on photographs.

| Diagnostic Modality | What It Detects | Added Yield (After Standard Testing) | When to Use |
|---|---|---|---|
| Trio WGS | SNVs, indels, SVs, CNVs across entire genome | Backbone of undiagnosed programs | First-line in undiagnosed programs |
| RNA-seq | Aberrant splicing, expression outliers, monoallelic expression | 7–17% beyond WES/WGS | Suspected splicing defect; VUS near splice sites; accessible tissue available |
| Optical genome mapping | Structural variants >500 bp; balanced rearrangements | Complements WGS for SVs | Suspected balanced rearrangement; complex karyotype |
| Long-read sequencing | Repeat expansions; complex SVs; phasing | Resolves specific blind spots | Suspected repeat expansion; pseudogene interference; phasing needed |
| Methylation array | Imprinting disorders; episignatures | Specific to epigenetic conditions | Suspected imprinting; unclassified ID/MCA with episignature |
| Metabolomics | Biochemical signatures of IEM | Supports novel gene validation | Suspected metabolic component; functional validation |
| Data reanalysis | New gene-disease associations; updated algorithms | 10–15% on re-analysis | Every 1–2 years for unsolved cases |

### Step 2: Genomic Analysis

Trio whole genome sequencing serves as the backbone of genomic analysis, complemented by several additional modalities. RNA sequencing from accessible tissues such as blood, fibroblasts, or muscle can reveal splicing abnormalities and expression outliers. Optical genome mapping detects structural variants, while long-read sequencing identifies repeat expansions and complex structural variants. Methylation arrays are used for imprinting disorders and epigenomic signatures. Importantly, iterative reanalysis of existing genomic data with updated gene-disease databases and new analytical tools is a key component of the process.

### Step 3: Variant Prioritization

Phenotype-driven analysis tools such as Exomiser, LIRICAL, and Phevor rank candidate genes based on HPO phenotype match. Family-based filtering applies de novo, compound heterozygous, homozygous, and X-linked inheritance models. Constraint analysis using LOEUF scores and missense Z-scores identifies genes intolerant to variation. Network and pathway analysis can highlight candidates in biological pathways related to the phenotype, and automated literature mining tools search for phenotypic overlap with published case reports.

### Step 4: Functional Validation

When candidate variants are identified, model organism studies in Drosophila, zebrafish, C. elegans, and mice are used to validate gene-disease associations. The UDN's Model Organisms Screening Center (MOSC) systematically tests candidate variants in Drosophila and zebrafish. Cell-based functional assays using patient fibroblasts, iPSC-derived cell types, or CRISPR-engineered cell lines provide additional evidence. Protein structural modeling through AlphaFold predictions can assess the impact of variants on protein structure and function. Untargeted metabolomic profiling can identify biochemical signatures that support a diagnosis.

![Flowchart showing the systematic approach to undiagnosed rare disease from deep phenotyping through multi-omic analysis, variant prioritization, and functional validation](images/undiagnosed-diagnostic-workflow.png)

## Data Sharing and Matchmaking

### GeneMatcher and Matchmaker Exchange

GeneMatcher is a web-based platform that allows clinicians and researchers to submit candidate genes and find other groups with patients carrying variants in the same gene. The Matchmaker Exchange is a federated network connecting GeneMatcher, DECIPHER, PhenomeCentral, MyGene2, and other databases. The identification of two or more unrelated patients with similar phenotypes and variants in the same novel gene is the standard for establishing new gene-disease associations, and these platforms have enabled the discovery of hundreds of new disease genes.

### Principles of Responsible Data Sharing

Responsible data sharing requires that patients consent to sharing their data for research purposes and that phenotypic and genomic data are shared with appropriate privacy protections and de-identification. The underlying rationale is community benefit, since data sharing accelerates diagnosis for all rare disease patients. The FAIR principles hold that data should be Findable, Accessible, Interoperable, and Reusable.

## Reanalysis of Existing Data

Systematic reanalysis of negative exome and genome data yields additional diagnoses in 10-15% of cases. Success on reanalysis occurs for several reasons: new gene-disease associations may have been published since the original analysis, bioinformatics algorithms for CNV calling, splice prediction, and structural variant detection may have improved, variant classification databases may have been updated, and population frequency databases may have expanded. Reanalysis is recommended every 1-2 years, and some programs now implement automated reanalysis pipelines that flag new candidate variants as databases are updated.

## Impact and Outcomes

Achieving a diagnosis ends the diagnostic odyssey and provides families with answers and closure. Approximately 30% of diagnoses lead to changes in clinical management. A molecular diagnosis enables accurate genetic counseling and recurrence risk assessment for the family, connects families to disease-specific support communities, and identifies potential therapeutic targets and eligibility for clinical trials. Novel gene discoveries advance scientific understanding and benefit future patients.

![Pie chart showing outcomes following diagnosis through undiagnosed diseases programs including changes in management, genetic counseling impact, and research contributions](images/undiagnosed-outcomes.png)

## Clinical Pearls

Approximately 25-35% of patients evaluated through undiagnosed diseases programs receive a molecular diagnosis, even after prior negative genetic testing, demonstrating the value of comprehensive multi-omic evaluation. Reanalysis of existing exome and genome data every 1-2 years is one of the most cost-effective strategies for solving previously undiagnosed cases. GeneMatcher and the Matchmaker Exchange have been transformative for novel gene discovery by connecting researchers and clinicians studying the same candidate genes worldwide. Deep phenotyping using standardized HPO terminology is as important as genomic testing itself, because phenotype-driven analysis tools rely on accurate and comprehensive clinical description to function effectively.

## References

1. Splinter K, Adams DR, Bacino CA, et al. Effect of genetic diagnosis on patients with previously undiagnosed disease. *New England Journal of Medicine*. 2018;379(22):2131-2139.
2. Sobreira N, Schiettecatte F, Valle D, Hamosh A. GeneMatcher: a matching tool for connecting investigators with an interest in the same gene. *Human Mutation*. 2015;36(10):928-930.
3. Wenger AM, Guturu H, Bernstein JA, Bejerano G. Systematic reanalysis of clinical exome data yields additional diagnoses: implications for providers. *Genetics in Medicine*. 2017;19(2):209-214.
4. Boycott KM, Rath A, Chong JX, et al. International cooperation to enable the diagnosis of all rare genetic diseases. *American Journal of Human Genetics*. 2017;100(5):695-705.
