When Whole Exome Sequencing does not provide an answer
From developmental delay or intellectual disability, with or without autistic features, to neuromuscular disorders, suspected metabolic conditions, isolated or syndromic epilepsy and complex multisystem phenotypes, many patients remain without a molecular diagnosis after whole exome sequencing (WES) or trio WES.
Despite the high diagnostic potential of WES—and its ability to detect not only single-nucleotide variants and small insertions or deletions, but increasingly also copy-number variants (CNVs)—a substantial proportion of patients remain without a molecular diagnosis. Diagnostic yield varies considerably according to phenotype, patient selection, family structure, sequencing quality and analytical strategy, and can be substantially higher in selected cohorts.
When WES is non-diagnostic, one of the first questions is whether the causative variant could lie outside the genomic regions effectively investigated by the test. More specifically: could the patient carry a deep intronic variant that was not adequately captured or assessed by WES?
The possible explanation, however, extends beyond deep intronic variants alone. A recent second-line WGS study in patients with unexplained dystonia identified elusive variants involving exons poorly covered by WES, RNA genes, mitochondrial DNA, small copy-number variants, complex genomic rearrangements and short tandem repeats. Complementary transcriptomic and proteomic analyses also revealed pathogenic transcript disruption caused by deep intronic variants.
Here, we focus on this specific possibility: when should a deep intronic variant be considered, what can whole genome sequencing add, and why identifying the variant may be only the first step towards demonstrating its pathogenicity?
→ If this situation sounds familiar, the diagnostic pathway can start here.
What are deep intronic variants?
Most variants identified by Whole Genome Sequencing (WGS) lie outside protein-coding exons. Intronic variants occur within the introns, i.e. the intervening sequences of a gene: the regions located between one exon and the next.
Deep intronic variants are generally defined as intronic changes located more than 100 nucleotides away from the nearest exon-intron boundary. This threshold is conventional rather than absolute, but it distinguishes them from the regions immediately adjacent to exons that are more often included in the design and analysis of whole exome sequencing.
Unlike intergenic variants, deep intronic variants lie within a gene and may therefore interfere with its normal function. Most are expected to be benign. A small but clinically important subset, however, can alter RNA processing and cause monogenic disease. Their likelihood of being pathogenic depends on their rarity in population databases, their predicted effect on splicing or gene regulation, the clinical phenotype, segregation data when available, and supportive functional evidence.
How can deep intronic variants cause disease?
Most pathogenic deep intronic variants affect RNA splicing: the process through which introns are removed and exons are joined to produce the final messenger RNA. A deep intronic change can create or activate a cryptic splice site, or disrupt regulatory sequences that normally guide the splicing machinery. The resulting transcript may contain extra sequence, lose part of an exon or use an abnormal exon boundary.
The main consequences include:
- Pseudoexon inclusion
A segment of intronic DNA is mistakenly recognised as an exon and inserted into the mature RNA. This may disrupt the reading frame or introduce a premature stop codon. - Partial intron retention
Part of an intron is retained in the mature transcript, potentially altering the encoded protein or destabilising the RNA. - Exon skipping or altered exon boundaries
A variant may interfere with splicing signals or regulatory motifs, causing an exon to be omitted or shortened or lengthened through the use of an abnormal splice site. - Altered balance of alternative transcripts
In some genes, a deep intronic variant may modify the relative proportion of naturally occurring RNA isoforms, with a disease-causing effect in the relevant tissue.
These mechanisms explain why the clinical relevance of a deep intronic variant cannot be inferred from its genomic position alone — and why even a variant predicted to affect splicing is not necessarily disease-causing. Its predicted effect on RNA must be considered together with the phenotype, inheritance pattern and, whenever possible, experimental evidence from RNA studies.
Selected examples of pathogenic deep intronic variants
| Mechanism | Illustrative examples |
|---|---|
| Pseudoexon inclusion |
β-thalassaemia — HBB IVS2+705G>T Leber congenital amaurosis — CEP290 c.2991+1655A>G Duchenne muscular dystrophy — DMD c.3787-843C>A Hereditary breast cancer — BRCA2 c.6937+594T>G |
|
Competition with natural splice sites |
Pompe disease — GAA c.2190-345A>G Barth syndrome — TAZ IVS3+110G>A Menkes disease — ATP7A c.2406+1117A>G |
|
Rare regulatory or structural mechanisms |
Charcot-Marie-Tooth disease type 1B — MPZ c.126-1086T>A Intronic structural rearrangements have also been described in DMD and IQSEC2. |
Can Whole Exome Sequencing detect deep intronic variants?
Clinical whole exome sequencing (WES), including trio WES, remains a first-line genomic test for many suspected monogenic disorders in both children and adults. It provides broad analysis of protein-coding genes while maintaining a more accessible cost and a more focused analytical scope than whole genome sequencing.
By design, WES primarily targets protein-coding exons and a limited portion of the intronic sequence adjacent to exon-intron boundaries. The extent of this adjacent intronic coverage varies between capture kits, sequencing platforms and analytical pipelines. Some deeper intronic bases may occasionally be sequenced, but this coverage is neither complete nor uniform and such regions are not routinely assessed as part of a diagnostic WES analysis.
Deep intronic variants are generally defined as changes located more than 100 nucleotides from the nearest exon-intron boundary: regions that fall outside the intended target of whole exome sequencing. The practical answer is therefore clear: WES is not designed to reliably detect or exclude deep intronic variants. A non-diagnostic WES result does not rule out a pathogenic variant located within these regions.
Four possible consequences of altered splicing
Pseudoexon inclusion
An intronic sequence is incorrectly included as an additional exon.
Partial intron retention
Part of an intron remains within the mature messenger RNA.
Exon skipping
An exon is omitted or an abnormal exon boundary is used.
Altered alternative transcripts
The balance between naturally occurring RNA isoforms is shifted.
When should a deep intronic variant be suspected?
A deep intronic variant should not be assumed after every non-diagnostic WES result. It becomes a more plausible diagnostic hypothesis in selected situations:
- A non-diagnostic WES or trio WES result, particularly when the phenotype remains strongly suggestive of a monogenic disorder.
- Identification of a single pathogenic or likely pathogenic variant in a phenotype-matched recessive gene, when a second variant has not been found through conventional coding-region analysis.
- Disorders in which pathogenic deep intronic variants are already recognised as a relevant part of the mutational spectrum. Examples include Duchenne muscular dystrophy, inherited retinal disorders and selected metabolic conditions.
Depending on the phenotype, the previous diagnostic pathway, parental availability and budget, whole genome sequencing may be performed as a solo test or in trio. WGS expands the analysis beyond protein-coding exons to include intronic and intergenic regions, as well as mitochondrial DNA and other classes of variation that may be incompletely assessed by WES.
Most pathogenic variants causing monogenic disease are located within coding exons or canonical splice regions. However, a clinically meaningful minority lies elsewhere in the genome. When the clinical context makes this possibility credible, WGS can make deep intronic variants searchable and allow them to be prioritised alongside other non-coding, structural or difficult-to-detect findings.
A practical illustration is provided in our capsule from practice, “When even the exome is just the beginning”, where a negative exome did not represent the end of the diagnostic pathway, and WGS ultimately enabled a diagnosis.
Whole genome sequencing is not a universal solution, and the interpretation of deep intronic and other non-coding variants remains challenging. Nevertheless, WGS is currently the most comprehensive genomic test available in routine diagnostics. Where budget and clinical circumstances permit, WGS — and especially trio WGS — offers the broadest and least restrictive starting point for the investigation of a suspected monogenic disorder, including as a first-line approach.
For further considerations on choosing between WES trio and WGS solo, read our dedicated Diagnostic Approach.
For Patients
Access through Genetic Consultation
Request an appointment through our online form.
For Physicians
Discuss your Clinical Case
Contact us for pricing, TAT and sample submission.