
The Simons Foundation Autism Research Initiative (SFARI) is pleased to announce the release of long-read whole genome sequencing (WGS) data for 110 individuals across 37 families participating in the SPARK autism research study. This is SFARI’s first release of long-read WGS data from SPARK. Long-read WGS data have demonstrated the ability to uncover new genetic contributors to autism, but its full potential remains underexplored.
This effort was part of a pilot to assess the feasibility of performing long-read sequencing on DNA extracted from saliva. In the pilot, samples from SPARK participants were sequenced using PacBio SPRQ chemistry and Revio sequencing machines. Sequencing was performed in two phases: first using banked, previously extracted DNA, then using DNA extracted from saliva using PacBio’s high-molecular-weight NanoBind extraction protocol.
Long-read WGS was performed on families containing autism probands for whom no high-impact autism-related variants were previously found using internal analysis, despite being phenotypically enriched to have a genetic finding. These data provide an opportunity to identify novel autism-related variants that could be discovered only with long-read WGS. This and related research will also help explore the overall utility of long-read WGS (versus short-read WGS) in understanding the genetics of autism.
Previous sequencing efforts in SPARK include short-read whole exome sequencing (WES) or WGS on over 178,000 individuals. DNA is extracted from donated saliva and banked for future research.
Interested researchers can request access to this dataset (DS0000134) through SFARI Base. The dataset includes detailed release notes, individual and sample metadata tables, aligned sequence reads (.bam) and raw HiFi reads (.fastq.gz).


