SPARC and dbGaP: Expanding Human Subjects Research Data Sharing

SPARC × dbGaP: open & controlled components of human subjects studies are now formally linked across both repositories — improving discoverability and supporting NIH data sharing compliance.

Sparc news

Published Date

June 4, 2026

Share

SPARC and dbGaP: Expanding Human Subjects Research Data Sharing

Investigators working with human subjects can now share their data through a dual-repository workflow that connects SPARC with the NIH database of Genotypes and Phenotypes (dbGaP). dbGaP hosts controlled-access human genotype and phenotype data, typically raw sequencing reads and individual-level phenotype data, and makes them available to qualified researchers through a Data Access Request (DAR) process. SPARC complements this by hosting the multi-modal experimental and derived research datasets organized according to a standardized structure, file naming, and metadata scheme, none of which require controlled access for personally identifiable health information or sensitive data. Together, the two repositories give researchers a comprehensive and robust view of the underlying science.

Human subjects research often produces two kinds of data: openly shareable derived results (expression counts, images, processed measurements) and privacy-sensitive raw data requiring controlled access. The SPARC ↔ dbGaP workflow establishes formal links between these two repositories so a single study's open and controlled components can be discovered, attributed, and accessed as a coherent whole.

Studies generating human sequencing data follow a two-repository workflow:

  • Open-access components: de-identified derived data, processed outputs, images, phenotypic and clinical data, protocols, and analysis code are deposited in SPARC.
  • Controlled-access components: identifiable phenotypic raw sequence reads (fastq, bam, cram) and read-level genomic data are deposited in dbGaP.

Each repository points to the other so users can navigate between them. From SPARC, datasets carry machine-readable identifiers in their metadata that link to the corresponding dbGaP study, ensuring the connection is preserved for automated harvesters and aggregators. From dbGaP, each linked submission includes enhanced documentation, a cross-repository Study Document navigation guide, plus a notice that points researchers to the corresponding SPARC dataset.

To aid this workflow, SPARC has developed an automated conversion pipeline that translates SPARC metadata templates directly into the dbGaP submission format. Thus investigators no longer need to manually reformat their subject and sample metadata for the second repository. The pipeline reads files in your Pennsieve workspace and produces the corresponding dbGaP submission files in seconds, saves significant time and effort and reduces the risk of formatting errors.

Information for data contributors

If you are an investigator preparing to share human subjects data with SPARC:

  • Begin by submitting a dataset submiusion request form if you are not a member of an existing SPARC-supported consortia
  • Next, contact the SPARC Curation Team (curation@sparc.science) they will help you determine whether any of your files require dbGaP submission and walk you through the workflow.
  • Register your study with dbGaP and obtain a study accession (phs#).
  • Prepare your SPARC metadata as usual, with the dbGaP accession recorded in your dataset description.
  • Run the automated SPARC → dbGaP conversion pipeline on Pennsieve to generate dbGaP submission files directly from your SPARC metadata.

SPARC publication and dbGaP submission can happen in any order — they do not block each other.

→ Join SPARC Curation office hours to learn how to apply this workflow to your data.

Information for data users

Researchers interested in SPARC-linked dbGaP datasets can:

  • browse datasets on the SPARC Portal at sparc.science. Datasets with controlled-access companion data display the linked dbGaP study accession in the dataset description.
  • access the open-access component directly from SPARC under a CC-BY 4.0 license. There is no cost for data access from SPARC.
  • request access to the controlled component by following the standard dbGaP Authorized Access process. Each linked SPARC dataset includes a cross-repository navigation guide explaining how the data is split and how to request the dbGaP component.
View All News >