Residency · Residency · Medical Genetics Genomics
Genomic Data Privacy, Consent, and Biobanking
Introduction
Genomic data is uniquely identifying, inherently familial, and immutable, creating distinct privacy challenges that existing regulatory frameworks were not designed to address. As genomic medicine expands and biobanks grow to population scale, the frameworks governing data privacy, informed consent, and sample governance must evolve to protect individuals and communities while enabling scientific advancement.
Unique Properties of Genomic Data
Genomic data possesses several properties that distinguish it from other forms of health information. It is uniquely identifying: even de-identified datasets can be re-identified using genealogical databases or reference panels. It is immutable, meaning that unlike passwords or credit card numbers, genomic data cannot be changed if breached. It has inherent familial implications, as an individual's genomic data reveals information about biological relatives who have not consented to disclosure. It has predictive power, revealing future disease risk, carrier status, ancestry, and traits. Finally, its significance evolves over time as scientific knowledge advances; today's variant of uncertain significance may become tomorrow's pathogenic variant, changing the implications of previously generated data.
Regulatory Framework
HIPAA (Health Insurance Portability and Accountability Act)
HIPAA classifies genetic information as protected health information (PHI). It applies to covered entities including healthcare providers, health plans, and clearinghouses, along with their business associates. HIPAA permits use and disclosure for treatment, payment, and healthcare operations without specific authorization. De-identified data, achieved through Safe Harbor or Expert Determination methods, is not subject to HIPAA protections. However, a significant limitation exists: genomic data is inherently difficult to truly de-identify given its uniqueness, making traditional de-identification approaches less effective than they are for other data types.
GINA (Genetic Information Nondiscrimination Act)
GINA prohibits the use of genetic information in health insurance and employment decisions, as detailed in the discussion of GINA and its gaps. While GINA does not address data privacy per se, it limits the discriminatory use of genetic information by specific entities.
Common Rule (45 CFR 46)
The Common Rule governs federally funded human subjects research. It requires Institutional Review Board (IRB) approval and informed consent for research involving identifiable specimens. The 2018 revision broadened the definition of human subjects to include identifiable biospecimens and introduced the concept of broad consent for future unspecified research use of stored specimens, acknowledging the practical challenges of obtaining study-specific consent for every future use of banked samples.
State Laws
State-level protections vary considerably. The California Consumer Privacy Act (CCPA/CPRA) provides California residents with rights over their personal data including genetic data, encompassing the right to know, delete, and opt out of data sharing. Several states have enacted genetic privacy statutes with varying degrees of protection. The Illinois Genetic Information Privacy Act requires informed consent before collection, disclosure, or use of genetic information, representing one of the more protective state frameworks.
International Frameworks
The European Union's General Data Protection Regulation (GDPR) classifies genetic data as a special category requiring explicit consent, with strong individual rights including the right to access, rectification, erasure, and data portability. The UNESCO International Declaration on Human Genetic Data (2003) provides a framework for the ethical collection, processing, storage, and use of genetic data at an international level.
Informed Consent for Genomic Testing and Research
Clinical Consent
Clinical consent for genomic testing must address the purpose of testing, the types of results that may be generated (diagnostic findings, secondary findings, VUS), implications for family members, data storage policies, and the potential for recontact. Opt-in or opt-out options for secondary findings analysis should be clearly presented. The discussion should cover privacy protections and their limitations honestly. Consent for data sharing with variant databases such as ClinVar to improve future interpretation should be addressed, as this sharing contributes to the common good of variant classification.
Research Consent Models
Several models exist for research consent in genomics. Study-specific consent is the traditional model where consent is obtained for a defined research protocol. Broad consent permits future, unspecified research use of specimens and data and is now permitted under the revised Common Rule. Tiered consent allows participants to specify categories of acceptable research use, such as consenting to cancer research but declining behavioral research. Dynamic consent is a technology-enabled model allowing participants to update their consent preferences over time via digital platforms. Community consent involves engagement with community representatives, and is particularly important for indigenous populations and minority communities whose collective interests may differ from individual participant preferences.
| Consent Model | Description | Advantages | Limitations |
|---|---|---|---|
| Study-specific | Consent for defined research protocol | Most informative for participant | Impractical for biobanks; limits future use |
| Broad consent | Permits future unspecified research use | Enables large-scale research; permitted by revised Common Rule | Participant cannot anticipate all future uses |
| Tiered consent | Participant specifies categories of acceptable research | Respects preferences while enabling flexibility | Complex to administer; categories may not fit future studies |
| Dynamic consent | Digital platform allowing ongoing preference updates | Empowers participants; responsive to evolving science | Requires technology infrastructure; engagement burden |
| Community consent | Engagement with community representatives | Protects collective interests; builds trust | May not represent all individuals; complex governance |
Consent Challenges
Genomic research consent forms are often written above the average reading level, limiting true informed participation. Participants cannot fully anticipate future uses of their data as science progresses, creating inherent limitations on how "informed" broad consent can be. The individual's consent does not cover the interests of relatives whose genetic information is indirectly revealed by the participant's data. Power dynamics in clinical settings may create implicit pressure for patients to consent to research participation.
Biobanking
Major Biobank Initiatives
Major biobanks now operate at population scale. UK Biobank has enrolled over 500,000 participants with genomic, health, and lifestyle data and operates an open-access model for approved researchers. The All of Us Research Program, funded by the NIH, aims to enroll over 1 million US participants reflecting national diversity and returns genetic results to participants. FinnGen has enrolled over 500,000 Finnish participants, leveraging Finland's unique population structure and comprehensive health registries. The Million Veteran Program (MVP) through the US Veterans Affairs system has enrolled over 900,000 veterans. Biobank Japan has enrolled over 260,000 participants focused on generating East Asian genomic data to address the diversity gap.
Governance Models
Biobank governance models vary in their approach to data access. Open access models make data available to any approved researcher, as exemplified by UK Biobank. Controlled access models require formal application and data use agreements, representing the most common approach for genomic data. Federated analysis keeps data at the source institution while allowing analyses to be run remotely without data transfer, reducing privacy risk. Data access committees review applications for data use, balancing the goals of open science with participant protections.
Return of Results from Biobanks
The practice of returning results from biobanks is evolving. Historically, most biobanks did not return individual results to participants. The All of Us program now returns medically actionable genetic results and pharmacogenomic information to participants. The Geisinger MyCode Community Health Initiative returns clinically confirmed pathogenic variants in ACMG SF genes to participants within their health system. These developments raise fundamental questions about the ethical obligation to return medically actionable findings when they are identified in research contexts.
Data Sharing and Re-Identification Risk
Re-Identification Concerns
Multiple pathways exist for re-identifying supposedly anonymous genomic data. Surname inference uses Y-chromosome STR profiles linked to surnames via genealogical databases. Cross-referencing combines genomic data with demographic or geographic data to enable identification. Forensic genealogy, in which law enforcement uses consumer genealogy databases to identify criminal suspects, has demonstrated that re-identification is technically feasible at scale. The Homer attack demonstrated that an individual's presence in a genomic dataset can be detected with as few as 75 independent SNPs, establishing a remarkably low threshold for identifiability.
Mitigation Strategies
Several strategies exist to mitigate re-identification risk. Controlled data access with data use agreements prohibiting re-identification attempts provides a legal deterrent. Cryptographic approaches including homomorphic encryption and secure multi-party computation allow computation on encrypted data. Differential privacy adds statistical noise to prevent individual identification while preserving aggregate patterns. Data enclaves allow researchers to analyze data in a secure environment without downloading it. Synthetic data generation creates artificial datasets that preserve statistical properties without containing real individual data.
Emerging Issues
Several emerging issues will shape the future of genomic data governance. DTC companies hold massive genetic databases with variable privacy policies, and several have experienced financial difficulties raising concerns about what happens to data assets during corporate sales or bankruptcy. Law enforcement access to genealogical databases for forensic identification raises questions about consent scope and privacy expectations that participants may not have anticipated. International data transfer of genomic data crossing jurisdictional boundaries must comply with multiple regulatory frameworks simultaneously. Pediatric biobanking involving consent by proxy raises questions about the child's future autonomy regarding their stored data. Indigenous data sovereignty is an increasingly asserted principle, with indigenous communities claiming governance rights over genetic data derived from their members.
Clinical Pearls
Genomic data is uniquely identifying and cannot be meaningfully de-identified; patients should be counseled that privacy protections reduce but do not eliminate re-identification risk. Broad consent for future research use of biospecimens is permitted under the revised Common Rule but requires clear communication about the open-ended nature of the consent. Return of medically actionable results from biobanks is becoming standard practice, blurring the line between research and clinical care and creating new obligations for both researchers and health systems. Clinicians should be aware that genomic data shared with variant databases such as ClinVar contributes to the common good of variant interpretation but requires appropriate consent from the patient whose data is being shared.
References
- Clayton EW, Evans BJ, Hazel JW, Rothstein MA. The law of genetic privacy: applications, implications, and limitations. Journal of Law and the Biosciences. 2019;6(1):1-36.
- Gymrek M, McGuire AL, Golan D, et al. Identifying personal genomes by surname inference. Science. 2013;339(6117):321-324.
- Garrison NA, Sathe NA, Antommaria AHM, et al. A systematic literature review of individuals' perspectives on broad consent and data sharing in the United States. Genetics in Medicine. 2016;18(7):663-671.
- All of Us Research Program Investigators. The "All of Us" Research Program. New England Journal of Medicine. 2019;381(7):668-676.