NG00116 - Resolving Mutations in Challenging Genomic Regions to Test Association with Disease Phenotypes

To access this data, please log into DSS and submit an application.
Within the application, add this dataset (accession NG00116) in the “Choose a dataset” section.
Once approved, you will be able to log in and access the data within the DARM portal.

Description

Many regions of the human genome present challenges that prohibit scientists from discovering potential disease-causing mutations. We developed methods to characterize mutations in these regions to rescue mutations that are otherwise overlooked. PMID: 31104630

Provided here are variant calls in VCF format for 14,526 samples derived from the ADSP whole-exome and whole-genome sequencing dataset (available via DSS: NG00067). For phenotypic information for the participants in this dataset, please also apply for access to the ADSP. A crosswalk file that maps the IDs between this dataset and the ADSP IDs is provided within this dataset.

Sample Summary per Data Type

Sample Set	Accession	Data Type	Number of Samples
Camouflaged Variants	snd10074	Camouflaged Variants	14,526

Available Filesets

Name	Accession	Latest Release	Description
Camouflaged Variants in VCF Format	fsa000090	NG00116.v1	Camouflaged Variants in VCF Format, README

View the File Manifest for a full list of files released in this dataset.

Data Dictionary Files

This sample set includes a .vcf file containing camouflaged variants from the 14,526 ADSP samples that were described in Ebbert et al. 2019. PMID: 31104630 This dataset was originally derived from the primary ADSP data available in NG00067.

Sample Set	Accession Number	Number of Subjects	Number of Samples
Camouflaged Variants	snd10074	14,314	14,526

Consent Level	Number of Subjects
DS-ADRD-IRB-PUB	391
DS-ADRD-IRB-PUB-NPU	607
DS-ADRDAGE-IRB-PUB	929
DS-ADRDMEM-IRB-PUB-NPU	104
DS-AGEADLT-IRB-PUB	644
DS-ND-IRB-PUB	312
DS-ND-IRB-PUB-MDS	18
DS-ND-IRB-PUB-NPU	929
DS-NEURO-IRB-PUB	91
DS-NEURO-IRB-PUB-NPU	145
GRU-IRB-PUB	5428
GRU-IRB-PUB-NPU	101
HMB-IRB-PUB	1086
HMB-IRB-PUB-GSO	745
HMB-IRB-PUB-MDS	1260
HMB-IRB-PUB-NPU	1250
HMB-IRB-PUB-NPU-MDS	274

Visit the Data Use Limitations page for definitions of the consent levels above.

Total number of approved DARs: 6

Investigator:
Belloy, Michael
Institution:
Washington University in St Louis
Project Title:
Elucidating sex-specific risk for Alzheimer's disease through state-of-the-art genetics and multi-omics
Date of Approval:
March 31, 2026
Request status:
Approved
Research use statements:
Show statements
Technical Research Use Statement:
• Objectives: In this project, we seek to holistically investigate the genetic and molecular drivers of sex dimorphism in Alzheimer’s disease across ancestries. • Study design: This study integrates large-scale population genetics with multi-omics and endophenotype analyses. We are integrating all data available from ADGC and ADSP, together with other data from AMP-AD and biobanks such as UKB, FinnGen, and MVP to conduct large-scale multi-ancestry GWAS, rare-variant gene aggregation analyses, QTL studies, PWAS, TWAS, etc. We also particularly focus on X chromosome association studies. The study design also interrogates interactions with ancestry, hormone exposures, and with APOE*4, as well as comparisons to non-stratified GWAS/XWAS of Alzheimer’s disease. Further, we will also employ genetic correlation analyses, mendelian randomization, colocalization, and pleiotropy analyses, to interrogate overlap with other complex traits to better understand the mechanisms underlying sex dimorphism in Alzheimer’s disease. • Analysis plan, including the phenotypic characteristics that will be evaluated in association with genetic variants: Our phenotypes will include Alzheimer’s disease risk, conversion risk, various endophenotypes (including amyloid/tau biomarkers, brain imaging metrics, etc.) as well as molecular traits. As noted above, we will conduct large-scale multi-ancestry GWAS, XWAS, rare-variant gene aggregation analyses, QTL studies, PWAS, TWAS, etc. Specific aims include interrogating these question and analyses on (1) the autosomes, (2) the X chromosome, and (3) leveraging sex stratified QTL studies to drive discovery of risk genes.
Non-Technical Research Use Statement:
Alzheimer’s disease (AD) manifests itself differently across men and women, but the genetic and molecular factors that drive this remain elusive. AD is the most common cause of dementia and till today remains largely untreatable. It is thus crucial to study the genetics of AD in a sex-specific manner, as this will help the field gain important insights into disease pathophysiology, identify novel sex-specific risk factors relevant to personalized genetic medicine, and uncover potential new AD drug targets that may benefit both sexes. This project uses large-scale genomics and multi-omics to elucidate novel sex agnostic and sex-specific AD risk genes. We will interrogate sex dimorphism for AD risk on the autosomes and the sex chromosomes. We similarly interrogate sex dimorphism in the genetic regulation of gene expression and protein levels, which we will integrate with genetic risk for Alzheimer’s disease to further discovery risk genes. Throughout, we will also interrogate how sex-specific risk for AD interactions with hormone exposures, ancestry, and the APOE*4 risk allele.
Investigator:
Cruchaga, Carlos
Institution:
Washington University School of Medicine
Project Title:
The Familial Alzheimer Sequencing (FASe) Project
Date of Approval:
January 21, 2026
Request status:
Approved
Research use statements:
Show statements
Technical Research Use Statement:
The goal of this study is to identify new genes and mutations that cause or increase risk for Alzheimer disease (AD), as well as protective factors. Individuals and families were selected from the Knight-ADRC (Washington University) and the NIA-LOAD study. Only families with at least three first-degree affected individuals were included. Families with pathogenic variants in the known AD or FTD genes, or in which APOE4 segregated with disease were excluded. At least two cases and one control were selected per family. Cases had an age at onset (AAO) after 65 yo and controls had a larger age at last assessment than the latest AAO within the family. Whole exome (WES) and whole genome sequencing (WGS) was generated for 1,235 individuals (285 families) that together with data from our collaborators and the ADSP family-based cohort (3,449 individuals and 757 families) will provide enough statistical power to identify new genes for AD. Dr. Tanzi (Harvard Medical School) will provide WGS from 400 families from the NIMH Alzheimer disease genetics initiative study. We will perform single variant and gene-based analyses to identify genes and variants that increase risk for disease in AD families. Single variant analysis will consist of a combination of association and segregation analyses. We will run family-based gene-based methods to identify genes that show and overall enrichment of variants in AD cases. We will also look for protective and modifier variants. To do this we will identify families loaded with AD cases, that also include individuals with a high burden of known risk variants but that do not develop the disease (escapees). We will use the sequence data and the family structure to identify variants that segregate with the escapee phenotype. The most promising variants and genes will be replicated in independent datasets (ADSP case-control, ADNI, Knight-ADRC, NIA-LOAD ). We will perform single variant and gene-based analyses to replicate the initial findings, and survival analysis to replicate the protective variants. We will select the most promising variants/genes for functional studies
Non-Technical Research Use Statement:
Family-based approaches led to the identification of disease-causing Alzheimer’s Disease (AD) variants in the genes encoding APP, PSEN1 and PSEN2. The identification of these genes led to the A?-cascade hypothesis and to the development of drugs that target this pathway. Recently, we have identified rare coding variants in TREM2, ABCA7, PLD3 and SORL1 with large effect sizes for risk for AD, confirming that rare coding variants play a role in the etiology of AD. In this proposal, we will identify rare risk and protective alleles using sequence data from families densely affected by AD. We hypothesize that these families are enriched for genetic risk factors. We already have sequence data from 695 families (2,462 individuals), that combined with the ADSP and the NIMH dataset will lead to a dataset of more than 1,042 families (4,684 individuals). Our preliminary results support the flexibility of this approach and strongly suggest that protective and risk variants with large effect size will be found, which will lead to a better understanding of the biology of the disease.
Investigator:
Kamboh, M. Ilyas
Institution:
University of Pittsburgh
Project Title:
Genetics of Alzheimer's Disease and Endophenotypes
Date of Approval:
March 31, 2026
Request status:
Approved
Research use statements:
Show statements
Technical Research Use Statement:
Objectives: We are requesting access to the NIAGADS datasets to augment our ongoing studies on the genetics of Alzheimer’s disease (AD) and AD-related endophenotypes being carried out by Kamboh and his group since 1995. We are doing GWAS using array genotypes, whole-exome sequencing and whole-genome sequencing on datasets derived from University of Pittsburgh ADRC and ancillary population-based longitudinal studies on dementia and biomarkers. Different available phenotypes include AD and non-AD dementia, age-at-set, disease progression and survival, neuroimaging, cognitive decline, plasma biomarkers for the core ATN and non-ATN pathologies. We also plan to expand on gene-gene interaction and sex-stratified analyses which require the actual genotype data. The NIAGADS datasets will be used for replication and meta-analysis, and for gene-gene interaction and sex-stratified analyses. Study Design: A case-control design will incorporate a diverse cohort of individuals with AD and age-matched controls. For quantitative traits (neuroimaging and plasma biomarkers, cognitive performance measures, indicators of disease progression), linear regression analyses will be performed to identify genetic loci. To ensure the findings are robust and inclusive, participants from diverse demographic backgrounds will be included, enabling the exploration of potential genetic variations across populations. Analysis Plan: We will conduct GWAS and targeted analyses on candidate genes on different AD and AD-related phenotypes. Primary phenotypic variables include AD disease status, age-at-onset, last age for controls, APOE genotype, cognitive decline trajectories, sex, and race. Analyses will evaluate the influence of specific genetic variants on disease risk, cognitive performance, and biomarker levels, considering both individual and interactive effects of the APOE genotype. Results will be adjusted for potential confounders, such as demographic factors, to ensure valid associations. Detail analytical methods are described in our published papers for case-control (PMID: 32651314;35694926), quantitative traits (PMID: 30361487;37666928), and cognitive decline (PMID: 37089073; 30954325).
Non-Technical Research Use Statement:
Our research group at the University of Pittsburgh (Pitt), has been working on the genetics of Alzheimer’s disease (AD) and AD-related endophenotypes for almost three decades, on data derived largely from the University of Pittsburgh Alzheimer’s Disease Research Center and ancillary dementia studies. We are requesting access to the NIAGADS genotype and phenotype datasets to augment our sample size to increase power to detect novel genetic associations with AD and related endophenotypes.
Investigator:
Konermann, Silvana
Institution:
Arc institute
Project Title:
Modeling Alzheimer’s disease risk and associated molecular phenotypes
Date of Approval:
August 8, 2025
Request status:
Approved
Research use statements:
Show statements
Technical Research Use Statement:
The objective of the proposed research is to determine the relationship between Alzheimer’s disease (AD) genetic risk and associated molecular phenotypes. Genotype data will be used to compute a polygenic risk score (PRS) for disease-affected and control (non-disease-affected) participants. Statistical regression and mediation analyses will be used to model variation of molecular phenotypes with respect to PRS and, where available, pathology stage or cognitive impairment. Molecular phenotypes to be analyzed include bulk/single-cell/single-nucleus transcriptome, epigenome, proteome, metabolome, lipidome, amyloid, and tau. Molecular phenotypes of participants, including controls, will be matched with molecular phenotypes of in vitro cellular models, informing the design of in vitro perturbation experiments that recapitulate the genetic drivers of AD risk.
Non-Technical Research Use Statement:
Our goal is to determine the relationship between human genetic profiles associated with Alzheimer’s disease (AD) risk and specific measurable characteristics of human cells. Using multiple statistical analysis methods, we will build quantitative models that describe how those characteristics vary as a function of AD genetic risk. The models we build will help us design in vitro cellular systems that reflect different levels of AD risk, enabling experiments that inform new strategies for treating or preventing AD.
Investigator:
Seshadri, Sudha
Institution:
Glenn Biggs Institute for Alzheimer's and Neurodegenerative Diseases, University of Texas Health Sciences Center, San Antonio, TX
Project Title:
Therapeutic target discovery in ADSP data via comprehensive whole-genome analysis incorporating ethnic diversity and systems approaches
Date of Approval:
August 12, 2025
Request status:
Approved
Research use statements:
Show statements
Technical Research Use Statement:
Objective: Utilize ADSP data sets to identify genes & specific genetic variants that confer risk for or protection from Alzheimer disease. Aim 1: Using combined WGS/WES across the ADSP Discovery, Disc-Ext, and FUS Phases, including single nucleotide variants, small insertion/deletions, and structural variants. We will: Aim 1a. Perform whole genome single variant and rare variant case/control association analyses of AD using ADSP and other available data; Aim 1b. Target protective variant identification via association analysis using selected controls within the ADSP data and performing meta analysis across association results based on selected controls from non-ADSP data sets. Aim 1c. Perform endophenotype analyses including cognitive function measures, hippocampal volume and circulation beta-amyloid ADSP data in subjects for which these measures are available. Meta analysis will be conducted across ADSP and non-ADSP analysis results. Aim 2: To leverage ethnically-diverse and admixed populations to identify AD variants we will: Aim 2a. Estimate and account for global and local ancestry in all analyses; Aim 2b. Perform admixture mapping in samples of admixed ancestry; and Aim 2c. Perform ethnicity-specific and trans-ethnic meta-analyses. Aim 3: To identify putative therapeutic targets through functional characterization of genes and networks via bioinformatics, integrative ‘omics analyses. We will: Aim 3a. Annotate variants with their functional consequences using bioinformatic tools and publicly available “omics” data. Aim 3b. Prioritize results, group variants with shared function, and identify key genes functionally related to AD via weighted association analyses and network approaches. Analyses will be performed in coordination with the following PIs. Coordination will involve sharing expertise, analysis plans or analysis results. No individual level data will be shared across institutions. Philip De Jager, Columbia University; Eric Boerwinkle & Myriam Fornage, U of Texas Health Science Center, Houston; Sudha Seshadri, U of Texas, San Antonio; Ellen Wijsman, U of Washington. William Salerno, Baylor College of Medicine
Non-Technical Research Use Statement:
This proposal seeks to analyze existing genetic sequencing data generated as part of the Alzheimer’s Disease Sequencing Project (ADSP) including the ADSP Follow-up Study (FUS) with the goal of identifying genes and specific changes within those genes that either confer risk for Alzheimer’s Disease or provide protection from Alzheimer’s Disease. Analytic challenges include analysis of whole genome sequencing data, appropriately accounting for population structure across European ancestry, Hispanic, and African American participants, and interpreting results in the context of other genomic data available.
Investigator:
Zhan, Huixin
Institution:
New Mexico Institute of Mining and Technology
Project Title:
AI-Driven Analysis of Genetic and Transcriptomic Data in Alzheimer’s Disease
Date of Approval:
March 30, 2026
Request status:
Approved
Research use statements:
Show statements
Technical Research Use Statement:
Objectives: This study aims to improve understanding of the genetic and molecular mechanisms underlying Alzheimer’s disease (AD) by applying advanced computational and deep learning models to existing genomic and transcriptomic datasets. Specifically, we seek to identify and characterize genetic variants associated with AD risk, progression, and related phenotypes, contributing to precision medicine approaches for neurodegenerative disorders. Study Design: This project involves secondary analysis of de-identified, controlled-access datasets from NIAGADS (NG00067, NG00116, NG00174, NG00027, NG00075). No new data will be collected. The data will be securely downloaded and analyzed on institutional servers at New Mexico Tech under an approved IRB and Data Use Agreement. Analysis Plan: We will integrate genomic, transcriptomic, and phenotypic data to develop and evaluate machine learning models—such as large language model–based architectures and disease-specific neural networks—to predict variant pathogenicity and gene-level associations. Phenotypic characteristics evaluated will include Alzheimer’s disease diagnosis, cognitive performance measures, neuropathological burden, and biomarker profiles (e.g., amyloid and tau levels). Statistical and model-based analyses will assess associations between genetic variants and these phenotypes, with results reported in aggregate, non-identifiable form. Collaborations (if applicable): N/A
Non-Technical Research Use Statement:
This project uses advanced artificial intelligence and statistical tools to study the genetic and molecular factors that contribute to Alzheimer’s disease. By analyzing existing, de-identified research data from the National Institute on Aging’s NIAGADS repository, we aim to identify genetic variants and biological pathways linked to disease risk and progression. The study will combine information from DNA and gene-expression data to build computer models that can better predict how certain genetic changes affect brain health. Our ultimate goal is to improve scientific understanding of Alzheimer’s disease and support future efforts in early detection and personalized treatment.

Total number of samples: 14,314

Female 8,589 60.0 %

Male 5,703 39.8 %

Sex not reported: 22 (0.2%)

American Indian/Alaska Native	2
Native Hawaiian or Other Pacific Islander	1
Black or African American	1,020
White	11,776
Other	1,490
NA	25

AD
Control	5,909	41.3%
Case	7,654	53.5%
Other	9	0.1%
Unknown	720	5.0%

NG00116 – Resolving Mutations in Challenging Genomic Regions to Test Association with Disease Phenotypes

Overview

Description

Sample Summary per Data Type

Available Filesets

Data Dictionary Files

Sample Information

Data Releases

Related Studies

Cohorts

Phenotype Harmonization

Consent Levels

Acknowledgement

Acknowledgment statement for any data distributed by NIAGADS:

For investigators using any data from this dataset:

For investigators using Resolving mutations in challenging genomic regions to test association with disease phenotypes (sa000042) data:

Publications

Third-Party Access

Approved Users

Total number of samples: 14,314

Source Datasets