Showing posts with label Credible sets. Show all posts
Showing posts with label Credible sets. Show all posts

Monday, October 8, 2018

DIAMANTE GWAS dataset adds close to a million samples along with fine-mapping to the T2DKP

In a groundbreaking paper published today, Anubha Mahajan and colleagues (Mahajan et al., Nature Genetics 2018) report on a meta-analysis of unprecedented size for genetic associations with type 2 diabetes (T2D) along with fine-mapping analyses to identify causal variants that can suggest new therapeutic targets. We are pleased to provide access to the summary results as well as the results of the fine-mapping today in the T2D Knowledge Portal (T2DKP).

Working as part of the DIAGRAM (DIAbetes Genetics Replication And Meta-analysis) and DIAMANTE (DIAbetes Meta-ANalysis of Trans-Ethnic association studies) consortia, the researchers aggregated and meta-analyzed genome-wide association studies for about 900,000 individuals of European ancestry (about 74,000 T2D cases and 824,000 controls). The studies were imputed using the most comprehensive reference panels possible, and in all, the analysis considered about 27 million genotyped or imputed variants.

After performing T2D association analysis (both unadjusted and adjusted for body mass index) 243 loci were seen to be associated with T2D at genome-wide significance or better (p-value for association ≤ 5 x 10-8). Of these, 135 were novel--not detected previously in any T2D association analysis to date.

Within these loci, each of which included multiple significantly associated variants, the researchers performed approximate conditional analysis to determine whether the associations were independent of each other. They found surprising complexity within some loci; for example, the well-known TCF7L2 locus appears to include as many as 8 distinct association signals!

All of the T2D associations from this study may be viewed in the T2DKP. They are represented in two datasets, named "DIAMANTE (European) T2D GWAS" and "UK Biobank T2D GWAS (DIAMANTE-Europeans Sept 2018)."  Manhattan plots showing the distribution of the associations across the genome may be seen by selecting either the "Type 2 diabetes" or "Type 2 diabetes adj BMI" phenotypes from the phenotype selection menu on the T2DKP home page. On Gene pages of the T2DKP, the results may be viewed in tables of variant associations and in the interactive LocusZoom visualization (see below). Results from this study are also displayed on Variant pages of the T2DKP.


LocusZoom plot on the PPARG Gene page


The credible set analysis performed in this study is also incorporated into the T2DKP. On the "Credible sets" tab of Gene pages, you may choose to visualize any of the credible sets available for the region. Epigenomic annotations that overlap the positions of the variants in the credible set are presented in an interactive display that allows you to select particular chromatin states or tissues to view. In the example shown below, one of the credible sets in the TCF7L2 region includes just two variants, and the one with the highest posterior probability overlaps active enhancer regions in adipose and liver tissue--both of which are important for T2D.


Detail of the Credible sets tab of the TCF7L2 Gene page

The multiple causal variants identified in this study support previous investigations on the biological mechanisms behind T2D and suggest new hypotheses that will likely lead to therapeutic insights. After reading the paper and a blog post from the authors, we invite you to explore the results in the T2DKP and to contact us with any suggestions or questions!

Monday, January 22, 2018

GWAS data re-analysis yields novel results about T2D risk



"Waste not, want not." The old proverb is about frugality, but a study published today gives it a whole new dimension. Lead author Sílvia Bonàs, directed by Josep Mercader and David Torrents and collaborating with many colleagues at the Barcelona Supercomputing Center, the Broad Institute, and other institutions (Bonàs-Guarch et al. (2018), Nature Communications 9), decided to investigate variants associated with type 2 diabetes (T2D) by re-analyzing existing GWAS data rather than initiating a new study.

This was a frugal strategy, conserving both time and resources. But the benefits of this approach went way beyond frugality. By aggregating multiple datasets and using unified, current methods for quality control, imputation, and association analysis, the researchers discovered nuggets of significant information that were not apparent in the original analyses of the individual sets. And all of these nuggets are freely available for browsing and searching in the T2D Knowledge Portal (T2DKP).

To amass these data, the researchers combined all of the individual-level T2D case-control GWAS data that were available from the European Genome-Phenome Archive (EGA) and the database of Genotypes and Phenotypes (dbGaP). After harmonization and quality control, data from 70,127 subjects (12,931 cases and 57,196 controls) remained, inspiring them to name the project "70KforT2D".

In the time since the original studies had been performed, better and more comprehensive reference panels for imputation had been generated by the 1000 Genomes and UK10K projects. By using both of these panels for imputation, the researchers were able to substantially increase the number of variants that could be imputed. They ended up with more than 15 million variants, including more than 5 million rare variants and over 1.3 million indels, which have previously been difficult to impute.

In performing association analysis, the authors took advantage of existing large datasets of T2D association summary statistics for meta-analysis, being careful to only combine non-overlapping samples. They also took advantage of the T2D Knowledge Portal to verify some associations for low-frequency variants that were located in coding regions and had suggestive, but not unambiguously significant, p-values. The significance of the T2D associations of these variants was confirmed by meta-analysis along with the associations seen in two large studies in the T2DKP (GoT2D exome chip analysis, with nearly 80,000 samples, and the 17K exome sequence analysis dataset with 17,000 samples).

The association analysis identified 57 loci associated with T2D risk at the genome-wide significance level or better (p-value ≤ 5x10e-8), seven of which had not previously been associated with T2D. The high quality of the data made it possible to fine-map the variants at each of these loci and construct credible sets. Many of the putative causal variants—including those in previously identified loci—were indels rather than single-nucleotide polymorphisms, underscoring the importance of an imputation procedure that discovers indels.

The T2D-associated loci discovered in this study give some tantalizing hints about genes potentially involved in T2D, and suggest new avenues for detailed wet-lab investigation. We can’t review all of them in this space, but one association is particularly interesting for the generalizable lessons it teaches us about case-control GWAS for T2D.

This association, which the authors validated and replicated using additional datasets, involves the X chromosome variant rs146662075. The risk allele confers a 2-fold elevated risk of developing T2D, in males. The variant appears to affect an enhancer that could regulate expression of AGTR2, a gene known to be involved in modulating insulin sensitivity—making it a very interesting subject for investigation with regard to T2D. More work is needed to figure out whether this is really a male-specific effect, or whether it was only detectable in males because imputation for the X chromosome is more accurate in males, who have only one copy of the chromosome.

The first lesson learned from this association is that the X chromosome harbors important loci, and deserves attention in association studies. While this seems obvious, since the X chromosome comprises 5% of the genome, it has been neglected in most studies to date.

The second lesson is that for an adult-onset disease like T2D, it’s very important to pay attention to the details of case-control classification. If there are young people in the control group, they may actually be future T2D cases, destined to develop the disease later in life. When the authors tried to replicate the initial discovery for this variant in different datasets, the associations were not as significant as expected. But after digging deeper into the experimental cohorts, they found that most of the replication datasets had many subjects younger than 55, which was the average age for T2D onset for these cohorts. Re-running the analysis after excluding controls younger than 55 and also excluding those who appeared to be pre-diabetic, based on an oral glucose tolerance test, brought the replication results into concordance with the discovery results and confirmed the significance of the rs146662075 association.

In keeping with the spirit of open access, the authors provided the summary statistics from this work to the T2DKP even before publication. These results are incorporated into the T2DKP and are visible on Gene and Variant pages as well as searchable via the Variant Finder. The authors have also made the full summary statistics available for public download.

The novel and important findings from this study strongly reaffirm the value of data sharing. Not only are data sharing and re-analysis the right things to do for reasons of fairness, equity, and frugality; they can also spark new insights and move science forward in unexpected ways.

Monday, June 12, 2017

T2D Knowledge Portal now distills and summarizes genetic information for individual genes

The Type 2 Diabetes (T2D) Knowledge Portal presents genetic data relevant to T2D on two major types of page: Variant pages for individual variants, or SNPs; and Gene pages focusing on individual genes. Visual displays on Variant pages provide an immediate indication of the possible significance of each variant for T2D. But until now, Gene pages have presented large amounts of information from disparate sources without much integration or interpretation to guide the viewer.

Now, that has all changed with our release of the new Gene page. It guides researchers through an organized workflow that can help them take advantage of the aggregated data in the Portal to move from a variant of interest, to a gene of interest, to an assessment of the potential involvement of that gene’s product in T2D.

The central feature of the new Gene page is an at-a-glance display that summarizes the strength of the evidence for associations of the gene with T2D or related traits. An algorithm scans the comprehensive collection of datasets within the Portal to find data on variants in the gene, and the overall conclusion is shown by a “traffic light” icon. A green light indicates that there is strong evidence for association of at least one variant in the gene with at least one phenotype; a yellow light indicates that there is suggestive evidence, and a red light indicates that the data aggregated in the Portal contain no evidence for associations of variants within this gene.

Figure 1. Traffic light display for MTNR1B


Several sections of the page below the traffic light allow the user to drill down to much more information about the variants within the gene, their individual associations, and their collective impact on the disease burden of the gene. An interactive LocusZoom plot allows users to view the linkage disequilibrium relationships and associations from multiple datasets, with a wide variety of phenotypes, for common variants. The plot also displays the location of chromatin states, which can indicate the regulatory role of a region, in multiple tissues.


Figure 2. LocusZoom plot of the credible set of T2D-associated variants in MTNR1B (above) and chromatin state annotations for the region (below).

In the example shown above, the traffic light (Fig. 1) shows that variants in the MTNR1B gene encoding the melatonin receptor have one or more strong phenotypic associations (view the MTNR1B Gene page in the T2D Knowledge Portal). The table of common variants for MTNR1B (not shown) tells us that the most significantly associated variant is rs10830963. And a view of the LocusZoom plot for the credible set of variants associated with T2D (Fig. 2, top) shows that in fact the credible set for this region contains only rs10830963, further supporting its significance. The chromatin state annotations for this region (Fig. 2, bottom) provide evidence for a regulatory effect in pancreatic islets, consistent with a potential role in T2D. This information, easily found in the Portal today, replicates the results of a 2015 genetic analysis that required over 100 authors (Gaulton, KJ, et al. (2015) Nature Genetics 47:1415).

The new Gene page presents a lot of information and we can't cover it all in this space. But don't worry, we've created a guide to the page that explains every feature in detail. It's linked from the top of the page, or you can download it here.

With the inclusion of the new Gene page, the Portal now enables the rapid generation of testable hypotheses, by integrating, interpreting, and presenting information that previously could only be generated by coordinated research across a consortium. This new development brings the T2D Knowledge Portal project one step closer to informing the discovery of new targets and treatments for T2D.