Showing posts with label T2DKP. Show all posts
Showing posts with label T2DKP. Show all posts

Wednesday, November 7, 2018

Meet the Knowledge Portal team at AHA

This weekend, cardiovascular researchers from around the globe will be meeting in Chicago for the 2018 Scientific Sessions of the American Heart Association. Members of the Knowledge Portal Network team will be there to meet and talk with geneticists and biologists who use the Portals and get your input on how we can improve them.

Please come visit us at booth #2249 in the Exhibit Hall! We'll be there on Saturday, Nov. 10 from 11am-5pm; on Sunday, Nov. 11 from 10am-4:30pm; and on Monday, Nov. 12 from 10am-3pm.

Wednesday, September 26, 2018

New datasets and many new phenotypes in the T2DKP

Today we release several new datasets, including associations for many new phenotypes and individual-level data for secure interactive analysis, to the Type 2 Diabetes Knowledge Portal.

The AAGILE GWAS dataset, from the African American Glucose and Insulin Genetic Epidemiology (AAGILE) Consortium, brings more diversity of ancestry to the T2DKP, with meta-analysis of fasting glucose and BMI-adjusted fasting insulin associations from over 20,000 African American individuals. These results were combined with associations for over 57,000 individuals of European ancestry from the Meta-Analyses of Glucose and Insulin-related traits Consortium (MAGIC) in a trans-ethnic meta-analysis.

This release also adds two new diabetic kidney disease datasets from the SUMMIT (SUrrogate markers for Micro- and Macro-vascular hard endpoints for Innovative diabetes Tools) consortium. All of the more than 40,000 subjects in the "Diabetic Kidney Disease GWAS: subjects with T1D or T2D" dataset had either type 1 or type 2 diabetes. The study measured seven different renal phenotypes in these subjects, including four that are new to the T2DKP. Summary association results are available for the entire group and for sub-cohorts that separate T1D from T2D and European from Asian ancestry. A separate dataset from SUMMIT, "Diabetic Kidney Disease GWAS: subjects with T1D or T2D, ESRD vs. controls" is comprised of more than 5,600 diabetics, nearly 1,200 of whom had end-stage renal disease. These two datasets greatly expand the range of diabetic complications for which genetic association data are available in the T2DKP.

The T2DKP is federated, meaning that in addition to the Data Coordinating Center at the Broad Institute, some results are drawn from a sister site at the European Bioinformatics Institute (EMBL-EBI). This system allows data that may not leave Europe to be represented in the T2DKP. Six of the new datasets in this release are housed at the T2DKP Federated Node at EMBL-EBI.

The Hoorn Diabetes Care System (DCS) dataset includes associations for 12 different anthropometric, blood lipid, blood pressure, and liver and kidney function measures for a cohort of over 3,400 type 2 diabetics in the Netherlands.





The GoDarts project (Genetics of Diabetes Audit and Research in Tayside Scotland) recruits type 2 diabetics and matching controls in the Tayside region of Scotland. This release includes five new datasets from GoDarts, representing experiments performed using different arrays. Each experiment determined genetic associations for a wide variety of phenotypes, including two that are new to the T2DKP: levels of adiponectin and leptin, hormones that are associated with risk of T2D and obesity.

Results from all of these datasets may be searched using the Variant Finder tool and may be browsed:

• On Gene Pages in the Common variants and High-impact variants tables and in LocusZoom plots;

• On Variant Pages in the Associations at a glance section, the Associations across all datasets section, and in LocusZoom plots;

• From the View full genetic association results for a phenotype search on the home page: first select a phenotype, then select a dataset on the resulting page.


Individual-level data from the Hoorn DCS and GoDarts datasets also power secure interactive analyses using the Genetic Association Interactive Tool (GAIT) on Variant Pages. With the new additional data, nearly 61,000 individual-level samples are now available for custom association analysis.

Please take a look at the new results and contact us any time with questions or suggestions!

Friday, May 11, 2018

T2DKP Spring Newsletter

The latest issue of our quarterly newsletter is now available. Download it here and get the latest!

Friday, April 27, 2018

New T2DKP release adds individual-level data for interactive analysis

With the April release of the Type 2 Diabetes Knowledge Portal, we are increasing the number of datasets and samples available for interactive analysis via the LocusZoom and GAIT tools. These tools now access individual-level data from three additional datasets, all of which were quality controlled and analyzed at the Accelerating Medicines Partnership in Type 2 Diabetes (AMP T2D) Data Coordinating Center (DCC):
  • CAMP GWAS: 3,628 multi-ancestry samples from the MGH Cardiology and Metabolic Patient cohort, generated by a public-private partnership between Pfizer Inc. and Massachusetts General Hospital;
  • METSIM GWAS: 8,791 European ancestry samples from the Metabolic Syndrome in Men study.
These individual-level data are available as "dynamic" datasets, powered by Hail software, in LocusZoom on Gene pages and Variant pages of the T2DKP, for the following phenotypes: 
  • BioMe AMP T2D GWAS: type 2 diabetes, BMI, diastolic blood pressure, fasting glucose, HbA1c, HDL cholesterol, LDL cholesterol, systolic blood pressure
  • CAMP GWAS: type 2 diabetes, BMI, fasting glucose, fasting insulin
  • METSIM GWAS: type 2 diabetes, BMI, diastolic blood pressure, fasting glucose, fasting insulin, HbA1c, HDL cholesterol, LDL cholesterol, systolic blood pressure
To perform interactive analyses on these data in LocusZoom, select one of the available phenotypes in step 1 and then choose a "dynamic" dataset in step 2.


When you click on a variant in the resulting LocusZoom plot, the option to condition on that variant appears in the tooltip:


Clicking on that link starts on-the-fly association analysis for the region while conditioning on that variant, which can reveal whether association signals are independent of each other. You can choose to condition on multiple variants. The variants of your choice are listed in the upper left-hand corner of the plot, and the list may be edited:



Individual-level data from these three datasets are also available for interactive analysis via the Genetic Association Interactive Tool (GAIT) on Variant Pages. After selecting one of the datasets, you will be able to choose a phenotype for association analysis, filter the sample pool by specifying a range of values for one or more phenotypes, choose custom covariates, and then run on-the-fly association analysis for your chosen subset of samples. Find all of the details about how to use this tool in our GAIT guide.

We hope that the increased ability to interact with individual-level data in the T2DKP will be helpful to your research. As always, we are happy to answer any questions about these or other data and tools; please contact us for help.

Tuesday, April 17, 2018

Developing a model for collaborative science: a mid-term perspective on the AMP T2D Partnership

In 2011, Dr. Francis Collins, Director of the National Institutes of Health (NIH), met with leaders in biomedical research to discuss a frustrating problem. Continual improvements in molecular biological and genomic techniques were generating an avalanche of data relevant to complex diseases, yet the translation of these data into insights about disease mechanisms and drug targets was unacceptably slow. It was clear that an entirely new paradigm for collaborative research would be needed to speed up the extraction of knowledge from data.

The result of these discussions was the creation of the Accelerating Medicines Partnership (AMP), one branch of which focuses on type 2 diabetes (T2D)—a life-threatening disease that affects hundreds of millions of people worldwide, whose incidence is growing, and whose progression cannot yet be effectively stopped or reversed. AMP T2D, a five-year project, includes the National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK); the pharmaceutical companies Janssen Pharmaceuticals, Eli Lilly and Company, Merck, Pfizer, and Sanofi; the University of Michigan; the University of Oxford; the Broad Institute; and other researchers around the globe. The Foundation for the National Institutes of Health (FNIH) also provides funding and coordination for the project.

Drawing on the strengths of both academia and industry, this public-private partnership brings together all stakeholders in a pre-competitive space to share data and combine resources, with the goal of validating new drug targets faster. Now in Spring 2018, roughly mid-way through the funding period, it is evident that this collaboration has resulted in remarkable progress on both scientific and collaborative fronts.

Genetic association data: the foundation of AMP T2D

Genetic association studies interrogate the genomes of individuals at millions of specific genomic positions to discover sequence variants that are correlated with the incidence of disease. From the outset, AMP T2D aimed to support the generation of unprecedented amounts of new genome-wide association study (GWAS), exome sequencing, and whole-genome sequencing data within the project as well as their aggregation with all relevant publicly available data. Originally, 5 sites were funded by the NIDDK to generate new data and deposit them into the AMP T2D Data Coordinating Center (DCC) at the Broad Institute. As the project evolved, another site was funded by the NIDDK and 8 more sites were funded by the FNIH. Additionally, an Opportunity Pool of funds from the NIDDK was created, allowing the AMP T2D Steering Committee to award smaller grants for complementary research projects in a flexible, science-driven manner.  Currently 10 Opportunity Pool projects are in progress, and more awards will be given in the future.

Not only has the number of genetic association studies increased since the inception of AMP T2D, but also the number of samples surveyed in each has grown dramatically, from typically under 100,000 to approaching 1 million today. The increased statistical power conferred by these large sample sizes has led to a huge increase in the number of loci found to be significantly associated with T2D, from about 70 at the start of the project to nearly 430.

Improvements in genomic technologies in the past few years have allowed AMP T2D collaborators to generate increasing amounts of sequencing data, which make it possible to comprehensively interrogate all alleles and to uncover rare variation. At the project’s start, T2D associations with exome sequences (covering the protein-coding regions of the genome) were available for about 13,000 samples, and no whole-genome sequencing studies had been published. Now, more than 2,600 whole genomes are available, and analysis of a set of 50,000 exomes—the largest disease-specific aggregation of exome sequencing data to date—is nearly complete. Importantly, many of the associations that have been newly discovered in sequencing studies involve relatively rare variants that affect protein-coding regions. It is often more straightforward to develop hypotheses about the impact of such variants than it is for variants outside of coding regions.

As the AMP T2D partnership has grown in prominence in the diabetes field, the DCC has been approached by investigators outside the project who want to contribute their data in order to aggregate and display them in the context of AMP T2D data. In early 2017, researchers in the 70kforT2D project, which found novel T2D associations by re-analyzing existing GWAS data, offered their results for integration into the DCC and display in the Type 2 Diabetes Knowledge Portal (T2DKP; see below) before publication.

70kforT2D GWAS was first pre-publication dataset to be added to the T2DKP from outside the AMP T2D partnership, and it was particularly appropriate that these scientists, whose results illustrate the value of data sharing, themselves chose to freely share their results. Incorporation of datasets into the AMP T2D DCC and T2DKP offers investigators the chance to take advantage of the expertise of the AMP DCC analysis team, apply cutting-edge analysis tools to their data, and display their results broadly to the T2D research community in the context of multiple datasets. The AMP T2D DCC is open to incorporating T2D-relevant datasets from all investigators (find details on contributing data here).

In addition to the datasets generated by AMP T2D partners and other T2D researchers, which focus on associations with T2D, glycemic measures, and T2D complications, the AMP T2D DCC also collects publicly available genetic association datasets for traits relevant to T2D, such as anthropometric measures, blood pressure and lipid levels, and heart and kidney disease.

Orthogonal data types to help identify and prioritize causal variants and genes

Finding genetic variants that are associated with T2D risk is critically important to understanding the genetics of T2D, but it is only a first step. The most significantly associated variant in a genomic region may not be the causal variant that is responsible for altered T2D risk. Researchers perform fine mapping to analyze genetic associations in specific regions of the genome and generate credible sets—that is, sets of variants that are predicted to include the causal variant. Mid-way through the AMP T2D funding period, emphasis among the data-generating partners is beginning to shift from simply generating association data to performing fine mapping and credible set analysis.

But even after predicting which sequence variations are responsible for altered risk, finding clues about how they affect risk requires integration with additional data types. Information about the functional importance of the genomic region where a variant is located—its relevance to gene expression, protein function, networks and pathways, metabolite levels, and more, all determined on a tissue-specific basis—can help prioritize genes and pathways for in-depth experimental investigation. These kinds of research were built into AMP T2D from the beginning, and as the importance of these data types became even clearer, several Opportunity Pool awards were given to projects focusing on complementary data types that shed light on the significance of genetic associations.

Several of these projects focus on generating tissue-specific epigenomic data: histone modifications, DNA methylation, chromatin conformation, transcription factor binding, 3-dimensional chromosome structure, and other data types. Epigenomic data can provide important clues about the mechanisms by which sequence variation affects T2D risk, particularly for variants that lie outside of protein-coding regions. For example, if a risk-associated variant is seen to disrupt a transcription factor binding site, this would support the hypothesis that the transcription factor and its target genes are relevant to T2D.

To make these data accessible to researchers, one Opportunity Pool award supports the creation of the Diabetes Epigenome Atlas, which collects and displays epigenomic datasets relevant to T2D. In the near future, these data will be fully integrated with genetic association data in the Type 2 Diabetes Knowledge Portal (see below).

Other Opportunity Pool projects are concerned with processes downstream of gene expression. Discovering interactions between proteins implicated in T2D risk, for example, could help to uncover all of the players in pathways important for the development of T2D, increasing the number of potential drug targets. Determining the effects of variants on the levels of key metabolites can illuminate the metabolic pathways that change during the development of T2D. 

In addition to generating all of these orthogonal data types, AMP T2D partners are developing algorithms and using machine learning to classify and prioritize variants on the basis of the functional annotations that accompany them. Finally, other Opportunity Pool projects will use model organisms to test and validate drug targets that are suggested by these analyses.

Tools and methods to speed analysis and interpretation

At the inception of AMP T2D it was also clear that the development of new methods and tools would need to accompany the generation of data, and support for these activities was built into the program. One major technical effort has addressed an obstacle to global data aggregation: because of institutional and national privacy regulations, some datasets may not leave their site of origin to be aggregated with other datasets at the AMP T2D DCC. A group at the European Bioinformatics Institute has built a technical replicate of the DCC and knowledgebase, such that data stored there are equally as accessible for browsing, searching, and interactive analysis as are the data stored at the AMP T2D DCC at the Broad Institute. This federation mechanism allows global data accessibility even when data aggregation is not permitted.

Other efforts supported by AMP T2D are aimed at improving the speed and efficiency at which data can be taken in and analyzed. In one project, a data intake system is being developed that will streamline the process for both data submitters and for the DCC team, and will be applicable to data submission both at the Broad DCC and at other federated sites. Another project has created a software pipeline, LoamStream, that will largely automate quality control and association analysis of incoming data. Currently, LoamStream is in use for quality control of genotype data, and this has already greatly reduced the time required to process new datasets. Future work will extend the pipeline to association analysis and will also allow it to take in sequence data as well as genotype data.

A genetic association of a variant with T2D gains credibility if multiple independent studies replicate the association. Thus, it is important for researchers to be able to evaluate the weight of available evidence. But currently this is difficult to assess from the association datasets in the AMP T2D DCC, because many are based on overlapping sets of subjects. AMP T2D partners at the University of Michigan and University of Oxford are working on a method to take these overlaps into account and synthesize associations from multiple datasets into a “bottom-line” significance for association of a variant with T2D, which will aid in prioritizing variants for future work.

Multiple AMP T2D projects for analysis, interpretation, and custom interactive analysis of variant-phenotype associations are ongoing at the Universities of Michigan, Chicago, and Oxford, Vanderbilt University, and the Broad Institute. These projects are aimed at facilitating, in various ways, the path from variant associations to functional knowledge, and all have been or will be integrated into the T2D Knowledge Portal (see below).

Hail software offers a pipeline that speeds up the analysis of huge genomic datasets, while the gnomAD resource aggregates and harmonizes exome and genome sequences to provide a catalog of genetic diversity, in more than 100,000 humans, that aids in interpretation of variant associations with disease. A tool under development in the gnomAD project will display the effects of variants on protein structures as another way to deduce their potential impact.

Other analysis modules include gene-based association methods for using expression data to predict genes that may impact a phenotype (PrediXcan and MetaXcan), and a phenome-wide association study (PheWAS) method for visualization of the associations of a variant across multiple phenotypes, which is a crucial consideration during drug development. 

The interactive visualization tool LocusZoom will integrate many of these methods to display variant associations and credible sets, epigenomic and functional annotations, and phenotype associations across a genomic region as well as offering custom association analysis.

An example LocusZoom plot


AMP T2D Knowledge Portal: democratizing T2D genetic results for researchers world-wide

AMP T2D was founded on the idea that in order to truly accelerate progress, genomic information must be freely accessible to all scientists and presented in a way that is understandable by a broad range of researchers working on T2D biology, not only by human geneticists and bioinformaticians with special computational skills. So the roadmap for the project included not only data generation and analysis, but also the production of a publicly available web resource that would integrate data types, interpret the evidence, and present of all these results. 

While it is under continuous development, mid-way through the initial funding period the T2D Knowledge Portal (T2DKP) is already a well-established resource. Other web resources collect genetic association data, but the T2DKP is unusual in providing harmonized datasets to which a consistent analysis pipeline has been applied. Rather than simply cataloging datasets, it offers distilled and synthesized results along with their interpretation, to guide more detailed exploration of the evidence. And, unlike any other extant resource, it offers researchers the ability to perform interactive queries on protected individual-level data. 

T2DKP home page

The Gene page of the T2DKP (see an example) illustrates the presentation of immediately understandable summary information along with the opportunity to drill down to the details. An algorithm considers the associations of all variants across a gene, for all phenotypes and in all datasets aggregated at the DCC, and calculates from them a “traffic light” signal for the gene: green to indicate that there is a significant association for at least one phenotype; yellow to indicate suggestive, if not highly significant, associations; and red to indicate that there is no evidence for association for any of the phenotypes considered in the T2DKP. Below this, tables and graphics invite users to explore all variants across the gene, their impacts on the encoded protein, and their associations, as well as their positions relative to epigenomic marks across the region in multiple tissues.

The T2DKP currently offers the ability to run custom, interactive association analyses using two different tools. In the LocusZoom visualization, users may choose one or more variants as covariates before performing association analysis. The Genetic Association Interactive Tool (GAIT) for single variant associations, which also powers the custom burden test for gene-level associations, is even more versatile, presenting the distributions of different characteristics of the sample set (age, sex, BMI, glycemic measures, blood lipid levels, and many more) and allowing users to filter the set by multiple criteria and to choose custom covariates before performing association analysis. Both of these tools allow analytical access to the individual-level data, whether housed at the Broad DCC or at the EBI federated node, in a secure environment so that data privacy is always protected.

Evolution of a collaborative environment

AMP T2D organization


The AMP T2D partnership is a multifaceted project (illustrated above) that embraces several aspects of basic research and combines them with building a product, the T2DKP. In connecting scientists both within and outside of consortia, in academia and in industry, working on genetic associations or functional studies, it is becoming the nexus of the T2D genetics community. Researchers are finding the T2DKP helpful for accessing even their own results and for viewing them in the context of multiple phenotypic associations and other complementary data types. Pharmaceutical partners are finding help via the Target Prioritization project, in which the tools and methods developed within AMP T2D are being used to prioritize a list of genes of mutual interest for further investigation.

Perhaps most importantly, AMP T2D has made researchers—both within and outside of the project—aware of the value of sharing data for representation in the context of all other relevant data. Only by compiling and interpreting all available information will we be able to make the best hypotheses about genes and pathways that are possible drug targets and prioritize them for in-depth functional investigation.

AMP T2D and beyond

In the remainder of the initial AMP T2D funding period, we expect continued progress in each of the areas discussed above. The data intake and analysis pipelines will be improved, and new data will be incorporated at an increasing pace—including data from the UK Biobank, which has generated association results for 500,000 genotyped subjects and more than 2,500 traits. Associations will be added for many more phenotypes related to T2D, including diabetic complications and longitudinal phenotype data that connect the development of various traits to the timeline of incident T2D.  Much more T2D-relevant epigenomic data will be available for query as well as for browsing, via dynamic connection with the Diabetes Epigenome Atlas. And entirely new data types (for example, metabolomic and proteomic data) arising from Opportunity Pool projects will be added to the T2DKP.

Ongoing work on tools and methods will result in the addition of many more interactive modules to the T2DKP. Researchers will be able to view PheWAS data; prune lists of variants by their linkage disequilibrium relationships; calculate credible sets and genetic risk scores with custom parameters; perform more versatile interactive burden tests; prioritize genes by pre-calculated association scores; overlay the positions of coding variants on protein structures to help assess their impact; and perform enrichment analysis on sets of loci to suggest pathways implicated in disease processes.

The Knowledge Portal platform developed for AMP T2D has already proved extensible to other complex diseases: in 2017, both the Cerebrovascular Disease and Cardiovascular Disease Knowledge Portals were launched. In the future, connections within the ecosystem formed by the T2D, Cerebrovascular, and Cardiovascular Portals will be improved, so that researchers can easily assess the impact of a variant or involvement of a gene for all of these related diseases. If funding and collaboration considerations allow, perhaps one day these Portals will merge into a single cardiometabolic disease genetics Knowledge Portal to accelerate the development of new therapeutics in this broader area.

Finally, the ultimate goal of this funding period is that by its end, the data generation, analysis, and interpretation will have facilitated the validation of multiple promising drug targets for further investigation. Given the rate of progress on multiple fronts, this seems a realistic goal. We hope that this unique collaborative environment will continue to accelerate T2D genetic research and will become a paradigm for other research communities.

Tuesday, March 6, 2018

T2DKP Winter Newsletter

The latest issue of our quarterly newsletter is now available. Download it here to find out what we've been up to!

Thursday, March 1, 2018

New release today, as the KPN moves to a regular release schedule

At the Knowledge Portal Network (consisting of the Type 2 Diabetes, Cardiovascular Disease, and Cerebrovascular Disease Knowledge Portals), we are establishing a regular bimonthly release schedule. Every other month, new data and features will be incorporated into the Portals. Today, we are pleased to announce the first of these releases.

New data in the Type 2 Diabetes Knowledge Portal

This release adds two new datasets to the T2DKP. The Diabetic Cohort - Singapore Prospective Study Program is a T2D case-control study to identify genetic and environmental risk factors for diabetes in Singapore Chinese. The DC-SP2 GWAS set, a meta-analysis of summary level T2D associations from 3,951 individuals, was contributed by Drs. Rob Martinus Van Dam, E Shyong Tai, and Xueling Sim from the National University of Singapore. They have also submitted individual-level data from this study to the Accelerating Medicines Partnership Data Coordinating Center (AMP DCC), and these data will be incorporated into the T2DKP after quality control and analysis are complete.

In addition to this set, we have incorporated the publicly available summary statistics from the DIAGRAM 1000G GWAS. This dataset, from the DIAGRAM (DIAbetes Genetics Replication And Meta-analysis) consortium, is a meta-analysis of 26,676 T2D cases and 132,532 control participants from 18 GWAS (Scott RA, et al. An Expanded Genome-Wide Association Study of Type 2 Diabetes in Europeans. (2017) Diabetes 66:2888). Samples were imputed using the all ancestries 1000 Genomes Project reference panel.

More details about both of these datasets are available on our Data page.

New features specific to the Type 2 Diabetes Knowledge Portal

We have expanded the range of data available for interactive analysis by adding individual-level data from the CAMP GWAS, BioMe AMP T2D GWAS, and METSIM GWAS datasets to the dynamic analysis modules LocusZoom and GAIT (Genetic Association Interactive Tool). LocusZoom, powered by the Hail software developed at the Broad Institute as part of the AMP T2D project, allows you to perform custom association analysis while conditioning on specific variants or sets of variants.

GAIT offers alternative options for custom association analysis, such as filtering samples by their phenotypic characteristics (e.g., age, BMI, cholesterol levels) and choosing specific covariates. To date, seven different datasets comprised of over 67,000 samples are available for dynamic analysis in GAIT. These include datasets housed both at the AMP DCC (19k exome sequence analysis; CAMP GWAS; BioMe AMP T2D GWAS; METSIM GWAS) and at the EBI Federated node (EXTEND GWAS; Oxford Biobank exome chip analysis; GoDARTS Affymetrix GWAS).

We have also taken an initial step towards integration of the T2DKP with a new federated node, the T2DREAM database of epigenomic data relevant to T2D. In the near future, epigenomic data displayed in the T2DKP will be drawn dynamically from T2DREAM. In the meantime, we have added gene- and variant-specific links to T2DREAM from the re-styled External Resources section at the bottom of Gene and Variant pages.

New features for all Knowledge Portals

Some of the improvements in this release are visible in all the Portals of the Knowledge Portal Network. One of the most significant affects LocusZoom, the dynamic plot that displays variant associations along with their genomic coordinates, linkage disequilibrium, and other information. Previously, the only way to select a phenotype was to scroll through a long list. Now, a new phenotype filter lets you enter one or more search criteria and filter the list by those criteria. Once you have selected a phenotype, the datasets that include associations for that phenotype are presented for selection. Previously, only one dataset (the one with the largest sample size) was available for each phenotype; now, associations from all relevant datasets may be viewed in LocusZoom.


Portion of the updated LocusZoom interface, showing phenotype filtering capability.


The sample filtering panel of the user interface for the custom burden test and GAIT (Genetic Association Interactive Tool) has also been improved to make it more intuitive to use. The External Resources sections of Gene and Variant pages have been re-styled, and gene- and variant-specific links to PheWeb have been added. PheWeb displays phenotypes most significantly associated with the gene or variant, based on a GWAS for over 2,400 phenotypes in UK Biobank data that was performed by Ben Neale's group. Finally, the home pages of all the Portals have been redesigned to make the appearance of the disease-specific portals more distinct.


Please browse these new data and features, and let us know what you think!

Monday, January 22, 2018

GWAS data re-analysis yields novel results about T2D risk



"Waste not, want not." The old proverb is about frugality, but a study published today gives it a whole new dimension. Lead author Sílvia Bonàs, directed by Josep Mercader and David Torrents and collaborating with many colleagues at the Barcelona Supercomputing Center, the Broad Institute, and other institutions (Bonàs-Guarch et al. (2018), Nature Communications 9), decided to investigate variants associated with type 2 diabetes (T2D) by re-analyzing existing GWAS data rather than initiating a new study.

This was a frugal strategy, conserving both time and resources. But the benefits of this approach went way beyond frugality. By aggregating multiple datasets and using unified, current methods for quality control, imputation, and association analysis, the researchers discovered nuggets of significant information that were not apparent in the original analyses of the individual sets. And all of these nuggets are freely available for browsing and searching in the T2D Knowledge Portal (T2DKP).

To amass these data, the researchers combined all of the individual-level T2D case-control GWAS data that were available from the European Genome-Phenome Archive (EGA) and the database of Genotypes and Phenotypes (dbGaP). After harmonization and quality control, data from 70,127 subjects (12,931 cases and 57,196 controls) remained, inspiring them to name the project "70KforT2D".

In the time since the original studies had been performed, better and more comprehensive reference panels for imputation had been generated by the 1000 Genomes and UK10K projects. By using both of these panels for imputation, the researchers were able to substantially increase the number of variants that could be imputed. They ended up with more than 15 million variants, including more than 5 million rare variants and over 1.3 million indels, which have previously been difficult to impute.

In performing association analysis, the authors took advantage of existing large datasets of T2D association summary statistics for meta-analysis, being careful to only combine non-overlapping samples. They also took advantage of the T2D Knowledge Portal to verify some associations for low-frequency variants that were located in coding regions and had suggestive, but not unambiguously significant, p-values. The significance of the T2D associations of these variants was confirmed by meta-analysis along with the associations seen in two large studies in the T2DKP (GoT2D exome chip analysis, with nearly 80,000 samples, and the 17K exome sequence analysis dataset with 17,000 samples).

The association analysis identified 57 loci associated with T2D risk at the genome-wide significance level or better (p-value ≤ 5x10e-8), seven of which had not previously been associated with T2D. The high quality of the data made it possible to fine-map the variants at each of these loci and construct credible sets. Many of the putative causal variants—including those in previously identified loci—were indels rather than single-nucleotide polymorphisms, underscoring the importance of an imputation procedure that discovers indels.

The T2D-associated loci discovered in this study give some tantalizing hints about genes potentially involved in T2D, and suggest new avenues for detailed wet-lab investigation. We can’t review all of them in this space, but one association is particularly interesting for the generalizable lessons it teaches us about case-control GWAS for T2D.

This association, which the authors validated and replicated using additional datasets, involves the X chromosome variant rs146662075. The risk allele confers a 2-fold elevated risk of developing T2D, in males. The variant appears to affect an enhancer that could regulate expression of AGTR2, a gene known to be involved in modulating insulin sensitivity—making it a very interesting subject for investigation with regard to T2D. More work is needed to figure out whether this is really a male-specific effect, or whether it was only detectable in males because imputation for the X chromosome is more accurate in males, who have only one copy of the chromosome.

The first lesson learned from this association is that the X chromosome harbors important loci, and deserves attention in association studies. While this seems obvious, since the X chromosome comprises 5% of the genome, it has been neglected in most studies to date.

The second lesson is that for an adult-onset disease like T2D, it’s very important to pay attention to the details of case-control classification. If there are young people in the control group, they may actually be future T2D cases, destined to develop the disease later in life. When the authors tried to replicate the initial discovery for this variant in different datasets, the associations were not as significant as expected. But after digging deeper into the experimental cohorts, they found that most of the replication datasets had many subjects younger than 55, which was the average age for T2D onset for these cohorts. Re-running the analysis after excluding controls younger than 55 and also excluding those who appeared to be pre-diabetic, based on an oral glucose tolerance test, brought the replication results into concordance with the discovery results and confirmed the significance of the rs146662075 association.

In keeping with the spirit of open access, the authors provided the summary statistics from this work to the T2DKP even before publication. These results are incorporated into the T2DKP and are visible on Gene and Variant pages as well as searchable via the Variant Finder. The authors have also made the full summary statistics available for public download.

The novel and important findings from this study strongly reaffirm the value of data sharing. Not only are data sharing and re-analysis the right things to do for reasons of fairness, equity, and frugality; they can also spark new insights and move science forward in unexpected ways.

Wednesday, November 15, 2017

T2DKP Fall Newsletter

The latest issue of our quarterly newsletter is now available. Download it here to find out what we've been up to!

Tuesday, November 14, 2017

Announcing the Cardiovascular Disease Knowledge Portal

We are pleased to announce the launch of the Cardiovascular Disease Knowledge Portal (CVDKP). Our collaboration with Dr. Patrick Ellinor, Dr. Sek Kathiresan, and their colleagues in the Atrial Fibrillation, Global Lipids Genetics, Myocardial Infarction Genetics, and CARDIoGRAMPlusC4D consortia has created a resource that offers world-wide open access to genetic and genomic information about atrial fibrillation, myocardial infarction, and related traits, with the goal of democratizing access to genomic data and accelerating cardiovascular genomics research.


CVDKP home page

The CVDKP is constructed on a software architecture originally developed for the Type 2 Diabetes Knowledge Portal (T2DKP), which is the central product of the Accelerating Medicines Partnership in Type 2 Diabetes (AMP T2D). AMP T2D is a public-private partnership between the National Institutes of Health, the U.S. Food and Drug Administration, biopharmaceutical companies, and non-profit organizations that is managed through the Foundation for the NIH. AMP seeks to harness collective capabilities, scale, and resources toward improving current efforts to develop new therapies for complex, heterogeneous diseases.

The ultimate goal of AMP T2D is to increase the number of new diagnostics and therapies for patients while reducing the time and cost of developing them, by jointly identifying and validating promising biological targets for type 2 diabetes. The T2DKP furthers that goal by aggregating, harmonizing, and displaying genetic association and epigenomic results along with user-friendly analysis tools, allowing research biologists who would not individually be able to amass and manipulate these large datasets to glean insights from the data.

We are working towards these same goals for other complex diseases, by extending the platform and analysis tools constructed for the T2DKP. In partnership with the International Stroke Genetics Consortium, we recently created a Knowledge Portal for cerebrovascular disease (CDKP) based on the same infrastructure. Now, with the advent of the Cardiovascular Disease Knowledge Portal, we have a three-member Knowledge Portal Network for the genetics of cardiometabolic and cerebrovascular disease.

Data in the CVDKP directly relevant to heart disease include genetic associations with atrial fibrillation, electrocardiogram traits, plasma lipid levels, and myocardial infarction. Additional association datasets are available for type 2 diabetes and glycemic traits, anthropometric traits, measures of kidney function, and psychiatric traits. You may browse the complete list of datasets and their descriptions on the CVDKP Data page.

As for the Cerebrovascular Disease Knowledge Portal, in the CVDKP we also continue to work with the American Heart Association Precision Medicine Platform (PMP) to provide an additional avenue for accessing cardiovascular genetic data. Currently, summary statistics from the AFGen GWAS and AFGen exome chip analysis datasets are deposited in the PMP.

We welcome all suggestions, comments, questions, and submission of relevant datasets for the CVDKP. Please contact us at help@cvdgenetics.org!

Monday, October 16, 2017

Learn about complex disease knowledge portals at ASHG 2017

Members of the Knowledge Portal team will be attending the American Society of Human Genetics meeting this week in Orlando, FL.

We'll be talking about the continuing progress of the Type 2 Diabetes Knowledge Portal, which has grown dramatically since ASHG 2016, with loads of new data and many new features. We'll also present our work towards expanding the T2DKP framework to other complex diseases, with the recent release of a new sibling portal for stroke genetics, the Cerebrovascular Disease Knowledge Portal.

You can catch us nearly every day of the meeting:

Wednesday 10/18

10 AM - 5 PM: Find us in the exhibit hall at booth #863. We’ll be there to answer your questions and give tours and tutorials on the Knowledge Portal Network.

10-10:30 AM: Demonstration of the Type 2 Diabetes Knowledge Portal at our booth, #863.
10:30-11 AM: Demonstration of the Cerebrovascular Disease Knowledge Portal at our booth, #863.

2 PM - 4 PM: Ben Alexander will present poster #1186: The Type 2 Diabetes Knowledge Portal: Clearing a path from genetic associations to disease biology.

Thursday 10/19

10 AM - 5 PM: We will again be in the exhibit hall at booth #863.

10:30 - 11:30 AM: Portal team members will be available at the Broad Institute booth (#1037) for demonstrations and tutorials.

2-2:30 PM: Demonstration of the Type 2 Diabetes Knowledge Portal at our booth, #863.
2:30-3 PM: Demonstration of the Cerebrovascular Disease Knowledge Portal at our booth, #863.

4:15 PM–6:15 PM: Portal team members will be participating in Concurrent Invited Session #49:

Data Sharing, Analysis, and Tools to Catalyze Translation from Genomic to Clinical Knowledge
Room 330C, Level 3, Convention Center
Moderators: Benjamin Neale and Noël Burtt
Talks:
Serving genetic data and tools to the world - Jason Flannick.
The EGA as a platform for effective data sharing of human genetic and phenotype data -Thomas Keane.
Converting sequence data from over 140,000 people into rare disease diagnoses - Daniel MacArthur.
Assessing the phenome-wide consequences of genetically regulated molecular traits - Hae Kyung Im.

Friday 10/20

10 AM - 2:30 PM: This is our last day in the exhibit hall at booth #863.


We look forward to meeting you at ASHG! If you have questions and cannot meet us any of these times, or if you won’t be at ASHG, our mailbox is always open at help@type2diabetesgenetics.orghelp@type2diabetesgenetics.org.