Showing posts with label Accelerating Medicines Partnership. Show all posts
Showing posts with label Accelerating Medicines Partnership. Show all posts

Friday, June 22, 2018

New data release brings new phenotypes and huge sample sizes to the T2DKP

Progressing towards the goal of the Accelerating Medicines Partnership in Type 2 Diabetes (AMP T2D) to aggregate, analyze, and present comprehensive genetic data relative to T2D in order to speed up the validation of new drug targets, today we release 10 new datasets to the Type 2 Diabetes Knowledge Portal. These datasets contain variant associations for 17 phenotypes, including 7 that are new to the T2DKP, from over 1.4 million samples.

Four of the new datasets were generated by collaborators in AMP T2D, the parent organization of the T2DKP. AMP T2D is a pre-competitive partnership among the National Institutes of Health, industry, and not-for-profit organizations, managed by the Foundation for the National Institutes of Health, that supports the generation of genetic association data and many other kinds of genomic data as well as providing access to these data in the T2DKP, to facilitate the translation of these data into biological knowledge about T2D.

For all four of these datasets, quality control and association analysis were performed by the Analysis Team of the AMP Data Coordinating Center (AMP DCC) at the Broad Institute, using standard, state-of-the-art methods. These processes are completely transparent and fully documented: the experimental design and analysis are summarized on our Data page, and detailed reports are available for download. In this first phase of analysis, associations were determined for type 2 diabetes, fasting glucose levels, and fasting insulin levels--both unadjusted, and adjusted for body mass index. Future analyses will add more phenotypes.

One of these datasets,  Diabetic Cohort - Singapore Prospective Study GWAS, was contributed by collaborators at the National University of Singapore. Consisting of 3,864 samples, it is a T2D case-control study to identify genetic and environmental risk factors for diabetes in Singapore Chinese. The other three new sets that were analyzed at the AMP DCC, contributed by collaborators at the University of Michigan, are from the Finland-United States Investigation of NIDDM Genetics (FUSION) Study that seeks to to map and identify genetic variants that predispose to type 2 diabetes or affect variability in diabetes-related traits. The three FUSION datasets include FUSION GWAS, with 1,681 samples; FUSION Metabochip, with 2,163 samples, and FUSION exome chip analysis, with 3,485 samples.

All four of these datasets now have “Early Access Phase 1” status, which is assigned to new data. This status denotes that although analysis and quality control checks have been performed, the data are not yet considered to be in their final state. During the early access period, users may analyze the data but may not submit the results of these analyses for publication. Find the full details about the different phases of data release on our Policies page.

In addition to the datasets from AMP T2D partners, we have also added or updated 6 new sets of publicly-available association summary statistics for phenotypes relevant to T2D:

  • The previous CKDGen GWAS dataset for chronic kidney disease has been replaced with a newer study from the CKDGen consortium, imputed to the 1000 Genomes reference set (Gorski et al., 2017), with 110,517 samples;
  • Early Growth Genetics Consortium GWAS associations for childhood obesity (Bradfield et al., 2012), with 13,848 samples;
  • Body fat distribution associations (Shungin et al., 2015), with 245,749 samples, have been added to the existing GIANT GWAS dataset;

Results from all the new datasets may be viewed at these locations in the T2D Knowledge Portal:

• On Gene Pages (e.g., GCKR) in the Common variants and High-impact variants tables and in LocusZoom plots;

• On Variant Pages (e.g.rs1260326) in the Associations at a glance section, the Association statistics across traits table, and in LocusZoom static plots;

• From the View full genetic association results for a phenotype search on the home page: select a phenotype and view the top variants in a Manhattan plot and table;

• Using the Variant Finder tool: specify multiple criteria and retrieve the set of variants meeting those criteria from any of these datasets.

Additionally, individual-level data from the Diabetic Cohort - Singapore Prospective Study GWAS and FUSION GWAS datasets are available for secure custom interactive analyses using these tools in the T2DKP:

• Using the Genetic Association Interactive Tool (GAIT) on Variant Pages, you may choose a phenotype for association analysis, choose custom covariates, filter the sample pool by specifying a range of values for one or more phenotypes, then run on-the-fly analysis.

• Dynamic LocusZoom plots on Gene and Variant pages allow you to run association analysis using one or more variants of your choice as covariates, in order to test whether associations are independent.

With today's release, the T2DKP includes genetic associations for 68 phenotypes from a total of 35 datasets. We welcome submissions of new datasets for incorporation into the T2DKP. Find information about collaboration here, and please contact us with questions.

Tuesday, April 17, 2018

Developing a model for collaborative science: a mid-term perspective on the AMP T2D Partnership

In 2011, Dr. Francis Collins, Director of the National Institutes of Health (NIH), met with leaders in biomedical research to discuss a frustrating problem. Continual improvements in molecular biological and genomic techniques were generating an avalanche of data relevant to complex diseases, yet the translation of these data into insights about disease mechanisms and drug targets was unacceptably slow. It was clear that an entirely new paradigm for collaborative research would be needed to speed up the extraction of knowledge from data.

The result of these discussions was the creation of the Accelerating Medicines Partnership (AMP), one branch of which focuses on type 2 diabetes (T2D)—a life-threatening disease that affects hundreds of millions of people worldwide, whose incidence is growing, and whose progression cannot yet be effectively stopped or reversed. AMP T2D, a five-year project, includes the National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK); the pharmaceutical companies Janssen Pharmaceuticals, Eli Lilly and Company, Merck, Pfizer, and Sanofi; the University of Michigan; the University of Oxford; the Broad Institute; and other researchers around the globe. The Foundation for the National Institutes of Health (FNIH) also provides funding and coordination for the project.

Drawing on the strengths of both academia and industry, this public-private partnership brings together all stakeholders in a pre-competitive space to share data and combine resources, with the goal of validating new drug targets faster. Now in Spring 2018, roughly mid-way through the funding period, it is evident that this collaboration has resulted in remarkable progress on both scientific and collaborative fronts.

Genetic association data: the foundation of AMP T2D

Genetic association studies interrogate the genomes of individuals at millions of specific genomic positions to discover sequence variants that are correlated with the incidence of disease. From the outset, AMP T2D aimed to support the generation of unprecedented amounts of new genome-wide association study (GWAS), exome sequencing, and whole-genome sequencing data within the project as well as their aggregation with all relevant publicly available data. Originally, 5 sites were funded by the NIDDK to generate new data and deposit them into the AMP T2D Data Coordinating Center (DCC) at the Broad Institute. As the project evolved, another site was funded by the NIDDK and 8 more sites were funded by the FNIH. Additionally, an Opportunity Pool of funds from the NIDDK was created, allowing the AMP T2D Steering Committee to award smaller grants for complementary research projects in a flexible, science-driven manner.  Currently 10 Opportunity Pool projects are in progress, and more awards will be given in the future.

Not only has the number of genetic association studies increased since the inception of AMP T2D, but also the number of samples surveyed in each has grown dramatically, from typically under 100,000 to approaching 1 million today. The increased statistical power conferred by these large sample sizes has led to a huge increase in the number of loci found to be significantly associated with T2D, from about 70 at the start of the project to nearly 430.

Improvements in genomic technologies in the past few years have allowed AMP T2D collaborators to generate increasing amounts of sequencing data, which make it possible to comprehensively interrogate all alleles and to uncover rare variation. At the project’s start, T2D associations with exome sequences (covering the protein-coding regions of the genome) were available for about 13,000 samples, and no whole-genome sequencing studies had been published. Now, more than 2,600 whole genomes are available, and analysis of a set of 50,000 exomes—the largest disease-specific aggregation of exome sequencing data to date—is nearly complete. Importantly, many of the associations that have been newly discovered in sequencing studies involve relatively rare variants that affect protein-coding regions. It is often more straightforward to develop hypotheses about the impact of such variants than it is for variants outside of coding regions.

As the AMP T2D partnership has grown in prominence in the diabetes field, the DCC has been approached by investigators outside the project who want to contribute their data in order to aggregate and display them in the context of AMP T2D data. In early 2017, researchers in the 70kforT2D project, which found novel T2D associations by re-analyzing existing GWAS data, offered their results for integration into the DCC and display in the Type 2 Diabetes Knowledge Portal (T2DKP; see below) before publication.

70kforT2D GWAS was first pre-publication dataset to be added to the T2DKP from outside the AMP T2D partnership, and it was particularly appropriate that these scientists, whose results illustrate the value of data sharing, themselves chose to freely share their results. Incorporation of datasets into the AMP T2D DCC and T2DKP offers investigators the chance to take advantage of the expertise of the AMP DCC analysis team, apply cutting-edge analysis tools to their data, and display their results broadly to the T2D research community in the context of multiple datasets. The AMP T2D DCC is open to incorporating T2D-relevant datasets from all investigators (find details on contributing data here).

In addition to the datasets generated by AMP T2D partners and other T2D researchers, which focus on associations with T2D, glycemic measures, and T2D complications, the AMP T2D DCC also collects publicly available genetic association datasets for traits relevant to T2D, such as anthropometric measures, blood pressure and lipid levels, and heart and kidney disease.

Orthogonal data types to help identify and prioritize causal variants and genes

Finding genetic variants that are associated with T2D risk is critically important to understanding the genetics of T2D, but it is only a first step. The most significantly associated variant in a genomic region may not be the causal variant that is responsible for altered T2D risk. Researchers perform fine mapping to analyze genetic associations in specific regions of the genome and generate credible sets—that is, sets of variants that are predicted to include the causal variant. Mid-way through the AMP T2D funding period, emphasis among the data-generating partners is beginning to shift from simply generating association data to performing fine mapping and credible set analysis.

But even after predicting which sequence variations are responsible for altered risk, finding clues about how they affect risk requires integration with additional data types. Information about the functional importance of the genomic region where a variant is located—its relevance to gene expression, protein function, networks and pathways, metabolite levels, and more, all determined on a tissue-specific basis—can help prioritize genes and pathways for in-depth experimental investigation. These kinds of research were built into AMP T2D from the beginning, and as the importance of these data types became even clearer, several Opportunity Pool awards were given to projects focusing on complementary data types that shed light on the significance of genetic associations.

Several of these projects focus on generating tissue-specific epigenomic data: histone modifications, DNA methylation, chromatin conformation, transcription factor binding, 3-dimensional chromosome structure, and other data types. Epigenomic data can provide important clues about the mechanisms by which sequence variation affects T2D risk, particularly for variants that lie outside of protein-coding regions. For example, if a risk-associated variant is seen to disrupt a transcription factor binding site, this would support the hypothesis that the transcription factor and its target genes are relevant to T2D.

To make these data accessible to researchers, one Opportunity Pool award supports the creation of the Diabetes Epigenome Atlas, which collects and displays epigenomic datasets relevant to T2D. In the near future, these data will be fully integrated with genetic association data in the Type 2 Diabetes Knowledge Portal (see below).

Other Opportunity Pool projects are concerned with processes downstream of gene expression. Discovering interactions between proteins implicated in T2D risk, for example, could help to uncover all of the players in pathways important for the development of T2D, increasing the number of potential drug targets. Determining the effects of variants on the levels of key metabolites can illuminate the metabolic pathways that change during the development of T2D. 

In addition to generating all of these orthogonal data types, AMP T2D partners are developing algorithms and using machine learning to classify and prioritize variants on the basis of the functional annotations that accompany them. Finally, other Opportunity Pool projects will use model organisms to test and validate drug targets that are suggested by these analyses.

Tools and methods to speed analysis and interpretation

At the inception of AMP T2D it was also clear that the development of new methods and tools would need to accompany the generation of data, and support for these activities was built into the program. One major technical effort has addressed an obstacle to global data aggregation: because of institutional and national privacy regulations, some datasets may not leave their site of origin to be aggregated with other datasets at the AMP T2D DCC. A group at the European Bioinformatics Institute has built a technical replicate of the DCC and knowledgebase, such that data stored there are equally as accessible for browsing, searching, and interactive analysis as are the data stored at the AMP T2D DCC at the Broad Institute. This federation mechanism allows global data accessibility even when data aggregation is not permitted.

Other efforts supported by AMP T2D are aimed at improving the speed and efficiency at which data can be taken in and analyzed. In one project, a data intake system is being developed that will streamline the process for both data submitters and for the DCC team, and will be applicable to data submission both at the Broad DCC and at other federated sites. Another project has created a software pipeline, LoamStream, that will largely automate quality control and association analysis of incoming data. Currently, LoamStream is in use for quality control of genotype data, and this has already greatly reduced the time required to process new datasets. Future work will extend the pipeline to association analysis and will also allow it to take in sequence data as well as genotype data.

A genetic association of a variant with T2D gains credibility if multiple independent studies replicate the association. Thus, it is important for researchers to be able to evaluate the weight of available evidence. But currently this is difficult to assess from the association datasets in the AMP T2D DCC, because many are based on overlapping sets of subjects. AMP T2D partners at the University of Michigan and University of Oxford are working on a method to take these overlaps into account and synthesize associations from multiple datasets into a “bottom-line” significance for association of a variant with T2D, which will aid in prioritizing variants for future work.

Multiple AMP T2D projects for analysis, interpretation, and custom interactive analysis of variant-phenotype associations are ongoing at the Universities of Michigan, Chicago, and Oxford, Vanderbilt University, and the Broad Institute. These projects are aimed at facilitating, in various ways, the path from variant associations to functional knowledge, and all have been or will be integrated into the T2D Knowledge Portal (see below).

Hail software offers a pipeline that speeds up the analysis of huge genomic datasets, while the gnomAD resource aggregates and harmonizes exome and genome sequences to provide a catalog of genetic diversity, in more than 100,000 humans, that aids in interpretation of variant associations with disease. A tool under development in the gnomAD project will display the effects of variants on protein structures as another way to deduce their potential impact.

Other analysis modules include gene-based association methods for using expression data to predict genes that may impact a phenotype (PrediXcan and MetaXcan), and a phenome-wide association study (PheWAS) method for visualization of the associations of a variant across multiple phenotypes, which is a crucial consideration during drug development. 

The interactive visualization tool LocusZoom will integrate many of these methods to display variant associations and credible sets, epigenomic and functional annotations, and phenotype associations across a genomic region as well as offering custom association analysis.

An example LocusZoom plot


AMP T2D Knowledge Portal: democratizing T2D genetic results for researchers world-wide

AMP T2D was founded on the idea that in order to truly accelerate progress, genomic information must be freely accessible to all scientists and presented in a way that is understandable by a broad range of researchers working on T2D biology, not only by human geneticists and bioinformaticians with special computational skills. So the roadmap for the project included not only data generation and analysis, but also the production of a publicly available web resource that would integrate data types, interpret the evidence, and present of all these results. 

While it is under continuous development, mid-way through the initial funding period the T2D Knowledge Portal (T2DKP) is already a well-established resource. Other web resources collect genetic association data, but the T2DKP is unusual in providing harmonized datasets to which a consistent analysis pipeline has been applied. Rather than simply cataloging datasets, it offers distilled and synthesized results along with their interpretation, to guide more detailed exploration of the evidence. And, unlike any other extant resource, it offers researchers the ability to perform interactive queries on protected individual-level data. 

T2DKP home page

The Gene page of the T2DKP (see an example) illustrates the presentation of immediately understandable summary information along with the opportunity to drill down to the details. An algorithm considers the associations of all variants across a gene, for all phenotypes and in all datasets aggregated at the DCC, and calculates from them a “traffic light” signal for the gene: green to indicate that there is a significant association for at least one phenotype; yellow to indicate suggestive, if not highly significant, associations; and red to indicate that there is no evidence for association for any of the phenotypes considered in the T2DKP. Below this, tables and graphics invite users to explore all variants across the gene, their impacts on the encoded protein, and their associations, as well as their positions relative to epigenomic marks across the region in multiple tissues.

The T2DKP currently offers the ability to run custom, interactive association analyses using two different tools. In the LocusZoom visualization, users may choose one or more variants as covariates before performing association analysis. The Genetic Association Interactive Tool (GAIT) for single variant associations, which also powers the custom burden test for gene-level associations, is even more versatile, presenting the distributions of different characteristics of the sample set (age, sex, BMI, glycemic measures, blood lipid levels, and many more) and allowing users to filter the set by multiple criteria and to choose custom covariates before performing association analysis. Both of these tools allow analytical access to the individual-level data, whether housed at the Broad DCC or at the EBI federated node, in a secure environment so that data privacy is always protected.

Evolution of a collaborative environment

AMP T2D organization


The AMP T2D partnership is a multifaceted project (illustrated above) that embraces several aspects of basic research and combines them with building a product, the T2DKP. In connecting scientists both within and outside of consortia, in academia and in industry, working on genetic associations or functional studies, it is becoming the nexus of the T2D genetics community. Researchers are finding the T2DKP helpful for accessing even their own results and for viewing them in the context of multiple phenotypic associations and other complementary data types. Pharmaceutical partners are finding help via the Target Prioritization project, in which the tools and methods developed within AMP T2D are being used to prioritize a list of genes of mutual interest for further investigation.

Perhaps most importantly, AMP T2D has made researchers—both within and outside of the project—aware of the value of sharing data for representation in the context of all other relevant data. Only by compiling and interpreting all available information will we be able to make the best hypotheses about genes and pathways that are possible drug targets and prioritize them for in-depth functional investigation.

AMP T2D and beyond

In the remainder of the initial AMP T2D funding period, we expect continued progress in each of the areas discussed above. The data intake and analysis pipelines will be improved, and new data will be incorporated at an increasing pace—including data from the UK Biobank, which has generated association results for 500,000 genotyped subjects and more than 2,500 traits. Associations will be added for many more phenotypes related to T2D, including diabetic complications and longitudinal phenotype data that connect the development of various traits to the timeline of incident T2D.  Much more T2D-relevant epigenomic data will be available for query as well as for browsing, via dynamic connection with the Diabetes Epigenome Atlas. And entirely new data types (for example, metabolomic and proteomic data) arising from Opportunity Pool projects will be added to the T2DKP.

Ongoing work on tools and methods will result in the addition of many more interactive modules to the T2DKP. Researchers will be able to view PheWAS data; prune lists of variants by their linkage disequilibrium relationships; calculate credible sets and genetic risk scores with custom parameters; perform more versatile interactive burden tests; prioritize genes by pre-calculated association scores; overlay the positions of coding variants on protein structures to help assess their impact; and perform enrichment analysis on sets of loci to suggest pathways implicated in disease processes.

The Knowledge Portal platform developed for AMP T2D has already proved extensible to other complex diseases: in 2017, both the Cerebrovascular Disease and Cardiovascular Disease Knowledge Portals were launched. In the future, connections within the ecosystem formed by the T2D, Cerebrovascular, and Cardiovascular Portals will be improved, so that researchers can easily assess the impact of a variant or involvement of a gene for all of these related diseases. If funding and collaboration considerations allow, perhaps one day these Portals will merge into a single cardiometabolic disease genetics Knowledge Portal to accelerate the development of new therapeutics in this broader area.

Finally, the ultimate goal of this funding period is that by its end, the data generation, analysis, and interpretation will have facilitated the validation of multiple promising drug targets for further investigation. Given the rate of progress on multiple fronts, this seems a realistic goal. We hope that this unique collaborative environment will continue to accelerate T2D genetic research and will become a paradigm for other research communities.

Tuesday, February 6, 2018

Federation brings three new datasets to the T2DKP

Our mission at the Type 2 Diabetes Knowledge Portal (T2DKP) is to aggregate and analyze genetic association data relevant to T2D, and to make the knowledge that can be gleaned from these data available to researchers around the world. But it isn't possible to aggregate all of the relevant data in one place: privacy regulations at the institutional, regional, and national levels determine how these data are handled, and whether or where they can be transferred.

The T2DKP is supported by the Accelerating Medicines Partnership in Type 2 Diabetes (AMP T2D),  a pre-competitive partnership among the National Institutes of Health, industry, and not-for-profit organizations, managed by the Foundation for the National Institutes of Health. Because AMP T2D seeks to facilitate discovery of new targets for T2D treatment by making as much data as possible available via the T2DKP, it funded the development of a mechanism for establishing interconnected federated nodes of the T2DKP that would enable researchers to interact with all of the data regardless of where they are located.

This goal was realized with the creation, by a team led by Thomas Keane and Dylan Spalding, of a federated node of the T2DKP at the European Bioinformatics Institute (EBI).  Data housed at the EBI node are stored in such a way that their specific privacy requirements are met, but they are made available for remote queries via T2DKP tools and interfaces. Results from such queries are served up alongside results from all of the datasets housed in the AMP T2D Data Coordinating Center (DCC) at the Broad Institute. Researchers may browse and query data from any location without even needing to know where they reside. This federation mechanism represents both an important technical advance in handling and protecting data, and a significant step forward in democratizing and improving access to genetic association results.

The first dataset to be incorporated into the Portal via the EBI federated node was the Oxford BioBank exome chip analysis dataset, which contains association data for glycemic, lipid, and blood pressure traits from over 7,100 subjects in Oxfordshire, U.K. The EBI Federated Node has now added three more datasets:

  • The EXTEND GWAS dataset, generated by Drs. Timothy Frayling and Andrew Wood and their colleagues, is comprised of 7,159 samples (1,395 T2D cases and 5,764 controls) from the Exeter EXTEND Biobank. It includes associations for a wealth of glycemic, anthropometric, cardiovascular, renal, and hepatic phenotypes--including many that are new to the T2DKP.
  • The GoDARTS Affymetrix GWAS dataset, from Dr. Colin Palmer and colleagues, includes summary-level statistics for associations with BMI and blood lipid levels from 3,307 diabetic participants in the Genetics of Diabetes Audit and Research Study in Tayside Scotland. In addition, individual-level data from over 17,000 subjects (including the set from which summary statistics were calculated) are available via the GAIT tool (see below). 
  • The Oxford BioBank Axiom GWAS dataset, from Dr. Fredrik Karpe and colleagues, includes associations for BMI and blood lipid levels from 7,193 participants, all healthy men and women between 30 and 50 years of age. It represents an additional analysis of the same samples contained in the Oxford BioBank exome chip analysis dataset.
These datasets are described in detail on our Data page. Summary results from all three sets are integrated into Gene and Variant pages in the T2DKP, and may also be viewed in the Manhattan plots accessible by searching for a phenotype from the T2DKP home page. The Variant Finder also queries these datasets.

The individual-level data behind all three of these datasets is accessible for custom association analysis in our Genetic Association Interactive Tool (GAIT) on Variant pages. Using this tool, researchers can filter samples to create a custom subset with defined characteristics such as age, gender, BMI, and other measures, and then run on-the-fly association analysis within that sample subset. Now, GAIT queries datasets both at the DCC and at the Federated node, using the same methodology for each, in a way that is transparent to users of the tool. The new Federated datasets bring the total number of individual-level samples available for custom analysis in the T2DKP to 67,768.

Friday, June 9, 2017

Providing data access, ensuring data protection

Readers of this post probably don’t need to be convinced that genetic association data have enormous potential for helping us to understand and treat complex diseases like type 2 diabetes. Significant associations between variants and diseases can suggest genes, or regions of the genome, that could be important for disease risk or progression—and this knowledge could help us identify new drug targets.

The Accelerating Medicines Partnership in Type 2 Diabetes (AMP T2D) is a pre-competitive partnership among the National Institutes of Health, industry and not-for-profit organizations, which is managed by the Foundation for the National Institutes of Health. Its mission is to make genetic association data accessible to the worldwide biomedical research community via the Type 2 Diabetes Knowledge Portal, in order to facilitate discovery of new targets for T2D treatment. But it can be a challenge to aggregate genetic data. The privacy of the individuals who contributed their health status and genomic sequences must always be protected, and there are many layers of regulation to ensure this. Restrictions at the institutional, regional, and national levels determine how data are handled and whether they can be transferred.

Until now, all of the results displayed in the Portal have been derived from data housed at the AMP T2D Data Coordinating Center (DCC) at the Broad Institute, where the Portal website resides. But some of the valuable data generated outside the U.S. cannot be transferred to the DCC. To address this issue, AMP T2D funded the development of a mechanism that enables researchers to interact with all of the data: federation. 

Federation means that data are housed at a site (a “federated node”) that meets their specific privacy requirements, but are made available for remote queries via the Portal. Results from such queries are served up alongside results from all of the datasets housed in the AMP T2D DCC. Researchers may browse and query data from any location without even needing to know where they reside.

A federated node has now been created at the European Bioinformatics Institute (EBI) and may be accessed via the T2D Knowledge Portal. Today, Portal tools and interfaces can query both data housed at the AMP T2D DCC at the Broad Institute and data at the EBI federated node. 

According to Paul Flicek, a Senior Scientist and Team Leader of Vertebrate Genomics at EMBL-EBI, “A key mission of EMBL-EBI is to make data available to the widest possible community. Seamlessly accessing stored in multiple locations via a single portal helps ensure that the data we store from many projects are maximally useful for additional research.”

The first dataset to be incorporated into the Portal via the EBI federated node is the Oxford BioBank exome chip analysis dataset, which contains association data for glycemic, lipid, and blood pressure traits from over 7,100 healthy subjects in Oxfordshire, U.K. The dataset is described on our Data page. Portal users can interact with this dataset in the same way (and with the same speed) as with other datasets. 

“Diabetes is a global problem, and it will take research and innovation on a global scale if we are to tackle it effectively,” says Mark McCarthy, Robert Turner Professor of Diabetic Medicine at University of Oxford. “The success of our research on the genetics of diabetes depends on access to data generated by groups around the world. The federated portal provides an additional set of tools that will allow us to jointly analyse those data sets wherever they happen to be based.” 

Federation represents both an important technical advance in handling and protecting data, and a significant step forward in democratizing and improving access to genetic association results. And because it is generally applicable to any kind of genetic association data, it has the potential to have an impact beyond T2D research, facilitating the study of other complex diseases and traits.

Wednesday, May 31, 2017

See you in San Diego!

Members of the T2D Knowledge Portal team are gearing up for the 77th Scientific Sessions of the American Diabetes Association, June 9-13 in San Diego, CA. We'll be releasing exciting new features of the Portal just before the conference, and we have a wide variety of presentations planned for each day.

On the opening day of the conference (Friday, June 9), join us for a mini-symposium that will present a comprehensive guide to the T2D Knowledge Portal and how you can use it to further your type 2 diabetes research. We will be exhibiting at booth #2452 on Saturday, Sunday, and Monday, and each day, genetics experts will be available at the booth to answer questions and discuss both the Portal and the genetics of T2D. On Saturday, members of the Portal team will participate in a moderated poster session, and posters will also be displayed on Monday. And on Sunday morning, our principal investigator, Dr. Jose C. Florez, will give a symposium presentation on "Mining the Genome for Therapeutic Targets."

Find the full details in the schedule below and follow us on Twitter (@T2DKP) for up-to-the-minute news throughout the conference. We're looking forward to meeting you!

Friday, June 9, 2017

Mini-Symposium: A Researcher’s Guide to Exploring Diabetes Genetic Data in the Type 2 Diabetes Knowledge Portal
Chair: Mark McCarthy
11:30am - 12:30pm, Room 28

11:30-11:50am        NoĂ«l Burtt: Data, Analysis, and Tools in the Type 2 Diabetes Knowledge Portal
11:50am-12:10pm   Jason Flannick: Demonstration of Questions that Can Be Addressed Using the Portal

12:10-12:30pm       Question and Discussion Period


Saturday, June 10, 2017

  • Exhibiting at booth #2452, 10am - 4pm
  • Moderated Poster Session: Genetic Data, Pathways, and Variants for Type 2 Diabetes and Related Traits. 12:30-1:30pm, Hall B
Poster 1765-P
The Type 2 Diabetes Knowledge Portal: Accelerating Type 2 Diabetes Research through Community Access to Human Genetic Information and Tools
Presenter: Maria C. Costanzo

Poster 1766-P
Key Biological Pathways for Type 2 Diabetes Determined by Genetic Cluster Analysis on Related Traits
Presenter: Miriam S. Udler


Sunday, June 11, 2017

  • Symposium presentation: Mining the Genome for Therapeutic Targets. 
Dr. Jose C. Florez
9:20-9:55am, Ballroom 20D

  • Exhibiting at booth #2452, 10am - 4pm


Monday, June 12, 2017

  • Exhibiting at booth #2452, 10am - 2pm
  • Poster session, 12-1pm, Hall B
Poster 1765-P
The Type 2 Diabetes Knowledge Portal: Accelerating Type 2 Diabetes Research through Community Access to Human Genetic Information and Tools
Presenter: Maria C. Costanzo

Poster 1795-P
Type 2 Diabetes Gene Bioinformatically Identified by Variants Mapping to Amino-Acid Changes in Three-Dimensional Protein Space
Presenter: Marcin von Grotthuss

Wednesday, May 3, 2017

Explore new datasets and phenotypes in the T2D Knowledge Portal

We are releasing multiple new datasets in the Portal and have updated existing sets with associations for new phenotypes. To make it even easier to browse and explore these sets, we've also updated our Data page and added new functionality. Here's an overview of what's new in the Portal today.

17K exome sequence analysis dataset has grown to 19K
The Data Coordinating Center (DCC) of the Accelerating Medicines Partnership in Type 2 Diabetes (AMP T2D) analyzes exome sequence data contributed by AMP T2D consortium members to find variant associations with T2D and related traits. The exome sequencing dataset available in the Portal has until now consisted of exome sequences from about 17,000 individuals. Today, we have added exome sequencing performed on 2,000 Danish subjects by the LuCamp (Lubeck Foundation Centre for Applied Medical Genomics in Personalised Disease Prediction, Prevention and Care) consortium, making a total of nearly 19,000 exomes. This is just a taste of things to come: at the AMP T2D DCC we are currently analyzing additional exome sequences that will bring the total up to 52,000!

New community-contributed datasets: GENESIS GWAS and 70KforT2D GWAS

We are grateful to two groups from the larger T2D research community who have shared data that will make the T2D Knowledge Portal even more valuable to worldwide T2D researchers.

The GENEticS of Insulin Sensitivity (GENESIS) consortium performed GWAS on over 2,700 nondiabetic participants, finding genetic associations with direct measures of insulin sensitivity.

The 70KforT2D project collected, harmonized, and re-analyzed public GWAS data from over 70,000 individuals to find T2D genetic associations.

New public dataset: VATGen GWAS

The VATGen GWAS consortium performed meta-analysis of GWAS data from a mixed-ancestry group of more than 18,000 people to identify genetic associations with the localization of body fat deposition, leading to insights into adipocyte development.

Updated dataset: glucose-stimulated insulin secretion phenotypes in MAGIC GWAS

A study by Prokopenko et al. analyzed genetic associations with insulin secretion. Associations of variants with nine different measures of insulin secretion, among them corrected insulin response (CIR) and disposition index (DI), have now been added to the MAGIC GWAS dataset.

ExAC updated to gnomAD exomes and whole genomes
The Exome Aggregation Consortium (ExAC) has more than doubled in size and has morphed into the Genome Aggregation Database (gnomAD). More than 120,000 exome sequences and 15,000 whole genome sequences are now available, and these data are accessible via several tools and interfaces in the T2D Knowledge Portal.

New Data page: explore datasets using new filters

As our collection of data grows, it becomes more difficult to understand the differences between datasets and to find those of interest. To address this challenge, we've reorganized and streamlined our Data page.


A section of the Data page, expanded to show phenotype selection.

At the top of the Data page, you can choose to filter the dataset table by data type, phenotype category, or both. When you click on a phenotype category, the phenotypes within that category are available for selection. Clicking on the name of any dataset expands a section with details and references for each. 

In the coming days, watch this space for more details about each of these new developments. And as always, please contact us if you have any comments or questions.


Tuesday, January 17, 2017

New Year, New Data: BioMe AMP T2D GWAS

We’re happy to announce the first addition of data to the Type 2 Diabetes Knowledge Portal in 2017: the BioMe AMP T2D GWAS data set. The generation of these data was funded by the Accelerating Medicines Partnership in Type 2 Diabetes (AMP T2D), a collaboration between multiple stakeholders that aims to catalyze the clinical translation of genetic discoveries by producing and aggregating data, developing and implementing novel analytical methods and tools, and building infrastructure for data storage and presentation.

The BioMe AMP T2D GWAS data set is the first set to be entirely produced by the AMP T2D project, which supplied the funding and carried out every step of its production, from data generation to analysis, quality control, and presentation. Its immediate availability in the Portal, prior to publication, fulfills the mission of AMP T2D to speed up access to and utilization of new data.

These data were generated at the Charles Bronfman Institute for Personalized Medicine BioMe BioBank, a biorepository located at the Mount Sinai Medical Center (MSMC) in the upper Manhattan area of New York City. MSMC serves a diverse population of over 800,000 outpatients each year. Importantly, since many BioMe participants are African American or Hispanic Latino, this data set adds significant ethnic diversity to the Portal’s genetic association data.

The BioMe AMP T2D GWAS data set is comprised of about 13,000 unique individuals, 41.5% of whom are admixed American, 38% African American, and 20% European. Subjects were genotyped using at least one of three platforms: the Illumina Exome Array, the Illumina GWAS array, or the Affymetrix GWAS array. Their T2D status was assessed by an algorithm, and many additional traits were also measured.

The data were subjected to quality control and association analysis by the Analysis Team at the AMP Data Coordinating Center (DCC) at the Broad Institute. Variant associations with T2D, fasting glucose levels, and HbA1c levels were analyzed. The top results included both previously known and novel variants, with only a single variant reaching genome-wide significance: T2D association of the variant rs7903146, within the well-established T2D risk gene TCF7L2. Now that these results are available in the T2D Knowledge Portal, the ability to analyze them further in the context of all other available T2D association data may lead to additional insights.

The BioMe AMP T2D GWAS data currently has the “Early Access Phase 1” status that is assigned to new data. This status denotes that although analysis and quality control checks have been performed, the data are not yet considered to be in their final state. During the early access period, users may analyze the data but may not submit the results of these analyses for publication. Find the full details about the different phases of data release on our Policies page. More information about the data set, along with links to download even more detailed reports on its quality control and analysis, may be found in the BioMe AMP T2D GWAS section of our Data page.

BioMe AMP T2D GWAS data are available at these locations in the Portal:

  • On Gene Pages (see an example) in the Variants & Associations table and the Minor allele frequencies across data sets table
  • On Variant Pages  (see an example) in the Associations at a glance section and in the Association statistics across traits table
  • Via the Variant Finder tool, for these phenotypes: type 2 diabetes; fasting glucose adjusted for age and sex; HbA1c adjusted for age and sex; and HbA1c adjusted for age, sex, and body mass index
  • A "Manhattan plot" of associations across the genome may be seen by selecting one of the phenotypes above in the View full genetic association results for a phenotype scroll box on the Portal home page, and then selecting the BioMe AMP T2D GWAS data set.

As always, please contact us with any questions, comments, or suggestions.

Friday, October 14, 2016

See you at ASHG 2016!

Members of the Type 2 Diabetes Knowledge Portal team will be attending the American Society of Human Genetics meeting next week in Vancouver, BC. You can catch us nearly every day of the meeting:

Tuesday 10/18

3 PM: Nöel Burtt will be one of the speakers in an informational session on the T2D Knowledge Portal and new funding opportunities offered by the Foundation for the NIH. Complimentary snacks, beer, and wine will be served! Please pre-register here.

Wednesday 10/19

10 AM - 4 PM: Find us in the exhibit hall at booth #428. We’ll be there to answer your questions and give tours and tutorials on the Portal.

Thursday 10/20

10 AM - 4 PM: We will again be in the exhibit hall at booth #428.

2 - 3 PM: Ryan Koesterer will present his poster on an automatic, scaleable quality control method for genetic association data that improves on current “gold-standard” methods (program #1943T).

2 - 3 PM: Maria Costanzo will present her poster giving an overview of data in the Portal and the global collaborative efforts behind its aggregation (program #329T).

Friday 10/21

10 AM - 4 PM: This is our last day in the exhibit hall at booth #428.

2 - 3 PM: Marcin von Grotthuss will present his poster on improving predictions of significant variants by taking protein structure into account (program #489F).

3 - 4 PM: Ben Alexander will present his poster on the software platform that powers the T2D Knowledge Portal user interface and custom analysis tools (program #1650F).

T2D Knowledge Portal staff attending ASHG

We look forward to meeting you at ASHG! If you have questions and cannot meet us any of these times, or if you won’t be at ASHG, our mailbox is always open at help@type2diabetesgenetics.org.

Monday, October 3, 2016

Come to a T2D Knowledge Portal information session at ASHG

The American Society of Human Genetics meeting is happening in Vancouver, B.C. in a little over two weeks! The Portal team will be presenting and exhibiting at multiple venues at ASHG, and the first event will take place immediately before the conference starts: an information session including an overview of the Accelerating Medicines Partnership in Type 2 Diabetes, a progress update on T2D Knowledge Portal functionality, and information on new funding opportunities. Complimentary hors d'oeuvres, beer and wine will be served!

Information session
Tuesday - October 18, 2016
3:00 pm - 4:00 pm PDT
Fairmont Waterfront

900 Canada Place Way

Vancouver, British Columbia


Please register here for this free event, hosted by FNIH.  Contact Nicole Spear at Nspear@fnih.org with any questions.

Watch this space over the next two weeks for a complete listing of opportunities to learn about the Portal and talk with the Portal team at ASHG!




Monday, June 20, 2016

Report from New Orleans: 76th Scientific Sessions of the American Diabetes Association


Members of the T2D Knowledge Portal team braved extreme heat and humidity, as well as icy air conditioning, to attend the American Diabetes Association conference in New Orleans, LA. Our booth in the conference exhibit hall was a great way to interact personally with conference attendees and showcase the Portal. Many genetics researchers stopped by for one-on-one tutorials on our new tools and features. And clinicians and diabetes patients, even if they had no immediate use for genetic information, were happy to hear the goals of the project—to accelerate the identification of genes involved in T2D and, ultimately, to find new treatments and better understand the disease mechanism. 

We were pleased to welcome some special visitors to our booth: National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK) Director Griffin Rodgers and Deputy Division Director Philip Smith. NIDDK is a major supporter of the T2D Knowledge Portal project.

Drs. Philip Smith (left) and Griffin Rodgers visit the Portal booth


Dr. Smith also made an video statement as part of the media coverage at ADA, eloquently explaining the rationale behind the Portal and the needs that it can address.





If you missed us at ADA, come visit us at our booth at the American Society for Human Genetics meeting next October! And if you can’t meet us in person, please feel free to email us at any time. We’re happy to answer questions or provide help in understanding the Portal data and tools.

Type 2 Diabetes Knowledge Portal team at ADA

Wednesday, October 7, 2015

Type 2 Diabetes knowledge portal goes live

Human genetic data contains valuable clues for solving the mystery of how human disease works, and the amount of data is growing as major projects around the world engage patients in large genomic studies. These projects would be most powerful if researchers worldwide collaborated with each other, sharing their findings and making results available to scientists from many disciplines including molecular biology and drug development. However, there are few frameworks for sharing the data in a way that properly credits researchers and protects patient privacy -- and as a result, today, only a relative handful of highly specialized geneticists can analyze the data or access the full results of studies.

What if anyone could?

Making that kind of sharing possible is the ultimate goal of a user-friendly web portal and knowledgebase launched October 7 at a special session of the American Society of Human Genetics annual meeting, by an international team that hopes to dramatically expand the number of researchers who can use human genetic data to study type 2 diabetes (T2D).

Research on other common diseases faces the same challenges and could therefore benefit from similar web portals.

“The investigators involved in the Accelerating Medicines Partnership in Type 2 Diabetes believe that genomic data should benefit all of humankind, as long as proper confidentiality protections are in place,” said Jose Florez, Chief of the Diabetes Unit and an Institute Member at the Broad Institute, who is one of the lead scientists in the project. “In this way the collective brainpower of all those with bright ideas, whether they come from academia, big Pharma, the biotechnology sector or government can be brought to bear on these difficult problems.”

The T2D portal (www.type2diabetesgenetics.org) -- which is open to the world free of charge -- is being developed by a team of scientists and software engineers at the Broad Institute, the University of Michigan, Oxford University, and many other collaborators as part of a worldwide scientific consortium with contributors from academia, industry, and non-profit organizations. Financial support is provided by the Accelerating Medicines Partnership in Type 2 Diabetes -- a collaboration of the National Institutes of Health, five major pharmaceutical companies, and three large non-profits -- as well as the Carlos Slim Foundation.

“Through AMP, we have an unprecedented opportunity to advance international research in type 2 diabetes,” said NIH Director Francis S. Collins in a press release. “Our hope is that this portal – and this partnership – will lead to better disease targets and a shorter, less expensive drug development process, enabling companies to get safe and effective medications to patients who need them faster.”

The current knowledgebase contains genetic and clinical data from dozens of academic institutions and partners worldwide, and researchers hope to add greatly to it over the next few years. The data already include detailed information from more than a hundred thousand participants in studies of T2D, as well as results from some of the largest genetic studies ever conducted on related traits such as obesity and glucose and insulin levels. Some of the datasets were mined by collaborating investigators in recent years to identify the genetic variants that influence T2D risk in new or unexpected ways, such as a set of variants near the gene SLC16A11 that are present in almost half of Mexicans with Native American ancestry, and another set of variants that inactivate the gene SLC30A8 and protect variant carriers against T2D.

The studies represented in the portal -- and those that researchers hope to add -- use different technologies and formats. As a result, they must be not only aggregated but also "harmonized" -- rendered compatible with each other. Those two challenges represent major work needed to make publicly accessible knowledgebases a reality for T2D or other complex traits, said Jason Flannick, a Research Associate at Massachusetts General Hospital and Technical Lead at the Broad Institute, who leads the Broad portal development team.

“There is a wealth of genetic data that has been and continues to be generated by investigators around the world that could be hugely valuable to many of the people focused on downstream biological or therapeutic research,” said Flannick. “But enabling these heterogeneous and scattered datasets to be accessible via one place, and building a sustainable model that will scale to the much larger datasets that are being produced as next-generation sequencing becomes more and more routine, requires a major investment in software and bioinformatics.”

Scientists can use the portal to easily browse results from many studies simultaneously, a process that has previously required specialized software and analysts. In addition, the portal is designed to anticipate and answer biological questions from users who may not already be familiar with statistics or genomics, enabling a broader community to learn from these results for the first time.

Users can:
  • Retrieve all genetic variants identified within a gene of interest -- learning which types of variants are present in different populations, what molecular effects are predicted for them, and how strongly those variants have been associated with T2D or related conditions in human studies.
  • Explore data for an entire chromosomal region, visualizing the effects across datasets of hundreds of variants via easy-to-use modules including a custom, web-only version of the Broad's Integrative Genomics Viewer.
  • Identify sets of variants that are associated with different combinations of traits, using a query builder that can construct custom searches across multiple datasets.
  • Investigate whether deactivating a gene in humans is likely to increase or decrease T2D risk by using an “on-the-fly” custom analysis feature that has recently been released for beta-testing.

At no point are users allowed to see protected genetic data that could be used to identify research participants -- they see only the results of analyses.

Over the coming years, the team hopes to contribute to a federated network of portals that would house many more datasets representing other studies, phenotypes, and data types -- for instance, results from large studies of gene expression and function across different human cell types and tissues, such as GTEx and ENCODE. Collaborators, including Daniel MacArthur and Benjamin Neale at the Broad as well a dedicated team led by Michael Boehnke and Gonçalo Abecasis at the University of Michigan, are already developing new methods and software tools destined for the knowledgebase.

“Combining the most extensive T2D genetic data possible with cutting-edge tools for data analysis, reporting, and visualization will accelerate genetic discovery and enable scientists to ask and answer the questions needed to identify new targets for treatment,” said Boehnke.

Ultimately, the foundation of the portal is the data; it will need collaborators from many countries and projects to succeed, said Mark McCarthy from Oxford University. Already data from over a dozen countries are in the portal. “The aspiration is to engage with researchers around the world," he said. "We will continue to reach out so that the portal contains data that are relevant to researchers and patients everywhere."