Showing posts with label new data. Show all posts
Showing posts with label new data. Show all posts

Monday, October 8, 2018

DIAMANTE GWAS dataset adds close to a million samples along with fine-mapping to the T2DKP

In a groundbreaking paper published today, Anubha Mahajan and colleagues (Mahajan et al., Nature Genetics 2018) report on a meta-analysis of unprecedented size for genetic associations with type 2 diabetes (T2D) along with fine-mapping analyses to identify causal variants that can suggest new therapeutic targets. We are pleased to provide access to the summary results as well as the results of the fine-mapping today in the T2D Knowledge Portal (T2DKP).

Working as part of the DIAGRAM (DIAbetes Genetics Replication And Meta-analysis) and DIAMANTE (DIAbetes Meta-ANalysis of Trans-Ethnic association studies) consortia, the researchers aggregated and meta-analyzed genome-wide association studies for about 900,000 individuals of European ancestry (about 74,000 T2D cases and 824,000 controls). The studies were imputed using the most comprehensive reference panels possible, and in all, the analysis considered about 27 million genotyped or imputed variants.

After performing T2D association analysis (both unadjusted and adjusted for body mass index) 243 loci were seen to be associated with T2D at genome-wide significance or better (p-value for association ≤ 5 x 10-8). Of these, 135 were novel--not detected previously in any T2D association analysis to date.

Within these loci, each of which included multiple significantly associated variants, the researchers performed approximate conditional analysis to determine whether the associations were independent of each other. They found surprising complexity within some loci; for example, the well-known TCF7L2 locus appears to include as many as 8 distinct association signals!

All of the T2D associations from this study may be viewed in the T2DKP. They are represented in two datasets, named "DIAMANTE (European) T2D GWAS" and "UK Biobank T2D GWAS (DIAMANTE-Europeans Sept 2018)."  Manhattan plots showing the distribution of the associations across the genome may be seen by selecting either the "Type 2 diabetes" or "Type 2 diabetes adj BMI" phenotypes from the phenotype selection menu on the T2DKP home page. On Gene pages of the T2DKP, the results may be viewed in tables of variant associations and in the interactive LocusZoom visualization (see below). Results from this study are also displayed on Variant pages of the T2DKP.


LocusZoom plot on the PPARG Gene page


The credible set analysis performed in this study is also incorporated into the T2DKP. On the "Credible sets" tab of Gene pages, you may choose to visualize any of the credible sets available for the region. Epigenomic annotations that overlap the positions of the variants in the credible set are presented in an interactive display that allows you to select particular chromatin states or tissues to view. In the example shown below, one of the credible sets in the TCF7L2 region includes just two variants, and the one with the highest posterior probability overlaps active enhancer regions in adipose and liver tissue--both of which are important for T2D.


Detail of the Credible sets tab of the TCF7L2 Gene page

The multiple causal variants identified in this study support previous investigations on the biological mechanisms behind T2D and suggest new hypotheses that will likely lead to therapeutic insights. After reading the paper and a blog post from the authors, we invite you to explore the results in the T2DKP and to contact us with any suggestions or questions!

Friday, May 11, 2018

T2DKP Spring Newsletter

The latest issue of our quarterly newsletter is now available. Download it here and get the latest!

Tuesday, March 6, 2018

T2DKP Winter Newsletter

The latest issue of our quarterly newsletter is now available. Download it here to find out what we've been up to!

Tuesday, February 6, 2018

Federation brings three new datasets to the T2DKP

Our mission at the Type 2 Diabetes Knowledge Portal (T2DKP) is to aggregate and analyze genetic association data relevant to T2D, and to make the knowledge that can be gleaned from these data available to researchers around the world. But it isn't possible to aggregate all of the relevant data in one place: privacy regulations at the institutional, regional, and national levels determine how these data are handled, and whether or where they can be transferred.

The T2DKP is supported by the Accelerating Medicines Partnership in Type 2 Diabetes (AMP T2D),  a pre-competitive partnership among the National Institutes of Health, industry, and not-for-profit organizations, managed by the Foundation for the National Institutes of Health. Because AMP T2D seeks to facilitate discovery of new targets for T2D treatment by making as much data as possible available via the T2DKP, it funded the development of a mechanism for establishing interconnected federated nodes of the T2DKP that would enable researchers to interact with all of the data regardless of where they are located.

This goal was realized with the creation, by a team led by Thomas Keane and Dylan Spalding, of a federated node of the T2DKP at the European Bioinformatics Institute (EBI).  Data housed at the EBI node are stored in such a way that their specific privacy requirements are met, but they are made available for remote queries via T2DKP tools and interfaces. Results from such queries are served up alongside results from all of the datasets housed in the AMP T2D Data Coordinating Center (DCC) at the Broad Institute. Researchers may browse and query data from any location without even needing to know where they reside. This federation mechanism represents both an important technical advance in handling and protecting data, and a significant step forward in democratizing and improving access to genetic association results.

The first dataset to be incorporated into the Portal via the EBI federated node was the Oxford BioBank exome chip analysis dataset, which contains association data for glycemic, lipid, and blood pressure traits from over 7,100 subjects in Oxfordshire, U.K. The EBI Federated Node has now added three more datasets:

  • The EXTEND GWAS dataset, generated by Drs. Timothy Frayling and Andrew Wood and their colleagues, is comprised of 7,159 samples (1,395 T2D cases and 5,764 controls) from the Exeter EXTEND Biobank. It includes associations for a wealth of glycemic, anthropometric, cardiovascular, renal, and hepatic phenotypes--including many that are new to the T2DKP.
  • The GoDARTS Affymetrix GWAS dataset, from Dr. Colin Palmer and colleagues, includes summary-level statistics for associations with BMI and blood lipid levels from 3,307 diabetic participants in the Genetics of Diabetes Audit and Research Study in Tayside Scotland. In addition, individual-level data from over 17,000 subjects (including the set from which summary statistics were calculated) are available via the GAIT tool (see below). 
  • The Oxford BioBank Axiom GWAS dataset, from Dr. Fredrik Karpe and colleagues, includes associations for BMI and blood lipid levels from 7,193 participants, all healthy men and women between 30 and 50 years of age. It represents an additional analysis of the same samples contained in the Oxford BioBank exome chip analysis dataset.
These datasets are described in detail on our Data page. Summary results from all three sets are integrated into Gene and Variant pages in the T2DKP, and may also be viewed in the Manhattan plots accessible by searching for a phenotype from the T2DKP home page. The Variant Finder also queries these datasets.

The individual-level data behind all three of these datasets is accessible for custom association analysis in our Genetic Association Interactive Tool (GAIT) on Variant pages. Using this tool, researchers can filter samples to create a custom subset with defined characteristics such as age, gender, BMI, and other measures, and then run on-the-fly association analysis within that sample subset. Now, GAIT queries datasets both at the DCC and at the Federated node, using the same methodology for each, in a way that is transparent to users of the tool. The new Federated datasets bring the total number of individual-level samples available for custom analysis in the T2DKP to 67,768.

Friday, January 19, 2018

New METSIM dataset adds individual-level GWAS data to the T2DKP

The Finnish population is a valuable genetic resource. Having undergone multiple population bottlenecks, this relatively homogeneous population is enriched in low-frequency and loss-of-function variants. Even better, Finns are generally willing to participate in research studies, and many measures of their health are detailed in comprehensive electronic health records.

To take advantage of these characteristics, the METSIM (Metabolic Syndrome in Men) study (Laakso et al. 2017, J. Lipid Res. 58, 481-493) was initiated in 2005. Over 10,000 Finnish men were examined between 2005 and 2010. All of the subjects were phenotyped extensively, with an emphasis on traits associated with type 2 diabetes (T2D), cardiovascular disease, and insulin resistance, and their genotypes and exome sequences were determined. Subsets of the group have been characterized in more detail, with whole-genome sequencing and detailed analyses of transcripts and gene expression, DNA methylation, gut microbiome composition, and other phenotypes.

Now, you can easily access results from the METSIM cohort in the T2D Knowledge Portal. Variant associations with T2D, fasting glucose levels, and fasting insulin levels are available, both unadjusted or adjusted for body mass index. The individual-level data are also available for interactive analyses using our Genetic Association Interactive Tool (GAIT; see below), which allows you to design and run custom association analyses using custom subsets of the samples, while always protecting patient privacy. The addition of METSIM data brings to nearly 68,000 the number of samples available for analysis in GAIT.

The Foundation for the NIH and the Accelerating Medicines Partnership in Type 2 Diabetes were instrumental in bringing these data, generated by researchers in Finland and the U.S., to the T2DKP. Individual-level genotype data from 1,185 T2D cases and 7,357 controls were deposited into the Data Coordinating Center (AMP T2D DCC), and analysis and quality control were performed by the DCC analysis team. The experiment design and analysis are summarized on our Data page, and detailed reports that fully document the analysis are available for download.

The METSIM GWAS dataset currently has "Early Access Phase 1" status in the T2DKP, which is assigned to new data. This status denotes that although analysis and quality control checks have been performed, the data are not yet considered to be in their final state. During the early access period, users may analyze the data but may not submit the results of these analyses for publication. Find full details about the different phases of data release on our Policies page.

Results from METSIM GWAS may be viewed at these locations in the T2D Knowledge Portal:

• On Gene Pages (e.g., MTNR1B) in the Common variants and High-impact variants tables and in LocusZoom static plots, for the phenotypes T2D, T2D adjusted for BMI, fasting glucose, fasting glucose adjusted for BMI, fasting insulin, and fasting insulin adjusted for BMI;

• On Variant Pages (e.g.rs579060) in the Associations at a glance section, the Association statistics across traits table, and in LocusZoom static plots;

• From the View full genetic association results for a phenotype search on the home page: first select one of the phenotypes listed above, and then on the resulting page, select the METSIM GWAS dataset.

Individual-level METSIM GWAS data may be used for custom interactive analyses using these tools in the T2DKP:

• Using the Variant Finder tool, you may specify multiple criteria and retrieve the set of variants meeting those criteria;

• Using the Genetic Association Interactive Tool (GAIT) on Variant Pages, you may select the METSIM GWAS dataset, choose one of 5 phenotypes for association analysis, choose custom covariates, and filter the sample pool by specifying a range of values for one or more of 8 different phenotypes, then run on-the-fly analysis.

Phenotypes available for association analysis of METSIM GWAS data in GAIT


Covariates available for selection when analyzing METSIM GWAS data in GAIT


Samples may be filtered by setting ranges for one or more of 8 phenotypes for the METSIM GWAS dataset


Wednesday, October 25, 2017

New phenotypes and physical activity stratification available in the T2DKP

We’ve recently updated one dataset and added another in the Type 2 Diabetes Knowledge Portal. Associations with multiple new phenotypes are now available for the BioMe AMP T2D GWAS dataset, and the new dataset "GIANT GWAS - stratified by physical activity" adds associations with anthropometric traits for cohorts stratified by gender and physical activity levels.

The BioMe AMP T2D GWAS dataset was first added to the T2DKP in early 2017, initially with three phenotypes (T2D, fasting glucose levels, and HbA1c levels). Deposition and analysis of these data was funded by the Accelerating Medicines Partnership in Type 2 Diabetes (AMP T2D), a collaboration between multiple stakeholders that aims to catalyze the clinical translation of genetic discoveries by producing and aggregating data, developing and implementing novel analytical methods and tools, and building infrastructure for data storage and presentation. This dataset was the first to be entirely produced within the AMP T2D project, including the deposition, analysis, quality control, and presentation of the data.

The data were generated at the Charles Bronfman Institute for Personalized Medicine BioMe BioBank, a biorepository located at the Mount Sinai Medical Center (MSMC) in the upper Manhattan area of New York City. MSMC serves a diverse population of over 800,000 outpatients each year. Importantly, since many BioMe participants are African American or Hispanic Latino, this dataset adds significant ethnic diversity to the Portal’s genetic association data.

The data were subjected to quality control and association analysis by the Analysis Team at the AMP Data Coordinating Center (DCC) at the Broad Institute. In this second phase of analysis, associations with seven traits were calculated: systolic and diastolic blood pressure; HDL and LDL cholesterol levels; creatinine levels and eGFR-creat; and BMI. A detailed analysis report for these associations may be downloaded from the BioMe AMP T2D GWAS section of our Data page.

The new GIANT dataset was generated by the GIANT (Genetic Investigation of Anthropometric Traits) consortium via a meta-analysis of genetic associations for BMI, waist-hip ratio, and waist circumference from more than 200,000 adults. Samples are stratified by sex, ancestry, and physical activity level (active or inactive). This work was published in a recent paper by Graff et al.

Data from both the BioMe and GIANT studies are available at these locations in the Portal:
  • On Gene pages (see an example) in the Common variants and High-impact variants tables and in LocusZoom static plots
  • On Variant pages  (see an example) in the Associations at a glance section and in the Association statistics across traits table, and in LocusZoom static plots 
  • Via the Variant Finder tool
  • "Manhattan plots" of associations across the genome may be seen by selecting one of the phenotypes analyzed in these datasets in the View full genetic association results for a phenotype scroll box on the Portal home page
  • Additionally, the BioMe data are available for sample filtering and custom association analysis via the Genetic Association Interactive Tool (GAIT) on Variant pages.

Please check out the new data and contact us with any questions, comments, or suggestions.

Friday, June 9, 2017

Providing data access, ensuring data protection

Readers of this post probably don’t need to be convinced that genetic association data have enormous potential for helping us to understand and treat complex diseases like type 2 diabetes. Significant associations between variants and diseases can suggest genes, or regions of the genome, that could be important for disease risk or progression—and this knowledge could help us identify new drug targets.

The Accelerating Medicines Partnership in Type 2 Diabetes (AMP T2D) is a pre-competitive partnership among the National Institutes of Health, industry and not-for-profit organizations, which is managed by the Foundation for the National Institutes of Health. Its mission is to make genetic association data accessible to the worldwide biomedical research community via the Type 2 Diabetes Knowledge Portal, in order to facilitate discovery of new targets for T2D treatment. But it can be a challenge to aggregate genetic data. The privacy of the individuals who contributed their health status and genomic sequences must always be protected, and there are many layers of regulation to ensure this. Restrictions at the institutional, regional, and national levels determine how data are handled and whether they can be transferred.

Until now, all of the results displayed in the Portal have been derived from data housed at the AMP T2D Data Coordinating Center (DCC) at the Broad Institute, where the Portal website resides. But some of the valuable data generated outside the U.S. cannot be transferred to the DCC. To address this issue, AMP T2D funded the development of a mechanism that enables researchers to interact with all of the data: federation. 

Federation means that data are housed at a site (a “federated node”) that meets their specific privacy requirements, but are made available for remote queries via the Portal. Results from such queries are served up alongside results from all of the datasets housed in the AMP T2D DCC. Researchers may browse and query data from any location without even needing to know where they reside.

A federated node has now been created at the European Bioinformatics Institute (EBI) and may be accessed via the T2D Knowledge Portal. Today, Portal tools and interfaces can query both data housed at the AMP T2D DCC at the Broad Institute and data at the EBI federated node. 

According to Paul Flicek, a Senior Scientist and Team Leader of Vertebrate Genomics at EMBL-EBI, “A key mission of EMBL-EBI is to make data available to the widest possible community. Seamlessly accessing stored in multiple locations via a single portal helps ensure that the data we store from many projects are maximally useful for additional research.”

The first dataset to be incorporated into the Portal via the EBI federated node is the Oxford BioBank exome chip analysis dataset, which contains association data for glycemic, lipid, and blood pressure traits from over 7,100 healthy subjects in Oxfordshire, U.K. The dataset is described on our Data page. Portal users can interact with this dataset in the same way (and with the same speed) as with other datasets. 

“Diabetes is a global problem, and it will take research and innovation on a global scale if we are to tackle it effectively,” says Mark McCarthy, Robert Turner Professor of Diabetic Medicine at University of Oxford. “The success of our research on the genetics of diabetes depends on access to data generated by groups around the world. The federated portal provides an additional set of tools that will allow us to jointly analyse those data sets wherever they happen to be based.” 

Federation represents both an important technical advance in handling and protecting data, and a significant step forward in democratizing and improving access to genetic association results. And because it is generally applicable to any kind of genetic association data, it has the potential to have an impact beyond T2D research, facilitating the study of other complex diseases and traits.

Wednesday, June 7, 2017

New clues about variant effects: epigenomic data now available in the Portal

The T2D Knowledge Portal aggregates a wealth of genetic association data identifying variants that are associated with type 2 diabetes and related traits. These identifications show us that something within these genomic regions contributes to the risk of developing T2D. That’s an important first step, but in order to make use of this information to develop new T2D treatments, we need to figure out exactly what is causing the effect and how it relates to the disease process.

If a variant lies within a gene and changes a protein sequence, it can be relatively straightforward to formulate testable hypotheses about its effects. But most of the variants that are significantly associated with T2D—and with complex diseases in general—lie within noncoding regions of the genome and are likely to affect regulation of genes that could be far removed from the chromosomal position of the variant. It can be difficult to find clues about which genes are affected by these distant, noncoding changes, but now, we present a new type of data in the T2DKP that can help address this challenge.

The pattern of epigenetic modifications within a genomic region can provide important clues about its regulatory role. The distribution of these position-specific and tissue-specific marks—for example, covalent modifications of the histone proteins that package DNA—is characteristic of elements such as enhancers or transcription start sites. The Roadmap Epigenomics Consortium has developed methods for detecting these modifications genome-wide (hence the term “epigenomics”) and integrating their positional data, using the ChromHMM algorithm, to categorize genomic regions into “chromatin states”. The presence of these states in a given genomic region in different tissue types can give hints about whether that region might be involved in regulation of specific genes or pathways.

Now, you can view the tissue-specific chromatin states spanning the position of each variant on Variant pages within the Portal. We have incorporated epigenomic data from a study (Varshney et al., 2017) in which the locations of 13 distinct chromatin states were determined across a diverse set of cell lines and tissues, including pancreatic islets. The new “Epigenomic annotations” section of each Variant page (see an example) presents information about chromatin states in three different ways.

1. An interactive table listing chromatin states in this region, the tissue or cell line in which they were observed, and their genomic coordinates. Filter the table by chromatin state or by tissue to find states of particular interest.



2. A matrix displaying chromatin states by tissue type. This graphic gives a quick indication of chromatin states that are present in this region, across the whole panel of tissues.



3. A graphic showing the positions of chromatin states relative to the position of the variant.


These new features represent only the first phase of incorporating this new data type into the Portal. In the future, we will be adding more of these data along with more versatile interfaces for exploring them. Please check out our new epigenomic annotations and send us your feedback!

Wednesday, May 3, 2017

Explore new datasets and phenotypes in the T2D Knowledge Portal

We are releasing multiple new datasets in the Portal and have updated existing sets with associations for new phenotypes. To make it even easier to browse and explore these sets, we've also updated our Data page and added new functionality. Here's an overview of what's new in the Portal today.

17K exome sequence analysis dataset has grown to 19K
The Data Coordinating Center (DCC) of the Accelerating Medicines Partnership in Type 2 Diabetes (AMP T2D) analyzes exome sequence data contributed by AMP T2D consortium members to find variant associations with T2D and related traits. The exome sequencing dataset available in the Portal has until now consisted of exome sequences from about 17,000 individuals. Today, we have added exome sequencing performed on 2,000 Danish subjects by the LuCamp (Lubeck Foundation Centre for Applied Medical Genomics in Personalised Disease Prediction, Prevention and Care) consortium, making a total of nearly 19,000 exomes. This is just a taste of things to come: at the AMP T2D DCC we are currently analyzing additional exome sequences that will bring the total up to 52,000!

New community-contributed datasets: GENESIS GWAS and 70KforT2D GWAS

We are grateful to two groups from the larger T2D research community who have shared data that will make the T2D Knowledge Portal even more valuable to worldwide T2D researchers.

The GENEticS of Insulin Sensitivity (GENESIS) consortium performed GWAS on over 2,700 nondiabetic participants, finding genetic associations with direct measures of insulin sensitivity.

The 70KforT2D project collected, harmonized, and re-analyzed public GWAS data from over 70,000 individuals to find T2D genetic associations.

New public dataset: VATGen GWAS

The VATGen GWAS consortium performed meta-analysis of GWAS data from a mixed-ancestry group of more than 18,000 people to identify genetic associations with the localization of body fat deposition, leading to insights into adipocyte development.

Updated dataset: glucose-stimulated insulin secretion phenotypes in MAGIC GWAS

A study by Prokopenko et al. analyzed genetic associations with insulin secretion. Associations of variants with nine different measures of insulin secretion, among them corrected insulin response (CIR) and disposition index (DI), have now been added to the MAGIC GWAS dataset.

ExAC updated to gnomAD exomes and whole genomes
The Exome Aggregation Consortium (ExAC) has more than doubled in size and has morphed into the Genome Aggregation Database (gnomAD). More than 120,000 exome sequences and 15,000 whole genome sequences are now available, and these data are accessible via several tools and interfaces in the T2D Knowledge Portal.

New Data page: explore datasets using new filters

As our collection of data grows, it becomes more difficult to understand the differences between datasets and to find those of interest. To address this challenge, we've reorganized and streamlined our Data page.


A section of the Data page, expanded to show phenotype selection.

At the top of the Data page, you can choose to filter the dataset table by data type, phenotype category, or both. When you click on a phenotype category, the phenotypes within that category are available for selection. Clicking on the name of any dataset expands a section with details and references for each. 

In the coming days, watch this space for more details about each of these new developments. And as always, please contact us if you have any comments or questions.


Monday, January 23, 2017

Insulin Sensitivity Index data added to the Portal

The loss of sensitivity to insulin, often termed insulin resistance, is characteristic of type 2 diabetes. Since this sensitivity is difficult to measure directly, researchers have developed an index that reflects it: the modified Stumvoll Insulin Sensitivity Index (ISI). The index is derived by a formula that combines fasting insulin levels with glucose and insulin levels measured two hours after a glucose load.

Now, the results of a study of genetic associations of variants with ISI are available in the T2D Knowledge Portal. These results are from a recent paper in Diabetes by co-first authors Geoffrey Walford, Stefan Gustafsson, Denis Rybin, and fellow members of the Meta-Analyses of Glucose and Insulin-related traits Consortium (MAGIC). (For an overview of the results, see our blog post about the paper.)

In this study, ISI was calculated for 16,753 non-diabetic individuals, and associations of their variants with ISI values were analyzed. The associations were adjusted in one of three ways: for age and sex; for age, sex, and body mass index (BMI); or according to a model that analyzed the combined influence of the genotype effect adjusted for BMI and the interaction effect between the genotype and BMI on ISI. More details about this data set and others from MAGIC may be found on our Data page.

ISI associations are a subset of the MAGIC GWAS data set. They may be viewed in the Portal by selecting one of these phenotypes:
  • ISI adjusted for age-sex
  • ISI adjusted for age-sex-BMI
  • ISI adjusted for genotype-BMI interaction
Associations with these phenotypes can be found in these locations on Portal pages:
  • On Gene Pages (see an example) in the Variants & Associations table
  • On Variant Pages (see an example) in the Associations at a glance section and in the Association statistics across traits table
  • Via the Variant Finder tool, for the phenotypes listed above
  • A "Manhattan plot" of associations across the genome may be seen by selecting one of the phenotypes listed above in the View full genetic association results for a phenotype scroll box on the Portal home page.

Tuesday, January 17, 2017

New Year, New Data: BioMe AMP T2D GWAS

We’re happy to announce the first addition of data to the Type 2 Diabetes Knowledge Portal in 2017: the BioMe AMP T2D GWAS data set. The generation of these data was funded by the Accelerating Medicines Partnership in Type 2 Diabetes (AMP T2D), a collaboration between multiple stakeholders that aims to catalyze the clinical translation of genetic discoveries by producing and aggregating data, developing and implementing novel analytical methods and tools, and building infrastructure for data storage and presentation.

The BioMe AMP T2D GWAS data set is the first set to be entirely produced by the AMP T2D project, which supplied the funding and carried out every step of its production, from data generation to analysis, quality control, and presentation. Its immediate availability in the Portal, prior to publication, fulfills the mission of AMP T2D to speed up access to and utilization of new data.

These data were generated at the Charles Bronfman Institute for Personalized Medicine BioMe BioBank, a biorepository located at the Mount Sinai Medical Center (MSMC) in the upper Manhattan area of New York City. MSMC serves a diverse population of over 800,000 outpatients each year. Importantly, since many BioMe participants are African American or Hispanic Latino, this data set adds significant ethnic diversity to the Portal’s genetic association data.

The BioMe AMP T2D GWAS data set is comprised of about 13,000 unique individuals, 41.5% of whom are admixed American, 38% African American, and 20% European. Subjects were genotyped using at least one of three platforms: the Illumina Exome Array, the Illumina GWAS array, or the Affymetrix GWAS array. Their T2D status was assessed by an algorithm, and many additional traits were also measured.

The data were subjected to quality control and association analysis by the Analysis Team at the AMP Data Coordinating Center (DCC) at the Broad Institute. Variant associations with T2D, fasting glucose levels, and HbA1c levels were analyzed. The top results included both previously known and novel variants, with only a single variant reaching genome-wide significance: T2D association of the variant rs7903146, within the well-established T2D risk gene TCF7L2. Now that these results are available in the T2D Knowledge Portal, the ability to analyze them further in the context of all other available T2D association data may lead to additional insights.

The BioMe AMP T2D GWAS data currently has the “Early Access Phase 1” status that is assigned to new data. This status denotes that although analysis and quality control checks have been performed, the data are not yet considered to be in their final state. During the early access period, users may analyze the data but may not submit the results of these analyses for publication. Find the full details about the different phases of data release on our Policies page. More information about the data set, along with links to download even more detailed reports on its quality control and analysis, may be found in the BioMe AMP T2D GWAS section of our Data page.

BioMe AMP T2D GWAS data are available at these locations in the Portal:

  • On Gene Pages (see an example) in the Variants & Associations table and the Minor allele frequencies across data sets table
  • On Variant Pages  (see an example) in the Associations at a glance section and in the Association statistics across traits table
  • Via the Variant Finder tool, for these phenotypes: type 2 diabetes; fasting glucose adjusted for age and sex; HbA1c adjusted for age and sex; and HbA1c adjusted for age, sex, and body mass index
  • A "Manhattan plot" of associations across the genome may be seen by selecting one of the phenotypes above in the View full genetic association results for a phenotype scroll box on the Portal home page, and then selecting the BioMe AMP T2D GWAS data set.

As always, please contact us with any questions, comments, or suggestions.

Monday, November 7, 2016

New MGH Cardiology and Metabolic Patient Cohort data in the T2D Knowledge Portal

We are pleased to announce a new data set in the T2D Knowledge Portal, from the MGH Cardiology and Metabolic Patient Cohort (CAMP). These data were contributed by Pfizer, Inc. as part of a public-private partnership to generate genotype data for a cardiometabolic and prediabetic cohort. This data set adds individual-level genetic association data for type 2 diabetes (T2D), fasting glucose levels, and fasting insulin levels from more than 3,500 samples to the Portal knowledgebase. Association data for additional phenotypes from this cohort will be incorporated in the future.

The inclusion of this data set in the T2D Knowledge Portal illustrates the uniqueness of the Accelerating Medicines Partnership, which brings together pharmaceutical companies and non-profit institutions with the goal of speeding up the discovery of new targets for treatment of T2D. The pharmaceutical partners in this collaboration have committed not only to providing funding, but also to sharing the data they generate. The CAMP data set contributed by Pfizer is the first set from a pharmaceutical partner to be made available in the Portal.

Another unique aspect of this data set is that it is the first to be included in the Portal with “Early Access Phase 1” status, which is assigned to new data. This status denotes that although analysis and quality control checks have been performed, the data are not yet considered to be in their final state. During the early access period, users may analyze the data but may not submit the results of these analyses for publication. Find the full details about the different phases of data release on our Policies page.

The CAMP cohort consists of 3,857 subjects who were recruited at the Massachusetts General Hospital Heart Center between 2008 and 2012. In addition to genotyping, the subjects had either vascular reactivity measurements (for T2D patients) or an oral glucose tolerance test (for patients not known to have T2D), and samples of their plasma and serum were analyzed. Most of the subjects were of European ancestry; about 10% were African American.

The analysis and quality control processes for this data set were performed by the Analysis Team of the Accelerating Medicines Partnership Data Coordinating Center (AMP-DCC) at the Broad Institute, and are completely transparent and fully documented. The experiment design and analysis are summarized on our Data page, and detailed reports are available for download. Going forward, all new data sets added to the Portal will be fully documented in this manner.

One intriguing—and somewhat puzzling—result from the analysis highlights the utility of incorporating data sets like this one into the Portal. The variant most strongly associated with T2D (at genome-wide significance) in this set is located in the major histocompatibility complex region near the HLA-C gene.

Known associations of genes in this region with type 1 diabetes, along with a high local recombination rate, make it challenging to interpret the meaning of this association. However, it certainly merits further investigation because of its genome-wide significance. The inclusion of this data set in the Portal, in the context of all other available data about T2D associations in the region, greatly facilitates the further analysis of this and other associations in the set.

The CAMP data may be accessed via multiple interfaces in the Portal. They are shown in tables of summary statistics and accessible in variant searches using the Variant Finder. Importantly, since the data are individual-level, samples may be filtered by various parameters and used for custom association analysis in our Genetic Association Interactive Tool (GAIT).

Find CAMP data at all of these locations in the Portal:

On Gene Pages (e.g.,  HLA-C) in the Variants & Associations table.
On Variant Pages (e.g., rs9468919) in the Associations at a glance section and in the Association statistics across traits table.
Via the Variant Finder tool, for the phenotypes T2D, fasting glucose, and fasting insulin.
Via the Genetic Association Interactive Tool (GAIT), which enables custom association analysis for either single variants (available on Variant Pages) or for the set of variants in and near a gene (Interactive burden test, available on Gene Pages).

Tuesday, October 18, 2016

New and updated data in the T2D Knowledge Portal

As members of the T2D Knowledge Portal team arrive in Vancouver for the American Society of Human Genetics meeting, we are pleased to announce that we have added a new data set to the Portal and made extensive updates to existing data sets. 

The new data set, named “CAMP GWAS” in the Portal, comes from the MGH Cardiology and Metabolic Patient Cohort (CAMP). These data were contributed by Pfizer, Inc. as part of a public-private partnership to generate genotype data for a cardiometabolic and prediabetic cohort, and were analyzed by the Analysis Team of the Accelerating Medicines Partnership Data Coordinating Center (AMP-DCC) at the Broad Institute. The set adds individual-level genetic association data for type 2 diabetes (T2D), fasting glucose levels, and fasting insulin levels from nearly 3,500 samples to the Portal knowledgebase, and association data for more phenotypes will be added in the future.

CAMP data may be accessed on Gene and Variant pages in the Portal and via the Variant Finder, and may also be filtered and queried using the Genetic Association Interactive tool (GAIT).

Several other data sets in the Portal have been updated and improved:
  • The size of the CARDIoGRAM GWAS data set has nearly doubled, now consisting of 184,305 samples, and the data analysis has been updated.
  • The size of the CKDGen GWAS data set has also nearly doubled, to 133,814 samples; the data analysis has been updated; new subsets have been added that stratify serum creatinine associations by African American ancestry and stratify both serum creatinine and urinary albumin-to-creatinine ratio by the presence or absence of T2D.
  • The data set previously named “DIAGRAM GWAS” in the Portal has been updated and re-named “DIAGRAM Trans-ethnic meta-analysis;” its sample size has increased to 149,821. Several new subsets have been added, including gender-stratified, MetaboChip, and fine mapping data.
  • The GIANT GWAS data have been updated and European cohorts have been added for BMI and height traits.
  • The GLGC GWAS data set has increased in size to 188,577 samples and has been updated.
  • The number of samples in the MAGIC GWAS dataset has more than doubled, to 133,010; the data have been updated, and associations with 2 hour glucose, fasting glucose, and fasting insulin have been added for MetaboChip data.
Full details about all of these data sets are available on our Data page.

Because of compatibility issues with the updated data, we have temporarily removed the “GWAS results summary” section from Gene pages of the Portal. This feature will be restored within the next week.

As always with major updates, issues or bugs may have been introduced and we may not have found all of them during our routine testing. We encourage you to let us know of any problems that you encounter in using the Portal, and we welcome your questions and suggestions.

Wednesday, May 18, 2016

Expanding the landscape of human genetic variation data in the Type 2 Diabetes Knowledge Portal

With the addition of four new sequence data sets to our database, the number of variants and associations accessible via the Portal pages and tools has increased by millions.

Two of the new data sets are from projects that have obtained sequence data from a wide range of individuals. The ExAC data set, comprising exome sequences collected and harmonized by the Exome Aggregation Consortium, includes sequence data from 60,706 unrelated people of multiple ancestries. The 1000 Genomes data set, from the International Genome Sample Resource project (IGSR), is composed of whole-genome sequences from 2,504 individuals in four different ethnic groups. 


The allele frequencies of variants in the different ethnic groups surveyed in the 1000 Genomes data set can be seen in the “How common is…?” section on the Variant pages (view an example). And both the ExAC and 1000 Genomes data sets can be queried using the Variant Finder tool. You can select them via a new tab on the interface, “Additional search options”, where you can choose these data sets and also add more criteria to your search. 

The Data set pull-down menu on the "Additional Search Options" tab of the Variant Finder lets you specify 1000 Genomes or ExAC data.

Available selections in the Data set pull-down menu.


The other two new data sets in the Portal were both generated by the GoT2D consortium. A whole-genome sequence data set (GoT2D WGS) adds data from 2,657 individuals, including the associations of noncoding variants that were not present in the previous whole-exome sequence data set from the GoT2D project. This new data set brings T2D association data across 30 million variants to the Portal. The GoT2D WGS + replication data set adds imputation to that set, bringing the sample size to over 47,000 and including most low-frequency and common variants.  

The new GoT2D data can be seen in multiple sections of the Portal’s Gene and Variant pages, and may also be accessed by selecting these data sets in the Variant Finder.

In addition to these major new additions, today’s release of data also includes some bug fixes and data harmonization.

Get out there and explore the new data landscape in the Portal, and let us know what you think!