Showing posts with label human genetics. Show all posts
Showing posts with label human genetics. Show all posts

Monday, October 8, 2018

DIAMANTE GWAS dataset adds close to a million samples along with fine-mapping to the T2DKP

In a groundbreaking paper published today, Anubha Mahajan and colleagues (Mahajan et al., Nature Genetics 2018) report on a meta-analysis of unprecedented size for genetic associations with type 2 diabetes (T2D) along with fine-mapping analyses to identify causal variants that can suggest new therapeutic targets. We are pleased to provide access to the summary results as well as the results of the fine-mapping today in the T2D Knowledge Portal (T2DKP).

Working as part of the DIAGRAM (DIAbetes Genetics Replication And Meta-analysis) and DIAMANTE (DIAbetes Meta-ANalysis of Trans-Ethnic association studies) consortia, the researchers aggregated and meta-analyzed genome-wide association studies for about 900,000 individuals of European ancestry (about 74,000 T2D cases and 824,000 controls). The studies were imputed using the most comprehensive reference panels possible, and in all, the analysis considered about 27 million genotyped or imputed variants.

After performing T2D association analysis (both unadjusted and adjusted for body mass index) 243 loci were seen to be associated with T2D at genome-wide significance or better (p-value for association ≤ 5 x 10-8). Of these, 135 were novel--not detected previously in any T2D association analysis to date.

Within these loci, each of which included multiple significantly associated variants, the researchers performed approximate conditional analysis to determine whether the associations were independent of each other. They found surprising complexity within some loci; for example, the well-known TCF7L2 locus appears to include as many as 8 distinct association signals!

All of the T2D associations from this study may be viewed in the T2DKP. They are represented in two datasets, named "DIAMANTE (European) T2D GWAS" and "UK Biobank T2D GWAS (DIAMANTE-Europeans Sept 2018)."  Manhattan plots showing the distribution of the associations across the genome may be seen by selecting either the "Type 2 diabetes" or "Type 2 diabetes adj BMI" phenotypes from the phenotype selection menu on the T2DKP home page. On Gene pages of the T2DKP, the results may be viewed in tables of variant associations and in the interactive LocusZoom visualization (see below). Results from this study are also displayed on Variant pages of the T2DKP.


LocusZoom plot on the PPARG Gene page


The credible set analysis performed in this study is also incorporated into the T2DKP. On the "Credible sets" tab of Gene pages, you may choose to visualize any of the credible sets available for the region. Epigenomic annotations that overlap the positions of the variants in the credible set are presented in an interactive display that allows you to select particular chromatin states or tissues to view. In the example shown below, one of the credible sets in the TCF7L2 region includes just two variants, and the one with the highest posterior probability overlaps active enhancer regions in adipose and liver tissue--both of which are important for T2D.


Detail of the Credible sets tab of the TCF7L2 Gene page

The multiple causal variants identified in this study support previous investigations on the biological mechanisms behind T2D and suggest new hypotheses that will likely lead to therapeutic insights. After reading the paper and a blog post from the authors, we invite you to explore the results in the T2DKP and to contact us with any suggestions or questions!

Wednesday, August 30, 2017

Bringing the power of epigenomics to the T2DKP

Until recently, all of the results displayed in the Type 2 Diabetes Knowledge Portal (T2DKP) were based on genetic association data: the significance with which variants, or SNPs, occur in people’s genomes in conjunction with a disease or trait.

This information is hugely important for pinpointing regions of the genome that contribute to disease risk. It is now relatively straightforward to identify these regions, but it is still a large challenge to discover the mechanisms by which they act—especially for variants that are outside of coding sequences, without an obvious effect on the sequence of a particular protein. These non-coding variants, the most commonly seen in genetic association studies, are likely to affect tissue-specific gene regulation that could potentially be important to the disease process.

How can we overcome this challenge to find clues about the effects of these non-coding variants? Epigenomic data to the rescue!

Dr. Kyle Gaulton of the University of California at San Diego researches the transcriptional regulatory networks involved in type 2 diabetes by using epigenomic data in concert with genetic association data. He explains, "Regulatory elements control gene production and function, and are often highly specialized across cell and tissues and located far away from the genes they regulate. Molecular epigenomic hallmarks of gene regulation such as histone and DNA modifications, nucleosome depletion, chromatin conformation and DNA-protein interactions can pinpoint the precise genomic locations of regulatory elements. High-resolution epigenome maps of regulatory elements in pancreatic islets, liver, muscle, adipose and many other human tissues can then enable annotation of non-coding genetic variants and their potential gene regulatory functions. These maps are thus an invaluable component of determining how type 2 diabetes associated non-coding variants influence disease pathogenesis."

A recent paper from Dr. Gaulton and colleagues (Gaulton, KJ, et al. (2015) Nat Genet. 47:1415) illustrates the power of integrating these two data types. By combining information on transcription factor binding sites and tissue-specific chromatin states with genetic fine-mapping of T2D-associated loci, the authors elicidated the molecular mechanisms behind the effects of some T2D-associated variants, uncovering the role of the FOXA2 transcription factor in glucose homeostasis in T2D-relevant tissues.

Now, the T2DKP facilitates this type of analysis by presenting both genetic association and epigenomic data on Gene and Variant pages. We described the display of epigenomic data on Variant pages in a recent blog post. On Gene pages, epigenomic data are integrated into the LocusZoom display.

Locations of variants associated with T2D and chromatin states in pancreatic islets, across the SLC30A8 gene (partial view)


Below the plot of variant associations, chromatin states are displayed by default for the major T2D-relevant tissues. Using the pull-down menu at the top of the plot, you can choose from a diverse set to display other tissues and cell types. All of the details on how to use this interactive plot are included in our Gene Page guide.

This is only the first step for epigenomic data in the T2DKP. In the future, we plan to include additional types of epigenomic data that indicate chromatin accessibility and conformation. We will also add functionality; for example, for any given variant, you will be able to search for the tissues in which enhancer regions overlap the location of that variant.

As we actively develop this aspect of the T2DKP, we welcome your suggestions!

Friday, June 9, 2017

Providing data access, ensuring data protection

Readers of this post probably don’t need to be convinced that genetic association data have enormous potential for helping us to understand and treat complex diseases like type 2 diabetes. Significant associations between variants and diseases can suggest genes, or regions of the genome, that could be important for disease risk or progression—and this knowledge could help us identify new drug targets.

The Accelerating Medicines Partnership in Type 2 Diabetes (AMP T2D) is a pre-competitive partnership among the National Institutes of Health, industry and not-for-profit organizations, which is managed by the Foundation for the National Institutes of Health. Its mission is to make genetic association data accessible to the worldwide biomedical research community via the Type 2 Diabetes Knowledge Portal, in order to facilitate discovery of new targets for T2D treatment. But it can be a challenge to aggregate genetic data. The privacy of the individuals who contributed their health status and genomic sequences must always be protected, and there are many layers of regulation to ensure this. Restrictions at the institutional, regional, and national levels determine how data are handled and whether they can be transferred.

Until now, all of the results displayed in the Portal have been derived from data housed at the AMP T2D Data Coordinating Center (DCC) at the Broad Institute, where the Portal website resides. But some of the valuable data generated outside the U.S. cannot be transferred to the DCC. To address this issue, AMP T2D funded the development of a mechanism that enables researchers to interact with all of the data: federation. 

Federation means that data are housed at a site (a “federated node”) that meets their specific privacy requirements, but are made available for remote queries via the Portal. Results from such queries are served up alongside results from all of the datasets housed in the AMP T2D DCC. Researchers may browse and query data from any location without even needing to know where they reside.

A federated node has now been created at the European Bioinformatics Institute (EBI) and may be accessed via the T2D Knowledge Portal. Today, Portal tools and interfaces can query both data housed at the AMP T2D DCC at the Broad Institute and data at the EBI federated node. 

According to Paul Flicek, a Senior Scientist and Team Leader of Vertebrate Genomics at EMBL-EBI, “A key mission of EMBL-EBI is to make data available to the widest possible community. Seamlessly accessing stored in multiple locations via a single portal helps ensure that the data we store from many projects are maximally useful for additional research.”

The first dataset to be incorporated into the Portal via the EBI federated node is the Oxford BioBank exome chip analysis dataset, which contains association data for glycemic, lipid, and blood pressure traits from over 7,100 healthy subjects in Oxfordshire, U.K. The dataset is described on our Data page. Portal users can interact with this dataset in the same way (and with the same speed) as with other datasets. 

“Diabetes is a global problem, and it will take research and innovation on a global scale if we are to tackle it effectively,” says Mark McCarthy, Robert Turner Professor of Diabetic Medicine at University of Oxford. “The success of our research on the genetics of diabetes depends on access to data generated by groups around the world. The federated portal provides an additional set of tools that will allow us to jointly analyse those data sets wherever they happen to be based.” 

Federation represents both an important technical advance in handling and protecting data, and a significant step forward in democratizing and improving access to genetic association results. And because it is generally applicable to any kind of genetic association data, it has the potential to have an impact beyond T2D research, facilitating the study of other complex diseases and traits.

Wednesday, May 3, 2017

Explore new datasets and phenotypes in the T2D Knowledge Portal

We are releasing multiple new datasets in the Portal and have updated existing sets with associations for new phenotypes. To make it even easier to browse and explore these sets, we've also updated our Data page and added new functionality. Here's an overview of what's new in the Portal today.

17K exome sequence analysis dataset has grown to 19K
The Data Coordinating Center (DCC) of the Accelerating Medicines Partnership in Type 2 Diabetes (AMP T2D) analyzes exome sequence data contributed by AMP T2D consortium members to find variant associations with T2D and related traits. The exome sequencing dataset available in the Portal has until now consisted of exome sequences from about 17,000 individuals. Today, we have added exome sequencing performed on 2,000 Danish subjects by the LuCamp (Lubeck Foundation Centre for Applied Medical Genomics in Personalised Disease Prediction, Prevention and Care) consortium, making a total of nearly 19,000 exomes. This is just a taste of things to come: at the AMP T2D DCC we are currently analyzing additional exome sequences that will bring the total up to 52,000!

New community-contributed datasets: GENESIS GWAS and 70KforT2D GWAS

We are grateful to two groups from the larger T2D research community who have shared data that will make the T2D Knowledge Portal even more valuable to worldwide T2D researchers.

The GENEticS of Insulin Sensitivity (GENESIS) consortium performed GWAS on over 2,700 nondiabetic participants, finding genetic associations with direct measures of insulin sensitivity.

The 70KforT2D project collected, harmonized, and re-analyzed public GWAS data from over 70,000 individuals to find T2D genetic associations.

New public dataset: VATGen GWAS

The VATGen GWAS consortium performed meta-analysis of GWAS data from a mixed-ancestry group of more than 18,000 people to identify genetic associations with the localization of body fat deposition, leading to insights into adipocyte development.

Updated dataset: glucose-stimulated insulin secretion phenotypes in MAGIC GWAS

A study by Prokopenko et al. analyzed genetic associations with insulin secretion. Associations of variants with nine different measures of insulin secretion, among them corrected insulin response (CIR) and disposition index (DI), have now been added to the MAGIC GWAS dataset.

ExAC updated to gnomAD exomes and whole genomes
The Exome Aggregation Consortium (ExAC) has more than doubled in size and has morphed into the Genome Aggregation Database (gnomAD). More than 120,000 exome sequences and 15,000 whole genome sequences are now available, and these data are accessible via several tools and interfaces in the T2D Knowledge Portal.

New Data page: explore datasets using new filters

As our collection of data grows, it becomes more difficult to understand the differences between datasets and to find those of interest. To address this challenge, we've reorganized and streamlined our Data page.


A section of the Data page, expanded to show phenotype selection.

At the top of the Data page, you can choose to filter the dataset table by data type, phenotype category, or both. When you click on a phenotype category, the phenotypes within that category are available for selection. Clicking on the name of any dataset expands a section with details and references for each. 

In the coming days, watch this space for more details about each of these new developments. And as always, please contact us if you have any comments or questions.


Wednesday, March 15, 2017

The Portal’s interactive burden test: now more versatile than ever

Significant associations between genes and T2D or related phenotypes can provide powerful insights into disease mechanisms and possible therapies. The T2D Knowledge Portal includes results from pre-computed analyses of genetic associations for a large, and growing, number of datasets. But what if you want to do a more fine-grained analysis? You might want to test whether the disease burden for a gene differs between groups of people with specific characteristics—for example, lean people with T2D versus obese people without T2D. Or you might want to test the aggregate effect of a specific subset of variants, such as those that are likely to knock out the function of a protein of interest.

Our interactive burden test on Gene pages, powered by the Genetic Association Analysis Tool (GAIT), allows you to do all that and more. The burden test considers a gene as the unit of inquiry, including all the variants it contains in a statistical test of disease association. We described the basics of the burden test and GAIT in a recent blog post. Now, we’ve added some options for selecting variants in the interactive burden test that make this tool even more versatile.

The variant selection step of the burden test on a Gene page is pre-populated with all of the variants present in the selected dataset that are located within the gene and its 100 kb up- and downstream flanking regions. You can create a specific subset of these by checking or un-checking individual variants. The table may be sorted by multiple criteria in order to find variants of interest: chromosomal coordinate; minor allele count; predictions of the effect allele’s impact on the encoded protein; and the protein change or type of mutation caused by the effect allele.


Section of the interactive burden test interface showing the default list of variants for the SLC30A8 gene. Options for customizing the list are located above the variant table.

The table of variants may be filtered so that the test considers only certain categories of variants, with varying predicted impacts on the encoded protein. Previously, the burden test offered filters based on an unpublished method. Now, we have replaced those filters with the set that was used in a recent major publication: The genetic architecture of type 2 diabetes, by Fuchsberger, Flannick, Teslovich, Mahajan, Agarwala, Gaulton, et al.

Variant filters in the interactive burden test

All coding variants--selects variants within the coding sequence, from the dataset that was initially selected for the burden test

Protein-truncating + missense with MAF<1%--selects variants in both of these categories:
  • protein-truncating (predicted to cause a truncated protein to be generated, either by creating a premature stop codon or by causing a frameshift) 
  • cause a missense mutation AND have minor allele frequency (MAF) of less than 1%. The MAF limit eliminates common variants, which would not be expected to have very deleterious effects. 

Protein-truncating + possibly deleterious missense with MAF<1%--selects variants in both of these categories:

Protein-truncating + probably deleterious missense--selects variants in both of these categories:

Protein-truncating only--selects variants predicted to cause a truncated protein to be generated, either by creating a premature stop codon or by causing a frameshift.

Using these filters, you can tailor the list of variants to those with specific impact on the encoded protein. If you would like to customize the list even further by adding variants that were not present in the default list, there is now an option to add single or multiple variants, using dbSNP IDs (e.g., rs112881768) or identifiers in the format “chromosome_coordinate_reference-nucleotide_variant-nucleotide” (e.g., 8_112881768_G_A).

When “single variant” is selected, once you begin typing, variant IDs that match your entry are suggested. When “multiple” is selected, you may type or paste in a list of variant IDs, separated by commas or returns. Note that any added variants are not subject to the filters, which act only on the default list of variants for a gene.

Our GAIT User Guide (download PDF) that summarizes all the details of the interface has been updated with the latest changes. Please check out our new, improved interactive burden test and let us know if you have comments or suggestions.

Sunday, February 5, 2017

Introductory guide to genetic association analysis now available

P-values. Odds scores and betas. GWAS. Linkage disequilibrium. What does it all mean?

Human geneticists are, of course, intimately familiar with these concepts. But for people who are not human geneticists, just getting past the terminology can be frustrating. So we’ve written a basic primer and reference guide that can help users of the T2D Knowledge Portal understand the information presented in our interfaces and tools.

Our Introduction to genetic association analysis guide is available from our Resources page. Or download it here (PDF).

This guide provides a basic introduction to the rationale behind applying human genetic association studies to complex diseases like T2D, explains some of the parameters of genetic associations such as p-values and odds ratios, and describes the different types of experiment used to determine genetic associations.

Many thanks to Andrew Morris, University of Oxford, for his thoughtful review and helpful comments on this guide.

We would be happy to hear your suggestions for improvements and additions!

Monday, October 3, 2016

Come to a T2D Knowledge Portal information session at ASHG

The American Society of Human Genetics meeting is happening in Vancouver, B.C. in a little over two weeks! The Portal team will be presenting and exhibiting at multiple venues at ASHG, and the first event will take place immediately before the conference starts: an information session including an overview of the Accelerating Medicines Partnership in Type 2 Diabetes, a progress update on T2D Knowledge Portal functionality, and information on new funding opportunities. Complimentary hors d'oeuvres, beer and wine will be served!

Information session
Tuesday - October 18, 2016
3:00 pm - 4:00 pm PDT
Fairmont Waterfront

900 Canada Place Way

Vancouver, British Columbia


Please register here for this free event, hosted by FNIH.  Contact Nicole Spear at Nspear@fnih.org with any questions.

Watch this space over the next two weeks for a complete listing of opportunities to learn about the Portal and talk with the Portal team at ASHG!




Thursday, September 15, 2016

New funding opportunities for T2D genetic research


The Foundation for the National Institutes of Health (FNIH) has released three new funding opportunities that aim to add to the growing body of data housed in the T2D Knowledge Portal. The new Request for Proposals (RFPs) solicit data on T2D related complications and individual level and whole exome sequencing data related to T2D.

FNIH awards will provide successful applicants with up to $200,000 per individual award for proposals to harmonize and transfer existing datasets and up to $500,000 per individual award for proposals that include the generation of new genotyping data. Awards will be made over two years and aim to enhance the NIH-funded T2D Knowledge Portal hosted by the Broad Institute at the Massachusetts Institute of Technology (MIT).

Responses to the FNIH Requests for Proposals are due by December 31, 2016. Details on the new funding opportunities can be found here


Wednesday, August 10, 2016

Insulin sensitivity comes into focus

Many different things can be seen in any landscape, depending on your focal point.
Image by Nicooo76 via Pixabay.
When photographing a landscape, different photographers choose different perspectives. Some capture a wide-angle view, while others focus on particular details.

It’s no different for researchers who use genome-wide association studies (GWAS) to investigate the genetic landscape of type 2 diabetes (T2D). A common perspective is to study the wide range of variants that are significantly associated with the presence of T2D in patients. But it can also be very informative to concentrate on individual traits related to the physiology of T2D. In a new paper in Diabetes, co-first authors Geoffrey Walford, Stefan Gustafsson, Denis Rybin, and fellow members of the Meta-Analyses of Glucose and Insulin-related traits Consortium (MAGIC) took this focused perspective to discover associations of genetic variants with insulin sensitivity.

Along with reduced insulin levels, the loss of insulin sensitivity (often termed insulin resistance) is a major hallmark of T2D. When muscle, liver, and fat cells become less able to respond to insulin, blood glucose levels rise. Since this can contribute to development of T2D and exacerbate its symptoms, knowing which genetic variants are associated with sensitivity to insulin could be informative for understanding pathways that contribute to T2D risk.

But insulin sensitivity is difficult to measure. Earlier GWAS have used simple estimates of insulin sensitivity, such as fasting levels of insulin, and have discovered a handful of genetic variants that influence insulin sensitivity. The “gold standard” test, the euglycemic clamp, involves giving patients continuous infusions of insulin and glucose and monitoring their blood glucose every few minutes. It’s expensive and time-consuming—not a test that is practical to perform on the tens of thousands of subjects that are commonly used in GWAS.

The authors wondered whether they could instead use an index that combines several measurements, each relatively easy to make. It’s an index with a long name: the modified Stumvoll Insulin Sensitivity Index (ISI). Developed by Stumvoll and colleagues in 2001, this index can be derived in a variety of ways. The authors chose the ISI requiring just three measurements: fasting insulin levels; glucose levels two hours after a glucose load; and insulin levels two hours after a glucose load. This ISI is as good as or better than other estimates of insulin sensitivity and correlates well with the euglycemic clamp.

So the researchers looked for variants associated with the Stumvoll ISI in nearly 17,000 participants in the discovery phase of the work. They added another 13,300 in the replication phase, adding up to about 30,000 in the combined meta-analysis. Since obesity, measured by body mass index (BMI), can affect insulin sensitivity, the authors added BMI to some of their statistical models.

First, the authors found associations between the ISI and other variants already known to affect simple measures of insulin sensitivity. This provided reassurance that the ISI was properly detecting genetic influences on insulin sensitivity. After discovery, replication, and meta-analysis, two novel genetic variants were associated with ISI at genome-wide significance (P-value < 5.0 ×10-8) in a model that tested the effect of the variant, age, sex, and the interaction between the variant and BMI: variant rs12454712, near the gene BCL2, and variant rs10506418, near the gene FAM19A2.

How might these variants affect insulin sensitivity? There’s a lot more work to be done before that question can be answered. Additional studies will need to clarify whether these variants, which are near BCL2 and FAM19A2, affect these or other genes, and then how these variants actually cause changes in insulin sensitivity. 

There are some clues already in the published literature. The variant rs12454712 near BCL2 has previously been found to be associated with T2D, supporting the hypothesis that this region of the genome contributes to T2D risk through reducing insulin sensitivity. And the gene itself (BCL2) has already been implicated in glycemic metabolism: inhibiting bcl2 improves glucose tolerance in a mouse model, while a drug that inhibits the protein product of the gene (BCL2) increases blood glucose levels in certain chronic lymphocytic leukemia patients. So there’s even more reason to suspect that the rs12454712 variant might affect insulin sensitivity via BCL2.

There is as yet no evidence linking the protein FAM19A2 function to glucose metabolism, so the jury is out on whether the variant rs10506418 affects FAM19A2 or some other nearby gene. 

By focusing on a detail of the T2D-related genetic landscape, this study has teased out two variants that may give us clues about the physiology of insulin sensitivity and the development of T2D. And that’s a valuable addition to our overall picture of T2D genetics!

Friday, June 10, 2016

Come meet the Portal team at ADA, booth #1762!

Today’s news comes to you from the Big Easy—New Orleans, LA, where the 76th Scientific Sessions of the American Diabetes Association are in full swing this weekend. Members of the Knowledge Portal team have traveled here to talk to researchers about how the Portal can become even more useful in helping to generate hypotheses that spark insights into the mechanism of T2D and the development of new therapies. Starting at 10am on Saturday June 11, we’ll be at booth #1762 in the exhibit hall, ready to hear your suggestions and give you an individual tutorial on the Portal’s tools and features. There just might be a gift waiting for you, too!

We’ve been working hard and we have an incredible number of new features to show off at #2016ADA. We’ll be featuring them individually in this space in the coming weeks, with in-depth explanation of each. To list some of the highlights:

  • a collaborative project between software engineers at the University of Michigan and the Broad Institute has come to fruition with the integration of LocusZoom into the Portal. This interactive visualization looks, superficially, like a Manhattan plot—but it’s so much more. It shows the significance of variant associations with any of several phenotypes and also displays linkage disequilibrium among nearby variants, and you can choose to do conditional analysis based on any variant.
  • engineers at the Broad Institute have developed a completely new tool, called Genetic Association Interactive Tool (GAIT), that offers a multitude of options allowing you to compute custom association statistics for a variant. You can specify the phenotype to test for association, stratify samples by ancestry, choose a subset of samples to analyze based on specific phenotypic criteria, and control for specific covariates. 
  • we’ve also redesigned and augmented many of the displays of pre-computed information that are available in the Portal
  • finally, we’ve added a lot of new, informative content: a Data page with a complete description of each data set in the Portal, more background about the AMP-T2D project that supports the Portal, and more help text to guide you as you use the Portal’s interfaces



Come to the booth and let us give you a tour of these new features—or, if you're not at ADA, take a look and let us know what you think. And take a look at this great press release from NIH about the project!

Wednesday, May 18, 2016

Expanding the landscape of human genetic variation data in the Type 2 Diabetes Knowledge Portal

With the addition of four new sequence data sets to our database, the number of variants and associations accessible via the Portal pages and tools has increased by millions.

Two of the new data sets are from projects that have obtained sequence data from a wide range of individuals. The ExAC data set, comprising exome sequences collected and harmonized by the Exome Aggregation Consortium, includes sequence data from 60,706 unrelated people of multiple ancestries. The 1000 Genomes data set, from the International Genome Sample Resource project (IGSR), is composed of whole-genome sequences from 2,504 individuals in four different ethnic groups. 


The allele frequencies of variants in the different ethnic groups surveyed in the 1000 Genomes data set can be seen in the “How common is…?” section on the Variant pages (view an example). And both the ExAC and 1000 Genomes data sets can be queried using the Variant Finder tool. You can select them via a new tab on the interface, “Additional search options”, where you can choose these data sets and also add more criteria to your search. 

The Data set pull-down menu on the "Additional Search Options" tab of the Variant Finder lets you specify 1000 Genomes or ExAC data.

Available selections in the Data set pull-down menu.


The other two new data sets in the Portal were both generated by the GoT2D consortium. A whole-genome sequence data set (GoT2D WGS) adds data from 2,657 individuals, including the associations of noncoding variants that were not present in the previous whole-exome sequence data set from the GoT2D project. This new data set brings T2D association data across 30 million variants to the Portal. The GoT2D WGS + replication data set adds imputation to that set, bringing the sample size to over 47,000 and including most low-frequency and common variants.  

The new GoT2D data can be seen in multiple sections of the Portal’s Gene and Variant pages, and may also be accessed by selecting these data sets in the Variant Finder.

In addition to these major new additions, today’s release of data also includes some bug fixes and data harmonization.

Get out there and explore the new data landscape in the Portal, and let us know what you think!

Tuesday, March 29, 2016

New graphics and table summarize variant associations at a glance

Our variant information pages now contain two new sections that make it easy to see quickly whether a variant is associated with type 2 diabetes or related traits, and just how significant those associations are.

At the top of the Variant page for one particular variant, the section titled “Is (variant name) associated with disease?” opens to show the associations of that variant with T2D in all datasets that are currently available via the Portal (view an example). Click the link “expand associations for all traits” to see significant associations with other T2D-related traits.

Each box represents an association between this variant and a trait as detected in one data set, and the color of the box indicates the significance of the association. Dark green shows genome-wide significance (p-value < 5 x 10e-8); medium green shows locus-wide significance (p-value < 5 x 10e-4); and light green denotes nominal significance (p-value < 0.05). Associations that do not meet the threshold for significance are shown in a white box.


These new graphics make it easy to see quickly that the variant rs13266634 is strongly linked to T2D, fasting glucose levels, and proinsulin levels.

Just below this section, the “Association statistics across traits” table gives complete details about the associations between the variant and multiple traits. The same shades of green show the most significant associations.


In this table with more details about the associations of this variant, the consistent color scheme highlights significance levels.

Information shown in this table for the variant-trait associations may include p-value, direction of effect, odds ratio, minor allele frequency, and effect size. The table can be sorted by trait name. Where a variant-trait association was detected in more than one study, the most significant result is shown; plus signs allow you to expand the table and view results from additional studies. Some datasets can also be expanded to show associations in different ancestries or cohorts.

We’re still developing these new features, and your feedback could help us make them even better. Please explore them and let us know what you think!

Thursday, February 11, 2016

Spanish-language version of the T2D Knowledge Portal launched

Members of the Portal team were in Mexico City on January 28 to unveil a Spanish translation of the T2D Knowledge Portal, which represents a major step towards democratizing genomic research on diabetes. “We want to harness the brainpower of the entire scientific community to crack open the genes, molecules and pathways that cause type 2 diabetes”, said Jose Florez, Principal Investigator of the Portal team at the Broad Institute. Better accessibility of the Portal for Spanish speakers will help expand the brainpower that can be applied to this global health problem.

During our visit to the National Autonomous University of Mexico, Broad Institute Director Eric Lander presented a talk on the Portal. Afterwards, we led a hands-on workshop for about 30 Mexican researchers, clinicians, and students on how to use the Portal for scientific discovery.

Poster advertising our workshop
Guidance from the Carlos Slim Foundation and its Slim Initiative for Genomic Medicine in the Americas (SIGMA), one of the funders of the T2D Knowledge Portal, inspired the creation of the Spanish-language version of this web resource. “Type 2 diabetes impacts up to 14 percent of the adult population in Mexico—clearly, tackling the disease is one of our biggest public health imperatives,” said Dr. Roberto Tapia Conyer, the executive director of the Carlos Slim Foundation. “Our investment in the portal, and in SIGMA, is an investment in the health of future generations in Mexico.”

The Portal contains genetic data from more than 100,000 individuals. It currently comprises information from 28 large genetic association studies performed by several major international networks, including recent SIGMA studies that uncovered major genetic risk factors for T2D among Latin American populations.

Previously, it was difficult for anyone other than select specialists to fully access and analyze these large-scale genomic data sets. Now, the Portal provides worldwide access to these studies for all students and researchers, and the Spanish-language version expands its accessibility even further.

Lander commented, “The open sharing of data is fundamental to scientific progress, but it isn’t always easy to achieve. This portal doesn’t just help us overcome that barrier—by tapping into data on patients from around the world, it’s also going to help lead scientists to the genetic risk factors most prevalent among Latin American populations and others that have been underrepresented in large-scale genomic studies.”


Members of the visiting Portal team, happy after a successful workshop

Wednesday, October 7, 2015

Type 2 Diabetes knowledge portal goes live

Human genetic data contains valuable clues for solving the mystery of how human disease works, and the amount of data is growing as major projects around the world engage patients in large genomic studies. These projects would be most powerful if researchers worldwide collaborated with each other, sharing their findings and making results available to scientists from many disciplines including molecular biology and drug development. However, there are few frameworks for sharing the data in a way that properly credits researchers and protects patient privacy -- and as a result, today, only a relative handful of highly specialized geneticists can analyze the data or access the full results of studies.

What if anyone could?

Making that kind of sharing possible is the ultimate goal of a user-friendly web portal and knowledgebase launched October 7 at a special session of the American Society of Human Genetics annual meeting, by an international team that hopes to dramatically expand the number of researchers who can use human genetic data to study type 2 diabetes (T2D).

Research on other common diseases faces the same challenges and could therefore benefit from similar web portals.

“The investigators involved in the Accelerating Medicines Partnership in Type 2 Diabetes believe that genomic data should benefit all of humankind, as long as proper confidentiality protections are in place,” said Jose Florez, Chief of the Diabetes Unit and an Institute Member at the Broad Institute, who is one of the lead scientists in the project. “In this way the collective brainpower of all those with bright ideas, whether they come from academia, big Pharma, the biotechnology sector or government can be brought to bear on these difficult problems.”

The T2D portal (www.type2diabetesgenetics.org) -- which is open to the world free of charge -- is being developed by a team of scientists and software engineers at the Broad Institute, the University of Michigan, Oxford University, and many other collaborators as part of a worldwide scientific consortium with contributors from academia, industry, and non-profit organizations. Financial support is provided by the Accelerating Medicines Partnership in Type 2 Diabetes -- a collaboration of the National Institutes of Health, five major pharmaceutical companies, and three large non-profits -- as well as the Carlos Slim Foundation.

“Through AMP, we have an unprecedented opportunity to advance international research in type 2 diabetes,” said NIH Director Francis S. Collins in a press release. “Our hope is that this portal – and this partnership – will lead to better disease targets and a shorter, less expensive drug development process, enabling companies to get safe and effective medications to patients who need them faster.”

The current knowledgebase contains genetic and clinical data from dozens of academic institutions and partners worldwide, and researchers hope to add greatly to it over the next few years. The data already include detailed information from more than a hundred thousand participants in studies of T2D, as well as results from some of the largest genetic studies ever conducted on related traits such as obesity and glucose and insulin levels. Some of the datasets were mined by collaborating investigators in recent years to identify the genetic variants that influence T2D risk in new or unexpected ways, such as a set of variants near the gene SLC16A11 that are present in almost half of Mexicans with Native American ancestry, and another set of variants that inactivate the gene SLC30A8 and protect variant carriers against T2D.

The studies represented in the portal -- and those that researchers hope to add -- use different technologies and formats. As a result, they must be not only aggregated but also "harmonized" -- rendered compatible with each other. Those two challenges represent major work needed to make publicly accessible knowledgebases a reality for T2D or other complex traits, said Jason Flannick, a Research Associate at Massachusetts General Hospital and Technical Lead at the Broad Institute, who leads the Broad portal development team.

“There is a wealth of genetic data that has been and continues to be generated by investigators around the world that could be hugely valuable to many of the people focused on downstream biological or therapeutic research,” said Flannick. “But enabling these heterogeneous and scattered datasets to be accessible via one place, and building a sustainable model that will scale to the much larger datasets that are being produced as next-generation sequencing becomes more and more routine, requires a major investment in software and bioinformatics.”

Scientists can use the portal to easily browse results from many studies simultaneously, a process that has previously required specialized software and analysts. In addition, the portal is designed to anticipate and answer biological questions from users who may not already be familiar with statistics or genomics, enabling a broader community to learn from these results for the first time.

Users can:
  • Retrieve all genetic variants identified within a gene of interest -- learning which types of variants are present in different populations, what molecular effects are predicted for them, and how strongly those variants have been associated with T2D or related conditions in human studies.
  • Explore data for an entire chromosomal region, visualizing the effects across datasets of hundreds of variants via easy-to-use modules including a custom, web-only version of the Broad's Integrative Genomics Viewer.
  • Identify sets of variants that are associated with different combinations of traits, using a query builder that can construct custom searches across multiple datasets.
  • Investigate whether deactivating a gene in humans is likely to increase or decrease T2D risk by using an “on-the-fly” custom analysis feature that has recently been released for beta-testing.

At no point are users allowed to see protected genetic data that could be used to identify research participants -- they see only the results of analyses.

Over the coming years, the team hopes to contribute to a federated network of portals that would house many more datasets representing other studies, phenotypes, and data types -- for instance, results from large studies of gene expression and function across different human cell types and tissues, such as GTEx and ENCODE. Collaborators, including Daniel MacArthur and Benjamin Neale at the Broad as well a dedicated team led by Michael Boehnke and Gonçalo Abecasis at the University of Michigan, are already developing new methods and software tools destined for the knowledgebase.

“Combining the most extensive T2D genetic data possible with cutting-edge tools for data analysis, reporting, and visualization will accelerate genetic discovery and enable scientists to ask and answer the questions needed to identify new targets for treatment,” said Boehnke.

Ultimately, the foundation of the portal is the data; it will need collaborators from many countries and projects to succeed, said Mark McCarthy from Oxford University. Already data from over a dozen countries are in the portal. “The aspiration is to engage with researchers around the world," he said. "We will continue to reach out so that the portal contains data that are relevant to researchers and patients everywhere."