Ian N.M. Day

New genetic loci implicated in fasting glucose homeostasis and their impact on type 2 diabetes risk

Josée Dupuis, the MAGIC investigators, Claudia Langenberg et al.|Nature Genetics|2010

Cited by 2.2kOpen Access

Predicting the Functional, Molecular, and Phenotypic Consequences of Amino Acid Substitutions using Hidden Markov Models

Hashem A. Shihab, Julian Gough, D.N. Cooper et al.|Human Mutation|2012

Cited by 1.4kOpen Access

The rate at which nonsynonymous single nucleotide polymorphisms (nsSNPs) are being identified in the human genome is increasing dramatically owing to advances in whole-genome/whole-exome sequencing technologies. Automated methods capable of accurately and reliably distinguishing between pathogenic and functionally neutral nsSNPs are therefore assuming ever-increasing importance. Here, we describe the Functional Analysis Through Hidden Markov Models (FATHMM) software and server: a species-independent method with optional species-specific weightings for the prediction of the functional effects of protein missense variants. Using a model weighted for human mutations, we obtained performance accuracies that outperformed traditional prediction methods (i.e., SIFT, PolyPhen, and PANTHER) on two separate benchmarks. Furthermore, in one benchmark, we achieve performance accuracies that outperform current state-of-the-art prediction methods (i.e., SNPs&GO and MutPred). We demonstrate that FATHMM can be efficiently applied to high-throughput/large-scale human and nonhuman genome sequencing projects with the added benefit of phenotypic outcome associations. To illustrate this, we evaluated nsSNPs in wheat (Triticum spp.) to identify some of the important genetic variants responsible for the phenotypic differences introduced by intense selection during domestication. A Web-based implementation of FATHMM, including a high-throughput batch facility and a downloadable standalone package, is available at http://fathmm.biocompute.org.uk.

Common variants near MC4R are associated with fat mass, weight and risk of obesity

The Prostate, Lung, Colorectal, and Ovarian (PLCO) Cancer Screening Trial, Ruth J. F. Loos, KORA et al.|Nature Genetics|2008

Cited by 1.3kOpen Access

The UK10K project identifies rare variants in health and disease

Writing group, Klaudia Walter, Josine L. Min et al.|Nature|2015

Cited by 1.2kOpen Access

The contribution of rare and low-frequency variants to human traits is largely unexplored. Here we describe insights from sequencing whole genomes (low read depth, 7×) or exomes (high read depth, 80×) of nearly 10,000 individuals from population-based and disease collections. In extensively phenotyped cohorts we characterize over 24 million novel sequence variants, generate a highly accurate imputation reference panel and identify novel alleles associated with levels of triglycerides (APOB), adiponectin (ADIPOQ) and low-density lipoprotein cholesterol (LDLR and RGAG1) from single-marker and rare variant aggregation tests. We describe population structure and functional annotation of rare and low-frequency variants, use the data to estimate the benefits of sequencing for association studies, and summarize lessons from disease-specific collections. Finally, we make available an extensive resource, including individual-level genetic and phenotypic data and web-based tools to facilitate the exploration of association results. Low read depth sequencing of whole genomes and high read depth exomes of nearly 10,000 extensively phenotyped individuals are combined to help characterize novel sequence variants, generate a highly accurate imputation reference panel and identify novel alleles associated with lipid-related traits; in addition to describing population structure and providing functional annotation of rare and low-frequency variants the authors use the data to estimate the benefits of sequencing for association studies. This paper, combining data and initial findings from the different arms of the UK10K project, describes insights from low-read-depth sequencing of whole genomes or high-read-depth exome sequencing of nearly 10,000 individuals sampled from a range of disease collections, as well as participants from healthy population based cohorts. The authors characterize novel sequence variants, generate a highly accurate imputation reference panel and identify novel alleles associated with lipid-related traits. In addition to describing population structure and providing functional annotation of rare and low frequency variants, they use the data to estimate the benefits of sequencing for association studies.

An integrative approach to predicting the functional effects of non-coding and coding sequence variation

Hashem A. Shihab, Mark F. Rogers, Julian Gough et al.|Bioinformatics|2015

Cited by 741Open Access

Abstract Motivation: Technological advances have enabled the identification of an increasingly large spectrum of single nucleotide variants within the human genome, many of which may be associated with monogenic disease or complex traits. Here, we propose an integrative approach, named FATHMM-MKL, to predict the functional consequences of both coding and non-coding sequence variants. Our method utilizes various genomic annotations, which have recently become available, and learns to weight the significance of each component annotation source. Results: We show that our method outperforms current state-of-the-art algorithms, CADD and GWAVA, when predicting the functional consequences of non-coding variants. In addition, FATHMM-MKL is comparable to the best of these algorithms when predicting the impact of coding variants. The method includes a confidence measure to rank order predictions. Availability and implementation: The FATHMM-MKL webserver is available at: http://fathmm.biocompute.org.uk Contact: H.Shihab@bristol.ac.uk or Mark.Rogers@bristol.ac.uk or C.Campbell@bristol.ac.uk Supplementary information: Supplementary data are available at Bioinformatics online.

Is this you? Claim your profile.

Top publicationsby citations