Showing posts with label GWAS. Show all posts
Showing posts with label GWAS. Show all posts

Saturday, September 18, 2010

Accounting for Environmental Effects in GWAS identified loci

This is an excellent paper on accounting for environmental effects in genetic variants. Ruth Loos actually talked a little about this at the CSHL course I attended in July - at the time I thought most of this was already published.

Physical Activity Attenuates the Genetic Predisposition to Obesity in 20,000 Men and Women from EPIC-Norfolk Prospective Population Study

Background

We have previously shown that multiple genetic loci identified by genome-wide association studies (GWAS) increase the susceptibility to obesity in a cumulative manner. It is, however, not known whether and to what extent this genetic susceptibility may be attenuated by a physically active lifestyle. We aimed to assess the influence of a physically active lifestyle on the genetic predisposition to obesity in a large population-based study.

Methods and Findings

We genotyped 12 SNPs in obesity-susceptibility loci in a population-based sample of 20,430 individuals (aged 39–79 y) from the European Prospective Investigation of Cancer (EPIC)-Norfolk cohort with an average follow-up period of 3.6 y. A genetic predisposition score was calculated for each individual by adding the body mass index (BMI)-increasing alleles across the 12 SNPs. Physical activity was assessed using a self-administered questionnaire. Linear and logistic regression models were used to examine main effects of the genetic predisposition score and its interaction with physical activity on BMI/obesity risk and BMI change over time, assuming an additive effect for each additional BMI-increasing allele carried. Each additional BMI-increasing allele was associated with 0.154 (standard error [SE] 0.012) kg/m2 (p = 6.73×10−37) increase in BMI (equivalent to 445 g in body weight for a person 1.70 m tall). This association was significantly (pinteraction = 0.005) more pronounced in inactive people (0.205 [SE 0.024] kg/m2 [p = 3.62×10−18; 592 g in weight]) than in active people (0.131 [SE 0.014] kg/m2 [p = 7.97×10−21; 379 g in weight]). Similarly, each additional BMI-increasing allele increased the risk of obesity 1.116-fold (95% confidence interval [CI] 1.093–1.139, p = 3.37×10−26) in the whole population, but significantly (pinteraction = 0.015) more in inactive individuals (odds ratio [OR] = 1.158 [95% CI 1.118–1.199; p = 1.93×10−16]) than in active individuals (OR = 1.095 (95% CI 1.068–1.123; p = 1.15×10−12]). Consistent with the cross-sectional observations, physical activity modified the association between the genetic predisposition score and change in BMI during follow-up (pinteraction = 0.028).

Conclusions

Our study shows that living a physically active lifestyle is associated with a 40% reduction in the genetic predisposition to common obesity, as estimated by the number of risk alleles carried for any of the 12 recently GWAS-identified loci.

 

Friday, April 30, 2010

Talking Points Thursday - Putting the Complex back in Complex Diseases

So each Thursday, I thought I would blog about some current talk in genetic or anthropology or the combination of them both. As I blogged about the Havasupai Indians and the finalization of their case against Arizona State University, I thought genetics would be more appropro this week.

There is a recent article being disected on the blogoverse (or at least the one I frequent) that rehashes the argument (which is now getting old) regarding common vs. rare variants in GWAS studies. Other people have done a much better job of disecting this article and this post is incredibly informative. My argument is that everyone knows GWAS works to an extent but that it is still not identifying the missing heritabilty or the lack of variation described by significant SNPs. This is however because these SNPs represent only a small portion of the whole interaction occuring between genes that are involved in pathways. The genome is part of a biological system that is intergrated and inherently complex.

The basic idea regarding complex disease genetics is that the phenotypes involved are "complex"! This means that they are made of several genes that interact to create a protein and that they are also affected by the environment. A problem with candidate gene studies is that focus on variation within a single gene, without regard to the other genes or the biological pathways involved. So when a GWAS reports a gene to be involved in a complex disease it is only the tip of the iceberg. I'm not going to argue that GWAS is not informative, but it is an explatory statistical method. The idea proposed in the Mckellan and King article that the SNPs are not functional identfied in GWAS, totally misinterpets the method. The idea of GWAS is to identify regions of the genome that may be involved in a phenotype at higher resolution. GWAS works on the basis of linkage disequilbrium and provides 1 centimorgan region around where association occurs. This is ten-fold increase over linkage studies based on STRs. Besides, there are certain genes we know are involved in complex disease genetics because experimental work has been conducted on them in mice, rabbits, E. coli, etc., like hepatic lipase with HDL-C and variants with this gene show up in GWAS studies.

An alternative approach is to intergrate gene expression data and SNP data in order to identify the biological pathway involved in your trait of intrest. A recent paper in the American Journal of Human Genetics entitled "Integrating Pathway Analysis and Genetics of Gene Expression for Genome-Wide Studies" does a good job of describing this along with a certain type of methodology involved. One of the current problems associated with the joint type of analysis is that there is no set methodology and so different people analyze this in different ways. This paper represents a good start at this and may help in better identfying genetic variants that are impacting protein expression, which are then altering biological pathways that lead to chronic diseases. This is what will lead to a better understanding of the complexity of diseases, not denser chip set.

Friday, January 29, 2010

Friday Five #12 - Rare Variants, Linkage good or bad, Indels, biological age, and the last of the Neanderthals

So January is whizzing by here in South Texas - it is almost spring - our 2 1/2 weeks of winter are over. It has also been a pretty interesting week in Science as well as elsewhere if you want an ipad or listened to the SOTU (honestly is everything a an acronym now). So here is my current FF12 for this week.

1) Rare Variants Create Synthetic Genome-Wide Associations: One of the most cited criticisms for GWAS studies is that they ignore rare variants in the genome (generally base pairs that occur in a frequency (seen here). However, this study suggest these common variants may not be the actual cause of the disease rather they are in high linkage disequilbrium with the rarer functional variant. So these authors conducted a simulation study to suggest that this is actually the case and make the argument that in future follow up studies of these regions - either deep or complete sequencing of the region should be conducted. Here is the abstract.
Genome-wide association studies (GWAS) have now identified at least 2,000 common variants that appear associated with common diseases or related traits, hundreds of which have been convincingly replicated. It is generally thought that the associated markers reflect the effect of a nearby common (minor allele frequency >0.05) causal site, which is associated with the marker, leading to extensive resequencing efforts to find causal sites. We propose as an alternative explanation that variants much less common than the associated one may create “synthetic associations” by occurring, stochastically, more often in association with one of the alleles at the common site versus the other allele. Although synthetic associations are an obvious theoretical possibility, they have never been systematically explored as a possible explanation for GWAS findings. Here, we use simple computer simulations to show the conditions under which such synthetic associations will arise and how they may be recognized. We show that they are not only possible, but inevitable, and that under simple but reasonable genetic models, they are likely to account for or contribute to many of the recently identified signals reported in genome-wide association studies. We also illustrate the behavior of synthetic associations in real datasets by showing that rare causal mutations responsible for both hearing loss and sickle cell anemia create genome-wide significant synthetic associations, in the latter case extending over a 2.5-Mb interval encompassing scores of “blocks” of associated variants. In conclusion, uncommon or rare genetic variants can easily create synthetic associations that are credited to common variants, and this possibility requires careful consideration in the interpretation and follow up of GWAS signals.
2) Were Genome-Wide Linkage Studies a Waste of Time? Exploiting Candidate Regions Within Genome-Wide Association Studies: Prior to GWAS - the most common way to identify genes was through linkage. The only drawbacks being that they identified large regions of the genome, which made it difficult to determine single genes that may of interest to whichever phenotype being investigated and you need to use families. A particular problem with GWAS is stratification, which may occur when you are using a number of indiviuals from different populations. This article basically suggest that if a researcher uses linkage it will help avoid the stratification problem and also allow for higher resolution in the SNPs identified. Here is the abstact.
A central issue in genome-wide association (GWA) studies is assessing statistical significance while adjusting for multiple hypothesis testing. An equally important question is the statistical efficiency of the GWA design as compared to the traditional sequential approach in which genome-wide linkage analysis is followed by region-wise association mapping. Nevertheless, GWA is becoming more popular due in part to cost efficiency: commercially available 1M chips are nearly as inexpensive as a custom-designed 10 K chip. It is becoming apparent, however, that most of the on-going GWA studies with 2,000-5,000 samples are in fact underpowered. As a means to improve power, we emphasize the importance of utilizing prior information such as results of previous linkage studies via a stratified false discovery rate (FDR) control. The essence of the stratified FDR control is to prioritize the genome and maintain power to interrogate candidate regions within the GWA study. These candidate regions can be defined as, but are by no means limited to, linkage-peak regions. Furthermore, we theoretically unify the stratified FDR approach and the weighted P-value method, and we show that stratified FDR can be formulated as a robust version of weighted FDR. Finally, we demonstrate the utility of the methods in two GWA datasets: Type 2 diabetes (FUSION) and an on-going study of long-term diabetic complications (DCCT/EDIC). The methods are implemented as a user-friendly software package, SFDR. The same stratification framework can be readily applied to other type of studies, for example, using GWA results to improve the power of sequencing data analyses
3) Insertion and Deletion Processes in Recent Human History: This article looks at insertions and deletions in the human genome in order to determine if they are governed by the same biological processes. This is interesting, however, often this type of data is hard to determine in populations. They basically come to the conclusion that deletions are more deleterious than insertions (although insertion have been implicated for cystic fibrosis and other Mendelian disorders) but more importantly that they are governed by the same processes.

Background

Although insertions and deletions (indels) account for a sizable portion of genetic changes within and among species, they have received little attention because they are difficult to type, are alignment dependent and their underlying mutational process is poorly understood. A fundamental question in this respect is whether insertions and deletions are governed by similar or different processes and, if so, what these differences are.

Methodology/Principal Findings

We use published resequencing data from Seattle SNPs and NIEHS human polymorphism databases to construct a genomewide data set of short polymorphic insertions and deletions in the human genome (n = 6228). We contrast these patterns of polymorphism with insertions and deletions fixed in the same regions since the divergence of human and chimpanzee (n = 10546). The macaque genome is used to resolve all indels into insertions and deletions. We find that the ratio of deletions to insertions is greater within humans than between human and chimpanzee. Deletions segregate at lower frequency in humans, providing evidence for deletions being under stronger purifying selection than insertions. The insertion and deletion rates correlate with several genomic features and we find evidence that both insertions and deletions are associated with point mutations. Finally, we find no evidence for a direct effect of the local recombination rate on the insertion and deletion rate.

Conclusions/Significance

Our data strongly suggest that deletions are more deleterious than insertions but that insertions and deletions are otherwise generally governed by the same genomic factors.

4) An empirical comparative study on biological age estimation algorithms with an application of Work Ability Index (WAI): This article looks at different estimates of biological age in order to determine which of these is the most robust. They find the Klemura and Doubal approach is the best at estimating biological age. This differs from other estimates as it includes chronological age as a covariate in the model.
In this study, we described the characteristics of five different biological age (BA) estimation algorithms, including (i) multiple linear regression, (ii) principal component analysis, and somewhat unique methods developed by (iii) Hochschild, (iv) Klemera and Doubal, and (v) a variant of Klemera and Doubal's method. The objective of this study is to find the most appropriate method of BA estimation by examining the association between Work Ability Index (WAI) and the differences of each algorithm's estimates from chronological age (CA). The WAI was found to be a measure that reflects an individual's current health status rather than the deterioration caused by a serious dependency with the age. Experiments were conducted on 200 Korean male participants using a BA estimation system developed principally under the concept of non-invasive, simple to operate and human function-based. Using the empirical data, BA estimation as well as various analyses including correlation analysis and discriminant function analysis was performed. As a result, it had been confirmed by the empirical data that Klemera and Doubal's method with uncorrelated variables from principal component analysis produces relatively reliable and acceptable BA estimates.
5) Last Neanderthals in Europe Died out 37,000 Years Ago: A Science Daily article that suggest the last best place for Neanderthals was in the Iberian peninsula.

Friday, January 22, 2010

Friday Five #11: GWAS, Ebola, babies, and Snoop

1) Prioritizing GWAS Results: A Review of Statistical Methods and Recommendations for Their Application (Need access to read the whole article). This is an excellent review article about what to do with GWAS results or how to go about taking the next step. The authors make a good point that GWAS is the tip of the iceberg and what is really needed is a fundamental understanding of the biology involved not just the genetic variants.
Genome-wide association studies (GWAS) have rapidly become a standard method for disease gene discovery. A substantial number of recent GWAS indicate that for most disorders, only a few common variants are implicated and the associated SNPs explain only a small fraction of the genetic risk. This review is written from the viewpoint that findings from the GWAS provide preliminary genetic information that is available for additional analysis by statistical procedures that accumulate evidence, and that these secondary analyses are very likely to provide valuable information that will help prioritize the strongest constellations of results. We review and discuss three analytic methods to combine preliminary GWAS statistics to identify genes, alleles, and pathways for deeper investigations. Meta-analysis seeks to pool information from multiple GWAS to increase the chances of finding true positives among the false positives and provides a way to combine associations across GWAS, even when the original data are unavailable. Testing for epistasis within a single GWAS study can identify the stronger results that are revealed when genes interact. Pathway analysis of GWAS results is used to prioritize genes and pathways within a biological context. Following a GWAS, association results can be assigned to pathways and tested in aggregate with computational tools and pathway databases. Reviews of published methods with recommendations for their application are provided within the framework for each approach.
2) Genome Study Provides a Census of Early Humans A NYT article regarding a PNAS study that looks at the number of ancient humans around and suggest the effective population size was small, around 18,500. Basically, they looked at Alu inserts, which are a type of retroelement and are unique to primate genomes.

3) Researcher Discovers Ebola's Deadly Secret: A cool (yes I'm a nerd) Science Daily article that demonstrates how Ebola is able to fool the host cell defense mechanism and identifies the protein involved (VP35).

4) Birth Weights Fell From 1990 to 2005: Looks at a downward secular trend in the birthweight of babies. They noted that "The lower-birth-weight trend couldn't be explained by common factors such as how much weight mothers gained during pregnancy, whether the delivery was induced or by cesarean section, the amount of prenatal care, or maternal-health issues such as smoking and hypertension, researchers said". However, I wonder if they accounted for the age of the mother, which wouldn't suprise me to much if that is the case as demonstrated by this article.

5) Snoop Dogg Shocked by Genealogy Results: From Eastman's Online Genealogy Newsletter (best viewed in Firefox of Chrome). Funny but I wonder if his results will change.