Friday, January 29, 2010

Friday Five #12 - Rare Variants, Linkage good or bad, Indels, biological age, and the last of the Neanderthals

So January is whizzing by here in South Texas - it is almost spring - our 2 1/2 weeks of winter are over. It has also been a pretty interesting week in Science as well as elsewhere if you want an ipad or listened to the SOTU (honestly is everything a an acronym now). So here is my current FF12 for this week.

1) Rare Variants Create Synthetic Genome-Wide Associations: One of the most cited criticisms for GWAS studies is that they ignore rare variants in the genome (generally base pairs that occur in a frequency (seen here). However, this study suggest these common variants may not be the actual cause of the disease rather they are in high linkage disequilbrium with the rarer functional variant. So these authors conducted a simulation study to suggest that this is actually the case and make the argument that in future follow up studies of these regions - either deep or complete sequencing of the region should be conducted. Here is the abstract.
Genome-wide association studies (GWAS) have now identified at least 2,000 common variants that appear associated with common diseases or related traits, hundreds of which have been convincingly replicated. It is generally thought that the associated markers reflect the effect of a nearby common (minor allele frequency >0.05) causal site, which is associated with the marker, leading to extensive resequencing efforts to find causal sites. We propose as an alternative explanation that variants much less common than the associated one may create “synthetic associations” by occurring, stochastically, more often in association with one of the alleles at the common site versus the other allele. Although synthetic associations are an obvious theoretical possibility, they have never been systematically explored as a possible explanation for GWAS findings. Here, we use simple computer simulations to show the conditions under which such synthetic associations will arise and how they may be recognized. We show that they are not only possible, but inevitable, and that under simple but reasonable genetic models, they are likely to account for or contribute to many of the recently identified signals reported in genome-wide association studies. We also illustrate the behavior of synthetic associations in real datasets by showing that rare causal mutations responsible for both hearing loss and sickle cell anemia create genome-wide significant synthetic associations, in the latter case extending over a 2.5-Mb interval encompassing scores of “blocks” of associated variants. In conclusion, uncommon or rare genetic variants can easily create synthetic associations that are credited to common variants, and this possibility requires careful consideration in the interpretation and follow up of GWAS signals.
2) Were Genome-Wide Linkage Studies a Waste of Time? Exploiting Candidate Regions Within Genome-Wide Association Studies: Prior to GWAS - the most common way to identify genes was through linkage. The only drawbacks being that they identified large regions of the genome, which made it difficult to determine single genes that may of interest to whichever phenotype being investigated and you need to use families. A particular problem with GWAS is stratification, which may occur when you are using a number of indiviuals from different populations. This article basically suggest that if a researcher uses linkage it will help avoid the stratification problem and also allow for higher resolution in the SNPs identified. Here is the abstact.
A central issue in genome-wide association (GWA) studies is assessing statistical significance while adjusting for multiple hypothesis testing. An equally important question is the statistical efficiency of the GWA design as compared to the traditional sequential approach in which genome-wide linkage analysis is followed by region-wise association mapping. Nevertheless, GWA is becoming more popular due in part to cost efficiency: commercially available 1M chips are nearly as inexpensive as a custom-designed 10 K chip. It is becoming apparent, however, that most of the on-going GWA studies with 2,000-5,000 samples are in fact underpowered. As a means to improve power, we emphasize the importance of utilizing prior information such as results of previous linkage studies via a stratified false discovery rate (FDR) control. The essence of the stratified FDR control is to prioritize the genome and maintain power to interrogate candidate regions within the GWA study. These candidate regions can be defined as, but are by no means limited to, linkage-peak regions. Furthermore, we theoretically unify the stratified FDR approach and the weighted P-value method, and we show that stratified FDR can be formulated as a robust version of weighted FDR. Finally, we demonstrate the utility of the methods in two GWA datasets: Type 2 diabetes (FUSION) and an on-going study of long-term diabetic complications (DCCT/EDIC). The methods are implemented as a user-friendly software package, SFDR. The same stratification framework can be readily applied to other type of studies, for example, using GWA results to improve the power of sequencing data analyses
3) Insertion and Deletion Processes in Recent Human History: This article looks at insertions and deletions in the human genome in order to determine if they are governed by the same biological processes. This is interesting, however, often this type of data is hard to determine in populations. They basically come to the conclusion that deletions are more deleterious than insertions (although insertion have been implicated for cystic fibrosis and other Mendelian disorders) but more importantly that they are governed by the same processes.

Background

Although insertions and deletions (indels) account for a sizable portion of genetic changes within and among species, they have received little attention because they are difficult to type, are alignment dependent and their underlying mutational process is poorly understood. A fundamental question in this respect is whether insertions and deletions are governed by similar or different processes and, if so, what these differences are.

Methodology/Principal Findings

We use published resequencing data from Seattle SNPs and NIEHS human polymorphism databases to construct a genomewide data set of short polymorphic insertions and deletions in the human genome (n = 6228). We contrast these patterns of polymorphism with insertions and deletions fixed in the same regions since the divergence of human and chimpanzee (n = 10546). The macaque genome is used to resolve all indels into insertions and deletions. We find that the ratio of deletions to insertions is greater within humans than between human and chimpanzee. Deletions segregate at lower frequency in humans, providing evidence for deletions being under stronger purifying selection than insertions. The insertion and deletion rates correlate with several genomic features and we find evidence that both insertions and deletions are associated with point mutations. Finally, we find no evidence for a direct effect of the local recombination rate on the insertion and deletion rate.

Conclusions/Significance

Our data strongly suggest that deletions are more deleterious than insertions but that insertions and deletions are otherwise generally governed by the same genomic factors.

4) An empirical comparative study on biological age estimation algorithms with an application of Work Ability Index (WAI): This article looks at different estimates of biological age in order to determine which of these is the most robust. They find the Klemura and Doubal approach is the best at estimating biological age. This differs from other estimates as it includes chronological age as a covariate in the model.
In this study, we described the characteristics of five different biological age (BA) estimation algorithms, including (i) multiple linear regression, (ii) principal component analysis, and somewhat unique methods developed by (iii) Hochschild, (iv) Klemera and Doubal, and (v) a variant of Klemera and Doubal's method. The objective of this study is to find the most appropriate method of BA estimation by examining the association between Work Ability Index (WAI) and the differences of each algorithm's estimates from chronological age (CA). The WAI was found to be a measure that reflects an individual's current health status rather than the deterioration caused by a serious dependency with the age. Experiments were conducted on 200 Korean male participants using a BA estimation system developed principally under the concept of non-invasive, simple to operate and human function-based. Using the empirical data, BA estimation as well as various analyses including correlation analysis and discriminant function analysis was performed. As a result, it had been confirmed by the empirical data that Klemera and Doubal's method with uncorrelated variables from principal component analysis produces relatively reliable and acceptable BA estimates.
5) Last Neanderthals in Europe Died out 37,000 Years Ago: A Science Daily article that suggest the last best place for Neanderthals was in the Iberian peninsula.

No comments: