Showing posts with label Statistics. Show all posts
Showing posts with label Statistics. Show all posts

Monday, August 16, 2010

Fall Meeting Abstracts

Here are the Abstracts for two meetings I will be attending this Fall.

International Genetic Epidemiology Society - Boston
Genomic, transcriptomic, and pathway-based analysis of Carotid Intima-Media Thickness
Detection of subclinical atherosclerosis through measurement of carotid intima-media thickness (IMT) is an important tool for assessment of cardiovascular disease risk. Previous GWAS of IMT have not produced any significant associations. This study used GWAS results to investigate biological pathways influencing IMT phenotypes in 750 Mexican American participants from the San Antonio Family Heart Study. A GWAS (using Illumina’s HumanHap 550K BeadChip) was performed using internal and common carotid IMT near and far wall measurements under an additive measured genotype association model. As with past GWAS, no SNP was significant after correction for multiple testing. Genes within ranges 10kb, 25kb, 50kb, 100kb,150 kb, and 250kb upstream and downstream of marginally significant (p <0.01) SNPs were entered into Ariadne Pathway Studio to identify relevant biological pathways and to investigate how the assignment of intergenic SNPs to genes might impact these results. The WNT-signaling (p=1.04 x 10-7) pathway remained significant after correction for multiple testing. We then investigated the relationship between lympochyte RNA expression data and IMT. The top gene was glutamate-cysteine ligase, catalytic subunit (GCLC), whose transcript is significantly correlated (p=7.8 x 10-5) with internal carotid IMT. We conclude that through the use of transcript data and the investigation of biological pathways that we are able to detect genetic signals not identified by GWAS alone.
American Society of Human Genetics - Washington D.C.
Bayesian analysis of hepatic lipase (LIPC) genetic variation suggests different variants influence HDL-C level and HDL particle size
The gene for hepatic lipase (LIPC) is known to have a major role in lipoprotein metabolism and has been implicated as a regulator of HDL-C concentration and HDL particle size reduction through the reverse cholesterol transport pathway. However, the relationship between LIPC and HDL-C levels and HDL particle size phenotypes remains unclear. One LIPC promoter variant, -514C>T (rs1800588), has shown association with HDL-C and related phenotypes in a number of studies and is presumed to be functional. This study investigated the biological relationship between LIPC genetic variation and HDL-C and five HDL size measures, median diameter of particles containing apoA1 (A1), apoA2 (A2), unesterified cholesterol (UC), esterified cholesterol (EC), and ΔHDL, the difference in fractional absorbance patterns between large and small lipoprotein particles. We deep sequenced 28.7 kb coding and conserved non-coding regions of LIPC in 182 founders from 42 extended pedigrees and examined the resulting 625 single nucleotide polymorphism (SNP) genotypes in 1,336 Mexican American participants from the San Antonio Family Heart Study. Genetic association was assessed with measured genotype analysis (MGA) and SNPs with nominal p-values (p <0.1) were used in Bayesian Quantitative Trait Nucleotide (BQTN) analyses. SNPs in high LD (r2>0.9) were grouped, with the single member with the lowest MGA p-value representing the isocorrelated redundant variant (IRV) set. BQTN identified 18 (five novel) potentially functional SNPs with substantial posterior probabilities (pps > 0.50) associated with one or more of the six investigated phenotypes. Seven SNPs demonstrated pps >0.95: two for HDL-C levels (rs1077834, rs16940299), two for A1 (rs375372, rs12914232), one for A2 and UC (rs11071387), and two for ΔHDL (rs485538, rs8023503). The SNP rs11071387 was associated with four size phenotypes (A1, A2, UC, EC). The promoter SNP, rs1077834, is in an IRV set with three other promoter SNPs, rs1800588, -763A>G (rs10077835) and -250G>A (rs2070895). Promoter variants did not demonstrate substantial pps with any of the five HDL-C size measures. Minor allele frequency ranged from 0.0008 for +106481 G>A (HDL-EC, HDL-UC) to 0.4596 for rs375372 (A1). This comprehensive investigation of LIPC genetic variation provides strong evidence that variants in addition to rs1800588 influence HDL-C levels and also suggests that different LIPC variants influence HDL-C levels and HDL particle size differentiation.

 

Wednesday, October 3, 2007

BioMed Central | Full text | Compression-based classification of biological sequences and structures via the Universal Similarity Metric: experimental assessment

BMC Bioinfomatics has an interesting article on alternatives to aligning sequences with genomic data.

BioMed Central Full text Compression-based classification of biological sequences and structures via the Universal Similarity Metric: experimental assessment

Compression-based classification of biological sequences and structures via the Universal Similarity Metric: experimental assessment.
Paolo Ferragina1 , Raffaele Giancarlo2 , Valentina Greco2 , Giovanni Manzini3 and Gabriel Valiente4

Background
Similarity of sequences is a key mathematical notion for Classification and Phylogenetic studies in Biology. It is currently primarily handled using alignments. However, the alignment methods seem inadequate for post-genomic studies since they do not scale well with data set size and they seem to be confined only to genomic and proteomic sequences. Therefore, alignment-free similarity measures are actively pursued. Among those, USM (Universal Similarity Metric) has gained prominence. It is based on the deep theory of Kolmogorov Complexity and universality is its most novel striking feature. Since it can only be approximated via data compression, USM is a methodology rather than a formula quantifying the similarity of two strings. Three approximations of USM are available, namely UCD (Universal Compression Dissimilarity), NCD (Normalized Compression Dissimilarity) and CD (Compression Dissimilarity). Their applicability and robustness is tested on various data sets yielding a first massive quantitative estimate that the USM methodology and its approximations are of value. Despite the rich theory developed around USM, its experimental assessment has limitations: only a few data compressors have been tested in conjunction with USM and mostly at a qualitative level, no comparison among UCD, NCD and CD is available and no comparison of USM with existing methods, both based on alignments and not, seems to be available.

Tuesday, August 14, 2007

US Health 2006

On Sunday Yahoo news had an article on Life Expectancy in the US and how we have slipped from 11th to 42nd on the list over the last 20 years. While 77.9 years is a record I became curious as to what the other countries in the list were and I came across the CDC National Health Statistics for 2006 (can be accessed here). This is a data heavy report, while light on the analytical side, it has a number of interesting highlights. Below are some of those ones I found interesting.

1) Large disparities in infant mortality rates among racial and ethnic groups continue to exist. In 2003, infant mortality rates were highest for infants of non-Hispanic black mothers (13.6 deaths per 1,000 live births), American Indian mothers (8.7 per 1,000), and Puerto Rican mothers (8.2 per 1,000); and lowest for infants of Cuban mothers (4.6 per 1,000 live births) and Asian or Pacific Islander mothers (4.8 per 1,000) (Table 19).
(2) From 1950 to 2005, the total resident population of the United States increased from 151 million to 296 million, representing an average annual growth rate of 1.2% (Figure 1). During the same period, the population 65 years of age and over grew on average 2.0% per year, increasing from 12 to 37 million persons. The population 75 years of age and over grew the fastest (on average, 2.8% per year),increasing from 4 to 18 million persons.

Projections indicate that the rate of growth for the total population from now to 2050 will be slower, but older age groups will continue to grow more rapidly than the total population (1). By 2029, all of the baby boomers (those born in the post World War II period 1946–1964) will be age 65 years and over. As a result, the population age 65–74 years will increase from 6% to 10% of the total population between 2005 and 2030 (data table for Figure 1). As the baby boomers age, the population 75 years and over will also rise from 6% to 9% of the population by 2030 and continue to grow to 12% in 2050. By 2040 the population age 75 years and over will exceed the population 65–74 years of age.

(3) In 2004, the United States spent 16% (up from 14% in 2000) of its Gross Domestic Product (GDP) on health care, a greater share than any other developed country for which data are collected by the Organisation of Economic Co-operation and Development (Figure 8 and Tables 119 and 120).

(4) In 2003, the age-adjusted death rate for heart disease, the leading cause of death, was 60% lower than the rate in 1950 (Table 36). The age-adjusted death rate for stroke, the third leading cause of death, declined 70% since 1950 (Table 37).Heart disease and stroke mortality are associated with risk factors such as high cholesterol, high blood pressure, smoking, and dietary factors. Other important factors include socioeconomic status, obesity, and physical inactivity. Factors contributing to the decline in heart disease and stroke mortality include better control of risk factors, improved access to early detection, and better treatment and care, including new drugs and expanded uses for existing drugs (1).

(5) In 2003, 96% of persons 65 years of age and over in the civilian noninstitutionalized population reported medical expenses that averaged about $8,210 per person with expense. Nineteen percent of expenses were paid out-of-pocket, 16% by private insurance, and 63% by public programs (primarily Medicare and Medicaid) (Tables 125 and 126).

Just a few facts to chew on.