Open Access. Powered by Scholars. Published by Universities.®

Biostatistics Commons™

Open Access. Powered by Scholars. Published by Universities.®

Genetics and Genomics

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 61 - 90 of 105

Full-Text Articles in Biostatistics

Family-Based Association Studies Of Autism In Boys Via Facial-Feature Clusters, Luke Andrew Settles Jan 2017

Family-Based Association Studies Of Autism In Boys Via Facial-Feature Clusters, Luke Andrew Settles

Masters Theses

"Autism spectrum disorder (ASD) refers to a set of developmental disorders with varied attributes. Due to its substantial heterogeneity in terms of behavioral and clinical phenotypes, it is challenging to discern the genetic biomarkers behind ASD, even though the disease is known to be genetic in nature. This serves as a motivation to detect relationships between single nucleotide polymorphisms (SNPs) and a causal autism disease susceptibility locus (DSL) within more homogeneous subgroups. Recently, clinically meaningful subclassifications of ASD have been discovered utilizing facial features of prepubescent boys. Therefore, through the employment of data from 44 prepubertal Caucasian boys with ASD …


Novel Models Of Visual Topographic Map Alignment In The Superior Colliculus., Ruben A Tikidji-Hamburyan, Tarek A El-Ghazawi, Jason W. Triplett Dec 2016

Novel Models Of Visual Topographic Map Alignment In The Superior Colliculus., Ruben A Tikidji-Hamburyan, Tarek A El-Ghazawi, Jason W. Triplett

Pediatrics Faculty Publications

The establishment of precise neuronal connectivity during development is critical for sensing the external environment and informing appropriate behavioral responses. In the visual system, many connections are organized topographically, which preserves the spatial order of the visual scene. The superior colliculus (SC) is a midbrain nucleus that integrates visual inputs from the retina and primary visual cortex (V1) to regulate goal-directed eye movements. In the SC, topographically organized inputs from the retina and V1 must be aligned to facilitate integration. Previously, we showed that retinal input instructs the alignment of V1 inputs in the SC in a manner dependent on …


High-Throughput Allele-Specific Expression Across 250 Environmental Conditions, Gregory A. Moyerbrailean, Allison L. Richards, Daniel Kurtz, Cynthia A. Kalita, Gordon O. Davis, Chris T. Harvey, Adnan Alazizi, Donovan Watza, Yoram Sorokin, Nancy J. Hauff, Xiang Zhou, Xiaoquan Wen, Roger Pique-Regi, Francesca Luca Oct 2016

High-Throughput Allele-Specific Expression Across 250 Environmental Conditions, Gregory A. Moyerbrailean, Allison L. Richards, Daniel Kurtz, Cynthia A. Kalita, Gordon O. Davis, Chris T. Harvey, Adnan Alazizi, Donovan Watza, Yoram Sorokin, Nancy J. Hauff, Xiang Zhou, Xiaoquan Wen, Roger Pique-Regi, Francesca Luca

Center for Molecular Medicine and Genetics

Gene-by-environment (GxE) interactions determine common disease risk factors and biomedically relevant complex traits. However, quantifying how the environment modulates genetic effects on human quantitative phenotypes presents unique challenges. Environmental covariates are complex and difficult to measure and control at the organismal level, as found in GWAS and epidemiological studies. An alternative approach focuses on the cellular environment using in vitro treatments as a proxy for the organismal environment. These cellular environments simplify the organism-level environmental exposures to provide a tractable influence on subcellular phenotypes, such as gene expression. Expression quantitative trait loci (eQTL) mapping studies identified GxE interactions in response …


Causal Effect Estimation In Sequencing Studies: A Bayesian Method To Account For Confounder Adjustment Uncertainty, Chi Wang, Jinpeng Liu, David W. Fardo Oct 2016

Causal Effect Estimation In Sequencing Studies: A Bayesian Method To Account For Confounder Adjustment Uncertainty, Chi Wang, Jinpeng Liu, David W. Fardo

Biostatistics Faculty Publications

Estimating the causal effect of a single nucleotide variant (SNV) on clinical phenotypes is of interest in many genetic studies. The effect estimation may be confounded by other SNVs as a result of linkage disequilibrium as well as demographic and clinical characteristics. Because a large number of these other variables, which we call potential confounders, are collected, it is challenging to select and adjust for the variables that truly confound the causal effect. The Bayesian adjustment for confounding (BAC) method has been proposed as a general method to estimate the average causal effect in the presence of a large number …


Comparing Performance Of Non-Tree-Based And Tree-Based Association Mapping Methods, Katherine L. Thompson, David W. Fardo Oct 2016

Comparing Performance Of Non-Tree-Based And Tree-Based Association Mapping Methods, Katherine L. Thompson, David W. Fardo

Statistics Faculty Publications

A central goal in the biomedical and biological sciences is to link variation in quantitative traits to locations along the genome (single nucleotide polymorphisms). Sequencing technology has rapidly advanced in recent decades, along with the statistical methodology to analyze genetic data. Two classes of association mapping methods exist: those that account for the evolutionary relatedness among individuals, and those that ignore the evolutionary relationships among individuals. While the former methods more fully use implicit information in the data, the latter methods are more flexible in the types of data they can handle. This study presents a comparison of the 2 …


Weighted-Samgsr: Combining Significance Analysis Of Microarray-Gene Set Reduction Algorithm With Pathway Topology-Based Weights To Select Relevant Genes, Suyan Tian, Howard H. Chang, Chi Wang Sep 2016

Weighted-Samgsr: Combining Significance Analysis Of Microarray-Gene Set Reduction Algorithm With Pathway Topology-Based Weights To Select Relevant Genes, Suyan Tian, Howard H. Chang, Chi Wang

Biostatistics Faculty Publications

Background: It has been demonstrated that a pathway-based feature selection method that incorporates biological information within pathways during the process of feature selection usually outperforms a gene-based feature selection algorithm in terms of predictive accuracy and stability. Significance analysis of microarray-gene set reduction algorithm (SAMGSR), an extension to a gene set analysis method with further reduction of the selected pathways to their respective core subsets, can be regarded as a pathway-based feature selection method.

Methods: In SAMGSR, whether a gene is selected is mainly determined by its expression difference between the phenotypes, and partially by the number of pathways to …


Identification Of Biomarkers For The Overall Survival Of Ovarian Cancer Patients, Kristi Mai May 2016

Identification Of Biomarkers For The Overall Survival Of Ovarian Cancer Patients, Kristi Mai

Graduate Theses and Dissertations

Rapid advance in sequencing technology has led to genome-wide analysis of genetic and epigenetic features simultaneously, making it possible to understand the biological mechanisms underlying cancer initiation and progression. However, how to identify important prognostic features poses a great challenge for both statistical modeling and computing. In this thesis, a network-based approach is applied to the Cancer Genome Atlas (TCGA) ovarian cancer data to identify important genes related to the overall survival of ovarian cancer patients. In the first step, a stepwise correlation-based selector is used to reduce the dimensionality of TCGA data, by filtering out a large number of …


Conditional Screening For Ultra-High Dimensional Covariates With Survival Outcomes, Hyokyoung Grace Hong, Jian Kang, Yi Li Mar 2016

Conditional Screening For Ultra-High Dimensional Covariates With Survival Outcomes, Hyokyoung Grace Hong, Jian Kang, Yi Li

The University of Michigan Department of Biostatistics Working Paper Series

Identifying important biomarkers that are predictive for cancer patients' prognosis is key in gaining better insights into the biological influences on the disease and has become a critical component of precision medicine. The emergence of large-scale biomedical survival studies, which typically involve excessive number of biomarkers, has brought high demand in designing efficient screening tools for selecting predictive biomarkers. The vast amount of biomarkers defies any existing variable selection methods via regularization. The recently developed variable screening methods, though powerful in many practical setting, fail to incorporate prior information on the importance of each biomarker and are less powerful in …


Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang Feb 2016

Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang

COBRA Preprint Series

Non-negative matrix factorization (NMF) is a widely used machine learning algorithm for dimension reduction of large-scale data. It has found successful applications in a variety of fields such as computational biology, neuroscience, natural language processing, information retrieval, image processing and speech recognition. In bioinformatics, for example, it has been used to extract patterns and profiles from genomic and text-mining data as well as in protein sequence and structure analysis. While the scientific performance of NMF is very promising in dealing with high dimensional data sets and complex data structures, its computational cost is high and sometimes could be critical for …


Models For Hsv Shedding Must Account For Two Levels Of Overdispersion, Amalia Magaret Jan 2016

Models For Hsv Shedding Must Account For Two Levels Of Overdispersion, Amalia Magaret

UW Biostatistics Working Paper Series

We have frequently implemented crossover studies to evaluate new therapeutic interventions for genital herpes simplex virus infection. The outcome measured to assess the efficacy of interventions on herpes disease severity is the viral shedding rate, defined as the frequency of detection of HSV on the genital skin and mucosa. We performed a simulation study to ascertain whether our standard model, which we have used previously, was appropriately considering all the necessary features of the shedding data to provide correct inference. We simulated shedding data under our standard, validated assumptions and assessed the ability of 5 different models to reproduce the …


Importance Of Hereditary And Selected Environmental Risk Factors In The Etiology Of Inflammatory Breast Cancer: A Case-Comparison Study., Roxana Moslehi, Elizabeth Freedman, Nur Zeinomar, Carmela Veneroso, Paul H. Levine Jan 2016

Importance Of Hereditary And Selected Environmental Risk Factors In The Etiology Of Inflammatory Breast Cancer: A Case-Comparison Study., Roxana Moslehi, Elizabeth Freedman, Nur Zeinomar, Carmela Veneroso, Paul H. Levine

Epidemiology Faculty Publications

BACKGROUND: To assess the importance of heredity in the etiology of inflammatory breast cancer (IBC), we compared IBC patients to several carefully chosen comparison groups with respect to the prevalence of first-degree family history of breast cancer.

METHODS: IBC cases (n = 141) were compared to non-inflammatory breast cancer cases (n = 178) ascertained through George Washington University (GWU) with respect to the prevalence of first-degree family history of breast cancer and selected environmental/lifestyle risk factors for breast cancer. Similar comparisons were conducted with subjects from three case-control studies: breast cancer cases (n = 1145) and unaffected controls (n = …


Sample Size Estimation For Genomics Experiments With Dependent End Points, Desmond Koomson Jan 2016

Sample Size Estimation For Genomics Experiments With Dependent End Points, Desmond Koomson

Open Access Theses & Dissertations

In typical genomics studies involving numerous association tests of gene mutations with a disease, error rate control via multiplicity adjustment is paramount because even if all genes were to be non-differentially associated, we would still make some false positives. Many methods exist that incorporate the control of multiplicity for normally distributed endpoints in sample size estimation, but none addresses the issue for non-normally correlated endpoints.

One common practice in the literature is to assume an equal correlation among all differentially associated or expressed genes, thereby using the generalized binomial or beta-binomial model to compute the comparison-wise power of detecting these …


Meta-Analysis Of Genome-Wide Association Studies With Correlated Individuals: Application To The Hispanic Community Health Study/Study Of Latinos (Hchs/Sol), Tamar Sofer, John R. Shaffer, Misa Graff, Qibin Qi, Adrienne M. Stilp, Stephanie M. Gogarten, Kari E. North, Carmen R. Isasi, Cathy C. Laurie, Adam A. Szpiro Nov 2015

Meta-Analysis Of Genome-Wide Association Studies With Correlated Individuals: Application To The Hispanic Community Health Study/Study Of Latinos (Hchs/Sol), Tamar Sofer, John R. Shaffer, Misa Graff, Qibin Qi, Adrienne M. Stilp, Stephanie M. Gogarten, Kari E. North, Carmen R. Isasi, Cathy C. Laurie, Adam A. Szpiro

UW Biostatistics Working Paper Series

Investigators often meta-analyze multiple genome-wide association studies (GWASs) to increase the power to detect associations of single nucleotide polymorphisms (SNPs) with a trait. Meta-analysis is also performed within a single cohort that is stratified by, e.g., sex or ancestry group. Having correlated individuals among the strata may complicate meta-analyses, limit power, and inflate Type 1 error. For example, in the Hispanic Community Health Study/Study of Latinos (HCHS/SOL), sources of correlation include genetic relatedness, shared household, and shared community. We propose a novel mixed-effect model for meta-analysis, “MetaCor", which accounts for correlation between stratum-specific effect estimates. Simulations show that MetaCor controls …


Germline Mutation Detection In Next Generation Sequencing Data And Tp53 Mutation Carrier Probability Estimation For Li-Fraumeni Syndrome, Gang Peng Aug 2015

Germline Mutation Detection In Next Generation Sequencing Data And Tp53 Mutation Carrier Probability Estimation For Li-Fraumeni Syndrome, Gang Peng

Dissertations and Theses (Open Access)

Next generation sequencing technology has been widely used in genomic analysis, but its application has been compromised by the missing true variants, especially when these variants are rare. We proposed a family-based variant calling method, FamSeq, integrating Mendelian transmission information with de novo mutation and sequencing data to improve the variant calling accuracy. We investigated the factors impacting the improvement of family-based variant calling in simulation data and validated it in real sequencing data. In both simulation and real data, FamSeq works better than the single individual based method.

In FamSeq, we implemented four different methods for the Mendelian genetic …


Genetics Of Obesity In Starr County, Texas Mexican Americans, Heather M. Highland May 2015

Genetics Of Obesity In Starr County, Texas Mexican Americans, Heather M. Highland

Dissertations and Theses (Open Access)

Currently, over two-thirds of Americans are classified as over-weight or obese. Obesity increases risk for many other diseases including type 2 diabetes, heart disease, stroke, and cancer, making obesity the largest public health problem in America and most other Westernized nations. Hispanics have a higher rate of both obesity and type 2 diabetes, making them a particularly interesting population in which to study obesity. For the last 33 years, the Starr County Health Studies has collected an array of phenotypes and biological samples from residents of Starr County, along Texas-Mexico border. This study includes 825 subjects who were not known …


Spectral Gene Set Enrichment (Sgse), H Robert Frost, Zhigang Li, Jason H. Moore Mar 2015

Spectral Gene Set Enrichment (Sgse), H Robert Frost, Zhigang Li, Jason H. Moore

Dartmouth Scholarship

Gene set testing is typically performed in a supervised context to quantify the association between groups of genes and a clinical phenotype. In many cases, however, a gene set-based interpretation of genomic data is desired in the absence of a phenotype variable. Although methods exist for unsupervised gene set testing, they predominantly compute enrichment relative to clusters of the genomic variables with performance strongly dependent on the clustering algorithm and number of clusters. We propose a novel method, spectral gene set enrichment (SGSE), for unsupervised competitive testing of the association between gene sets and empirical data sources. SGSE first computes …


Proof-Of-Concept Of Environmental Dna Tools For Atlantic Sturgeon Management, Jameson Hinkle Jan 2015

Proof-Of-Concept Of Environmental Dna Tools For Atlantic Sturgeon Management, Jameson Hinkle

Theses and Dissertations

Abstract

The Atlantic Sturgeon (Acipenser oxyrinchus oxyrinchus, Mitchell) is an anadromous species that spawns in tidal freshwater rivers from Canada to Florida. Overfishing, river sedimentation and alteration of the river bottom have decreased Atlantic Sturgeon populations, and NOAA lists the species as endangered. Ecologists sometimes find it difficult to locate individuals of a species that is rare, endangered or invasive. The need for methods less invasive that can create more resolution of cryptic species presence is necessary. Environmental DNA (eDNA) is a non-invasive means of detecting rare, endangered, or invasive species by isolating nuclear or mitochondrial DNA (mtDNA) from the …


A Meta-Analysis Of Association Between One-Carbon Metabolism Gene Polymorphisms And Risk Of Prostate Cancer, Mahmood Tazari Jan 2015

A Meta-Analysis Of Association Between One-Carbon Metabolism Gene Polymorphisms And Risk Of Prostate Cancer, Mahmood Tazari

Walden Dissertations and Doctoral Studies

Prostate cancer is the most common cancer among men. The purpose of this quantitative, meta-analysis study was to examine one-carbon metabolism gene polymorphisms in a group of genes to determine their association with prostate cancer risk. The genetic epidemiology theory provided the framework for the study. The data collected were from published articles. From over 2,800 individual studies, 20 articles were retained for results and data abstraction, following the title, abstract screen, and full text screening in the second phase. The data were analyzed by a meta-analysis statistical method, combining the results from selected studies to estimate the overall association. …


Testing Gene-Environment Interactions In The Presence Of Measurement Error, Chongzhi Di, Li Hsu, Charles Kooperberg, Alex Reiner, Ross Prentice Nov 2014

Testing Gene-Environment Interactions In The Presence Of Measurement Error, Chongzhi Di, Li Hsu, Charles Kooperberg, Alex Reiner, Ross Prentice

UW Biostatistics Working Paper Series

Complex diseases result from an interplay between genetic and environmental risk factors, and it is of great interest to study the gene-environment interaction (GxE) to understand the etiology of complex diseases. Recent developments in genetics field allows one to study GxE systematically. However, one difficulty with GxE arises from the fact that environmental exposures are often measured with error. In this paper, we focus on testing GxE when the environmental exposure E is subject to measurement error. Surprisingly, contrast to the well-established results that the naive test ignoring measurement error is valid in testing the main effects, we find that …


Genetic Predictors Of Metabolic Side Effects Of Diuretic Therapy, Jorge L. Del Aguila Aug 2014

Genetic Predictors Of Metabolic Side Effects Of Diuretic Therapy, Jorge L. Del Aguila

Dissertations and Theses (Open Access)

Thiazide diuretics are a recommended first-line monotherapy for hypertension (i.e.SBP>140 mmHg or DBP>90 mmHg). Even so, diuretics are associated with adverse metabolic side effects, such as hyperlipidemia, hyperglycemia and hypokalemia which increase the risk of developing type II diabetes. This thesis used three analytical strategies to identify and quantify genetic factors that contribute to the development of adverse metabolic effects due to thiazide diuretic treatment. I performed a genome-wide association study (GWAS) and meta-analysis of the change in fasting plasma glucose and triglycerides in response to HCTZ from two different clinical trials: the Pharmacogenomic Evaluation of Antihypertensive Responses …


The Association Between The Il-1 Pathway, Isaac C. Wun May 2014

The Association Between The Il-1 Pathway, Isaac C. Wun

Dissertations and Theses (Open Access)

Cutaneous malignant melanoma (CMM) is a potentially lethal malignancy that warrants attention and further research, as it is known to that there is an increasing rate of incidence in theUnited States, and it is also known that exposure to UV light is its most crucial risk factor, and family history of melanoma is also an important risk factor. Melanoma is an aggressive and lethal cancer in humans. There are an estimated new 132,000 melanoma cases annually worldwide, and the trend has doubled in the past 20 years. However, attempts to treat melanoma have encountered considerable resistance and remained ineffective. The …


Five Fundamental Gaps In Nature-Nurture Science, Peter J. Taylor Mar 2014

Five Fundamental Gaps In Nature-Nurture Science, Peter J. Taylor

Working Papers on Science in a Changing World

Difficulties identifying causally relevant genetic variants underlying patterns of human variation have been given competing interpretations. The debate is illuminated in this article by drawing attention to the issue of underlying heterogeneity—the possibility that genetic and environmental factors or entities underlying a trait are heterogeneous—as well as four other fundamental gaps in the methods and interpretation of classical quantitative genetics: "Genetic" and "environmental" fractions of variation in traits are distinct from measurable genetic and environmental factors underlying the traits’ development; Standard formulas for partitioning variation in human traits are unreliable; Methods for translation from fractions of variation to measurable …


Set-Based Tests For Genetic Association In Longitudinal Studies, Zihuai He, Min Zhang, Seunggeun Lee, Jennifer A. Smith, Xiuqing Guo, Walter Palmas, Sharon L.R. Kardia, Ana V. Diez Roux, Bhramar Mukherjee Jan 2014

Set-Based Tests For Genetic Association In Longitudinal Studies, Zihuai He, Min Zhang, Seunggeun Lee, Jennifer A. Smith, Xiuqing Guo, Walter Palmas, Sharon L.R. Kardia, Ana V. Diez Roux, Bhramar Mukherjee

The University of Michigan Department of Biostatistics Working Paper Series

Genetic association studies with longitudinal markers of chronic diseases (e.g., blood pressure, body mass index) provide a valuable opportunity to explore how genetic variants affect traits over time by utilizing the full trajectory of longitudinal outcomes. Since these traits are likely influenced by the joint effect of multiple variants in a gene, a joint analysis of these variants considering linkage disequilibrium (LD) may help to explain additional phenotypic variation. In this article, we propose a longitudinal genetic random field model (LGRF), to test the association between a phenotype measured repeatedly during the course of an observational study and a set …


Methods For Integrative Analysis Of Genomic Data, Paul Manser Jan 2014

Methods For Integrative Analysis Of Genomic Data, Paul Manser

Theses and Dissertations

In recent years, the development of new genomic technologies has allowed for the investigation of many regulatory epigenetic marks besides expression levels, on a genome-wide scale. As the price for these technologies continues to decrease, study sizes will not only increase, but several different assays are beginning to be used for the same samples. It is therefore desirable to develop statistical methods to integrate multiple data types that can handle the increased computational burden of incorporating large data sets. Furthermore, it is important to develop sound quality control and normalization methods as technical errors can compound when integrating multiple genomic …


Genetic Susceptibility To Type 2 Diabetes: A Global Meta-Analysis Studying The Genetic Differences In Tunisian Populations, Rym Berhouma, S. Kouidhi, M. Ammar, H. Abid, T. Baroudi, H. Ennafaa, A. Benammar-Elgaaied Aug 2012

Genetic Susceptibility To Type 2 Diabetes: A Global Meta-Analysis Studying The Genetic Differences In Tunisian Populations, Rym Berhouma, S. Kouidhi, M. Ammar, H. Abid, T. Baroudi, H. Ennafaa, A. Benammar-Elgaaied

Human Biology Open Access Pre-Prints

The present study is the first meta-analysis to evaluate type 2 diabetes (T2D) - associated polymorphisms in cohorts originated from several Tunisian regions. In fact, we evaluated the effect of seven polymorphisms in the following genes; PPARg ( Pro12Ala), TNFα (-308A/G), ENPP1(K121Q), TCF7L2(rs7903146 C/T), MTHFR( C677T), ACE(I/D), CAPN10(3R/2R) on T2D risk, through a meta-analysis combining data of previous studies performed on Tunisian populations originating from the north, centre or south of the country. R statistics version 2.12.1 software was used to estimate the heterogeneity between studies. Pooled ORs were computed by the fixed-effects method of Mantel-Haenszel if no heterogeneity between …


Dna Methylation Arrays As Surrogate Measures Of Cell Mixture Distribution, Eugene Houseman, William P. Accomando, Devin C. Koestler, Brock C. Christensen, Carmen J. Marsit May 2012

Dna Methylation Arrays As Surrogate Measures Of Cell Mixture Distribution, Eugene Houseman, William P. Accomando, Devin C. Koestler, Brock C. Christensen, Carmen J. Marsit

Dartmouth Scholarship

There has been a long-standing need in biomedical research for a method that quantifies the normally mixed composition of leukocytes beyond what is possible by simple histological or flow cytometric assessments. The latter is restricted by the labile nature of protein epitopes, requirements for cell processing, and timely cell analysis. In a diverse array of diseases and following numerous immune-toxic exposures, leukocyte composition will critically inform the underlying immuno-biology to most chronic medical conditions. Emerging research demonstrates that DNA methylation is responsible for cellular differentiation, and when measured in whole peripheral blood, serves to distinguish cancer cases from controls.


Why Odds Ratio Estimates Of Gwas Are Almost Always Close To 1.0, Yutaka Yasui May 2012

Why Odds Ratio Estimates Of Gwas Are Almost Always Close To 1.0, Yutaka Yasui

COBRA Preprint Series

“Missing heritability” in genome-wide association studies (GWAS) refers to the seeming inability for GWAS data to capture the great majority of genetic causes of a disease in comparison to the known degree of heritability for the disease, in spite of GWAS’ genome-wide measures of genetic variations. This paper presents a simple mathematical explanation for this phenomenon, assuming that the heritability information exists in GWAS data. Specifically, it focuses on the fact that the great majority of association measures (in the form of odds ratios) from GWAS are consistently close to the value that indicates no association, explains why this occurs, …


Estimation Of A Non-Parametric Variable Importance Measure Of A Continuous Exposure, Chambaz Antoine, Pierre Neuvial, Mark J. Van Der Laan Oct 2011

Estimation Of A Non-Parametric Variable Importance Measure Of A Continuous Exposure, Chambaz Antoine, Pierre Neuvial, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

We define a new measure of variable importance of an exposure on a continuous outcome, accounting for potential confounders. The exposure features a reference level x0 with positive mass and a continuum of other levels. For the purpose of estimating it, we fully develop the semi-parametric estimation methodology called targeted minimum loss estimation methodology (TMLE) [van der Laan & Rubin, 2006; van der Laan & Rose, 2011]. We cover the whole spectrum of its theoretical study (convergence of the iterative procedure which is at the core of the TMLE methodology; consistency and asymptotic normality of the estimator), practical implementation, simulation …


Gene By Bmi Interactions Influencing C-Reactive Protein Levels In European-Americans, Sarah Tudor Aug 2011

Gene By Bmi Interactions Influencing C-Reactive Protein Levels In European-Americans, Sarah Tudor

Dissertations and Theses (Open Access)

C-Reactive Protein (CRP) is a biomarker indicating tissue damage, inflammation, and infection. High-sensitivity CRP (hsCRP) is an emerging biomarker often used to estimate an individual’s risk for future coronary heart disease (CHD). hsCRP levels falling below 1.00 mg/l indicate a low risk for developing CHD, levels ranging between 1.00 mg/l and 3.00 mg/l indicate an elevated risk, and levels exceeding 3.00 mg/l indicate high risk. Multiple Genome-Wide Association Studies (GWAS) have identified a number of genetic polymorphisms which influence CRP levels. SNPs implicated in such studies have been found in or near genes of interest including: CRP, APOE, APOC, IL-6, …


Minimum Description Length Measures Of Evidence For Enrichment, Zhenyu Yang, David R. Bickel Dec 2010

Minimum Description Length Measures Of Evidence For Enrichment, Zhenyu Yang, David R. Bickel

COBRA Preprint Series

In order to functionally interpret differentially expressed genes or other discovered features, researchers seek to detect enrichment in the form of overrepresentation of discovered features associated with a biological process. Most enrichment methods treat the p-value as the measure of evidence using a statistical test such as the binomial test, Fisher's exact test or the hypergeometric test. However, the p-value is not interpretable as a measure of evidence apart from adjustments in light of the sample size. As a measure of evidence supporting one hypothesis over the other, the Bayes factor (BF) overcomes this drawback of the p-value but lacks …