Open Access. Powered by Scholars. Published by Universities.®

Biostatistics Commons™

Open Access. Powered by Scholars. Published by Universities.®

2010

Discipline
Institution
Keyword
Publication
Publication Type

Articles 1 - 30 of 52

Full-Text Articles in Biostatistics

Minimum Description Length Measures Of Evidence For Enrichment, Zhenyu Yang, David R. Bickel Dec 2010

Minimum Description Length Measures Of Evidence For Enrichment, Zhenyu Yang, David R. Bickel

COBRA Preprint Series

In order to functionally interpret differentially expressed genes or other discovered features, researchers seek to detect enrichment in the form of overrepresentation of discovered features associated with a biological process. Most enrichment methods treat the p-value as the measure of evidence using a statistical test such as the binomial test, Fisher's exact test or the hypergeometric test. However, the p-value is not interpretable as a measure of evidence apart from adjustments in light of the sample size. As a measure of evidence supporting one hypothesis over the other, the Bayes factor (BF) overcomes this drawback of the p-value but lacks …


Predicting Treatment Efficacy Via Quantitative Mri: A Bayesian Joint Model, Jincao Wu, Tim Johnson Dec 2010

Predicting Treatment Efficacy Via Quantitative Mri: A Bayesian Joint Model, Jincao Wu, Tim Johnson

The University of Michigan Department of Biostatistics Working Paper Series

The prognosis for patients with high-grade gliomas is poor, with a median survival of one year. Treatment efficacy assessment is typically unavailable until 5{6 months post diagnosis. Investigators hypothesize that quantitative MRI (qMRI) can assess treatment efficacy three weeks after therapy starts, thereby allowing salvage treatments to begin earlier. The purpose of this work is to build a predictive model of treatment efficacy using qMRI data and to assess its performance. The outcome is one-year survival status. We propose a joint, two-stage Bayesian model. In stage I, we smooth the image data with a multivariate spatio-temporal pairwise dierence prior. We …


Coronary Heart Disease Mortality And Long-Term Exposure To Ambient Particulate Air Pollutants In Elderly Nonsmoking California Residents, Lie Hong Chen Dec 2010

Coronary Heart Disease Mortality And Long-Term Exposure To Ambient Particulate Air Pollutants In Elderly Nonsmoking California Residents, Lie Hong Chen

Loma Linda University Electronic Theses, Dissertations & Projects

The purpose of this study is to assess the effect of long-term concentrations of ambient PM on risks of all causes, cardiopulmonary, coronary heart disease (CHD), total cancer, and any mention of nonmalignant respiratory disease (NMRD) mortality.

The health effects of long-term ambient air pollution have been studied with up to 30 years of follow-up in the AHSMOG cohort, a cohort of 6,338 nonsmoking white California adults. Monthly concentrations of ambient air pollutants [particulate matter(PMio), Ozone (O3), sulfur dioxide (SO2), nitrogen dioxide (NO2) or particulate matter

In the AHSMOG cohort, each increment of 10 |ig/m3 in PMio in two-pollutant models …


Spatial Epidemiology Of Birth Defects In The United States And The State Of Utah Using Geographic Information Systems And Spatial Statistics, Samson Y. Gebreab Dec 2010

Spatial Epidemiology Of Birth Defects In The United States And The State Of Utah Using Geographic Information Systems And Spatial Statistics, Samson Y. Gebreab

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

Oral clefts are the most common form of birth defects in the United States (US) and the State of Utah has among the highest prevalence of oral clefts in the nation. The overall objective of this dissertation was to examine the spatial distribution of oral clefts and their linkage with a broad range of demographic, behavioral, social, economic, and environmental risk factors through the application of Geographic Information Systems (GIS) and spatial statistics. Using innovative linked micromaps plots, we investigated the geographic patterns of oral clefts occurrence from 1998 to 2002 and their relationships with maternal smoking rates and proportion …


The Determinants Of Colorectal Cancer Survival Disparities In Nevada, Lucas N. Wassira Dec 2010

The Determinants Of Colorectal Cancer Survival Disparities In Nevada, Lucas N. Wassira

UNLV Theses, Dissertations, Professional Papers, and Capstones

Different population groups across Nevada and throughout the United States suffer disproportionately from colorectal cancer and its after-effects. Overcoming cancer health disparities is important for lessening the burden of cancer. There has been an overall decline in the incidence of and mortality from colorectal cancer (CRC). This is likely due, in part, to the increasing use of screening procedures such as Fecal Occult Blood Test (FOBT) and/or endoscopy, which can reduce the risk of CRC mortality by fifty percent. Nevertheless, screening procedures are routinely used by only fifty percent of Americans aged fifty years and older. Despite overall mortality decreasing …


Asymptotic Theory For Cross-Validated Targeted Maximum Likelihood Estimation, Wenjing Zheng, Mark J. Van Der Laan Nov 2010

Asymptotic Theory For Cross-Validated Targeted Maximum Likelihood Estimation, Wenjing Zheng, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

We consider a targeted maximum likelihood estimator of a path-wise differentiable parameter of the data generating distribution in a semi-parametric model based on observing n independent and identically distributed observations. The targeted maximum likelihood estimator (TMLE) uses V-fold sample splitting for the initial estimator in order to make the TMLE maximally robust in its bias reduction step. We prove a general theorem that states asymptotic efficiency (and thereby regularity) of the targeted maximum likelihood estimator when the initial estimator is consistent and a second order term converges to zero in probability at a rate faster than the square root of …


Observational Study And Individualized Antiretroviral Therapy Initiation Rules For Reducing Cancer Incidence In Hiv-Infected Patients, Romain Neugebauer, Michael J. Silverberg, Mark J. Van Der Laan Nov 2010

Observational Study And Individualized Antiretroviral Therapy Initiation Rules For Reducing Cancer Incidence In Hiv-Infected Patients, Romain Neugebauer, Michael J. Silverberg, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Targeted Maximum Likelihood Learning (TMLL) has been proposed as a general estimation methodology that can, in particular, be applied to draw causal inferences based on marginal structural modeling with observational data using either a point treatment approach (all confounders are assumed not to be affected by the exposure(s) of interest) or a longitudinal data approach (some confounders may be affected by one of the exposures of interest). While formal development of TMLL has included road maps for applications in longitudinal data approaches, real-life implementations have been restricted to studies based on a point treatment approach. In this article, we illustrate …


Inferential Methods For High-Throughput Methylation Data, Maria Capparuccini Nov 2010

Inferential Methods For High-Throughput Methylation Data, Maria Capparuccini

Theses and Dissertations

The role of abnormal DNA methylation in the progression of disease is a growing area of research that relies upon the establishment of sound statistical methods. The common method for declaring there is differential methylation between two groups at a given CpG site, as summarized by the difference between proportions methylated db=b1-b2, has been through use of a Filtered Two Sample t-test, using the recommended filter of 0.17 (Bibikova et al., 2006b). In this dissertation, we performed a re-analysis of the data used in recommending the threshold by fitting a mixed-effects ANOVA model. It was determined that the 0.17 filter …


Power And Sample Size For Three-Level Cluster Designs, Tina Cunningham Nov 2010

Power And Sample Size For Three-Level Cluster Designs, Tina Cunningham

Theses and Dissertations

Over the past few decades, Cluster Randomized Trials (CRT) have become a design of choice in many research areas. One of the most critical issues in planning a CRT is to ensure that the study design is sensitive enough to capture the intervention effect. The assessment of power and sample size in such studies is often faced with many challenges due to several methodological difficulties. While studies on power and sample size for cluster designs with one and two levels are abundant, the evaluation of required sample size for three-level designs has been generally overlooked. First, the nesting effect introduces …


Stereotype Logit Models For High Dimensional Data, Andre Williams Oct 2010

Stereotype Logit Models For High Dimensional Data, Andre Williams

Theses and Dissertations

Gene expression studies are of growing importance in the field of medicine. In fact, subtypes within the same disease have been shown to have differing gene expression profiles (Golub et al., 1999). Often, researchers are interested in differentiating a disease by a categorical classification indicative of disease progression. For example, it may be of interest to identify genes that are associated with progression and to accurately predict the state of progression using gene expression data. One challenge when modeling microarray gene expression data is that there are more genes (variables) than there are observations. In addition, the genes usually demonstrate …


Mengukur Kualitas Hidup Anak, Toha Muhaimin Oct 2010

Mengukur Kualitas Hidup Anak, Toha Muhaimin

Kesmas

Kata kualitas hidup sering dihubungkan dengan pembangunan, khususnya pembangunan manusia, yang sering dikaitkan dengan kondisi seseorang baik dalam keadaan sehat maupun sakit, untuk menunjukkan aktivitas fisik, atau kondisi seseorang dalam hidup sehari-harinya. Sebagian orang mengkaitkan istilah kualitas hidup dengan kondisi sejauh mana terpenuhinya kebutuhan dasar untuk hidup seperti sandang, pangan, papan dan pendidikan pada seseorang. Oleh karena itu, banyak penelitian mengukur kualitas hidup dengan instrumen yang berbeda-beda, termasuk mengukur kualitas hidup anak dan banyak instrumen yang telah dikembangkan. Tulisan ini mencoba membahas pengertian kualitas hidup dan cara mengukurnya, terutama pada anak. Belum ada konsensus mengukur atau menggambarkan definisi konseptual kualitas …


Modeling Functional Data With Spatially Heterogeneous Shape Characteristics, Ana-Maria Staicu, Ciprian M. Crainiceanu, Daniel S. Reich, David Ruppert Oct 2010

Modeling Functional Data With Spatially Heterogeneous Shape Characteristics, Ana-Maria Staicu, Ciprian M. Crainiceanu, Daniel S. Reich, David Ruppert

Johns Hopkins University, Dept. of Biostatistics Working Papers

We propose a novel class of models for functional data exhibiting skewness or other shape characteristics that vary with spatial or temporal location. We use copulas so that the marginal distributions and the dependence structure can be modeled independently. Dependence is modeled with a Gaussian or t-copula, so that there is an underlying latent Gaussian process. We model the marginal distributions using the skew t family. The mean, variance, and shape parameters are modeled nonparametrically as functions of location. A computationally tractable inferential framework for estimating heterogeneous asymmetric or heavy-tailed marginal distributions is introduced. This framework provides a new set …


A Maximum Pseudo-Likelihood Approach For Estimating Species Trees Under The Coalescent Model, Liang Liu, Lili Yu, Scott V. Edwards Oct 2010

A Maximum Pseudo-Likelihood Approach For Estimating Species Trees Under The Coalescent Model, Liang Liu, Lili Yu, Scott V. Edwards

Biostatistics: Faculty Publications

Background

Several phylogenetic approaches have been developed to estimate species trees from collections of gene trees. However, maximum likelihood approaches for estimating species trees under the coalescent model are limited. Although the likelihood of a species tree under the multispecies coalescent model has already been derived by Rannala and Yang, it can be shown that the maximum likelihood estimate (MLE) of the species tree (topology, branch lengths, and population sizes) from gene trees under this formula does not exist. In this paper, we develop a pseudo-likelihood function of the species tree to obtain maximum pseudo-likelihood estimates (MPE) of species trees, …


Population Value Decomposition, A Framework For The Analysis Of Image Populations, Ciprian M. Crainiceanu, Brian S. Caffo, Sheng Luo, Vadim Zipunnikov Oct 2010

Population Value Decomposition, A Framework For The Analysis Of Image Populations, Ciprian M. Crainiceanu, Brian S. Caffo, Sheng Luo, Vadim Zipunnikov

Johns Hopkins University, Dept. of Biostatistics Working Papers

Images, often stored in multidimensional arrays are fast becoming ubiquitous in medical and public health research. Analyzing populations of images is a statistical problem that raises a host of daunting challenges. The most severe challenge is that data sets incorporating images recorded for hundreds or thousands of subjects at multiple visits are massive. We introduce the population value decomposition (PVD), a general method for simultaneous dimensionality reduction of large populations of massive images. We show how PVD can seamlessly be incorporated into statistical modeling and lead to a new, transparent and fast inferential framework. Our methodology was motivated by and …


Targeted Bayesian Learning, Ivan Diaz Munoz, Alan E. Hubbard, Mark J. Van Der Laan Oct 2010

Targeted Bayesian Learning, Ivan Diaz Munoz, Alan E. Hubbard, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Targeted maximum likelihood estimation (van der Laan & Rubin 2006) is a loss-based semi-parametric estimation method that yields a substitution estimator of a target parameter of the probability distribution of the data that solves the efficient influence curve estimating equation, and thereby yields a double robust locally efficient estimator of the parameter of interest, under regularity conditions. The Bayesian paradigm is concerned with including the researcher’s prior uncertainty about the parameter through a prior distribution, which combined with the likelihood yields a posterior distribution for the parameter that reflects the researcher’s posterior uncertainty. In this paper, we present a way …


The Pathways To Mental Health Care Of First-Episode Psychosis Patients: A Systematic Review., Kelly K. Anderson, Rebecca Fuhrer, Ashok K. Malla Oct 2010

The Pathways To Mental Health Care Of First-Episode Psychosis Patients: A Systematic Review., Kelly K. Anderson, Rebecca Fuhrer, Ashok K. Malla

Epidemiology and Biostatistics Publications

BACKGROUND: Although there is agreement on the association between delay in treatment of psychosis and outcome, less is known regarding the pathways to care of patients suffering from a first psychotic episode. Pathways are complex, involve a diverse range of contacts, and are likely to influence delay in treatment. We conducted a systematic review on the nature and determinants of the pathway to care of patients experiencing a first psychotic episode.

METHOD: We searched four databases (Medline, HealthStar, EMBASE, PsycINFO) to identify articles published between 1985 and 2009. We manually searched reference lists and relevant journals and used forward citation …


Landmark Prediction Of Survival, Layla Parast, Tianxi Cai Sep 2010

Landmark Prediction Of Survival, Layla Parast, Tianxi Cai

Harvard University Biostatistics Working Paper Series

No abstract provided.


Diagnosing And Responding To Violations In The Positivity Assumption, Maya L. Petersen, Kristin Porter, Susan Gruber, Yue Wang, Mark J. Van Der Laan Sep 2010

Diagnosing And Responding To Violations In The Positivity Assumption, Maya L. Petersen, Kristin Porter, Susan Gruber, Yue Wang, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

The assumption of positivity or experimental treatment assignment requires that observed treatment levels vary within confounder strata. This article discusses the positivity assumption in the context of assessing model and parameter-specific identifiability of causal effects. Positivity violations occur when certain subgroups in a sample rarely or never receive some treatments of interest. The resulting sparsity in the data may increase bias with or without an increase in variance and can threaten valid inference. The parametric bootstrap is presented as a tool to assess the severity of such threats and its utility as a diagnostic is explored using simulated data. Several …


Stratifying Subjects For Treatment Selection With Censored Event Time Data From A Comparative Study, Lihui Zhao, Tianxi Cai, Lu Tian, Hajime Uno, Scott D. Solomon, L. J. Wei Sep 2010

Stratifying Subjects For Treatment Selection With Censored Event Time Data From A Comparative Study, Lihui Zhao, Tianxi Cai, Lu Tian, Hajime Uno, Scott D. Solomon, L. J. Wei

Harvard University Biostatistics Working Paper Series

No abstract provided.


Spousal Concordance In Academic Achievements And Intelligence And Family-Based Association Studies Identified Novel Loci Associated With Intelligence., Yue Pan Aug 2010

Spousal Concordance In Academic Achievements And Intelligence And Family-Based Association Studies Identified Novel Loci Associated With Intelligence., Yue Pan

Electronic Theses and Dissertations

Assortative Mating, the tendency for mate selection to occur on the basis of similar traits, plays an essential role in understanding the genetic variation on academic achievements and intelligence (IQ). It is an important mechanism explaining spousal concordance. We used principal component analysis (PCA) for spousal correlation. There is a significant positive correlation between spouses by the new variable PC1 (correlation coefficient=0.515, p<0.0001). We further research the genetic factor that affects IQ by using the same data. We performed a low density genome-wide association (GWA) analysis with a family-based association test to identify genetic variants that associated with intelligence as measured by WAIS full-score IQ (FSIQ). NTM at 11q25 (rs411280, p=0.000764) and NR3C2 at 4q31.23 (rs3846329, p=0.000675) were 2 novel genes that haven't been associated with IQ from other studies. This study may serve as a resource for replication in other populations and a foundation for future investigations.


A Bayesian Approach To Dose-Response Assessment And Drug-Drug Interaction Analysis: Application To In Vitro Studies, Violeta G. Hennessey Aug 2010

A Bayesian Approach To Dose-Response Assessment And Drug-Drug Interaction Analysis: Application To In Vitro Studies, Violeta G. Hennessey

Dissertations and Theses (Open Access)

The considerable search for synergistic agents in cancer research is motivated by the therapeutic benefits achieved by combining anti-cancer agents. Synergistic agents make it possible to reduce dosage while maintaining or enhancing a desired effect. Other favorable outcomes of synergistic agents include reduction in toxicity and minimizing or delaying drug resistance. Dose-response assessment and drug-drug interaction analysis play an important part in the drug discovery process, however analysis are often poorly done. This dissertation is an effort to notably improve dose-response assessment and drug-drug interaction analysis.

The most commonly used method in published analysis is the Median-Effect Principle/Combination Index method …


Principled Sure Independence Screening For Cox Models With Ultra-High-Dimensional Covariates, Sihai Dave Zhao, Yi Li Jul 2010

Principled Sure Independence Screening For Cox Models With Ultra-High-Dimensional Covariates, Sihai Dave Zhao, Yi Li

Harvard University Biostatistics Working Paper Series

No abstract provided.


An Inferential Framework For Network Hypothesis Tests: With Applications To Biological Networks, Phillip Yates Jun 2010

An Inferential Framework For Network Hypothesis Tests: With Applications To Biological Networks, Phillip Yates

Theses and Dissertations

The analysis of weighted co-expression gene sets is gaining momentum in systems biology. In addition to substantial research directed toward inferring co-expression networks on the basis of microarray/high-throughput sequencing data, inferential methods are being developed to compare gene networks across one or more phenotypes. Common gene set hypothesis testing procedures are mostly confined to comparing average gene/node transcription levels between one or more groups and make limited use of additional network features, e.g., edges induced by significant partial correlations. Ignoring the gene set architecture disregards relevant network topological comparisons and can result in familiar n<


Estimation Of Causal Effects Of Community Based Interventions, Mark J. Van Der Laan Jun 2010

Estimation Of Causal Effects Of Community Based Interventions, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Suppose one assigns two interventions to a small number K of different populations or communities, and one measures covariates and outcomes on a random sample of independent individuals from each of the K populations. We investigate the problem of identification and estimation of the causal effect of the choice of intervention assigned at the community level, and, if the intervention is time-dependent, the causal effect of the changes in the intervention at time t, on the outcome. The challenge one is confronted with is that different populations have different environmental factors and that the intervention and environment are assigned to …


Multi-State Life Tables, Equilibrium Prevalence, And Baseline Selection Bias, Paula Diehr, David Yanez Jun 2010

Multi-State Life Tables, Equilibrium Prevalence, And Baseline Selection Bias, Paula Diehr, David Yanez

UW Biostatistics Working Paper Series

Consider a 3-state system with one absorbing state, such as Healthy, Sick, and Dead. If the system satisfies the 1-step Markov conditions, the prevalence of the Healthy state will converge to a value that is independent of the initial distribution. This equilibrium prevalence and its variance are known under the assumption of time homogeneity, and provided reasonable estimates in the time non-homogeneous systems studied. Here, we derived the equilibrium prevalence for a system with more than three states. Under time homogeneity, the equilibrium prevalence distribution was shown to be an eigenvector of a partition of the matrix of transition probabilities. …


An Empirical Approach To Evaluating Sufficient Similarity: Utilization Of Euclidean Distance As A Similarity Measure, Scott Marshall May 2010

An Empirical Approach To Evaluating Sufficient Similarity: Utilization Of Euclidean Distance As A Similarity Measure, Scott Marshall

Theses and Dissertations

Individuals are exposed to chemical mixtures while carrying out everyday tasks, with unknown risk associated with exposure. Given the number of resulting mixtures it is not economically feasible to identify or characterize all possible mixtures. When complete dose-response data are not available on a (candidate) mixture of concern, EPA guidelines define a similar mixture based on chemical composition, component proportions and expert biological judgment (EPA, 1986, 2000). Current work in this literature is by Feder et al. (2009), evaluating sufficient similarity in exposure to disinfection by-products of water purification using multivariate statistical techniques and traditional hypothesis testing. The work of …


Cost And Accuracy Comparisons In Medical Testing Using Sequential Testing Strategies, Anwar Ahmed May 2010

Cost And Accuracy Comparisons In Medical Testing Using Sequential Testing Strategies, Anwar Ahmed

Theses and Dissertations

The practice of sequential testing is followed by the evaluation of accuracy, but often not by the evaluation of cost. This research described and compared three sequential testing strategies: believe the negative (BN), believe the positive (BP) and believe the extreme (BE), the latter being a less-examined strategy. All three strategies were used to combine results of two medical tests to diagnose a disease or medical condition. Descriptions of these strategies were provided in terms of accuracy (using the maximum receiver operating curve or MROC) and cost of testing (defined as the proportion of subjects who need 2 tests to …


Powerful Snp Set Analysis For Case-Control Genome Wide Association Studies, Michael C. Wu, Peter Kraft, Michael P. Epstein, Deanne M. Taylor, Stephen J. Chanock, David J. Hunter, Xihong Lin May 2010

Powerful Snp Set Analysis For Case-Control Genome Wide Association Studies, Michael C. Wu, Peter Kraft, Michael P. Epstein, Deanne M. Taylor, Stephen J. Chanock, David J. Hunter, Xihong Lin

Harvard University Biostatistics Working Paper Series

No abstract provided.


Estimating Causal Effects In Trials Involving Multi-Treatment Arms Subject To Non-Compliance: A Bayesian Frame-Work, Qi Long, Roderick J. Little, Xihong Lin May 2010

Estimating Causal Effects In Trials Involving Multi-Treatment Arms Subject To Non-Compliance: A Bayesian Frame-Work, Qi Long, Roderick J. Little, Xihong Lin

Harvard University Biostatistics Working Paper Series

No abstract provided.


A Targeted Maximum Likelihood Estimator Of A Causal Effect On A Bounded Continuous Outcome, Susan Gruber, Mark J. Van Der Laan May 2010

A Targeted Maximum Likelihood Estimator Of A Causal Effect On A Bounded Continuous Outcome, Susan Gruber, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Targeted maximum likelihood estimation of a parameter of a data generating distribution, known to be an element of a semiparametric model, involves constructing a parametric model through an initial density estimator with parameter epsilon representing an amount of fluctuation of the initial density estimator, where the score of this fluctuation model at epsilon=0 equals the efficient influence curve/canonical gradient. The latter constraint can be satisfied by many parametric fluctuation models, since it represents only a local constraint of its behavior at zero fluctuation. However, it is very important that the fluctuations stay within the semiparametric model for the observed data …