Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Biostatistics

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 2191 - 2220 of 2512

Full-Text Articles in Statistics and Probability

An Inferential Framework For Network Hypothesis Tests: With Applications To Biological Networks, Phillip Yates Jun 2010

An Inferential Framework For Network Hypothesis Tests: With Applications To Biological Networks, Phillip Yates

Theses and Dissertations

The analysis of weighted co-expression gene sets is gaining momentum in systems biology. In addition to substantial research directed toward inferring co-expression networks on the basis of microarray/high-throughput sequencing data, inferential methods are being developed to compare gene networks across one or more phenotypes. Common gene set hypothesis testing procedures are mostly confined to comparing average gene/node transcription levels between one or more groups and make limited use of additional network features, e.g., edges induced by significant partial correlations. Ignoring the gene set architecture disregards relevant network topological comparisons and can result in familiar n<


Estimation Of Causal Effects Of Community Based Interventions, Mark J. Van Der Laan Jun 2010

Estimation Of Causal Effects Of Community Based Interventions, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Suppose one assigns two interventions to a small number K of different populations or communities, and one measures covariates and outcomes on a random sample of independent individuals from each of the K populations. We investigate the problem of identification and estimation of the causal effect of the choice of intervention assigned at the community level, and, if the intervention is time-dependent, the causal effect of the changes in the intervention at time t, on the outcome. The challenge one is confronted with is that different populations have different environmental factors and that the intervention and environment are assigned to …


Multi-State Life Tables, Equilibrium Prevalence, And Baseline Selection Bias, Paula Diehr, David Yanez Jun 2010

Multi-State Life Tables, Equilibrium Prevalence, And Baseline Selection Bias, Paula Diehr, David Yanez

UW Biostatistics Working Paper Series

Consider a 3-state system with one absorbing state, such as Healthy, Sick, and Dead. If the system satisfies the 1-step Markov conditions, the prevalence of the Healthy state will converge to a value that is independent of the initial distribution. This equilibrium prevalence and its variance are known under the assumption of time homogeneity, and provided reasonable estimates in the time non-homogeneous systems studied. Here, we derived the equilibrium prevalence for a system with more than three states. Under time homogeneity, the equilibrium prevalence distribution was shown to be an eigenvector of a partition of the matrix of transition probabilities. …


An Empirical Approach To Evaluating Sufficient Similarity: Utilization Of Euclidean Distance As A Similarity Measure, Scott Marshall May 2010

An Empirical Approach To Evaluating Sufficient Similarity: Utilization Of Euclidean Distance As A Similarity Measure, Scott Marshall

Theses and Dissertations

Individuals are exposed to chemical mixtures while carrying out everyday tasks, with unknown risk associated with exposure. Given the number of resulting mixtures it is not economically feasible to identify or characterize all possible mixtures. When complete dose-response data are not available on a (candidate) mixture of concern, EPA guidelines define a similar mixture based on chemical composition, component proportions and expert biological judgment (EPA, 1986, 2000). Current work in this literature is by Feder et al. (2009), evaluating sufficient similarity in exposure to disinfection by-products of water purification using multivariate statistical techniques and traditional hypothesis testing. The work of …


Cost And Accuracy Comparisons In Medical Testing Using Sequential Testing Strategies, Anwar Ahmed May 2010

Cost And Accuracy Comparisons In Medical Testing Using Sequential Testing Strategies, Anwar Ahmed

Theses and Dissertations

The practice of sequential testing is followed by the evaluation of accuracy, but often not by the evaluation of cost. This research described and compared three sequential testing strategies: believe the negative (BN), believe the positive (BP) and believe the extreme (BE), the latter being a less-examined strategy. All three strategies were used to combine results of two medical tests to diagnose a disease or medical condition. Descriptions of these strategies were provided in terms of accuracy (using the maximum receiver operating curve or MROC) and cost of testing (defined as the proportion of subjects who need 2 tests to …


Powerful Snp Set Analysis For Case-Control Genome Wide Association Studies, Michael C. Wu, Peter Kraft, Michael P. Epstein, Deanne M. Taylor, Stephen J. Chanock, David J. Hunter, Xihong Lin May 2010

Powerful Snp Set Analysis For Case-Control Genome Wide Association Studies, Michael C. Wu, Peter Kraft, Michael P. Epstein, Deanne M. Taylor, Stephen J. Chanock, David J. Hunter, Xihong Lin

Harvard University Biostatistics Working Paper Series

No abstract provided.


Estimating Causal Effects In Trials Involving Multi-Treatment Arms Subject To Non-Compliance: A Bayesian Frame-Work, Qi Long, Roderick J. Little, Xihong Lin May 2010

Estimating Causal Effects In Trials Involving Multi-Treatment Arms Subject To Non-Compliance: A Bayesian Frame-Work, Qi Long, Roderick J. Little, Xihong Lin

Harvard University Biostatistics Working Paper Series

No abstract provided.


A Targeted Maximum Likelihood Estimator Of A Causal Effect On A Bounded Continuous Outcome, Susan Gruber, Mark J. Van Der Laan May 2010

A Targeted Maximum Likelihood Estimator Of A Causal Effect On A Bounded Continuous Outcome, Susan Gruber, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Targeted maximum likelihood estimation of a parameter of a data generating distribution, known to be an element of a semiparametric model, involves constructing a parametric model through an initial density estimator with parameter epsilon representing an amount of fluctuation of the initial density estimator, where the score of this fluctuation model at epsilon=0 equals the efficient influence curve/canonical gradient. The latter constraint can be satisfied by many parametric fluctuation models, since it represents only a local constraint of its behavior at zero fluctuation. However, it is very important that the fluctuations stay within the semiparametric model for the observed data …


Super Learner In Prediction, Eric C. Polley, Mark J. Van Der Laan May 2010

Super Learner In Prediction, Eric C. Polley, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Super learning is a general loss based learning method that has been proposed and analyzed theoretically in van der Laan et al. (2007). In this article we consider super learning for prediction. The super learner is a prediction method designed to find the optimal combination of a collection of prediction algorithms. The super learner algorithm finds the combination of algorithms minimizing the cross-validated risk. The super learner framework is built on the theory of cross-validation and allows for a general class of prediction algorithms to be considered for the ensemble. Due to the previously established oracle results for the cross-validation …


Assessing Noninferiority In A Three-Arm Trial Using The Bayesian Approach, Pulak Ghosh, Farouk S. Nathoo, Mithat Gonen, Ram C. Tiwari May 2010

Assessing Noninferiority In A Three-Arm Trial Using The Bayesian Approach, Pulak Ghosh, Farouk S. Nathoo, Mithat Gonen, Ram C. Tiwari

Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series

Non-inferiority trials, which aim to demonstrate that a test product is not worse than a competitor by more than a pre-specified small amount, are of great importance to the pharmaceutical community. As a result, methodology for designing and analyzing such trials is required, and developing new methods for such analysis is an important area of statistical research. The three-arm clinical trial is usually recommended for non-inferiority trials by the Food and Drug Administration (FDA). The three-arm trial consists of a placebo, a reference, and an experimental treatment, and simultaneously tests the superiority of the reference over the placebo along with …


Survival Prediction For Brain Tumor Patients Using Gene Expression Data, Vinicius Bonato May 2010

Survival Prediction For Brain Tumor Patients Using Gene Expression Data, Vinicius Bonato

Dissertations and Theses (Open Access)

Brain tumor is one of the most aggressive types of cancer in humans, with an estimated median survival time of 12 months and only 4% of the patients surviving more than 5 years after disease diagnosis. Until recently, brain tumor prognosis has been based only on clinical information such as tumor grade and patient age, but there are reports indicating that molecular profiling of gliomas can reveal subgroups of patients with distinct survival rates. We hypothesize that coupling molecular profiling of brain tumors with clinical information might improve predictions of patient survival time and, consequently, better guide future treatment decisions. …


A Numerical Method For Estimating The Variance Of Age At Maximum Growth Rate In Growth Models, Semhar Ogbagaber Apr 2010

A Numerical Method For Estimating The Variance Of Age At Maximum Growth Rate In Growth Models, Semhar Ogbagaber

Theses and Dissertations

Most studies on maturation and body composition using the Fels Longitudinal data mention peak height velocity (PHV) as an important outcome measure. The PHV is often derived from growth models such as the triple logistic model fitted to the stature (height) data. The age at PHV is sometimes ordinalized to designate an individual as an early, average or late maturer. In theory, age at PHV is the age at which the rate of growth reaches the maximum. Theoretically, for a well behaved growth function, this could be obtained by setting the second derivative of the growth function to zero and …


Nonparametric Regression With Missing Outcomes Using Weighted Kernel Estimating Equations, Lu Wang, Andrea Rotnitzky, Xihong Lin Apr 2010

Nonparametric Regression With Missing Outcomes Using Weighted Kernel Estimating Equations, Lu Wang, Andrea Rotnitzky, Xihong Lin

Harvard University Biostatistics Working Paper Series

No abstract provided.


Utilizing The Integrated Difference Of Two Survival Functions To Quantify The Treatment Contrast For Designing, Monitoring And Analyzing A Comparative Clinical Study, Lihui Zhao, Lu Tian, Hajime Uno, Scott D. Solomon, Marc A. Pfeffer, J. S. Schindler, L. J. Wei Apr 2010

Utilizing The Integrated Difference Of Two Survival Functions To Quantify The Treatment Contrast For Designing, Monitoring And Analyzing A Comparative Clinical Study, Lihui Zhao, Lu Tian, Hajime Uno, Scott D. Solomon, Marc A. Pfeffer, J. S. Schindler, L. J. Wei

Harvard University Biostatistics Working Paper Series

No abstract provided.


Likelihood Ratio Testing For Admixture Models With Application To Genetic Linkage Analysis, Chong-Zhi Di, Kung-Yee Liang Mar 2010

Likelihood Ratio Testing For Admixture Models With Application To Genetic Linkage Analysis, Chong-Zhi Di, Kung-Yee Liang

Johns Hopkins University, Dept. of Biostatistics Working Papers

We consider likelihood ratio tests (LRT) and their modifications for homogeneity in admixture models. The admixture model is a special case of two component mixture model, where one component is indexed by an unknown parameter while the parameter value for the other component is known. It has been widely used in genetic linkage analysis under heterogeneity, in which the kernel distribution is binomial. For such models, it is long recognized that testing for homogeneity is nonstandard and the LRT statistic does not converge to a conventional 2 distribution. In this paper, we investigate the asymptotic behavior of the LRT for …


Bayesian And Frequentist Approaches For The Analysis Of Multiple Endpoints Data Resulting From Exposure To Multiple Health Stressors., Epiphanie Nyirabahizi Mar 2010

Bayesian And Frequentist Approaches For The Analysis Of Multiple Endpoints Data Resulting From Exposure To Multiple Health Stressors., Epiphanie Nyirabahizi

Theses and Dissertations

In risk analysis, Benchmark dose (BMD)methodology is used to quantify the risk associated with exposure to stressors such as environmental chemicals. It consists of fitting a mathematical model to the exposure data and the BMD is the dose expected to result in a pre-specified response or benchmark response (BMR). Most available exposure data are from single chemical exposure, but living objects are exposed to multiple sources of hazards. Furthermore, in some studies, researchers may observe multiple endpoints on one subject. Statistical approaches to address multiple endpoints problem can be partitioned into a dimension reduction group and a dimension preservative group. …


Targeted Maximum Likelihood Method For Repeated Measures Semiparametric Regression: Discovery For Transcription Factor Activity, Catherine Tuglus, Mark J. Van Der Laan Mar 2010

Targeted Maximum Likelihood Method For Repeated Measures Semiparametric Regression: Discovery For Transcription Factor Activity, Catherine Tuglus, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

In longitudinal and repeated measures data analysis, often the goal is to determine the effect of a treatment or aspect on a particular outcome (e.g. disease progression). We consider semiparametric repeated measures regression model, where the parametric component models effect of the variable of interest and any modification by other covariates. The expectation of this parametric component over the other covariates is a measure of variable importance. Here we present a targeted maximum likelihood estimator of the finite dimensional regression parameter, which is easily estimated using standard software for generalized estimating equations. The targeted maximum likelihood method provides double robust …


Graphical Procedures For Evaluating Overall And Subject-Specific Incremental Values From New Predictors With Censored Event Time Data, Hajime Uno, Tianxi Cai, Lu Tian, L. J. Wei Mar 2010

Graphical Procedures For Evaluating Overall And Subject-Specific Incremental Values From New Predictors With Censored Event Time Data, Hajime Uno, Tianxi Cai, Lu Tian, L. J. Wei

Harvard University Biostatistics Working Paper Series

No abstract provided.


Doubly Regularized Reml For Estimation And Selection Of Fixed And Random Effects In Linear Mixed-Effects Models, Sijian Wang, Peter Xuewin Song, Ji Zhu Mar 2010

Doubly Regularized Reml For Estimation And Selection Of Fixed And Random Effects In Linear Mixed-Effects Models, Sijian Wang, Peter Xuewin Song, Ji Zhu

The University of Michigan Department of Biostatistics Working Paper Series

The linear mixed effects model (LMM) is widely used in the analysis of clustered or longitudinal data. In the practice of LMM, the inference on the structure of the random effects component is of great importance, not only to yield proper interpretation of subject-specific effects but also to draw valid statistical conclusions. This task of inference becomes significantly challenging when a large number of fixed effects and random effects are involved in the analysis. The difficulty of variable selection arises from the need of simultaneously regularizing both mean model and covariance structures, with possible parameter constraints between the two. In …


Multilevel Sparse Functional Principal Component Analysis, Chong-Zhi Di, Ciprian M. Crainiceanu Mar 2010

Multilevel Sparse Functional Principal Component Analysis, Chong-Zhi Di, Ciprian M. Crainiceanu

Johns Hopkins University, Dept. of Biostatistics Working Papers

The basic observational unit in this paper is a function. Data are assumed to have a natural hierarchy of basic units. A simple example is when functions are recorded at multiple visits for the same subject. Di et al. (2009) proposed Multilevel Functional Principal Component Analysis (MFPCA) for this type of data structure when functions are densely sampled. Here we consider the case when functions are sparsely sampled and may contain as few as 2 or 3 observations per function. As with MFPCA, we exploit the multilevel structure of covariance operators and data reduction induced by the use of principal …


Targeting The Optimal Design In Randomized Clinical Trials With Binary Outcomes And No Covariate, Antoine Chambaz, Mark J. Van Der Laan Feb 2010

Targeting The Optimal Design In Randomized Clinical Trials With Binary Outcomes And No Covariate, Antoine Chambaz, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

This article is devoted to the asymptotic study of adaptive group sequential designs in the case of randomized clinical trials with binary treatment, binary outcome and no covariate. By adaptive design, we mean in this setting a clinical trial design that allows the investigator to dynamically modify its course through data-driven adjustment of the randomization probability based on data accrued so far, without negatively impacting on the statistical integrity of the trial. By adaptive group sequential design, we refer to the fact that group sequential testing methods can be equally well applied on top of adaptive designs. Prior to collection …


Targeted Maximum Likelihood Based Causal Inference, Mark J. Van Der Laan Feb 2010

Targeted Maximum Likelihood Based Causal Inference, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Given causal graph assumptions, intervention-specific counterfactual distributions of the data can be defined by the so called G-computation formula, which is obtained by carrying out these interventions on the likelihood of the data factorized according to the causal graph. The obtained G-computation formula represents the counterfactual distribution the data would have had if this intervention would have been enforced on the system generating the data. A causal effect of interest can now be defined as some difference between these counterfactual distributions indexed by different interventions. For example, the interventions can represent static treatment regimens or individualized treatment rules that assign …


Bio-Creep In Non-Inferiority Clinical Trials, Siobhan P. Everson-Stewart, Scott S. Emerson Feb 2010

Bio-Creep In Non-Inferiority Clinical Trials, Siobhan P. Everson-Stewart, Scott S. Emerson

UW Biostatistics Working Paper Series

After a non-inferiority clinical trial, a new therapy may be accepted as effective, even if its treatment effect is slightly smaller than the current standard. It is therefore possible that, after a series of trials where the new therapy is slightly worse than the preceding drugs, an ineffective or harmful therapy might be incorrectly declared efficacious; this is known as “bio-creep.” Several factors may influence the rate at which bio-creep occurs, including the distribution of the effects of the new agents being tested and how that changes over time, the choice of active comparator, the method used to model the …


Estimates Of Information Growth In Longitudinal Clinical Trials, Abigail Shoben, Kyle Rudser, Scott S. Emerson Feb 2010

Estimates Of Information Growth In Longitudinal Clinical Trials, Abigail Shoben, Kyle Rudser, Scott S. Emerson

UW Biostatistics Working Paper Series

In group sequential clinical trials, it is necessary to estimate the amount of information present at interim analysis times relative to the amount of information that would be present at the final analysis. If only one measurement is made per individual, this is often the ratio of sample sizes available at the interim and final analyses. However, as discussed by Wu and Lan (1992), when the statistic of interest is a change over time, as with longitudinal data, such an approach overstates the information. In this paper, we discuss other problems that can result in overestimating the information, such as …


Simple, Efficient Estimators Of Treatment Effects In Randomized Trials Using Generalized Linear Models To Leverage Baseline Variables, Michael Rosenblum, Mark J. Van Der Laan Jan 2010

Simple, Efficient Estimators Of Treatment Effects In Randomized Trials Using Generalized Linear Models To Leverage Baseline Variables, Michael Rosenblum, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Models, such as logistic regression and Poisson regression models, are often used to estimate treatment effects in randomized trials. These models leverage information in variables collected before randomization, in order to obtain more precise estimates of treatment effects. However, there is the danger that model misspecification will lead to bias. We show that certain easy to compute, model-based estimators are asymptotically unbiased even when the working model used is arbitrarily misspecified. Furthermore, these estimators are locally efficient. As a special case of our main result, we consider a simple Poisson working model containing only main terms; in this case, we …


Targeted Maximum Likelihood Estimation Of The Parameter Of A Marginal Structural Model, Michael Rosenblum, Mark J. Van Der Laan Jan 2010

Targeted Maximum Likelihood Estimation Of The Parameter Of A Marginal Structural Model, Michael Rosenblum, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Targeted maximum likelihood estimation is a versatile tool for estimating parameters in semiparametric and nonparametric models. We work through an example applying targeted maximum likelihood methodology to estimate the parameter of a marginal structural model. In the case we consider, we show how this can be easily done by clever use of standard statistical software. We point out differences between targeted maximum likelihood estimation and other approaches (including estimating function based methods). The application we consider is to estimate the effect of adherence to antiretroviral medications on virologic failure in HIV positive individuals.


Identification Of Neuroblastoma And Its Prognostic Markers Using Raman Spectroscopy, Rachel Kast Jan 2010

Identification Of Neuroblastoma And Its Prognostic Markers Using Raman Spectroscopy, Rachel Kast

Wayne State University Dissertations

Introduction: Neuroblastoma is the most common cancer of infancy. It is one of several peripheral nervous system tumors, including ganglioneuroma, peripheral nerve sheath tumor, and pheochromocytoma. It is commonly situated on the adrenal gland. It displays similar histology to other small round blue cell tumors, including non-Hodgkin lymphoma, rhabdomyosarcoma, and Ewing sarcoma. One method of judging neuroblastoma aggressiveness uses tumor histology factors, including mitosis-karyorrhexis index, Schwannian stromal development, degree of differentiation, and patient age. Tumor aggressiveness can also be judged based on the amplification of certain genes, including MYCN. Raman spectroscopy is a physics-based method which identifies the biochemical …


Detecting Outliers And Influential Observations In Survival Model., Nor Akmal Md Noh Jan 2010

Detecting Outliers And Influential Observations In Survival Model., Nor Akmal Md Noh

Student Works (2010-2019)

This study proposes outlier and influential observation detection procedures for Cox proportional hazard model. In the estimation process, the parameters for Cox proportional hazard model are estimated using partial likelihood method, while the baseline hazard estimates are obtained using Nelson-Aalen method. The procedure of outlier detection is based on three types of residuals; deviance, log-odd and normal deviate residuals. We study their properties and compare their performance in detecting outliers via simulation. On the other hand, we propose a procedure of identifying influential observation using forward search method. The method has been shown to be effective in detecting influential observations …


On The Eigenstructures Of Functional K-Potent Matrices And Their Integral Forms, Yan Wu, Daniel F. Linder Jan 2010

On The Eigenstructures Of Functional K-Potent Matrices And Their Integral Forms, Yan Wu, Daniel F. Linder

Biostatistics: Faculty Publications

In this paper, a functional k-potent matrix satisfies the equation, where k and r are positive integers, and are real numbers. This class of matrices includes idempotent, Nilpotent, and involutary matrices, and more. It turns out that the matrices in this group are best distinguished by their associated eigen-structures. The spectral properties of the matrices are exploited to construct integral k-potent matrices, which have special roles in digital image encryption.


A Markov Transition Model To Dementia With Death As A Competing Event, Liou Xu Jan 2010

A Markov Transition Model To Dementia With Death As A Competing Event, Liou Xu

University of Kentucky Doctoral Dissertations

The research on multi-state Markov transition model is motivated by the nature of the longitudinal data from the Nun Study (Snowdon, 1997), and similar information on the BRAiNS cohort (Salazar, 2004). Our goal is to develop a flexible methodology for handling the categorical longitudinal responses and competing risks time-to-event that characterizes the features of the data for research on dementia. To do so, we treat the survival from death as a continuous variable rather than defining death as a competing absorbing state to dementia. We assume that within each subject the survival component and the Markov process are linked by …