Open Access. Powered by Scholars. Published by Universities.®

Biostatistics Commons™

Open Access. Powered by Scholars. Published by Universities.®

Theses and Dissertations

Discipline
Institution
Keyword
Publication Year

Articles 121 - 150 of 162

Full-Text Articles in Biostatistics

Advanced Methodology Developments In Mixture Cure Models, Chao Cai Jan 2013

Advanced Methodology Developments In Mixture Cure Models, Chao Cai

Theses and Dissertations

Modern medical treatments have substantially improved cure rates for many chronic diseases and have generated increasing interest in appropriate statistical models to handle survival data with non-negligible cure fractions. The mixture cure models are designed to model such data set, which assume that studied population is a mixture of being cured and uncured. In this dissertation, I will develop two programs named smcure and NPHMC in R. The first program aims to facilitate estimating two popular mixture cure models: the proportional hazards (PH) mixture cure model and accelerated failure time (AFT) mixture cure model. The second program focuses on designing …


A Comparison Of Methods Of Analysis To Control For Confounding In A Cohort Study Of A Dietary Intervention, Esinhart Hali Jul 2012

A Comparison Of Methods Of Analysis To Control For Confounding In A Cohort Study Of A Dietary Intervention, Esinhart Hali

Theses and Dissertations

Comparing samples from different populations can be biased by confounding. There are several statistical methods that can be used to control for confounding. These include; multiple linear regression, propensity score matching, propensity score/logit of propensity score as a single covariate in a linear regression model, stratified analysis using propensity score quintiles, weighted analysis using propensity scores or trimmed scores. The data were from two studies of a dietary intervention (FIBERR and RNP). The outcome variable was change from baseline to one month for eight outcome measures; fat, fiber, and fruits/ vegetables behavior, fat, fiber, and fruits/vegetables intentions, fat and fruits/vegetables …


The Effect Of Baseline Cluster Stratification On The Power Of Pre-Post Analysis, Fengjiao Hu Jul 2012

The Effect Of Baseline Cluster Stratification On The Power Of Pre-Post Analysis, Fengjiao Hu

Theses and Dissertations

The purpose of study is to check whether the power of detecting the effect of intervention versus control in a pre- and post-study can be increased by using a stratified randomized controlled design. A stratified randomized controlled design with two study arms and two time points, where strata are determined by clustering on baseline outcomes of the primary measure, is considered. A modified hierarchical clustering algorithm is developed which guarantees optimality as well as requiring each cluster to have at least one subject per study arm. The power is calculated based on simulated bivariate normal distributed primary measures with mixture …


Does Pair-Matching On Ordered Baseline Measures Increase Power: A Simulation Study, Yan Jin Jul 2012

Does Pair-Matching On Ordered Baseline Measures Increase Power: A Simulation Study, Yan Jin

Theses and Dissertations

It has been shown that pair-matching on an ordered baseline with normally distributed measures reduces the variance of the estimated treatment effect (Park and Johnson, 2006). The main objective of this study is to examine if pair-matching improves the power when the distribution is a mixture of two normal distributions. Multiple scenarios with a combination of different sample sizes and parameters are simulated. The power curves are provided for three cases, with and without matching, as follows: analysis of post-intervention data only, adding baseline as a covariate, and classic pre-post comparison. The study shows that the additional variance reduction provided …


Unbiased Estimation For The Contextual Effect Of Duration Of Adolescent Height Growth On Adulthood Obesity And Health Outcomes Via Hierarchical Linear And Nonlinear Models, Robert Carrico May 2012

Unbiased Estimation For The Contextual Effect Of Duration Of Adolescent Height Growth On Adulthood Obesity And Health Outcomes Via Hierarchical Linear And Nonlinear Models, Robert Carrico

Theses and Dissertations

This dissertation has multiple aims in studying hierarchical linear models in biomedical data analysis. In Chapter 1, the novel idea of studying the durations of adolescent growth spurts as a predictor of adulthood obesity is defined, established, and illustrated. The concept of contextual effects modeling is introduced in this first section as we study secular trend of adulthood obesity and how this trend is mitigated by the durations of individual adolescent growth spurts and the secular average length of adolescent growth spurts. It is found that individuals with longer periods of fast height growth in adolescence are more prone to …


Statistical Methods For Normalization And Analysis Of High-Throughput Genomic Data, Tobias Guennel Jan 2012

Statistical Methods For Normalization And Analysis Of High-Throughput Genomic Data, Tobias Guennel

Theses and Dissertations

High-throughput genomic datasets obtained from microarray or sequencing studies have revolutionized the field of molecular biology over the last decade. The complexity of these new technologies also poses new challenges to statisticians to separate biological relevant information from technical noise. Two methods are introduced that address important issues with normalization of array comparative genomic hybridization (aCGH) microarrays and the analysis of RNA sequencing (RNA-Seq) studies. Many studies investigating copy number aberrations at the DNA level for cancer and genetic studies use comparative genomic hybridization (CGH) on oligo arrays. However, aCGH data often suffer from low signal to noise ratios resulting …


Inferential Methods For High-Throughput Methylation Data, Maria Capparuccini Nov 2010

Inferential Methods For High-Throughput Methylation Data, Maria Capparuccini

Theses and Dissertations

The role of abnormal DNA methylation in the progression of disease is a growing area of research that relies upon the establishment of sound statistical methods. The common method for declaring there is differential methylation between two groups at a given CpG site, as summarized by the difference between proportions methylated db=b1-b2, has been through use of a Filtered Two Sample t-test, using the recommended filter of 0.17 (Bibikova et al., 2006b). In this dissertation, we performed a re-analysis of the data used in recommending the threshold by fitting a mixed-effects ANOVA model. It was determined that the 0.17 filter …


Power And Sample Size For Three-Level Cluster Designs, Tina Cunningham Nov 2010

Power And Sample Size For Three-Level Cluster Designs, Tina Cunningham

Theses and Dissertations

Over the past few decades, Cluster Randomized Trials (CRT) have become a design of choice in many research areas. One of the most critical issues in planning a CRT is to ensure that the study design is sensitive enough to capture the intervention effect. The assessment of power and sample size in such studies is often faced with many challenges due to several methodological difficulties. While studies on power and sample size for cluster designs with one and two levels are abundant, the evaluation of required sample size for three-level designs has been generally overlooked. First, the nesting effect introduces …


Stereotype Logit Models For High Dimensional Data, Andre Williams Oct 2010

Stereotype Logit Models For High Dimensional Data, Andre Williams

Theses and Dissertations

Gene expression studies are of growing importance in the field of medicine. In fact, subtypes within the same disease have been shown to have differing gene expression profiles (Golub et al., 1999). Often, researchers are interested in differentiating a disease by a categorical classification indicative of disease progression. For example, it may be of interest to identify genes that are associated with progression and to accurately predict the state of progression using gene expression data. One challenge when modeling microarray gene expression data is that there are more genes (variables) than there are observations. In addition, the genes usually demonstrate …


An Inferential Framework For Network Hypothesis Tests: With Applications To Biological Networks, Phillip Yates Jun 2010

An Inferential Framework For Network Hypothesis Tests: With Applications To Biological Networks, Phillip Yates

Theses and Dissertations

The analysis of weighted co-expression gene sets is gaining momentum in systems biology. In addition to substantial research directed toward inferring co-expression networks on the basis of microarray/high-throughput sequencing data, inferential methods are being developed to compare gene networks across one or more phenotypes. Common gene set hypothesis testing procedures are mostly confined to comparing average gene/node transcription levels between one or more groups and make limited use of additional network features, e.g., edges induced by significant partial correlations. Ignoring the gene set architecture disregards relevant network topological comparisons and can result in familiar n<


An Empirical Approach To Evaluating Sufficient Similarity: Utilization Of Euclidean Distance As A Similarity Measure, Scott Marshall May 2010

An Empirical Approach To Evaluating Sufficient Similarity: Utilization Of Euclidean Distance As A Similarity Measure, Scott Marshall

Theses and Dissertations

Individuals are exposed to chemical mixtures while carrying out everyday tasks, with unknown risk associated with exposure. Given the number of resulting mixtures it is not economically feasible to identify or characterize all possible mixtures. When complete dose-response data are not available on a (candidate) mixture of concern, EPA guidelines define a similar mixture based on chemical composition, component proportions and expert biological judgment (EPA, 1986, 2000). Current work in this literature is by Feder et al. (2009), evaluating sufficient similarity in exposure to disinfection by-products of water purification using multivariate statistical techniques and traditional hypothesis testing. The work of …


Cost And Accuracy Comparisons In Medical Testing Using Sequential Testing Strategies, Anwar Ahmed May 2010

Cost And Accuracy Comparisons In Medical Testing Using Sequential Testing Strategies, Anwar Ahmed

Theses and Dissertations

The practice of sequential testing is followed by the evaluation of accuracy, but often not by the evaluation of cost. This research described and compared three sequential testing strategies: believe the negative (BN), believe the positive (BP) and believe the extreme (BE), the latter being a less-examined strategy. All three strategies were used to combine results of two medical tests to diagnose a disease or medical condition. Descriptions of these strategies were provided in terms of accuracy (using the maximum receiver operating curve or MROC) and cost of testing (defined as the proportion of subjects who need 2 tests to …


A Numerical Method For Estimating The Variance Of Age At Maximum Growth Rate In Growth Models, Semhar Ogbagaber Apr 2010

A Numerical Method For Estimating The Variance Of Age At Maximum Growth Rate In Growth Models, Semhar Ogbagaber

Theses and Dissertations

Most studies on maturation and body composition using the Fels Longitudinal data mention peak height velocity (PHV) as an important outcome measure. The PHV is often derived from growth models such as the triple logistic model fitted to the stature (height) data. The age at PHV is sometimes ordinalized to designate an individual as an early, average or late maturer. In theory, age at PHV is the age at which the rate of growth reaches the maximum. Theoretically, for a well behaved growth function, this could be obtained by setting the second derivative of the growth function to zero and …


Bayesian And Frequentist Approaches For The Analysis Of Multiple Endpoints Data Resulting From Exposure To Multiple Health Stressors., Epiphanie Nyirabahizi Mar 2010

Bayesian And Frequentist Approaches For The Analysis Of Multiple Endpoints Data Resulting From Exposure To Multiple Health Stressors., Epiphanie Nyirabahizi

Theses and Dissertations

In risk analysis, Benchmark dose (BMD)methodology is used to quantify the risk associated with exposure to stressors such as environmental chemicals. It consists of fitting a mathematical model to the exposure data and the BMD is the dose expected to result in a pre-specified response or benchmark response (BMR). Most available exposure data are from single chemical exposure, but living objects are exposed to multiple sources of hazards. Furthermore, in some studies, researchers may observe multiple endpoints on one subject. Statistical approaches to address multiple endpoints problem can be partitioned into a dimension reduction group and a dimension preservative group. …


Nonlinear Models In Multivariate Population Bioequivalence Testing, Bassam Dahman Nov 2009

Nonlinear Models In Multivariate Population Bioequivalence Testing, Bassam Dahman

Theses and Dissertations

In this dissertation a methodology is proposed for simultaneously evaluating the population bioequivalence (PBE) of a generic drug to a pre-licensed drug, or the bioequivalence of two formulations of a drug using multiple correlated pharmacokinetic metrics. The univariate criterion that is accepted by the food and drug administration (FDA) for testing population bioequivalence is generalized. Very few approaches for testing multivariate extensions of PBE have appeared in the literature. One method uses the trace of the covariance matrix as a measure of total variability, and another uses a pooled variance instead of the reference variance. The former ignores the correlation …


Joint Mixed-Effects Models For Longitudinal Data Analysis: An Application For The Metabolic Syndrome, John Thorp Iii Nov 2009

Joint Mixed-Effects Models For Longitudinal Data Analysis: An Application For The Metabolic Syndrome, John Thorp Iii

Theses and Dissertations

Mixed-effects models are commonly used to model longitudinal data as they can appropriately account for within and between subject sources of variability. Univariate mixed effect modeling strategies are well developed for a single outcome (response) variable that may be continuous (e.g. Gaussian) or categorical (e.g. binary, Poisson) in nature. Only recently have extensions been discussed for jointly modeling multiple outcome variables measures longitudinally. Many diseases processes are a function of several factors that are correlated. For example, the metabolic syndrome, a constellation of cardiovascular risk factors associated with an increased risk of cardiovascular disease and type 2 diabetes, is often …


Deriving Optimal Composite Scores: Relating Observational/Longitudinal Data With A Primary Endpoint, Rhonda Ellis Sep 2009

Deriving Optimal Composite Scores: Relating Observational/Longitudinal Data With A Primary Endpoint, Rhonda Ellis

Theses and Dissertations

In numerous clinical/experimental studies, multiple endpoints are measured on each subject. It is often not clear which of these endpoints should be designated as of primary importance. The desirability function approach is a way of combining multiple responses into a single unitless composite score. The response variables may include multiple types of data: binary, ordinal, count, interval data. Each response variable is transformed to a 0 to1 unitless scale with zero representing a completely undesirable response and one representing the ideal value. In desirability function methodology, weights on individual components can be incorporated to allow different levels of importance to …


A Sequential Algorithm To Identify The Mixing Endpoints In Liquids In Pharmaceutical Applications, Akriti Saxena Jul 2009

A Sequential Algorithm To Identify The Mixing Endpoints In Liquids In Pharmaceutical Applications, Akriti Saxena

Theses and Dissertations

The objective of this thesis is to develop a sequential algorithm to determine accurately and quickly, at which point in time a product is well mixed or reaches a steady state plateau, in terms of the Refractive Index (RI). An algorithm using sequential non-linear model fitting and prediction is proposed. A simulation study representing typical scenarios in a liquid manufacturing process in pharmaceutical industries was performed to evaluate the proposed algorithm. The data simulated included autocorrelated normal errors and used the Gompertz model. A set of 27 different combinations of the parameters of the Gompertz function were considered. The results …


Comparing Bootstrap And Jackknife Variance Estimation Methods For Area Under The Roc Curve Using One-Stage Cluster Survey Data, Allison Dunning Jun 2009

Comparing Bootstrap And Jackknife Variance Estimation Methods For Area Under The Roc Curve Using One-Stage Cluster Survey Data, Allison Dunning

Theses and Dissertations

The purpose of this research is to examine the bootstrap and jackknife as methods for estimating the variance of the AUC from a study using a complex sampling design and to determine which characteristics of the sampling design effects this estimation. Data from a one-stage cluster sampling design of 10 clusters was examined. Factors included three true AUCs (.60, .75, and .90), three prevalence levels (50/50, 70/30, 90/10) (non-disease/disease), and finally three number of clusters sampled (2, 5, or 7). A simulated sample was constructed for each of the 27 combinations of AUC, prevalence and number of clusters. Estimates of …


Tolerance Intervals In Random-Effects Models, Kakotan Sanogo Dec 2008

Tolerance Intervals In Random-Effects Models, Kakotan Sanogo

Theses and Dissertations

In the pharmaceutical setting, it is often necessary to establish the shelf life of a drug product and sometimes suitable to assess the risk of product failure at the desired expiry period. The current statistical methodology use confidence intervals for the predicted mean to establish the expiry period and prediction intervals for a predicted new assay value or a tolerance interval for a proportion of the population for use in a risk assessment. A major concern is that most methodology treat a homogeneous subpopulation, say batch, either as a fixed effect and therefore uses a fixed-effects regression model (Graybill, 1976) …


Applications Of The Bivariate Gamma Distribution In Nutritional Epidemiology And Medical Physics, Jolene Barker Sep 2008

Applications Of The Bivariate Gamma Distribution In Nutritional Epidemiology And Medical Physics, Jolene Barker

Theses and Dissertations

In this thesis the utility of a bivariate gamma distribution is explored. In the field of nutritional epidemiology a nutrition density transformation is used to reduce collinearity. This phenomenon will be shown to result due to the independent variables following a bivariate gamma model. In the field of radiation oncology paired comparison of variances is often performed. The bivariate gamma model is also appropriate for fitting correlated variances. A method for simulating bivariate gamma random variables is presented. This method is used to generate data from several bivariate gamma models and the asymptotic properties of a test statistic, suggested for …


Variable Selection In Competing Risks Using The L1-Penalized Cox Model, Xiangrong Kong Sep 2008

Variable Selection In Competing Risks Using The L1-Penalized Cox Model, Xiangrong Kong

Theses and Dissertations

One situation in survival analysis is that the failure of an individual can happen because of one of multiple distinct causes. Survival data generated in this scenario are commonly referred to as competing risks data. One of the major tasks, when examining survival data, is to assess the dependence of survival time on explanatory variables. In competing risks, as with ordinary univariate survival data, there may be explanatory variables associated with the risks raised from the different causes being studied. The same variable might have different degrees of influence on the risks due to different causes. Given a set of …


Probe Level Analysis Of Affymetrix Microarray Data, Richard Ellis Kennedy Jan 2008

Probe Level Analysis Of Affymetrix Microarray Data, Richard Ellis Kennedy

Theses and Dissertations

The analysis of Affymetrix GeneChip® data is a complex, multistep process. Most often, methodscondense the multiple probe level intensities into single probeset level measures (such as RobustMulti-chip Average (RMA), dChip and Microarray Suite version 5.0 (MAS5)), which are thenfollowed by application of statistical tests to determine which genes are differentially expressed. An alternative approach is a probe-level analysis, which tests for differential expression directly using the probe-level data. Probe-level models offer the potential advantage of more accurately capturing sources of variation in microarray experiments. However, this has not been thoroughly investigated, since current research efforts have largely focused on the …


An Adaptive Dose Finding Design (Dosefind) Using A Nonlinear Dose Response Model, James Michael Davenport Jan 2007

An Adaptive Dose Finding Design (Dosefind) Using A Nonlinear Dose Response Model, James Michael Davenport

Theses and Dissertations

First-in-man (FIM) Phase I clinical trials are part of the critical path in the development of a new compound entity (NCE). Since FIM clinical trials are the first time that an NCE is dosed in human subjects, the designs used in these trials are unique and geared toward patient safety. We develop a method for obtaining the desired response using an adaptive non-linear approach. This method is applicable for studies in which MTD, NOEL,NOAEL, PK, PD effects or other such endpoints are evaluated to determine the desired dose. The method has application whenever a measurable PD marker is an indicator …


Phase Ii Trials Powered To Detect Activity In Tumor Subsets With Retrospective (Or Prospective) Use Of Predictive Markers, Grishma S. Sheth Jan 2007

Phase Ii Trials Powered To Detect Activity In Tumor Subsets With Retrospective (Or Prospective) Use Of Predictive Markers, Grishma S. Sheth

Theses and Dissertations

Classical phase II trial designs assume a patient population with a homogeneous tumor type and yield an estimate of a stochastic probability of tumor response. Clinically, however, oncology is moving towards identifying patients who are likely to respond to therapy using tumor subtyping based upon predictive markers. Such designs are called targeted designs (Simon, 2004). For a given phase I1 trial predictive markers may be defined prospectively (on the basis of previous results) or identified retrospectively on the basis of analysis of responding and non-responding tumors. For the prospective case we propose two Phase I1 targeted designs in which a) …


Quantifying The Effects Of Correlated Covariates On Variable Importance Estimates From Random Forests, Ryan Vincent Kimes Jan 2006

Quantifying The Effects Of Correlated Covariates On Variable Importance Estimates From Random Forests, Ryan Vincent Kimes

Theses and Dissertations

Recent advances in computing technology have lead to the development of algorithmic modeling techniques. These methods can be used to analyze data which are difficult to analyze using traditional statistical models. This study examined the effectiveness of variable importance estimates from the random forest algorithm in identifying the true predictor among a large number of candidate predictors. A simulation study was conducted using twenty different levels of association among the independent variables and seven different levels of association between the true predictor and the response. We conclude that the random forest method is an effective classification tool when the goals …


Assessing, Modifying, And Combining Data Fields From The Virginia Office Of The Chief Medical Examiner (Ocme) Dataset And The Virginia Department Of Forensic Science (Dfs) Datasets In Order To Compare Concentrations Of Selected Drugs, Amy Elizabeth Herrin Jan 2006

Assessing, Modifying, And Combining Data Fields From The Virginia Office Of The Chief Medical Examiner (Ocme) Dataset And The Virginia Department Of Forensic Science (Dfs) Datasets In Order To Compare Concentrations Of Selected Drugs, Amy Elizabeth Herrin

Theses and Dissertations

The Medical Examiner of Virginia (ME) dataset and the Virginia Department of Forensic Science Driving Under the Influence of Drugs (DUI) datasets were used to determine whether people have the potential to develop tolerances to diphenhydramine, cocaine, oxycodone, hydrocodone, methadone, and morphine. These datasets included the years 2000-2004 and were used to compare the concentrations of these six drugs between people who died from a drug-related cause of death (of the drug of interest) and people who were pulled over for driving under the influence. Three drug pattern groups were created to divide each of the six drug-specific datasets in …


Optimal Clustering: Genetic Constrained K-Means And Linear Programming Algorithms, Jianmin Zhao Jan 2006

Optimal Clustering: Genetic Constrained K-Means And Linear Programming Algorithms, Jianmin Zhao

Theses and Dissertations

Methods for determining clusters of data under- specified constraints have recently gained popularity. Although general constraints may be used, we focus on clustering methods with the constraint of a minimal cluster size. In this dissertation, we propose two constrained k-means algorithms: Linear Programming Algorithm (LPA) and Genetic Constrained K-means Algorithm (GCKA). Linear Programming Algorithm modifies the k-means algorithm into a linear programming problem with constraints requiring that each cluster have m or more subjects. In order to achieve an acceptable clustering solution, we run the algorithm with a large number of random sets of initial seeds, and choose the solution …


A Comparison For Longitudinal Data Missing Due To Truncation, Rong Liu Jan 2006

A Comparison For Longitudinal Data Missing Due To Truncation, Rong Liu

Theses and Dissertations

Many longitudinal clinical studies suffer from patient dropout. Often the dropout is nonignorable and the missing mechanism needs to be incorporated in the analysis. The methods handling missing data make various assumptions about the missing mechanism, and their utility in practice depends on whether these assumptions apply in a specific application. Ramakrishnan and Wang (2005) proposed a method (MDT) to handle nonignorable missing data, where missing is due to the observations exceeding an unobserved threshold. Assuming that the observations arise from a truncated normal distribution, they suggested an EM algorithm to simplify the estimation.In this dissertation the EM algorithm is …


Statistical Methods And Experimental Design For Inference Regarding Dose And/Or Interaction Thresholds Along A Fixed-Ratio Ray, Sharon Dziuba Yeatts Jan 2006

Statistical Methods And Experimental Design For Inference Regarding Dose And/Or Interaction Thresholds Along A Fixed-Ratio Ray, Sharon Dziuba Yeatts

Theses and Dissertations

An alternative to the full factorial design, the ray design is appropriate for investigating a mixture of c chemicals, which are present according to a fixed mixing ratio, called the mixture ray. Using single chemical and mixture ray data, we can investigate interaction among the chemicals in a particular mixture. Statistical models have been used to describe the dose-response relationship of the single agents and the mixture; additivity is tested through the significance of model parameters associated with the coincidence of the additivity and mixture models.It is often assumed that a chemical or mixture must be administered above an unknown …