Open Access. Powered by Scholars. Published by Universities.®

Statistical Models Commons

Open Access. Powered by Scholars. Published by Universities.®

Biostatistics

Institution
Keyword
Publication Year
Publication
Publication Type

Articles 121 - 150 of 192

Full-Text Articles in Statistical Models

Multilevel Models For Longitudinal Data, Aastha Khatiwada Aug 2016

Multilevel Models For Longitudinal Data, Aastha Khatiwada

Electronic Theses and Dissertations

Longitudinal data arise when individuals are measured several times during an ob- servation period and thus the data for each individual are not independent. There are several ways of analyzing longitudinal data when different treatments are com- pared. Multilevel models are used to analyze data that are clustered in some way. In this work, multilevel models are used to analyze longitudinal data from a case study. Results from other more commonly used methods are compared to multilevel models. Also, comparison in output between two software, SAS and R, is done. Finally a method consisting of fitting individual models for each …


Integration Of Multi-Platform High-Dimensional Omic Data, Xuebei An May 2016

Integration Of Multi-Platform High-Dimensional Omic Data, Xuebei An

Dissertations and Theses (Open Access)

The development of high-throughput biotechnologies have made data accessible from different platforms, including RNA sequencing, copy number variation, DNA methylation, protein lysate arrays, etc. The high-dimensional omic data derived from different technological platforms have been extensively used to facilitate comprehensive understanding of disease mechanisms and to determine personalized health treatments. Although vital to the progress of clinical research, the high dimensional multi-platform data impose new challenges for data analysis. Numerous studies have been proposed to integrate multi-platform omic data; however, few have efficiently and simultaneously addressed the problems that arise from high dimensionality and complex correlations.

In my dissertation, I …


Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang Feb 2016

Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang

COBRA Preprint Series

Non-negative matrix factorization (NMF) is a widely used machine learning algorithm for dimension reduction of large-scale data. It has found successful applications in a variety of fields such as computational biology, neuroscience, natural language processing, information retrieval, image processing and speech recognition. In bioinformatics, for example, it has been used to extract patterns and profiles from genomic and text-mining data as well as in protein sequence and structure analysis. While the scientific performance of NMF is very promising in dealing with high dimensional data sets and complex data structures, its computational cost is high and sometimes could be critical for …


Models For Hsv Shedding Must Account For Two Levels Of Overdispersion, Amalia Magaret Jan 2016

Models For Hsv Shedding Must Account For Two Levels Of Overdispersion, Amalia Magaret

UW Biostatistics Working Paper Series

We have frequently implemented crossover studies to evaluate new therapeutic interventions for genital herpes simplex virus infection. The outcome measured to assess the efficacy of interventions on herpes disease severity is the viral shedding rate, defined as the frequency of detection of HSV on the genital skin and mucosa. We performed a simulation study to ascertain whether our standard model, which we have used previously, was appropriately considering all the necessary features of the shedding data to provide correct inference. We simulated shedding data under our standard, validated assumptions and assessed the ability of 5 different models to reproduce the …


Statistical Methods For Environmental Exposure Data Subject To Detection Limits, Yuchen Yang Jan 2016

Statistical Methods For Environmental Exposure Data Subject To Detection Limits, Yuchen Yang

Theses and Dissertations--Statistics

In this dissertation, we develop unified and efficient nonparametric statistical methods for estimating and comparing environmental exposure distributions in presence of detection limits. In the first part, we propose a kernel-smoothed nonparametric estimator for the exposure distribution without imposing any independence assumption between the exposure level and detection limit. We show that the proposed estimator is consistent and asymptotically normal. Simulation studies demonstrate that the proposed estimator performs well in practical situations. A colon cancer study is provided for illustration. In the second part, we develop a class of test statistics to compare exposure distributions between two groups by using …


Improved Models For Differential Analysis For Genomic Data, Hong Wang Jan 2016

Improved Models For Differential Analysis For Genomic Data, Hong Wang

Theses and Dissertations--Statistics

This paper intend to develop novel statistical methods to improve genomic data analysis, especially for differential analysis. We considered two different data type: NanoString nCounter data and somatic mutation data. For NanoString nCounter data, we develop a novel differential expression detection method. The method considers a generalized linear model of the negative binomial family to characterize count data and allows for multi-factor design. Data normalization is incorporated in the model framework through data normalization parameters, which are estimated from control genes embedded in the nCounter system. For somatic mutation data, we develop beta-binomial model-based approaches to identify highly or lowly …


A Pairwise Likelihood Augmented Estimator For The Cox Model Under Left-Truncation, Fan Wu, Sehee Kim, Jing Qin, Rajiv Saran, Yi Li Sep 2015

A Pairwise Likelihood Augmented Estimator For The Cox Model Under Left-Truncation, Fan Wu, Sehee Kim, Jing Qin, Rajiv Saran, Yi Li

The University of Michigan Department of Biostatistics Working Paper Series

Survival data collected from prevalent cohorts are subject to left-truncation and the analysis is challenging. Conditional approaches for left-truncated data under the Cox model are inefficient as they typically ignore the information in the marginal likelihood of the truncation times. Length-biased sampling methods can improve the estimation efficiency but only when the stationarity assumption of the disease incidence holds, i.e., the truncation distribution is uniform; otherwise they may generate biased estimates. In this paper, we propose a semi-parametric method for the Cox model under general left-truncation, where the truncation distribution is unspecified. Our approach is to make inference based on …


Preparedness Of Hospitals In The Republic Of Ireland For An Influenza Pandemic, An Infection Control Perspective, Mary Reidy, Fiona Ryan, Dervla Hogan, Seán Lacey, Claire Buckley Sep 2015

Preparedness Of Hospitals In The Republic Of Ireland For An Influenza Pandemic, An Infection Control Perspective, Mary Reidy, Fiona Ryan, Dervla Hogan, Seán Lacey, Claire Buckley

Department of Mathematics Publications

When an influenza pandemic occurs most of the population is susceptible and attack rates can range as high as 40–50 %. The most important failure in pandemic planning is the lack of standards or guidelines regarding what it means to be ‘prepared’. The aim of this study was to assess the preparedness of acute hospitals in the Republic of Ireland for an influenza pandemic from an infection control perspective.


Using Capture-Mark-Recapture Techniques To Estimate Detection Probabilities & Fidelity Of Expression For The Critically Endangered James Spinymussel (Pleurobema Collina)., Alaina C. Esposito May 2015

Using Capture-Mark-Recapture Techniques To Estimate Detection Probabilities & Fidelity Of Expression For The Critically Endangered James Spinymussel (Pleurobema Collina)., Alaina C. Esposito

Masters Theses, 2010-2019

The critically endangered James Spinymussel (Pleurobema collina) is a species of freshwater mussel endemic to Virginia’s James and Dan River basins. In the last 20 years, P. collina has experienced a substantial decline in numbers and currently occupies approximately 10% of its original habitat; however, little information is known about this species to assist in conservation. A 230-meter reach of transitional habitat in Swift Run was selected for repeat observations to estimate detection probabilities using a Capture-Mark-Recapture framework. In June 2014, visual scouting began to locate and tag P. collina (including other mussels in the community) with PIT …


Meta-Analysis Of Gene Expression Studies, Umaporn Siangphoe Jan 2015

Meta-Analysis Of Gene Expression Studies, Umaporn Siangphoe

Theses and Dissertations

Combining effect sizes from individual studies using random-effects models are commonly applied in high-dimensional gene expression data. However, unknown study heterogeneity can arise from inconsistency of sample qualities and experimental conditions. High heterogeneity of effect sizes can reduce statistical power of the models. We proposed two new methods for random effects estimation and measurements for model variation and strength of the study heterogeneity. We then developed a statistical technique to test for significance of random effects and identify heterogeneous genes. We also proposed another meta-analytic approach that incorporates informative weights in the random effects meta-analysis models. We compared the proposed …


Using Graphs To Characterize Nationwide Physician Referral Networks, Ding Tong, Shu-Xia Li, Isuru Ranasinghe, Sudhakar Nuti, Hongyu Zhao, Harlan Krumholz Sep 2014

Using Graphs To Characterize Nationwide Physician Referral Networks, Ding Tong, Shu-Xia Li, Isuru Ranasinghe, Sudhakar Nuti, Hongyu Zhao, Harlan Krumholz

Yale Day of Data

AIM:

Evaluating physician referral network characteristics can help to understand how physicians and hospitals interact to provide patient services within the US healthcare system and ultimately how this may influence patient outcomes.

METHOD:

We used the 2012-2013 national Physician Referral data from the Centers for Medicare & Medicaid Services (CMS), which consists of 73,071,804 pairs of referrals from one health provider to another in calendar year 2012 and the first two quarters of year 2013 within 30 days of care. These referrals are from 642,144 national-wide physicians and 4,811 hospitals. We obtained information for each provider, physician or hospital, from …


A Study Of Joinpoint Models For Longitudinal Data, Libo Zhou Aug 2014

A Study Of Joinpoint Models For Longitudinal Data, Libo Zhou

UNLV Theses, Dissertations, Professional Papers, and Capstones

In many medical studies, data are collected simultaneously on multiple biomarkers from each individual. Levels of these biomarkers are measured periodically over certain time duration, giving rise to longitudinal trajectories. The subjects under study may also be subject to dropout due to several competing causes, the likelihood of which may be affected by the levels of these biomarkers. In this dissertation, we investigate flexible Bayesian modeling of such data, taking into account any available covariate information as well as possible censoring of the drop-out times. We propose joint models for multiple biomarkers with multiple causes of dropout. Our proposed models …


Meta-Analysis Of Social-Personality Psychological Research, Blair T. Johnson, Alice H. Eagly Jan 2014

Meta-Analysis Of Social-Personality Psychological Research, Blair T. Johnson, Alice H. Eagly

CHIP Documents

This publication provides a contemporary treatment of the subject of meta-analysis in relation to social-personality psychology. Meta-analysis literally refers to the statistical pooling of the results of independent studies on a given subject, although in practice it refers as well to other steps of research synthesis, including defining the question under investigation, gathering all available research reports, coding of information about the studies and their effects, and interpretation/dissemination of results. Discussed as well are the hallmarks of high-quality meta-analyses.


Genetic Association Testing Of Copy Number Variation, Yinglei Li Jan 2014

Genetic Association Testing Of Copy Number Variation, Yinglei Li

Theses and Dissertations--Statistics

Copy-number variation (CNV) has been implicated in many complex diseases. It is of great interest to detect and locate such regions through genetic association testings. However, the association testings are complicated by the fact that CNVs usually span multiple markers and thus such markers are correlated to each other. To overcome the difficulty, it is desirable to pool information across the markers. In this thesis, we propose a kernel-based method for aggregation of marker-level tests, in which first we obtain a bunch of p-values through association tests for every marker and then the association test involving CNV is based on …


Net Reclassification Index: A Misleading Measure Of Prediction Improvement, Margaret Sullivan Pepe, Holly Janes, Kathleen F. Kerr, Bruce M. Psaty Sep 2013

Net Reclassification Index: A Misleading Measure Of Prediction Improvement, Margaret Sullivan Pepe, Holly Janes, Kathleen F. Kerr, Bruce M. Psaty

UW Biostatistics Working Paper Series

The evaluation of biomarkers to improve risk prediction is a common theme in modern research. Since its introduction in 2008, the net reclassification index (NRI) (Pencina et al. 2008, Pencina et al. 2011) has gained widespread use as a measure of prediction performance with over 1,200 citations as of June 30, 2013. The NRI is considered by some to be more sensitive to clinically important changes in risk than the traditional change in the AUC (Delta AUC) statistic (Hlatky et al. 2009). Recent statistical research has raised questions, however, about the validity of conclusions based on the NRI. (Hilden and …


Attributing Effects To Interactions, Tyler J. Vanderweele, Eric J. Tchetgen Tchetgen Jul 2013

Attributing Effects To Interactions, Tyler J. Vanderweele, Eric J. Tchetgen Tchetgen

Harvard University Biostatistics Working Paper Series

A framework is presented which allows an investigator to estimate the portion of the effect of one exposure that is attributable to an interaction with a second exposure. We show that when the two exposures are independent, the total effect of one exposure can be decomposed into a conditional effect of that exposure and a component due to interaction. The decomposition applies on difference or ratio scales. We discuss how the components can be estimated using standard regression models, and how these components can be used to evaluate the proportion of the total effect of the primary exposure attributable to …


Statistical Inference For Data Adaptive Target Parameters, Mark J. Van Der Laan, Alan E. Hubbard, Sara Kherad Pajouh Jun 2013

Statistical Inference For Data Adaptive Target Parameters, Mark J. Van Der Laan, Alan E. Hubbard, Sara Kherad Pajouh

U.C. Berkeley Division of Biostatistics Working Paper Series

Consider one observes n i.i.d. copies of a random variable with a probability distribution that is known to be an element of a particular statistical model. In order to define our statistical target we partition the sample in V equal size sub-samples, and use this partitioning to define V splits in estimation-sample (one of the V subsamples) and corresponding complementary parameter-generating sample that is used to generate a target parameter. For each of the V parameter-generating samples, we apply an algorithm that maps the sample in a target parameter mapping which represent the statistical target parameter generated by that parameter-generating …


Targeted Maximum Likelihood Estimation For Dynamic And Static Longitudinal Marginal Structural Working Models, Maya L. Petersen, Joshua Schwab, Susan Gruber, Nello Blaser, Michael Schomaker, Mark J. Van Der Laan May 2013

Targeted Maximum Likelihood Estimation For Dynamic And Static Longitudinal Marginal Structural Working Models, Maya L. Petersen, Joshua Schwab, Susan Gruber, Nello Blaser, Michael Schomaker, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

This paper describes a targeted maximum likelihood estimator (TMLE) for the parameters of longitudinal static and dynamic marginal structural models. We consider a longitudinal data structure consisting of baseline covariates, time-dependent intervention nodes, intermediate time-dependent covariates, and a possibly time dependent outcome. The intervention nodes at each time point can include a binary treatment as well as a right-censoring indicator. Given a class of dynamic or static interventions, a marginal structural model is used to model the mean of the intervention specific counterfactual outcome as a function of the intervention, time point, and possibly a subset of baseline covariates. Because …


Estimating Effects On Rare Outcomes: Knowledge Is Power, Laura B. Balzer, Mark J. Van Der Laan May 2013

Estimating Effects On Rare Outcomes: Knowledge Is Power, Laura B. Balzer, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Many of the secondary outcomes in observational studies and randomized trials are rare. Methods for estimating causal effects and associations with rare outcomes, however, are limited, and this represents a missed opportunity for investigation. In this article, we construct a new targeted minimum loss-based estimator (TMLE) for the effect of an exposure or treatment on a rare outcome. We focus on the causal risk difference and statistical models incorporating bounds on the conditional risk of the outcome, given the exposure and covariates. By construction, the proposed estimator constrains the predicted outcomes to respect this model knowledge. Theoretically, this bounding provides …


Integrative Biomarker Identification And Classification Using High Throughput Assays, Pan Tong May 2013

Integrative Biomarker Identification And Classification Using High Throughput Assays, Pan Tong

Dissertations and Theses (Open Access)

It is well accepted that tumorigenesis is a multi-step procedure involving aberrant functioning of genes regulating cell proliferation, differentiation, apoptosis, genome stability, angiogenesis and motility. To obtain a full understanding of tumorigenesis, it is necessary to collect information on all aspects of cell activity. Recent advances in high throughput technologies allow biologists to generate massive amounts of data, more than might have been imagined decades ago. These advances have made it possible to launch comprehensive projects such as (TCGA) and (ICGC) which systematically characterize the molecular fingerprints of cancer cells using gene expression, methylation, copy number, microRNA and SNP microarrays …


Is Obesity Socially Contagious?, Ciani Jean Sparks Mar 2013

Is Obesity Socially Contagious?, Ciani Jean Sparks

Statistics

The main objective of this paper is to analyze three different articles that discuss whether obesity could be socially contagious. According to the World Health Organization in 2013, obesity is the fifth leading risk for deaths around the world. This disease has dramatically increased in the last decade, which has led scientists to believe there are other factors contributing to the epidemic besides genetics. The first article I analyzed, written by Nicholas Christakis and James Fowler, provided a logistic regression model to estimate the odds of a person becoming obese. The model included the explanatory variables: age, sex, education, smoking …


A Bayesian Regression Tree Approach To Identify The Effect Of Nanoparticles Properties On Toxicity Profiles, Cecile Low-Kam, Haiyuan Zhang, Zhaoxia Ji, Tian Xia, Jeffrey I. Zinc, Andre Nel, Donatello Telesca Mar 2013

A Bayesian Regression Tree Approach To Identify The Effect Of Nanoparticles Properties On Toxicity Profiles, Cecile Low-Kam, Haiyuan Zhang, Zhaoxia Ji, Tian Xia, Jeffrey I. Zinc, Andre Nel, Donatello Telesca

COBRA Preprint Series

We introduce a Bayesian multiple regression tree model to characterize relationships between physico-chemical properties of nanoparticles and their in-vitro toxicity over multiple doses and times of exposure. Unlike conventional models that rely on data summaries, our model solves the low sample size issue and avoids arbitrary loss of information by combining all measurements from a general exposure experiment across doses, times of exposure, and replicates. The proposed technique integrates Bayesian trees for modeling threshold effects and interactions, and penalized B-splines for dose and time-response surfaces smoothing. The resulting posterior distribution is sampled via a Markov Chain Monte Carlo algorithm. This …


Gulf-Wide Decreases In The Size Of Large Coastal Sharks Documented By Generations Of Fishermen, Sean P. Powers, F. Joel Frodrie, Steven B. Scyphers, J. Marcus Drymon, Robert L. Shipp, Gregory W. Stunz Jan 2013

Gulf-Wide Decreases In The Size Of Large Coastal Sharks Documented By Generations Of Fishermen, Sean P. Powers, F. Joel Frodrie, Steven B. Scyphers, J. Marcus Drymon, Robert L. Shipp, Gregory W. Stunz

University Faculty and Staff Publications

Large sharks are top predators in most coastal and marine ecosystems throughout the world, and evidence of their reduced prominence in marine ecosystems has been a serious concern for fisheries and ecosystem management. Unfortunately, quantitative data to document the extent, timing, and consequences of changes in shark populations are scarce, thwarting examination of long-term (decadal, century) trends, and reconstructions based on incomplete data sets have been the subject of debate. Absence of quantitative descriptors of past ecological conditions is a generic problem facing many fields of science but is particularly troublesome for fisheries scientists who must develop specific targets for …


An Analysis Of Risk Reduction Choices In Dcis Breast Cancer Patients, Lauren Soltesz Dec 2012

An Analysis Of Risk Reduction Choices In Dcis Breast Cancer Patients, Lauren Soltesz

Statistics

The main focus of this paper was to evaluate possible demographic and clinical characteristics associated with a woman’s choice of breast conserving surgery (BCS), unilateral mastectomy (ULM), or bilateral risk reduction mastectomy (BRRM). The cohort consisted of patients presenting to the City of Hope National Medical Center with ductal carcinoma in situ breast cancer who elected to have cancer directed surgery (N=305). Analyses to examine associations of patient characteristics with type of surgery were conducted using a multinomial logistic regression. Results showed that older women were more likely to choose breast conserving surgery over bilateral risk reduction mastectomy than younger …


Differential Patterns Of Interaction And Gaussian Graphical Models, Masanao Yajima, Donatello Telesca, Yuan Ji, Peter Muller Apr 2012

Differential Patterns Of Interaction And Gaussian Graphical Models, Masanao Yajima, Donatello Telesca, Yuan Ji, Peter Muller

COBRA Preprint Series

We propose a methodological framework to assess heterogeneous patterns of association amongst components of a random vector expressed as a Gaussian directed acyclic graph. The proposed framework is likely to be useful when primary interest focuses on potential contrasts characterizing the association structure between known subgroups of a given sample. We provide inferential frameworks as well as an efficient computational algorithm to fit such a model and illustrate its validity through a simulation. We apply the model to Reverse Phase Protein Array data on Acute Myeloid Leukemia patients to show the contrast of association structure between refractory patients and relapsed …


Alternatives To Mixture Model Analysis Of Correlated Binomial Data, N. Rao Chaganty, Roy Sabo, Yihao Deng Jan 2012

Alternatives To Mixture Model Analysis Of Correlated Binomial Data, N. Rao Chaganty, Roy Sabo, Yihao Deng

Mathematics & Statistics Faculty Publications

While univariate instances of binomial data are readily handled with generalized linear models, cases of multivariate or repeated measure binomial data are complicated by the possibility of correlated responses. Likelihood-based estimation can be applied by using mixture distribution models, though this approach can present computational challenges. The logistic transformation can be used to bypass these concerns and allow for alternative estimating procedures. One popular alternative is the generalized estimating equation (GEE) method, though systematic errors can lead to infeasible correlation estimates or nonconvergence problems. Our approach is the coupling of quasileast squares (QLSs) method with a rarely used matrix factorization, …


Analysis Of Binary Data Via Spatial-Temporal Autologistic Regression Models, Zilong Wang Jan 2012

Analysis Of Binary Data Via Spatial-Temporal Autologistic Regression Models, Zilong Wang

Theses and Dissertations--Statistics

Spatial-temporal autologistic models are useful models for binary data that are measured repeatedly over time on a spatial lattice. They can account for effects of potential covariates and spatial-temporal statistical dependence among the data. However, the traditional parametrization of spatial-temporal autologistic model presents difficulties in interpreting model parameters across varying levels of statistical dependence, where its non-negative autocovariates could bias the realizations toward 1. In order to achieve interpretable parameters, a centered spatial-temporal autologistic regression model has been developed. Two efficient statistical inference approaches, expectation-maximization pseudo-likelihood approach (EMPL) and Monte Carlo expectation-maximization likelihood approach (MCEML), have been proposed. Also, Bayesian …


Flexible Distributed Lag Models Using Random Functions With Application To Estimating Mortality Displacement From Heat-Related Deaths, Roger D. Peng Dec 2011

Flexible Distributed Lag Models Using Random Functions With Application To Estimating Mortality Displacement From Heat-Related Deaths, Roger D. Peng

Johns Hopkins University, Dept. of Biostatistics Working Papers

No abstract provided.


Development Of A Bayesian Joint Logistic Model To Better Study The Association Between Haplotypes And Disease, Anthony M. D'Amelio Jr Dec 2011

Development Of A Bayesian Joint Logistic Model To Better Study The Association Between Haplotypes And Disease, Anthony M. D'Amelio Jr

Dissertations and Theses (Open Access)

In 2011, there will be an estimated 1,596,670 new cancer cases and 571,950 cancer-related deaths in the US. With the ever-increasing applications of cancer genetics in epidemiology, there is great potential to identify genetic risk factors that would help identify individuals with increased genetic susceptibility to cancer, which could be used to develop interventions or targeted therapies that could hopefully reduce cancer risk and mortality.

In this dissertation, I propose to develop a new statistical method to evaluate the role of haplotypes in cancer susceptibility and development. This model will be flexible enough to handle not only haplotypes of any …


Depicting Estimates Using The Intercept In Meta-Regression Models: The Moving Constant Technique, Blair T. Johnson Dr., Tania B. Huedo-Medina Dr. Oct 2011

Depicting Estimates Using The Intercept In Meta-Regression Models: The Moving Constant Technique, Blair T. Johnson Dr., Tania B. Huedo-Medina Dr.

CHIP Documents

In any scientific discipline, the ability to portray research patterns graphically often aids greatly in interpreting a phenomenon. In part to depict phenomena, the statistics and capabilities of meta-analytic models have grown increasingly sophisticated. Accordingly, this article details how to move the constant in weighted meta-analysis regression models (viz. “meta-regression”) to illuminate the patterns in such models across a range of complexities. Although it is commonly ignored in practice, the constant (or intercept) in such models can be indispensible when it is not relegated to its usual static role. The moving constant technique makes possible estimates and confidence intervals at …