Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

2016

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 331 - 360 of 616

Full-Text Articles in Statistics and Probability

New Flexible Regression Models Generated By Gamma Random Variables With Censored Data, Elizabeth M. Hashimoto, Gauss M. Cordeiro, Edwin M. M. Ortega, Gholamhossein G. Hamedani May 2016

New Flexible Regression Models Generated By Gamma Random Variables With Censored Data, Elizabeth M. Hashimoto, Gauss M. Cordeiro, Edwin M. M. Ortega, Gholamhossein G. Hamedani

Mathematics, Statistics and Computer Science Faculty Research and Publications

We propose and study a new log-gamma Weibull regression model. We obtain explicit expressions for the raw and incomplete moments, quantile and generating functions and mean deviations of the log-gamma Weibull distribution. We demonstrate that the new regression model can be applied to censored data since it represents a parametric family of models which includes as sub-models several widely-known regression models and therefore can be used more effectively in the analysis of survival data. We obtain the maximum likelihood estimates of the model parameters by considering censored data and evaluate local influence on the estimates of the parameters by taking …


Statistical Modeling Of The Temporal Dynamics In A Large Scale-Citation Network, Luis Javier Ek Jr. May 2016

Statistical Modeling Of The Temporal Dynamics In A Large Scale-Citation Network, Luis Javier Ek Jr.

Graduate Theses and Dissertations

Citation Networks of papers are vast networks that grow over time. The manner or the form a citation network grows is not entirely a random process, but a preferential attachment relationship; highly cited papers are more likely to be cited by newly published papers. The result is a network whose degree distribution follows a power law. This growth of citation network of papers will be modeled with a negative binomial regression coupled with logistic growth and/or Cauchy distribution curve. Then a Barabasi-Albert model, based on the negative binomial models, and a combination of the Dirichlet distribution and multinomial will be …


Generalized Singular Value Decomposition With Additive Components, Stan Lipovetsky May 2016

Generalized Singular Value Decomposition With Additive Components, Stan Lipovetsky

Journal of Modern Applied Statistical Methods

The singular value decomposition (SVD) technique is extended to incorporate the additive components for approximation of a rectangular matrix by the outer products of vectors. While dual vectors of the regular SVD can be expressed one via linear transformation of the other, the modified SVD corresponds to the general linear transformation with the additive part. The method obtained can be related to the family of principal component and correspondence analyses, and can be reduced to an eigenproblem of a specific transformation of a data matrix. This technique is applied to constructing dual eigenvectors for data visualizing in a two dimensional …


Almost Unbiased Estimator Using Known Value Of Population Parameter(S) In Sample Surveys, Rajesh Singh, S.B. Gupta, Sachin Malik May 2016

Almost Unbiased Estimator Using Known Value Of Population Parameter(S) In Sample Surveys, Rajesh Singh, S.B. Gupta, Sachin Malik

Journal of Modern Applied Statistical Methods

An almost unbiased estimator using known value of some population parameter(s) is proposed. A class of estimators is defined which includes Singh and Solanki (2012) and Sahai and Ray (1980), Sisodiya and Dwivedi (1981), Singh, Cauhan, Sawan, and Smarandache (2007), Upadhyaya and Singh (1984), Singh and Tailor (2003) estimators. Under simple random sampling without replacement (SRSWOR) scheme the expressions for bias and mean square error (MSE) are derived. Numerical illustrations are given.


A Comparison Of Estimation Methods For Nonlinear Mixed-Effects Models Under Model Misspecification And Data Sparseness: A Simulation Study, Jeffrey R. Harring, Junhui Liu May 2016

A Comparison Of Estimation Methods For Nonlinear Mixed-Effects Models Under Model Misspecification And Data Sparseness: A Simulation Study, Jeffrey R. Harring, Junhui Liu

Journal of Modern Applied Statistical Methods

A Monte Carlo simulation is employed to investigate the performance of five estimation methods of nonlinear mixed effects models in terms of parameter recovery and efficiency of both regression coefficients and variance/covariance parameters under varying levels of data sparseness and model misspecification.


Variable Selection In Regression Using Multilayer Feedforward Network, Tejaswi S. Kamble, Dattatraya N. Kashid May 2016

Variable Selection In Regression Using Multilayer Feedforward Network, Tejaswi S. Kamble, Dattatraya N. Kashid

Journal of Modern Applied Statistical Methods

The selection of relevant variables in the model is one of the important problems in regression analysis. Recently, a few methods were developed based on a model free approach. A multilayer feedforward neural network model was proposed for developing variable selection in regression. A simulation study and real data were used for evaluating the performance of proposed method in the presence of outliers, and multicollinearity.


Jmasm39: Algorithm For Combining Robust And Bootstrap In Multiple Linear Model Regression (Sas), Wan Muhamad Amir, Mohamad Shafiq, Hanafi A.Rahim, Puspa Liza, Azlida Aleng, Zailani Abdullah May 2016

Jmasm39: Algorithm For Combining Robust And Bootstrap In Multiple Linear Model Regression (Sas), Wan Muhamad Amir, Mohamad Shafiq, Hanafi A.Rahim, Puspa Liza, Azlida Aleng, Zailani Abdullah

Journal of Modern Applied Statistical Methods

The aim of bootstrapping is to approximate the sampling distribution of some estimator. An algorithm for combining method is given in SAS, along with applications and visualizations.


Jmasm35: A Percentile-Based Power Method: Simulating Multivariate Non-Normal Continuous Distributions (Sas), Jennifer Koran, Todd C. Headrick May 2016

Jmasm35: A Percentile-Based Power Method: Simulating Multivariate Non-Normal Continuous Distributions (Sas), Jennifer Koran, Todd C. Headrick

Journal of Modern Applied Statistical Methods

The conventional power method transformation is a moment-matching technique that simulates non-normal distributions with controlled measures of skew and kurtosis. The percentile-based power method is an alternative that uses the percentiles of a distribution in lieu of moments. This article presents a SAS/IML macro that implements the percentile-based power method.


Some Contributions To Nonparametric And Semiparametric Inference For Clustered And Multistate Data., Sandipan Dutta May 2016

Some Contributions To Nonparametric And Semiparametric Inference For Clustered And Multistate Data., Sandipan Dutta

Electronic Theses and Dissertations

This dissertation is composed of research projects that involve methods which can be broadly classified as either nonparametric or semiparametric. Chapter 1 provides an introduction of the problems addressed in these projects, a brief review of the related works that have done so far, and an outline of the methods developed in this dissertation. Chapter 2 describes in details the first project which aims at developing a rank-sum test for clustered data where an outcome from group in a cluster is associated with the number of observations belonging to that group in that cluster. Chapter 3 proposes the use of …


Identification Of Biomarkers For The Overall Survival Of Ovarian Cancer Patients, Kristi Mai May 2016

Identification Of Biomarkers For The Overall Survival Of Ovarian Cancer Patients, Kristi Mai

Graduate Theses and Dissertations

Rapid advance in sequencing technology has led to genome-wide analysis of genetic and epigenetic features simultaneously, making it possible to understand the biological mechanisms underlying cancer initiation and progression. However, how to identify important prognostic features poses a great challenge for both statistical modeling and computing. In this thesis, a network-based approach is applied to the Cancer Genome Atlas (TCGA) ovarian cancer data to identify important genes related to the overall survival of ovarian cancer patients. In the first step, a stepwise correlation-based selector is used to reduce the dimensionality of TCGA data, by filtering out a large number of …


Self-Monitoring Practices, Attitudes, And Needs Of Individuals With Bipolar Disorder: Implications For The Design Of Technologies To Manage Mental Health, Elizabeth L. Murnane, Dan Cosley, Pamara Chang, Shion Guha, Ellen Frank, Geri K. Gay, Mark Matthews May 2016

Self-Monitoring Practices, Attitudes, And Needs Of Individuals With Bipolar Disorder: Implications For The Design Of Technologies To Manage Mental Health, Elizabeth L. Murnane, Dan Cosley, Pamara Chang, Shion Guha, Ellen Frank, Geri K. Gay, Mark Matthews

Mathematics, Statistics and Computer Science Faculty Research and Publications

Objective To understand self-monitoring strategies used independently of clinical treatment by individuals with bipolar disorder (BD), in order to recommend technology design principles to support mental health management.

Materials and Methods Participants with BD (N = 552) were recruited through the Depression and Bipolar Support Alliance, the International Bipolar Foundation, and WeSearchTogether.org to complete a survey of closed- and open-ended questions. In this study, we focus on descriptive results and qualitative analyses.

Results Individuals reported primarily self-monitoring items related to their bipolar disorder (mood, sleep, finances, exercise, and social interactions), with an increasing trend towards the use of digital …


Bases For Mckay Centralizer Algebras, Lucas Gagnon May 2016

Bases For Mckay Centralizer Algebras, Lucas Gagnon

Mathematics, Statistics, and Computer Science Honors Projects

The finite subgroups of the special unitary group SU2 have been classified to be isomorphic to one of the following groups: cyclic, binary dihedral, binary tetrahedral, binary octahedral, and binary icosahedral, of order n, 4n, 24, 48, and 120, respectively. Associated to each group is a representation graph, which by the McKay correspondence is a Dynkin diagram of type Aˆ n−1, Dˆ n+2, Eˆ 6, Eˆ 7, or Eˆ 8. The centralizer algebra Zk(G) = EndG(V ⊗k ) is the algebra of transformations that commute with G acting on the k-fold tensor product of the defining representation V = C …


Building Voters: Exploring Interdependent Preferences In Binary Contexts, Ian Calaway May 2016

Building Voters: Exploring Interdependent Preferences In Binary Contexts, Ian Calaway

Mathematics, Statistics, and Computer Science Honors Projects

In this thesis we develop a new method for constructing binary preference orders for given interdependent structures, called characters. We introduce the preference space, which is a vector space of preference vectors. The preference vectors correspond to binary preference orders. We show that the hyperoctahedral group, Z2 o Sn, describes the symmetries of binary preferences orders and then define an action of Z2 o Sn on our preference vectors. We find a natural basis for a preference space. These basis vectors are indexed by subsets of proposals. We show that when completely separable binary preference vectors are decomposed using this …


Takens Theorem With Singular Spectrum Analysis Applied To Noisy Time Series, Thomas K. Torku May 2016

Takens Theorem With Singular Spectrum Analysis Applied To Noisy Time Series, Thomas K. Torku

Electronic Theses and Dissertations

The evolution of big data has led to financial time series becoming increasingly complex, noisy, non-stationary and nonlinear. Takens theorem can be used to analyze and forecast nonlinear time series, but even small amounts of noise can hopelessly corrupt a Takens approach. In contrast, Singular Spectrum Analysis is an excellent tool for both forecasting and noise reduction. Fortunately, it is possible to combine the Takens approach with Singular Spectrum analysis (SSA), and in fact, estimation of key parameters in Takens theorem is performed with Singular Spectrum Analysis. In this thesis, we combine the denoising abilities of SSA with the Takens …


Integration Of Multi-Platform High-Dimensional Omic Data, Xuebei An May 2016

Integration Of Multi-Platform High-Dimensional Omic Data, Xuebei An

Dissertations and Theses (Open Access)

The development of high-throughput biotechnologies have made data accessible from different platforms, including RNA sequencing, copy number variation, DNA methylation, protein lysate arrays, etc. The high-dimensional omic data derived from different technological platforms have been extensively used to facilitate comprehensive understanding of disease mechanisms and to determine personalized health treatments. Although vital to the progress of clinical research, the high dimensional multi-platform data impose new challenges for data analysis. Numerous studies have been proposed to integrate multi-platform omic data; however, few have efficiently and simultaneously addressed the problems that arise from high dimensionality and complex correlations.

In my dissertation, I …


A Log Rank Test For Clustered Data Under Informative Within-Cluster Group Size., Mary Elizabeth Gregg May 2016

A Log Rank Test For Clustered Data Under Informative Within-Cluster Group Size., Mary Elizabeth Gregg

Electronic Theses and Dissertations

The log rank test is a popular nonparametric test for comparing the marginal survival distribution of two groups. When data are organized within clusters and the size of clusters or the distribution of group membership within a cluster is related to an outcome of interest, traditional methods of data analysis can be biased. In this thesis, we develop a within-cluster group weighted log rank test to compare marginal survival time distributions between groups from clustered data, correcting for cluster size and intra-cluster group size informativeness. The performance of this new test is compared with the unweighted and cluster-weighted log rank …


Integrated Analysis Of Mirna/Mrna Expression And Gene Methylation Using Sparse Canonical Correlation Analysis., Dake Yang May 2016

Integrated Analysis Of Mirna/Mrna Expression And Gene Methylation Using Sparse Canonical Correlation Analysis., Dake Yang

Electronic Theses and Dissertations

MicroRNAs (miRNAs) are a large number of small endogenous non-coding RNA molecules (18-25 nucleotides in length) which regulate expression of genes post-transcriptionally. While a variety of algorithms exist for determining the targets of miRNAs, they are generally based on sequence information and frequently produce lists consisting of thousands of genes. Canonical correlation analysis (CCA) is a multivariate statistical method that can be used to find linear relationships between two data sets, and here we apply CCA to find the linear combination of differentially expressed miRNAs and their corresponding target genes having maximal negative correlation. Due to the high dimensionality, sparse …


Propensity Score Methods : A Simulation And Case Study Involving Breast Cancer Patients., John Craycroft May 2016

Propensity Score Methods : A Simulation And Case Study Involving Breast Cancer Patients., John Craycroft

Electronic Theses and Dissertations

Observational data presents unique challenges for analysis that are not encountered with experimental data resulting from carefully designed randomized controlled trials. Selection bias and unbalanced treatment assignments can obscure estimations of treatment effects, making the process of causal inference from observational data highly problematic. In 1983, Paul Rosenbaum and Donald Rubin formalized an approach for analyzing observational data that adjusts treatment effect estimates for the set of non-treatment variables that are measured at baseline. The propensity score is the conditional probability of assignment to a treatment group given the covariates. Using this score, one may balance the covariates across treatment …


Semi-Parametric Methods For Personalized Treatment Selection And Multi-State Models., Chathura K. Siriwardhana May 2016

Semi-Parametric Methods For Personalized Treatment Selection And Multi-State Models., Chathura K. Siriwardhana

Electronic Theses and Dissertations

This dissertation contains three research projects on personalized medicine and a project on multi-state modelling. The idea behind personalized medicine is selecting the best treatment that maximizes interested clinical outcomes of an individual based on his or her genetic and genomic information. We propose a method for treatment assignment based on individual covariate information for a patient. Our method covers more than two treatments and it can be applied with a broad set of models and it has very desirable large sample properties. An empirical study using simulations and a real data analysis show the applicability of the proposed procedure. …


Inference For A Zero-Inflated Conway-Maxwell-Poisson Regression For Clustered Count Data., Hyoyoung Choo-Wosoba May 2016

Inference For A Zero-Inflated Conway-Maxwell-Poisson Regression For Clustered Count Data., Hyoyoung Choo-Wosoba

Electronic Theses and Dissertations

This dissertation is directed toward developing a statistical methodology with applications of the Conway-Maxwell-Poisson (CMP) distribution (Conway, R. W., and Maxwell, W. L., 1962) to count data. The count data for this dissertation exhibit three different characteristics: clustering, zero inflation, and dispersion. Clustering suggests that observations within clusters are correlated, and the zero inflation phenomenon occurs when the data exhibit excessive zero counts. Dispersion implies that the mean is greater/smaller than the variance unlike a Poisson distribution. The dissertation starts with an introduction of inference for a zero-inflated clustered count data in the first chapter. Then, it presents novel methodologies …


Risk Estimation Toward A Natural History Model For Low Grade Glioma Patients, Anh Thi Hoang Pham May 2016

Risk Estimation Toward A Natural History Model For Low Grade Glioma Patients, Anh Thi Hoang Pham

Graduate Theses and Dissertations

Glioma is a common type of primary brain tumor that represents 28% of all brain tumors and 80% of malignant tumors. According to a recent study by the Centers for Disease Control and Prevention (CDC), gliomas account for 53%, 35% and 29% of all brain tumors (68%, 74% and 81% of malignant brain tumors) among children (aged 0-14), teenagers (aged 15-19) and young adults, respectively. Gliomas are often diagnosed through radiological imaging and histopathology. There are two main groups of gliomas following World Health Organization’s classification: Low grade gliomas (LGG), or grade I and II gliomas; and high grade gliomas …


Spread Trading In Corn Futures Market, Ryan D. Napier May 2016

Spread Trading In Corn Futures Market, Ryan D. Napier

Graduate Theses and Dissertations

The non-linear relationship between old crop – new crop year spreads in corn futures market and stock-to-use (S-U) ratios published by the United States Department of Agriculture is analyzed. Using a non-linear logarithmic smooth transition regression (LSTR) model, we capture asymmetric market behaviors in high and low S-U regimes. Capturing this relationship and understanding the non-linear aspects of the relationship is of interest of grain merchandizers and speculators in the market. A spread trading strategy is simulated for the sample period, January 1985 through April 2015, to determine if the non-linear relationship is a profitable arbitrage opportunity in the market.


Separation Of Points And Interval Estimation In Mixed Dose-Response Curves With Selective Component Labeling, Darl D. Flake Ii May 2016

Separation Of Points And Interval Estimation In Mixed Dose-Response Curves With Selective Component Labeling, Darl D. Flake Ii

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

Dose-response experiments are those that involve giving subjects different amounts of a treatment and observing the outcome. For example, plants may be given fertilizer and their growth could be measured or cancer patients could be given different doses of chemotherapy and their response could be monitored. These experiments are used to understand the relationship between the amount of, and response to, the treatment. Logistic regression models are often used to summarize data from these types of experiments. The dose-response experiment that motivated this dissertation involved treating a grain-pest with a pesticide. Some of the beetles had genes that made them …


Multivariate Thinking In An Intro Stats Course – Is It Possible?, Beverly Wood May 2016

Multivariate Thinking In An Intro Stats Course – Is It Possible?, Beverly Wood

Publications

Many of our students have an intuitive sense that there is more to the story than univariate or bivariate data can tell us. We can acknowledge and encourage that habit of digging deeper by demonstrating some ways to look at additional variables. Simpson’s paradox and side-by-side scatter plots are ways to provide a glimpse of more complex analysis that are accessible to students in an introductory course with or without strong quantitative skills.


Differences In Perceived Importance Of Preventative Services And Healthcare Provider Trust Among Hispanics, Jonathan James Gore May 2016

Differences In Perceived Importance Of Preventative Services And Healthcare Provider Trust Among Hispanics, Jonathan James Gore

UNLV Theses, Dissertations, Professional Papers, and Capstones

The Hispanic population varies greatly in their risk factors, health outcomes and access to care by country of origin, level of education and language dominance (Vega & Amaro, 1994) (Fiscella, Franks, Doescher, & Saver, 2002b). The differences within the Hispanic population also extend to their knowledge and attitudes toward health choices and maintenance, where they receive their health information, and what they access to meet their health care needs. Subpopulations within the Hispanic community as defined by language dominance and nativity must be understood as separate and distinct so that the health needs of each can be adequately addressed. The …


To Dot Product Graphs And Beyond, Sean Bailey May 2016

To Dot Product Graphs And Beyond, Sean Bailey

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

We will introduce three new vector representations of graphs. These representations are based on relationships between the vectors that are used. Specifically, we will examine scenarios where we ignore specific relationships, where we consider if information is missing, and where we look for when the information in common is not of a specified amount.


Bayesian Models For Repeated Measures Data Using Markov Chain Monte Carlo Methods, Yuanzhi Li May 2016

Bayesian Models For Repeated Measures Data Using Markov Chain Monte Carlo Methods, Yuanzhi Li

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

Bayesian models for repeated measures data are fitted to three different data an analysis projects. Markov Chain Monte Carlo (MCMC) methodology is applied to each case with Gibbs sampling and/or an adaptive Metropolis-Hastings (MH) algorithm used to simulate the posterior distribution of parameters. We implement a Bayesian model with different variance-covariance structures to an audit fee data set. Block structures and linear models for variances are used to examine the linear trend and different behaviors before and after regulatory change during year 2004-2005. We proposed a Bayesian hierarchical model with latent teacher effects, to determine whether teacher professional development (PD) …


Methods For Dealing With Death And Missing Data, And For Standardizing Different Health Variables In Longitudinal Datasets: The Cardiovascular Health Study, Paula Diehr Apr 2016

Methods For Dealing With Death And Missing Data, And For Standardizing Different Health Variables In Longitudinal Datasets: The Cardiovascular Health Study, Paula Diehr

UW Biostatistics Working Paper Series

Longitudinal studies of older adults usually need to account for deaths and missing data. The study databases often include multiple health-related variables, whose trends over time are hard to compare because they were measured on different scales. Here we present a unified approach to these three problems that was developed and used in the Cardiovascular Health Study. Data were first transformed to a new scale that had integer/ratio properties, and on which “dead” logically takes the value zero. Missing data were then imputed on this new scale, using each person’s own data over time. Imputation could thus be informed by …


Blossom: A Language Built To Grow, Jeffrey Lyman Apr 2016

Blossom: A Language Built To Grow, Jeffrey Lyman

Mathematics, Statistics, and Computer Science Honors Projects

No abstract provided.


Aberrant Dna Methylation: Implications In Racial Health Disparity, Xuefeng Wang, Ping Ji, Yuanhao Zhang, Joseph F. Lacomb, Xinyu Tian, Ellen Li, Jennie L. Williams Apr 2016

Aberrant Dna Methylation: Implications In Racial Health Disparity, Xuefeng Wang, Ping Ji, Yuanhao Zhang, Joseph F. Lacomb, Xinyu Tian, Ellen Li, Jennie L. Williams

Department of Biomedical Engineering Faculty Publications

Background Incidence and mortality rates of colorectal carcinoma (CRC) are higher in African Americans (AAs) than in Caucasian Americans (CAs). Deficient micronutrient intake due to dietary restrictions in racial/ethnic populations can alter genetic and molecular profiles leading to dysregulated methylation patterns and the inheritance of somatic to germline mutations. Materials and Methods Total DNA and RNA samples of paired tumor and adjacent normal colon tissues were prepared from AA and CA CRC specimens. Reduced Representation Bisulfite Sequencing (RRBS) and RNA sequencing were employed to evaluate total genome methylation of 5’-regulatory regions and dysregulation of gene expression, respectively. Robust analysis was …