Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 6691 - 6720 of 12832

Full-Text Articles in Statistics and Probability

New Procedures Of Estimating Proportion And Sensitivity Using Randomized Response In A Dichotomous Finite Population, Tanveer A. Tarray, Housila P. Singh May 2016

New Procedures Of Estimating Proportion And Sensitivity Using Randomized Response In A Dichotomous Finite Population, Tanveer A. Tarray, Housila P. Singh

Journal of Modern Applied Statistical Methods

The problem of estimating the population proportion possessing a sensitive attribute using simple random sampling with replacement (SRSWR) is advocated. Two new procedures are proposed. The suggested models are more efficient than the Huang (2004) randomized response technique under some realistic conditions. Numerical and graphic illustrations are given.


Construction Of Pair-Wise Balanced Design, Rajarathinam Arunachalam, Mahalakshmi Sivasubramanian, Dilip Kumar Ghosh May 2016

Construction Of Pair-Wise Balanced Design, Rajarathinam Arunachalam, Mahalakshmi Sivasubramanian, Dilip Kumar Ghosh

Journal of Modern Applied Statistical Methods

A new procedure for construction of pair wise balanced design with equal replication and un-equal block sizes based on factorial design have been evolved. Numerical illustration also provided. It was found that the constructed pair wise balanced design was found to be universal optimal.


A Log Rank Test For Clustered Data Under Informative Within-Cluster Group Size., Mary Elizabeth Gregg May 2016

A Log Rank Test For Clustered Data Under Informative Within-Cluster Group Size., Mary Elizabeth Gregg

Electronic Theses and Dissertations

The log rank test is a popular nonparametric test for comparing the marginal survival distribution of two groups. When data are organized within clusters and the size of clusters or the distribution of group membership within a cluster is related to an outcome of interest, traditional methods of data analysis can be biased. In this thesis, we develop a within-cluster group weighted log rank test to compare marginal survival time distributions between groups from clustered data, correcting for cluster size and intra-cluster group size informativeness. The performance of this new test is compared with the unweighted and cluster-weighted log rank …


Integrated Analysis Of Mirna/Mrna Expression And Gene Methylation Using Sparse Canonical Correlation Analysis., Dake Yang May 2016

Integrated Analysis Of Mirna/Mrna Expression And Gene Methylation Using Sparse Canonical Correlation Analysis., Dake Yang

Electronic Theses and Dissertations

MicroRNAs (miRNAs) are a large number of small endogenous non-coding RNA molecules (18-25 nucleotides in length) which regulate expression of genes post-transcriptionally. While a variety of algorithms exist for determining the targets of miRNAs, they are generally based on sequence information and frequently produce lists consisting of thousands of genes. Canonical correlation analysis (CCA) is a multivariate statistical method that can be used to find linear relationships between two data sets, and here we apply CCA to find the linear combination of differentially expressed miRNAs and their corresponding target genes having maximal negative correlation. Due to the high dimensionality, sparse …


Propensity Score Methods : A Simulation And Case Study Involving Breast Cancer Patients., John Craycroft May 2016

Propensity Score Methods : A Simulation And Case Study Involving Breast Cancer Patients., John Craycroft

Electronic Theses and Dissertations

Observational data presents unique challenges for analysis that are not encountered with experimental data resulting from carefully designed randomized controlled trials. Selection bias and unbalanced treatment assignments can obscure estimations of treatment effects, making the process of causal inference from observational data highly problematic. In 1983, Paul Rosenbaum and Donald Rubin formalized an approach for analyzing observational data that adjusts treatment effect estimates for the set of non-treatment variables that are measured at baseline. The propensity score is the conditional probability of assignment to a treatment group given the covariates. Using this score, one may balance the covariates across treatment …


Semi-Parametric Methods For Personalized Treatment Selection And Multi-State Models., Chathura K. Siriwardhana May 2016

Semi-Parametric Methods For Personalized Treatment Selection And Multi-State Models., Chathura K. Siriwardhana

Electronic Theses and Dissertations

This dissertation contains three research projects on personalized medicine and a project on multi-state modelling. The idea behind personalized medicine is selecting the best treatment that maximizes interested clinical outcomes of an individual based on his or her genetic and genomic information. We propose a method for treatment assignment based on individual covariate information for a patient. Our method covers more than two treatments and it can be applied with a broad set of models and it has very desirable large sample properties. An empirical study using simulations and a real data analysis show the applicability of the proposed procedure. …


Inference For A Zero-Inflated Conway-Maxwell-Poisson Regression For Clustered Count Data., Hyoyoung Choo-Wosoba May 2016

Inference For A Zero-Inflated Conway-Maxwell-Poisson Regression For Clustered Count Data., Hyoyoung Choo-Wosoba

Electronic Theses and Dissertations

This dissertation is directed toward developing a statistical methodology with applications of the Conway-Maxwell-Poisson (CMP) distribution (Conway, R. W., and Maxwell, W. L., 1962) to count data. The count data for this dissertation exhibit three different characteristics: clustering, zero inflation, and dispersion. Clustering suggests that observations within clusters are correlated, and the zero inflation phenomenon occurs when the data exhibit excessive zero counts. Dispersion implies that the mean is greater/smaller than the variance unlike a Poisson distribution. The dissertation starts with an introduction of inference for a zero-inflated clustered count data in the first chapter. Then, it presents novel methodologies …


Data Driven Sample Generator Model With Application To Classification, Alvaro Emilio Ulloa Cerna May 2016

Data Driven Sample Generator Model With Application To Classification, Alvaro Emilio Ulloa Cerna

Mathematics & Statistics ETDs

Despite the rapidly growing interest, progress in the study of relations between physiological abnormalities and mental disorders is hampered by complexity of the human brain and high costs of data collection. The complexity can be captured by machine learning approaches, but they still may require significant amounts of data. In this thesis, we seek to mitigate the latter challenge by developing a data driven sample generator model for the generation of synthetic realistic training data. Our method greatly improves generalization in classification of schizophrenia patients and healthy controls from their structural magnetic resonance images. A feed forward neural network trained …


Generalized Linear Model Analyses For Treatment Group Equality When Data Are Non-Normal, Harvey J. Kesleman, Abdul R. Othman, Rand R. Wilcox May 2016

Generalized Linear Model Analyses For Treatment Group Equality When Data Are Non-Normal, Harvey J. Kesleman, Abdul R. Othman, Rand R. Wilcox

Journal of Modern Applied Statistical Methods

One of the validity conditions of classical test statistics (e.g., Student’s t-test, the ANOVA and MANOVA F-tests) is that data be normally distributed in the populations. When this and/or other derivational assumptions do not hold the classical test statistic can be prone to too many Type I errors (i.e., falsely rejecting too often) and/or have low power (i.e., failing to reject when the null hypothesis is false) to detect treatment effects when they are present. However, alternative procedures are available for assessing equality of treatment group effects when data are non-normal. For example, researchers can use robust estimators …


Liu-Type Logistic Estimators With Optimal Shrinkage Parameter, Yasin Asar May 2016

Liu-Type Logistic Estimators With Optimal Shrinkage Parameter, Yasin Asar

Journal of Modern Applied Statistical Methods

Multicollinearity in logistic regression affects the variance of the maximum likelihood estimator negatively. In this study, Liu-type estimators are used to reduce the variance and overcome the multicollinearity by applying some existing ridge regression estimators to the case of logistic regression model. A Monte Carlo simulation is given to evaluate the performances of these estimators when the optimal shrinkage parameter is used in the Liu-type estimators, along with an application of real case data.


The Xgamma Distribution: Statistical Properties And Application, Subhradev Sen, Sudhansu S. Maiti, N. Chandra May 2016

The Xgamma Distribution: Statistical Properties And Application, Subhradev Sen, Sudhansu S. Maiti, N. Chandra

Journal of Modern Applied Statistical Methods

A new probability distribution, the xgamma distribution, is proposed and studied. The distribution is generated as a special finite mixture of exponential and gamma distributions and hence the name proposed. Various mathematical, structural, and survival properties of the xgamma distribution are derived, and it is found that in many cases the xgamma has more flexibility than the exponential distribution. To evaluate the comparative behavior, stochastic ordering of the distribution is studied. To estimate the model parameter, the method of moment and the method of maximum likelihood estimation are proposed. A simulation algorithm to generate random samples from the xgamma distribution …


Analysis And Modeling Of Statistical Properties Of Fmdfb Subband Coefficients, E. Jebamalar Leavline, Sutha Shunmugam May 2016

Analysis And Modeling Of Statistical Properties Of Fmdfb Subband Coefficients, E. Jebamalar Leavline, Sutha Shunmugam

Journal of Modern Applied Statistical Methods

Fast Multiscale Directional Filter Bank (FMDFB) is an image representation scheme used in several image processing applications. The statistical nature of the FMDFB subbands is analyzed, and a mathematical model of FMDFB coefficients is proposed. Experimental results are justified by goodness-of-fit tests.


Jmasm37: Simple Response Surface Methodology Using Rsreg (Sas), Wan Muhamad Amir, Mohamad Shafiq, Kasypi Mokhtar, Nor Azlida Aleng, Hanafi A.Rahim, Zalila Ali May 2016

Jmasm37: Simple Response Surface Methodology Using Rsreg (Sas), Wan Muhamad Amir, Mohamad Shafiq, Kasypi Mokhtar, Nor Azlida Aleng, Hanafi A.Rahim, Zalila Ali

Journal of Modern Applied Statistical Methods

Response surface methodology (RSM) can be used when the response variable, y, is influenced by several variables, x’s. When treatments take the form of quantitative values, then the true relationship between response variables and independent variables might be known. Examples are given in SAS.


Statistical Modeling Of The Temporal Dynamics In A Large Scale-Citation Network, Luis Javier Ek Jr. May 2016

Statistical Modeling Of The Temporal Dynamics In A Large Scale-Citation Network, Luis Javier Ek Jr.

Graduate Theses and Dissertations

Citation Networks of papers are vast networks that grow over time. The manner or the form a citation network grows is not entirely a random process, but a preferential attachment relationship; highly cited papers are more likely to be cited by newly published papers. The result is a network whose degree distribution follows a power law. This growth of citation network of papers will be modeled with a negative binomial regression coupled with logistic growth and/or Cauchy distribution curve. Then a Barabasi-Albert model, based on the negative binomial models, and a combination of the Dirichlet distribution and multinomial will be …


Spread Trading In Corn Futures Market, Ryan D. Napier May 2016

Spread Trading In Corn Futures Market, Ryan D. Napier

Graduate Theses and Dissertations

The non-linear relationship between old crop – new crop year spreads in corn futures market and stock-to-use (S-U) ratios published by the United States Department of Agriculture is analyzed. Using a non-linear logarithmic smooth transition regression (LSTR) model, we capture asymmetric market behaviors in high and low S-U regimes. Capturing this relationship and understanding the non-linear aspects of the relationship is of interest of grain merchandizers and speculators in the market. A spread trading strategy is simulated for the sample period, January 1985 through April 2015, to determine if the non-linear relationship is a profitable arbitrage opportunity in the market.


Generalized Singular Value Decomposition With Additive Components, Stan Lipovetsky May 2016

Generalized Singular Value Decomposition With Additive Components, Stan Lipovetsky

Journal of Modern Applied Statistical Methods

The singular value decomposition (SVD) technique is extended to incorporate the additive components for approximation of a rectangular matrix by the outer products of vectors. While dual vectors of the regular SVD can be expressed one via linear transformation of the other, the modified SVD corresponds to the general linear transformation with the additive part. The method obtained can be related to the family of principal component and correspondence analyses, and can be reduced to an eigenproblem of a specific transformation of a data matrix. This technique is applied to constructing dual eigenvectors for data visualizing in a two dimensional …


Almost Unbiased Estimator Using Known Value Of Population Parameter(S) In Sample Surveys, Rajesh Singh, S.B. Gupta, Sachin Malik May 2016

Almost Unbiased Estimator Using Known Value Of Population Parameter(S) In Sample Surveys, Rajesh Singh, S.B. Gupta, Sachin Malik

Journal of Modern Applied Statistical Methods

An almost unbiased estimator using known value of some population parameter(s) is proposed. A class of estimators is defined which includes Singh and Solanki (2012) and Sahai and Ray (1980), Sisodiya and Dwivedi (1981), Singh, Cauhan, Sawan, and Smarandache (2007), Upadhyaya and Singh (1984), Singh and Tailor (2003) estimators. Under simple random sampling without replacement (SRSWOR) scheme the expressions for bias and mean square error (MSE) are derived. Numerical illustrations are given.


A Comparison Of Estimation Methods For Nonlinear Mixed-Effects Models Under Model Misspecification And Data Sparseness: A Simulation Study, Jeffrey R. Harring, Junhui Liu May 2016

A Comparison Of Estimation Methods For Nonlinear Mixed-Effects Models Under Model Misspecification And Data Sparseness: A Simulation Study, Jeffrey R. Harring, Junhui Liu

Journal of Modern Applied Statistical Methods

A Monte Carlo simulation is employed to investigate the performance of five estimation methods of nonlinear mixed effects models in terms of parameter recovery and efficiency of both regression coefficients and variance/covariance parameters under varying levels of data sparseness and model misspecification.


Variable Selection In Regression Using Multilayer Feedforward Network, Tejaswi S. Kamble, Dattatraya N. Kashid May 2016

Variable Selection In Regression Using Multilayer Feedforward Network, Tejaswi S. Kamble, Dattatraya N. Kashid

Journal of Modern Applied Statistical Methods

The selection of relevant variables in the model is one of the important problems in regression analysis. Recently, a few methods were developed based on a model free approach. A multilayer feedforward neural network model was proposed for developing variable selection in regression. A simulation study and real data were used for evaluating the performance of proposed method in the presence of outliers, and multicollinearity.


Jmasm39: Algorithm For Combining Robust And Bootstrap In Multiple Linear Model Regression (Sas), Wan Muhamad Amir, Mohamad Shafiq, Hanafi A.Rahim, Puspa Liza, Azlida Aleng, Zailani Abdullah May 2016

Jmasm39: Algorithm For Combining Robust And Bootstrap In Multiple Linear Model Regression (Sas), Wan Muhamad Amir, Mohamad Shafiq, Hanafi A.Rahim, Puspa Liza, Azlida Aleng, Zailani Abdullah

Journal of Modern Applied Statistical Methods

The aim of bootstrapping is to approximate the sampling distribution of some estimator. An algorithm for combining method is given in SAS, along with applications and visualizations.


Jmasm35: A Percentile-Based Power Method: Simulating Multivariate Non-Normal Continuous Distributions (Sas), Jennifer Koran, Todd C. Headrick May 2016

Jmasm35: A Percentile-Based Power Method: Simulating Multivariate Non-Normal Continuous Distributions (Sas), Jennifer Koran, Todd C. Headrick

Journal of Modern Applied Statistical Methods

The conventional power method transformation is a moment-matching technique that simulates non-normal distributions with controlled measures of skew and kurtosis. The percentile-based power method is an alternative that uses the percentiles of a distribution in lieu of moments. This article presents a SAS/IML macro that implements the percentile-based power method.


New Flexible Regression Models Generated By Gamma Random Variables With Censored Data, Elizabeth M. Hashimoto, Gauss M. Cordeiro, Edwin M. M. Ortega, Gholamhossein G. Hamedani May 2016

New Flexible Regression Models Generated By Gamma Random Variables With Censored Data, Elizabeth M. Hashimoto, Gauss M. Cordeiro, Edwin M. M. Ortega, Gholamhossein G. Hamedani

Mathematics, Statistics and Computer Science Faculty Research and Publications

We propose and study a new log-gamma Weibull regression model. We obtain explicit expressions for the raw and incomplete moments, quantile and generating functions and mean deviations of the log-gamma Weibull distribution. We demonstrate that the new regression model can be applied to censored data since it represents a parametric family of models which includes as sub-models several widely-known regression models and therefore can be used more effectively in the analysis of survival data. We obtain the maximum likelihood estimates of the model parameters by considering censored data and evaluate local influence on the estimates of the parameters by taking …


Self-Monitoring Practices, Attitudes, And Needs Of Individuals With Bipolar Disorder: Implications For The Design Of Technologies To Manage Mental Health, Elizabeth L. Murnane, Dan Cosley, Pamara Chang, Shion Guha, Ellen Frank, Geri K. Gay, Mark Matthews May 2016

Self-Monitoring Practices, Attitudes, And Needs Of Individuals With Bipolar Disorder: Implications For The Design Of Technologies To Manage Mental Health, Elizabeth L. Murnane, Dan Cosley, Pamara Chang, Shion Guha, Ellen Frank, Geri K. Gay, Mark Matthews

Mathematics, Statistics and Computer Science Faculty Research and Publications

Objective To understand self-monitoring strategies used independently of clinical treatment by individuals with bipolar disorder (BD), in order to recommend technology design principles to support mental health management.

Materials and Methods Participants with BD (N = 552) were recruited through the Depression and Bipolar Support Alliance, the International Bipolar Foundation, and WeSearchTogether.org to complete a survey of closed- and open-ended questions. In this study, we focus on descriptive results and qualitative analyses.

Results Individuals reported primarily self-monitoring items related to their bipolar disorder (mood, sleep, finances, exercise, and social interactions), with an increasing trend towards the use of digital …


Differences In Perceived Importance Of Preventative Services And Healthcare Provider Trust Among Hispanics, Jonathan James Gore May 2016

Differences In Perceived Importance Of Preventative Services And Healthcare Provider Trust Among Hispanics, Jonathan James Gore

UNLV Theses, Dissertations, Professional Papers, and Capstones

The Hispanic population varies greatly in their risk factors, health outcomes and access to care by country of origin, level of education and language dominance (Vega & Amaro, 1994) (Fiscella, Franks, Doescher, & Saver, 2002b). The differences within the Hispanic population also extend to their knowledge and attitudes toward health choices and maintenance, where they receive their health information, and what they access to meet their health care needs. Subpopulations within the Hispanic community as defined by language dominance and nativity must be understood as separate and distinct so that the health needs of each can be adequately addressed. The …


Some Contributions To Nonparametric And Semiparametric Inference For Clustered And Multistate Data., Sandipan Dutta May 2016

Some Contributions To Nonparametric And Semiparametric Inference For Clustered And Multistate Data., Sandipan Dutta

Electronic Theses and Dissertations

This dissertation is composed of research projects that involve methods which can be broadly classified as either nonparametric or semiparametric. Chapter 1 provides an introduction of the problems addressed in these projects, a brief review of the related works that have done so far, and an outline of the methods developed in this dissertation. Chapter 2 describes in details the first project which aims at developing a rank-sum test for clustered data where an outcome from group in a cluster is associated with the number of observations belonging to that group in that cluster. Chapter 3 proposes the use of …


Identification Of Biomarkers For The Overall Survival Of Ovarian Cancer Patients, Kristi Mai May 2016

Identification Of Biomarkers For The Overall Survival Of Ovarian Cancer Patients, Kristi Mai

Graduate Theses and Dissertations

Rapid advance in sequencing technology has led to genome-wide analysis of genetic and epigenetic features simultaneously, making it possible to understand the biological mechanisms underlying cancer initiation and progression. However, how to identify important prognostic features poses a great challenge for both statistical modeling and computing. In this thesis, a network-based approach is applied to the Cancer Genome Atlas (TCGA) ovarian cancer data to identify important genes related to the overall survival of ovarian cancer patients. In the first step, a stepwise correlation-based selector is used to reduce the dimensionality of TCGA data, by filtering out a large number of …


Bases For Mckay Centralizer Algebras, Lucas Gagnon May 2016

Bases For Mckay Centralizer Algebras, Lucas Gagnon

Mathematics, Statistics, and Computer Science Honors Projects

The finite subgroups of the special unitary group SU2 have been classified to be isomorphic to one of the following groups: cyclic, binary dihedral, binary tetrahedral, binary octahedral, and binary icosahedral, of order n, 4n, 24, 48, and 120, respectively. Associated to each group is a representation graph, which by the McKay correspondence is a Dynkin diagram of type Aˆ n−1, Dˆ n+2, Eˆ 6, Eˆ 7, or Eˆ 8. The centralizer algebra Zk(G) = EndG(V ⊗k ) is the algebra of transformations that commute with G acting on the k-fold tensor product of the defining representation V = C …


Building Voters: Exploring Interdependent Preferences In Binary Contexts, Ian Calaway May 2016

Building Voters: Exploring Interdependent Preferences In Binary Contexts, Ian Calaway

Mathematics, Statistics, and Computer Science Honors Projects

In this thesis we develop a new method for constructing binary preference orders for given interdependent structures, called characters. We introduce the preference space, which is a vector space of preference vectors. The preference vectors correspond to binary preference orders. We show that the hyperoctahedral group, Z2 o Sn, describes the symmetries of binary preferences orders and then define an action of Z2 o Sn on our preference vectors. We find a natural basis for a preference space. These basis vectors are indexed by subsets of proposals. We show that when completely separable binary preference vectors are decomposed using this …


Takens Theorem With Singular Spectrum Analysis Applied To Noisy Time Series, Thomas K. Torku May 2016

Takens Theorem With Singular Spectrum Analysis Applied To Noisy Time Series, Thomas K. Torku

Electronic Theses and Dissertations

The evolution of big data has led to financial time series becoming increasingly complex, noisy, non-stationary and nonlinear. Takens theorem can be used to analyze and forecast nonlinear time series, but even small amounts of noise can hopelessly corrupt a Takens approach. In contrast, Singular Spectrum Analysis is an excellent tool for both forecasting and noise reduction. Fortunately, it is possible to combine the Takens approach with Singular Spectrum analysis (SSA), and in fact, estimation of key parameters in Takens theorem is performed with Singular Spectrum Analysis. In this thesis, we combine the denoising abilities of SSA with the Takens …


Integration Of Multi-Platform High-Dimensional Omic Data, Xuebei An May 2016

Integration Of Multi-Platform High-Dimensional Omic Data, Xuebei An

Dissertations and Theses (Open Access)

The development of high-throughput biotechnologies have made data accessible from different platforms, including RNA sequencing, copy number variation, DNA methylation, protein lysate arrays, etc. The high-dimensional omic data derived from different technological platforms have been extensively used to facilitate comprehensive understanding of disease mechanisms and to determine personalized health treatments. Although vital to the progress of clinical research, the high dimensional multi-platform data impose new challenges for data analysis. Numerous studies have been proposed to integrate multi-platform omic data; however, few have efficiently and simultaneously addressed the problems that arise from high dimensionality and complex correlations.

In my dissertation, I …