Open Access. Powered by Scholars. Published by Universities.®

Statistical Theory Commons

Open Access. Powered by Scholars. Published by Universities.®

1,633 Full-Text Articles 2,146 Authors 1,950,602 Downloads 69 Institutions

All Articles in Statistical Theory

Faceted Search

1,633 full-text articles. Page 41 of 45.

A Generalized Approach For Testing The Association Of A Set Of Predictors With An Outcome: A Gene Based Test, Benjamin A. Goldstein, Alan E. Hubbard, Lisa F. Barcellos 2011 University of California - Berkeley

A Generalized Approach For Testing The Association Of A Set Of Predictors With An Outcome: A Gene Based Test, Benjamin A. Goldstein, Alan E. Hubbard, Lisa F. Barcellos

U.C. Berkeley Division of Biostatistics Working Paper Series

In many analyses, one has data on one level but desires to draw inference on another level. For example, in genetic association studies, one observes units of DNA referred to as SNPs, but wants to determine whether genes that are comprised of SNPs are associated with disease. While there are some available approaches for addressing this issue, they usually involve making parametric assumptions and are not easily generalizable. A statistical test is proposed for testing the association of a set of variables with an outcome of interest. No assumptions are made about the functional form relating the variables to the …


Flipping The Winner Of A Poset Game, Adam O. Kalinich '12 2011 Illinois Mathematics and Science Academy

Flipping The Winner Of A Poset Game, Adam O. Kalinich '12

Student Publications & Research

Partially-ordered set games, also called poset games, are a class of two-player combinatorial games. The playing field consists of a set of elements, some of which are greater than other elements. Two players take turns removing an element and all elements greater than it, and whoever takes the last element wins. Examples of poset games include Nim and Chomp. We investigate the complexity of computing which player of a poset game has a winning strategy. We give an inductive procedure that modifies poset games to change the nim-value which informally captures the winning strategies in the game. For a generic …


Parametric Estimation In Competing Risks And Multi-State Models, Yushun Lin 2011 University of Kentucky

Parametric Estimation In Competing Risks And Multi-State Models, Yushun Lin

Theses and Dissertations--Statistics

The typical research of Alzheimer's disease includes a series of cognitive states. Multi-state models are often used to describe the history of disease evolvement. Competing risks models are a sub-category of multi-state models with one starting state and several absorbing states.

Analyses for competing risks data in medical papers frequently assume independent risks and evaluate covariate effects on these events by modeling distinct proportional hazards regression models for each event. Jeong and Fine (2007) proposed a parametric proportional sub-distribution hazard (SH) model for cumulative incidence functions (CIF) without assumptions about the dependence among the risks. We modified their model to …


Oracle And Multiple Robustness Properties Of Survey Calibration Estimator In Missing Response Problem, Kwun Chuen Gary Chan 2010 University of Washington

Oracle And Multiple Robustness Properties Of Survey Calibration Estimator In Missing Response Problem, Kwun Chuen Gary Chan

UW Biostatistics Working Paper Series

In the presence of missing response, reweighting the complete case subsample by the inverse of nonmissing probability is both intuitive and easy to implement. However, inverse probability weighting is not efficient in general and is not robust against misspecification of the missing probability model. Calibration was developed by survey statisticians for improving efficiency of inverse probability weighting estimators when population totals of auxiliary variables are known and when inclusion probability is known by design. In missing data problem we can calibrate auxiliary variables in the complete case subsample to the full sample. However, the inclusion probability is unknown in general …


Modification And Improvement Of Empirical Likelihood For Missing Response Problem, Kwun Chuen Gary Chan 2010 University of Washington - Seattle Campus

Modification And Improvement Of Empirical Likelihood For Missing Response Problem, Kwun Chuen Gary Chan

UW Biostatistics Working Paper Series

An empirical likelihood (EL) estimator was proposed by Qin and Zhang (2007) for a missing response problem under a missing at random assumption. They showed by simulation studies that the finite sample performance of EL estimator is better than some existing estimators. However, the empirical likelihood estimator does not have a uniformly smaller asymptotic variance than other estimators in general. We consider several modifications to the empirical likelihood estimator and show that the proposed estimator dominates the empirical likelihood estimator and several other existing estimators in terms of asymptotic efficiencies. The proposed estimator also attains the minimum asymptotic variance among …


Modification And Improvement Of Empirical Liklihood For Missing Response Problem, Gary Chan 2010 University of Washington

Modification And Improvement Of Empirical Liklihood For Missing Response Problem, Gary Chan

UW Biostatistics Working Paper Series

An empirical likelihood (EL) estimator was proposed by Qin and Zhang (2007) for a missing response problem under a missing at random assumption. They showed by simulation studies that the finite sample performance of EL estimator is better than some existing estimators. However, the empirical likelihood estimator does not have a uniformly smaller asymptotic variance than other estimators in general. We consider several modifications to the empirical likelihood estimator and show that the proposed estimator dominates the empirical likelihood estimator and several other existing estimators in terms of asymptotic efficiencies. The proposed estimator also attains the minimum asymptotic variance among …


Minimum Description Length Measures Of Evidence For Enrichment, Zhenyu Yang, David R. Bickel 2010 Ottawa Institute of Systems Biology, Department of Biochemistry, Microbiology, and Immunology, University of Ottawa

Minimum Description Length Measures Of Evidence For Enrichment, Zhenyu Yang, David R. Bickel

COBRA Preprint Series

In order to functionally interpret differentially expressed genes or other discovered features, researchers seek to detect enrichment in the form of overrepresentation of discovered features associated with a biological process. Most enrichment methods treat the p-value as the measure of evidence using a statistical test such as the binomial test, Fisher's exact test or the hypergeometric test. However, the p-value is not interpretable as a measure of evidence apart from adjustments in light of the sample size. As a measure of evidence supporting one hypothesis over the other, the Bayes factor (BF) overcomes this drawback of the p-value but lacks …


Efficient Measurement Error Correction With Spatially Misaligned Data, Adam A. Szpiro, Lianne Sheppard, Thomas Lumley 2010 University of Washington

Efficient Measurement Error Correction With Spatially Misaligned Data, Adam A. Szpiro, Lianne Sheppard, Thomas Lumley

UW Biostatistics Working Paper Series

Association studies in environmental statistics often involve exposure and outcome data that are misaligned in space. A common strategy is to employ a spatial model such as universal kriging to predict exposures at locations with outcome data and then estimate a regression parameter of interest using the predicted exposures. This results in measurement error because the predicted exposures do not correspond exactly to the true values. We characterize the measurement error by decomposing it into Berkson-like and classical-like components. One correction approach is the parametric bootstrap, which is effective but computationally intensive since it requires solving a nonlinear optimization problem …


Minimum Description Length And Empirical Bayes Methods Of Identifying Snps Associated With Disease, Ye Yang, David R. Bickel 2010 Ottawa Institute of Systems Biology, Department of Biochemistry, Microbiology, and Immunology

Minimum Description Length And Empirical Bayes Methods Of Identifying Snps Associated With Disease, Ye Yang, David R. Bickel

COBRA Preprint Series

The goal of determining which of hundreds of thousands of SNPs are associated with disease poses one of the most challenging multiple testing problems. Using the empirical Bayes approach, the local false discovery rate (LFDR) estimated using popular semiparametric models has enjoyed success in simultaneous inference. However, the estimated LFDR can be biased because the semiparametric approach tends to overestimate the proportion of the non-associated single nucleotide polymorphisms (SNPs). One of the negative consequences is that, like conventional p-values, such LFDR estimates cannot quantify the amount of information in the data that favors the null hypothesis of no disease-association.

We …


Improving The Power Of Chronic Disease Surveillance By Incorporating Residential History, Justin Manjourides, Marcello Pagano 2010 Harvard School of Public Health

Improving The Power Of Chronic Disease Surveillance By Incorporating Residential History, Justin Manjourides, Marcello Pagano

Harvard University Biostatistics Working Paper Series

No abstract provided.


Generalized Variances Ratio Test For Comparing K Covariance Matrices From Dependent Normal Populations, Marcelo Angelo Cirillo, Daniel Furtado Ferreira, Thelma Sáfadi, Eric Batista Ferreira 2010 Federal University of Lavras, Brazil

Generalized Variances Ratio Test For Comparing K Covariance Matrices From Dependent Normal Populations, Marcelo Angelo Cirillo, Daniel Furtado Ferreira, Thelma Sáfadi, Eric Batista Ferreira

Journal of Modern Applied Statistical Methods

New tests based on the ratio of generalized variances are presented to compare covariance matrices from dependent normal populations. Monte Carlo simulation concluded that the tests considered controlled the Type I error, providing empirical probabilities that were consistent with the nominal level stipulated.


A Ga-Based Sales Forecasting Model Incorporating Promotion Factors, Li-Chih Wang, Chin-Lien Wang 2010 Tunghai University, Taichung, Taiwan ROC

A Ga-Based Sales Forecasting Model Incorporating Promotion Factors, Li-Chih Wang, Chin-Lien Wang

Journal of Modern Applied Statistical Methods

Because promotions are critical factors highly related to product sales of consumer packaged goods (CPG) companies, predictors concerning sales forecast of CPG products must take promotions into consideration. Decomposition regression incorporating contextual factors offers a method for exploiting both reliability of statistical forecasting and flexibility of judgmental forecasting employing domain knowledge. However, it suffers from collinearity causing poor performance in variable identification and parameter estimation with traditional ordinary least square (OLS). Empirical research evidence shows that - in the case of collinearity - in variable identification, parameter estimation, and out of sample forecasting, genetic algorithms (GA) as an estimator outperform …


Estimating The Non-Existent Mean And Variance Of The F-Distribution By Simulation, Hamid Reza Kamali, Parisa Shahnazari-Shahrezaei 2010 Private Scholar

Estimating The Non-Existent Mean And Variance Of The F-Distribution By Simulation, Hamid Reza Kamali, Parisa Shahnazari-Shahrezaei

Journal of Modern Applied Statistical Methods

In theory, all moments of some probability distributions do not necessarily exist. In the other words, they may be infinite or undefined. One of these distributions is the F-distribution whose mean and variance have not been defined for the second degree of freedom less than 3 and 5, respectively. In some cases, a large statistical population having an F-distribution may exist and the aim is to obtain its mean and variance which are an estimation of the non-existent mean and variance of F-distribution. This article considers a large sample F-distribution to estimate its non-existent mean and variance using Simul8 simulation …


Ridge Regression Based On Some Robust Estimators, Hatice Samkar, Ozlem Alpu 2010 Eskisehir Osmangazi University, Turkey

Ridge Regression Based On Some Robust Estimators, Hatice Samkar, Ozlem Alpu

Journal of Modern Applied Statistical Methods

Robust ridge methods based on M, S, MM and GM estimators are examined in the presence of multicollinearity and outliers. GMWalker, using the LS estimator as the initial estimator is used. S and MM estimators are also used as initial estimators with the aim of evaluating the two alternatives as biased robust methods.


A Flexible Method For Testing Independence In Two-Way Contingency Tables, Peyman Jafari, Noori Akhtar-Danesh, Zahra Bagheri 2010 Shiraz University of Medical Sciences, Shiraz, Iran

A Flexible Method For Testing Independence In Two-Way Contingency Tables, Peyman Jafari, Noori Akhtar-Danesh, Zahra Bagheri

Journal of Modern Applied Statistical Methods

A flexible approach for testing association in two-way contingency tables is presented. It is simple, does not assume a specific form for the association and is applicable to tables with nominal-by-nominal, nominal-by-ordinal, and ordinal-by-ordinal classifications.


Statistical And Mathematical Modeling Versus Nhst? There’S No Competition!, Joseph Lee Rodgers 2010 University of Oklahoma

Statistical And Mathematical Modeling Versus Nhst? There’S No Competition!, Joseph Lee Rodgers

Journal of Modern Applied Statistical Methods

Some of Robinson & Levin’s critique of Rodgers (2010) is cogent, helpful, and insightful – although limiting. Recent methodology has advanced through the development of structural equation modeling, multi-level modeling, missing data methods, hierarchical linear modeling, categorical data analysis, as well as the development of many dedicated and specific behavioral models. These methodological approaches are based on a revised epistemological system, and have emerged naturally, without the need for task forces, or even much self-conscious discussion. The original goal was neither to develop nor promote a modeling revolution. That has occurred; I documented its development and its status. Two organizing …


Effect Of Measurement Errors On The Separate And Combined Ratio And Product Estimators In Stratified Random Sampling, Housila P. Singh, Namrata Karpe 2010 Vikram University, Ujjain, India

Effect Of Measurement Errors On The Separate And Combined Ratio And Product Estimators In Stratified Random Sampling, Housila P. Singh, Namrata Karpe

Journal of Modern Applied Statistical Methods

Separate and combined ratio, product and difference estimators are introduced for population mean μY of a study variable Y using auxiliary variable X in stratified sampling when the observations are contaminated with measurement errors. The bias and mean squared error of the proposed estimators have been derived under large sample approximation and their properties are analyzed. Generalized versions of these estimators are given along with their properties.


Recommended Sample Size For Conducting Exploratory Factor Analysis On Dichotomous Data, Robert H. Pearson, Daniel J. Mundform 2010 University of Northern Colorado

Recommended Sample Size For Conducting Exploratory Factor Analysis On Dichotomous Data, Robert H. Pearson, Daniel J. Mundform

Journal of Modern Applied Statistical Methods

Minimum sample sizes are recommended for conducting exploratory factor analysis on dichotomous data. A Monte Carlo simulation was conducted, varying the level of communalities, number of factors, variable-to-factor ratio and dichotomization threshold. Sample sizes were identified based on congruence between rotated population and sample factor loadings.


Incidence And Prevalence For A Triply Censored Data, Hilmi F. Kittani 2010 The Hashemite University, Jordan

Incidence And Prevalence For A Triply Censored Data, Hilmi F. Kittani

Journal of Modern Applied Statistical Methods

The model introduced for the natural history of a progressive disease has four disease states which are expressed as a joint distribution of three survival random variables. Covariates are included in the model using Cox’s proportional hazards model with necessary assumptions needed. Effects of the covariates are estimated and tested. Formulas for incidence in the preclinical, clinical and death states are obtained, and prevalence formulas are obtained for the preclinical and clinical states. Estimates of the sojourn times in the preclinical and clinical states are obtained.


Robust Estimators In Logistic Regression: A Comparative Simulation Study, Sanizah Ahmad, Norazan Mohamed Ramli, Habshah Midi 2010 [email protected]

Robust Estimators In Logistic Regression: A Comparative Simulation Study, Sanizah Ahmad, Norazan Mohamed Ramli, Habshah Midi

Journal of Modern Applied Statistical Methods

The maximum likelihood estimator (MLE) is commonly used to estimate the parameters of logistic regression models due to its efficiency under a parametric model. However, evidence has shown the MLE has an unduly effect on the parameter estimates in the presence of outliers. Robust methods are put forward to rectify this problem. This article examines the performance of the MLE and four existing robust estimators under different outlier patterns, which are investigated by real data sets and Monte Carlo simulation.


Digital Commons powered by bepress