A Generalized Approach For Testing The Association Of A Set Of Predictors With An Outcome: A Gene Based Test,
2011
University of California - Berkeley
A Generalized Approach For Testing The Association Of A Set Of Predictors With An Outcome: A Gene Based Test, Benjamin A. Goldstein, Alan E. Hubbard, Lisa F. Barcellos
U.C. Berkeley Division of Biostatistics Working Paper Series
In many analyses, one has data on one level but desires to draw inference on another level. For example, in genetic association studies, one observes units of DNA referred to as SNPs, but wants to determine whether genes that are comprised of SNPs are associated with disease. While there are some available approaches for addressing this issue, they usually involve making parametric assumptions and are not easily generalizable. A statistical test is proposed for testing the association of a set of variables with an outcome of interest. No assumptions are made about the functional form relating the variables to the …
Flipping The Winner Of A Poset Game,
2011
Illinois Mathematics and Science Academy
Flipping The Winner Of A Poset Game, Adam O. Kalinich '12
Student Publications & Research
Partially-ordered set games, also called poset games, are a class of two-player combinatorial games. The playing field consists of a set of elements, some of which are greater than other elements. Two players take turns removing an element and all elements greater than it, and whoever takes the last element wins. Examples of poset games include Nim and Chomp. We investigate the complexity of computing which player of a poset game has a winning strategy. We give an inductive procedure that modifies poset games to change the nim-value which informally captures the winning strategies in the game. For a generic …
Parametric Estimation In Competing Risks And Multi-State Models,
2011
University of Kentucky
Parametric Estimation In Competing Risks And Multi-State Models, Yushun Lin
Theses and Dissertations--Statistics
The typical research of Alzheimer's disease includes a series of cognitive states. Multi-state models are often used to describe the history of disease evolvement. Competing risks models are a sub-category of multi-state models with one starting state and several absorbing states.
Analyses for competing risks data in medical papers frequently assume independent risks and evaluate covariate effects on these events by modeling distinct proportional hazards regression models for each event. Jeong and Fine (2007) proposed a parametric proportional sub-distribution hazard (SH) model for cumulative incidence functions (CIF) without assumptions about the dependence among the risks. We modified their model to …
Oracle And Multiple Robustness Properties Of Survey Calibration Estimator In Missing Response Problem,
2010
University of Washington
Oracle And Multiple Robustness Properties Of Survey Calibration Estimator In Missing Response Problem, Kwun Chuen Gary Chan
UW Biostatistics Working Paper Series
In the presence of missing response, reweighting the complete case subsample by the inverse of nonmissing probability is both intuitive and easy to implement. However, inverse probability weighting is not efficient in general and is not robust against misspecification of the missing probability model. Calibration was developed by survey statisticians for improving efficiency of inverse probability weighting estimators when population totals of auxiliary variables are known and when inclusion probability is known by design. In missing data problem we can calibrate auxiliary variables in the complete case subsample to the full sample. However, the inclusion probability is unknown in general …
Modification And Improvement Of Empirical Likelihood For Missing Response Problem,
2010
University of Washington - Seattle Campus
Modification And Improvement Of Empirical Likelihood For Missing Response Problem, Kwun Chuen Gary Chan
UW Biostatistics Working Paper Series
An empirical likelihood (EL) estimator was proposed by Qin and Zhang (2007) for a missing response problem under a missing at random assumption. They showed by simulation studies that the finite sample performance of EL estimator is better than some existing estimators. However, the empirical likelihood estimator does not have a uniformly smaller asymptotic variance than other estimators in general. We consider several modifications to the empirical likelihood estimator and show that the proposed estimator dominates the empirical likelihood estimator and several other existing estimators in terms of asymptotic efficiencies. The proposed estimator also attains the minimum asymptotic variance among …
Modification And Improvement Of Empirical Liklihood For Missing Response Problem,
2010
University of Washington
Modification And Improvement Of Empirical Liklihood For Missing Response Problem, Gary Chan
UW Biostatistics Working Paper Series
An empirical likelihood (EL) estimator was proposed by Qin and Zhang (2007) for a missing response problem under a missing at random assumption. They showed by simulation studies that the finite sample performance of EL estimator is better than some existing estimators. However, the empirical likelihood estimator does not have a uniformly smaller asymptotic variance than other estimators in general. We consider several modifications to the empirical likelihood estimator and show that the proposed estimator dominates the empirical likelihood estimator and several other existing estimators in terms of asymptotic efficiencies. The proposed estimator also attains the minimum asymptotic variance among …
Minimum Description Length Measures Of Evidence For Enrichment,
2010
Ottawa Institute of Systems Biology, Department of Biochemistry, Microbiology, and Immunology, University of Ottawa
Minimum Description Length Measures Of Evidence For Enrichment, Zhenyu Yang, David R. Bickel
COBRA Preprint Series
In order to functionally interpret differentially expressed genes or other discovered features, researchers seek to detect enrichment in the form of overrepresentation of discovered features associated with a biological process. Most enrichment methods treat the p-value as the measure of evidence using a statistical test such as the binomial test, Fisher's exact test or the hypergeometric test. However, the p-value is not interpretable as a measure of evidence apart from adjustments in light of the sample size. As a measure of evidence supporting one hypothesis over the other, the Bayes factor (BF) overcomes this drawback of the p-value but lacks …
Efficient Measurement Error Correction With Spatially Misaligned Data,
2010
University of Washington
Efficient Measurement Error Correction With Spatially Misaligned Data, Adam A. Szpiro, Lianne Sheppard, Thomas Lumley
UW Biostatistics Working Paper Series
Association studies in environmental statistics often involve exposure and outcome data that are misaligned in space. A common strategy is to employ a spatial model such as universal kriging to predict exposures at locations with outcome data and then estimate a regression parameter of interest using the predicted exposures. This results in measurement error because the predicted exposures do not correspond exactly to the true values. We characterize the measurement error by decomposing it into Berkson-like and classical-like components. One correction approach is the parametric bootstrap, which is effective but computationally intensive since it requires solving a nonlinear optimization problem …
Minimum Description Length And Empirical Bayes Methods Of Identifying Snps Associated With Disease,
2010
Ottawa Institute of Systems Biology, Department of Biochemistry, Microbiology, and Immunology
Minimum Description Length And Empirical Bayes Methods Of Identifying Snps Associated With Disease, Ye Yang, David R. Bickel
COBRA Preprint Series
The goal of determining which of hundreds of thousands of SNPs are associated with disease poses one of the most challenging multiple testing problems. Using the empirical Bayes approach, the local false discovery rate (LFDR) estimated using popular semiparametric models has enjoyed success in simultaneous inference. However, the estimated LFDR can be biased because the semiparametric approach tends to overestimate the proportion of the non-associated single nucleotide polymorphisms (SNPs). One of the negative consequences is that, like conventional p-values, such LFDR estimates cannot quantify the amount of information in the data that favors the null hypothesis of no disease-association.
We …
Improving The Power Of Chronic Disease Surveillance By Incorporating Residential History,
2010
Harvard School of Public Health
Improving The Power Of Chronic Disease Surveillance By Incorporating Residential History, Justin Manjourides, Marcello Pagano
Harvard University Biostatistics Working Paper Series
No abstract provided.
Generalized Variances Ratio Test For Comparing K Covariance Matrices From Dependent Normal Populations,
2010
Federal University of Lavras, Brazil
Generalized Variances Ratio Test For Comparing K Covariance Matrices From Dependent Normal Populations, Marcelo Angelo Cirillo, Daniel Furtado Ferreira, Thelma Sáfadi, Eric Batista Ferreira
Journal of Modern Applied Statistical Methods
New tests based on the ratio of generalized variances are presented to compare covariance matrices from dependent normal populations. Monte Carlo simulation concluded that the tests considered controlled the Type I error, providing empirical probabilities that were consistent with the nominal level stipulated.
A Ga-Based Sales Forecasting Model Incorporating Promotion Factors,
2010
Tunghai University, Taichung, Taiwan ROC
A Ga-Based Sales Forecasting Model Incorporating Promotion Factors, Li-Chih Wang, Chin-Lien Wang
Journal of Modern Applied Statistical Methods
Because promotions are critical factors highly related to product sales of consumer packaged goods (CPG) companies, predictors concerning sales forecast of CPG products must take promotions into consideration. Decomposition regression incorporating contextual factors offers a method for exploiting both reliability of statistical forecasting and flexibility of judgmental forecasting employing domain knowledge. However, it suffers from collinearity causing poor performance in variable identification and parameter estimation with traditional ordinary least square (OLS). Empirical research evidence shows that - in the case of collinearity - in variable identification, parameter estimation, and out of sample forecasting, genetic algorithms (GA) as an estimator outperform …
Estimating The Non-Existent Mean And Variance Of The F-Distribution By Simulation,
2010
Private Scholar
Estimating The Non-Existent Mean And Variance Of The F-Distribution By Simulation, Hamid Reza Kamali, Parisa Shahnazari-Shahrezaei
Journal of Modern Applied Statistical Methods
In theory, all moments of some probability distributions do not necessarily exist. In the other words, they may be infinite or undefined. One of these distributions is the F-distribution whose mean and variance have not been defined for the second degree of freedom less than 3 and 5, respectively. In some cases, a large statistical population having an F-distribution may exist and the aim is to obtain its mean and variance which are an estimation of the non-existent mean and variance of F-distribution. This article considers a large sample F-distribution to estimate its non-existent mean and variance using Simul8 simulation …
Ridge Regression Based On Some Robust Estimators,
2010
Eskisehir Osmangazi University, Turkey
Ridge Regression Based On Some Robust Estimators, Hatice Samkar, Ozlem Alpu
Journal of Modern Applied Statistical Methods
Robust ridge methods based on M, S, MM and GM estimators are examined in the presence of multicollinearity and outliers. GMWalker, using the LS estimator as the initial estimator is used. S and MM estimators are also used as initial estimators with the aim of evaluating the two alternatives as biased robust methods.
A Flexible Method For Testing Independence In Two-Way Contingency Tables,
2010
Shiraz University of Medical Sciences, Shiraz, Iran
A Flexible Method For Testing Independence In Two-Way Contingency Tables, Peyman Jafari, Noori Akhtar-Danesh, Zahra Bagheri
Journal of Modern Applied Statistical Methods
A flexible approach for testing association in two-way contingency tables is presented. It is simple, does not assume a specific form for the association and is applicable to tables with nominal-by-nominal, nominal-by-ordinal, and ordinal-by-ordinal classifications.
Statistical And Mathematical Modeling Versus Nhst? There’S No Competition!,
2010
University of Oklahoma
Statistical And Mathematical Modeling Versus Nhst? There’S No Competition!, Joseph Lee Rodgers
Journal of Modern Applied Statistical Methods
Some of Robinson & Levin’s critique of Rodgers (2010) is cogent, helpful, and insightful – although limiting. Recent methodology has advanced through the development of structural equation modeling, multi-level modeling, missing data methods, hierarchical linear modeling, categorical data analysis, as well as the development of many dedicated and specific behavioral models. These methodological approaches are based on a revised epistemological system, and have emerged naturally, without the need for task forces, or even much self-conscious discussion. The original goal was neither to develop nor promote a modeling revolution. That has occurred; I documented its development and its status. Two organizing …
Effect Of Measurement Errors On The Separate And Combined Ratio And Product Estimators In Stratified Random Sampling,
2010
Vikram University, Ujjain, India
Effect Of Measurement Errors On The Separate And Combined Ratio And Product Estimators In Stratified Random Sampling, Housila P. Singh, Namrata Karpe
Journal of Modern Applied Statistical Methods
Separate and combined ratio, product and difference estimators are introduced for population mean μY of a study variable Y using auxiliary variable X in stratified sampling when the observations are contaminated with measurement errors. The bias and mean squared error of the proposed estimators have been derived under large sample approximation and their properties are analyzed. Generalized versions of these estimators are given along with their properties.
Recommended Sample Size For Conducting Exploratory Factor Analysis On Dichotomous Data,
2010
University of Northern Colorado
Recommended Sample Size For Conducting Exploratory Factor Analysis On Dichotomous Data, Robert H. Pearson, Daniel J. Mundform
Journal of Modern Applied Statistical Methods
Minimum sample sizes are recommended for conducting exploratory factor analysis on dichotomous data. A Monte Carlo simulation was conducted, varying the level of communalities, number of factors, variable-to-factor ratio and dichotomization threshold. Sample sizes were identified based on congruence between rotated population and sample factor loadings.
Incidence And Prevalence For A Triply Censored Data,
2010
The Hashemite University, Jordan
Incidence And Prevalence For A Triply Censored Data, Hilmi F. Kittani
Journal of Modern Applied Statistical Methods
The model introduced for the natural history of a progressive disease has four disease states which are expressed as a joint distribution of three survival random variables. Covariates are included in the model using Cox’s proportional hazards model with necessary assumptions needed. Effects of the covariates are estimated and tested. Formulas for incidence in the preclinical, clinical and death states are obtained, and prevalence formulas are obtained for the preclinical and clinical states. Estimates of the sojourn times in the preclinical and clinical states are obtained.
Robust Estimators In Logistic Regression: A Comparative Simulation Study,
2010
[email protected]
Robust Estimators In Logistic Regression: A Comparative Simulation Study, Sanizah Ahmad, Norazan Mohamed Ramli, Habshah Midi
Journal of Modern Applied Statistical Methods
The maximum likelihood estimator (MLE) is commonly used to estimate the parameters of logistic regression models due to its efficiency under a parametric model. However, evidence has shown the MLE has an unduly effect on the parameter estimates in the presence of outliers. Robust methods are put forward to rectify this problem. This article examines the performance of the MLE and four existing robust estimators under different outlier patterns, which are investigated by real data sets and Monte Carlo simulation.
