Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (93)
- Social and Behavioral Sciences (93)
- Statistical Methodology (41)
- Statistical Models (13)
- Biostatistics (12)
-
- Life Sciences (8)
- Genetics and Genomics (7)
- Survival Analysis (7)
- Multivariate Analysis (6)
- Bioinformatics (5)
- Computational Biology (5)
- Microarrays (5)
- Longitudinal Data Analysis and Time Series (4)
- Clinical Trials (3)
- Epidemiology (3)
- Genetics (3)
- Laboratory and Basic Science Research (3)
- Medicine and Health Sciences (3)
- Public Health (3)
- Applied Mathematics (2)
- Numerical Analysis and Computation (2)
- Design of Experiments and Sample Surveys (1)
- Institution
- Keyword
-
- Bootstrap (7)
- Power (5)
- Type I error rate (4)
- Factor analysis (3)
- Logistic regression (3)
-
- Nonnormality (3)
- Null distribution (3)
- P-value (3)
- Permutation test (3)
- Robustness (3)
- Sample size (3)
- T test (3)
- Adjusted p-value (2)
- Average run length (ARL) (2)
- Bayes factor (2)
- Bayesian (2)
- Classification (2)
- Confidence intervals (2)
- Edgeworth expansion (2)
- Effect size (2)
- Empirical likelihood (2)
- Gamma distribution (2)
- Information criteria (2)
- Interactions (2)
- Interval estimation (2)
- Linear regression (2)
- Lognormal distribution (2)
- Maximum test (2)
- Missing data (2)
- Monte Carlo (2)
- Publication
- Publication Type
Articles 1 - 30 of 134
Full-Text Articles in Statistical Theory
Alleviating Linear Ecological Bias And Optimal Design With Subsample Data, Adam Glynn, Jon Wakefield, Mark Handcock, Thomas Richardson
Alleviating Linear Ecological Bias And Optimal Design With Subsample Data, Adam Glynn, Jon Wakefield, Mark Handcock, Thomas Richardson
UW Biostatistics Working Paper Series
In this paper, we illustrate that combining ecological data with subsample data in situations in which a linear model is appropriate provides three main benefits. First, by including the individual level subsample data, the biases associated with linear ecological inference can be eliminated. Second, by supplementing the subsample data with ecological data, the information about parameters will be increased. Third, we can use readily available ecological data to design optimal subsampling schemes, so as to further increase the information about parameters. We present an application of this methodology to the classic problem of estimating the effect of a college degree …
Empirical Likelihood Inference For The Area Under The Roc Curve, Gengsheng Qin, Xiao-Hua Zhou
Empirical Likelihood Inference For The Area Under The Roc Curve, Gengsheng Qin, Xiao-Hua Zhou
UW Biostatistics Working Paper Series
For a continuous-scale diagnostic test, the most commonly used summary index of the receiver operating characteristic (ROC) curve is the area under the curve (AUC) that measures the accuracy of the diagnostic test. In this paper we propose an empirical likelihood approach for the inference of AUC. We first define an empirical likelihood ratio for AUC and show that its limiting distribution is a scaled chi-square distribution. We then obtain an empirical likelihood based confidence interval for AUC using the scaled chi-square distribution. This empirical likelihood inference for AUC can be extended to stratified samples and the resulting limiting distribution …
Interval Estimation For The Ratio And Difference Of Two Lognormal Means, Yea-Hung Chen, Xiao-Hua Zhou
Interval Estimation For The Ratio And Difference Of Two Lognormal Means, Yea-Hung Chen, Xiao-Hua Zhou
UW Biostatistics Working Paper Series
Health research often gives rise to data that follow lognormal distributions. In two sample situations, researchers are likely to be interested in estimating the difference or ratio of the population means. Several methods have been proposed for providing confidence intervals for these parameters. However, it is not clear which techniques are most appropriate, or how their performance might vary. Additionally, methods for the difference of means have not been adequately explored. We discuss in the present article five methods of analysis. These include two methods based on the log-likelihood ratio statistic and a generalized pivotal approach. Additionally, we provide and …
Inferences In Censored Cost Regression Models With Empirical Likelihood, Xiao-Hua Zhou, Gengsheng Qin, Huazhen Lin, Gang Li
Inferences In Censored Cost Regression Models With Empirical Likelihood, Xiao-Hua Zhou, Gengsheng Qin, Huazhen Lin, Gang Li
UW Biostatistics Working Paper Series
In many studies of health economics, we are interested in the expected total cost over a certain period for a patient with given characteristics. Problems can arise if cost estimation models do not account for distributional aspects of costs. Two such problems are 1) the skewed nature of the data and 2) censored observations. In this paper we propose an empirical likelihood (EL) method for constructing a confidence region for the vector of regression parameters and a confidence interval for the expected total cost of a patient with the given covariates. We show that this new method has good theoretical …
Confidence Intervals For Predictive Values Using Data From A Case Control Study, Nathaniel David Mercaldo, Xiao-Hua Zhou, Kit F. Lau
Confidence Intervals For Predictive Values Using Data From A Case Control Study, Nathaniel David Mercaldo, Xiao-Hua Zhou, Kit F. Lau
UW Biostatistics Working Paper Series
The accuracy of a binary-scale diagnostic test can be represented by sensitivity (Se), specificity (Sp) and positive and negative predictive values (PPV and NPV). Although Se and Sp measure the intrinsic accuracy of a diagnostic test that does not depend on the prevalence rate, they do not provide information on the diagnostic accuracy of a particular patient. To obtain this information we need to use PPV and NPV. Since PPV and NPV are functions of both the intrinsic accuracy and the prevalence of the disease, constructing confidence intervals for PPV and NPV for a particular patient in a population with …
Issues Of Processing And Multiple Testing Of Seldi-Tof Ms Proteomic Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan, Christine F. Skibola, Christine M. Hegedus, Martyn T. Smith
Issues Of Processing And Multiple Testing Of Seldi-Tof Ms Proteomic Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan, Christine F. Skibola, Christine M. Hegedus, Martyn T. Smith
U.C. Berkeley Division of Biostatistics Working Paper Series
A new data filtering method for SELDI-TOF MS proteomic spectra data is described. We examined technical repeats (2 per subject) of intensity versus m/z (mass/charge) of bone marrow cell lysate for two groups of childhood leukemia patients: acute myeloid leukemia (AML) and acute lymphoblastic leukemia (ALL). As others have noted, the type of data processing as well as experimental variability can have a disproportionate impact on the list of "interesting" proteins (see Baggerly et al. (2004)). We propose a list of processing and multiple testing techniques to correct for 1) background drift; 2) filtering using smooth regression and cross-validated bandwidth …
Quantile-Function Based Null Distribution In Resampling Based Multiple Testing, Mark J. Van Der Laan, Alan E. Hubbard
Quantile-Function Based Null Distribution In Resampling Based Multiple Testing, Mark J. Van Der Laan, Alan E. Hubbard
U.C. Berkeley Division of Biostatistics Working Paper Series
Simultaneously testing a collection of null hypotheses about a data generating distribution based on a sample of independent and identically distributed observations is a fundamental and important statistical problem involving many applications. Methods based on marginal null distributions (i.e., marginal p-values) are attractive since the marginal p-values can be based on a user supplied choice of marginal null distributions and they are computationally trivial, but they, by necessity, are known to either be conservative or to rely on assumptions about the dependence structure between the test-statistics. Resampling based multiple testing (Westfall and Young, 1993) involves sampling from a joint null …
Optimal Feature Selection For Nearest Centroid Classifiers, With Applications To Gene Expression Microarrays, Alan R. Dabney, John D. Storey
Optimal Feature Selection For Nearest Centroid Classifiers, With Applications To Gene Expression Microarrays, Alan R. Dabney, John D. Storey
UW Biostatistics Working Paper Series
Nearest centroid classifiers have recently been successfully employed in high-dimensional applications. A necessary step when building a classifier for high-dimensional data is feature selection. Feature selection is typically carried out by computing univariate statistics for each feature individually, without consideration for how a subset of features performs as a whole. For subsets of a given size, we characterize the optimal choice of features, corresponding to those yielding the smallest misclassification rate. Furthermore, we propose an algorithm for estimating this optimal subset in practice. Finally, we investigate the applicability of shrinkage ideas to nearest centroid classifiers. We use gene-expression microarrays for …
A New Approach To Intensity-Dependent Normalization Of Two-Channel Microarrays, Alan R. Dabney, John D. Storey
A New Approach To Intensity-Dependent Normalization Of Two-Channel Microarrays, Alan R. Dabney, John D. Storey
UW Biostatistics Working Paper Series
A two-channel microarray measures the relative expression levels of thousands of genes from a pair of biological samples. In order to reliably compare gene expression levels between and within arrays, it is necessary to remove systematic errors that distort the biological signal of interest. The standard for accomplishing this is smoothing "MA-plots" to remove intensity-dependent dye bias and array-specific effects. However, MA methods require strong assumptions. We review these assumptions and derive several practical scenarios in which they fail. The "dye-swap" normalization method has been much less frequently used because it requires two arrays per pair of samples. We show …
A General Imputation Methodology For Nonparametric Regression With Censored Data, Dan Rubin, Mark J. Van Der Laan
A General Imputation Methodology For Nonparametric Regression With Censored Data, Dan Rubin, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
We consider the random design nonparametric regression problem when the response variable is subject to a general mode of missingness or censoring. A traditional approach to such problems is imputation, in which the missing or censored responses are replaced by well-chosen values, and then the resulting covariate/response data are plugged into algorithms designed for the uncensored setting. We present a general methodology for imputation with the property of double robustness, in that the method works well if either a parameter of the full data distribution (covariate and response distribution) or a parameter of the censoring mechanism is well approximated. These …
Estimating A Treatment Effect With Repeated Measurements Accounting For Varying Effectiveness Duration, Ying Qing Chen, Jingrong Yang, Su-Chun Cheng
Estimating A Treatment Effect With Repeated Measurements Accounting For Varying Effectiveness Duration, Ying Qing Chen, Jingrong Yang, Su-Chun Cheng
UW Biostatistics Working Paper Series
To assess treatment efficacy in clinical trials, certain clinical outcomes are repeatedly measured for same subject over time. They can be regarded as function of time. The difference in their mean functions between the treatment arms usually characterises a treatment effect. Due to the potential existence of subject-specific treatment effectiveness lag and saturation times, erosion of treatment effect in the difference may occur during the observation period of time. Instead of using ad hoc parametric or purely nonparametric time-varying coefficients in statistical modeling, we first propose to model the treatment effectiveness durations, which are the varying time intervals between the …
Comparison Of Statistical Tests In Logistic Regression: The Case Of Hypernatreamia, Stylianos Katsaragakis, Christos Koukouvinos, Stella Stylianou, Eleni-Maria Theodoraki, Eleni-Maria Theodoraki
Comparison Of Statistical Tests In Logistic Regression: The Case Of Hypernatreamia, Stylianos Katsaragakis, Christos Koukouvinos, Stella Stylianou, Eleni-Maria Theodoraki, Eleni-Maria Theodoraki
Journal of Modern Applied Statistical Methods
The logistic regression has become an integral component of any medical data analysis concerning binary responses. The main issue rising after the adaptation of the final model is its goodness-of-fit. The fit of the model is assessed via the overall measures and summary statistics and comparing them in the case of hypernateamia.
An Estimator Of Intervention Effect On Disease Severity, David Siev
An Estimator Of Intervention Effect On Disease Severity, David Siev
Journal of Modern Applied Statistical Methods
When a medical intervention prevents a dichotomous outcome, the size of its effect is often estimated with the prevented fraction. Some interventions may reduce the severity of an outcome without entirely preventing it. To quantify the effect of a severity-moderating intervention, a measure termed the mitigated fraction (MF) is proposed. MF has broad applicability, because it measures the overlap of two empirical distributions based on their stochastic ordering. It is also useful in the specific context of medical interventions, because it shares certain structural and functional features with the prevented fraction. The two measures may be applied together …
Bootstrap Intervals Of The Parameters Of Lognormal Distribution Using Power Rule Model And Accelerated Life Tests, Mohammed Al-Haj Ebrahem
Bootstrap Intervals Of The Parameters Of Lognormal Distribution Using Power Rule Model And Accelerated Life Tests, Mohammed Al-Haj Ebrahem
Journal of Modern Applied Statistical Methods
Assumed that the distribution of the lifetime of any unit follows a lognormal distribution with parameters μ and σ . Also, assume that the relationship between μ and the stress level V is given by the power rule model. Several types of bootstrap intervals of the parameters were studied and their performance was studied using simulations and compared in term of attainment of the nominal confidence level, symmetry of lower and upper error rates and the expected width. Conclusions and recommendations are given.
Large Sample And Bootstrap Intervals For The Gamma Scale Parameter Based On Grouped Data, Ayman Baklizi, Amjad Al-Nasser
Large Sample And Bootstrap Intervals For The Gamma Scale Parameter Based On Grouped Data, Ayman Baklizi, Amjad Al-Nasser
Journal of Modern Applied Statistical Methods
Interval estimation of the scale parameter of the gamma distribution using grouped data is considered in this article. Exact intervals do not exist and approximate intervals are needed Recently, Chen and Mi (2001) proposed alternative approximate intervals. In this article, some bootstrap and jackknife type intervals are proposed. The performance of these intervals is investigated and compared. The results show that some of the suggested intervals have a satisfactory statistical performance in situations where the sample size is small with heavy proportion of censoring.
A Comparison Of The Spearman-Brown And Flanagan-Rulon Formulas For Split Half Reliability Under Various Variance Parameter Conditions, David A. Walker
A Comparison Of The Spearman-Brown And Flanagan-Rulon Formulas For Split Half Reliability Under Various Variance Parameter Conditions, David A. Walker
Journal of Modern Applied Statistical Methods
Differences between the Spearman-Brown and Flanagan-Rulon formulas are examined when the variance parameters for two halves of a test had the following ratios: 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0 and also had a correlation between the two halves of a test at 1.00, .95, .90, .80, .70, .60, .50, .40, .30, .20, .10, .05. It was found that use of the Spearman-Brown formula to estimate the population ρ when the ratio between the standard deviations on two halves of a test is disparate, or beyond .9 to 1.1, was not warranted. Applied and theoretical examples …
Restricted Quasi-Independent Model Resolves Paradoxical Behaviors Of Cohen’S Kappa, Vicki Stover Hertzberg, Frank Xu, Michael Haber
Restricted Quasi-Independent Model Resolves Paradoxical Behaviors Of Cohen’S Kappa, Vicki Stover Hertzberg, Frank Xu, Michael Haber
Journal of Modern Applied Statistical Methods
Cohen’s kappa, an index of inter-rater agreement, behaves paradoxically in 2×2 tables. λA is derived, an index from the restricted quasi-independent model for 2×2 tables. Simulation studies are used to demonstrate λA has superior performance compared to Scott’s pi. Moreover, λA does not show paradoxical behavior for 2×2 tables.
Maximum Tests Are Adaptive Permutation Tests, Markus Neuhäeuser, Ludwig A. Hothorn
Maximum Tests Are Adaptive Permutation Tests, Markus Neuhäeuser, Ludwig A. Hothorn
Journal of Modern Applied Statistical Methods
In some areas, e.g., statistical genetics, it is common to apply a maximum test, where the maximum of several competing test statistics is used as a new statistic, and the permutation distribution of the maximum is used for inference. Here, it is shown that maximum tests are special cases of adaptive permutation tests. The 30-year old idea of adaptive statistical tests is more flexible than previously thought when permutation tests are used, and the selector statistic is calculated for every permutation. Because the independence between the selector and the test statistics is no longer needed, the test statistics themselves can …
Inferences About The Components Of A Generalized Additive Model, Rand R. Wilcox
Inferences About The Components Of A Generalized Additive Model, Rand R. Wilcox
Journal of Modern Applied Statistical Methods
A method for making inferences about the components of a generalized additive model is described. It is found that a variation of the method, based on means, performs well in simulations. Unlike many other inferential methods, switching from a mean to a 20% trimmed mean was found to offer little or no advantage in terms of both power and controlling the probability of a Type I error.
The Individuals Control Chart In Case Of Non-Normality, BetüL Kan, Berna Yazici
The Individuals Control Chart In Case Of Non-Normality, BetüL Kan, Berna Yazici
Journal of Modern Applied Statistical Methods
This article examines the effects of non-normality as measured by skewness and provides an alternative method of designing individuals control chart with non-normal distributions. A skewness correction method for constructing the individuals control chart is provided. An example of thickness of biscuit process is presented to illustrate the individuals control chart limits.
An Alternative To Warner’S Randomized Response Model, Sat Gupta, Javid Shabbir
An Alternative To Warner’S Randomized Response Model, Sat Gupta, Javid Shabbir
Journal of Modern Applied Statistical Methods
A modification to Warner’s (1965) Randomized Response Model is suggested. The suggested model is more efficient than the original model.
Inference For P(Y, Vee Ming Ng
Inference For P(Y, Vee Ming Ng
Journal of Modern Applied Statistical Methods
Some tests and confidence bounds for the reliability parameter R=P(Y
Determining Parallel Analysis Criteria, Marley W. Watkins
Determining Parallel Analysis Criteria, Marley W. Watkins
Journal of Modern Applied Statistical Methods
Determining the number of factors to extract is a critical decision in exploratory factor analysis. Simulation studies have found the Parallel Analysis criterion to be accurate, but it is computationally intensive. Two freeware programs that implement Parallel Analysis on Macintosh and Windows operating systems are presented.
Change Point Estimation Of Bilevel Functions, Leming Qu, Yi-Cheng Tu
Change Point Estimation Of Bilevel Functions, Leming Qu, Yi-Cheng Tu
Journal of Modern Applied Statistical Methods
Reconstruction of a bilevel function such as a bar code signal in a partially blind deconvolution problem is an important task in industrial processes. Existing methods are based on either the local approach or the regularization approach with a total variation penalty. This article reformulated the problem explicitly in terms of change points of the 0-1 step function. The bilevel function is then reconstructed by solving the nonlinear least squares problem subject to linear inequality constraints, with starting values provided by the local extremas of the derivative of the convolved signal from discrete noisy data. Simulation results show a considerable …
Correlation Between The Number Of Epileptic And Healthy Children In Family Size That Follows A Size-Biased Modified Power Series Distribution, Ramalingam Shanmugam, Anwar Hassan, Peer Bilal Ahmad
Correlation Between The Number Of Epileptic And Healthy Children In Family Size That Follows A Size-Biased Modified Power Series Distribution, Ramalingam Shanmugam, Anwar Hassan, Peer Bilal Ahmad
Journal of Modern Applied Statistical Methods
An expression for the correlation between the random number of epileptic and healthy children in family whose size follows a size-biased Modified Power Series Distribution (SBMPSD) is obtained and illustrated. As special cases, results are extracted for size biased Modified Negative Binomial Distribution (SBGNBD), size biased Modified Poisson Distribution (SBGPD) and size biased Modified Logarithmic Series Distribution (SBGLSD).
Simulation Of Non-Normal Autocorrelated Variables, H.E.T. Holgersson
Simulation Of Non-Normal Autocorrelated Variables, H.E.T. Holgersson
Journal of Modern Applied Statistical Methods
All statistical methods rely on assumptions to some extent. Two assumptions frequently met in statistical analyses are those of normal distribution and independence. When examining robustness properties of such assumptions by Monte Carlo simulations it is therefore crucial that the possible effects of autocorrelation and non-normality are not confounded so that their separate effects may be investigated. This article presents a number of non-normal variables with non-confounded autocorrelation, thus allowing the analyst to specify autocorrelation or shape properties while keeping the other effect fixed.
Interval Estimation Of Risk Difference In Simple Compliance Randomized Trials, Kung-Jong Lui
Interval Estimation Of Risk Difference In Simple Compliance Randomized Trials, Kung-Jong Lui
Journal of Modern Applied Statistical Methods
Consider the simple compliance randomized trial, in which patients randomly assigned to the experimental treatment may switch to receive the standard treatment, while patients randomly assigned to the standard treatment are all assumed to receive their assigned treatment. Six asymptotic interval estimators for the risk difference in probabilities of response among patients who would accept the experimental treatment were developed. Monte Carlo methods were employed to evaluate and compare the finite-sample performance of these estimators. An example studying the effect of vitamin A supplementation on reducing mortality in preschool children was included to illustrate their practical use.
A Robust Exponentially Weighted Moving Average Control Chart For The Process Mean, Michael B. C. Khoo, S. Y. Sim
A Robust Exponentially Weighted Moving Average Control Chart For The Process Mean, Michael B. C. Khoo, S. Y. Sim
Journal of Modern Applied Statistical Methods
To date, numerous extensions of the exponentially weighted moving average, EWMA charts have been made. A new robust EWMA chart for the process mean is proposed. It enables easier detection of outliers and increase sensitivity to other forms of out-of-control situation when outliers are present.
A Comparison Of Risk Classification Methods For Claim Severity Data, Noriszura Ismail, Abdul Aziz Jemain
A Comparison Of Risk Classification Methods For Claim Severity Data, Noriszura Ismail, Abdul Aziz Jemain
Journal of Modern Applied Statistical Methods
The objective of this article is to compare several risk classification methods for claim severity data by using weighted equation which is written as a weighted difference between the observed and fitted values. The weighted equation will be applied to estimate claim severities which is equivalent to the total claim costs divided by the number of claims.
Supporting And Preparing Future Decision-Makers With The Needed Tools, Michael Wolf-Branigin
Supporting And Preparing Future Decision-Makers With The Needed Tools, Michael Wolf-Branigin
Journal of Modern Applied Statistical Methods
Supporting and Preparing Future Decision-makers with the Needed ToolsEducational and social service researchers and evaluators continue to develop advanced statistical methods. To ensure that our students have the essential skills as they enter direct service, the focus must be on assuring that they learn readily understandable methods that are appropriate for small samples and use repeated measures.