Open Access. Powered by Scholars. Published by Universities.®

Statistical Theory Commons

Open Access. Powered by Scholars. Published by Universities.®

2005

Discipline
Institution
Keyword
Publication
Publication Type

Articles 1 - 30 of 134

Full-Text Articles in Statistical Theory

Alleviating Linear Ecological Bias And Optimal Design With Subsample Data, Adam Glynn, Jon Wakefield, Mark Handcock, Thomas Richardson Dec 2005

Alleviating Linear Ecological Bias And Optimal Design With Subsample Data, Adam Glynn, Jon Wakefield, Mark Handcock, Thomas Richardson

UW Biostatistics Working Paper Series

In this paper, we illustrate that combining ecological data with subsample data in situations in which a linear model is appropriate provides three main benefits. First, by including the individual level subsample data, the biases associated with linear ecological inference can be eliminated. Second, by supplementing the subsample data with ecological data, the information about parameters will be increased. Third, we can use readily available ecological data to design optimal subsampling schemes, so as to further increase the information about parameters. We present an application of this methodology to the classic problem of estimating the effect of a college degree …


Empirical Likelihood Inference For The Area Under The Roc Curve, Gengsheng Qin, Xiao-Hua Zhou Dec 2005

Empirical Likelihood Inference For The Area Under The Roc Curve, Gengsheng Qin, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

For a continuous-scale diagnostic test, the most commonly used summary index of the receiver operating characteristic (ROC) curve is the area under the curve (AUC) that measures the accuracy of the diagnostic test. In this paper we propose an empirical likelihood approach for the inference of AUC. We first define an empirical likelihood ratio for AUC and show that its limiting distribution is a scaled chi-square distribution. We then obtain an empirical likelihood based confidence interval for AUC using the scaled chi-square distribution. This empirical likelihood inference for AUC can be extended to stratified samples and the resulting limiting distribution …


Interval Estimation For The Ratio And Difference Of Two Lognormal Means, Yea-Hung Chen, Xiao-Hua Zhou Dec 2005

Interval Estimation For The Ratio And Difference Of Two Lognormal Means, Yea-Hung Chen, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

Health research often gives rise to data that follow lognormal distributions. In two sample situations, researchers are likely to be interested in estimating the difference or ratio of the population means. Several methods have been proposed for providing confidence intervals for these parameters. However, it is not clear which techniques are most appropriate, or how their performance might vary. Additionally, methods for the difference of means have not been adequately explored. We discuss in the present article five methods of analysis. These include two methods based on the log-likelihood ratio statistic and a generalized pivotal approach. Additionally, we provide and …


Inferences In Censored Cost Regression Models With Empirical Likelihood, Xiao-Hua Zhou, Gengsheng Qin, Huazhen Lin, Gang Li Dec 2005

Inferences In Censored Cost Regression Models With Empirical Likelihood, Xiao-Hua Zhou, Gengsheng Qin, Huazhen Lin, Gang Li

UW Biostatistics Working Paper Series

In many studies of health economics, we are interested in the expected total cost over a certain period for a patient with given characteristics. Problems can arise if cost estimation models do not account for distributional aspects of costs. Two such problems are 1) the skewed nature of the data and 2) censored observations. In this paper we propose an empirical likelihood (EL) method for constructing a confidence region for the vector of regression parameters and a confidence interval for the expected total cost of a patient with the given covariates. We show that this new method has good theoretical …


Confidence Intervals For Predictive Values Using Data From A Case Control Study, Nathaniel David Mercaldo, Xiao-Hua Zhou, Kit F. Lau Dec 2005

Confidence Intervals For Predictive Values Using Data From A Case Control Study, Nathaniel David Mercaldo, Xiao-Hua Zhou, Kit F. Lau

UW Biostatistics Working Paper Series

The accuracy of a binary-scale diagnostic test can be represented by sensitivity (Se), specificity (Sp) and positive and negative predictive values (PPV and NPV). Although Se and Sp measure the intrinsic accuracy of a diagnostic test that does not depend on the prevalence rate, they do not provide information on the diagnostic accuracy of a particular patient. To obtain this information we need to use PPV and NPV. Since PPV and NPV are functions of both the intrinsic accuracy and the prevalence of the disease, constructing confidence intervals for PPV and NPV for a particular patient in a population with …


Issues Of Processing And Multiple Testing Of Seldi-Tof Ms Proteomic Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan, Christine F. Skibola, Christine M. Hegedus, Martyn T. Smith Dec 2005

Issues Of Processing And Multiple Testing Of Seldi-Tof Ms Proteomic Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan, Christine F. Skibola, Christine M. Hegedus, Martyn T. Smith

U.C. Berkeley Division of Biostatistics Working Paper Series

A new data filtering method for SELDI-TOF MS proteomic spectra data is described. We examined technical repeats (2 per subject) of intensity versus m/z (mass/charge) of bone marrow cell lysate for two groups of childhood leukemia patients: acute myeloid leukemia (AML) and acute lymphoblastic leukemia (ALL). As others have noted, the type of data processing as well as experimental variability can have a disproportionate impact on the list of "interesting" proteins (see Baggerly et al. (2004)). We propose a list of processing and multiple testing techniques to correct for 1) background drift; 2) filtering using smooth regression and cross-validated bandwidth …


Quantile-Function Based Null Distribution In Resampling Based Multiple Testing, Mark J. Van Der Laan, Alan E. Hubbard Nov 2005

Quantile-Function Based Null Distribution In Resampling Based Multiple Testing, Mark J. Van Der Laan, Alan E. Hubbard

U.C. Berkeley Division of Biostatistics Working Paper Series

Simultaneously testing a collection of null hypotheses about a data generating distribution based on a sample of independent and identically distributed observations is a fundamental and important statistical problem involving many applications. Methods based on marginal null distributions (i.e., marginal p-values) are attractive since the marginal p-values can be based on a user supplied choice of marginal null distributions and they are computationally trivial, but they, by necessity, are known to either be conservative or to rely on assumptions about the dependence structure between the test-statistics. Resampling based multiple testing (Westfall and Young, 1993) involves sampling from a joint null …


Optimal Feature Selection For Nearest Centroid Classifiers, With Applications To Gene Expression Microarrays, Alan R. Dabney, John D. Storey Nov 2005

Optimal Feature Selection For Nearest Centroid Classifiers, With Applications To Gene Expression Microarrays, Alan R. Dabney, John D. Storey

UW Biostatistics Working Paper Series

Nearest centroid classifiers have recently been successfully employed in high-dimensional applications. A necessary step when building a classifier for high-dimensional data is feature selection. Feature selection is typically carried out by computing univariate statistics for each feature individually, without consideration for how a subset of features performs as a whole. For subsets of a given size, we characterize the optimal choice of features, corresponding to those yielding the smallest misclassification rate. Furthermore, we propose an algorithm for estimating this optimal subset in practice. Finally, we investigate the applicability of shrinkage ideas to nearest centroid classifiers. We use gene-expression microarrays for …


A New Approach To Intensity-Dependent Normalization Of Two-Channel Microarrays, Alan R. Dabney, John D. Storey Nov 2005

A New Approach To Intensity-Dependent Normalization Of Two-Channel Microarrays, Alan R. Dabney, John D. Storey

UW Biostatistics Working Paper Series

A two-channel microarray measures the relative expression levels of thousands of genes from a pair of biological samples. In order to reliably compare gene expression levels between and within arrays, it is necessary to remove systematic errors that distort the biological signal of interest. The standard for accomplishing this is smoothing "MA-plots" to remove intensity-dependent dye bias and array-specific effects. However, MA methods require strong assumptions. We review these assumptions and derive several practical scenarios in which they fail. The "dye-swap" normalization method has been much less frequently used because it requires two arrays per pair of samples. We show …


A General Imputation Methodology For Nonparametric Regression With Censored Data, Dan Rubin, Mark J. Van Der Laan Nov 2005

A General Imputation Methodology For Nonparametric Regression With Censored Data, Dan Rubin, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

We consider the random design nonparametric regression problem when the response variable is subject to a general mode of missingness or censoring. A traditional approach to such problems is imputation, in which the missing or censored responses are replaced by well-chosen values, and then the resulting covariate/response data are plugged into algorithms designed for the uncensored setting. We present a general methodology for imputation with the property of double robustness, in that the method works well if either a parameter of the full data distribution (covariate and response distribution) or a parameter of the censoring mechanism is well approximated. These …


Estimating A Treatment Effect With Repeated Measurements Accounting For Varying Effectiveness Duration, Ying Qing Chen, Jingrong Yang, Su-Chun Cheng Nov 2005

Estimating A Treatment Effect With Repeated Measurements Accounting For Varying Effectiveness Duration, Ying Qing Chen, Jingrong Yang, Su-Chun Cheng

UW Biostatistics Working Paper Series

To assess treatment efficacy in clinical trials, certain clinical outcomes are repeatedly measured for same subject over time. They can be regarded as function of time. The difference in their mean functions between the treatment arms usually characterises a treatment effect. Due to the potential existence of subject-specific treatment effectiveness lag and saturation times, erosion of treatment effect in the difference may occur during the observation period of time. Instead of using ad hoc parametric or purely nonparametric time-varying coefficients in statistical modeling, we first propose to model the treatment effectiveness durations, which are the varying time intervals between the …


Comparison Of Statistical Tests In Logistic Regression: The Case Of Hypernatreamia, Stylianos Katsaragakis, Christos Koukouvinos, Stella Stylianou, Eleni-Maria Theodoraki, Eleni-Maria Theodoraki Nov 2005

Comparison Of Statistical Tests In Logistic Regression: The Case Of Hypernatreamia, Stylianos Katsaragakis, Christos Koukouvinos, Stella Stylianou, Eleni-Maria Theodoraki, Eleni-Maria Theodoraki

Journal of Modern Applied Statistical Methods

The logistic regression has become an integral component of any medical data analysis concerning binary responses. The main issue rising after the adaptation of the final model is its goodness-of-fit. The fit of the model is assessed via the overall measures and summary statistics and comparing them in the case of hypernateamia.


An Estimator Of Intervention Effect On Disease Severity, David Siev Nov 2005

An Estimator Of Intervention Effect On Disease Severity, David Siev

Journal of Modern Applied Statistical Methods

When a medical intervention prevents a dichotomous outcome, the size of its effect is often estimated with the prevented fraction. Some interventions may reduce the severity of an outcome without entirely preventing it. To quantify the effect of a severity-moderating intervention, a measure termed the mitigated fraction (MF) is proposed. MF has broad applicability, because it measures the overlap of two empirical distributions based on their stochastic ordering. It is also useful in the specific context of medical interventions, because it shares certain structural and functional features with the prevented fraction. The two measures may be applied together …


Bootstrap Intervals Of The Parameters Of Lognormal Distribution Using Power Rule Model And Accelerated Life Tests, Mohammed Al-Haj Ebrahem Nov 2005

Bootstrap Intervals Of The Parameters Of Lognormal Distribution Using Power Rule Model And Accelerated Life Tests, Mohammed Al-Haj Ebrahem

Journal of Modern Applied Statistical Methods

Assumed that the distribution of the lifetime of any unit follows a lognormal distribution with parameters μ and σ . Also, assume that the relationship between μ and the stress level V is given by the power rule model. Several types of bootstrap intervals of the parameters were studied and their performance was studied using simulations and compared in term of attainment of the nominal confidence level, symmetry of lower and upper error rates and the expected width. Conclusions and recommendations are given.


Large Sample And Bootstrap Intervals For The Gamma Scale Parameter Based On Grouped Data, Ayman Baklizi, Amjad Al-Nasser Nov 2005

Large Sample And Bootstrap Intervals For The Gamma Scale Parameter Based On Grouped Data, Ayman Baklizi, Amjad Al-Nasser

Journal of Modern Applied Statistical Methods

Interval estimation of the scale parameter of the gamma distribution using grouped data is considered in this article. Exact intervals do not exist and approximate intervals are needed Recently, Chen and Mi (2001) proposed alternative approximate intervals. In this article, some bootstrap and jackknife type intervals are proposed. The performance of these intervals is investigated and compared. The results show that some of the suggested intervals have a satisfactory statistical performance in situations where the sample size is small with heavy proportion of censoring.


A Comparison Of The Spearman-Brown And Flanagan-Rulon Formulas For Split Half Reliability Under Various Variance Parameter Conditions, David A. Walker Nov 2005

A Comparison Of The Spearman-Brown And Flanagan-Rulon Formulas For Split Half Reliability Under Various Variance Parameter Conditions, David A. Walker

Journal of Modern Applied Statistical Methods

Differences between the Spearman-Brown and Flanagan-Rulon formulas are examined when the variance parameters for two halves of a test had the following ratios: 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0 and also had a correlation between the two halves of a test at 1.00, .95, .90, .80, .70, .60, .50, .40, .30, .20, .10, .05. It was found that use of the Spearman-Brown formula to estimate the population ρ when the ratio between the standard deviations on two halves of a test is disparate, or beyond .9 to 1.1, was not warranted. Applied and theoretical examples …


Restricted Quasi-Independent Model Resolves Paradoxical Behaviors Of Cohen’S Kappa, Vicki Stover Hertzberg, Frank Xu, Michael Haber Nov 2005

Restricted Quasi-Independent Model Resolves Paradoxical Behaviors Of Cohen’S Kappa, Vicki Stover Hertzberg, Frank Xu, Michael Haber

Journal of Modern Applied Statistical Methods

Cohen’s kappa, an index of inter-rater agreement, behaves paradoxically in 2×2 tables. λA is derived, an index from the restricted quasi-independent model for 2×2 tables. Simulation studies are used to demonstrate λA has superior performance compared to Scott’s pi. Moreover, λA does not show paradoxical behavior for 2×2 tables.


Maximum Tests Are Adaptive Permutation Tests, Markus Neuhäeuser, Ludwig A. Hothorn Nov 2005

Maximum Tests Are Adaptive Permutation Tests, Markus Neuhäeuser, Ludwig A. Hothorn

Journal of Modern Applied Statistical Methods

In some areas, e.g., statistical genetics, it is common to apply a maximum test, where the maximum of several competing test statistics is used as a new statistic, and the permutation distribution of the maximum is used for inference. Here, it is shown that maximum tests are special cases of adaptive permutation tests. The 30-year old idea of adaptive statistical tests is more flexible than previously thought when permutation tests are used, and the selector statistic is calculated for every permutation. Because the independence between the selector and the test statistics is no longer needed, the test statistics themselves can …


Inferences About The Components Of A Generalized Additive Model, Rand R. Wilcox Nov 2005

Inferences About The Components Of A Generalized Additive Model, Rand R. Wilcox

Journal of Modern Applied Statistical Methods

A method for making inferences about the components of a generalized additive model is described. It is found that a variation of the method, based on means, performs well in simulations. Unlike many other inferential methods, switching from a mean to a 20% trimmed mean was found to offer little or no advantage in terms of both power and controlling the probability of a Type I error.


The Individuals Control Chart In Case Of Non-Normality, BetüL Kan, Berna Yazici Nov 2005

The Individuals Control Chart In Case Of Non-Normality, BetüL Kan, Berna Yazici

Journal of Modern Applied Statistical Methods

This article examines the effects of non-normality as measured by skewness and provides an alternative method of designing individuals control chart with non-normal distributions. A skewness correction method for constructing the individuals control chart is provided. An example of thickness of biscuit process is presented to illustrate the individuals control chart limits.


An Alternative To Warner’S Randomized Response Model, Sat Gupta, Javid Shabbir Nov 2005

An Alternative To Warner’S Randomized Response Model, Sat Gupta, Javid Shabbir

Journal of Modern Applied Statistical Methods

A modification to Warner’s (1965) Randomized Response Model is suggested. The suggested model is more efficient than the original model.


Inference For P(Y, Vee Ming Ng Nov 2005

Inference For P(Y, Vee Ming Ng

Journal of Modern Applied Statistical Methods

Some tests and confidence bounds for the reliability parameter R=P(Y


Determining Parallel Analysis Criteria, Marley W. Watkins Nov 2005

Determining Parallel Analysis Criteria, Marley W. Watkins

Journal of Modern Applied Statistical Methods

Determining the number of factors to extract is a critical decision in exploratory factor analysis. Simulation studies have found the Parallel Analysis criterion to be accurate, but it is computationally intensive. Two freeware programs that implement Parallel Analysis on Macintosh and Windows operating systems are presented.


Change Point Estimation Of Bilevel Functions, Leming Qu, Yi-Cheng Tu Nov 2005

Change Point Estimation Of Bilevel Functions, Leming Qu, Yi-Cheng Tu

Journal of Modern Applied Statistical Methods

Reconstruction of a bilevel function such as a bar code signal in a partially blind deconvolution problem is an important task in industrial processes. Existing methods are based on either the local approach or the regularization approach with a total variation penalty. This article reformulated the problem explicitly in terms of change points of the 0-1 step function. The bilevel function is then reconstructed by solving the nonlinear least squares problem subject to linear inequality constraints, with starting values provided by the local extremas of the derivative of the convolved signal from discrete noisy data. Simulation results show a considerable …


Correlation Between The Number Of Epileptic And Healthy Children In Family Size That Follows A Size-Biased Modified Power Series Distribution, Ramalingam Shanmugam, Anwar Hassan, Peer Bilal Ahmad Nov 2005

Correlation Between The Number Of Epileptic And Healthy Children In Family Size That Follows A Size-Biased Modified Power Series Distribution, Ramalingam Shanmugam, Anwar Hassan, Peer Bilal Ahmad

Journal of Modern Applied Statistical Methods

An expression for the correlation between the random number of epileptic and healthy children in family whose size follows a size-biased Modified Power Series Distribution (SBMPSD) is obtained and illustrated. As special cases, results are extracted for size biased Modified Negative Binomial Distribution (SBGNBD), size biased Modified Poisson Distribution (SBGPD) and size biased Modified Logarithmic Series Distribution (SBGLSD).


Simulation Of Non-Normal Autocorrelated Variables, H.E.T. Holgersson Nov 2005

Simulation Of Non-Normal Autocorrelated Variables, H.E.T. Holgersson

Journal of Modern Applied Statistical Methods

All statistical methods rely on assumptions to some extent. Two assumptions frequently met in statistical analyses are those of normal distribution and independence. When examining robustness properties of such assumptions by Monte Carlo simulations it is therefore crucial that the possible effects of autocorrelation and non-normality are not confounded so that their separate effects may be investigated. This article presents a number of non-normal variables with non-confounded autocorrelation, thus allowing the analyst to specify autocorrelation or shape properties while keeping the other effect fixed.


Interval Estimation Of Risk Difference In Simple Compliance Randomized Trials, Kung-Jong Lui Nov 2005

Interval Estimation Of Risk Difference In Simple Compliance Randomized Trials, Kung-Jong Lui

Journal of Modern Applied Statistical Methods

Consider the simple compliance randomized trial, in which patients randomly assigned to the experimental treatment may switch to receive the standard treatment, while patients randomly assigned to the standard treatment are all assumed to receive their assigned treatment. Six asymptotic interval estimators for the risk difference in probabilities of response among patients who would accept the experimental treatment were developed. Monte Carlo methods were employed to evaluate and compare the finite-sample performance of these estimators. An example studying the effect of vitamin A supplementation on reducing mortality in preschool children was included to illustrate their practical use.


A Robust Exponentially Weighted Moving Average Control Chart For The Process Mean, Michael B. C. Khoo, S. Y. Sim Nov 2005

A Robust Exponentially Weighted Moving Average Control Chart For The Process Mean, Michael B. C. Khoo, S. Y. Sim

Journal of Modern Applied Statistical Methods

To date, numerous extensions of the exponentially weighted moving average, EWMA charts have been made. A new robust EWMA chart for the process mean is proposed. It enables easier detection of outliers and increase sensitivity to other forms of out-of-control situation when outliers are present.


A Comparison Of Risk Classification Methods For Claim Severity Data, Noriszura Ismail, Abdul Aziz Jemain Nov 2005

A Comparison Of Risk Classification Methods For Claim Severity Data, Noriszura Ismail, Abdul Aziz Jemain

Journal of Modern Applied Statistical Methods

The objective of this article is to compare several risk classification methods for claim severity data by using weighted equation which is written as a weighted difference between the observed and fitted values. The weighted equation will be applied to estimate claim severities which is equivalent to the total claim costs divided by the number of claims.


Supporting And Preparing Future Decision-Makers With The Needed Tools, Michael Wolf-Branigin Nov 2005

Supporting And Preparing Future Decision-Makers With The Needed Tools, Michael Wolf-Branigin

Journal of Modern Applied Statistical Methods

Supporting and Preparing Future Decision-makers with the Needed ToolsEducational and social service researchers and evaluators continue to develop advanced statistical methods. To ensure that our students have the essential skills as they enter direct service, the focus must be on assuring that they learn readily understandable methods that are appropriate for small samples and use repeated measures.