Open Access. Powered by Scholars. Published by Universities.®

Statistical Theory Commons™

Open Access. Powered by Scholars. Published by Universities.®

2005

Discipline
Institution
Keyword
Publication
Publication Type

Articles 61 - 90 of 134

Full-Text Articles in Statistical Theory

Estimation Of Process Variances In Robust Parameter Designs, T. K. Mak, Fassil Nebebe Nov 2005

Estimation Of Process Variances In Robust Parameter Designs, T. K. Mak, Fassil Nebebe

Journal of Modern Applied Statistical Methods

The modeling of variation through interactions is appealing in crossed array design as it leads to greater robustness to certain type of model misspecification. As an alternative to signal-to-noise analysis, a new, systematic method based on Taguchi type crossed array design is given. It is shown in this article that when fractional factorial design is used for the outer array, the crossed array design is not robust to the presence of noise-noise interactions and a method of rectifying the problem is suggested.


Simulation Procedure In Periodic Cancer Screening Trials, Ioana Barnicescu, Ricolindo L. Cariño Nov 2005

Simulation Procedure In Periodic Cancer Screening Trials, Ioana Barnicescu, Ricolindo L. Cariño

Journal of Modern Applied Statistical Methods

A general simulation procedure is described to validate model fitting algorithms for complex likelihood functions that are utilized in periodic cancer screening trials. Although screening programs have existed for a few decades, there are still many unsolved problems, such as how age or hormone affects the screening sensitivity, the sojourn time in the preclinical state, and the transition probability from diseasefree state to the preclinical state. Simulations are needed to check reliability or validity of the likelihood function combined with the associated effect functions. One bottleneck in the simulation procedure is the very time consuming calculations of the maximum likelihood …


Inference On (Y < X) In A Pareto Distribution, M. Masoom Ali, Jungsoo Woo Nov 2005

Inference On (Y < X) In A Pareto Distribution, M. Masoom Ali, Jungsoo Woo

Journal of Modern Applied Statistical Methods

Inference on the reliability R = P(Y < X) in a Pareto distribution with a known scale parameter is considered. Point estimates and confidence intervals of R are obtained a test of hypothesis is also considered.


Nonparametric Pooling And Testing Of Preference Ratings For Full-Profile Conjoint Analysis Experiments, Rosa Arboretti G., Marco Marozzi, Luigi Salmaso Nov 2005

Nonparametric Pooling And Testing Of Preference Ratings For Full-Profile Conjoint Analysis Experiments, Rosa Arboretti G., Marco Marozzi, Luigi Salmaso

Journal of Modern Applied Statistical Methods

The problem of pooling customer preference ratings within a conjoint analysis experiment has been addressed. A method based on the nonparametric combination of rankings has been proposed to compete with the usual method based on the arithmetic mean. This method is nonparametric with respect to the underlying dependence structure and so no dependence model must be assumed. The two methods have been compared using Spearman’s rank correlation coefficient and related test. Moreover, a further nonparametric testing method has been considered and proposed; this method takes both correlation and distance between ranks into account. By means of a simulation study it …


Statistical Pronouncements Iv, Jmasm Editors Nov 2005

Statistical Pronouncements Iv, Jmasm Editors

Journal of Modern Applied Statistical Methods

No abstract provided.


Jmasm20: Exact Permutation Critical Values For The Kruskal-Wallis One-Way Anova, Justice I. Odiase, Sunday M. Ogbonmwan Nov 2005

Jmasm20: Exact Permutation Critical Values For The Kruskal-Wallis One-Way Anova, Justice I. Odiase, Sunday M. Ogbonmwan

Journal of Modern Applied Statistical Methods

The exhaustive enumeration of all the permutations of the observations in an experiment is the only possible way of truly constructing exact tests of significance. The permutation paradigm requires no distributional assumptions and works well with values that are normal, almost normal and non-normally distributed. The Kruskal-Wallis test does not require the assumptions that the samples are from normal populations and that the samples have the same standard deviation. In this article, the exact permutation distribution of the Kruskal-Wallis test statistic is generated empirically by actually obtaining all the distinct permutations of an experiment. The tables of exact critical values …


Statistical Model And Estimation Of The Optimum Price For A Chain Of Price Setting Firms, Chengjie Xiong, Kejun Zhu Nov 2005

Statistical Model And Estimation Of The Optimum Price For A Chain Of Price Setting Firms, Chengjie Xiong, Kejun Zhu

Journal of Modern Applied Statistical Methods

A stochastic approach is used to model the economics of a chain of price setting firms. It is assumed that these firms have fixed capacities in their products, but random demands for their products. The optimum price, the optimum revenue, and the expected marginal revenue at a given price are investigated. The method of maximum likelihood is used to provide both point and confidence interval estimates. The coverage probabilities of confidence interval estimates based on a simulation study are presented.


The Influence Of Reliability On Four Rules For Determining The Number Of Components To Retain, Gibbs Y. Kanyongo Nov 2005

The Influence Of Reliability On Four Rules For Determining The Number Of Components To Retain, Gibbs Y. Kanyongo

Journal of Modern Applied Statistical Methods

Imperfectly reliable scores impact the performance of factor analytic procedures. A series of Monte Carlo studies was conducted to generate scores with known component structure from population matrices with varying levels of reliability. The scores were submitted to four procedures: Kaiser rule, scree plot, parallel analysis, and modified Horn’s parallel analysis to find if each procedure accurately determines the number of components at the different reliability levels. The performance of each procedure was judged by the percentage of the number of times that the procedure was correct and the mean components that each procedure extracted in each cell. Generally, the …


Corrections For Type I Error In Social Science Research: A Disconnect Between Theory And Practice, Kenneth Lachlan, Patric R. Spence Nov 2005

Corrections For Type I Error In Social Science Research: A Disconnect Between Theory And Practice, Kenneth Lachlan, Patric R. Spence

Journal of Modern Applied Statistical Methods

Type I errors are a common problem in factorial ANOVA and ANOVA based analyses. Despite decades of literature offering solutions to the Type I error problems associated with multiple significance tests, simple solutions such as Bonferroni corrections have been largely ignored by social scientists. To examine this discontinuity between theory and practice, a content analysis was performed on 5 flagship social science journals. Results indicate that corrections for Type I error are seldom utilized, even in designs so complicated as to almost guarantee erroneous rejection of null hypotheses.


Model Selection Of Meat Demand System Using The Rotterdam Model And The Almost Ideal Demand System (Aids), Maria Divina S. Paraguas, Anton Abdulbasah Kamil Nov 2005

Model Selection Of Meat Demand System Using The Rotterdam Model And The Almost Ideal Demand System (Aids), Maria Divina S. Paraguas, Anton Abdulbasah Kamil

Journal of Modern Applied Statistical Methods

Aggregated time series data for differentiated meat products namely, beef, pork, poultry, and mutton were used to estimate and analyze Malaysian market demand for meats. The study aimed to select the most appropriate demand model between the equally popular Rotterdam model and the first difference Linear Approximate Almost Ideal Demand System (LA/AIDS) model by using a non-nested test. Both models were accepted, but further diagnostic tests revealed that the first difference LA/AIDS represents more appropriately the Malaysian market demand for meat than the Rotterdam model. Also, the elasticities from the first difference LA/AIDS were found to be more reliable than …


Statistical Methods And Artificial Neural Networks, Mammadagha Mammadov, Berna Yazici, Şenay Yolaçan, Atilla Aslanargun, Ali Fuat YüZer, Embiya Ağaoğlu Nov 2005

Statistical Methods And Artificial Neural Networks, Mammadagha Mammadov, Berna Yazici, Şenay Yolaçan, Atilla Aslanargun, Ali Fuat YüZer, Embiya Ağaoğlu

Journal of Modern Applied Statistical Methods

Artificial Neural Networks and statistical methods are applied on real data sets for forecasting, classification, and clustering problems. Hybrid models for two components are examined on different data sets; tourist arrival forecasting to Turkey, macro-economic problem on rescheduling of the countries’ international debts, and grouping twenty-five European Union member and four candidate countries according to macro-economic indicators.


Jmasm25: Computing Percentiles Of Skew-Normal Distributions, Sikha Bagui, Subhash Bagui Nov 2005

Jmasm25: Computing Percentiles Of Skew-Normal Distributions, Sikha Bagui, Subhash Bagui

Journal of Modern Applied Statistical Methods

An algorithm and code is provided for computing percentiles of skew-normal distributions with parameter λ using Monte Carlo methods. A critical values table was created for various parameter values of λ at various probability levels of α . The table will be useful to practitioners as it is not available in the literature.


Applications Of Some Improved Estimators In Linear Regression, B. M. Golam Kibria Nov 2005

Applications Of Some Improved Estimators In Linear Regression, B. M. Golam Kibria

Journal of Modern Applied Statistical Methods

The problem of estimation of the regression coefficients under multicollinearity situation for the restricted linear model is discussed. Some improve estimators are considered, including the unrestricted ridge regression estimator (URRE), restricted ridge regression estimator (RRRE), shrinkage restricted ridge regression estimator (SRRRE), preliminary test ridge regression estimator (PTRRE), and restricted Liu estimator (RLIUE). The were compared based on the sampling variance-covariance criterion. The RRRE dominates other ridge estimators when the restriction does or does not hold. A numerical example was provided. The RRRE performed equivalently or better than the RLIUE in the sense of having smaller sampling variance.


A Fine-Scale Linkage Disequilibrium Measure Based On Length Of Haplotype Sharing, Yan Wang, Lue Ping Zhao, Sandrine Dudoit Oct 2005

A Fine-Scale Linkage Disequilibrium Measure Based On Length Of Haplotype Sharing, Yan Wang, Lue Ping Zhao, Sandrine Dudoit

U.C. Berkeley Division of Biostatistics Working Paper Series

High-throughput genotyping technologies for single nucleotide polymorphisms (SNP) have enabled the recent completion of the International HapMap Project (Phase I), which has stimulated much interest in studying genome-wide linkage disequilibrium (LD) patterns. Conventional LD measures, such as D' and r-square, are two-point measurements, and their relationship with physical distance is highly noisy. We propose a new LD measure, defined in terms of the correlation coefficient for shared haplotype lengths around two loci, thereby borrowing information from multiple loci. A U-statistic-based estimator of the new LD measure, which takes into consideration the dependence structure of the observed data, is developed and …


Population Intervention Models In Causal Inference, Alan E. Hubbard, Mark J. Van Der Laan Oct 2005

Population Intervention Models In Causal Inference, Alan E. Hubbard, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Marginal structural models (MSM) provide a powerful tool for estimating the causal effect of a] treatment variable or risk variable on the distribution of a disease in a population. These models, as originally introduced by Robins (e.g., Robins (2000a), Robins (2000b), van der Laan and Robins (2002)), model the marginal distributions of treatment-specific counterfactual outcomes, possibly conditional on a subset of the baseline covariates, and its dependence on treatment. Marginal structural models are particularly useful in the context of longitudinal data structures, in which each subject's treatment and covariate history are measured over time, and an outcome is recorded at …


Designed Extension Of Survival Studies: Application To Clinical Trials With Unrecognized Heterogeneity, Yi Li, Mei-Chiung Shih, Rebecca A. Betensky Oct 2005

Designed Extension Of Survival Studies: Application To Clinical Trials With Unrecognized Heterogeneity, Yi Li, Mei-Chiung Shih, Rebecca A. Betensky

Harvard University Biostatistics Working Paper Series

It is well known that unrecognized heterogeneity among patients, such as is conferred by genetic subtype, can undermine the power of randomized trial, designed under the assumption of homogeneity, to detect a truly beneficial treatment. We consider the conditional power approach to allow for recovery of power under unexplained heterogeneity. While Proschan and Hunsberger (1995) confined the application of conditional power design to normally distributed observations, we consider more general and difficult settings in which the data are in the framework of continuous time and are subject to censoring. In particular, we derive a procedure appropriate for the analysis of …


A Pseudolikelihood Approach For Simultaneous Analysis Of Array Comparative Genomic Hybridizations (Acgh), David A. Engler, Gayatry Mohapatra, David N. Louis, Rebecca Betensky Sep 2005

A Pseudolikelihood Approach For Simultaneous Analysis Of Array Comparative Genomic Hybridizations (Acgh), David A. Engler, Gayatry Mohapatra, David N. Louis, Rebecca Betensky

Harvard University Biostatistics Working Paper Series

DNA sequence copy number has been shown to be associated with cancer development and progression. Array-based Comparative Genomic Hybridization (aCGH) is a recent development that seeks to identify the copy number ratio at large numbers of markers across the genome. Due to experimental and biological variations across chromosomes and across hybridizations, current methods are limited to analyses of single chromosomes. We propose a more powerful approach that borrows strength across chromosomes and across hybridizations. We assume a Gaussian mixture model, with a hidden Markov dependence structure, and with random effects to allow for intertumoral variation, as well as intratumoral clonal …


Semiparametric Estimation In General Repeated Measures Problems, Xihong Lin, Raymond J. Carroll Sep 2005

Semiparametric Estimation In General Repeated Measures Problems, Xihong Lin, Raymond J. Carroll

Harvard University Biostatistics Working Paper Series

This paper considers a wide class of semiparametric problems with a parametric part for some covariate effects and repeated evaluations of a nonparametric function. Special cases in our approach include marginal models for longitudinal/clustered data, conditional logistic regression for matched case-control studies, multivariate measurement error models, generalized linear mixed models with a semiparametric component, and many others. We propose profile-kernel and backfitting estimation methods for these problems, derive their asymptotic distributions, and show that in likelihood problems the methods are semiparametric efficient. While generally not true, with our methods profiling and backfitting are asymptotically equivalent. We also consider pseudolikelihood methods …


The Optimal Discovery Procedure: A New Approach To Simultaneous Significance Testing, John D. Storey Sep 2005

The Optimal Discovery Procedure: A New Approach To Simultaneous Significance Testing, John D. Storey

UW Biostatistics Working Paper Series

Significance testing is one of the main objectives of statistics. The Neyman-Pearson lemma provides a simple rule for optimally testing a single hypothesis when the null and alternative distributions are known. This result has played a major role in the development of significance testing strategies that are used in practice. Most of the work extending single testing strategies to multiple tests has focused on formulating and estimating new types of significance measures, such as the false discovery rate. These methods tend to be based on p-values that are calculated from each test individually, ignoring information from the other tests. As …


Mixture Cure Survival Models With Dependent Censoring, Yi Li, Ram C. Tiwari, Subharup Guha Sep 2005

Mixture Cure Survival Models With Dependent Censoring, Yi Li, Ram C. Tiwari, Subharup Guha

Harvard University Biostatistics Working Paper Series

A number of authors have studies the mixture survival model to analyze survival data with nonnegligible cure fractions. A key assumption made by these authors is the independence between the survival time and the censoring time. To our knowledge, no one has studies the mixture cure model in the presence of dependent censoring. To account for such dependence, we propose a more general cure model which allows for dependent censoring. In particular, we derive the cure models from the perspective of competing risks and model the dependence between the censoring time and the survival time using a class of Archimedean …


Semiparametric Normal Transformation Models For Spatially Correlated Survival Data, Yi Li, Xihong Lin Sep 2005

Semiparametric Normal Transformation Models For Spatially Correlated Survival Data, Yi Li, Xihong Lin

Harvard University Biostatistics Working Paper Series

There is an emerging interest in modeling spatially correlated survival data in biomedical and epidemiological studies. In this paper, we propose a new class of semiparametric normal transformation models for right censored spatially correlated survival data. This class of models assumes that survival outcomes marginally follow a Cox proportional hazard model with unspecified baseline hazard, and their joint distribution is obtained by transforming survival outcomes to normal random variables, whose joint distribution is assumed to be multivariate normal with a spatial correlation structure. A key feature of the class of semiparametric normal transformation models is that it provides a rich …


Inference On Survival Data With Covariate Measurement Error - An Imputation-Based Approach, Yi Li, Louise Ryan Sep 2005

Inference On Survival Data With Covariate Measurement Error - An Imputation-Based Approach, Yi Li, Louise Ryan

Harvard University Biostatistics Working Paper Series

We propose a new method for fitting proportional hazards models with error-prone covariates. Regression coefficients are estimated by solving an estimating equation that is the average of the partial likelihood scores based on imputed true covariates. For the purpose of imputation, a linear spline model is assumed on the baseline hazard. We discuss consistency and asymptotic normality of the resulting estimators, and propose a stochastic approximation scheme to obtain the estimates. The algorithm is easy to implement, and reduces to the ordinary Cox partial likelihood approach when the measurement error has a degenerative distribution. Simulations indicate high efficiency and robustness. …


The Optimal Discovery Procedure For Large-Scale Significance Testing, With Applications To Comparative Microarray Experiments, John D. Storey, James Y. Dai, Jeffrey T. Leek Sep 2005

The Optimal Discovery Procedure For Large-Scale Significance Testing, With Applications To Comparative Microarray Experiments, John D. Storey, James Y. Dai, Jeffrey T. Leek

UW Biostatistics Working Paper Series

As much of the focus of genetics and molecular biology has shifted toward the systems level, it has become increasingly important to accurately extract biologically relevant signal from thousands of related measurements. The common property among these high-dimensional biological studies is that the measured features have a rich and largely unknown underlying structure. One example of much recent interest is identifying differentially expressed genes in comparative microarray experiments. We propose a new approach aimed at optimally performing many hypothesis tests in a high-dimensional study. This approach estimates the Optimal Discovery Procedure (ODP), which has recently been introduced and theoretically shown …


Direct Effect Models, Mark J. Van Der Laan, Maya L. Petersen Aug 2005

Direct Effect Models, Mark J. Van Der Laan, Maya L. Petersen

U.C. Berkeley Division of Biostatistics Working Paper Series

The causal effect of a treatment on an outcome is generally mediated by several intermediate variables. Estimation of the component of the causal effect of a treatment that is mediated by a given intermediate variable (the indirect effect of the treatment), and the component that is not mediated by that intermediate variable (the direct effect of the treatment) is often relevant to mechanistic understanding and to the design of clinical and public health interventions. Under the assumption of no-unmeasured confounders for treatment and the intermediate variable, Robins & Greenland (1992) define an individual direct effect as the counterfactual effect of …


Application Of A Multiple Testing Procedure Controlling The Proportion Of False Positives To Protein And Bacterial Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan Aug 2005

Application Of A Multiple Testing Procedure Controlling The Proportion Of False Positives To Protein And Bacterial Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Simultaneously testing multiple hypotheses is important in high-dimensional biological studies. In these situations, one is often interested in controlling the Type-I error rate, such as the proportion of false positives to total rejections (TPPFP) at a specific level, alpha. This article will present an application of the E-Bayes/Bootstrap TPPFP procedure, presented in van der Laan et al. (2005), which controls the tail probability of the proportion of false positives (TPPFP), on two biological datasets. The two data applications include firstly, the application to a mass-spectrometry dataset of two leukemia subtypes, AML and ALL. The protein data measurements include intensity and …


Cross-Validating And Bagging Partitioning Algorithms With Variable Importance, Annette M. Molinaro, Mark J. Van Der Laan Aug 2005

Cross-Validating And Bagging Partitioning Algorithms With Variable Importance, Annette M. Molinaro, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

We present a cross-validated bagging scheme in the context of partitioning algorithms. To explore the benefits of the various bagging scheme, we compare via simulations the predictive ability of single Classification and Regression (CART) Tree with several previously suggested bagging schemes and with our proposed approach. Additionally, a variable importance measure is explained and illustrated.


Test Statistics Null Distributions In Multiple Testing: Simulation Studies And Applications To Genomics, Katherine S. Pollard, Merrill D. Birkner, Mark J. Van Der Laan, Sandrine Dudoit Jul 2005

Test Statistics Null Distributions In Multiple Testing: Simulation Studies And Applications To Genomics, Katherine S. Pollard, Merrill D. Birkner, Mark J. Van Der Laan, Sandrine Dudoit

U.C. Berkeley Division of Biostatistics Working Paper Series

Multiple hypothesis testing problems arise frequently in biomedical and genomic research, for instance, when identifying differentially expressed or co-expressed genes in microarray experiments. We have developed generally applicable resampling-based single-step and stepwise multiple testing procedures (MTP) for control of a broad class of Type I error rates, defined as tail probabilities and expected values for arbitrary functions of the numbers of false positives and rejected hypotheses (Dudoit and van der Laan, 2005; Dudoit et al., 2004a,b; Pollard and van der Laan, 2004; van der Laan et al., 2005, 2004a,b). As argued in the early article of Pollard and van der …


Linear Regression Of Censored Length-Biased Lifetimes, Ying Qing Chen, Yan Wang Jul 2005

Linear Regression Of Censored Length-Biased Lifetimes, Ying Qing Chen, Yan Wang

UW Biostatistics Working Paper Series

Length-biased lifetimes may be collected in observational studies or sample surveys due to biased sampling scheme. In this article, we use a linear regression model, namely, the accelerated failure time model, for the population lifetime distributions in regression analysis of the length-biased lifetimes. It is discovered that the associated regression parameters are invariant under the length-biased sampling scheme. According to this discovery, we propose the quasi partial score estimating equations to estimate the population regression parameters. The proposed methodologies are evaluated and demonstrated by simulation studies and an application to actual data set.


A Note On The Construction Of Counterfactuals And The G-Computation Formula, Zhuo Yu, Mark J. Van Der Laan Jun 2005

A Note On The Construction Of Counterfactuals And The G-Computation Formula, Zhuo Yu, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Robins' causal inference theory assumes existence of treatment specific counterfactual variables so that the observed data augmented by the counterfactual data will satisfy a consistency and a randomization assumption. In this paper we provide an explicit function that maps the observed data into a counterfactual variable which satisfies the consistency and randomization assumptions. This offers a practically useful imputation method for counterfactuals. Gill & Robins [2001]'s construction of counterfactuals can be used as an imputation method in principle, but it is very hard to implement in practice. Robins [1987] shows that the counterfactual distribution can be identified from the observed …


Cross-Validated Bagged Learning, Mark J. Van Der Laan, Sandra E. Sinisi, Maya L. Petersen Jun 2005

Cross-Validated Bagged Learning, Mark J. Van Der Laan, Sandra E. Sinisi, Maya L. Petersen

U.C. Berkeley Division of Biostatistics Working Paper Series

Many applications aim to learn a high dimensional parameter of a data generating distribution based on a sample of independent and identically distributed observations. For example, the goal might be to estimate the conditional mean of an outcome given a list of input variables. In this prediction context, Breiman (1996a) introduced bootstrap aggregating (bagging) as a method to reduce the variance of a given estimator at little cost to bias. Bagging involves applying the estimator to multiple bootstrap samples, and averaging the result across bootstrap samples. In order to deal with the curse of dimensionality, typical practice has been to …