Open Access. Powered by Scholars. Published by Universities.®

Statistical Theory Commons™

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type

Articles 1501 - 1530 of 1633

Full-Text Articles in Statistical Theory

Inference For The Population Total From Probability-Proportional-To-Size Samples Based On Predictions From A Penalized Spline Nonparametric Model, Hui Zheng, Rod Little Aug 2003

Inference For The Population Total From Probability-Proportional-To-Size Samples Based On Predictions From A Penalized Spline Nonparametric Model, Hui Zheng, Rod Little

The University of Michigan Department of Biostatistics Working Paper Series

Inference about the finite population total from probability-proportional-to-size (PPS) samples is considered. In previous work (Zheng and Little, 2003), penalized spline (p-spline) nonparametric model-based estimators were shown to generally outperform the Horvitz-Thompson (HT) and generalized regression (GR) estimators in terms of the root mean squared error. In this article we develop model-based, jackknife and balanced repeated replicate variance estimation methods for the p-spline based estimators. Asymptotic properties of the jackknife method are discussed. Simulations show that p-spline point estimators and their jackknife standard errors lead to inferences that are superior to HT or GR based inferences. This suggests that nonparametric …


Locally Efficient Estimation Of Nonparametric Causal Effects On Mean Outcomes In Longitudinal Studies, Romain Neugebauer, Mark J. Van Der Laan Jul 2003

Locally Efficient Estimation Of Nonparametric Causal Effects On Mean Outcomes In Longitudinal Studies, Romain Neugebauer, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Marginal Structural Models (MSM) have been introduced by Robins (1998a) as a powerful tool for causal inference as they directly model causal curves of interest, i.e. mean treatment-specific outcomes possibly adjusted for baseline covariates. Two estimators of the corresponding MSM parameters of interest have been proposed, see van der Laan and Robins (2002): the Inverse Probability of Treatment Weighted (IPTW) and the Double Robust (DR) estimators. A parametric MSM approach to causal inference has been favored since the introduction of MSM. It relies on correct specification of a parametric MSM to consistently estimate the parameter of interest using the IPTW …


Resampling-Based Multiple Testing: Asymptotic Control Of Type I Error And Applications To Gene Expression Data, Katherine S. Pollard, Mark J. Van Der Laan Jun 2003

Resampling-Based Multiple Testing: Asymptotic Control Of Type I Error And Applications To Gene Expression Data, Katherine S. Pollard, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

We define a general statistical framework for multiple hypothesis testing and show that the correct null distribution for the test statistics is obtained by projecting the true distribution of the test statistics onto the space of mean zero distributions. For common choices of test statistics (based on an asymptotically linear parameter estimator), this distribution is asymptotically multivariate normal with mean zero and the covariance of the vector influence curve for the parameter estimator. This test statistic null distribution can be estimated by applying the non-parametric or parametric bootstrap to correctly centered test statistics. We prove that this bootstrap estimated null …


Maximization By Parts In Likelihood Inference, Peter Xuekun Song, Yanqin Fan, Jack Kalbfleisch Jun 2003

Maximization By Parts In Likelihood Inference, Peter Xuekun Song, Yanqin Fan, Jack Kalbfleisch

The University of Michigan Department of Biostatistics Working Paper Series

This paper presents and examines a new algorithm for solving a score equation for the maximum likelyhood estimate in certain problems of practical interest. The method circumvents the need to compute second order derivaties of the full likelihood function. It exploits the structure of certain models that yield a natural decomposition of a very complicated likelihood function. In this decomposition, the first part is a log likelihood from a simply analyzed model and the second part is used to update estimates from the first. Convergence properties of this fixed point algorithm are examined and asymptotics are derived for estimators obtained …


Double Robust Estimation In Longitudinal Marginal Structural Models, Zhuo Yu, Mark J. Van Der Laan Jun 2003

Double Robust Estimation In Longitudinal Marginal Structural Models, Zhuo Yu, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Consider estimation of causal parameters in a marginal structural model for the discrete intensity of the treatment specific counting process (e.g. hazard of a treatment specific survival time) based on longitudinal observational data on treatment, covariates and survival. We assume the sequential randomization assumption (SRA) on the treatment assignment mechanism and the so called experimental treatment assignment assumption which is needed to identify the causal parameters from the observed data distribution. Under SRA, the likelihood of the observed data structure factorizes in the auxiliary treatment mechanism and the partial likelihood consisting of the product over time of conditional distributions of …


A New Confidence Interval For The Difference Between Two Binomial Proportions Of Paired Data, Xiao-Hua Zhou, Gengsheng Qin Jun 2003

A New Confidence Interval For The Difference Between Two Binomial Proportions Of Paired Data, Xiao-Hua Zhou, Gengsheng Qin

UW Biostatistics Working Paper Series

Motivated by a study on comparing sensitivities and specificities of two diagnostic tests in a paired design when the sample size is small, we first derived an Edgeworth expansion for the studentized difference between two binomial proportions of paired data. The Edgeworth expansion can help us understand why the usual Wald interval for the difference has poor coverage performance in the small sample size. Based on the Edgeworth expansion, we then derived a transformation based confidence interval for the difference. The new interval removes the skewness in the Edgeworth expansion; the new interval is easy to compute, and its coverage …


Supervised Detection Of Regulatory Motifs In Dna Sequences, Sunduz Keles, Mark J. Van Der Laan, Sandrine Dudoit, Biao Xing, Michael B. Eisen May 2003

Supervised Detection Of Regulatory Motifs In Dna Sequences, Sunduz Keles, Mark J. Van Der Laan, Sandrine Dudoit, Biao Xing, Michael B. Eisen

U.C. Berkeley Division of Biostatistics Working Paper Series

Identification of transcription factor binding sites (regulatory motifs) is a major interest in contemporary biology. We propose a new likelihood based method, COMODE, for identifying structural motifs in DNA sequences. Commonly used methods (e.g. MEME, Gibbs sampler) model binding sites as families of sequences described by a position weight matrix (PWM) and identify PWMs that maximize the likelihood of observed sequence data under a simple multinomial mixture model. This model assumes that the positions of the PWM correspond to independent multinomial distributions with four cell probabilities. We address supervising the search for DNA binding sites using the information derived from …


Improved Confidence Intervals For The Sensitivity At A Fixed Level Of Specificity Of A Continuous-Scale Diagnostic Test, Xiao-Hua Zhou, Gengsheng Qin May 2003

Improved Confidence Intervals For The Sensitivity At A Fixed Level Of Specificity Of A Continuous-Scale Diagnostic Test, Xiao-Hua Zhou, Gengsheng Qin

UW Biostatistics Working Paper Series

For a continuous-scale test, it is an interest to construct a confidence interval for the sensitivity of the diagnostic test at the cut-off that yields a predetermined level of its specificity (eg. 80%, 90%, or 95%). IN this paper we proposed two new intervals for the sensitivity of a continuous-scale diagnostic test at a fixed level of specificity. We then conducted simulation studies to compare the relative performance of these two intervals with the best existing BCa bootstrap interval, proposed by Platt et al. (2000). Our simulation results showed that the newly proposed intervals are better than the BCa bootstrap …


Bootstrap Confidence Intervals For Medical Costs With Censored Observations, Hongyu Jiang, Xiao-Hua Zhou May 2003

Bootstrap Confidence Intervals For Medical Costs With Censored Observations, Hongyu Jiang, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

Medical costs data with administratively censored observations often arise in cost-effectiveness studies of treatments for life threatening diseases. Mean of medical costs incurred from the start of a treatment till death or certain timepoint after the implementation of treatment is frequently of interest. In many situations, due to the skewed nature of the cost distribution and non-uniform rate of cost accumulation over time, the currently available normal approximation confidence interval has poor coverage accuracy. In this paper, we proposed a bootstrap confidence interval for the mean of medical costs with censored observations. In simulation studies, we showed that the proposed …


New Intervals For The Difference Between Two Independent Binomial Proportions, Xiao-Hua Zhou, Min Tsao, Gengsheng Qin May 2003

New Intervals For The Difference Between Two Independent Binomial Proportions, Xiao-Hua Zhou, Min Tsao, Gengsheng Qin

UW Biostatistics Working Paper Series

In this paper we gave an Edgeworth expansion for the studentized difference of two binomial proportions. We then proposed two new intervals by correcting the skewness in the Edgeworth expansion in a direct and an indirect way. Such the bias-correct confidence intervals are easy to compute, and their coverage probabilities converge to the nominal level at a rate of O(n-½), where n is the size of the combined samples. Our simulation results suggest tat in finite samples the new interval based on the indirect method have the similar performance to the two best existing intervals in terms of coverage accuracy …


A Bootstrap Confidence Interval Procedure For The Treatment Effect Using Propensity Score Subclassification, Wanzhu Tu, Xiao-Hua Zhou May 2003

A Bootstrap Confidence Interval Procedure For The Treatment Effect Using Propensity Score Subclassification, Wanzhu Tu, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

In the analysis of observational studies, propensity score subclassification has been shown to be a powerful method for adjusting unbalanced covariates for the purpose of causal inferences. One practical difficulty in carrying out such an analysis is to obtain a correct variance estimate for such inferences, while reducing bias in the estimate of the treatment effect due to an imbalance in the measured covariates. In this paper, we propose a bootstrap procedure for the inferences concerning the average treatment effect; our bootstrap method is based on an extension of Efron’s bias-corrected accelerated (BCa) bootstrap confidence interval to a two-sample problem. …


Semiparametric Regression Models With Missing Data: The Mathematics In The Work Of Robins Et Al., Menggang Yu, Bin Nan May 2003

Semiparametric Regression Models With Missing Data: The Mathematics In The Work Of Robins Et Al., Menggang Yu, Bin Nan

The University of Michigan Department of Biostatistics Working Paper Series

This review is an attempt to understand the landmark papers of Robins, Rotnitzky, and Zhao (1994) and Robins and Rotnitzky (1992). We revisit their main results and corresponding proofs using the theory outlined in the monograph by Bickel, Klaassen, Ritov, and Wellner (1993). We also discuss an illustrative example to show the details of applying these theoretical results.


Random Number Generators, George Marsaglia May 2003

Random Number Generators, George Marsaglia

Journal of Modern Applied Statistical Methods

The quasi-negative-binomial distribution was applied to queuing theory for determining the distribution of total number of customers served before the queue vanishes under certain assumptions. Some structural properties (probability generating function, convolution, mode and recurrence relation) for the moments of quasi-negative-binomial distribution are discussed. The distribution’s characterization and its relation with other distributions were investigated. A computer program was developed using R to obtain ML estimates and the distribution was fitted to some observed sets of data to test its goodness of fit.


Without Supporting Statistical Evidence, Where Would Reported Measures Of Substantive Importance Lead? To No Good Effect, Anthony J. Onwuegbuzie, Joel R. Levin May 2003

Without Supporting Statistical Evidence, Where Would Reported Measures Of Substantive Importance Lead? To No Good Effect, Anthony J. Onwuegbuzie, Joel R. Levin

Journal of Modern Applied Statistical Methods

Although estimating substantive importance (in the form of reporting effect sizes) has recently received widespread endorsement, its use has not been subjected to the same degree of scrutiny as has statistical hypothesis testing. As such, many researchers do not seem to be aware that certain of the same criticisms launched against the latter can also be aimed at the former. Our purpose here is to highlight major concerns about effect sizes and their estimation. In so doing, we argue that effect size measures per se are not the hoped-for panaceas for interpreting empirical research findings. Further, we contend that if …


Bayesian Analysis Of Poverty Rates: The Case Of Vietnamese Provinces, Dominique Haughton, Nguyen Phong May 2003

Bayesian Analysis Of Poverty Rates: The Case Of Vietnamese Provinces, Dominique Haughton, Nguyen Phong

Journal of Modern Applied Statistical Methods

This paper presents a Bayesian analysis of poverty rates in urban Ho Chi Minh City and rural Nghe An province in Vietnam. Using mixtures of beta distributions as priors for the poverty rates, we find that, when the prior is reasonably informative, our approach yields more accurate estimated poverty rates than a frequentist approach. On the other hand, we find that, in the presence of poor/non-poor misclassification, average probabilities of posterior credible intervals for poverty rates can fall well short of .95 even with sample sizes such as 2000 or 3000 when the width of the interval is for example …


Not All Effects Are Created Equal: A Rejoinder To Sawilowsky, J. Kyle Roberts, Robin K. Henson May 2003

Not All Effects Are Created Equal: A Rejoinder To Sawilowsky, J. Kyle Roberts, Robin K. Henson

Journal of Modern Applied Statistical Methods

In the continuing debate over the use and utility of effect sizes, more discussion often helps to both clarify and syncretize methodological views. Here, further defense is given of Roberts & Henson (2002) in terms of measuring bias in Cohen’s d, and a rejoinder to Sawilowsky (2003) is presented.


Comparison Of Estimates Of Proprietary And Syndicated Methods In Auto Industry Surveys, Daniel X. Wang May 2003

Comparison Of Estimates Of Proprietary And Syndicated Methods In Auto Industry Surveys, Daniel X. Wang

Journal of Modern Applied Statistical Methods

Proprietary and syndicate surveys are often used in assessing appeal and initial quality of new vehicles for automobile manufactures. This study discusses the difference between the two types of studies, and proposes a computer simulation based method for checking the appropriateness of the comparisons.


Steady State Analysis Of An M/D/2 Queue With Bernoulli Schedule Server Vacations, Kailash C. Madan, Walid Abu-Dayyeh, Firas Tayyan May 2003

Steady State Analysis Of An M/D/2 Queue With Bernoulli Schedule Server Vacations, Kailash C. Madan, Walid Abu-Dayyeh, Firas Tayyan

Journal of Modern Applied Statistical Methods

We examine an M/D/2 queue with Bernoulli schedules and a single vacation policy. We have assumed Poisson arrivals waiting in a single queue and two parallel servers who provide identical deterministic service to customers on first-come, first-served basis. We consider two models; in one we assume that after completion of a service both servers can take a vacation while in the other we assume that only one may take a vacation. The vacation periods in both models are assumed to be exponential. We obtain steady state probability generating functions of system size for various states of the servers.


A Different Future For Social And Behavioral Science Research, Shlomo S. Sawilowsky May 2003

A Different Future For Social And Behavioral Science Research, Shlomo S. Sawilowsky

Journal of Modern Applied Statistical Methods

The dissemination of intervention and treatment outcomes as effect sizes bounded by conf idence intervals in order to think meta-analytically was promoted in a recent article in Educational Researcher. I raise concerns with unfettered reporting of effect sizes, point out the con in confidence interval, and caution against thinking meta-analytically. Instead, cataloging effect sizes is recommended for sample size estimation and power analysis to improve social and behavioral science research.


The Trouble With Interpreting Statistically Nonsignificant Effect Sizes In Single-Study Investigations, Joel R. Levin, Daniel H. Robinson May 2003

The Trouble With Interpreting Statistically Nonsignificant Effect Sizes In Single-Study Investigations, Joel R. Levin, Daniel H. Robinson

Journal of Modern Applied Statistical Methods

In this commentary, we offer a perspective on the problem of authors reporting and interpreting effect sizes in the absence of formal statistical tests of their chanceness. The perspective reinforces our previous distinction between single-study investigations and multiple-study syntheses.


A Recursive Algorithm For Fractionally Differencing Long Data Series, Joseph Mccarthy, Robert Disario, Hakan Saraoglu May 2003

A Recursive Algorithm For Fractionally Differencing Long Data Series, Joseph Mccarthy, Robert Disario, Hakan Saraoglu

Journal of Modern Applied Statistical Methods

We propose a recursive algorithm to fractionally difference time series data. The algorithm eliminates the need to evaluate the gamma function directly, and hence avoids the overflow problem that arises when fractionally differencing a long data series. The proposed algorithm can be implemented using any general matrix programming language. An implementation using SAS is presented. The algorithm and the code provide a practical approach to including fractional differencing as part of a time series data analysis.


Modeling Correlated Time-Varying Covariate Effects In A Cox-Type Regression Model, Mourad Tighiouart May 2003

Modeling Correlated Time-Varying Covariate Effects In A Cox-Type Regression Model, Mourad Tighiouart

Journal of Modern Applied Statistical Methods

In this paper, I extend the proposed model by McKeague and Tighiouart (2000) to handle time-varying correlated covariate effects for the analysis of survival data. I use the conditional predictive ordinates (CPO’s) for model comparison and the methodology is illustrated by an application to nasopharynx cancer survival data. A reversible jump MCMC sampler to estimate the CPO’s will be presented.


Performing Two-Way Analysis Of Variance Under Variance Heterogeneity, Scott J. Richter, Mark E. Payton May 2003

Performing Two-Way Analysis Of Variance Under Variance Heterogeneity, Scott J. Richter, Mark E. Payton

Journal of Modern Applied Statistical Methods

Small sample properties of the method proposed by Brunner et al. (1997) for performing two-way analysis of variance are compared to those of the normal based ANOVA method for factorial arrangements. Different effect sizes, sample sizes, and error structures are utilized in a simulation study to compare type I error rates and power of the two methods. An SAS program is also presented to assist those wishing to implement the Brunner method to real data.


The Way Ahead In Qualitative Computing, Tom Richards, Lyn Richards May 2003

The Way Ahead In Qualitative Computing, Tom Richards, Lyn Richards

Journal of Modern Applied Statistical Methods

Specialized computer programs for Qualitative Research in social sciences have greatly changed ways of doing QR, the reliability and comprehensiveness of results, the ability to inspect and challenge a researcher’s working, and the relationship with quantitative methods in social research. This article explores these claims in the context of N6 (NUD*IST) and NVivo, the two programs designed by the authors; and considers possible future developments in the field.


Homogeneous Markov Processes For Breast Cancer Analysis, Ricardo Ocaña-Rilola, Emilio Sanchez-Cantalejo, Carmen Martinez-Garcia May 2003

Homogeneous Markov Processes For Breast Cancer Analysis, Ricardo Ocaña-Rilola, Emilio Sanchez-Cantalejo, Carmen Martinez-Garcia

Journal of Modern Applied Statistical Methods

Sometimes, the introduction of covariates in stochastic processes is required to study their effect on disease history events. However these types of models increase the complexity of analysis, even for simpler processes, and standard software to analyse stochastic processes is limited. In this paper, a method for fitting homogeneous Markov models with covariates is proposed for analysing breast cancer data. Specific software for this purpose has been implemented.


Incorporating Sampling Weights Into The Generalizability Theory For Large-Scale Analyses, Christopher W. T. Chiu, Ronald S. Fesco May 2003

Incorporating Sampling Weights Into The Generalizability Theory For Large-Scale Analyses, Christopher W. T. Chiu, Ronald S. Fesco

Journal of Modern Applied Statistical Methods

Large scale studies frequently use complex sampling procedures, disproportionate sampling weights, and adjustment techniques to account for potential bias due to nonresponses and to ensure that results from the sample can be generalized to a larger population. Survey researchers are concerned about measurement error and the use of weights in developing models. Consequently, multiple weighting factors are used and these weighting factors are manifested as a final survey (composite) weight available for analysis. We developed a method to incorporate an external weighting factor like this for analyses of measurement errors in the theory of generalizability to provide researchers with a …


Using Multinomial Logistic Models To Predict Adolescent Behavioral Risk, Chao-Ying Joanne Peng, Rebecca Naegle Nichols May 2003

Using Multinomial Logistic Models To Predict Adolescent Behavioral Risk, Chao-Ying Joanne Peng, Rebecca Naegle Nichols

Journal of Modern Applied Statistical Methods

Multinomial logistic regression was applied to data comprising 432 adolescents’ self reports of engagement in risky behaviors. Results showed that gender, intention to drop from the school, family structure, self-esteem, and emotional risk were effective predictors collectively. Three methodological issues were highlighted: (1) the use of odds ratio, (2) the absence of an extension of the Hosmer and Lemeshow test for multinomial logistic models, and (3) the missing data problem. Psychologists and educators can utilize findings to plan prevention programs, as well as to apply the versatile and effective logistic technique in psychological, educational, and health research concerning adolescents.


Was Monte Carlo Necessary?, Thomas R. Knapp May 2003

Was Monte Carlo Necessary?, Thomas R. Knapp

Journal of Modern Applied Statistical Methods

In the critique that follows, I have attempted to summarize the principal disagreements between Sawilowsky and Roberts & Henson regarding the reporting and interpreting of statistically non-significant effect sizes, and to provide my own personal evaluations of their respective arguments.


Trivials: The Birth, Sale, And Final Production Of Meta-Analysis, Shlomo S. Sawilowsky May 2003

Trivials: The Birth, Sale, And Final Production Of Meta-Analysis, Shlomo S. Sawilowsky

Journal of Modern Applied Statistical Methods

The structure of the first invited debate in JMASM is to present a target article (Sawilowsky, 2003), provide an opportunity for a response (Roberts & Henson, 2003), and to follow with independent comments from noted scholars in the field (Knapp, 2003; Levin & Robinson, 2003). In this rejoinder, I provide a correction and a clarification in an effort to bring some closure to the debate. The intension, however, is not to rehash previously made points, even where I disagree with the response of Roberts & Henson (2003).


The Importance Of Fortran In The 21st Century, Walt Brainerd May 2003

The Importance Of Fortran In The 21st Century, Walt Brainerd

Journal of Modern Applied Statistical Methods

A brief discussion on the history and purpose of Fortran for scientific and engineering computing is given. This leads to the role Fortran, in its various environments, will likely play well into the 21st century.