Open Access. Powered by Scholars. Published by Universities.®

Statistical Theory Commons

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type

Articles 1111 - 1140 of 1633

Full-Text Articles in Statistical Theory

Generalized Linear Mixed-Effects Models For The Analysis Of Odor Detection Data, Sandra Hall, Matthew S. Mayo, Xu-Feng Niu, James C. Walker Nov 2007

Generalized Linear Mixed-Effects Models For The Analysis Of Odor Detection Data, Sandra Hall, Matthew S. Mayo, Xu-Feng Niu, James C. Walker

Journal of Modern Applied Statistical Methods

Olfactory detection has become a science of interest. Seven individuals’ odor detection abilities are explored and an attempt is made to characterize all subjects with one generalized linear mixed effects model. Two methods of fitting the models were used and simulations were conducted to discover which method yielded the best results.


Operating Characteristics Of The Dif Mimic Approach Using Jöreskog’S Covariance Matrix With Ml And Wls Estimation For Short Scales, Michaela N. Gelin, Bruno D. Zumbo Nov 2007

Operating Characteristics Of The Dif Mimic Approach Using Jöreskog’S Covariance Matrix With Ml And Wls Estimation For Short Scales, Michaela N. Gelin, Bruno D. Zumbo

Journal of Modern Applied Statistical Methods

Type I error rate of a structural equation modeling (SEM) approach for investigating differential item functioning (DIF) in short scales was studied. Muthén’s SEM model for DIF was examined using a covariance matrix (Jöreskog, 2002). It is conditioned on the latent variable, while testing the effect of the grouping variable over-and-above the underlying latent variable. Thus, it is a multiple-indicators, multiple-causes (MIMIC) DIF model. Type I error rates were determined using data reflective of short scales with ordinal item response formats typically found in the social and behavioral sciences. Results indicate Type I error rates for the DIF MIMIC model, …


A Note On Targeted Maximum Likelihood And Right Censored Data, Mark J. Van Der Laan, Daniel Rubin Oct 2007

A Note On Targeted Maximum Likelihood And Right Censored Data, Mark J. Van Der Laan, Daniel Rubin

U.C. Berkeley Division of Biostatistics Working Paper Series

A popular way to estimate an unknown parameter is with substitution, or evaluating the parameter at a likelihood based fit of the data generating density. In many cases, such estimators have substantial bias and can fail to converge at the parametric rate. van der Laan and Rubin (2006) introduced targeted maximum likelihood learning, removing these shackles from substitution estimators, which were made in full agreement with the locally efficient estimating equation procedures as presented in Robins and Rotnitzsky (1992) and van der Laan and Robins (2003). This note illustrates how targeted maximum likelihood can be applied in right censored data …


Detailed Version: Analyzing Direct Effects In Randomized Trials With Secondary Interventions: An Application To Hiv Prevention Trials, Michael A. Rosenblum, Nicholas P. Jewell, Mark J. Van Der Laan, Stephen Shiboski, Ariane Van Der Straten, Nancy Padian Oct 2007

Detailed Version: Analyzing Direct Effects In Randomized Trials With Secondary Interventions: An Application To Hiv Prevention Trials, Michael A. Rosenblum, Nicholas P. Jewell, Mark J. Van Der Laan, Stephen Shiboski, Ariane Van Der Straten, Nancy Padian

U.C. Berkeley Division of Biostatistics Working Paper Series

This is the detailed technical report that accompanies the paper “Analyzing Direct Effects in Randomized Trials with Secondary Interventions: An Application to HIV Prevention Trials” (an unpublished, technical report version of which is available online at http://www.bepress.com/ucbbiostat/paper223).

The version here gives full details of the models for the time-dependent analysis, and presents further results in the data analysis section. The Methods for Improving Reproductive Health in Africa (MIRA) trial is a recently completed randomized trial that investigated the effect of diaphragm and lubricant gel use in reducing HIV infection among susceptible women. 5,045 women were randomly assigned to either the …


Optimal Propensity Score Stratification, Jessica A. Myers, Thomas A. Louis Oct 2007

Optimal Propensity Score Stratification, Jessica A. Myers, Thomas A. Louis

Johns Hopkins University, Dept. of Biostatistics Working Papers

Stratifying on propensity score in observational studies of treatment is a common technique used to control for bias in treatment assignment; however, there have been few studies of the relative efficiency of the various ways of forming those strata. The standard method is to use the quintiles of propensity score to create subclasses, but this choice is not based on any measure of performance either observed or theoretical. In this paper, we investigate the optimal subclassification of propensity scores for estimating treatment effect with respect to mean squared error of the estimate. We consider the optimal formation of subclasses within …


Multiple Model Evaluation Absent The Gold Standard Via Model Combination, Edwin J. Iversen, Jr., Giovanni Parmigiani, Sining Chen Oct 2007

Multiple Model Evaluation Absent The Gold Standard Via Model Combination, Edwin J. Iversen, Jr., Giovanni Parmigiani, Sining Chen

Johns Hopkins University, Dept. of Biostatistics Working Papers

We describe a method for evaluating an ensemble of predictive models given a sample of observations comprising the model predictions and the outcome event measured with error. Our formulation allows us to simultaneously estimate measurement error parameters, true outcome — aka the gold standard — and a relative weighting of the predictive scores. We describe conditions necessary to estimate the gold standard and for these estimates to be calibrated and detail how our approach is related to, but distinct from, standard model combination techniques. We apply our approach to data from a study to evaluate a collection of BRCA1/BRCA2 gene …


Analyzing Direct Effects In Randomized Trials With Secondary Interventions , Michael Rosenblum, Nicholas P. Jewell, Mark J. Van Der Laan, Stephen Shiboski, Ariane Van Der Straten, Nancy Padian Sep 2007

Analyzing Direct Effects In Randomized Trials With Secondary Interventions , Michael Rosenblum, Nicholas P. Jewell, Mark J. Van Der Laan, Stephen Shiboski, Ariane Van Der Straten, Nancy Padian

U.C. Berkeley Division of Biostatistics Working Paper Series

The Methods for Improving Reproductive Health in Africa (MIRA) trial is a recently completed randomized trial that investigated the effect of diaphragm and lubricant gel use in reducing HIV infection among susceptible women. 5,045 women were randomly assigned to either the active treatment arm or not. Additionally, all subjects in both arms received intensive condom counselling and provision, the "gold standard" HIV prevention barrier method. There was much lower reported condom use in the intervention arm than in the control arm, making it difficult to answer important public health questions based solely on the intention-to-treat analysis. We adapt an analysis …


Comparing Trends In Cancer Rates Across Overlapping Regions, Yi Li, Ram C. Tiwari Aug 2007

Comparing Trends In Cancer Rates Across Overlapping Regions, Yi Li, Ram C. Tiwari

Harvard University Biostatistics Working Paper Series

No abstract provided.


Correcting Instrumental Variables Estimators For Systematic Measurement Error, Stijn Vansteelandt, Manoochehr Babanezhad, Els Goetghebeur Aug 2007

Correcting Instrumental Variables Estimators For Systematic Measurement Error, Stijn Vansteelandt, Manoochehr Babanezhad, Els Goetghebeur

Harvard University Biostatistics Working Paper Series

No abstract provided.


A Statistical Evaluation Of Algorithms For Independently Seeding Pseudo-Random Number Generators Of Type Multiplicative Congruential (Lehmer-Class)., Robert Grisham Stewart Aug 2007

A Statistical Evaluation Of Algorithms For Independently Seeding Pseudo-Random Number Generators Of Type Multiplicative Congruential (Lehmer-Class)., Robert Grisham Stewart

Electronic Theses and Dissertations

To be effective, a linear congruential random number generator (LCG) should produce values that are (a) uniformly distributed on the unit interval (0,1) excluding endpoints and (b) substantially free of serial correlation. It has been found that many statistical methods produce inflated Type I error rates for correlated observations. Theoretically, independently seeding an LCG under the following conditions attenuates serial correlation: (a) simple random sampling of seeds, (b) non-replicate streams, (c) non-overlapping streams, and (d) non-adjoining streams. Accordingly, 4 algorithms (each satisfying at least 1 condition) were developed: (a) zero-leap, (b) fixed-leap, (c) scaled random-leap, and (d) unscaled random-leap. Note …


Empirical Efficiency Maximization, Daniel B. Rubin, Mark J. Van Der Laan Jul 2007

Empirical Efficiency Maximization, Daniel B. Rubin, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

It has long been recognized that covariate adjustment can increase precision, even when it is not strictly necessary. The phenomenon is particularly emphasized in clinical trials, whether using continuous, categorical, or censored time-to-event outcomes. Adjustment is often straightforward when a discrete covariate partitions the sample into a handful of strata, but becomes more involved when modern studies collect copious amounts of baseline information on each subject.

The dilemma helped motivate locally efficient estimation for coarsened data structures, as surveyed in the books of van der Laan and Robins (2003) and Tsiatis (2006). Here one fits a relatively small working model …


Assessment Of A Cgh-Based Genetic Instability, David A. Engler, Yiping Shen, J F. Gusella, Rebecca A. Betensky Jul 2007

Assessment Of A Cgh-Based Genetic Instability, David A. Engler, Yiping Shen, J F. Gusella, Rebecca A. Betensky

Harvard University Biostatistics Working Paper Series

No abstract provided.


Survival Analysis With Large Dimensional Covariates: An Application In Microarray Studies, David A. Engler, Yi Li Jul 2007

Survival Analysis With Large Dimensional Covariates: An Application In Microarray Studies, David A. Engler, Yi Li

Harvard University Biostatistics Working Paper Series

Use of microarray technology often leads to high-dimensional and low- sample size data settings. Over the past several years, a variety of novel approaches have been proposed for variable selection in this context. However, only a small number of these have been adapted for time-to-event data where censoring is present. Among standard variable selection methods shown both to have good predictive accuracy and to be computationally efficient is the elastic net penalization approach. In this paper, adaptation of the elastic net approach is presented for variable selection both under the Cox proportional hazards model and under an accelerated failure time …


Super Learner, Mark J. Van Der Laan, Eric C. Polley, Alan E. Hubbard Jul 2007

Super Learner, Mark J. Van Der Laan, Eric C. Polley, Alan E. Hubbard

U.C. Berkeley Division of Biostatistics Working Paper Series

Previous articles (van der Laan and Dudoit (2003); van der Laan et al. (2006); Sinisi et al. (2007)) advertised and theoretically validated the use of cross-validation to select among many candidate estimators to compute a so called super learner which outperforms any of the given candidate estimators. The theoretical basis was provided for this super learner based on oracle results for the cross-validation selector (e.g., van der Laan and Dudoit (2003); van der Laan et al. (2006)) and in Sinisi et al. (2007). In addition, these papers contained a practical demonstration of the adaptivity of this so called super learner …


Evaluating The Roc Performance Of Markers For Future Events, Margaret Pepe, Yingye Zheng, Yuying Jin May 2007

Evaluating The Roc Performance Of Markers For Future Events, Margaret Pepe, Yingye Zheng, Yuying Jin

UW Biostatistics Working Paper Series

Receiver operating characteristic (ROC) curves play a central role in the evaluation of biomarkers and tests for disease diagnosis. Predictors for event time outcomes can also be evaluated with ROC curves, but the time lag between marker measurement and event time must be acknowledged. We discuss different definitions of time-dependent ROC curves in the context of real applications. Several approaches have been proposed for estimation. We contrast retrospective versus prospective methods in regards to assumptions and flexibility, including their capacities to incorporate censored data, competing risks and different sampling schemes. Applications to two datasets are presented.


Review Of The Maximum Likelihood Functions For Right Censored Data. A New Elementary Derivation., Stefano Patti, Elia Biganzoli, Patrizia Boracchi May 2007

Review Of The Maximum Likelihood Functions For Right Censored Data. A New Elementary Derivation., Stefano Patti, Elia Biganzoli, Patrizia Boracchi

COBRA Preprint Series

Censoring is a well known feature recurrent in the analysis of lifetime data, occurring in the model when exact lifetimes can be collected for only a representative portion of the surveyed individuals. If lifetimes are known only to exceed some given values, it is referred to as right censoring. In this paper we propose a systematization and a new derivation of the likelihood function for right censored sampling schemes; calculations are reported and assumptions are carefully stated. The sampling schemes considered (Type I, II and Random Censoring) give rise to the same ML function. Only the knowledge of elementary probability …


Ordinal Versions Of Coefficients Alpha And Theta For Likert Rating Scales, Bruno D. Zumbo, Anne M. Gadermann, Cornelia Zeisser May 2007

Ordinal Versions Of Coefficients Alpha And Theta For Likert Rating Scales, Bruno D. Zumbo, Anne M. Gadermann, Cornelia Zeisser

Journal of Modern Applied Statistical Methods

Two new reliability indices, ordinal coefficient alpha and ordinal coefficient theta, are introduced. A simulation study was conducted in order to compare the new ordinal reliability estimates to each other and to coefficient alpha with Likert data. Results indicate that ordinal coefficients alpha and theta are consistently suitable estimates of the theoretical reliability, regardless of the magnitude of the theoretical reliability, the number of scale points, and the skewness of the scale point distributions. In contrast, coefficient alpha is in general a negatively biased estimate of reliability. The use of ordinal coefficients alpha and theta as alternatives to coefficient alpha …


Lq-Moments For Statistical Analysis Of Extreme Events, Ani Shabri, Abdul Aziz Jemain May 2007

Lq-Moments For Statistical Analysis Of Extreme Events, Ani Shabri, Abdul Aziz Jemain

Journal of Modern Applied Statistical Methods

Statistical analysis of extremes is conducted for predicting large return periods events. LQ-moments that are based on linear combinations are reviewed for characterizing the upper quantiles of distributions and larger events in data. The LQ-moments method is presented based on a new quick estimator using five points quantiles and the weighted kernel estimator to estimate the parameters of the generalized extreme value (GEV) distribution. Monte Carlo methods illustrate the performance of LQ-moments in fitting the GEV distribution to both GEV and non-GEV samples. The proposed estimators of the GEV distribution were compared with conventional L-moments and LQ-moments based on linear …


Jmasm 26: Hettmansperger And Mckean Linear Model Aligned Rank Test For The Single Covariate And One-Way Ancova Case (Sas), Paul A. Nakonezny, Robert D. Shull May 2007

Jmasm 26: Hettmansperger And Mckean Linear Model Aligned Rank Test For The Single Covariate And One-Way Ancova Case (Sas), Paul A. Nakonezny, Robert D. Shull

Journal of Modern Applied Statistical Methods

A SAS program (SAS 9.1.3 release, SAS Institute, Cary, N.C.) is presented to implement the Hettmansperger and McKean (1983) linear model aligned rank test (nonparametric ANCOVA) for the single covariate and one-way ANCOVA case. As part of this program, SAS code is also provided to derive the residuals from the regression of Y on X (which is step 1 in the Hettmansperger and McKean procedure) using either ordinary least squares regression (proc reg in SAS) or robust regression with MM estimation (proc robustreg in SAS).


Reliability And Statistical Power: How Measurement Fallibility Affects Power And Required Sample Sizes For Several Parametric And Nonparametric Statistics, Gibbs Y. Kanyongo, Gordon P. Brook, Lydia Kyei-Blankson, Gulsah Gocmen May 2007

Reliability And Statistical Power: How Measurement Fallibility Affects Power And Required Sample Sizes For Several Parametric And Nonparametric Statistics, Gibbs Y. Kanyongo, Gordon P. Brook, Lydia Kyei-Blankson, Gulsah Gocmen

Journal of Modern Applied Statistical Methods

The relationship between reliability and statistical power is considered, and tables that account for reduced reliability are presented. A series of Monte Carlo experiments were conducted to determine the effect of changes in reliability on parametric and nonparametric statistical methods, including the paired samples dependent t test, pooled-variance independent t test, one-way analysis of variance with three levels, Wilcoxon signed-rank test for paired samples, and Mann-Whitney-Wilcoxon test for independent groups. Power tables were created that illustrate the reduction in statistical power from decreased reliability for given sample sizes. Sample size tables were created to provide the approximate sample sizes required …


On Flexible Tests Of Independence And Homoscedasticity, Rand R. Wilcox May 2007

On Flexible Tests Of Independence And Homoscedasticity, Rand R. Wilcox

Journal of Modern Applied Statistical Methods

Consider the nonparametric regression model Y = m(X) + τ(X)ε , where X and ε are independent random variables, ε has a mean of zero and variance σ2, τ is some unknown function used to model heteroscedasticity, and m(X) is an unknown function reflecting some conditional measure of location associated with Y, given X. Detecting dependence, by testing the hypothesis that m(X) does not vary with X, has the potential of being more sensitive to a wider range of associations compared to using Pearson's correlation. This note has two goals. The first is to point …


Using The Fractional Imputation Methodology To Evaluate Variance Due To Hot Deck Imputation In Survey Data, Adriana Pérez May 2007

Using The Fractional Imputation Methodology To Evaluate Variance Due To Hot Deck Imputation In Survey Data, Adriana Pérez

Journal of Modern Applied Statistical Methods

This article examines empirically the effect on the variance estimate due to the use of hot deck imputation with a nearest neighbor donor in comparison with the pairwise fractional hot deck imputation methodology in the 1999 Survey of Doctorate Recipients.


A Comparison Of Eight Shrinkage Formulas Under Extreme Conditions, David A. Walker May 2007

A Comparison Of Eight Shrinkage Formulas Under Extreme Conditions, David A. Walker

Journal of Modern Applied Statistical Methods

The performance of various shrinkage formulas for estimating the population squared multiple correlation coefficient (ρ2) were compared under extreme conditions often found in educational research with small sample sizes of 10, 15, 20, 25, 30 and regressor variates ranging from 2 to 4. A new formula for estimating ρ2, Adj R2 DW, was examined in terms of its performance under various conditions of N, p, ρ2, along with its bias properties and standard error estimates. The two shrinkage formulas that performed most consistently were the Claudy (Adj R2 C) and Walker (Adj R2 DW)


Examining Cronbach Alpha, Theta, Omega Reliability Coefficients According To Sample Size, Ilker Ercan, Berna Yazici, Deniz Sigirli, Bulent Ediz, Ismet Kan May 2007

Examining Cronbach Alpha, Theta, Omega Reliability Coefficients According To Sample Size, Ilker Ercan, Berna Yazici, Deniz Sigirli, Bulent Ediz, Ismet Kan

Journal of Modern Applied Statistical Methods

Differentiations according to the sample size of different reliability coefficients are examined. It is concluded that the estimates obtained by Cronbach alpha and teta coefficients are not related with the sample size, even the estimates obtained from the small samples can represent the population parameter. However, the Omega coefficient requires large sample sizes.


Analyses Of Unbalanced Groups-Versus-Individual Research Designs Using Three Alternative Approximate Degrees Of Freedom Tests: Test Development And Type I Error Rates, Stephanie Wehry, James Algina May 2007

Analyses Of Unbalanced Groups-Versus-Individual Research Designs Using Three Alternative Approximate Degrees Of Freedom Tests: Test Development And Type I Error Rates, Stephanie Wehry, James Algina

Journal of Modern Applied Statistical Methods

Three approximate degrees of freedom quasi-F tests of treatment effectiveness were developed for use in research designs when one treatment is individually delivered and the other is delivered to individuals nested in groups of unequal size. Imbalance in the data was studied from the prospective of subject attrition. The results indicated the test that best controls the Type I error rate depends on the number of groups in the group-administered treatment but does not depend on the subject attrition rates included in the study.


Another Look At The Confidence Intervals For The Noncentral T Distribution, Bruno Lecoutre May 2007

Another Look At The Confidence Intervals For The Noncentral T Distribution, Bruno Lecoutre

Journal of Modern Applied Statistical Methods

An alternative approach to the computation of confidence intervals for the noncentrality parameter of the Noncentral t distribution is proposed. It involves the percent points of a statistical distribution. This conceptual improvement renders the technical process for deriving the limits more comprehensible. Accurate approximations can be derived and easily used.


Better Binomial Confidence Intervals, James F. Reed Iii May 2007

Better Binomial Confidence Intervals, James F. Reed Iii

Journal of Modern Applied Statistical Methods

The construction of a confidence interval for a binomial parameter is a basic analysis in statistical inference. Most introductory statistics textbook authors present the binomial confidence interval based on the asymptotic normality of the sample proportion and estimating the standard error - the Wald method. For the one sample binomial confidence interval the Clopper-Pearson exact method has been regarded as definitive as it eliminates both overshoot and zero width intervals. The Clopper-Pearson exact method is the most conservative and is unquestionably a better alternative to the Wald method. Other viable alternatives include Wilson's Score, the Agresti-Coull method, and the Borkowf …


A Spline-Based Lack-Of-Fit Test For Independent Variable Effect, Chin-Shang Li, Wanzhu Tu May 2007

A Spline-Based Lack-Of-Fit Test For Independent Variable Effect, Chin-Shang Li, Wanzhu Tu

Journal of Modern Applied Statistical Methods

In regression analysis of count data, independent variables are often modeled by their linear effects under the assumption of log-linearity. In reality, the validity of such an assumption is rarely tested, and its use is at times unjustifiable. A lack-of-fit test is proposed for the adequacy of a postulated functional form of an independent variable within the framework of semiparametric Poisson regression models based on penalized splines. It offers added flexibility in accommodating the potentially non-loglinear effect of the independent variable. A likelihood ratio test is constructed for the adequacy of the postulated parametric form, for example log-linearity, of the …


A Comparison Of One-High-Threshold And Two-High-Threshold Multinomial Models Of Source Monitoring, Mahesh Menon, Todd S. Woodward May 2007

A Comparison Of One-High-Threshold And Two-High-Threshold Multinomial Models Of Source Monitoring, Mahesh Menon, Todd S. Woodward

Journal of Modern Applied Statistical Methods

A data simulation study comparing the one-high-threshold (1HT) and two-high-threshold (2HT) multinomial models suggested that 2HT models are more likely to misestimate the underlying parameter values, due to inflation of some parameters (b and d), and deflation of others (D).


Multinomial Logistic Regression Model For The Inferential Risk Age Groups For Infection Caused By Vibrio Cholerae In Kolkata, India, Krishnan Rajendran, Thandavarayan Ramamurthy, Dipika Sur May 2007

Multinomial Logistic Regression Model For The Inferential Risk Age Groups For Infection Caused By Vibrio Cholerae In Kolkata, India, Krishnan Rajendran, Thandavarayan Ramamurthy, Dipika Sur

Journal of Modern Applied Statistical Methods

Multinomial Logistic Regression (MLR) modeling is an effective approach for categorical outcomes, as compared with discriminant function analysis and log-linear models for profiling individual category of dependent variable. To explore the yearly change of inferential age groups of acute diarrhoeal patients infected with Vibrio cholerae during 1996-2000 by MLR, systematic sampling data were generated from an active surveillance study. Among 1330 V.cholerae infected cases, the predominant age category was up to 5 years accounting for 478 (30.5%) cases. The independent variables V.cholerae O1 (p<0.001) and non-O1 and non-O139 (p < 0.001) were significantly associated with children under 5 years age group. V.cholerae O139 inferential age group was > 40 years. The infection mediated by V.cholerae O1 had significantly decreasing trend Exp(B) year wise from …