Open Access. Powered by Scholars. Published by Universities.®

Statistical Theory Commons™

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type

Articles 931 - 960 of 1633

Full-Text Articles in Statistical Theory

Readings In Targeted Maximum Likelihood Estimation, Mark J. Van Der Laan, Sherri Rose, Susan Gruber Sep 2009

Readings In Targeted Maximum Likelihood Estimation, Mark J. Van Der Laan, Sherri Rose, Susan Gruber

U.C. Berkeley Division of Biostatistics Working Paper Series

This is a compilation of current and past work on targeted maximum likelihood estimation. It features the original targeted maximum likelihood learning paper as well as chapters on super (machine) learning using cross validation, randomized controlled trials, realistic individualized treatment rules in observational studies, biomarker discovery, case-control studies, and time-to-event outcomes with censored data, among others. We hope this collection is helpful to the interested reader and stimulates additional research in this important area.


Causal Inference For Nested Case-Control Studies Using Targeted Maximum Likelihood Estimation, Sherri Rose, Mark J. Van Der Laan Sep 2009

Causal Inference For Nested Case-Control Studies Using Targeted Maximum Likelihood Estimation, Sherri Rose, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

A nested case-control study is conducted within a well-defined cohort arising out of a population of interest. This design is often used in epidemiology to reduce the costs associated with collecting data on the full cohort; however, the case control sample within the cohort is a biased sample. Methods for analyzing case-control studies have largely focused on logistic regression models that provide conditional and not marginal causal estimates of the odds ratio. We previously developed a Case-Control Weighted Targeted Maximum Likelihood Estimation (TMLE) procedure for case-control study designs, which relies on the prevalence probability q0. We propose the use of …


Targeted Maximum Likelihood Estimation: A Gentle Introduction, Susan Gruber, Mark J. Van Der Laan Aug 2009

Targeted Maximum Likelihood Estimation: A Gentle Introduction, Susan Gruber, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

This paper provides a concise introduction to targeted maximum likelihood estimation (TMLE) of causal effect parameters. The interested analyst should gain sufficient understanding of TMLE from this introductory tutorial to be able to apply the method in practice. A program written in R is provided. This program implements a basic version of TMLE that can be used to estimate the effect of a binary point treatment on a continuous or binary outcome.


Comparing Risk Scoring Systems Beyond The Roc Paradigm In Survival Analysis, Hajime Uno, Lu Tian, Tianxi Cai, Isaac S. Kohane, L. J. Wei Aug 2009

Comparing Risk Scoring Systems Beyond The Roc Paradigm In Survival Analysis, Hajime Uno, Lu Tian, Tianxi Cai, Isaac S. Kohane, L. J. Wei

Harvard University Biostatistics Working Paper Series

No abstract provided.


Combinational Mixtures Of Multiparameter Distributions, Valeria Edefonti, Giovanni Parmigiani Aug 2009

Combinational Mixtures Of Multiparameter Distributions, Valeria Edefonti, Giovanni Parmigiani

Johns Hopkins University, Dept. of Biostatistics Working Papers

We introduce combinatorial mixtures - a flexible class of models for inference on mixture distributions whose component have multidimensional parameters. The key idea is to allow each element of the component-specific parameter vectors to be shared by a subset of other components. This approach allows for mixtures that range from very flexible to very parsimonious, and unifies inference on component-specific parameters with inference on the number of components. We develop Bayesian inference and computation approaches for this class of distributions, and illustrate them in an application. This work was originally motivated by the analysis of cancer subtypes: in terms of …


Shrinkage Estimation Of Expression Fold Change As An Alternative To Testing Hypotheses Of Equivalent Expression, Zahra Montazeri, Corey M. Yanofsky, David R. Bickel Aug 2009

Shrinkage Estimation Of Expression Fold Change As An Alternative To Testing Hypotheses Of Equivalent Expression, Zahra Montazeri, Corey M. Yanofsky, David R. Bickel

COBRA Preprint Series

Research on analyzing microarray data has focused on the problem of identifying differentially expressed genes to the neglect of the problem of how to integrate evidence that a gene is differentially expressed with information on the extent of its differential expression. Consequently, researchers currently prioritize genes for further study either on the basis of volcano plots or, more commonly, according to simple estimates of the fold change after filtering the genes with an arbitrary statistical significance threshold. While the subjective and informal nature of the former practice precludes quantification of its reliability, the latter practice is equivalent to using a …


The Effect Of Correlation In False Discovery Rate Estimation, Armin Schwartzman, Xihong Lin Jul 2009

The Effect Of Correlation In False Discovery Rate Estimation, Armin Schwartzman, Xihong Lin

Harvard University Biostatistics Working Paper Series

No abstract provided.


Spatial Cluster Detection For Repeatedly Measured Outcomes While Accounting For Residential History, Andrea J. Cook, Diane Gold, Yi Li Jun 2009

Spatial Cluster Detection For Repeatedly Measured Outcomes While Accounting For Residential History, Andrea J. Cook, Diane Gold, Yi Li

Harvard University Biostatistics Working Paper Series

No abstract provided.


Marginalized Frailty Models For Multivariate Survival Data, Megan Othus, Yi Li Jun 2009

Marginalized Frailty Models For Multivariate Survival Data, Megan Othus, Yi Li

Harvard University Biostatistics Working Paper Series

No abstract provided.


Spatial Cluster Detection For Weighted Outcomes Using Cumulative Geographic Residuals, Andrea J. Cook, Yi Li, David Arterburn, Ram C. Tiwari Jun 2009

Spatial Cluster Detection For Weighted Outcomes Using Cumulative Geographic Residuals, Andrea J. Cook, Yi Li, David Arterburn, Ram C. Tiwari

Harvard University Biostatistics Working Paper Series

No abstract provided.


On The C-Statistics For Evaluating Overall Adequacy Of Risk Prediction Procedures With Censored Survival Data, Hajime Uno, Tianxi Cai, Michael J. Pencina, Ralph B. D'Agostino, L. J. Wei Jun 2009

On The C-Statistics For Evaluating Overall Adequacy Of Risk Prediction Procedures With Censored Survival Data, Hajime Uno, Tianxi Cai, Michael J. Pencina, Ralph B. D'Agostino, L. J. Wei

Harvard University Biostatistics Working Paper Series

No abstract provided.


Estimating Subject-Specific Dependent Competing Risk Profile With Censored Event Time Observations, Yi Li, Lu Tian, L. J. Wei May 2009

Estimating Subject-Specific Dependent Competing Risk Profile With Censored Event Time Observations, Yi Li, Lu Tian, L. J. Wei

Harvard University Biostatistics Working Paper Series

No abstract provided.


On The Blue Of The Population Mean For Location And Scale Parameters Of Distributions Based On Moving Extreme Ranked Set Sampling, Walid A. Abu-Dayyeh, Lana Al-Rousan May 2009

On The Blue Of The Population Mean For Location And Scale Parameters Of Distributions Based On Moving Extreme Ranked Set Sampling, Walid A. Abu-Dayyeh, Lana Al-Rousan

Journal of Modern Applied Statistical Methods

The best linear unbiased estimator (BLUE) for the population mean under moving extreme ranked set sampling (MERSS) is derived for general location and scale parameters of distributions which generalizes Al-Odat and Al-Saleh (2001). It is compared with the sample mean of simple random sampling (SRS). The efficient sample size under the MERSS for which the BLUE estimator dominates the usual sample mean under SRS for estimating the population mean is also computed for several distributions.


Robustness To Non-Independence And Power Of The I Test For Trend In Construct Validity, John L. Cuzzocrea, Shlomo Sawilowsky May 2009

Robustness To Non-Independence And Power Of The I Test For Trend In Construct Validity, John L. Cuzzocrea, Shlomo Sawilowsky

Journal of Modern Applied Statistical Methods

The Multitrait-Multimethod Matrix is used to evaluate construct validity; Sawilowsky (2002) created the I test to analyze the matrix. This article examined the robustness and power of the Sawilowsky I test. Ad hoc critical values were determined to improve the statistical power of the technique for analyzing the Multitrait-Multimethod Matrix.


Least Absolute Value Vs. Least Squares Estimation And Inference Procedures In Regression Models With Asymmetric Error Distributions, Terry E. Dielman May 2009

Least Absolute Value Vs. Least Squares Estimation And Inference Procedures In Regression Models With Asymmetric Error Distributions, Terry E. Dielman

Journal of Modern Applied Statistical Methods

A Monte Carlo simulation is used to compare estimation and inference procedures in least absolute value (LAV) and least squares (LS) regression models with asymmetric error distributions. Mean square errors (MSE) of coefficient estimates are used to assess the relative efficiency of the estimators. Hypothesis tests for coefficients are compared on the basis of empirical level of significance and power.


Covariate-Adjusted Constrained Bayes Predictions Of Random Intercepts And Slopes. Sujit Ghosh Is A, Robert H. Lyles, Reneé H. Moore, Amita K. Manatunga, Kirk A. Easley May 2009

Covariate-Adjusted Constrained Bayes Predictions Of Random Intercepts And Slopes. Sujit Ghosh Is A, Robert H. Lyles, Reneé H. Moore, Amita K. Manatunga, Kirk A. Easley

Journal of Modern Applied Statistical Methods

No abstract provided.


On The Expected Values Of Distribution Of The Sample Range Of Order Statistics From The Geometric Distribution, Sinan Calik, Cemil Colak, Ayse Turan May 2009

On The Expected Values Of Distribution Of The Sample Range Of Order Statistics From The Geometric Distribution, Sinan Calik, Cemil Colak, Ayse Turan

Journal of Modern Applied Statistical Methods

The expected values of the distribution of the sample range of order statistics from the geometric distribution are presented. For n up to 10, algebraic expressions for the expected values are obtained. Using the algebraic expressions, expected values based on the p and n values can be easily computed.


Approximations To Power When Comparing Two Small Independent Proportions, Michael Vorburger, Breda Munoz May 2009

Approximations To Power When Comparing Two Small Independent Proportions, Michael Vorburger, Breda Munoz

Journal of Modern Applied Statistical Methods

No abstract provided.


The Bootstrap Method For The Selection Of A Shrinkage Factor In Two-Stage Estimation Of The Reliability Function Of An Exponential Distribution, Makarand V. Ratnaparkhi, Vasant B. Waikar, Fredrick J. Schuurmann May 2009

The Bootstrap Method For The Selection Of A Shrinkage Factor In Two-Stage Estimation Of The Reliability Function Of An Exponential Distribution, Makarand V. Ratnaparkhi, Vasant B. Waikar, Fredrick J. Schuurmann

Journal of Modern Applied Statistical Methods

An application of a bootstrap method for selecting a suitable shrinkage factor for the two-stage shrinkage estimator of a reliability function for the exponential distribution is discussed. The estimator obtained here has higher efficiency as compared to the one where the shrinkage factor is not subjected to bootstrapping.


Effects Of Population Distribution, Sample Size And Correlation Structure On Huberty’S Effect Size R, James B. Hittner May 2009

Effects Of Population Distribution, Sample Size And Correlation Structure On Huberty’S Effect Size R, James B. Hittner

Journal of Modern Applied Statistical Methods

Huberty’s (1994) R2 is derived by subtracting the expected value of R2 from an adjusted R2, and the square root of Huberty’s R2 is Huberty’s effect size R. The present study examined the effects of population distribution, sample size and population correlation structure on the statistical power of Huberty’s R.


Estimating Task Duration In Pert Using The Weibull Probability Distribution, Edward L. Mccombs, Matthew E. Elam, David B. Pratt May 2009

Estimating Task Duration In Pert Using The Weibull Probability Distribution, Edward L. Mccombs, Matthew E. Elam, David B. Pratt

Journal of Modern Applied Statistical Methods

The Weibull probability distribution can be used as an alternative model for task time estimates in the PERT estimating methodology. It has the same advantages as the traditional beta distribution for this application. It has additional benefits, however, that make it a preferred option.


Beyond Kappa: Estimating Inter-Rater Agreement With Nominal Classifications, Nol Bendermacher, Pierre Souren May 2009

Beyond Kappa: Estimating Inter-Rater Agreement With Nominal Classifications, Nol Bendermacher, Pierre Souren

Journal of Modern Applied Statistical Methods

Cohen’s Kappa and a number of related measures can all be criticized for their definition of correction for chance agreement. A measure is introduced that derives the corrected proportion of agreement directly from the data, thereby overcoming objections to Kappa and its related measures.


Some Estimators For The Population Mean Using Auxiliary Information Under Ranked Set Sampling, Walid A. Abu-Dayyeh, M. S. Ahmed, R. A. Ahmed, Hassen A. Muttlak May 2009

Some Estimators For The Population Mean Using Auxiliary Information Under Ranked Set Sampling, Walid A. Abu-Dayyeh, M. S. Ahmed, R. A. Ahmed, Hassen A. Muttlak

Journal of Modern Applied Statistical Methods

Auxiliary information is used along with ranking information to derive several classes of estimators to estimate the population mean of a variable of interest based on RSS (ranked set sample). The properties of these newly suggested estimators were examined. Comparisons between special cases of these estimators and other known estimators are made using a real data set. Some of the new estimators are superior to the old ones in terms of bias and mean square error.


A Heteroscedastic, Rank-Based Approach For Analyzing 2 X 2 Independent Groups Designs, Laura Mills, Robert A. Cribbie, Wei-Ming Luh May 2009

A Heteroscedastic, Rank-Based Approach For Analyzing 2 X 2 Independent Groups Designs, Laura Mills, Robert A. Cribbie, Wei-Ming Luh

Journal of Modern Applied Statistical Methods

The ANOVA F is a widely used statistic in psychological research despite its shortcomings when the assumptions of normality and variance heterogeneity are violated. A Monte Carlo investigation compared Type I error and power rates of the ANOVA F, Alexander-Govern with trimmed means and Johnson transformation, Welch-James with trimmed means and Johnson Transformation, Welch with trimmed means, and Welch on ranked data using Johansen’s interaction procedure. Results suggest that the ANOVA F is not appropriate when assumptions of normality and variance homogeneity are violated, and that the Welch/Johansen on ranks offers the best balance of empirical Type I error …


Efficiency Of Canonical Discriminant Function Versus Mahalanobis Distance In Differentiating Groups: Screening Ovarian Cancer In A Multivariate System Analysis Using Enzyme Markers, Chinmoy K. Bose May 2009

Efficiency Of Canonical Discriminant Function Versus Mahalanobis Distance In Differentiating Groups: Screening Ovarian Cancer In A Multivariate System Analysis Using Enzyme Markers, Chinmoy K. Bose

Journal of Modern Applied Statistical Methods

Due to its low prevalence, high mortality and uniquely hidden intrapelvic position, ovarian cancer remains a subject of intense interest to researchers. Statistical calculation and new technology both have major roles to play in the effort to screen this cancer at an early stage. Advanced statistics, such as multivariate analysis, remain at the root of screening endeavors. Multivariate analysis has the power to combine many tests and to produce better results in terms high specificity and positive predictive value. Multivariate analysis techniques include Mahalanobis distance (D2), canonical stepwise discriminant function (Z) and Posterior Probability. These may have varied …


Applying Census Data For Small Area Estimation In Community And Social Service Planning, Michael Wolf-Branigin, Hyon-Sook Suh, Star Muir, Emily S. Ihara May 2009

Applying Census Data For Small Area Estimation In Community And Social Service Planning, Michael Wolf-Branigin, Hyon-Sook Suh, Star Muir, Emily S. Ihara

Journal of Modern Applied Statistical Methods

Small area estimation provides a tool for community analysis. A procedure for accessing, selecting, joining and analyzing US Census data is provided. Skills acquired while completing the procedure include accessing census data, downloading boundary files and displaying themes. Such skills are valuable tools for students to possess as they enter the workforce.


A Comparative Study Of Bayesian Model Selection Criteria For Capture-Recapture Models For Closed Populations, Ross M. Gosky, Sujit K. Ghosh May 2009

A Comparative Study Of Bayesian Model Selection Criteria For Capture-Recapture Models For Closed Populations, Ross M. Gosky, Sujit K. Ghosh

Journal of Modern Applied Statistical Methods

Capture-Recapture models estimate unknown population sizes. Eight standard closed population models exist, allowing for time, behavioral, and heterogeneity effects. Bayesian versions of these models are presented and use of Akaike's Information Criterion (AIC) and the Deviance Information Criterion (DIC) are explored as model selection tools, through simulation and real dataset analysis.


Quantile Regression: On Inferences About The Slopes Corresponding To One, Two Or Three Quantiles, Rand R. Wilcox, Kathleen Costa May 2009

Quantile Regression: On Inferences About The Slopes Corresponding To One, Two Or Three Quantiles, Rand R. Wilcox, Kathleen Costa

Journal of Modern Applied Statistical Methods

The problem of testing hypotheses about the slope of a quantile regression line when the sample size is small is considered. A modified bootstrap method is suggested that is found to have certain advantages over the inverse rank method recommended by Koenker (1994). A method is suggested that simultaneously controls the probability of at least one Type I error when performing two or three tests corresponding to two or three specific quantiles. Using data from actual studies, it is illustrated that the new method can yield substantially shorter confidence intervals than the rank inverse method and, even with a large …


A Monte Carlo Comparison Of Regression Estimators When The Error Distribution Is Long-Tailed Symmetric, Oya Can Mutan, Birdal Şenoğlu May 2009

A Monte Carlo Comparison Of Regression Estimators When The Error Distribution Is Long-Tailed Symmetric, Oya Can Mutan, Birdal Şenoğlu

Journal of Modern Applied Statistical Methods

The performances of the ordinary least squares (OLS), modified maximum likelihood (MML), least absolute deviations (LAD), Winsorized least squares (WIN), trimmed least squares (TLS), Theil’s (Theil) and weighted Theil’s (Weighted Theil) estimators are compared under the simple linear regression model in terms of their bias and efficiency when the distribution of error terms is long-tailed symmetric.


Improved Confidence Intervals For The Difference Between Two Proportions, James F. Reed Iii May 2009

Improved Confidence Intervals For The Difference Between Two Proportions, James F. Reed Iii

Journal of Modern Applied Statistical Methods

Wald-z asymptotic methods, with and without a continuity correction, have less than nominal coverage probability characteristics but continue to be used. Newcombe's hybrid method and the Agresti-Caffo methods have coverage probabilities that are near nominal for either equal or unequal samples. Newcombe's hybrid and Agresti-Caffo methods demonstrate superior coverage properties.