Open Access. Powered by Scholars. Published by Universities.®

Statistical Theory Commons

Open Access. Powered by Scholars. Published by Universities.®

2008

Discipline
Institution
Keyword
Publication
Publication Type

Articles 31 - 60 of 90

Full-Text Articles in Statistical Theory

Evaluating Subject-Level Incremental Values Of New Markers For Risk Classification Rule, Tianxi Cai, Lu Tian, Donald M. Lloyd-Jones, L. J. Wei Oct 2008

Evaluating Subject-Level Incremental Values Of New Markers For Risk Classification Rule, Tianxi Cai, Lu Tian, Donald M. Lloyd-Jones, L. J. Wei

Harvard University Biostatistics Working Paper Series

No abstract provided.


Generalized Multilevel Functional Regression, Ciprian M. Crainiceanu, Ana-Maria Staicu, Chongzhi Di Sep 2008

Generalized Multilevel Functional Regression, Ciprian M. Crainiceanu, Ana-Maria Staicu, Chongzhi Di

Johns Hopkins University, Dept. of Biostatistics Working Papers

We introduce Generalized Multilevel Functional Linear Models (GMFLM), a novel statistical framework motivated by and applied to the Sleep Heart Health Study (SHHS), the largest community cohort study of sleep. The primary goal of SHHS is to study the association between sleep disrupted breathing (SDB) and adverse health effects. An exposure of primary interest is the sleep electroencephalogram (EEG), which was observed for thousands of individuals at two visits, roughly 5 years apart. This unique study design led to the development of models where the outcome, e.g. hypertension, is in an exponential family and the exposure, e.g. sleep EEG, is …


Measurement Error Caused By Spatial Misalignment In Environmental Epidemiology, Alexandros Gryparis, Christopher J. Paciorek, Ariana Zeka, Joel Schwartz, Brent A. Coull Sep 2008

Measurement Error Caused By Spatial Misalignment In Environmental Epidemiology, Alexandros Gryparis, Christopher J. Paciorek, Ariana Zeka, Joel Schwartz, Brent A. Coull

Harvard University Biostatistics Working Paper Series

No abstract provided.


Practical Large-Scale Spatio-Temporal Modeling Of Particulate Matter Concentrations, Christopher J. Paciorek, Jeff D. Yanosky, Robin C. Puett, Francine Laden, Helen H. Suh Sep 2008

Practical Large-Scale Spatio-Temporal Modeling Of Particulate Matter Concentrations, Christopher J. Paciorek, Jeff D. Yanosky, Robin C. Puett, Francine Laden, Helen H. Suh

Harvard University Biostatistics Working Paper Series

The last two decades have seen intense scientific and regulatory interest in the health effects of particulate matter (PM). Influential epidemiological studies that characterize chronic exposure of individuals rely on monitoring data that are sparse in space and time, so they often assign the same exposure to participants in large geographic areas and across time. We estimate monthly PM during 1988-2002 in a large spatial domain for use in studying health effects in the Nurses' Health Study. We develop a conceptually simple spatio-temporal model that uses a rich set of covariates. The model is used to estimate concentrations of PM10 …


Confidence Intervals For Negative Binomial Random Variables Of High Dispersion, David Shilane, Alan E. Hubbard, S N. Evans Aug 2008

Confidence Intervals For Negative Binomial Random Variables Of High Dispersion, David Shilane, Alan E. Hubbard, S N. Evans

U.C. Berkeley Division of Biostatistics Working Paper Series

This paper considers the problem of constructing confidence intervals for the mean of a Negative Binomial random variable based upon sampled data. When the sample size is large, we traditionally rely upon a Normal distribution approximation to construct these intervals. However, we demonstrate that the sample mean of highly dispersed Negative Binomials exhibits a slow convergence to the Normal in distribution as a function of the sample size. As a result, standard techniques (such as the Normal approximation and bootstrap) that construct confidence intervals for the mean will typically be too narrow and significantly undercover in the case of high …


A New Method For Constructing Exact Tests Without Making Any Assumptions, Karl H. Schlag Aug 2008

A New Method For Constructing Exact Tests Without Making Any Assumptions, Karl H. Schlag

COBRA Preprint Series

We present a new method for constructing exact distribution-free tests (and con…fidence intervals) for variables that can generate more than two possible outcomes. This method separates the search for an exact test from the goal to create a non- randomized test. Randomization is used to extend any exact test relating to means of variables with fi…nitely many outcomes to variables with outcomes belonging to a given bounded set. Tests in terms of variance and covariance are reduced to tests relating to means. Randomness is then eliminated in a separate step. This method is used to create con…fidence intervals for the …


Estimating The Difference Of Percentiles From Two Independent Populations., Romual Eloge Tchouta Aug 2008

Estimating The Difference Of Percentiles From Two Independent Populations., Romual Eloge Tchouta

Electronic Theses and Dissertations

We first consider confidence intervals for a normal percentile, an exponential percentile and a uniform percentile. Then we develop confidence intervals for a difference of percentiles from two independent normal populations, two independent exponential populations and two independent uniform populations. In our study, we mainly focus on the maximum likelihood to develop our confidence intervals. The efficiency of this method is examined via coverage rates obtained in a simulation study done with the statistical software R.


Interval Estimation For The Ratio Of Percentiles From Two Independent Populations., Pius Matheka Muindi Aug 2008

Interval Estimation For The Ratio Of Percentiles From Two Independent Populations., Pius Matheka Muindi

Electronic Theses and Dissertations

Percentiles are used everyday in descriptive statistics and data analysis. In real life, many quantities are normally distributed and normal percentiles are often used to describe those quantities. In life sciences, distributions like exponential, uniform, Weibull and many others are used to model rates, claims, pensions etc. The need to compare two or more independent populations can arise in data analysis. The ratio of percentiles is just one of the many ways of comparing populations. This thesis constructs a large sample confidence interval for the ratio of percentiles whose underlying distributions are known. A simulation study is conducted to evaluate …


Trading Bias For Precision: Decision Theory For Intervals And Sets, Kenneth M. Rice, Thomas Lumley, Adam A. Szpiro Aug 2008

Trading Bias For Precision: Decision Theory For Intervals And Sets, Kenneth M. Rice, Thomas Lumley, Adam A. Szpiro

UW Biostatistics Working Paper Series

Interval- and set-valued decisions are an essential part of statistical inference. Despite this, the justification behind them is often unclear, leading in practice to a great deal of confusion about exactly what is being presented. In this paper we review and attempt to unify several competing methods of interval-construction, within a formal decision-theoretic framework. The result is a new emphasis on interval-estimation as a distinct goal, and not as an afterthought to point estimation. We also see that representing intervals as trade-offs between measures of precision and bias unifies many existing approaches -- as well as suggesting interpretable criteria to …


Fdr Controlling Procedure For Multi-Stage Analyses, Catherine Tuglus, Mark J. Van Der Laan Jul 2008

Fdr Controlling Procedure For Multi-Stage Analyses, Catherine Tuglus, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Multiple testing has become an integral component in genomic analyses involving microarray experiments where large number of hypotheses are tested simultaneously. However before applying more computationally intensive methods, it is often desirable to complete an initial truncation of the variable set using a simpler and faster supervised method such as univariate regression. Once such a truncation is completed, multiple testing methods applied to any subsequent analysis no longer control the appropriate Type I error rates. Here we propose a modified marginal Benjamini \& Hochberg step-up FDR controlling procedure for multi-stage analyses (FDR-MSA), which correctly controls Type I error in terms …


Bringing Game Theory To Hypothesis Testing: Establishing Finite Sample Bounds On Inference, Karl H. Schlag Jun 2008

Bringing Game Theory To Hypothesis Testing: Establishing Finite Sample Bounds On Inference, Karl H. Schlag

COBRA Preprint Series

Small sample properties are of fundamental interest when only limited data is available. Exact inference is limited by constraints imposed by specific nonrandomized tests and of course also by lack of more data. These effects can be separated as we propose to evaluate a test by comparing its type II error to the minimal type II error among all tests for the given sample. Game theory is used to establish this minimal type II error, the associated randomized test is characterized as part of a Nash equilibrium of a fictitious game against nature. We use this method to investigate sequential …


Estimation And Testing For The Effect Of A Genetic Pathway On A Disease Outcome Using Logistic Kernel Machine Regression Via Logistic Mixed Models, Dawei Liu, Debashis Ghosh, Xihong Lin Jun 2008

Estimation And Testing For The Effect Of A Genetic Pathway On A Disease Outcome Using Logistic Kernel Machine Regression Via Logistic Mixed Models, Dawei Liu, Debashis Ghosh, Xihong Lin

Harvard University Biostatistics Working Paper Series

No abstract provided.


A Powerful And Flexible Multilocus Association Test For Quantitative Traits, Lydia Coulter Kwee, Dawei Liu, Xihong Lin, Debashis Ghosh, Michael P. Epstein Jun 2008

A Powerful And Flexible Multilocus Association Test For Quantitative Traits, Lydia Coulter Kwee, Dawei Liu, Xihong Lin, Debashis Ghosh, Michael P. Epstein

Harvard University Biostatistics Working Paper Series

No abstract provided.


Nonparametric Regression Using Local Kernel Estimating Equations For Correlated Failure Time Data, Zhangsheng Yu, Xihong Lin Jun 2008

Nonparametric Regression Using Local Kernel Estimating Equations For Correlated Failure Time Data, Zhangsheng Yu, Xihong Lin

Harvard University Biostatistics Working Paper Series

No abstract provided.


A Comparison Of Methods For Estimating The Causal Effect Of A Treatment In Randomized Clinical Trials Subject To Noncompliance, Rod Little, Qi Long, Xihong Lin Jun 2008

A Comparison Of Methods For Estimating The Causal Effect Of A Treatment In Randomized Clinical Trials Subject To Noncompliance, Rod Little, Qi Long, Xihong Lin

Harvard University Biostatistics Working Paper Series

No abstract provided.


Semiparametric Maximum Likelihood Estimation In Normal Transformation Models For Bivariate Survival Data, Yi Li, Ross L. Prentice, Xihong Lin Jun 2008

Semiparametric Maximum Likelihood Estimation In Normal Transformation Models For Bivariate Survival Data, Yi Li, Ross L. Prentice, Xihong Lin

Harvard University Biostatistics Working Paper Series

No abstract provided.


Accounting For Errors From Predicting Exposures In Environmental Epidemiology And Environmental Statistics, Adam A. Szpiro, Lianne Sheppard, Thomas Lumley Jun 2008

Accounting For Errors From Predicting Exposures In Environmental Epidemiology And Environmental Statistics, Adam A. Szpiro, Lianne Sheppard, Thomas Lumley

UW Biostatistics Working Paper Series

PLEASE NOTE THAT AN UPDATED VERSION OF THIS RESEARCH IS AVAILABLE AS WORKING PAPER 350 IN THE UNIVERSITY OF WASHINGTON BIOSTATISTICS WORKING PAPER SERIES (http://www.bepress.com/uwbiostat/paper350).

In environmental epidemiology and related problems in environmental statistics, it is typically not practical to directly measure the exposure for each subject. Environmental monitoring is employed with a statistical model to assign exposures to individuals. The result is a form of exposure misspecification that can result in complicated errors in the health effect estimates if the exposure is naively treated as known. The exposure error is neither “classical” nor “Berkson”, so standard regression calibration methods …


Supervised Distance Matrices: Theory And Applications To Genomics, Katherine S. Pollard, Mark J. Van Der Laan Jun 2008

Supervised Distance Matrices: Theory And Applications To Genomics, Katherine S. Pollard, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

We propose a new approach to studying the relationship between a very high dimensional random variable and an outcome. Our method is based on a novel concept, the supervised distance matrix, which quantifies pairwise similarity between variables based on their association with the outcome. A supervised distance matrix is derived in two stages. The first stage involves a transformation based on a particular model for association. In particular, one might regress the outcome on each variable and then use the residuals or the influence curve from each regression as a data transformation. In the second stage, a choice of distance …


Confidence Intervals For The Population Mean Tailored To Small Sample Sizes, With Applications To Survey Sampling, Michael Rosenblum, Mark J. Van Der Laan Jun 2008

Confidence Intervals For The Population Mean Tailored To Small Sample Sizes, With Applications To Survey Sampling, Michael Rosenblum, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

The validity of standard confidence intervals constructed in survey sampling is based on the central limit theorem. For small sample sizes, the central limit theorem may give a poor approximation, resulting in confidence intervals that are misleading. We discuss this issue and propose methods for constructing confidence intervals for the population mean tailored to small sample sizes.

We present a simple approach for constructing confidence intervals for the population mean based on tail bounds for the sample mean that are correct for all sample sizes. Bernstein's inequality provides one such tail bound. The resulting confidence intervals have guaranteed coverage probability …


Doubly Robust Ecological Inference, Daniel B. Rubin, Mark J. Van Der Laan May 2008

Doubly Robust Ecological Inference, Daniel B. Rubin, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

The ecological inference problem is a famous longstanding puzzle that arises in many disciplines. The usual formulation in epidemiology is that we would like to quantify an exposure-disease association by obtaining disease rates among the exposed and unexposed, but only have access to exposure rates and disease rates for several regions. The problem is generally intractable, but can be attacked under the assumptions of King's (1997) extended technique if we can correctly specify a model for a certain conditional distribution. We introduce a procedure that it is a valid approach if either this original model is correct or if we …


Confidence Intervals Based On Robust Estimators, Meral Cetin, Serpil Aktas May 2008

Confidence Intervals Based On Robust Estimators, Meral Cetin, Serpil Aktas

Journal of Modern Applied Statistical Methods

Classical estimation of confidence intervals based on the sample mean and variance is sensitive to outliers. Robust methods were proposed for reducing the influence of outliers. The Minimum Volume Ellipsoid estimator (MVE), having a high breakdown point, is one of the robust estimators for location and scale parameters. The robust confidence interval for location parameter is constructed based on the MVE, and compared with the proposed robust confidence interval estimation methods. The performance of the robust confidence interval based on MVE is illustrated with a simulation study. The lengths of 100(1-α)% confidence intervals were investigated.


Using Connectionist Models To Evaluate Examinees’ Response Patterns To Achievement Tests, Mark J. Gierl, Ying Cui, Steve Hunka May 2008

Using Connectionist Models To Evaluate Examinees’ Response Patterns To Achievement Tests, Mark J. Gierl, Ying Cui, Steve Hunka

Journal of Modern Applied Statistical Methods

The attribute hierarchy method (AHM) applied to assessment engineering is described. It is a psychometric method for classifying examinees’ test item responses into a set of attribute mastery patterns associated with different components in a cognitive model of task performance. Attribute probabilities, computed using a neural network, can be estimated for each examinee thereby providing specific information about the examinee’s attribute-mastery level. The pattern recognition approach described in this study relies on an explicit cognitive model to produce the expected response patterns. The expected response patterns serve as the input to the neural network. The model also yields the cognitive …


Coverage Performance Of The Non-Central F-Based And Percentile Bootstrap Confidence Intervals For Root Mean Square Standardized Effect Size In One-Way Fixed-Effects Anova, Guili Zhang, James Algina May 2008

Coverage Performance Of The Non-Central F-Based And Percentile Bootstrap Confidence Intervals For Root Mean Square Standardized Effect Size In One-Way Fixed-Effects Anova, Guili Zhang, James Algina

Journal of Modern Applied Statistical Methods

The coverage performance of the confidence intervals (CIs) for the Root Mean Square Standardized Effect Size (RMSSE) was investigated in a balanced, one-way, fixed-effects, between-subjects ANOVA design. The noncentral F distribution-based and the percentile bootstrap CI construction methods were compared. The results indicated that the coverage probabilities of the CIs for RMSSE were not adequate.


An Evaluation Of Standard, Alternative, And Robust Slope Test Strategies, Tim Moses, Alan Klockars May 2008

An Evaluation Of Standard, Alternative, And Robust Slope Test Strategies, Tim Moses, Alan Klockars

Journal of Modern Applied Statistical Methods

The robustness and power of nine strategies for testing the differences between two groups’ regression slopes under nonnormality and residual variance heterogeneity are compared. The results showed that three most robust slope test strategies were the combination of the trimmed and Winsorized slopes with the James second order test, the combination of Theil-Sen with James, and Theil-Sen with percentile bootstrapping. The slope tests based on Theil-Sen slopes were more powerful than those based on trimmed and Winsorized slopes.


Second-Order Latent Growth Models With Shifting Indicators, Gregory R. Hancock, Michelle M. Buehl May 2008

Second-Order Latent Growth Models With Shifting Indicators, Gregory R. Hancock, Michelle M. Buehl

Journal of Modern Applied Statistical Methods

Second-order latent growth models assess longitudinal change in a latent construct, typically employing identical manifest variables as indicators across time. However, the same indicators may be unavailable and/or inappropriate for all time points. This article details methods for second-order growth models in which constructs’ indicators shift over time.


Selection Of Non-Regular Fractional Factorial Designs When Some Two-Factor Interactions Are Important, Weiming Ke, Rui Yao May 2008

Selection Of Non-Regular Fractional Factorial Designs When Some Two-Factor Interactions Are Important, Weiming Ke, Rui Yao

Journal of Modern Applied Statistical Methods

A new method is proposed for selecting the optimal non-regular fractional factorial designs in the situation when some two-factor interactions are potentially important. Searching for the best designs according to this method is discussed and some results for the Plackett-Burman design of 12 runs are presented.


A Weighted Moving Average Process For Forecasting, Shou Hsing Shih, Chris P. Tsokos May 2008

A Weighted Moving Average Process For Forecasting, Shou Hsing Shih, Chris P. Tsokos

Journal of Modern Applied Statistical Methods

The object of the present study is to propose a forecasting model for a nonstationary stochastic realization. The subject model is based on modifying a given time series into a new k-time moving average time series to begin the development of the model. The study is based on the autoregressive integrated moving average process along with its analytical constrains. The analytical procedure of the proposed model is given. A stock XYZ selected from the Fortune 500 list of companies and its daily closing price constitute the time series. Both the classical and proposed forecasting models were developed and a comparison …


Comparing Different Methods For Multiple Testing In Reaction Time Data, Massimiliano Pastore, Massimo Nucci, Giovanni Galfano May 2008

Comparing Different Methods For Multiple Testing In Reaction Time Data, Massimiliano Pastore, Massimo Nucci, Giovanni Galfano

Journal of Modern Applied Statistical Methods

Reaction times were simulated for examining the power of six methods for multiple testing, as a function of sample size and departures from normality. Power estimates were low for all methods for non-normal distributions. With normal distributions, even for small sample sizes, satisfactory power estimates were observed, especially for FDR-based procedures.


Two-Stage Short-Run (X, Mr) Control Charts, Matthew E. Elam, Kenneth E. Case May 2008

Two-Stage Short-Run (X, Mr) Control Charts, Matthew E. Elam, Kenneth E. Case

Journal of Modern Applied Statistical Methods

This article is the first in a series of two articles that applies two-stage short-run control charting to (X, MR) charts. Theory is developed and then used to derive the control chart factor equations. In the sequel, the control chart factor calculations are computerized and an example is presented.


Confidence Intervals For The Squared Multiple Semipartial Correlation Coefficient, James Algina, H. J. Keselman, Randall D. Penfield May 2008

Confidence Intervals For The Squared Multiple Semipartial Correlation Coefficient, James Algina, H. J. Keselman, Randall D. Penfield

Journal of Modern Applied Statistical Methods

The squared multiple semipartial correlation coefficient is the increase in the squared multiple correlation coefficient that occurs when two or more predictors are added to a multiple regression model. Coverage probability was investigated for two variations of each of three methods for setting confidence intervals for the population squared multiple semipartial correlation coefficient. Results indicated that the procedure that provides coverage probability in the [.925, .975] interval for a 95% confidence interval depends primarily on the number of added predictors. Guidelines for selecting a procedure are presented.