Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

2005

Discipline
Institution
Keyword
Publication
Publication Type

Articles 121 - 150 of 279

Full-Text Articles in Statistics and Probability

The Optimal Discovery Procedure: A New Approach To Simultaneous Significance Testing, John D. Storey Sep 2005

The Optimal Discovery Procedure: A New Approach To Simultaneous Significance Testing, John D. Storey

UW Biostatistics Working Paper Series

Significance testing is one of the main objectives of statistics. The Neyman-Pearson lemma provides a simple rule for optimally testing a single hypothesis when the null and alternative distributions are known. This result has played a major role in the development of significance testing strategies that are used in practice. Most of the work extending single testing strategies to multiple tests has focused on formulating and estimating new types of significance measures, such as the false discovery rate. These methods tend to be based on p-values that are calculated from each test individually, ignoring information from the other tests. As …


Mixture Cure Survival Models With Dependent Censoring, Yi Li, Ram C. Tiwari, Subharup Guha Sep 2005

Mixture Cure Survival Models With Dependent Censoring, Yi Li, Ram C. Tiwari, Subharup Guha

Harvard University Biostatistics Working Paper Series

A number of authors have studies the mixture survival model to analyze survival data with nonnegligible cure fractions. A key assumption made by these authors is the independence between the survival time and the censoring time. To our knowledge, no one has studies the mixture cure model in the presence of dependent censoring. To account for such dependence, we propose a more general cure model which allows for dependent censoring. In particular, we derive the cure models from the perspective of competing risks and model the dependence between the censoring time and the survival time using a class of Archimedean …


Semiparametric Normal Transformation Models For Spatially Correlated Survival Data, Yi Li, Xihong Lin Sep 2005

Semiparametric Normal Transformation Models For Spatially Correlated Survival Data, Yi Li, Xihong Lin

Harvard University Biostatistics Working Paper Series

There is an emerging interest in modeling spatially correlated survival data in biomedical and epidemiological studies. In this paper, we propose a new class of semiparametric normal transformation models for right censored spatially correlated survival data. This class of models assumes that survival outcomes marginally follow a Cox proportional hazard model with unspecified baseline hazard, and their joint distribution is obtained by transforming survival outcomes to normal random variables, whose joint distribution is assumed to be multivariate normal with a spatial correlation structure. A key feature of the class of semiparametric normal transformation models is that it provides a rich …


Inference On Survival Data With Covariate Measurement Error - An Imputation-Based Approach, Yi Li, Louise Ryan Sep 2005

Inference On Survival Data With Covariate Measurement Error - An Imputation-Based Approach, Yi Li, Louise Ryan

Harvard University Biostatistics Working Paper Series

We propose a new method for fitting proportional hazards models with error-prone covariates. Regression coefficients are estimated by solving an estimating equation that is the average of the partial likelihood scores based on imputed true covariates. For the purpose of imputation, a linear spline model is assumed on the baseline hazard. We discuss consistency and asymptotic normality of the resulting estimators, and propose a stochastic approximation scheme to obtain the estimates. The algorithm is easy to implement, and reduces to the ordinary Cox partial likelihood approach when the measurement error has a degenerative distribution. Simulations indicate high efficiency and robustness. …


The Optimal Discovery Procedure For Large-Scale Significance Testing, With Applications To Comparative Microarray Experiments, John D. Storey, James Y. Dai, Jeffrey T. Leek Sep 2005

The Optimal Discovery Procedure For Large-Scale Significance Testing, With Applications To Comparative Microarray Experiments, John D. Storey, James Y. Dai, Jeffrey T. Leek

UW Biostatistics Working Paper Series

As much of the focus of genetics and molecular biology has shifted toward the systems level, it has become increasingly important to accurately extract biologically relevant signal from thousands of related measurements. The common property among these high-dimensional biological studies is that the measured features have a rich and largely unknown underlying structure. One example of much recent interest is identifying differentially expressed genes in comparative microarray experiments. We propose a new approach aimed at optimally performing many hypothesis tests in a high-dimensional study. This approach estimates the Optimal Discovery Procedure (ODP), which has recently been introduced and theoretically shown …


Comparison Of Affymetrix Genechip Expression Measures, Rafael A. Irizarry, Zhijin Wu, Harris A. Jaffee Sep 2005

Comparison Of Affymetrix Genechip Expression Measures, Rafael A. Irizarry, Zhijin Wu, Harris A. Jaffee

Johns Hopkins University, Dept. of Biostatistics Working Papers

Affymetrix GeneChip expression array technology has become a standard tool in medical science and basic biology research. In this system, preprocessing occurs before one obtains expression level measurements. Because the number of competing preprocessing methods was large and growing, in the summer of 2003 we developed a benchmark to help users of the technology identify the best method for their application. In conjunction with the release of a Bioconductor R package (affycomp), a webtool was made available for developers of preprocessing methods to submit them to a benchmark for comparison. There have now been over 30 methods compared via the …


The Outcome Of Mta As A Root End Filling Material: A Long Term Evaluation, Christopher M. Sechrist Sep 2005

The Outcome Of Mta As A Root End Filling Material: A Long Term Evaluation, Christopher M. Sechrist

Loma Linda University Electronic Theses, Dissertations & Projects

Periradicular surgery is a viable option to save natural teeth when non-surgical treatment fails or when endodontic retreatment is not feasible or contraindicated. Laboratory and animal studies have demonstrated that MTA is biocompatible, provides an excellent seal against penetrating bacteria, and promotes hard tissue healing. The purpose of this study was to provide long term (>3 years) clinical evidence for its use as a root-end filling material in endodontics. The clinical records of 294 patients who had MTA used during endodontic treatment from 1996 to 2001 were reviewed. From these, 75 patients whose root end cavities had been filled …


Laser And Led Effects On The Proliferation Rate Of Periodontal Ligament Fibroblasts, Allen J. Job Sep 2005

Laser And Led Effects On The Proliferation Rate Of Periodontal Ligament Fibroblasts, Allen J. Job

Loma Linda University Electronic Theses, Dissertations & Projects

PURPOSE: To compare the effectiveness of a Gallium Aluminum Arsenide (GaAlAs) diode laser and a light emitting diode (LED) on periodontal ligament fibroblast cell proliferative rates.

METHODS and MATERIALS: PDLF obtained from freshly extracted permanent teeth were cultured under standard conditions until a subconfluent monolayer was present. The next section took 5 days to complete. On day 1, the initial cell concentration of 700 uL/cm2 was plated on 96-well assay plates and placed in a CO2 incubator at 37° C for 24 hours. On day 2, cell counts were first verified using hemocytometry then were irradiated using an …


Sample Size And Power Calculations For Body Weight In Beef Cattle, Claudia Cristina Paro Paz, Alfredo Ribeiro De Freitas, Irineu Umberto Packer, Daniela Tambasco-Talhari, Luciana Correa De Almeida Regitano, Mauricio Mello Alencar Aug 2005

Sample Size And Power Calculations For Body Weight In Beef Cattle, Claudia Cristina Paro Paz, Alfredo Ribeiro De Freitas, Irineu Umberto Packer, Daniela Tambasco-Talhari, Luciana Correa De Almeida Regitano, Mauricio Mello Alencar

COBRA Preprint Series

Estimates of minimum sample sizes are calculated in order to test differences in rates of changes over time for longitudinal designs. In this study, body weight of crossbred beef cattle, considering 14 measurements on individuals, taken at birth, weaning (7 months of age) and monthly from 8 to 19 months of age, were analyzed by an usual mixed model for repeated measures. The number of individuals n required to detect significant differences (delta) between any two consecutive measurements on the individual, was obtained by a SAS program considering a t-variate normal distribution (t = 14), sample variance–covariance matrix among the …


A Survey On Intrusion Detection Approaches, A Murali M. Rao Aug 2005

A Survey On Intrusion Detection Approaches, A Murali M. Rao

International Conference on Information and Communication Technologies

Intrusion detection plays one of the key roles in computer security techniques and is one of the prime areas of research. Usages of computer network services are tremendously increasing day by day and at the same time intruders are also playing a major role to deny network services, compromising the crucial services for Email, FTP and Web. Realizing the importance of the problem due to intrusions, many researchers have taken up research in this area and have proposed several solutions. It has come to a stage to take a stock of the research results and project a comprehensive view so …


Direct Effect Models, Mark J. Van Der Laan, Maya L. Petersen Aug 2005

Direct Effect Models, Mark J. Van Der Laan, Maya L. Petersen

U.C. Berkeley Division of Biostatistics Working Paper Series

The causal effect of a treatment on an outcome is generally mediated by several intermediate variables. Estimation of the component of the causal effect of a treatment that is mediated by a given intermediate variable (the indirect effect of the treatment), and the component that is not mediated by that intermediate variable (the direct effect of the treatment) is often relevant to mechanistic understanding and to the design of clinical and public health interventions. Under the assumption of no-unmeasured confounders for treatment and the intermediate variable, Robins & Greenland (1992) define an individual direct effect as the counterfactual effect of …


Statistical Inference For Variable Importance, Mark J. Van Der Laan Aug 2005

Statistical Inference For Variable Importance, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Many statistical problems involve the learning of an importance/effect of a variable for predicting an outcome of interest based on observing a sample of n independent and identically distributed observations on a list of input variables and an outcome. For example, though prediction/machine learning is, in principle, concerned with learning the optimal unknown mapping from input variables to an outcome from the data, the typical reported output is a list of importance measures for each input variable. The typical approach in prediction has been to learn the unknown optimal predictor from the data and derive, for each of the input …


Computing The Total Sample Size When Group Sizes Are Not Fixed, Mithat Gonen Aug 2005

Computing The Total Sample Size When Group Sizes Are Not Fixed, Mithat Gonen

Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series

This article is concerned with computing the total sample size required for a two-sample comparison when the sizes of the two groups to be compared cannot be fixed in advance. This is frequently encountered when group membership depends on a variable which is observable only after the subject is enrolled to the study, such as a genetic or a biological marker. The most common way of circumventing this problem is assuming a fixed number for the prevalence of the condition that will determine the group membership and compute the required sample size conditionally. In this article this practice is formalized …


Survival Point Estimate Prediction In Matched And Non-Matched Case-Control Subsample Designed Studies, Annette M. Molinaro, Mark J. Van Der Laan, Dan H. Moore, Karla Kerlikowske Aug 2005

Survival Point Estimate Prediction In Matched And Non-Matched Case-Control Subsample Designed Studies, Annette M. Molinaro, Mark J. Van Der Laan, Dan H. Moore, Karla Kerlikowske

U.C. Berkeley Division of Biostatistics Working Paper Series

Providing information about the risk of disease and clinical factors that may increase or decrease a patient's risk of disease is standard medical practice. Although case-control studies can provide evidence of strong associations between diseases and risk factors, clinicians need to be able to communicate to patients the age-specific risks of disease over a defined time interval for a set of risk factors.

An estimate of absolute risk cannot be determined from case-control studies because cases are generally chosen from a population whose size is not known (necessary for calculation of absolute risk) and where duration of follow-up is not …


A User-Friendly Introduction To Link-Probit-Normal Models, Brian S. Caffo, Michael Griswold Aug 2005

A User-Friendly Introduction To Link-Probit-Normal Models, Brian S. Caffo, Michael Griswold

Johns Hopkins University, Dept. of Biostatistics Working Papers

Probit-normal models have attractive properties compared to logit-normal models. In particular, they allow for easy specification of marginal links of interest while permitting a conditional random effects structure. Moreover, programming fitting algorithms for probit-normal models can be trivial with the use of well-developed algorithms for approximating multivariate normal quantiles. In typical settings, the data cannot distinguish between probit and logit conditional link functions. Therefore, if marginal interpretations are desired, the default conditional link should be the most convenient one. We refer to models with a probit conditional link an arbitrary marginal link and a normal random effect distribution as link-probit-normal …


The Interquartile Range: Theory And Estimation., Dewey Lonzo Whaley Aug 2005

The Interquartile Range: Theory And Estimation., Dewey Lonzo Whaley

Electronic Theses and Dissertations

The interquartile range (IQR) is used to describe the spread of a distribution. In an introductory statistics course, the IQR might be introduced as simply the “range within which the middle half of the data points lie.” In other words, it is the distance between the two quartiles, IQR = Q3 - Q1. We will compute the population IQR, the expected value, and the variance of the sample IQR for various continuous distributions. In addition, a bootstrap confidence interval for the population IQR will be evaluated.


Using Box-Scores To Determine A Position's Contribution To Winning Basketball Games, Garritt L. Page Aug 2005

Using Box-Scores To Determine A Position's Contribution To Winning Basketball Games, Garritt L. Page

Theses and Dissertations

Basketball is a sport that has become increasingly popular world-wide. At the professional level it is a game in which each of the five positions has a specific responsibility that requires unique skills. It seems likely that it would be valuable for coaches to know which skills for each position are most conducive to winning. Knowing which skills to develop for each position could help coaches optimize each player's ability by customizing practice to contain drills that develop the most important skills for each position that would in turn improve the team's overall ability. Through the use of Bayesian hierarchical …


Semiparametric Inferences For Association With Semi-Competing Risks Data, Debashis Ghosh Aug 2005

Semiparametric Inferences For Association With Semi-Competing Risks Data, Debashis Ghosh

The University of Michigan Department of Biostatistics Working Paper Series

In many biomedical studies, it is of interest to assess dependence between bivariate failure time data. We focus here on a special type of such data, referred to as semi-competing risks data. In this article, we develop methods for making inferences regarding dependence of semi-competing risks data across strata of a discrete covariate Z. A class of rank statistics for testing constancy of association across strata are proposed; its asymptotic properties are also derived. We develop a novel resampling-based technique for calculating the variances of the proposed test statistics. In addition, we develop methods for combining test statistics for assessing …


Simultaneous Estimation Procedures And Multiple Testing: A Decision-Theoretic Framework, Debashis Ghosh Aug 2005

Simultaneous Estimation Procedures And Multiple Testing: A Decision-Theoretic Framework, Debashis Ghosh

The University of Michigan Department of Biostatistics Working Paper Series

There is recent tremendous interest in statistical methods regarding the false discovery rate (FDR). Two classes of literature on this topic exist. In the first, authors have proposed sequential testing procedures that control the false discovery rate. For the second, authors have studied the procedures involving FDR in a univariate mixture model setting. We consider a decision-theoretic approach to the assessment of FDR-based methods. In particular, we attempt to reconcile the current literature on false discovery rate procedures with more classical simultaneous estimation procedures. Formulation of the link will allow us to apply results from decision theory; we can then …


Shrunken P-Values For Assessing Differential Expression, With Applications To Genomic Data Analysis, Debashis Ghosh Aug 2005

Shrunken P-Values For Assessing Differential Expression, With Applications To Genomic Data Analysis, Debashis Ghosh

The University of Michigan Department of Biostatistics Working Paper Series

n many scientific problems involving high-throughput technology, inference must be made involving several hundreds or thousands of hypotheses. Recent attention has focused on how to address the multiple testing issue; much focus has been devoted towards use of the false discovery rate. In this article, we consider an alternative estimation procedure titled shrunken p-values for assessing differential expression (SPADE). The estimators are motivated by risk considerations from decision theory and lead to a completely new method for adjustment in the multiple testing problem. Some theoretical results are outlined. The proposed methodology is illustrated using simulation studies and with application to …


Specific Ige Response To Purified And Recombinant Allergens In Latex Allergy, Viswanath P. Kurup, Gordon L. Sussman, Hoong Y. Yeang, Nancy Elms, Heimo Breiteneder, Siti Am Arif, Kevin J. Kelly, Naveen K. Bansal, Jordan N. Fink Aug 2005

Specific Ige Response To Purified And Recombinant Allergens In Latex Allergy, Viswanath P. Kurup, Gordon L. Sussman, Hoong Y. Yeang, Nancy Elms, Heimo Breiteneder, Siti Am Arif, Kevin J. Kelly, Naveen K. Bansal, Jordan N. Fink

Mathematics, Statistics and Computer Science Faculty Research and Publications

Background

In recent years, allergy to natural rubber latex has emerged as a major allergy among certain occupational groups and patients with underlying diseases. The sensitization and development of latex allergy has been attributed to exposure to products containing residual latex proteins. Although improved manufacturing procedures resulted in a considerable reduction of new cases, the potential risk for some patient groups is still great. In addition the prevalent cross-reactivity of latex proteins with other food allergens poses a major concern. A number of purified allergens and a few commercial kits are currently available, but no concerted effort was undertaken to …


Application Of A Multiple Testing Procedure Controlling The Proportion Of False Positives To Protein And Bacterial Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan Aug 2005

Application Of A Multiple Testing Procedure Controlling The Proportion Of False Positives To Protein And Bacterial Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Simultaneously testing multiple hypotheses is important in high-dimensional biological studies. In these situations, one is often interested in controlling the Type-I error rate, such as the proportion of false positives to total rejections (TPPFP) at a specific level, alpha. This article will present an application of the E-Bayes/Bootstrap TPPFP procedure, presented in van der Laan et al. (2005), which controls the tail probability of the proportion of false positives (TPPFP), on two biological datasets. The two data applications include firstly, the application to a mass-spectrometry dataset of two leukemia subtypes, AML and ALL. The protein data measurements include intensity and …


Estimating The Discrepancy Between Computer Model Data And Field Data: Modeling Techniques For Deterministic And Stochastic Computer Simulators, Emily Joy Dastrup Aug 2005

Estimating The Discrepancy Between Computer Model Data And Field Data: Modeling Techniques For Deterministic And Stochastic Computer Simulators, Emily Joy Dastrup

Theses and Dissertations

Computer models have become useful research tools in many disciplines. In many cases a researcher has access to data from a computer simulator and from a physical system. This research discusses Bayesian models that allow for the estimation of the discrepancy between the two data sources. We fit two models to data in the field of electrical engineering. Using this data we illustrate ways of modeling both a deterministic and a stochastic simulator when specific parametric assumptions can be made about the discrepancy term.


Cross-Validating And Bagging Partitioning Algorithms With Variable Importance, Annette M. Molinaro, Mark J. Van Der Laan Aug 2005

Cross-Validating And Bagging Partitioning Algorithms With Variable Importance, Annette M. Molinaro, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

We present a cross-validated bagging scheme in the context of partitioning algorithms. To explore the benefits of the various bagging scheme, we compare via simulations the predictive ability of single Classification and Regression (CART) Tree with several previously suggested bagging schemes and with our proposed approach. Additionally, a variable importance measure is explained and illustrated.


Investigation Of Laser Beam Induced Current Techniques For Heterojunction Photodiode Characterization, Weifu Fang, David A. Redfern, Kazufumi Ito, G. Bahir, Charles A. Musca, John M. Dell, Lorenzo Faraone Aug 2005

Investigation Of Laser Beam Induced Current Techniques For Heterojunction Photodiode Characterization, Weifu Fang, David A. Redfern, Kazufumi Ito, G. Bahir, Charles A. Musca, John M. Dell, Lorenzo Faraone

Mathematics and Statistics Faculty Publications

A reduced model is developed that has significant advantages over the full drift-diffusion model for the simulation of laser beam-induced current (LBIC) signals in the presence of heterojunctions. The model determines the contribution to the LBIC signal that would occur from photogeneration at any position within the semiconductor, and is particularly useful for heterostructures where judicious choice of illumination wavelength can result in photogeneration at different depths within the device structure. The reduced model is used to examine the basic features of LBIC as applied to two types of planar P-n HgCdTe heterojunction photodiode structures. In particular, the question of …


Test Statistics Null Distributions In Multiple Testing: Simulation Studies And Applications To Genomics, Katherine S. Pollard, Merrill D. Birkner, Mark J. Van Der Laan, Sandrine Dudoit Jul 2005

Test Statistics Null Distributions In Multiple Testing: Simulation Studies And Applications To Genomics, Katherine S. Pollard, Merrill D. Birkner, Mark J. Van Der Laan, Sandrine Dudoit

U.C. Berkeley Division of Biostatistics Working Paper Series

Multiple hypothesis testing problems arise frequently in biomedical and genomic research, for instance, when identifying differentially expressed or co-expressed genes in microarray experiments. We have developed generally applicable resampling-based single-step and stepwise multiple testing procedures (MTP) for control of a broad class of Type I error rates, defined as tail probabilities and expected values for arbitrary functions of the numbers of false positives and rejected hypotheses (Dudoit and van der Laan, 2005; Dudoit et al., 2004a,b; Pollard and van der Laan, 2004; van der Laan et al., 2005, 2004a,b). As argued in the early article of Pollard and van der …


Performance Of Aic-Selected Spatial Covariance Structures For Fmri Data, David A. Stromberg Jul 2005

Performance Of Aic-Selected Spatial Covariance Structures For Fmri Data, David A. Stromberg

Theses and Dissertations

FMRI datasets allow scientists to assess functionality of the brain by measuring the response of blood flow to a stimulus. Since the responses from neighboring locations within the brain are correlated, simple linear models that assume independence of measurements across locations are inadequate. Mixed models can be used to model the spatial correlation between observations, however selecting the correct covariance structure is difficult. Information criteria, such as AIC are often used to choose among covariance structures. Once the covariance structure is selected, significance tests can be used to determine if a region of interest within the brain is significantly active. …


When Should One Substract Background Fluorescence In Two Color Microarrays?, Robert B. Scharpf, Christine A. Iacobuzio-Donahue, Julie B. Sneddon, Giovanni Parmigiani Jul 2005

When Should One Substract Background Fluorescence In Two Color Microarrays?, Robert B. Scharpf, Christine A. Iacobuzio-Donahue, Julie B. Sneddon, Giovanni Parmigiani

Johns Hopkins University, Dept. of Biostatistics Working Papers

Two color microarrays are a powerful tool for genomic analysis, but have noise components that make inferences regarding gene expression inefficient and potentially misleading. Background fluorescence,whether attributable to non-specific binding or other sources,is an important component of noise. The decision to subtract fluorescence surrounding spots of hybridization from spot fluorescence has been controversial, with no clear criteria for determining circumstances that may favor, or disfavor, background subtraction. While it is generally accepted that subtracting background reduces bias but increases variance in the estimates of the ratios of interest, no formal analysis of the bias-variance trade off of background subtraction has …


Does The Effect Of Micronutrient Supplementation On Neonatal Survival Vary With Respect To The Percentiles Of The Birth Weight Distribution?, Francesca Dominici, Scott L. Zeger, Giovanni Parmigiani, Joanne Katz, Parul Christian Jul 2005

Does The Effect Of Micronutrient Supplementation On Neonatal Survival Vary With Respect To The Percentiles Of The Birth Weight Distribution?, Francesca Dominici, Scott L. Zeger, Giovanni Parmigiani, Joanne Katz, Parul Christian

Johns Hopkins University, Dept. of Biostatistics Working Papers

Scientific Background: In developing countries, higher infant mortality is partially caused by poor maternal and fetal nutrition. Clinical trials of micronutrient supplementation are aimed at reducing the risk of infant mortality by increasing birth weight. Because infant mortality is greatest among the low birth weight infants (LBW) (less than or equal to 2500 grams), an effective intervention might be needed to increase birth weight among the smallest babies. Although it has been demonstrated that supplementation increases the birth weight in a trial conducted in Nepal, there is inconclusive evidence that the supplementation improves their survival. It has been hypothesized that …


Linear Regression Of Censored Length-Biased Lifetimes, Ying Qing Chen, Yan Wang Jul 2005

Linear Regression Of Censored Length-Biased Lifetimes, Ying Qing Chen, Yan Wang

UW Biostatistics Working Paper Series

Length-biased lifetimes may be collected in observational studies or sample surveys due to biased sampling scheme. In this article, we use a linear regression model, namely, the accelerated failure time model, for the population lifetime distributions in regression analysis of the length-biased lifetimes. It is discovered that the associated regression parameters are invariant under the length-biased sampling scheme. According to this discovery, we propose the quasi partial score estimating equations to estimate the population regression parameters. The proposed methodologies are evaluated and demonstrated by simulation studies and an application to actual data set.