Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

COBRA

Discipline
Keyword
Publication Year
Publication

Articles 1081 - 1108 of 1108

Full-Text Articles in Statistics and Probability

Why Prefer Double Robust Estimates? Illustration With Causal Point Treatment Studies, Romain Neugebauer, Mark J. Van Der Laan Sep 2002

Why Prefer Double Robust Estimates? Illustration With Causal Point Treatment Studies, Romain Neugebauer, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

In point treatment marginal structural models with treatment A, outcome Y and covariates W, causal parameters can be estimated under the assumption of no unobserved confounders. Three estimates can be used: the G-computation, Inverse Probability of Treatment Weighted (IPTW) or Double Robust (DR) estimates. The properties of the IPTW and DR estimates are known under an assumption on the treatment mechanism that we name "Experimental Treatment Assignment" (ETA) assumption. We show that the DR estimating function is unbiased when the ETA assumption is violated if the model used to regress Y on A and W is correctly specified. The practical …


Bivariate Current Status Data, Mark J. Van Der Laan, Nicholas P. Jewell Sep 2002

Bivariate Current Status Data, Mark J. Van Der Laan, Nicholas P. Jewell

U.C. Berkeley Division of Biostatistics Working Paper Series

In many applications, it is often of interest to estimate a bivariate distribution of two survival random variables. Complete observation of such random variables is often incomplete. If one only observes whether or not each of the individual survival times exceeds a common observed monitoring time C, then the data structure is referred to as bivariate current status data (Wang and Ding, 2000). For such data, we show that the identifiable part of the joint distribution is represented by three univariate cumulative distribution functions, namely the two marginal cumulative distribution functions, and the bivariate cumulative distribution function evaluated on the …


Current Status Data: Review, Recent Developments And Open Problems, Nicholas P. Jewell, Mark J. Van Der Laan Sep 2002

Current Status Data: Review, Recent Developments And Open Problems, Nicholas P. Jewell, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Researchers working with survival data are by now adept at handling issues associated with incomplete data, particular those associated with various forms of censoring. An extreme form of interval censoring, known as current status observation, refers to situations where the only available information on a survival random variable T is whether or not T exceeds a random independent monitoring time C. This article contains a brief review of the extensive literature on the analysis of current status data, discussing the implications of response-based sampling on these methods. The majority of the paper introduces some recent extensions of these ideas to …


Semiparametric Regression Analysis On Longitudinal Pattern Of Recurrent Gap Times, Ying Qing Chen, Mei-Cheng Wang, Yijian Huang Aug 2002

Semiparametric Regression Analysis On Longitudinal Pattern Of Recurrent Gap Times, Ying Qing Chen, Mei-Cheng Wang, Yijian Huang

U.C. Berkeley Division of Biostatistics Working Paper Series

In longitudinal studies, individual subjects may experience recurrent events of the same type over a relatively long period of time. The longitudinal pattern of the gaps between the successive recurrent events is often of great research interest. In this article, the probability structure of the recurrent gap times is first explored in the presence of censoring. According to the discovered structure, we introduce the proportional reverse-time hazards models with unspecified baseline functions to accommodate heterogeneous individual underlying distributions, when the ongitudinal pattern parameter is of main interest. Inference procedures are proposed and studied by way of proper riskset construction. The …


Multiple Hypothesis Testing In Microarray Experiments, Sandrine Dudoit, Juliet Popper Shaffer, Jennifer C. Boldrick Aug 2002

Multiple Hypothesis Testing In Microarray Experiments, Sandrine Dudoit, Juliet Popper Shaffer, Jennifer C. Boldrick

U.C. Berkeley Division of Biostatistics Working Paper Series

DNA microarrays are a new and promising biotechnology which allows the monitoring of expression levels in cells for thousands of genes simultaneously. An important and common question in microarray experiments is the identification of differentially expressed genes, i.e., genes whose expression levels are associated with a response or covariate of interest. The biological question of differential expression can be restated as a problem in multiple hypothesis testing: the simultaneous test for each gene of the null hypothesis of no association between the expression levels and the responses or covariates. As a typical microarray experiment measures expression levels for thousands of …


Estimation Of The Bivariate Survival Function With Generalized Bivariate Right Censored Data Structures, Sunduz Keles, Mark J. Van Der Laan, James M. Robins Aug 2002

Estimation Of The Bivariate Survival Function With Generalized Bivariate Right Censored Data Structures, Sunduz Keles, Mark J. Van Der Laan, James M. Robins

U.C. Berkeley Division of Biostatistics Working Paper Series

We propose a bivariate survival function estimator for a general right censored data structure that includes a time dependent covariate process. Firstly, an initial estimator that generalizes Dabrowska's (1988) estimator is introduced. We obtain this estimator by a general methodology of constructing estimating functions in censored data models. The initial estimator is guaranteed to improve on Dabrowska's estimator and remains consistent and asymptotically linear under informative censoring schemes if the censoring mechanism is estimated consistently. We then construct an orthogonalized estimating function which results in a more robust and efficient estimator than our initial estimator. A simulation study demonstrates the …


Inference For Proportional Mean Residual Life Model In The Presence Of Censoring, Ying Q. Chen, Nicholas P. Jewell May 2002

Inference For Proportional Mean Residual Life Model In The Presence Of Censoring, Ying Q. Chen, Nicholas P. Jewell

U.C. Berkeley Division of Biostatistics Working Paper Series

As a function of time t, mean residual life is defined as remaining life expectancy of a subject given its survival to t. It plays an important role in many research areas to characterise stochastic behavior of survival over time. Similar to the Cox proportional hazard model, the proportional mean residual life model were proposed in statistical literature to study association between the mean residual life and individual subject's explanatory covariates. In this article, we will study this model and develop appropriate inference procedures in presence of censoring. Numerical studies including simulation and real data analysis are presented as well.


Comparative Genomic Hybridization Array Analysis, Annette M. Molinaro, Mark J. Van Der Laan, Dan H. Moore Apr 2002

Comparative Genomic Hybridization Array Analysis, Annette M. Molinaro, Mark J. Van Der Laan, Dan H. Moore

U.C. Berkeley Division of Biostatistics Working Paper Series

At the present time, there is increasing evidence that cancer may be regulated by the number of copies of genes in tumor cells. Through microarray technology it is now possible to measure the number of copies of thousands of genes and gene segments in samples of chromosomal DNA. Microarray comparative genomic hybridization (array CGH) provides the opportunity to both measure DNA sequence copy number gains and losses and map these aberrations to the genomic sequence. Gains can signify the over-expression of oncogenes, genes which stimulate cell growth and have become hyperactive, while losses can signify under-expression of tumor suppressor genes, …


A Method To Identify Significant Clusters In Gene Expression Data, Katherine S. Pollard, Mark J. Van Der Laan Apr 2002

A Method To Identify Significant Clusters In Gene Expression Data, Katherine S. Pollard, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Clustering algorithms have been widely applied to gene expression data. For both hierarchical and partitioning clustering algorithms, selecting the number of significant clusters is an important problem and many methods have been proposed. Existing methods for selecting the number of clusters tend to find only the global patterns in the data (e.g.: the over and under expressed genes). We have noted the need for a better method in the gene expression context, where small, biologically meaningful clusters can be difficult to identify. In this paper, we define a new criteria, Mean Split Silhouette (MSS), which is a measure of cluster …


A New Partitioning Around Medoids Algorithm, Mark J. Van Der Laan, Katherine S. Pollard, Jennifer Bryan Feb 2002

A New Partitioning Around Medoids Algorithm, Mark J. Van Der Laan, Katherine S. Pollard, Jennifer Bryan

U.C. Berkeley Division of Biostatistics Working Paper Series

Kaufman & Rousseeuw (1990) proposed a clustering algorithm Partitioning Around Medoids (PAM) which maps a distance matrix into a specified number of clusters. A particularly nice property is that PAM allows clustering with respect to any specified distance metric. In addition, the medoids are robust representations of the cluster centers, which is particularly important in the common context that many elements do not belong well to any cluster. Based on our experience in clustering gene expression data, we have noticed that PAM does have problems recognizing relatively small clusters in situations where good partitions around medoids clearly exist. In this …


Regression Analysis Of Recurrent Gap Times With Time-Dependent Covariates, Ying Qing Chen, Mei-Cheng Wang, Yijian Huang Jan 2002

Regression Analysis Of Recurrent Gap Times With Time-Dependent Covariates, Ying Qing Chen, Mei-Cheng Wang, Yijian Huang

U.C. Berkeley Division of Biostatistics Working Paper Series

Individual subjects may experience recurrent events of same type over a relatively long period of time in a longitudinal study. Researchers are often interested in the distributional pattern of gaps between the successive recurrent events and their association with certain concomitant covariates as well. In this article, their probability structure is investigated in presence of censoring. According to the identified structure, we introduce the proportional reverse-time hazards models that allow arbitrary baseline function for every individual in the study, when the time-dependent covariates effect is of main interest. Appropriate inference procedures are proposed and studied to estimate the parameters of …


Estimating Causal Parameters In Marginal Structural Models With Unmeasured Confounders Using Instrumental Variables, Tanya A. Henneman, Mark Johannes Van Der Laan, Alan E. Hubbard Jan 2002

Estimating Causal Parameters In Marginal Structural Models With Unmeasured Confounders Using Instrumental Variables, Tanya A. Henneman, Mark Johannes Van Der Laan, Alan E. Hubbard

U.C. Berkeley Division of Biostatistics Working Paper Series

For statisticians analyzing medical data, a significant problem in determining the causal effect of a treatment on a particular outcome of interest, is how to control for unmeasured confounders. Techniques using instrumental variables (IV) have been developed to estimate causal parameters in the presence of unmeasured confounders. In this paper we apply IV methods to both linear and non-linear marginal structural models. We study a specific class of generalized estimating equations that is appropriate to these data, and compare the performance of the resulting estimator to the standard IV method, a two-stage least squares procedure. Our results are applied to …


Marginal Regression Of Gaps Between Recurrent Events, Yijian Huang, Ying Qing Chen Nov 2001

Marginal Regression Of Gaps Between Recurrent Events, Yijian Huang, Ying Qing Chen

U.C. Berkeley Division of Biostatistics Working Paper Series

Recurrent event data typically exhibit the phenomenon of intra-individual correlation, owing to not only observed covariates but also random effects. In many applications, the population can be reasonably postulated as a heterogeneous mixture of individual renewal processes, and the inference of interest is the effect of individual-level covariates. In this article, we suggest and investigate a marginal proportional hazards model for gaps between recurrent events. A connection is established between observed gap times and clustered survival data, however, with informative cluster size. We then derive a novel and general inference procedure for the latter, based on a functional formulation of …


Maximum Likelihood Estimation Of Ordered Multinomial Parameters, Nicholas P. Jewell, John D. Kalbfleisch Oct 2001

Maximum Likelihood Estimation Of Ordered Multinomial Parameters, Nicholas P. Jewell, John D. Kalbfleisch

U.C. Berkeley Division of Biostatistics Working Paper Series

The pool-adjacent violator-algorithm (Ayer, et al., 1955) has long been known to give the maximum likelihood estimator of a series of ordered binomial parameters, based on an independent observation from each distribution (see Barlow et al., 1972). This result has immediate application to estimation of a survival distribution based on current survival status at a set of monitoring times. This paper considers an extended problem of maximum likelihood estimation of a series of ‘ordered’ multinomial parameters. By making use of variants of the pool adjacent violator algorithm, we obtain a simple algorithm to compute the maximum likelihood estimator and demonstrate …


Identification Of Regulatory Elements Using A Feature Selection Method, Sunduz Keles, Mark J. Van Der Laan, Michael B. Eisen Sep 2001

Identification Of Regulatory Elements Using A Feature Selection Method, Sunduz Keles, Mark J. Van Der Laan, Michael B. Eisen

U.C. Berkeley Division of Biostatistics Working Paper Series

Many methods have been described to identify regulatory motifs in the transcription control regions of genes that exhibit similar patterns of gene expression across a variety of experimental conditions. Here we focus on a single experimental condition, and utilize gene expression data to identify sequence motifs associated with genes that are activated under this experimental condition. We use a linear model with two way interactions to model gene expression as a function of sequence features (words) present in presumptive transcription control regions. The most relevant features are selected by a feature selection method called stepwise selection with monte carlo cross …


Statistical Inference For Simultaneous Clustering Of Gene Expression Data, Katherine S. Pollard, Mark J. Van Der Laan Jul 2001

Statistical Inference For Simultaneous Clustering Of Gene Expression Data, Katherine S. Pollard, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Current methods for analysis of gene expression data are mostly based on clustering and classification of either genes or samples. We offer support for the idea that more complex patterns can be identified in the data if genes and samples are considered simultaneously. We formalize the approach and propose a statistical framework for two-way clustering. A simultaneous clustering parameter is defined as a function of the true data generating distribution, and an estimate is obtained by applying this function to the empirical distribution. We illustrate that a wide range of clustering procedures, including generalized hierarchical methods, can be defined as …


Mixture Hazards Models With Additive Random Effects Accounting For Treatment Effectiveness Lag Time, Ying Qing Chen, C. A. Rohde, M.-C. Wang Jan 2001

Mixture Hazards Models With Additive Random Effects Accounting For Treatment Effectiveness Lag Time, Ying Qing Chen, C. A. Rohde, M.-C. Wang

U.C. Berkeley Division of Biostatistics Working Paper Series

In many clinical trials to evaluate treatment efficacy, it is believed that there may exist latent treatment effectiveness lag times after which medical treatment procedure or chemical compound would be in full effect. In this article, semiparametric regression models are proposed and studied for estimating the treatment effect accounting for such latent lag times. The new models take advantage of the invariance property of the additive hazards model in marginalising over an additive latent variable; parameters in the models are thus easily estimated and interpreted, while the flexibility of not having to specify the baseline hazard function is preserved. Monte …


Probabilities Of Transition Among Health States For Older Adults, Paula Diehr, Donald L. Patrick Jan 2001

Probabilities Of Transition Among Health States For Older Adults, Paula Diehr, Donald L. Patrick

UW Biostatistics Working Paper Series

Goal: To estimate the probabilities of transition among self-rated health states for older adults, and examine how they vary by age and sex. Methods: We used self-rated health (Excellent, Very Good, Good, Fair, Poor, Dead) collected in two longitudinal studies of older adults (Mean age 75) to estimate the probability of transition in two years. We used the estimates to project future health for selected cohorts.

Findings: These older adults were most likely to be in the same health state 2 years later, but a substantial proportion changed in both directions. Transition probabilities varied by initial health state, age and …


A Class Of Semiparametric Scale-Change Hazards Regression Models And Its Adequacy For Censored Survival Data, Ying Qing Chen Oct 2000

A Class Of Semiparametric Scale-Change Hazards Regression Models And Its Adequacy For Censored Survival Data, Ying Qing Chen

U.C. Berkeley Division of Biostatistics Working Paper Series

A class of semiparametric hazards regression models called the accelerated hazards models was introduced to identify the covariate effect characterized by the scale-change between hazard functions. In this article, we compare the accelerated hazards models with several other popular classes of regression models in statistical literature for censored survival data. We also propose and study some test statistics to assess the models' adequacy. Simulation studies are conducted to evaluate the performance of the test statistics. Actual clinical trials data are analyzed to demonstrate the proposed models and test statistics.


Assessing The Accuracy Of A New Diagnostic Test When A Gold Standard Does Not Exist, Todd A. Alonzo, Margaret S. Pepe Oct 1998

Assessing The Accuracy Of A New Diagnostic Test When A Gold Standard Does Not Exist, Todd A. Alonzo, Margaret S. Pepe

UW Biostatistics Working Paper Series

Often the accuracy of a new diagnostic test must be assessed when a perfect gold standard does not exist. Use of an imperfect test biases the accuracy estimates of the new test. This paper reviews existing approaches to this problem including discrepant resolution and latent class analysis. Deficiencies with these approaches are identified. A new approach is proposed that combines the results of several imperfect reference tests to define a better reference standard. We call this the composite reference standard (CRS). Using the CRS, accuracy can be assessed using multistage sampling designs. Maximum likelihood estimates of accuracy and expressions for …


Multiple Outcomes In Health Services Research: Hypothesis Tests And Power, Donald C. Martin, Paula Diehr, Thomas D. Koepsell, Stephan D. Fihn Oct 1997

Multiple Outcomes In Health Services Research: Hypothesis Tests And Power, Donald C. Martin, Paula Diehr, Thomas D. Koepsell, Stephan D. Fihn

UW Biostatistics Working Paper Series

Health services research often is directed towards making small improvements in a number of outcomes that reflect many aspects of the patient’s life rather than a large improvement in a single well defined outcome. A researcher might choose five scales to measure different aspects of treatment outcomes and not expect any large treatment differences on any single outcome measure. O’Brien (1984) has proposed a nonparametric statistical procedure which is particularly well suited to this type of problem and that can result in considerable increases in statistical power. This paper will briefly review O’Brien’s pooled rank method and develop power calculations. …


Pooling Community Data For Community Interventions When The Number Of Pairs Is Small, Paula Diehr, Ted Lystig, Holly Andrilla, Ziding Feng May 1997

Pooling Community Data For Community Interventions When The Number Of Pairs Is Small, Paula Diehr, Ted Lystig, Holly Andrilla, Ziding Feng

UW Biostatistics Working Paper Series

There is considerable interest in community interventions for health promotion, where the community is the experimental unit. Because such interventions are expensive, the number of experimental units (communities) is usually very small, yielding a study with low power. We examined the ability of a process known as “pooling” or “preliminary significance testing” to improve the power of community variations. In this process, one first tests whether there is significant community variation, using type 1 error of perhaps 0.25. If there is significant variation, the usual community-level test is performed. If not, a person-level test is performed. We found through Monte …


An Empirical Study Of Small-Area Variation For Icd-9 Surgical Procedures, Paula Diehr, Kevin Cain, Zhan Ye, John Loeser Apr 1994

An Empirical Study Of Small-Area Variation For Icd-9 Surgical Procedures, Paula Diehr, Kevin Cain, Zhan Ye, John Loeser

UW Biostatistics Working Paper Series

Objective. Several measures of variation have been used in SAVA. One study of DRGs found that the coefficient of variation from analysis of variance (CVA) had superior performance. That work is replicated here for ICD-9 surgical procedures, and extended to age/sex-standardized rates. Results are compared with those in the literature, and recommendations are made for assessing small-area variation in future studies.

Data Sources. Data were taken from Washington State's "Episode of Illness" file of hospital discharges in the State in 1987. Up to three ICD-9 surgical procedures and a unique patient identifier were available for each discharge.

Study Design. We …


Breaking The Matches In A Paired T-Test For Community Interventions When The Number Of Pairs Is Small, Paula Diehr, Don C. Martin, Thomas D. Koepsell, Allen D. Cheadle Mar 1993

Breaking The Matches In A Paired T-Test For Community Interventions When The Number Of Pairs Is Small, Paula Diehr, Don C. Martin, Thomas D. Koepsell, Allen D. Cheadle

UW Biostatistics Working Paper Series

There is considerable interest in community interventions for health promotion, where the community is the experimental unit. Because such interventions are expensive, the number of experimental units (communities) is usually small. Because of the small number of communities involved, investigators often match treatment and control communities on demographic variables before randomization to minimize the possibility of a bad split. Unfortunately, matching has been shown to decrease the power of the design when the number of pairs is small, unless the matching variable is very highly correlated with the outcome variable (in this case, with change in the health behavior). We …


The Multiple Admission Factor (Maf) In Small Area Variation Analysis, Kevin Cain, Paula Diehr Dec 1992

The Multiple Admission Factor (Maf) In Small Area Variation Analysis, Kevin Cain, Paula Diehr

UW Biostatistics Working Paper Series

Small area variation analysis are often based on area-level data such as the total number of hospital admissions within an area, rather than person-level data. Such analysis often make the assumption that the number of admissions within a small area follow a Poisson distribution. This may not be a reasonable assumption when multiple admissions per person are possible. In this case, the multiple admission factor (MAF) can be used to adjust for the extra variance introduced by multiple admissions. In this article, data from Washington State are used to estimate the multiple admission rate and the MAF for each modifed …


Regression Models For Bivariate Binary Responses, Juni Palmgren Nov 1989

Regression Models For Bivariate Binary Responses, Juni Palmgren

UW Biostatistics Working Paper Series

We discuss maximum likelihood inference for the bivariate logistic model, specified in terms of the marginal logits and the log odds ratio. Using the exponential family nonlinear model formulation the model fitting can be done in GLIM. The procedure is illustrated by modelling survival of unilateral and bilateral total hip arthroplasties as function of patient specific and hip specific covariates. We compare maximum likelihood inference with inference obtained from solving likelihood equations under the assumption of within block independence and using robust standard errors for the estimates. Simulations indicate that the latter procedure is effcient for block specific covariates but …


Sample Size Calculations And Optimal Followup Time In Health Services Research Using Utilization Rates, Paula Diehr Aug 1980

Sample Size Calculations And Optimal Followup Time In Health Services Research Using Utilization Rates, Paula Diehr

UW Biostatistics Working Paper Series

It is not always possible to estimate the sample sizes needed in health services research because special formulas are needed, and the necessary data may not be available to use in the formulas. We provide some useful formulas for the sample size required in comparing the means of two groups. These include the special case where the two groups are not of equal size either because one is known to have a higher variability or because one group has already been chosen and its size is thus fixed. We also explore the relationship of the mean to the standard deviation …


Statistical Measures For Admission Rates, Paula Diehr Aug 1978

Statistical Measures For Admission Rates, Paula Diehr

UW Biostatistics Working Paper Series

Hospital admission rates are often shown and interpreted without consideration of their inherent variability, which may lead to faulty conclusions. This may be because theoretically correct variance estimates are not known for the type of estimates usually used; i.e., total admissions divided by total person-months of observation. Here, correct methods for testing and estimation are shown for situations where they exist. For other types of data, approximate procedures are proposed and their properties examined theoretically and empirically, yielding recommendations for exact and approximate estimation and testing methods for admission rates in common situations.