Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Biostatistics (567)
- Statistical Methodology (362)
- Statistical Theory (336)
- Statistical Models (242)
- Medicine and Health Sciences (176)
-
- Survival Analysis (147)
- Public Health (142)
- Epidemiology (99)
- Life Sciences (92)
- Longitudinal Data Analysis and Time Series (89)
- Genetics and Genomics (88)
- Clinical Trials (82)
- Microarrays (78)
- Multivariate Analysis (78)
- Applied Mathematics (57)
- Genetics (57)
- Numerical Analysis and Computation (57)
- Bioinformatics (50)
- Computational Biology (50)
- Categorical Data Analysis (49)
- Design of Experiments and Sample Surveys (39)
- Clinical Epidemiology (36)
- Diseases (30)
- Disease Modeling (28)
- Medical Specialties (23)
- Health Services Research (17)
- Applied Statistics (13)
- Vital and Health Statistics (11)
- Keyword
-
- Causal inference (30)
- Cross-validation (25)
- Prediction (23)
- Genetics (21)
- Longitudinal data (19)
-
- Survival analysis (16)
- Classification (14)
- Influence curve (14)
- Model selection (14)
- Sensitivity (14)
- Bootstrap (13)
- Gene expression (13)
- Clinical trials (12)
- Targeted maximum likelihood estimation (12)
- Counterfactual (11)
- Efficient influence curve (11)
- Multiple testing (11)
- Confounding (10)
- Loss function (10)
- Missing data (10)
- Variable selection (10)
- Causal effect (9)
- Estimating equation (9)
- Measurement error (9)
- Regression (9)
- Specificity (9)
- Adjusted p-value (8)
- Air pollution (8)
- Asymptotic linearity (8)
- Censoring (8)
- Publication Year
- Publication
-
- U.C. Berkeley Division of Biostatistics Working Paper Series (242)
- UW Biostatistics Working Paper Series (215)
- Harvard University Biostatistics Working Paper Series (212)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (178)
- The University of Michigan Department of Biostatistics Working Paper Series (111)
Articles 1081 - 1108 of 1108
Full-Text Articles in Statistics and Probability
Why Prefer Double Robust Estimates? Illustration With Causal Point Treatment Studies, Romain Neugebauer, Mark J. Van Der Laan
Why Prefer Double Robust Estimates? Illustration With Causal Point Treatment Studies, Romain Neugebauer, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
In point treatment marginal structural models with treatment A, outcome Y and covariates W, causal parameters can be estimated under the assumption of no unobserved confounders. Three estimates can be used: the G-computation, Inverse Probability of Treatment Weighted (IPTW) or Double Robust (DR) estimates. The properties of the IPTW and DR estimates are known under an assumption on the treatment mechanism that we name "Experimental Treatment Assignment" (ETA) assumption. We show that the DR estimating function is unbiased when the ETA assumption is violated if the model used to regress Y on A and W is correctly specified. The practical …
Bivariate Current Status Data, Mark J. Van Der Laan, Nicholas P. Jewell
Bivariate Current Status Data, Mark J. Van Der Laan, Nicholas P. Jewell
U.C. Berkeley Division of Biostatistics Working Paper Series
In many applications, it is often of interest to estimate a bivariate distribution of two survival random variables. Complete observation of such random variables is often incomplete. If one only observes whether or not each of the individual survival times exceeds a common observed monitoring time C, then the data structure is referred to as bivariate current status data (Wang and Ding, 2000). For such data, we show that the identifiable part of the joint distribution is represented by three univariate cumulative distribution functions, namely the two marginal cumulative distribution functions, and the bivariate cumulative distribution function evaluated on the …
Current Status Data: Review, Recent Developments And Open Problems, Nicholas P. Jewell, Mark J. Van Der Laan
Current Status Data: Review, Recent Developments And Open Problems, Nicholas P. Jewell, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Researchers working with survival data are by now adept at handling issues associated with incomplete data, particular those associated with various forms of censoring. An extreme form of interval censoring, known as current status observation, refers to situations where the only available information on a survival random variable T is whether or not T exceeds a random independent monitoring time C. This article contains a brief review of the extensive literature on the analysis of current status data, discussing the implications of response-based sampling on these methods. The majority of the paper introduces some recent extensions of these ideas to …
Semiparametric Regression Analysis On Longitudinal Pattern Of Recurrent Gap Times, Ying Qing Chen, Mei-Cheng Wang, Yijian Huang
Semiparametric Regression Analysis On Longitudinal Pattern Of Recurrent Gap Times, Ying Qing Chen, Mei-Cheng Wang, Yijian Huang
U.C. Berkeley Division of Biostatistics Working Paper Series
In longitudinal studies, individual subjects may experience recurrent events of the same type over a relatively long period of time. The longitudinal pattern of the gaps between the successive recurrent events is often of great research interest. In this article, the probability structure of the recurrent gap times is first explored in the presence of censoring. According to the discovered structure, we introduce the proportional reverse-time hazards models with unspecified baseline functions to accommodate heterogeneous individual underlying distributions, when the ongitudinal pattern parameter is of main interest. Inference procedures are proposed and studied by way of proper riskset construction. The …
Multiple Hypothesis Testing In Microarray Experiments, Sandrine Dudoit, Juliet Popper Shaffer, Jennifer C. Boldrick
Multiple Hypothesis Testing In Microarray Experiments, Sandrine Dudoit, Juliet Popper Shaffer, Jennifer C. Boldrick
U.C. Berkeley Division of Biostatistics Working Paper Series
DNA microarrays are a new and promising biotechnology which allows the monitoring of expression levels in cells for thousands of genes simultaneously. An important and common question in microarray experiments is the identification of differentially expressed genes, i.e., genes whose expression levels are associated with a response or covariate of interest. The biological question of differential expression can be restated as a problem in multiple hypothesis testing: the simultaneous test for each gene of the null hypothesis of no association between the expression levels and the responses or covariates. As a typical microarray experiment measures expression levels for thousands of …
Estimation Of The Bivariate Survival Function With Generalized Bivariate Right Censored Data Structures, Sunduz Keles, Mark J. Van Der Laan, James M. Robins
Estimation Of The Bivariate Survival Function With Generalized Bivariate Right Censored Data Structures, Sunduz Keles, Mark J. Van Der Laan, James M. Robins
U.C. Berkeley Division of Biostatistics Working Paper Series
We propose a bivariate survival function estimator for a general right censored data structure that includes a time dependent covariate process. Firstly, an initial estimator that generalizes Dabrowska's (1988) estimator is introduced. We obtain this estimator by a general methodology of constructing estimating functions in censored data models. The initial estimator is guaranteed to improve on Dabrowska's estimator and remains consistent and asymptotically linear under informative censoring schemes if the censoring mechanism is estimated consistently. We then construct an orthogonalized estimating function which results in a more robust and efficient estimator than our initial estimator. A simulation study demonstrates the …
Inference For Proportional Mean Residual Life Model In The Presence Of Censoring, Ying Q. Chen, Nicholas P. Jewell
Inference For Proportional Mean Residual Life Model In The Presence Of Censoring, Ying Q. Chen, Nicholas P. Jewell
U.C. Berkeley Division of Biostatistics Working Paper Series
As a function of time t, mean residual life is defined as remaining life expectancy of a subject given its survival to t. It plays an important role in many research areas to characterise stochastic behavior of survival over time. Similar to the Cox proportional hazard model, the proportional mean residual life model were proposed in statistical literature to study association between the mean residual life and individual subject's explanatory covariates. In this article, we will study this model and develop appropriate inference procedures in presence of censoring. Numerical studies including simulation and real data analysis are presented as well.
Comparative Genomic Hybridization Array Analysis, Annette M. Molinaro, Mark J. Van Der Laan, Dan H. Moore
Comparative Genomic Hybridization Array Analysis, Annette M. Molinaro, Mark J. Van Der Laan, Dan H. Moore
U.C. Berkeley Division of Biostatistics Working Paper Series
At the present time, there is increasing evidence that cancer may be regulated by the number of copies of genes in tumor cells. Through microarray technology it is now possible to measure the number of copies of thousands of genes and gene segments in samples of chromosomal DNA. Microarray comparative genomic hybridization (array CGH) provides the opportunity to both measure DNA sequence copy number gains and losses and map these aberrations to the genomic sequence. Gains can signify the over-expression of oncogenes, genes which stimulate cell growth and have become hyperactive, while losses can signify under-expression of tumor suppressor genes, …
A Method To Identify Significant Clusters In Gene Expression Data, Katherine S. Pollard, Mark J. Van Der Laan
A Method To Identify Significant Clusters In Gene Expression Data, Katherine S. Pollard, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Clustering algorithms have been widely applied to gene expression data. For both hierarchical and partitioning clustering algorithms, selecting the number of significant clusters is an important problem and many methods have been proposed. Existing methods for selecting the number of clusters tend to find only the global patterns in the data (e.g.: the over and under expressed genes). We have noted the need for a better method in the gene expression context, where small, biologically meaningful clusters can be difficult to identify. In this paper, we define a new criteria, Mean Split Silhouette (MSS), which is a measure of cluster …
A New Partitioning Around Medoids Algorithm, Mark J. Van Der Laan, Katherine S. Pollard, Jennifer Bryan
A New Partitioning Around Medoids Algorithm, Mark J. Van Der Laan, Katherine S. Pollard, Jennifer Bryan
U.C. Berkeley Division of Biostatistics Working Paper Series
Kaufman & Rousseeuw (1990) proposed a clustering algorithm Partitioning Around Medoids (PAM) which maps a distance matrix into a specified number of clusters. A particularly nice property is that PAM allows clustering with respect to any specified distance metric. In addition, the medoids are robust representations of the cluster centers, which is particularly important in the common context that many elements do not belong well to any cluster. Based on our experience in clustering gene expression data, we have noticed that PAM does have problems recognizing relatively small clusters in situations where good partitions around medoids clearly exist. In this …
Regression Analysis Of Recurrent Gap Times With Time-Dependent Covariates, Ying Qing Chen, Mei-Cheng Wang, Yijian Huang
Regression Analysis Of Recurrent Gap Times With Time-Dependent Covariates, Ying Qing Chen, Mei-Cheng Wang, Yijian Huang
U.C. Berkeley Division of Biostatistics Working Paper Series
Individual subjects may experience recurrent events of same type over a relatively long period of time in a longitudinal study. Researchers are often interested in the distributional pattern of gaps between the successive recurrent events and their association with certain concomitant covariates as well. In this article, their probability structure is investigated in presence of censoring. According to the identified structure, we introduce the proportional reverse-time hazards models that allow arbitrary baseline function for every individual in the study, when the time-dependent covariates effect is of main interest. Appropriate inference procedures are proposed and studied to estimate the parameters of …
Estimating Causal Parameters In Marginal Structural Models With Unmeasured Confounders Using Instrumental Variables, Tanya A. Henneman, Mark Johannes Van Der Laan, Alan E. Hubbard
Estimating Causal Parameters In Marginal Structural Models With Unmeasured Confounders Using Instrumental Variables, Tanya A. Henneman, Mark Johannes Van Der Laan, Alan E. Hubbard
U.C. Berkeley Division of Biostatistics Working Paper Series
For statisticians analyzing medical data, a significant problem in determining the causal effect of a treatment on a particular outcome of interest, is how to control for unmeasured confounders. Techniques using instrumental variables (IV) have been developed to estimate causal parameters in the presence of unmeasured confounders. In this paper we apply IV methods to both linear and non-linear marginal structural models. We study a specific class of generalized estimating equations that is appropriate to these data, and compare the performance of the resulting estimator to the standard IV method, a two-stage least squares procedure. Our results are applied to …
Marginal Regression Of Gaps Between Recurrent Events, Yijian Huang, Ying Qing Chen
Marginal Regression Of Gaps Between Recurrent Events, Yijian Huang, Ying Qing Chen
U.C. Berkeley Division of Biostatistics Working Paper Series
Recurrent event data typically exhibit the phenomenon of intra-individual correlation, owing to not only observed covariates but also random effects. In many applications, the population can be reasonably postulated as a heterogeneous mixture of individual renewal processes, and the inference of interest is the effect of individual-level covariates. In this article, we suggest and investigate a marginal proportional hazards model for gaps between recurrent events. A connection is established between observed gap times and clustered survival data, however, with informative cluster size. We then derive a novel and general inference procedure for the latter, based on a functional formulation of …
Maximum Likelihood Estimation Of Ordered Multinomial Parameters, Nicholas P. Jewell, John D. Kalbfleisch
Maximum Likelihood Estimation Of Ordered Multinomial Parameters, Nicholas P. Jewell, John D. Kalbfleisch
U.C. Berkeley Division of Biostatistics Working Paper Series
The pool-adjacent violator-algorithm (Ayer, et al., 1955) has long been known to give the maximum likelihood estimator of a series of ordered binomial parameters, based on an independent observation from each distribution (see Barlow et al., 1972). This result has immediate application to estimation of a survival distribution based on current survival status at a set of monitoring times. This paper considers an extended problem of maximum likelihood estimation of a series of ‘ordered’ multinomial parameters. By making use of variants of the pool adjacent violator algorithm, we obtain a simple algorithm to compute the maximum likelihood estimator and demonstrate …
Identification Of Regulatory Elements Using A Feature Selection Method, Sunduz Keles, Mark J. Van Der Laan, Michael B. Eisen
Identification Of Regulatory Elements Using A Feature Selection Method, Sunduz Keles, Mark J. Van Der Laan, Michael B. Eisen
U.C. Berkeley Division of Biostatistics Working Paper Series
Many methods have been described to identify regulatory motifs in the transcription control regions of genes that exhibit similar patterns of gene expression across a variety of experimental conditions. Here we focus on a single experimental condition, and utilize gene expression data to identify sequence motifs associated with genes that are activated under this experimental condition. We use a linear model with two way interactions to model gene expression as a function of sequence features (words) present in presumptive transcription control regions. The most relevant features are selected by a feature selection method called stepwise selection with monte carlo cross …
Statistical Inference For Simultaneous Clustering Of Gene Expression Data, Katherine S. Pollard, Mark J. Van Der Laan
Statistical Inference For Simultaneous Clustering Of Gene Expression Data, Katherine S. Pollard, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Current methods for analysis of gene expression data are mostly based on clustering and classification of either genes or samples. We offer support for the idea that more complex patterns can be identified in the data if genes and samples are considered simultaneously. We formalize the approach and propose a statistical framework for two-way clustering. A simultaneous clustering parameter is defined as a function of the true data generating distribution, and an estimate is obtained by applying this function to the empirical distribution. We illustrate that a wide range of clustering procedures, including generalized hierarchical methods, can be defined as …
Mixture Hazards Models With Additive Random Effects Accounting For Treatment Effectiveness Lag Time, Ying Qing Chen, C. A. Rohde, M.-C. Wang
Mixture Hazards Models With Additive Random Effects Accounting For Treatment Effectiveness Lag Time, Ying Qing Chen, C. A. Rohde, M.-C. Wang
U.C. Berkeley Division of Biostatistics Working Paper Series
In many clinical trials to evaluate treatment efficacy, it is believed that there may exist latent treatment effectiveness lag times after which medical treatment procedure or chemical compound would be in full effect. In this article, semiparametric regression models are proposed and studied for estimating the treatment effect accounting for such latent lag times. The new models take advantage of the invariance property of the additive hazards model in marginalising over an additive latent variable; parameters in the models are thus easily estimated and interpreted, while the flexibility of not having to specify the baseline hazard function is preserved. Monte …
Probabilities Of Transition Among Health States For Older Adults, Paula Diehr, Donald L. Patrick
Probabilities Of Transition Among Health States For Older Adults, Paula Diehr, Donald L. Patrick
UW Biostatistics Working Paper Series
Goal: To estimate the probabilities of transition among self-rated health states for older adults, and examine how they vary by age and sex. Methods: We used self-rated health (Excellent, Very Good, Good, Fair, Poor, Dead) collected in two longitudinal studies of older adults (Mean age 75) to estimate the probability of transition in two years. We used the estimates to project future health for selected cohorts.
Findings: These older adults were most likely to be in the same health state 2 years later, but a substantial proportion changed in both directions. Transition probabilities varied by initial health state, age and …
A Class Of Semiparametric Scale-Change Hazards Regression Models And Its Adequacy For Censored Survival Data, Ying Qing Chen
A Class Of Semiparametric Scale-Change Hazards Regression Models And Its Adequacy For Censored Survival Data, Ying Qing Chen
U.C. Berkeley Division of Biostatistics Working Paper Series
A class of semiparametric hazards regression models called the accelerated hazards models was introduced to identify the covariate effect characterized by the scale-change between hazard functions. In this article, we compare the accelerated hazards models with several other popular classes of regression models in statistical literature for censored survival data. We also propose and study some test statistics to assess the models' adequacy. Simulation studies are conducted to evaluate the performance of the test statistics. Actual clinical trials data are analyzed to demonstrate the proposed models and test statistics.
Assessing The Accuracy Of A New Diagnostic Test When A Gold Standard Does Not Exist, Todd A. Alonzo, Margaret S. Pepe
Assessing The Accuracy Of A New Diagnostic Test When A Gold Standard Does Not Exist, Todd A. Alonzo, Margaret S. Pepe
UW Biostatistics Working Paper Series
Often the accuracy of a new diagnostic test must be assessed when a perfect gold standard does not exist. Use of an imperfect test biases the accuracy estimates of the new test. This paper reviews existing approaches to this problem including discrepant resolution and latent class analysis. Deficiencies with these approaches are identified. A new approach is proposed that combines the results of several imperfect reference tests to define a better reference standard. We call this the composite reference standard (CRS). Using the CRS, accuracy can be assessed using multistage sampling designs. Maximum likelihood estimates of accuracy and expressions for …
Multiple Outcomes In Health Services Research: Hypothesis Tests And Power, Donald C. Martin, Paula Diehr, Thomas D. Koepsell, Stephan D. Fihn
Multiple Outcomes In Health Services Research: Hypothesis Tests And Power, Donald C. Martin, Paula Diehr, Thomas D. Koepsell, Stephan D. Fihn
UW Biostatistics Working Paper Series
Health services research often is directed towards making small improvements in a number of outcomes that reflect many aspects of the patient’s life rather than a large improvement in a single well defined outcome. A researcher might choose five scales to measure different aspects of treatment outcomes and not expect any large treatment differences on any single outcome measure. O’Brien (1984) has proposed a nonparametric statistical procedure which is particularly well suited to this type of problem and that can result in considerable increases in statistical power. This paper will briefly review O’Brien’s pooled rank method and develop power calculations. …
Pooling Community Data For Community Interventions When The Number Of Pairs Is Small, Paula Diehr, Ted Lystig, Holly Andrilla, Ziding Feng
Pooling Community Data For Community Interventions When The Number Of Pairs Is Small, Paula Diehr, Ted Lystig, Holly Andrilla, Ziding Feng
UW Biostatistics Working Paper Series
There is considerable interest in community interventions for health promotion, where the community is the experimental unit. Because such interventions are expensive, the number of experimental units (communities) is usually very small, yielding a study with low power. We examined the ability of a process known as “pooling” or “preliminary significance testing” to improve the power of community variations. In this process, one first tests whether there is significant community variation, using type 1 error of perhaps 0.25. If there is significant variation, the usual community-level test is performed. If not, a person-level test is performed. We found through Monte …
An Empirical Study Of Small-Area Variation For Icd-9 Surgical Procedures, Paula Diehr, Kevin Cain, Zhan Ye, John Loeser
An Empirical Study Of Small-Area Variation For Icd-9 Surgical Procedures, Paula Diehr, Kevin Cain, Zhan Ye, John Loeser
UW Biostatistics Working Paper Series
Objective. Several measures of variation have been used in SAVA. One study of DRGs found that the coefficient of variation from analysis of variance (CVA) had superior performance. That work is replicated here for ICD-9 surgical procedures, and extended to age/sex-standardized rates. Results are compared with those in the literature, and recommendations are made for assessing small-area variation in future studies.
Data Sources. Data were taken from Washington State's "Episode of Illness" file of hospital discharges in the State in 1987. Up to three ICD-9 surgical procedures and a unique patient identifier were available for each discharge.
Study Design. We …
Breaking The Matches In A Paired T-Test For Community Interventions When The Number Of Pairs Is Small, Paula Diehr, Don C. Martin, Thomas D. Koepsell, Allen D. Cheadle
Breaking The Matches In A Paired T-Test For Community Interventions When The Number Of Pairs Is Small, Paula Diehr, Don C. Martin, Thomas D. Koepsell, Allen D. Cheadle
UW Biostatistics Working Paper Series
There is considerable interest in community interventions for health promotion, where the community is the experimental unit. Because such interventions are expensive, the number of experimental units (communities) is usually small. Because of the small number of communities involved, investigators often match treatment and control communities on demographic variables before randomization to minimize the possibility of a bad split. Unfortunately, matching has been shown to decrease the power of the design when the number of pairs is small, unless the matching variable is very highly correlated with the outcome variable (in this case, with change in the health behavior). We …
The Multiple Admission Factor (Maf) In Small Area Variation Analysis, Kevin Cain, Paula Diehr
The Multiple Admission Factor (Maf) In Small Area Variation Analysis, Kevin Cain, Paula Diehr
UW Biostatistics Working Paper Series
Small area variation analysis are often based on area-level data such as the total number of hospital admissions within an area, rather than person-level data. Such analysis often make the assumption that the number of admissions within a small area follow a Poisson distribution. This may not be a reasonable assumption when multiple admissions per person are possible. In this case, the multiple admission factor (MAF) can be used to adjust for the extra variance introduced by multiple admissions. In this article, data from Washington State are used to estimate the multiple admission rate and the MAF for each modifed …
Regression Models For Bivariate Binary Responses, Juni Palmgren
Regression Models For Bivariate Binary Responses, Juni Palmgren
UW Biostatistics Working Paper Series
We discuss maximum likelihood inference for the bivariate logistic model, specified in terms of the marginal logits and the log odds ratio. Using the exponential family nonlinear model formulation the model fitting can be done in GLIM. The procedure is illustrated by modelling survival of unilateral and bilateral total hip arthroplasties as function of patient specific and hip specific covariates. We compare maximum likelihood inference with inference obtained from solving likelihood equations under the assumption of within block independence and using robust standard errors for the estimates. Simulations indicate that the latter procedure is effcient for block specific covariates but …
Sample Size Calculations And Optimal Followup Time In Health Services Research Using Utilization Rates, Paula Diehr
Sample Size Calculations And Optimal Followup Time In Health Services Research Using Utilization Rates, Paula Diehr
UW Biostatistics Working Paper Series
It is not always possible to estimate the sample sizes needed in health services research because special formulas are needed, and the necessary data may not be available to use in the formulas. We provide some useful formulas for the sample size required in comparing the means of two groups. These include the special case where the two groups are not of equal size either because one is known to have a higher variability or because one group has already been chosen and its size is thus fixed. We also explore the relationship of the mean to the standard deviation …
Statistical Measures For Admission Rates, Paula Diehr
Statistical Measures For Admission Rates, Paula Diehr
UW Biostatistics Working Paper Series
Hospital admission rates are often shown and interpreted without consideration of their inherent variability, which may lead to faulty conclusions. This may be because theoretically correct variance estimates are not known for the type of estimates usually used; i.e., total admissions divided by total person-months of observation. Here, correct methods for testing and estimation are shown for situations where they exist. For other types of data, approximate procedures are proposed and their properties examined theoretically and empirically, yielding recommendations for exact and approximate estimation and testing methods for admission rates in common situations.