Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- COBRA (331)
- Central Bank of Nigeria (25)
- University of Kentucky (9)
- Georgia Southern University (5)
- Southern Methodist University (4)
-
- Stephen F. Austin State University (4)
- University of Arkansas, Fayetteville (3)
- University of Louisville (3)
- University of North Florida (3)
- Virginia Commonwealth University (3)
- California Polytechnic State University, San Luis Obispo (2)
- Claremont Colleges (2)
- Marshall University (2)
- University of Central Florida (2)
- University of Nebraska - Lincoln (2)
- University of New Mexico (2)
- Utah State University (2)
- Chapman University (1)
- City University of New York (CUNY) (1)
- Clemson University (1)
- Department of Primary Industries and Regional Development, Western Australia (1)
- East Tennessee State University (1)
- Florida Institute of Technology (1)
- Kennesaw State University (1)
- Louisiana State University (1)
- Michigan Technological University (1)
- Old Dominion University (1)
- Portland State University (1)
- Rochester Institute of Technology (1)
- South Dakota State University (1)
- Keyword
-
- Prediction (13)
- Bootstrap (12)
- Causal inference (11)
- Genetics (11)
- Model selection (11)
-
- Statistics (9)
- Type I error rate (9)
- Adjusted p-value (8)
- Counterfactual (8)
- Cross-validation (8)
- Multiple testing (8)
- False discovery rate (7)
- Gene expression (7)
- Censored data (6)
- Classification (6)
- Counting process (6)
- Estimating equation (6)
- Loss function (6)
- Null distribution (6)
- Survival analysis (6)
- Asymptotic control (5)
- Current status data (5)
- Density estimation (5)
- Diagnostic tests (5)
- Generalized family-wise error rate (5)
- Longitudinal data (5)
- Microarray (5)
- Multiple hypothesis testing (5)
- Proportion of false positives (5)
- Regression (5)
- Publication Year
- Publication
-
- U.C. Berkeley Division of Biostatistics Working Paper Series (113)
- Harvard University Biostatistics Working Paper Series (73)
- UW Biostatistics Working Paper Series (55)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (42)
- CBN Journal of Applied Statistics (JAS) (25)
-
- The University of Michigan Department of Biostatistics Working Paper Series (23)
- COBRA Preprint Series (20)
- Theses and Dissertations--Statistics (9)
- Electronic Theses and Dissertations (7)
- College of Graduate Studies: Theses & Dissertations (5)
- Theses and Dissertations (4)
- Graduate Theses and Dissertations (3)
- Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series (3)
- Statistical Science Theses and Dissertations (3)
- UNF Graduate Theses and Dissertations (3)
- Data Science and Data Mining (2)
- Statistics (2)
- Theses, Dissertations and Capstones (2)
- All Dissertations (1)
- All Graduate Theses and Dissertations, Spring 1920 to Summer 2023 (1)
- All other publications (1)
- Articles (1)
- Bioconductor Project Working Papers (1)
- Branch Mathematics and Statistics Faculty and Staff Publications (1)
- Business and Economics Honors Papers (1)
- CHIP Documents (1)
- CMC Senior Theses (1)
- Community & Environmental Health Faculty Publications (1)
- Department of Management: Faculty Publications (1)
- Department of Statistics: Dissertations, Theses, and Student Research (1)
- Publication Type
Articles 391 - 420 of 426
Full-Text Articles in Statistical Theory
Supervised Detection Of Regulatory Motifs In Dna Sequences, Sunduz Keles, Mark J. Van Der Laan, Sandrine Dudoit, Biao Xing, Michael B. Eisen
Supervised Detection Of Regulatory Motifs In Dna Sequences, Sunduz Keles, Mark J. Van Der Laan, Sandrine Dudoit, Biao Xing, Michael B. Eisen
U.C. Berkeley Division of Biostatistics Working Paper Series
Identification of transcription factor binding sites (regulatory motifs) is a major interest in contemporary biology. We propose a new likelihood based method, COMODE, for identifying structural motifs in DNA sequences. Commonly used methods (e.g. MEME, Gibbs sampler) model binding sites as families of sequences described by a position weight matrix (PWM) and identify PWMs that maximize the likelihood of observed sequence data under a simple multinomial mixture model. This model assumes that the positions of the PWM correspond to independent multinomial distributions with four cell probabilities. We address supervising the search for DNA binding sites using the information derived from …
Improved Confidence Intervals For The Sensitivity At A Fixed Level Of Specificity Of A Continuous-Scale Diagnostic Test, Xiao-Hua Zhou, Gengsheng Qin
Improved Confidence Intervals For The Sensitivity At A Fixed Level Of Specificity Of A Continuous-Scale Diagnostic Test, Xiao-Hua Zhou, Gengsheng Qin
UW Biostatistics Working Paper Series
For a continuous-scale test, it is an interest to construct a confidence interval for the sensitivity of the diagnostic test at the cut-off that yields a predetermined level of its specificity (eg. 80%, 90%, or 95%). IN this paper we proposed two new intervals for the sensitivity of a continuous-scale diagnostic test at a fixed level of specificity. We then conducted simulation studies to compare the relative performance of these two intervals with the best existing BCa bootstrap interval, proposed by Platt et al. (2000). Our simulation results showed that the newly proposed intervals are better than the BCa bootstrap …
Bootstrap Confidence Intervals For Medical Costs With Censored Observations, Hongyu Jiang, Xiao-Hua Zhou
Bootstrap Confidence Intervals For Medical Costs With Censored Observations, Hongyu Jiang, Xiao-Hua Zhou
UW Biostatistics Working Paper Series
Medical costs data with administratively censored observations often arise in cost-effectiveness studies of treatments for life threatening diseases. Mean of medical costs incurred from the start of a treatment till death or certain timepoint after the implementation of treatment is frequently of interest. In many situations, due to the skewed nature of the cost distribution and non-uniform rate of cost accumulation over time, the currently available normal approximation confidence interval has poor coverage accuracy. In this paper, we proposed a bootstrap confidence interval for the mean of medical costs with censored observations. In simulation studies, we showed that the proposed …
New Intervals For The Difference Between Two Independent Binomial Proportions, Xiao-Hua Zhou, Min Tsao, Gengsheng Qin
New Intervals For The Difference Between Two Independent Binomial Proportions, Xiao-Hua Zhou, Min Tsao, Gengsheng Qin
UW Biostatistics Working Paper Series
In this paper we gave an Edgeworth expansion for the studentized difference of two binomial proportions. We then proposed two new intervals by correcting the skewness in the Edgeworth expansion in a direct and an indirect way. Such the bias-correct confidence intervals are easy to compute, and their coverage probabilities converge to the nominal level at a rate of O(n-½), where n is the size of the combined samples. Our simulation results suggest tat in finite samples the new interval based on the indirect method have the similar performance to the two best existing intervals in terms of coverage accuracy …
A Bootstrap Confidence Interval Procedure For The Treatment Effect Using Propensity Score Subclassification, Wanzhu Tu, Xiao-Hua Zhou
A Bootstrap Confidence Interval Procedure For The Treatment Effect Using Propensity Score Subclassification, Wanzhu Tu, Xiao-Hua Zhou
UW Biostatistics Working Paper Series
In the analysis of observational studies, propensity score subclassification has been shown to be a powerful method for adjusting unbalanced covariates for the purpose of causal inferences. One practical difficulty in carrying out such an analysis is to obtain a correct variance estimate for such inferences, while reducing bias in the estimate of the treatment effect due to an imbalance in the measured covariates. In this paper, we propose a bootstrap procedure for the inferences concerning the average treatment effect; our bootstrap method is based on an extension of Efron’s bias-corrected accelerated (BCa) bootstrap confidence interval to a two-sample problem. …
Semiparametric Regression Models With Missing Data: The Mathematics In The Work Of Robins Et Al., Menggang Yu, Bin Nan
Semiparametric Regression Models With Missing Data: The Mathematics In The Work Of Robins Et Al., Menggang Yu, Bin Nan
The University of Michigan Department of Biostatistics Working Paper Series
This review is an attempt to understand the landmark papers of Robins, Rotnitzky, and Zhao (1994) and Robins and Rotnitzky (1992). We revisit their main results and corresponding proofs using the theory outlined in the monograph by Bickel, Klaassen, Ritov, and Wellner (1993). We also discuss an illustrative example to show the details of applying these theoretical results.
Penalized Spline Nonparametric Mixed Models For Inference About A Finite Population Mean From Two-Stage Samples, Hui Zheng, Rod Little
Penalized Spline Nonparametric Mixed Models For Inference About A Finite Population Mean From Two-Stage Samples, Hui Zheng, Rod Little
The University of Michigan Department of Biostatistics Working Paper Series
Samplers often distrust model-based approaches to survey inference due to concerns about model misspecification when applied to large samples from complex populations. We suggest that the model-based paradigm can work very successfully in survey settings, provided models are chosen that take into account the sample design and avoid strong parametric assumptions. The Horvitz-Thompson (HT) estimator is a simple design-unbiased estimator of the finite population total in probability sampling designs. From a modeling perspective, the HT estimator performs well when the ratios of the outcome values and the inclusion probabilities are exchangeable. When this assumption is not met, the HT estimator …
A Semiparametric Model Selection Criterion With Applications To The Marginal Structural Model, M. Alan Brookhart, Mark J. Van Der Laan
A Semiparametric Model Selection Criterion With Applications To The Marginal Structural Model, M. Alan Brookhart, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Estimators for the parameter of interest in semiparametric models often depend on a guessed model for the nuisance parameter. The choice of the model for the nuisance parameter can affect both the finite sample bias and efficiency of the resulting estimator of the parameter of interest. In this paper we propose a finite sample criterion based on cross validation that can be used to select a nuisance parameter model from a list of candidate models. We show that expected value of this criterion is minimized by the nuisance parameter model that yields the estimator of the parameter of interest with …
Ibd Configuration Transition Matrices And Linkage Score Tests For Unilineal Relative Pairs, Sandrine Dudoit
Ibd Configuration Transition Matrices And Linkage Score Tests For Unilineal Relative Pairs, Sandrine Dudoit
U.C. Berkeley Division of Biostatistics Working Paper Series
Properties of transition matrices between IBD configurations are derived for four general classes of unilineal relative pairs obtained from the grand-parent/ grand-child, half-sib, avuncular, and cousin relationships. In this setting, IBD configurations are defined as orbits of groups acting on a set of inheritance vectors. Properties of the transition matrix between IBD configurations at two linked loci are derived by relating its infinitesimal generator to the adjacency matrix of a quotient graph. The second largest eigenvalue of the infinitesimal generator and its multiplicity are key in determining the form of the transition matrix and of likelihood-based linkage tests such as …
Asymptotic Optimality Of Likelihood Based Cross-Validation, Mark J. Van Der Laan, Sandrine Dudoit, Sunduz Keles
Asymptotic Optimality Of Likelihood Based Cross-Validation, Mark J. Van Der Laan, Sandrine Dudoit, Sunduz Keles
U.C. Berkeley Division of Biostatistics Working Paper Series
Likelihood-based cross-validation is a statistical tool for selecting a density estimate based on n i.i.d. observations from the true density among a collection of candidate density estimators. General examples are the selection of a model indexing a maximum likelihood estimator, and the selection of a bandwidth indexing a nonparametric (e.g. kernel) density estimator. In this article, we establish asymptotic optimality of a general class of likelihood based cross-validation procedures (as indexed by the type of sample splitting used, e.g. V-fold cross-validation), in the sense that the cross-validation selector performs asymptotically as well (w.r.t. to the Kullback-Leibler distance to the true …
Asymptotics Of Cross-Validated Risk Estimation In Estimator Selection And Performance Assessment, Sandrine Dudoit, Mark J. Van Der Laan
Asymptotics Of Cross-Validated Risk Estimation In Estimator Selection And Performance Assessment, Sandrine Dudoit, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Risk estimation is an important statistical question for the purposes of selecting a good estimator (i.e., model selection) and assessing its performance (i.e., estimating generalization error). This article introduces a general framework for cross-validation and derives distributional properties of cross-validated risk estimators in the context of estimator selection and performance assessment. Arbitrary classes of estimators are considered, including density estimators and predictors for both continuous and polychotomous outcomes. Results are provided for general full data loss functions (e.g., absolute and squared error, indicator, negative log density). A broad definition of cross-validation is used in order to cover leave-one-out cross-validation, V-fold …
Semiparametric Receiver Operating Characteristic Analysis To Evaluate Biomarkers For Disease, Tianxi Cai, Margaret S. Pepe
Semiparametric Receiver Operating Characteristic Analysis To Evaluate Biomarkers For Disease, Tianxi Cai, Margaret S. Pepe
UW Biostatistics Working Paper Series
The receiver operating characteristic (ROC) curve is a popular method for characterizing the accuracy of diagnostic tests when test results are not binary. Various methodologies for estimating and comparing ROC curves have been developed. One approach, due to Pepe, uses a parametric regression model with the baseline function specified up to a finite-dimensional parameter. In this article we extend the regression models by allowing arbitrary nonparametric baseline functions. We also provide asymptotic distribution theory and procedures for making statistical inference. We illustrate our approach with dataset from a prostate cancer biomarker study. Simulation studies suggest that the extra flexibility inherent …
Semi-Parametric Regression For The Area Under The Receiver Operating Characteristic Curve, Lori E. Dodd, Margaret S. Pepe
Semi-Parametric Regression For The Area Under The Receiver Operating Characteristic Curve, Lori E. Dodd, Margaret S. Pepe
UW Biostatistics Working Paper Series
Medical advances continue to provide new and potentially better means for detecting disease. Such is true in cancer, for example, where biomarkers are sought for early detection and where improvements in imaging methods may pick up the initial functional and molecular changes associated with cancer development. In other binary classification tasks, computational algorithms such as Neural Networks, Support Vector Machines and Evolutionary Algorithms have been applied to areas as diverse as credit scoring, object recognition, and peptide-binding prediction. Before a classifier becomes an accepted technology, it must undergo rigorous evaluation to determine its ability to discriminate between states. Characterization of …
Checking Assumptions In Latent Class Regression Models Via A Markov Chain Monte Carlo Estimation Approach: An Application To Depression And Socio-Economic Status, Elizabeth Garrett, Richard Miech, Pamela Owens, William W. Eaton, Scott L. Zeger
Checking Assumptions In Latent Class Regression Models Via A Markov Chain Monte Carlo Estimation Approach: An Application To Depression And Socio-Economic Status, Elizabeth Garrett, Richard Miech, Pamela Owens, William W. Eaton, Scott L. Zeger
Johns Hopkins University, Dept. of Biostatistics Working Papers
Latent class regression models are useful tools for assessing associations between covariates and latent variables. However, evaluation of key model assumptions cannot be performed using methods from standard regression models due to the unobserved nature of latent outcome variables. This paper presents graphical diagnostic tools to evaluate whether or not latent class regression models adhere to standard assumptions of the model: conditional independence and non-differential measurement. An integral part of these methods is the use of a Markov Chain Monte Carlo estimation procedure. Unlike standard maximum likelihood implementations for latent class regression model estimation, the MCMC approach allows us to …
Recurrent Events Analysis In The Presence Of Time Dependent Covariates And Dependent Censoring, Maja Miloslavsky, Sunduz Keles, Mark J. Van Der Laan, Steve Butler
Recurrent Events Analysis In The Presence Of Time Dependent Covariates And Dependent Censoring, Maja Miloslavsky, Sunduz Keles, Mark J. Van Der Laan, Steve Butler
U.C. Berkeley Division of Biostatistics Working Paper Series
Recurrent events models have lately received a lot of attention in the literature. The majority of approaches discussed show the consistency of parameter estimates under the assumption that censoring is independent of the recurrent events process of interest conditional on the covariates included into the model. We provide an overview of available recurrent events analysis methods, and present an inverse probability of censoring weighted estimator for the regression parameters in the Andersen-Gill model that is commonly used for recurrent event analysis. This estimator remains consistent under informative censoring if the censoring mechanism is estimated consistently, and generally improves on the …
Construction Of Counterfactuals And The G-Computation Formula, Zhuo Yu, Mark J. Van Der Laan
Construction Of Counterfactuals And The G-Computation Formula, Zhuo Yu, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Robins' causal inference theory assumes existence of treatment specific counterfactual variables so that the observed data augmented by the counterfactual data will satisfy a consistency and a randomization assumption. Gill and Robins [2001] show that the consistency and randomization assumptions do not add any restrictions to the observed data distribution. In particular, they provide a construction of counterfactuals as a function of the observed data distribution. In this paper we provide a construction of counterfactuals as a function of the observed data itself. Our construction provides a new statistical tool for estimation of counterfactual distributions. Robins [1987b] shows that the …
Locally Efficient Estimation With Bivariate Right Censored Data , Christopher M. Quale, Mark J. Van Der Laan, James M. Robins
Locally Efficient Estimation With Bivariate Right Censored Data , Christopher M. Quale, Mark J. Van Der Laan, James M. Robins
U.C. Berkeley Division of Biostatistics Working Paper Series
Estimation for bivariate right censored data is a problem that has had much study over the past 15 years. In this paper we propose a new class of estimators for the bivariate survivor function based on locally efficient estimation. The locally efficient estimator takes bivariate estimators Fn and Gn of the distributions of the time variables T1,T2 and the censoring variables C1,C2, respectively, and maps them to the resulting estimator. If Fn and Gn are consistent estimators of F and G, respectively, then the resulting estimator will be nonparametrically efficient (thus the term ``locally efficient''). However, if either Fn or …
The Analysis Of Placement Values For Evaluating Discriminatory Measures, Margaret S. Pepe, Tianxi Cai
The Analysis Of Placement Values For Evaluating Discriminatory Measures, Margaret S. Pepe, Tianxi Cai
UW Biostatistics Working Paper Series
The idea of using measurements such as biomarkers, clinical data, or molecular biology assays for classification and prediction is popular in modern medicine. The scientific evaluation of such measures includes assessing the accuracy with which they predict the outcome of interest. Receiver operating characteristic curves are commonly used for evaluating the accuracy of diagnostic tests. They can be applied more broadly, indeed to any problem involving classification to two states or populations (D = 0 or D = 1). We show that the ROC curve can be interpreted as a cumulative distribution function for the discriminatory measure Y in the …
Accelerated Hazards Model: Method, Theory And Applications, Ying Qing Chen, Nicholas P. Jewell, Jingrong Yang
Accelerated Hazards Model: Method, Theory And Applications, Ying Qing Chen, Nicholas P. Jewell, Jingrong Yang
U.C. Berkeley Division of Biostatistics Working Paper Series
In an accelerated hazards model, the hazard functions of a failure time are related through the time scale-change, which is often a function of covariates and associated parameters. When the hazard functions have special properties, such as monotonicity in time, the parameters may be clinically meaningful in measuring a treatment effect. This paper reviews methodological and theoretical development of this model. Applications of the accelerated hazards model including sample size calculation in clinical trials, are also explored.
Locally Efficient Estimation Of Regression Parameters Using Current Status Data, Chris Andrews, Mark J. Van Der Laan, James M. Robins
Locally Efficient Estimation Of Regression Parameters Using Current Status Data, Chris Andrews, Mark J. Van Der Laan, James M. Robins
U.C. Berkeley Division of Biostatistics Working Paper Series
In biostatistics applications interest often focuses on the estimation of the distribution of a time-variable T. If one only observes whether or not T exceeds an observed monitoring time C, then the data structure is called current status data, also known as interval censored data, case I. We consider this data structure extended to allow the presence of both time-independent covariates and time-dependent covariate processes that are observed until the monitoring time. We assume that the monitoring process satisfies coarsening at random.
Our goal is to estimate the regression parameter beta of the regression model T = Z*beta+epsilon where the …
Case-Control Current Status Data, Nicholas P. Jewell, Mark J. Van Der Laan
Case-Control Current Status Data, Nicholas P. Jewell, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Current status observation on survival times has recently been widely studied. An extreme form of interval censoring, this data structure refers to situations where the only available information on a survival random variable, T, is whether or not T exceeds a random independent monitoring time C, a binary random variable, Y. To date, nonparametric analyses of current status data have assumed the availability of i.i.d. random samples of the random variable (Y, C), or a similar random sample at each of a set of fixed monitoring times. In many situations, it is useful to consider a case-control sampling scheme. Here, …
Why Prefer Double Robust Estimates? Illustration With Causal Point Treatment Studies, Romain Neugebauer, Mark J. Van Der Laan
Why Prefer Double Robust Estimates? Illustration With Causal Point Treatment Studies, Romain Neugebauer, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
In point treatment marginal structural models with treatment A, outcome Y and covariates W, causal parameters can be estimated under the assumption of no unobserved confounders. Three estimates can be used: the G-computation, Inverse Probability of Treatment Weighted (IPTW) or Double Robust (DR) estimates. The properties of the IPTW and DR estimates are known under an assumption on the treatment mechanism that we name "Experimental Treatment Assignment" (ETA) assumption. We show that the DR estimating function is unbiased when the ETA assumption is violated if the model used to regress Y on A and W is correctly specified. The practical …
Bivariate Current Status Data, Mark J. Van Der Laan, Nicholas P. Jewell
Bivariate Current Status Data, Mark J. Van Der Laan, Nicholas P. Jewell
U.C. Berkeley Division of Biostatistics Working Paper Series
In many applications, it is often of interest to estimate a bivariate distribution of two survival random variables. Complete observation of such random variables is often incomplete. If one only observes whether or not each of the individual survival times exceeds a common observed monitoring time C, then the data structure is referred to as bivariate current status data (Wang and Ding, 2000). For such data, we show that the identifiable part of the joint distribution is represented by three univariate cumulative distribution functions, namely the two marginal cumulative distribution functions, and the bivariate cumulative distribution function evaluated on the …
Current Status Data: Review, Recent Developments And Open Problems, Nicholas P. Jewell, Mark J. Van Der Laan
Current Status Data: Review, Recent Developments And Open Problems, Nicholas P. Jewell, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Researchers working with survival data are by now adept at handling issues associated with incomplete data, particular those associated with various forms of censoring. An extreme form of interval censoring, known as current status observation, refers to situations where the only available information on a survival random variable T is whether or not T exceeds a random independent monitoring time C. This article contains a brief review of the extensive literature on the analysis of current status data, discussing the implications of response-based sampling on these methods. The majority of the paper introduces some recent extensions of these ideas to …
Semiparametric Regression Analysis On Longitudinal Pattern Of Recurrent Gap Times, Ying Qing Chen, Mei-Cheng Wang, Yijian Huang
Semiparametric Regression Analysis On Longitudinal Pattern Of Recurrent Gap Times, Ying Qing Chen, Mei-Cheng Wang, Yijian Huang
U.C. Berkeley Division of Biostatistics Working Paper Series
In longitudinal studies, individual subjects may experience recurrent events of the same type over a relatively long period of time. The longitudinal pattern of the gaps between the successive recurrent events is often of great research interest. In this article, the probability structure of the recurrent gap times is first explored in the presence of censoring. According to the discovered structure, we introduce the proportional reverse-time hazards models with unspecified baseline functions to accommodate heterogeneous individual underlying distributions, when the ongitudinal pattern parameter is of main interest. Inference procedures are proposed and studied by way of proper riskset construction. The …
Multiple Hypothesis Testing In Microarray Experiments, Sandrine Dudoit, Juliet Popper Shaffer, Jennifer C. Boldrick
Multiple Hypothesis Testing In Microarray Experiments, Sandrine Dudoit, Juliet Popper Shaffer, Jennifer C. Boldrick
U.C. Berkeley Division of Biostatistics Working Paper Series
DNA microarrays are a new and promising biotechnology which allows the monitoring of expression levels in cells for thousands of genes simultaneously. An important and common question in microarray experiments is the identification of differentially expressed genes, i.e., genes whose expression levels are associated with a response or covariate of interest. The biological question of differential expression can be restated as a problem in multiple hypothesis testing: the simultaneous test for each gene of the null hypothesis of no association between the expression levels and the responses or covariates. As a typical microarray experiment measures expression levels for thousands of …
Estimation Of The Bivariate Survival Function With Generalized Bivariate Right Censored Data Structures, Sunduz Keles, Mark J. Van Der Laan, James M. Robins
Estimation Of The Bivariate Survival Function With Generalized Bivariate Right Censored Data Structures, Sunduz Keles, Mark J. Van Der Laan, James M. Robins
U.C. Berkeley Division of Biostatistics Working Paper Series
We propose a bivariate survival function estimator for a general right censored data structure that includes a time dependent covariate process. Firstly, an initial estimator that generalizes Dabrowska's (1988) estimator is introduced. We obtain this estimator by a general methodology of constructing estimating functions in censored data models. The initial estimator is guaranteed to improve on Dabrowska's estimator and remains consistent and asymptotically linear under informative censoring schemes if the censoring mechanism is estimated consistently. We then construct an orthogonalized estimating function which results in a more robust and efficient estimator than our initial estimator. A simulation study demonstrates the …
Inference For Proportional Mean Residual Life Model In The Presence Of Censoring, Ying Q. Chen, Nicholas P. Jewell
Inference For Proportional Mean Residual Life Model In The Presence Of Censoring, Ying Q. Chen, Nicholas P. Jewell
U.C. Berkeley Division of Biostatistics Working Paper Series
As a function of time t, mean residual life is defined as remaining life expectancy of a subject given its survival to t. It plays an important role in many research areas to characterise stochastic behavior of survival over time. Similar to the Cox proportional hazard model, the proportional mean residual life model were proposed in statistical literature to study association between the mean residual life and individual subject's explanatory covariates. In this article, we will study this model and develop appropriate inference procedures in presence of censoring. Numerical studies including simulation and real data analysis are presented as well.
A Method To Identify Significant Clusters In Gene Expression Data, Katherine S. Pollard, Mark J. Van Der Laan
A Method To Identify Significant Clusters In Gene Expression Data, Katherine S. Pollard, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Clustering algorithms have been widely applied to gene expression data. For both hierarchical and partitioning clustering algorithms, selecting the number of significant clusters is an important problem and many methods have been proposed. Existing methods for selecting the number of clusters tend to find only the global patterns in the data (e.g.: the over and under expressed genes). We have noted the need for a better method in the gene expression context, where small, biologically meaningful clusters can be difficult to identify. In this paper, we define a new criteria, Mean Split Silhouette (MSS), which is a measure of cluster …
A New Partitioning Around Medoids Algorithm, Mark J. Van Der Laan, Katherine S. Pollard, Jennifer Bryan
A New Partitioning Around Medoids Algorithm, Mark J. Van Der Laan, Katherine S. Pollard, Jennifer Bryan
U.C. Berkeley Division of Biostatistics Working Paper Series
Kaufman & Rousseeuw (1990) proposed a clustering algorithm Partitioning Around Medoids (PAM) which maps a distance matrix into a specified number of clusters. A particularly nice property is that PAM allows clustering with respect to any specified distance metric. In addition, the medoids are robust representations of the cluster centers, which is particularly important in the common context that many elements do not belong well to any cluster. Based on our experience in clustering gene expression data, we have noticed that PAM does have problems recognizing relatively small clusters in situations where good partitions around medoids clearly exist. In this …