Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Biostatistics (567)
- Statistical Methodology (362)
- Statistical Theory (336)
- Statistical Models (242)
- Medicine and Health Sciences (176)
-
- Survival Analysis (147)
- Public Health (142)
- Epidemiology (99)
- Life Sciences (92)
- Longitudinal Data Analysis and Time Series (89)
- Genetics and Genomics (88)
- Clinical Trials (82)
- Microarrays (78)
- Multivariate Analysis (78)
- Applied Mathematics (57)
- Genetics (57)
- Numerical Analysis and Computation (57)
- Bioinformatics (50)
- Computational Biology (50)
- Categorical Data Analysis (49)
- Design of Experiments and Sample Surveys (39)
- Clinical Epidemiology (36)
- Diseases (30)
- Disease Modeling (28)
- Medical Specialties (23)
- Health Services Research (17)
- Applied Statistics (13)
- Vital and Health Statistics (11)
- Keyword
-
- Causal inference (30)
- Cross-validation (25)
- Prediction (23)
- Genetics (21)
- Longitudinal data (19)
-
- Survival analysis (16)
- Classification (14)
- Influence curve (14)
- Model selection (14)
- Sensitivity (14)
- Bootstrap (13)
- Gene expression (13)
- Clinical trials (12)
- Targeted maximum likelihood estimation (12)
- Counterfactual (11)
- Efficient influence curve (11)
- Multiple testing (11)
- Confounding (10)
- Loss function (10)
- Missing data (10)
- Variable selection (10)
- Causal effect (9)
- Estimating equation (9)
- Measurement error (9)
- Regression (9)
- Specificity (9)
- Adjusted p-value (8)
- Air pollution (8)
- Asymptotic linearity (8)
- Censoring (8)
- Publication Year
- Publication
-
- U.C. Berkeley Division of Biostatistics Working Paper Series (242)
- UW Biostatistics Working Paper Series (215)
- Harvard University Biostatistics Working Paper Series (212)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (178)
- The University of Michigan Department of Biostatistics Working Paper Series (111)
Articles 811 - 840 of 1108
Full-Text Articles in Statistics and Probability
A New Approach To Intensity-Dependent Normalization Of Two-Channel Microarrays, Alan R. Dabney, John D. Storey
A New Approach To Intensity-Dependent Normalization Of Two-Channel Microarrays, Alan R. Dabney, John D. Storey
UW Biostatistics Working Paper Series
A two-channel microarray measures the relative expression levels of thousands of genes from a pair of biological samples. In order to reliably compare gene expression levels between and within arrays, it is necessary to remove systematic errors that distort the biological signal of interest. The standard for accomplishing this is smoothing "MA-plots" to remove intensity-dependent dye bias and array-specific effects. However, MA methods require strong assumptions. We review these assumptions and derive several practical scenarios in which they fail. The "dye-swap" normalization method has been much less frequently used because it requires two arrays per pair of samples. We show …
Nonparametric Estimation Of Bivariate Failure Time Associations In The Presence Of A Competing Risk, Karen Bandeen-Roche, Jing Ning
Nonparametric Estimation Of Bivariate Failure Time Associations In The Presence Of A Competing Risk, Karen Bandeen-Roche, Jing Ning
Johns Hopkins University, Dept. of Biostatistics Working Papers
There has been much research on the study of associations among paired failure times. Most has either assumed time invariance of association or been based on complex measures or estimators. Little has accommodated failures arising amid competing risks. This paper targets the conditional cause specific hazard ratio, a recent modification of the conditional hazard ratio to accommodate competing risks data. Estimation is accomplished by an intuitive, nonparametric method that localizes Kendall’s tau. Time variance is accommodated through a partitioning of space into “bins” between which the strength of association may differ. Inferential procedures are researched, small sample performance evaluated, and …
Correspondences Between Regression Models For Complex Binary Outcomes And Those For Structured Multivariate Survival Analyses, Nicholas P. Jewell
Correspondences Between Regression Models For Complex Binary Outcomes And Those For Structured Multivariate Survival Analyses, Nicholas P. Jewell
U.C. Berkeley Division of Biostatistics Working Paper Series
Doksum and Gasko [5] described a one-to-one correspondence between regression models for binary outcomes and those for continuous time survival analyses. This correspondence has been exploited heavily in the analysis of current status data (Jewell and van der Laan [11], Shiboski [18]). Here, we explore similar correspondences for complex survival models and categorical regression models for polytomous data. We include discussion of competing risks and progressive multi-state survival random variables.
Application Of A Variable Importance Measure Method To Hiv-1 Sequence Data, Merrill D. Birkner, Mark J. Van Der Laan
Application Of A Variable Importance Measure Method To Hiv-1 Sequence Data, Merrill D. Birkner, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
van der Laan (2005) proposed a method to construct variable importance measures and provided the respective statistical inference. This technique involves determining the importance of a variable in predicting an outcome. This method can be applied as an inverse probability of treatment weighted (IPTW) or double robust inverse probability of treatment weighted (DR-IPTW) estimator. A respective significance of the estimator is determined by estimating the influence curve and hence determining the corresponding variance and p-value. This article applies the van der Laan (2005) variable importance measures and corresponding inference to HIV-1 sequence data. In this data application, protease and reverse …
Casual Mediation Analyses With Structural Mean Models, Thomas R. Tenhave, Marshall Joffe, Kevin Lynch, Greg Brown, Stephen Maisto
Casual Mediation Analyses With Structural Mean Models, Thomas R. Tenhave, Marshall Joffe, Kevin Lynch, Greg Brown, Stephen Maisto
UPenn Biostatistics Working Papers
We represent a linear structural mean model (SMM)approach for analyzing mediation of a randomized baseline intervention's effect on a univariate follow-up outcome. Unlike standard mediation analyses, our approach does not assume that the mediating factor is randomly assigned to individuals (i.e., sequential ignorability). Hence, a comparison of the results of the proposed and standard approaches in with respect to mediation offers a sensitivity analyses of the sequential ignorability assumption. The G-estimation procedure for the proposed SMM represents an extension of the work on direct effects of randomized treatment effects for survival outcomes by Robins and Greenland (1994) (Section 5.0 and …
A General Imputation Methodology For Nonparametric Regression With Censored Data, Dan Rubin, Mark J. Van Der Laan
A General Imputation Methodology For Nonparametric Regression With Censored Data, Dan Rubin, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
We consider the random design nonparametric regression problem when the response variable is subject to a general mode of missingness or censoring. A traditional approach to such problems is imputation, in which the missing or censored responses are replaced by well-chosen values, and then the resulting covariate/response data are plugged into algorithms designed for the uncensored setting. We present a general methodology for imputation with the property of double robustness, in that the method works well if either a parameter of the full data distribution (covariate and response distribution) or a parameter of the censoring mechanism is well approximated. These …
Estimating A Treatment Effect With Repeated Measurements Accounting For Varying Effectiveness Duration, Ying Qing Chen, Jingrong Yang, Su-Chun Cheng
Estimating A Treatment Effect With Repeated Measurements Accounting For Varying Effectiveness Duration, Ying Qing Chen, Jingrong Yang, Su-Chun Cheng
UW Biostatistics Working Paper Series
To assess treatment efficacy in clinical trials, certain clinical outcomes are repeatedly measured for same subject over time. They can be regarded as function of time. The difference in their mean functions between the treatment arms usually characterises a treatment effect. Due to the potential existence of subject-specific treatment effectiveness lag and saturation times, erosion of treatment effect in the difference may occur during the observation period of time. Instead of using ad hoc parametric or purely nonparametric time-varying coefficients in statistical modeling, we first propose to model the treatment effectiveness durations, which are the varying time intervals between the …
Modeling Differentiated Treatment Effects For Multiple Outcomes Data, Hongfei Guo, Karen Bandeen-Roche
Modeling Differentiated Treatment Effects For Multiple Outcomes Data, Hongfei Guo, Karen Bandeen-Roche
Johns Hopkins University, Dept. of Biostatistics Working Papers
Multiple outcomes data are commonly used to characterize treatment effects in medical research, for instance, multiple symptoms to characterize potential remission of a psychiatric disorder. Often either a global, i.e. symptom-invariant, treatment effect is evaluated. Such a treatment effect may over generalize the effect across the outcomes. On the other hand individual treatment effects, varying across all outcomes, are complicated to interpret, and their estimation may lose precision relative to a global summary. An effective compromise to summarize the treatment effect may be through patterns of the treatment effects, i.e. "differentiated effects." In this paper we propose a two-category model …
Model Evaluation Based On The Distribution Of Estimated Absolute Prediction Error, Lu Tian, Tianxi Cai, Els Goetghebeur, L. J. Wei
Model Evaluation Based On The Distribution Of Estimated Absolute Prediction Error, Lu Tian, Tianxi Cai, Els Goetghebeur, L. J. Wei
Harvard University Biostatistics Working Paper Series
The construction of a reliable, practically useful prediction rule for future response is heavily dependent on the "adequacy" of the fitted regression model. In this article, we consider the absolute prediction error, the expected value of the absolute difference between the future and predicted responses, as the model evaluation criterion. This prediction error is easier to interpret than the average squared error and is equivalent to the mis-classification error for the binary outcome. We show that the distributions of the apparent error and its cross-validation counterparts are approximately normal even under a misspecified fitted model. When the prediction rule is …
Efficacy Studies Of Malaria Treatments In Africa: Efficient Estimation With Missing Indicators Of Failure, Rhoderick N. Machekano, Grant Dorsey, Alan E. Hubbard
Efficacy Studies Of Malaria Treatments In Africa: Efficient Estimation With Missing Indicators Of Failure, Rhoderick N. Machekano, Grant Dorsey, Alan E. Hubbard
U.C. Berkeley Division of Biostatistics Working Paper Series
Efficacy studies of malaria treatments can be plagued by indeterminate outcomes for some patients. The study motivating this paper defines the outcome of interest (treatment failure) as recrudescence and for some subjects, it is unclear whether a recurrence of malaria is due to that or new infection. This results in a specific kind of missing data. The effect of missing data in causal inference problems is widely recognized. Methods that adjust for possible bias from missing data include a variety of imputation procedures (extreme case analysis, hot-deck, single and multiple imputation), inverse weighting methods, and likelihood based methods (data augmentation, …
Analyzing Panel Count Data With Informative Observation Times, Chiung-Yu Huang, Mei-Cheng Wang, Ying Zhang
Analyzing Panel Count Data With Informative Observation Times, Chiung-Yu Huang, Mei-Cheng Wang, Ying Zhang
Johns Hopkins University, Dept. of Biostatistics Working Papers
In this paper, we study panel count data with informative observation times. We assume nonparametric and semiparametric proportional rate models for the underlying recurrent event process, where the form of the baseline rate function is left unspecified and a subject-specific frailty variable inflates or deflates the rate function multiplicatively. The proposed models allow the recurrent event processes and observation times to be correlated through their connections with the unobserved frailty; moreover, the distributions of both the frailty variable and observation times are considered as nuisance parameters. The baseline rate function and the regression parameters are estimated by maximizing a conditional …
A Fine-Scale Linkage Disequilibrium Measure Based On Length Of Haplotype Sharing, Yan Wang, Lue Ping Zhao, Sandrine Dudoit
A Fine-Scale Linkage Disequilibrium Measure Based On Length Of Haplotype Sharing, Yan Wang, Lue Ping Zhao, Sandrine Dudoit
U.C. Berkeley Division of Biostatistics Working Paper Series
High-throughput genotyping technologies for single nucleotide polymorphisms (SNP) have enabled the recent completion of the International HapMap Project (Phase I), which has stimulated much interest in studying genome-wide linkage disequilibrium (LD) patterns. Conventional LD measures, such as D' and r-square, are two-point measurements, and their relationship with physical distance is highly noisy. We propose a new LD measure, defined in terms of the correlation coefficient for shared haplotype lengths around two loci, thereby borrowing information from multiple loci. A U-statistic-based estimator of the new LD measure, which takes into consideration the dependence structure of the observed data, is developed and …
Population Intervention Models In Causal Inference, Alan E. Hubbard, Mark J. Van Der Laan
Population Intervention Models In Causal Inference, Alan E. Hubbard, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Marginal structural models (MSM) provide a powerful tool for estimating the causal effect of a] treatment variable or risk variable on the distribution of a disease in a population. These models, as originally introduced by Robins (e.g., Robins (2000a), Robins (2000b), van der Laan and Robins (2002)), model the marginal distributions of treatment-specific counterfactual outcomes, possibly conditional on a subset of the baseline covariates, and its dependence on treatment. Marginal structural models are particularly useful in the context of longitudinal data structures, in which each subject's treatment and covariate history are measured over time, and an outcome is recorded at …
Gauss-Seidel Estimation Of Generalized Linear Mixed Models With Application To Poisson Modeling Of Spatially Varying Disease Rates, Subharup Guha, Louise Ryan
Gauss-Seidel Estimation Of Generalized Linear Mixed Models With Application To Poisson Modeling Of Spatially Varying Disease Rates, Subharup Guha, Louise Ryan
Harvard University Biostatistics Working Paper Series
Generalized linear mixed models (GLMMs) provide an elegant framework for the analysis of correlated data. Due to the non-closed form of the likelihood, GLMMs are often fit by computational procedures like penalized quasi-likelihood (PQL). Special cases of these models are generalized linear models (GLMs), which are often fit using algorithms like iterative weighted least squares (IWLS). High computational costs and memory space constraints often make it difficult to apply these iterative procedures to data sets with very large number of cases.
This paper proposes a computationally efficient strategy based on the Gauss-Seidel algorithm that iteratively fits sub-models of the GLMM …
Designed Extension Of Survival Studies: Application To Clinical Trials With Unrecognized Heterogeneity, Yi Li, Mei-Chiung Shih, Rebecca A. Betensky
Designed Extension Of Survival Studies: Application To Clinical Trials With Unrecognized Heterogeneity, Yi Li, Mei-Chiung Shih, Rebecca A. Betensky
Harvard University Biostatistics Working Paper Series
It is well known that unrecognized heterogeneity among patients, such as is conferred by genetic subtype, can undermine the power of randomized trial, designed under the assumption of homogeneity, to detect a truly beneficial treatment. We consider the conditional power approach to allow for recovery of power under unexplained heterogeneity. While Proschan and Hunsberger (1995) confined the application of conditional power design to normally distributed observations, we consider more general and difficult settings in which the data are in the framework of continuous time and are subject to censoring. In particular, we derive a procedure appropriate for the analysis of …
Computational Techniques For Spatial Logistic Regression With Large Datasets, Christopher J. Paciorek, Louise Ryan
Computational Techniques For Spatial Logistic Regression With Large Datasets, Christopher J. Paciorek, Louise Ryan
Harvard University Biostatistics Working Paper Series
In epidemiological work, outcomes are frequently non-normal, sample sizes may be large, and effects are often small. To relate health outcomes to geographic risk factors, fast and powerful methods for fitting spatial models, particularly for non-normal data, are required. We focus on binary outcomes, with the risk surface a smooth function of space. We compare penalized likelihood models, including the penalized quasi-likelihood (PQL) approach, and Bayesian models based on fit, speed, and ease of implementation.
A Bayesian model using a spectral basis representation of the spatial surface provides the best tradeoff of sensitivity and specificity in simulations, detecting real spatial …
Is The Number Of Sick Persons In A Cohort Constant Over Time?, Paula Diehr, Ann Derleth, Anne Newman, Liming Cai
Is The Number Of Sick Persons In A Cohort Constant Over Time?, Paula Diehr, Ann Derleth, Anne Newman, Liming Cai
UW Biostatistics Working Paper Series
Objectives: To estimate the number of persons in a cohort who are sick, over time.
Methods: We calculated the number of sick persons in the Cardiovascular Health Study (CHS), a cohort study of older adults followed up to 14 years, using eight definitions of “healthy” and “sick”. We projected the number in each health state over time for a birth cohort.
Results: The number of sick persons in CHS was approximately constant for 14 years, for all definitions of “sick”. The estimated number of sick persons in the birth cohort was approximately constant from ages 55-75, after which it decreased. …
On The Synthesis Of Microarray Experiments, Robert Gentleman, Markus Ruschhaupt, Wolfgang Huber
On The Synthesis Of Microarray Experiments, Robert Gentleman, Markus Ruschhaupt, Wolfgang Huber
Bioconductor Project Working Papers
With many different investigators studying the same disease and with a strong commitment to publish supporting data in the scientific community, there are often many different datasets available for any given disease. Hence there is substantial interest in finding methods for combining these datasets to provide better and more detailed understanding of the underlying biology. We consider the synthesis of different microarray data sets using a random effects paradigm and demonstrate how relatively standard statistical approaches yield good results. We identify a number of important and substantive areas which require further investigation.
Feature-Specific Penalized Latent Class Analysis For Genomic Data, E. Andres Houseman, Brent A. Coull, Rebecca A. Betensky
Feature-Specific Penalized Latent Class Analysis For Genomic Data, E. Andres Houseman, Brent A. Coull, Rebecca A. Betensky
Harvard University Biostatistics Working Paper Series
No abstract provided.
Marginal Regression Modeling Under Irregular, Biased Sampling, Petra Buzkova, Thomas Lumley
Marginal Regression Modeling Under Irregular, Biased Sampling, Petra Buzkova, Thomas Lumley
UW Biostatistics Working Paper Series
In longitudinal studies observations are often obtained at continuous subject-specific times. Frequently the availability of outcome data may be related to the outcome measure or other covariates that are related to the outcome measure. Under such biased sampling designs unadjusted regression analysis yield biased estimates. Building on the work of Lin & Ying (2001) that integrates counting processes techniques with longitudinal data settings we propose a class of estimators that can handle biased sampling. We call those estimators ``inverse--intensity--rate--ratio--weighted'' (IIRR) estimators. Of major focus is a mean--response model where we examine the marginal effect of the covariate X at time …
Longitudinal Data Analysis For Generalized Linear Models Under Irregular, Biased Sampling: Situations With Follow-Up Dependent On Outcome Or Auxiliary Outcome-Related Variables, Petra Buzkova, Thomas Lumley
Longitudinal Data Analysis For Generalized Linear Models Under Irregular, Biased Sampling: Situations With Follow-Up Dependent On Outcome Or Auxiliary Outcome-Related Variables, Petra Buzkova, Thomas Lumley
UW Biostatistics Working Paper Series
In longitudinal studies, observations are often obtained at subject-specific observation times. Those times can be continuous times, not at a set of prespecified times. Frequently the observation times may be related to the outcome measure or other auxiliary variables that are related to the outcome measure but undesirable to condition upon in the regression model for outcome. Regression analysis unadjusted for such sampling designs yield biased estimates. Based on estimating equations, we propose a class of estimators in generalized linear regression models that can handle biased sampling under continuous observation times. We call those estimators ``inverse--intensity rate--ratio--weighted'' (IIRR) estimators. The …
Semiparametric Loglinear Regression For Longitudinal Measurements Subject To Irregular, Biased Follow-Up, Petra Buzkova, Thomas Lumley
Semiparametric Loglinear Regression For Longitudinal Measurements Subject To Irregular, Biased Follow-Up, Petra Buzkova, Thomas Lumley
UW Biostatistics Working Paper Series
We propose a method for analysis of loglinear regression models for longitudinal data that are subject to continuous and irregular follow-up. Frequently, if the follow-up is irregular, the availability of outcome data may be related to the outcome measure or other covariates that are related to the outcome measure. Under such biased sampling designs unadjusted regression analysis yield biased estimates. We examine the marginal association of the covariates X at time t and the logarithm of the mean of response Y at time t. We focus on semiparametric regression with unspecified baseline function of time. To predict the follow-up times …
A Pseudolikelihood Approach For Simultaneous Analysis Of Array Comparative Genomic Hybridizations (Acgh), David A. Engler, Gayatry Mohapatra, David N. Louis, Rebecca Betensky
A Pseudolikelihood Approach For Simultaneous Analysis Of Array Comparative Genomic Hybridizations (Acgh), David A. Engler, Gayatry Mohapatra, David N. Louis, Rebecca Betensky
Harvard University Biostatistics Working Paper Series
DNA sequence copy number has been shown to be associated with cancer development and progression. Array-based Comparative Genomic Hybridization (aCGH) is a recent development that seeks to identify the copy number ratio at large numbers of markers across the genome. Due to experimental and biological variations across chromosomes and across hybridizations, current methods are limited to analyses of single chromosomes. We propose a more powerful approach that borrows strength across chromosomes and across hybridizations. We assume a Gaussian mixture model, with a hidden Markov dependence structure, and with random effects to allow for intertumoral variation, as well as intratumoral clonal …
A Nonstationary Negative Binomial Time Series With Time-Dependent Covariates: Enterococcus Counts In Boston Harbor, E. Andres Houseman, Brent Coull, James P. Shine
A Nonstationary Negative Binomial Time Series With Time-Dependent Covariates: Enterococcus Counts In Boston Harbor, E. Andres Houseman, Brent Coull, James P. Shine
Harvard University Biostatistics Working Paper Series
Boston Harbor has had a history of poor water quality, including contamination by enteric pathogens. We conduct a statistical analysis of data collected by the Massachusetts Water Resources Authority (MWRA) between 1996 and 2002 to evaluate the effects of court-mandated improvements in sewage treatment. Motivated by the ineffectiveness of standard Poisson mixture models and their zero-inflated counterparts, we propose a new negative binomial model for time series of Enterococcus counts in Boston Harbor, where nonstationarity and autocorrelation are modeled using a nonparametric smooth function of time in the predictor. Without further restrictions, this function is not identifiable in the presence …
Cross-Validated Bagged Prediction Of Survival, Sandra E. Sinisi, Romain Neugebauer, Mark J. Van Der Laan
Cross-Validated Bagged Prediction Of Survival, Sandra E. Sinisi, Romain Neugebauer, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
In this article, we show how to apply our previously proposed Deletion/Substitution/Addition algorithm in the context of right-censoring for the prediction of survival. Furthermore, we introduce how to incorporate bagging into the algorithm to obtain a cross-validated bagged estimator. The method is used for predicting the survival time of patients with diffuse large B-cell lymphoma based on gene expression variables.
Semiparametric Estimation In General Repeated Measures Problems, Xihong Lin, Raymond J. Carroll
Semiparametric Estimation In General Repeated Measures Problems, Xihong Lin, Raymond J. Carroll
Harvard University Biostatistics Working Paper Series
This paper considers a wide class of semiparametric problems with a parametric part for some covariate effects and repeated evaluations of a nonparametric function. Special cases in our approach include marginal models for longitudinal/clustered data, conditional logistic regression for matched case-control studies, multivariate measurement error models, generalized linear mixed models with a semiparametric component, and many others. We propose profile-kernel and backfitting estimation methods for these problems, derive their asymptotic distributions, and show that in likelihood problems the methods are semiparametric efficient. While generally not true, with our methods profiling and backfitting are asymptotically equivalent. We also consider pseudolikelihood methods …
The Optimal Discovery Procedure: A New Approach To Simultaneous Significance Testing, John D. Storey
The Optimal Discovery Procedure: A New Approach To Simultaneous Significance Testing, John D. Storey
UW Biostatistics Working Paper Series
Significance testing is one of the main objectives of statistics. The Neyman-Pearson lemma provides a simple rule for optimally testing a single hypothesis when the null and alternative distributions are known. This result has played a major role in the development of significance testing strategies that are used in practice. Most of the work extending single testing strategies to multiple tests has focused on formulating and estimating new types of significance measures, such as the false discovery rate. These methods tend to be based on p-values that are calculated from each test individually, ignoring information from the other tests. As …
Mixture Cure Survival Models With Dependent Censoring, Yi Li, Ram C. Tiwari, Subharup Guha
Mixture Cure Survival Models With Dependent Censoring, Yi Li, Ram C. Tiwari, Subharup Guha
Harvard University Biostatistics Working Paper Series
A number of authors have studies the mixture survival model to analyze survival data with nonnegligible cure fractions. A key assumption made by these authors is the independence between the survival time and the censoring time. To our knowledge, no one has studies the mixture cure model in the presence of dependent censoring. To account for such dependence, we propose a more general cure model which allows for dependent censoring. In particular, we derive the cure models from the perspective of competing risks and model the dependence between the censoring time and the survival time using a class of Archimedean …
Semiparametric Normal Transformation Models For Spatially Correlated Survival Data, Yi Li, Xihong Lin
Semiparametric Normal Transformation Models For Spatially Correlated Survival Data, Yi Li, Xihong Lin
Harvard University Biostatistics Working Paper Series
There is an emerging interest in modeling spatially correlated survival data in biomedical and epidemiological studies. In this paper, we propose a new class of semiparametric normal transformation models for right censored spatially correlated survival data. This class of models assumes that survival outcomes marginally follow a Cox proportional hazard model with unspecified baseline hazard, and their joint distribution is obtained by transforming survival outcomes to normal random variables, whose joint distribution is assumed to be multivariate normal with a spatial correlation structure. A key feature of the class of semiparametric normal transformation models is that it provides a rich …
Inference On Survival Data With Covariate Measurement Error - An Imputation-Based Approach, Yi Li, Louise Ryan
Inference On Survival Data With Covariate Measurement Error - An Imputation-Based Approach, Yi Li, Louise Ryan
Harvard University Biostatistics Working Paper Series
We propose a new method for fitting proportional hazards models with error-prone covariates. Regression coefficients are estimated by solving an estimating equation that is the average of the partial likelihood scores based on imputed true covariates. For the purpose of imputation, a linear spline model is assumed on the baseline hazard. We discuss consistency and asymptotic normality of the resulting estimators, and propose a stochastic approximation scheme to obtain the estimates. The algorithm is easy to implement, and reduces to the ordinary Cox partial likelihood approach when the measurement error has a degenerative distribution. Simulations indicate high efficiency and robustness. …