Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Statistical Methodology (99)
- Statistical Theory (74)
- Medicine and Health Sciences (65)
- Statistical Models (51)
- Public Health (44)
-
- Epidemiology (32)
- Survival Analysis (28)
- Clinical Trials (26)
- Life Sciences (26)
- Genetics and Genomics (23)
- Multivariate Analysis (23)
- Genetics (18)
- Microarrays (14)
- Bioinformatics (13)
- Computational Biology (13)
- Diseases (13)
- Disease Modeling (11)
- Longitudinal Data Analysis and Time Series (11)
- Applied Statistics (10)
- Clinical Epidemiology (10)
- Medical Specialties (9)
- Categorical Data Analysis (7)
- Design of Experiments and Sample Surveys (7)
- Applied Mathematics (6)
- Laboratory and Basic Science Research (6)
- Numerical Analysis and Computation (6)
- Social and Behavioral Sciences (6)
- Keyword
-
- Causal inference (16)
- Cross-validation (13)
- Targeted maximum likelihood estimation (12)
- Efficient influence curve (11)
- Genetics (10)
-
- Influence curve (10)
- Causal effect (8)
- Longitudinal data (8)
- Super-learning (8)
- Asymptotic linearity (7)
- Confounding (7)
- Empirical process (7)
- Measurement error (7)
- Missing data (7)
- Survival analysis (7)
- Biomarker (6)
- Interaction (6)
- Inverse probability weighting (6)
- Pathwise differentiable parameter (6)
- Semiparametric statistical model (6)
- Variable selection (6)
- Efficient estimator (5)
- Functional data analysis (5)
- Mediation (5)
- Optimal dynamic treatment (5)
- Sensitivity (5)
- Asymptotic linear estimator (4)
- Asymptotic linearity of an estimator (4)
- Biostatistics (4)
- Canonical gradient (4)
- Publication Year
- Publication
-
- Harvard University Biostatistics Working Paper Series (140)
- U.C. Berkeley Division of Biostatistics Working Paper Series (118)
- UW Biostatistics Working Paper Series (102)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (69)
- The University of Michigan Department of Biostatistics Working Paper Series (55)
Articles 541 - 567 of 567
Full-Text Articles in Biostatistics
Statistical Inference For Variable Importance, Mark J. Van Der Laan
Statistical Inference For Variable Importance, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Many statistical problems involve the learning of an importance/effect of a variable for predicting an outcome of interest based on observing a sample of n independent and identically distributed observations on a list of input variables and an outcome. For example, though prediction/machine learning is, in principle, concerned with learning the optimal unknown mapping from input variables to an outcome from the data, the typical reported output is a list of importance measures for each input variable. The typical approach in prediction has been to learn the unknown optimal predictor from the data and derive, for each of the input …
Semiparametric Inferences For Association With Semi-Competing Risks Data, Debashis Ghosh
Semiparametric Inferences For Association With Semi-Competing Risks Data, Debashis Ghosh
The University of Michigan Department of Biostatistics Working Paper Series
In many biomedical studies, it is of interest to assess dependence between bivariate failure time data. We focus here on a special type of such data, referred to as semi-competing risks data. In this article, we develop methods for making inferences regarding dependence of semi-competing risks data across strata of a discrete covariate Z. A class of rank statistics for testing constancy of association across strata are proposed; its asymptotic properties are also derived. We develop a novel resampling-based technique for calculating the variances of the proposed test statistics. In addition, we develop methods for combining test statistics for assessing …
Simultaneous Estimation Procedures And Multiple Testing: A Decision-Theoretic Framework, Debashis Ghosh
Simultaneous Estimation Procedures And Multiple Testing: A Decision-Theoretic Framework, Debashis Ghosh
The University of Michigan Department of Biostatistics Working Paper Series
There is recent tremendous interest in statistical methods regarding the false discovery rate (FDR). Two classes of literature on this topic exist. In the first, authors have proposed sequential testing procedures that control the false discovery rate. For the second, authors have studied the procedures involving FDR in a univariate mixture model setting. We consider a decision-theoretic approach to the assessment of FDR-based methods. In particular, we attempt to reconcile the current literature on false discovery rate procedures with more classical simultaneous estimation procedures. Formulation of the link will allow us to apply results from decision theory; we can then …
Shrunken P-Values For Assessing Differential Expression, With Applications To Genomic Data Analysis, Debashis Ghosh
Shrunken P-Values For Assessing Differential Expression, With Applications To Genomic Data Analysis, Debashis Ghosh
The University of Michigan Department of Biostatistics Working Paper Series
n many scientific problems involving high-throughput technology, inference must be made involving several hundreds or thousands of hypotheses. Recent attention has focused on how to address the multiple testing issue; much focus has been devoted towards use of the false discovery rate. In this article, we consider an alternative estimation procedure titled shrunken p-values for assessing differential expression (SPADE). The estimators are motivated by risk considerations from decision theory and lead to a completely new method for adjustment in the multiple testing problem. Some theoretical results are outlined. The proposed methodology is illustrated using simulation studies and with application to …
Application Of A Multiple Testing Procedure Controlling The Proportion Of False Positives To Protein And Bacterial Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan
Application Of A Multiple Testing Procedure Controlling The Proportion Of False Positives To Protein And Bacterial Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Simultaneously testing multiple hypotheses is important in high-dimensional biological studies. In these situations, one is often interested in controlling the Type-I error rate, such as the proportion of false positives to total rejections (TPPFP) at a specific level, alpha. This article will present an application of the E-Bayes/Bootstrap TPPFP procedure, presented in van der Laan et al. (2005), which controls the tail probability of the proportion of false positives (TPPFP), on two biological datasets. The two data applications include firstly, the application to a mass-spectrometry dataset of two leukemia subtypes, AML and ALL. The protein data measurements include intensity and …
Test Statistics Null Distributions In Multiple Testing: Simulation Studies And Applications To Genomics, Katherine S. Pollard, Merrill D. Birkner, Mark J. Van Der Laan, Sandrine Dudoit
Test Statistics Null Distributions In Multiple Testing: Simulation Studies And Applications To Genomics, Katherine S. Pollard, Merrill D. Birkner, Mark J. Van Der Laan, Sandrine Dudoit
U.C. Berkeley Division of Biostatistics Working Paper Series
Multiple hypothesis testing problems arise frequently in biomedical and genomic research, for instance, when identifying differentially expressed or co-expressed genes in microarray experiments. We have developed generally applicable resampling-based single-step and stepwise multiple testing procedures (MTP) for control of a broad class of Type I error rates, defined as tail probabilities and expected values for arbitrary functions of the numbers of false positives and rejected hypotheses (Dudoit and van der Laan, 2005; Dudoit et al., 2004a,b; Pollard and van der Laan, 2004; van der Laan et al., 2005, 2004a,b). As argued in the early article of Pollard and van der …
Linear Regression Of Censored Length-Biased Lifetimes, Ying Qing Chen, Yan Wang
Linear Regression Of Censored Length-Biased Lifetimes, Ying Qing Chen, Yan Wang
UW Biostatistics Working Paper Series
Length-biased lifetimes may be collected in observational studies or sample surveys due to biased sampling scheme. In this article, we use a linear regression model, namely, the accelerated failure time model, for the population lifetime distributions in regression analysis of the length-biased lifetimes. It is discovered that the associated regression parameters are invariant under the length-biased sampling scheme. According to this discovery, we propose the quasi partial score estimating equations to estimate the population regression parameters. The proposed methodologies are evaluated and demonstrated by simulation studies and an application to actual data set.
On Additive Regression Of Expectancy, Ying Qing Chen
On Additive Regression Of Expectancy, Ying Qing Chen
UW Biostatistics Working Paper Series
Regression models have been important tools to study the association between outcome variables and their covariates. The traditional linear regression models usually specify such an association by the expectations of the outcome variables as function of the covariates and some parameters. In reality, however, interests often focus on their expectancies characterized by the conditional means. In this article, a new class of additive regression models is proposed to model the expectancies. The model parameters carry practical implication, which may allow the models to be useful in applications such as treatment assessment, resource planning or short-term forecasting. Moreover, the new model …
A Partial Likelihood For Spatio-Temporal Point Processes, Peter J. Diggle
A Partial Likelihood For Spatio-Temporal Point Processes, Peter J. Diggle
Johns Hopkins University, Dept. of Biostatistics Working Papers
Spatio-temporal point process data arise in many fields of application. An intuitively natural way to specify a model for a spatio-temporal point process is through its conditional intensity at location x and time t, given the history of the process up to time t. Typically, this results in an analytically intractable likelihood. Likelihood-based inference therefore relies on Monte Carlo methods which are computationally intensive and require careful tuning to each application. We propose a partial likelihood alternative which is computationally straightforward and can be applied routinely. We apply the method to data from the 2001 foot-and-mouth epidemic in the UK, …
Polydesigns And Causal Inference, Fan Li, Constantine E. Frangakis
Polydesigns And Causal Inference, Fan Li, Constantine E. Frangakis
Johns Hopkins University, Dept. of Biostatistics Working Papers
In an increasingly common class of studies, the goal is to evaluate causal effects of treatments that are only partially controlled by the investigator. In such studies there are two conflicting features: (1) a model on the full cohort design and data can identify the causal effects of interest, but can be sensitive to extreme regions of that design's data, where model specification can have more impact; and (2) models on a reduced design (i.e., a subset of the full data), e.g., conditional likelihood on matched subsets of data, can avoid such sensitivity, but do not generally identify the causal …
A Linear Regression Framework For Receiver Operating Characteristic(Roc) Curve Analysis, Zheng Zhang, Margaret S. Pepe
A Linear Regression Framework For Receiver Operating Characteristic(Roc) Curve Analysis, Zheng Zhang, Margaret S. Pepe
UW Biostatistics Working Paper Series
In the field of medical diagnostic testing, the receiver operating characteristics(ROC) curve has long been used as a standard statistical tool to assess the accuracy of tests that yield continuous results. Although previous research in this area focused mostly on estimating the ROC curve, recently it has been recognized that the accuracy of a given test may fluctuate depending on certain factors, which motivates modelling covariate effects on the ROC curve. Comparing the corresponding ROC curves between two or more tests is a special case of covariate effect modelling. In this manuscript, we introduce a linear regression framework to model …
New Statistical Paradigms Leading To Web-Based Tools For Clinical/Translational Science, Knut M. Wittkowski
New Statistical Paradigms Leading To Web-Based Tools For Clinical/Translational Science, Knut M. Wittkowski
COBRA Preprint Series
As the field of functional genetics and genomics is beginning to mature, we become confronted with new challenges. The constant drop in price for sequencing and gene expression profiling as well as the increasing number of genetic and genomic variables that can be measured makes it feasible to address more complex questions. The success with rare diseases caused by single loci or genes has provided us with a proof-of-concept that new therapies can be developed based on functional genomics and genetics.
Common diseases, however, typically involve genetic epistasis, genomic pathways, and proteomic pattern. Moreover, to better understand the underlying biologi-cal …
Estimating Function Based Cross-Validation And Learning, Mark J. Van Der Laan, Daniel Rubin
Estimating Function Based Cross-Validation And Learning, Mark J. Van Der Laan, Daniel Rubin
U.C. Berkeley Division of Biostatistics Working Paper Series
Suppose that we observe a sample of independent and identically distributed realizations of a random variable. Given a model for the data generating distribution, assume that the parameter of interest can be characterized as the parameter value which makes the population mean of a possibly infinite dimensional estimating function equal to zero. Given a collection of candidate estimators of this parameter, and specification of the vector estimating function, we propose cross-validation criteria for selecting among these estimators. This cross-validation criteria is defined as the Euclidean norm of the empirical mean over the validation sample of the estimating function at the …
Multiple Imputation For Correcting Verification Bias, Ofer Harel, Xiao-Hua Zhou
Multiple Imputation For Correcting Verification Bias, Ofer Harel, Xiao-Hua Zhou
UW Biostatistics Working Paper Series
In the case in which all subjects are screened using a common test, and only a subset of these subjects are tested using a golden standard test, it is well documented that there is a risk for bias, called verification bias. When the test has only two levels (e.g. positive and negative) and we are trying to estimate the sensitivity and specificity of the test, one is actually constructing a confidence interval for a binomial proportion. Since it is well documented that this estimation is not trivial even with complete data, we adopt Multiple imputation (MI) framework for verification bias …
Causal Inference In Longitudinal Studies With History-Restricted Marginal Structural Models, Romain Neugebauer, Mark J. Van Der Laan, Ira B. Tager
Causal Inference In Longitudinal Studies With History-Restricted Marginal Structural Models, Romain Neugebauer, Mark J. Van Der Laan, Ira B. Tager
U.C. Berkeley Division of Biostatistics Working Paper Series
Causal Inference based on Marginal Structural Models (MSMs) is particularly attractive to subject-matter investigators because MSM parameters provide explicit representations of causal effects. We introduce History-Restricted Marginal Structural Models (HRMSMs) for longitudinal data for the purpose of defining causal parameters which may often be better suited for Public Health research. This new class of MSMs allows investigators to analyze the causal effect of a treatment on an outcome based on a fixed, shorter and user-specified history of exposure compared to MSMs. By default, the latter represents the treatment causal effect of interest based on a treatment history defined by the …
The Sensitivity And Specificity Of Markers For Event Times, Tianxi Cai, Margaret S. Pepe, Thomas Lumley, Yingye Zheng, Nancy Swords Jenny
The Sensitivity And Specificity Of Markers For Event Times, Tianxi Cai, Margaret S. Pepe, Thomas Lumley, Yingye Zheng, Nancy Swords Jenny
Harvard University Biostatistics Working Paper Series
No abstract provided.
New Confidence Intervals For The Difference Between Two Sensitivities At A Fixed Level Of Specificity, Gengsheng Qin, Yu-Sheng Hsu, Xiao-Hua Zhou
New Confidence Intervals For The Difference Between Two Sensitivities At A Fixed Level Of Specificity, Gengsheng Qin, Yu-Sheng Hsu, Xiao-Hua Zhou
UW Biostatistics Working Paper Series
For two continuous-scale diagnostic tests, it is of interest to compare their sensitivities at a predetermined level of specificity. In this paper we propose three new intervals for the difference between two sensitivities at a fixed level of specificity. These intervals are easy to compute. We also conduct simulation studies to compare the relative performance of the new intervals with the existing normal approximation based interval proposed by Wieand et al (1989). Our simulation results show that the newly proposed intervals perform better than the existing normal approximation based interval in terms of coverage accuracy and interval length.
A Causal Inference Approach For Constructing Transcriptional Regulatory Networks, Biao Xing, Mark J. Van Der Laan
A Causal Inference Approach For Constructing Transcriptional Regulatory Networks, Biao Xing, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Transcriptional regulatory networks specify the interactions among regulatory genes and between regulatory genes and their target genes. Discovering transcriptional regulatory networks helps us to understand the underlying mechanism of complex cellular processes and responses. In this paper, we describe a causal inference approach for constructing transcriptional regulatory networks using gene expression data, promoter sequences and information on transcription factor binding sites. The method rst identies active transcription factors under each individual experiment using a feature selection approach similar to Bussemaker et al. (2001), Keles et al. (2002) and Conlon et al. (2003). Transcription factors are viewed as `treatments' and gene …
Selective Multiple Imputation Of Keys For Statistical Disclosure Control In Microdata, Rod Little, Fang Liu
Selective Multiple Imputation Of Keys For Statistical Disclosure Control In Microdata, Rod Little, Fang Liu
The University of Michigan Department of Biostatistics Working Paper Series
The fundamental tension in statistical disclosure control (SDC) of microdata is the trade-off between the protection of individual respondents and the release of enough information for statistical inferences. We consider microdata that include key variables that contain identifying information and target variables that include sensitive information. Releasing the original data may expose some individuals in the sample to high risk of disclosure; deleting key variables is a common approach, but this loses information for some statistical analysis. This paper proposes selective multiple imputation of key variables (SMIKe) as an alternative SDC technique between those two extremes, and applies SMIKe to …
Multiple Outcomes In Health Services Research: Hypothesis Tests And Power, Donald C. Martin, Paula Diehr, Thomas D. Koepsell, Stephan D. Fihn
Multiple Outcomes In Health Services Research: Hypothesis Tests And Power, Donald C. Martin, Paula Diehr, Thomas D. Koepsell, Stephan D. Fihn
UW Biostatistics Working Paper Series
Health services research often is directed towards making small improvements in a number of outcomes that reflect many aspects of the patient’s life rather than a large improvement in a single well defined outcome. A researcher might choose five scales to measure different aspects of treatment outcomes and not expect any large treatment differences on any single outcome measure. O’Brien (1984) has proposed a nonparametric statistical procedure which is particularly well suited to this type of problem and that can result in considerable increases in statistical power. This paper will briefly review O’Brien’s pooled rank method and develop power calculations. …
Pooling Community Data For Community Interventions When The Number Of Pairs Is Small, Paula Diehr, Ted Lystig, Holly Andrilla, Ziding Feng
Pooling Community Data For Community Interventions When The Number Of Pairs Is Small, Paula Diehr, Ted Lystig, Holly Andrilla, Ziding Feng
UW Biostatistics Working Paper Series
There is considerable interest in community interventions for health promotion, where the community is the experimental unit. Because such interventions are expensive, the number of experimental units (communities) is usually very small, yielding a study with low power. We examined the ability of a process known as “pooling” or “preliminary significance testing” to improve the power of community variations. In this process, one first tests whether there is significant community variation, using type 1 error of perhaps 0.25. If there is significant variation, the usual community-level test is performed. If not, a person-level test is performed. We found through Monte …
An Empirical Study Of Small-Area Variation For Icd-9 Surgical Procedures, Paula Diehr, Kevin Cain, Zhan Ye, John Loeser
An Empirical Study Of Small-Area Variation For Icd-9 Surgical Procedures, Paula Diehr, Kevin Cain, Zhan Ye, John Loeser
UW Biostatistics Working Paper Series
Objective. Several measures of variation have been used in SAVA. One study of DRGs found that the coefficient of variation from analysis of variance (CVA) had superior performance. That work is replicated here for ICD-9 surgical procedures, and extended to age/sex-standardized rates. Results are compared with those in the literature, and recommendations are made for assessing small-area variation in future studies.
Data Sources. Data were taken from Washington State's "Episode of Illness" file of hospital discharges in the State in 1987. Up to three ICD-9 surgical procedures and a unique patient identifier were available for each discharge.
Study Design. We …
Breaking The Matches In A Paired T-Test For Community Interventions When The Number Of Pairs Is Small, Paula Diehr, Don C. Martin, Thomas D. Koepsell, Allen D. Cheadle
Breaking The Matches In A Paired T-Test For Community Interventions When The Number Of Pairs Is Small, Paula Diehr, Don C. Martin, Thomas D. Koepsell, Allen D. Cheadle
UW Biostatistics Working Paper Series
There is considerable interest in community interventions for health promotion, where the community is the experimental unit. Because such interventions are expensive, the number of experimental units (communities) is usually small. Because of the small number of communities involved, investigators often match treatment and control communities on demographic variables before randomization to minimize the possibility of a bad split. Unfortunately, matching has been shown to decrease the power of the design when the number of pairs is small, unless the matching variable is very highly correlated with the outcome variable (in this case, with change in the health behavior). We …
The Multiple Admission Factor (Maf) In Small Area Variation Analysis, Kevin Cain, Paula Diehr
The Multiple Admission Factor (Maf) In Small Area Variation Analysis, Kevin Cain, Paula Diehr
UW Biostatistics Working Paper Series
Small area variation analysis are often based on area-level data such as the total number of hospital admissions within an area, rather than person-level data. Such analysis often make the assumption that the number of admissions within a small area follow a Poisson distribution. This may not be a reasonable assumption when multiple admissions per person are possible. In this case, the multiple admission factor (MAF) can be used to adjust for the extra variance introduced by multiple admissions. In this article, data from Washington State are used to estimate the multiple admission rate and the MAF for each modifed …
Regression Models For Bivariate Binary Responses, Juni Palmgren
Regression Models For Bivariate Binary Responses, Juni Palmgren
UW Biostatistics Working Paper Series
We discuss maximum likelihood inference for the bivariate logistic model, specified in terms of the marginal logits and the log odds ratio. Using the exponential family nonlinear model formulation the model fitting can be done in GLIM. The procedure is illustrated by modelling survival of unilateral and bilateral total hip arthroplasties as function of patient specific and hip specific covariates. We compare maximum likelihood inference with inference obtained from solving likelihood equations under the assumption of within block independence and using robust standard errors for the estimates. Simulations indicate that the latter procedure is effcient for block specific covariates but …
Sample Size Calculations And Optimal Followup Time In Health Services Research Using Utilization Rates, Paula Diehr
Sample Size Calculations And Optimal Followup Time In Health Services Research Using Utilization Rates, Paula Diehr
UW Biostatistics Working Paper Series
It is not always possible to estimate the sample sizes needed in health services research because special formulas are needed, and the necessary data may not be available to use in the formulas. We provide some useful formulas for the sample size required in comparing the means of two groups. These include the special case where the two groups are not of equal size either because one is known to have a higher variability or because one group has already been chosen and its size is thus fixed. We also explore the relationship of the mean to the standard deviation …
Statistical Measures For Admission Rates, Paula Diehr
Statistical Measures For Admission Rates, Paula Diehr
UW Biostatistics Working Paper Series
Hospital admission rates are often shown and interpreted without consideration of their inherent variability, which may lead to faulty conclusions. This may be because theoretically correct variance estimates are not known for the type of estimates usually used; i.e., total admissions divided by total person-months of observation. Here, correct methods for testing and estimation are shown for situations where they exist. For other types of data, approximate procedures are proposed and their properties examined theoretically and empirically, yielding recommendations for exact and approximate estimation and testing methods for admission rates in common situations.