Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Statistical Methodology (99)
- Statistical Theory (74)
- Medicine and Health Sciences (65)
- Statistical Models (51)
- Public Health (44)
-
- Epidemiology (32)
- Survival Analysis (28)
- Clinical Trials (26)
- Life Sciences (26)
- Genetics and Genomics (23)
- Multivariate Analysis (23)
- Genetics (18)
- Microarrays (14)
- Bioinformatics (13)
- Computational Biology (13)
- Diseases (13)
- Disease Modeling (11)
- Longitudinal Data Analysis and Time Series (11)
- Applied Statistics (10)
- Clinical Epidemiology (10)
- Medical Specialties (9)
- Categorical Data Analysis (7)
- Design of Experiments and Sample Surveys (7)
- Applied Mathematics (6)
- Laboratory and Basic Science Research (6)
- Numerical Analysis and Computation (6)
- Social and Behavioral Sciences (6)
- Keyword
-
- Causal inference (16)
- Cross-validation (13)
- Targeted maximum likelihood estimation (12)
- Efficient influence curve (11)
- Genetics (10)
-
- Influence curve (10)
- Causal effect (8)
- Longitudinal data (8)
- Super-learning (8)
- Asymptotic linearity (7)
- Confounding (7)
- Empirical process (7)
- Measurement error (7)
- Missing data (7)
- Survival analysis (7)
- Biomarker (6)
- Interaction (6)
- Inverse probability weighting (6)
- Pathwise differentiable parameter (6)
- Semiparametric statistical model (6)
- Variable selection (6)
- Efficient estimator (5)
- Functional data analysis (5)
- Mediation (5)
- Optimal dynamic treatment (5)
- Sensitivity (5)
- Asymptotic linear estimator (4)
- Asymptotic linearity of an estimator (4)
- Biostatistics (4)
- Canonical gradient (4)
- Publication Year
- Publication
-
- Harvard University Biostatistics Working Paper Series (140)
- U.C. Berkeley Division of Biostatistics Working Paper Series (118)
- UW Biostatistics Working Paper Series (102)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (69)
- The University of Michigan Department of Biostatistics Working Paper Series (55)
Articles 391 - 420 of 567
Full-Text Articles in Biostatistics
Redefining Cpg Islands Using A Hideen Markov Model, Hao Wu, Brain Caffo, Harris A. Jaffee, Andrew P. Feinberg, Rafael A. Irizarry
Redefining Cpg Islands Using A Hideen Markov Model, Hao Wu, Brain Caffo, Harris A. Jaffee, Andrew P. Feinberg, Rafael A. Irizarry
Johns Hopkins University, Dept. of Biostatistics Working Papers
The DNA of most vertebrates is depleted in CpG dinucleotides; C followed by a G in the 5’ to 3’ direction. CpGs are the target for DNA methylation, a chemical modification of cytosine (C) heritable during cell division and the most well characterized epigenetic mechanism. The remaining CpGs tend to cluster in regions referred to as CpG islands (CGI). Knowing CGI locations is important because they mark functionally relevant epigenetic loci in development and disease. For various mammals, including human, a readily available and widely used list of CGI is available from the UCSC Genome Browser. This list was derived …
Comparing Risk Scoring Systems Beyond The Roc Paradigm In Survival Analysis, Hajime Uno, Lu Tian, Tianxi Cai, Isaac S. Kohane, L. J. Wei
Comparing Risk Scoring Systems Beyond The Roc Paradigm In Survival Analysis, Hajime Uno, Lu Tian, Tianxi Cai, Isaac S. Kohane, L. J. Wei
Harvard University Biostatistics Working Paper Series
No abstract provided.
Nonparametric Population Average Models: Deriving The Form Of Approximate Population Average Models Estimated Using Generalized Estimating Equations, Alan E. Hubbard, Mark J. Van Der Laan
Nonparametric Population Average Models: Deriving The Form Of Approximate Population Average Models Estimated Using Generalized Estimating Equations, Alan E. Hubbard, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
For estimating regressions for repeated measures outcome data, a popular choice is the population average models estimated by generalized estimating equations (GEE). We review in this report the derivation of the robust inference (sandwich-type estimator of the standard error). In addition, we present formally how the approximation of a misspecified working population average model relates to the true model and in turn how to interpret the results of such a misspecified model.
Marginalized Frailty Models For Multivariate Survival Data, Megan Othus, Yi Li
Marginalized Frailty Models For Multivariate Survival Data, Megan Othus, Yi Li
Harvard University Biostatistics Working Paper Series
No abstract provided.
"Implementation Of Quasi-Least Squares With The R Package Qlspack", Jichun Xie, Justine Shults
"Implementation Of Quasi-Least Squares With The R Package Qlspack", Jichun Xie, Justine Shults
UPenn Biostatistics Working Papers
Quasi-least squares (QLS) is an alternative method for estimating the correlation parameters within the framework of generalized estimating equations (GEE) that has two main advantages over the moment estimates that are typically applied for GEE: (1) It guarantees a consistent estimate of the correlation parameter and a positive definite estimated correlation matrix, for several correlation structures; and (2) It allows for easier implementation of some correlation structures that have not yet been implemented in the framework of GEE. Furthermore, because QLS is a method in the framework of GEE, existing software can be employed within the QLS algorithm for estimation …
Simple, Defensible Sample Sizes Based On Cost Efficiency -- With Discussion And Rejoinder, Peter Bacchetti, Charles E. Mcculloch, Mark R. Segal, Richard Simon, Peter Muller, Gary L. Rosner, James A. Hanley, Stan Shapiro
Simple, Defensible Sample Sizes Based On Cost Efficiency -- With Discussion And Rejoinder, Peter Bacchetti, Charles E. Mcculloch, Mark R. Segal, Richard Simon, Peter Muller, Gary L. Rosner, James A. Hanley, Stan Shapiro
COBRA Preprint Series
The conventional approach of choosing sample size to provide 80% or greater power ignores the cost implications of different sample size choices. Costs, however, are often impossible for investigators and funders to ignore in actual practice. Here, we propose and justify a new approach for choosing sample size based on cost efficiency, the ratio of a study’s projected scientific and/or practical value to its total cost. By showing that a study’s projected value exhibits diminishing marginal returns as a function of increasing sample size for a wide variety of definitions of study value, we are able to develop two simple …
Nonparametric And Semiparametric Estimation Of The Three Way Receiver Operating Characteristic Surface, Jialiang Li, Xiao-Hua Zhou
Nonparametric And Semiparametric Estimation Of The Three Way Receiver Operating Characteristic Surface, Jialiang Li, Xiao-Hua Zhou
UW Biostatistics Working Paper Series
In many situations the diagnostic decision is not limited to a binary choice. Binary statistical tools such as receiver operating characteristic (ROC) curve and area under the ROC curve (AUC) need to be expanded to address three-category classification problem. Previous authors have suggest various ways to model the extension of AUC but not the ROC surface. Only simple parametric approaches are proposed for modeling the ROC measure under the assumption that test results all follow normal distributions. We study the estimation methods of three dimensional ROC surfaces with nonparametric and semiparametric estimators. Asymptotical results are provided as a basis for …
Evaluating Markers For Treatment Selection Based On Survival Time, Xiao Song, Xiao-Hua Zhou
Evaluating Markers For Treatment Selection Based On Survival Time, Xiao Song, Xiao-Hua Zhou
UW Biostatistics Working Paper Series
For many medical conditions several treatment options may be available for treating patients. We consider evaluating markers based on a simple treatment selection policy that incorporates information on the patient's marker value exceeding a threshold. For example, colon cancer patients may be treated by surgery alone or surgery plus chemotherapy. The c-myc gene expression level may be used as a biomarker for treatment selection. Although traditional regression methods may assess the effect of the marker and treatment on outcomes, it is appealing to quantify more directly the potential impact on the population of using the marker to select treatment. A …
A Machine-Learning Algorithm For Estimating And Ranking The Impact Of Environmental Risk Factors In Exploratory Epidemiological Studies, Jessica G. Young, Alan E. Hubbard, B Eskenazi, Nicholas P. Jewell
A Machine-Learning Algorithm For Estimating And Ranking The Impact Of Environmental Risk Factors In Exploratory Epidemiological Studies, Jessica G. Young, Alan E. Hubbard, B Eskenazi, Nicholas P. Jewell
U.C. Berkeley Division of Biostatistics Working Paper Series
No abstract provided.
Estimating Subject-Specific Dependent Competing Risk Profile With Censored Event Time Observations, Yi Li, Lu Tian, L. J. Wei
Estimating Subject-Specific Dependent Competing Risk Profile With Censored Event Time Observations, Yi Li, Lu Tian, L. J. Wei
Harvard University Biostatistics Working Paper Series
No abstract provided.
A Class Of Semiparametric Mixture Cure Survival Models With Dependent Censoring, Megan Othus, Yi Li, Ram C. Tiwari
A Class Of Semiparametric Mixture Cure Survival Models With Dependent Censoring, Megan Othus, Yi Li, Ram C. Tiwari
Harvard University Biostatistics Working Paper Series
No abstract provided.
Interval Estimation For The Difference In Paired Areas Under The Roc Curves In The Absence Of A Gold Standard Test, Hsin-Neng Hsieh, Hsiu-Yuan Su, Xiao-Hua Zhou
Interval Estimation For The Difference In Paired Areas Under The Roc Curves In The Absence Of A Gold Standard Test, Hsin-Neng Hsieh, Hsiu-Yuan Su, Xiao-Hua Zhou
UW Biostatistics Working Paper Series
Receiver operating characteristic (ROC) curves can be used to assess the accuracy of tests measured on ordinal or continuous scales. The most commonly used measure for the overall diagnostic accuracy of diagnostic tests is the area under the ROC curve (AUC). A gold standard test on the true disease status is required to estimate the AUC. However, a gold standard test may sometimes be too expensive or infeasible. Therefore, in many medical research studies, the true disease status of the subjects may remain unknown. Under the normality assumption on test results from each disease group of subjects, using the expectation-maximization …
A Semi-Parametric Two-Part Mixed-Effects Heteroscedastic Transformation Model For Correlated Right-Skewed Semi-Continuous Data, Huazhen Lin, Xiao-Hua Zhou
A Semi-Parametric Two-Part Mixed-Effects Heteroscedastic Transformation Model For Correlated Right-Skewed Semi-Continuous Data, Huazhen Lin, Xiao-Hua Zhou
UW Biostatistics Working Paper Series
In longitudinal or hierarchical structure studies, we often encounter a semi-continuous variable that has a certain proportion of a single value and a continuous and skewed distribution among the rest of values. In the paper, we propose a new semi-parametric two-part mixed-effects transformation model to fit correlated skewed semi-continuous data. In our model, we allow the transformation to be non-parametric. Fitting the proposed model faces computational challenges due to intractable numerical integrations. We derive the estimates for the parameter and the transformation function based on an approximate likelihood, which has high order accuracy but less computational burden. We also propose …
Composite Likelihood Bayesian Information Criteria For Model Selection In High Dimensional Data, X Gao, Peter Xuekun Song
Composite Likelihood Bayesian Information Criteria For Model Selection In High Dimensional Data, X Gao, Peter Xuekun Song
The University of Michigan Department of Biostatistics Working Paper Series
For high-dimensional data set with complicated dependency structures, the full likelihood approach often renders to intractable computational complexity. This imposes di±culty on model selection as most of the traditionally used information criteria require the evaluation of the full likelihood. We propose a composite likelihood version of the Bayesian information criterion (BIC) and establish its consistency property for the selection of the true underlying model. Under some mild regularity conditions, the proposed BIC is shown to be selection consistent, where the number of potential model parameters is allowed to increase to in¯nity at a certain rate of the sample size. Simulation …
Longitudinal Image Analysis Of Tumor/Brain Change In Contrast Uptake Induced By Radiation, Xiaoxi Zhang, Tim Johnson, Rod Little, Yue Cao
Longitudinal Image Analysis Of Tumor/Brain Change In Contrast Uptake Induced By Radiation, Xiaoxi Zhang, Tim Johnson, Rod Little, Yue Cao
The University of Michigan Department of Biostatistics Working Paper Series
This work is motivated by a quantitative Magnetic Resonance Imaging study of the differential tumor/healthy tissue change in contrast uptake induced by radiation. The goal is to determine the time in which there is maximal contrast uptake, a surrogate for permeability, in the tumor relative to healthy tissue. A notable feature of the data is its spatial heterogeneity. Zhang, Johnson, Little, and Cao (2008a and 2008b) discuss two parallel approaches to “denoise” a single image of change in contrast uptake from baseline to a single follow-up visit of interest. In this work we explore the longitudinal profile of the tumor/healthy …
Joint Multiple Testing Procedures For Graphical Model Selection With Applications To Biological Networks, Houston N. Gilbert, Mark J. Van Der Laan, Sandrine Dudoit
Joint Multiple Testing Procedures For Graphical Model Selection With Applications To Biological Networks, Houston N. Gilbert, Mark J. Van Der Laan, Sandrine Dudoit
U.C. Berkeley Division of Biostatistics Working Paper Series
Gaussian graphical models have become popular tools for identifying relationships between genes when analyzing microarray expression data. In the classical undirected Gaussian graphical model setting, conditional independence relationships can be inferred from partial correlations obtained from the concentration matrix (= inverse covariance matrix) when the sample size n exceeds the number of parameters p which need to estimated. In situations where n < p, another approach to graphical model estimation may rely on calculating unconditional (zero-order) and first-order partial correlations. In these settings, the goal is to identify a lower-order conditional independence graph, sometimes referred to as a ‘0-1 graphs’. For either choice of graph, model selection may involve a multiple testing problem, in which edges in a graph are drawn only after rejecting hypotheses involving (saturated or lower-order) partial correlation parameters. Most multiple testing procedures applied in previously proposed graphical model selection algorithms rely on standard, marginal testing methods which do not take into account the joint distribution of the test statistics derived from (partial) correlations. We propose and implement a multiple testing framework useful when testing for edge inclusion during graphical model selection. Two features of our methodology include (i) a computationally efficient and asymptotically valid test statistics joint null distribution derived from influence curves for correlation-based parameters, and (ii) the application of empirical Bayes joint multiple testing procedures which can effectively control a variety of popular Type I error rates by incorpo- rating joint null distributions such as those described here (Dudoit and van der Laan, 2008). Using a dataset from Arabidopsis thaliana, we observe that the use of more sophisticated, modular approaches to multiple testing allows one to identify greater numbers of edges when approximating an undirected graphical model using a 0-1 graph. Our framework may also be extended to edge testing algorithms for other types of graphical models (e.g., for classical undirected, bidirected, and directed acyclic graphs).
Analysis Of Randomized Comparative Clinical Trial Data For Personalized Treatment Selections, Tianxi Cai, Lu Tian, Peggy H. Wong, L. J. Wei
Analysis Of Randomized Comparative Clinical Trial Data For Personalized Treatment Selections, Tianxi Cai, Lu Tian, Peggy H. Wong, L. J. Wei
Harvard University Biostatistics Working Paper Series
No abstract provided.
Semiparametric Two-Part Models With Proportionality Constraints: Analysis Of The Multi-Ethnic Study Of Atherosclerosis (Mesa), Anna Liu, Richard Kronmal, Xiao-Hua Zhou, Shuangge Ma
Semiparametric Two-Part Models With Proportionality Constraints: Analysis Of The Multi-Ethnic Study Of Atherosclerosis (Mesa), Anna Liu, Richard Kronmal, Xiao-Hua Zhou, Shuangge Ma
UW Biostatistics Working Paper Series
SUMMARY. In this article, we analyze the coronary artery calcium (CAC) score in the Multi-Ethnic Study of Atherosclerosis (MESA), where about half of the CAC scores are zero and the rest are continuously distributed. When the observed data has a mixture distribution, two-part models can be the natural choice. With a two-part model, there are two covariate effects, with one in each part of the model. Determination of whether the two covariate effects are proportional can provide more insights into the process underlying development and progression of CAC. In this study, we model the CAC score using a semiparametric two-part …
Pooled Nucleic Acid Testing To Identify Antiretroviral Treatment Failure During Hiv Infection, Susanne May, Anthony Gamst, Richard Haubrich, Constance Benson, Davey Smith
Pooled Nucleic Acid Testing To Identify Antiretroviral Treatment Failure During Hiv Infection, Susanne May, Anthony Gamst, Richard Haubrich, Constance Benson, Davey Smith
UW Biostatistics Working Paper Series
Abstract Background: Pooling strategies have been used to reduce the costs of polymerase chain reaction based screening for acute HIV infection in populations where the prevalence of acute infection is low (<1%). Only limited research has been done for conditions where the prevalence of screening positivity is higher (>1%). Methods and Results: We present data on a variety of pooling strategies that incorporate the use of PCR-based quantitative measures to monitor for virologic failure among HIV-infected patients receiving antiretroviral therapy. For a prevalence of virologic failure between 1% and 25%, we demonstrate relative efficiency and accuracy of various strategies. These results could be used to choose the best strategy based on the requirements of individual laboratory …1%).>
Analysis Of Adverse Events In Drug Safety: A Multivariate Approach Using Stratified Quasi-Least Squares, Hanjoo Kim, Justine Shults, Scott Patterson, Robert Goldberg-Alberts
Analysis Of Adverse Events In Drug Safety: A Multivariate Approach Using Stratified Quasi-Least Squares, Hanjoo Kim, Justine Shults, Scott Patterson, Robert Goldberg-Alberts
UPenn Biostatistics Working Papers
Safety assessment in drug development involves numerous statistical challenges, and yet statistical methodologies and their applications to safety data have not been fully developed, despite a recent increase of interest in this area. In practice, a conventional univariate approach for analysis of safety data involves application of the Fisher's exact test to compare the proportion of subjects who experience adverse events (AEs) between treatment groups; This approach ignores several common features of safety data, including the presence of multiple endpoints, longitudinal follow-up, and a possible relationship between the AEs within body systems. In this article, we propose various regression modeling …
Synthesis Analysis Of Regression Models With A Continuous Outcome, Andrew Zhou, Nan Hu, Guizhou Hu, Martin Root
Synthesis Analysis Of Regression Models With A Continuous Outcome, Andrew Zhou, Nan Hu, Guizhou Hu, Martin Root
UW Biostatistics Working Paper Series
Synthesis Analysis of Regression Models with a Continuous Outcome Xiao-Hua Zhou 1,2, Nan Hu 2, Guizhou Hu3, and Martin Root3 1 HSR&D Center of Excellence, VA Puget Sound Health Care System, Seattle, WA 98101. 2 Department of Biostatistics, University of Washington, Seattle, WA 98195. 3 BioSignia, Inc., 1822 East NC Highway 54, Suite 350, Durham, NC 27713 To estimate the multivariate regression model from multiple individual studies, it would be challenging to obtain results if the input from individual studies only provide univariate or incomplete multivariate regression information. Samsa et al [1] proposed a simple method to combine coefficients from …
A Small Sample Correction For Estimating Attributable Risk In Case-Control Studies, Daniel B. Rubin
A Small Sample Correction For Estimating Attributable Risk In Case-Control Studies, Daniel B. Rubin
U.C. Berkeley Division of Biostatistics Working Paper Series
The attributable risk, often called the population attributable risk, is in many epidemiological contexts a more relevant measure of exposure-disease association than the excess risk, relative risk, or odds ratio. When estimating attributable risk with case-control data and a rare disease, we present a simple correction to the standard approach making it essentially unbiased, and also less noisy. As with analogous corrections given in Jewell (1986) for other measures of association, the adjustment often won't make a substantial difference unless the sample size is very small or point estimates are desired within fine strata, but we discuss the possible utility …
Optimal Cutpoint Estimation With Censored Data, Mithat Gonen, Camelia Sima
Optimal Cutpoint Estimation With Censored Data, Mithat Gonen, Camelia Sima
Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series
We consider the problem of selecting an optimal cutpoint for a continuous marker when the outcome of interest is subject to right censoring. Maximal chi square methods and receiver operating characteristic (ROC) curves-based methods are commonly-used when the outcome is binary. In this article we show that selecting the cutpoint that maximizes the concordance, a metric similar to the area under an ROC curve, is equivalent to maximizing the Youden index, a popular criterion when the ROC curve is used to choose a threshold. We use this as a basis for proposing maximal concordance as a metric to use with …
A New Class Of Rank Tests For Interval-Censored Data, Guadalupe Gomez, Ramon Oller Pique
A New Class Of Rank Tests For Interval-Censored Data, Guadalupe Gomez, Ramon Oller Pique
Harvard University Biostatistics Working Paper Series
No abstract provided.
The Highest Confidence Density Region And Its Usage For Inferences About The Survival Function With Censored Data, Lu Tian, Rui Wang, Tianxi Cai, L. J. Wei
The Highest Confidence Density Region And Its Usage For Inferences About The Survival Function With Censored Data, Lu Tian, Rui Wang, Tianxi Cai, L. J. Wei
Harvard University Biostatistics Working Paper Series
No abstract provided.
Calibrating Parametric Subject-Specific Risk Estimation, Tianxi Cai, Lu Tian, Hajime Uno, Scott D. Solomon, L. J. Wei
Calibrating Parametric Subject-Specific Risk Estimation, Tianxi Cai, Lu Tian, Hajime Uno, Scott D. Solomon, L. J. Wei
Harvard University Biostatistics Working Paper Series
No abstract provided.
Multilevel Latent Class Models With Dirichlet Mixing Distribution, Chongzhi Di, Karen Bandeen-Roche
Multilevel Latent Class Models With Dirichlet Mixing Distribution, Chongzhi Di, Karen Bandeen-Roche
Johns Hopkins University, Dept. of Biostatistics Working Papers
Latent class analysis (LCA) and latent class regression (LCR) are widely used for modeling multivariate categorical outcomes in social sciences and biomedical studies. Standard analyses assume data of different respondents to be mutually independent, excluding application of the methods to familial and other designs in which participants are clustered. In this paper, we develop multilevel latent class model, in which subpopulation mixing probabilities are treated as random effects that vary among clusters according to a common Dirichlet distribution. We apply the Expectation-Maximization (EM) algorithm for model fitting by maximum likelihood (ML). This approach works well, but is computationally intensive when …
Evaluating Subject-Level Incremental Values Of New Markers For Risk Classification Rule, Tianxi Cai, Lu Tian, Donald M. Lloyd-Jones, L. J. Wei
Evaluating Subject-Level Incremental Values Of New Markers For Risk Classification Rule, Tianxi Cai, Lu Tian, Donald M. Lloyd-Jones, L. J. Wei
Harvard University Biostatistics Working Paper Series
No abstract provided.
A Note On Risk Prediction For Case-Control Studies, Sherri Rose, Mark J. Van Der Laan
A Note On Risk Prediction For Case-Control Studies, Sherri Rose, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
We introduce a new method for prediction in case-control study designs, which is a simple extension of the work by van der Laan (2008). Case-control samples are biased since the proportion of cases in the sample is not the same as the population of interest. The case-control weighting for prediction proposed in this paper relies on knowledge of the true incidence probability P(Y=1) to eliminate the bias of the sampling design. In many practical settings, case-control weighting will outperform an existing method for prediction, intercept adjustment.
Confidence Intervals For Negative Binomial Random Variables Of High Dispersion, David Shilane, Alan E. Hubbard, S N. Evans
Confidence Intervals For Negative Binomial Random Variables Of High Dispersion, David Shilane, Alan E. Hubbard, S N. Evans
U.C. Berkeley Division of Biostatistics Working Paper Series
This paper considers the problem of constructing confidence intervals for the mean of a Negative Binomial random variable based upon sampled data. When the sample size is large, we traditionally rely upon a Normal distribution approximation to construct these intervals. However, we demonstrate that the sample mean of highly dispersed Negative Binomials exhibits a slow convergence to the Normal in distribution as a function of the sample size. As a result, standard techniques (such as the Normal approximation and bootstrap) that construct confidence intervals for the mean will typically be too narrow and significantly undercover in the case of high …