Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Statistical Methodology (99)
- Statistical Theory (74)
- Medicine and Health Sciences (65)
- Statistical Models (51)
- Public Health (44)
-
- Epidemiology (32)
- Survival Analysis (28)
- Clinical Trials (26)
- Life Sciences (26)
- Genetics and Genomics (23)
- Multivariate Analysis (23)
- Genetics (18)
- Microarrays (14)
- Bioinformatics (13)
- Computational Biology (13)
- Diseases (13)
- Disease Modeling (11)
- Longitudinal Data Analysis and Time Series (11)
- Applied Statistics (10)
- Clinical Epidemiology (10)
- Medical Specialties (9)
- Categorical Data Analysis (7)
- Design of Experiments and Sample Surveys (7)
- Applied Mathematics (6)
- Laboratory and Basic Science Research (6)
- Numerical Analysis and Computation (6)
- Social and Behavioral Sciences (6)
- Keyword
-
- Causal inference (16)
- Cross-validation (13)
- Targeted maximum likelihood estimation (12)
- Efficient influence curve (11)
- Genetics (10)
-
- Influence curve (10)
- Causal effect (8)
- Longitudinal data (8)
- Super-learning (8)
- Asymptotic linearity (7)
- Confounding (7)
- Empirical process (7)
- Measurement error (7)
- Missing data (7)
- Survival analysis (7)
- Biomarker (6)
- Interaction (6)
- Inverse probability weighting (6)
- Pathwise differentiable parameter (6)
- Semiparametric statistical model (6)
- Variable selection (6)
- Efficient estimator (5)
- Functional data analysis (5)
- Mediation (5)
- Optimal dynamic treatment (5)
- Sensitivity (5)
- Asymptotic linear estimator (4)
- Asymptotic linearity of an estimator (4)
- Biostatistics (4)
- Canonical gradient (4)
- Publication Year
- Publication
-
- Harvard University Biostatistics Working Paper Series (140)
- U.C. Berkeley Division of Biostatistics Working Paper Series (118)
- UW Biostatistics Working Paper Series (102)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (69)
- The University of Michigan Department of Biostatistics Working Paper Series (55)
Articles 211 - 240 of 567
Full-Text Articles in Biostatistics
Challenges In Estimating The Causal Effect Of An Intervention With Pre-Post Data (Part 1): Definition & Identification Of The Causal Parameter, Ann M. Weber, Mark J. Van Der Laan, Maya L. Petersen
Challenges In Estimating The Causal Effect Of An Intervention With Pre-Post Data (Part 1): Definition & Identification Of The Causal Parameter, Ann M. Weber, Mark J. Van Der Laan, Maya L. Petersen
U.C. Berkeley Division of Biostatistics Working Paper Series
There is mixed evidence of the effectiveness of interventions operating on a large scale. Although the lack of consistent results is generally attributed to problems of implementation or governance of the program, the failure to find a statistically significant effect (or the success of finding one) may be due to choices made in the evaluation. To demonstrate the potential limitations and pitfalls of the usual analytic methods used for estimating causal effects, we apply the first half of a roadmap for causal inference to a pre-post evaluation of a community-level, national nutrition program. Selection into the program was non-random and …
Variable Importance And Prediction Methods For Longitudinal Problems With Missing Variables, Ivan Diaz, Alan E. Hubbard, Anna Decker, Mitchell Cohen
Variable Importance And Prediction Methods For Longitudinal Problems With Missing Variables, Ivan Diaz, Alan E. Hubbard, Anna Decker, Mitchell Cohen
U.C. Berkeley Division of Biostatistics Working Paper Series
In this paper we present prediction and variable importance (VIM) methods for longitudinal data sets containing both continuous and binary exposures subject to missingness. We demonstrate the use of these methods for prognosis of medical outcomes of severe trauma patients, a field in which current medical practice involves rules of thumb and scoring methods that only use a few variables and ignore the dynamic and high-dimensional nature of trauma recovery. Well-principled prediction and VIM methods can thus provide a tool to make care decisions informed by the high-dimensional patient’s physiological and clinical history. Our VIM parameters can be causally interpreted …
Targeted Learning Of An Optimal Dynamic Treatment, And Statistical Inference For Its Mean Outcome, Mark J. Van Der Laan
Targeted Learning Of An Optimal Dynamic Treatment, And Statistical Inference For Its Mean Outcome, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Suppose we observe n independent and identically distributed observations of a time-dependent random variable consisting of baseline covariates, initial treatment and censoring indicator, intermediate covariates, subsequent treatment and censoring indicator, and a final outcome. For example, this could be data generated by a sequentially randomized controlled trial, where subjects are sequentially randomized to a first line and second line treatment, possibly assigned in response to an intermediate biomarker, and are subject to right-censoring. In this article we consider estimation of an optimal dynamic multiple time-point treatment rule defined as the rule that maximizes the mean outcome under the dynamic treatment, …
Sparse Median Graphs Estimation In A High Dimensional Semiparametric Model, Fang Han, Han Liu, Brian Caffo
Sparse Median Graphs Estimation In A High Dimensional Semiparametric Model, Fang Han, Han Liu, Brian Caffo
Johns Hopkins University, Dept. of Biostatistics Working Papers
In this manuscript a unified framework for conducting inference on complex aggregated data in high dimensional settings is proposed. The data are assumed to be a collection of multiple non-Gaussian realizations with underlying undirected graphical structures. Utilizing the concept of median graphs in summarizing the commonality across these graphical structures, a novel semiparametric approach to modeling such complex aggregated data is provided along with robust estimation of the median graph, which is assumed to be sparse. The estimator is proved to be consistent in graph recovery and an upper bound on the rate of convergence is given. Experiments on both …
Adapting Data Adaptive Methods For Small, But High Dimensional Omic Data: Applications To Gwas/Ewas And More, Sara Kherad Pajouh, Alan E. Hubbard, Martyn T. Smith
Adapting Data Adaptive Methods For Small, But High Dimensional Omic Data: Applications To Gwas/Ewas And More, Sara Kherad Pajouh, Alan E. Hubbard, Martyn T. Smith
U.C. Berkeley Division of Biostatistics Working Paper Series
Exploratory analysis of high dimensional "omics" data has received much attention since the explosion of high-throughput technology allows simultaneous screening of tens of thousands of characteristics (genomics, metabolomics, proteomics, adducts, etc., etc.). Part of this trend has been an increase in the dimension of exposure data in studies of environmental exposure and associated biomarkers. Though some of the general approaches, such as GWAS, are transferable, what has received less focus is 1) how to derive estimation of independent associations in the context of many competing causes, without resorting to a misspecified model, and 2) how to derive accurate small-sample inference …
Regression Trees For Longitudinal Data, Madan Gopal Kundu, Jaroslaw Harezlak
Regression Trees For Longitudinal Data, Madan Gopal Kundu, Jaroslaw Harezlak
COBRA Preprint Series
Often when a longitudinal change is studied in a population of interest we find that changes over time are heterogeneous (in terms of time and/or covariates' effect) and a traditional linear mixed effect model [Laird and Ware, 1982] on the entire population assuming common parametric form for covariates and time may not be applicable to the entire population. This is usually the case in studies when there are many possible predictors influencing the response trajectory. For example, Raudenbush [2001] used depression as an example to argue that it is incorrect to assume that all the people in a given population …
Net Reclassification Index: A Misleading Measure Of Prediction Improvement, Margaret Sullivan Pepe, Holly Janes, Kathleen F. Kerr, Bruce M. Psaty
Net Reclassification Index: A Misleading Measure Of Prediction Improvement, Margaret Sullivan Pepe, Holly Janes, Kathleen F. Kerr, Bruce M. Psaty
UW Biostatistics Working Paper Series
The evaluation of biomarkers to improve risk prediction is a common theme in modern research. Since its introduction in 2008, the net reclassification index (NRI) (Pencina et al. 2008, Pencina et al. 2011) has gained widespread use as a measure of prediction performance with over 1,200 citations as of June 30, 2013. The NRI is considered by some to be more sensitive to clinically important changes in risk than the traditional change in the AUC (Delta AUC) statistic (Hlatky et al. 2009). Recent statistical research has raised questions, however, about the validity of conclusions based on the NRI. (Hilden and …
Normalization Techniques For Statistical Inference From Magnetic Resonance Imaging, Russell T. Shinohara, Elizabeth M. Sweeney, Jeff Goldsmith, Navid Shiee, Farrah J. Mateen, Peter A. Calabresi, Samson Jarso, Dzung L. Pham, Daniel S. Reich, Ciprian M. Crainiceanu
Normalization Techniques For Statistical Inference From Magnetic Resonance Imaging, Russell T. Shinohara, Elizabeth M. Sweeney, Jeff Goldsmith, Navid Shiee, Farrah J. Mateen, Peter A. Calabresi, Samson Jarso, Dzung L. Pham, Daniel S. Reich, Ciprian M. Crainiceanu
UPenn Biostatistics Working Papers
While computed tomography and other imaging techniques are measured in absolute units with physical meaning, magnetic resonance images are expressed in arbitrary units that are difficult to interpret and differ between study visits and subjects. Much work in the image processing literature on intensity normalization has focused on histogram matching and other histogram mapping techniques, with little emphasis on normalizing images to have biologically interpretable units. Furthermore, there are no formalized principles or goals for the crucial comparability of image intensities within and across subjects. To address this, we propose a set of criteria necessary for the normalization of images. …
Net Reclassification Indices For Evaluating Risk Prediction Instruments: A Critical Review, Kathleen F. Kerr, Zheyu Wang, Holly Janes, Robyn Mcclelland, Bruce M. Psaty, Margaret S. Pepe
Net Reclassification Indices For Evaluating Risk Prediction Instruments: A Critical Review, Kathleen F. Kerr, Zheyu Wang, Holly Janes, Robyn Mcclelland, Bruce M. Psaty, Margaret S. Pepe
UW Biostatistics Working Paper Series
Background Net Reclassification Indices (NRI) have recently become popular statistics for measuring the prediction increment of new biomarkers.
Methods In this review, we examine the various types of NRI statistics and their correct interpretations. We evaluate the advantages and disadvantages of the NRI approach. For pre-defined risk categories, we relate NRI to existing measures of the prediction increment. We also consider statistical methodology for constructing confidence intervals for NRI statistics and evaluate the merits of NRI-based hypothesis testing.
Conclusions Investigators using NRI statistics should report them separately for events (cases) and nonevents (controls). When there are two risk categories, the …
Testing The Relative Performance Of Data Adaptive Prediction Algorithms: A Generalized Test Of Conditional Risk Differences, Benjamin A. Goldstein, Eric Polley, Farren Briggs, Mark J. Van Der Laan
Testing The Relative Performance Of Data Adaptive Prediction Algorithms: A Generalized Test Of Conditional Risk Differences, Benjamin A. Goldstein, Eric Polley, Farren Briggs, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
In statistical medicine comparing the predictability or fit of two models can help to determine whether a set of prognostic variables contains additional information about medical outcomes, or whether one of two different model fits (perhaps based on different algorithms, or different set of variables) should be preferred for clinical use. Clinical medicine has tended to rely on comparisons of clinical metrics like C-statistics and more recently reclassification. Such metrics rely on the outcome being categorical and utilize a specific and often obscure loss function. In classical statistics one can use likelihood ratio tests and information based criterion if the …
Attributing Effects To Interactions, Tyler J. Vanderweele, Eric J. Tchetgen Tchetgen
Attributing Effects To Interactions, Tyler J. Vanderweele, Eric J. Tchetgen Tchetgen
Harvard University Biostatistics Working Paper Series
A framework is presented which allows an investigator to estimate the portion of the effect of one exposure that is attributable to an interaction with a second exposure. We show that when the two exposures are independent, the total effect of one exposure can be decomposed into a conditional effect of that exposure and a component due to interaction. The decomposition applies on difference or ratio scales. We discuss how the components can be estimated using standard regression models, and how these components can be used to evaluate the proportion of the total effect of the primary exposure attributable to …
Sample Size Considerations In The Design Of Cluster Randomized Trials Of Combination Hiv Prevention, Rui Wang, Ravi Goyal, Quanhong Lei, M. Essex, Victor Degruttola
Sample Size Considerations In The Design Of Cluster Randomized Trials Of Combination Hiv Prevention, Rui Wang, Ravi Goyal, Quanhong Lei, M. Essex, Victor Degruttola
Harvard University Biostatistics Working Paper Series
No abstract provided.
Fast Covariance Estimation For High-Dimensional Functional Data, Luo Xiao, David Ruppert, Vadim Zipunnikov, Ciprian Crainiceanu
Fast Covariance Estimation For High-Dimensional Functional Data, Luo Xiao, David Ruppert, Vadim Zipunnikov, Ciprian Crainiceanu
Johns Hopkins University, Dept. of Biostatistics Working Papers
For smoothing covariance functions, we propose two fast algorithms that scale linearly with the number of observations per function. Most available methods and software cannot smooth covariance matrices of dimension J x J with J>500; the recently introduced sandwich smoother is an exception, but it is not adapted to smooth covariance matrices of large dimensions such as J \ge 10,000. Covariance matrices of order J=10,000, and even J=100,000$ are becoming increasingly common, e.g., in 2- and 3-dimensional medical imaging and high-density wearable sensor data. We introduce two new algorithms that can handle very large covariance matrices: 1) FACE: a …
Soft Null Hypotheses: A Case Study Of Image Enhancement Detection In Brain Lesions, Haochang Shou, Russell T. Shinohara, Han Liu, Daniel Reich, Ciprian Crainiceanu
Soft Null Hypotheses: A Case Study Of Image Enhancement Detection In Brain Lesions, Haochang Shou, Russell T. Shinohara, Han Liu, Daniel Reich, Ciprian Crainiceanu
Johns Hopkins University, Dept. of Biostatistics Working Papers
This work is motivated by a study of a population of multiple sclerosis (MS) patients using dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) to identify active brain lesions. At each visit, a contrast agent is administered intravenously to a subject and a series of images is acquired to reveal the location and activity of MS lesions within the brain. Our goal is to identify and quantify lesion enhancement location at the subject level and lesion enhancement patterns at the population level. With this example, we aim to address the difficult problem of transforming a qualitative scientific null hypothesis, such as "this …
Phylogenetic Linkage Among Hiv-Infected Village Residents In Botswana: Estimation Of Clustering Rates In The Presence Of Missing Data, Nicole Bohme Carnegie, Rui Wang, Vladimir Novitsky, Victor G. Degruttola
Phylogenetic Linkage Among Hiv-Infected Village Residents In Botswana: Estimation Of Clustering Rates In The Presence Of Missing Data, Nicole Bohme Carnegie, Rui Wang, Vladimir Novitsky, Victor G. Degruttola
Harvard University Biostatistics Working Paper Series
No abstract provided.
Statistical Inference For Data Adaptive Target Parameters, Mark J. Van Der Laan, Alan E. Hubbard, Sara Kherad Pajouh
Statistical Inference For Data Adaptive Target Parameters, Mark J. Van Der Laan, Alan E. Hubbard, Sara Kherad Pajouh
U.C. Berkeley Division of Biostatistics Working Paper Series
Consider one observes n i.i.d. copies of a random variable with a probability distribution that is known to be an element of a particular statistical model. In order to define our statistical target we partition the sample in V equal size sub-samples, and use this partitioning to define V splits in estimation-sample (one of the V subsamples) and corresponding complementary parameter-generating sample that is used to generate a target parameter. For each of the V parameter-generating samples, we apply an algorithm that maps the sample in a target parameter mapping which represent the statistical target parameter generated by that parameter-generating …
Restricted Likelihood Ratio Tests For Functional Effects In The Functional Linear Model, Bruce J. Swihart, Jeff Goldsmith, Ciprian M. Crainiceanu
Restricted Likelihood Ratio Tests For Functional Effects In The Functional Linear Model, Bruce J. Swihart, Jeff Goldsmith, Ciprian M. Crainiceanu
Johns Hopkins University, Dept. of Biostatistics Working Papers
The goal of our article is to provide a transparent, robust, and computationally feasible statistical approach for testing in the context of scalar-on-function linear regression models. In particular, we are interested in testing for the necessity of functional effects against standard linear models. Our methods are motivated by and applied to a large longitudinal study involving diffusion tensor imaging of intracranial white matter tracts in a susceptible cohort. In the context of this study, we conduct hypothesis tests that are motivated by anatomical knowledge and which support recent findings regarding the relationship between cognitive impairment and white matter demyelination. R-code …
Augmentation Of Propensity Scores For Medical Records-Based Research, Mikel Aickin
Augmentation Of Propensity Scores For Medical Records-Based Research, Mikel Aickin
COBRA Preprint Series
Therapeutic research based on electronic medical records suffers from the possibility of various kinds of confounding. Over the past 30 years, propensity scores have increasingly been used to try to reduce this possibility. In this article a gap is identified in the propensity score methodology, and it is proposed to augment traditional treatment-propensity scores with outcome-propensity scores, thereby removing all other aspects of common causes from the analysis of treatment effects.
A Versatile Test For Equality Of Two Survival Functions Based On Weighted Differences Of Kaplan-Meier Curves, Hajime Uno, Lu Tian, Brian Claggett, L. J. Wei
A Versatile Test For Equality Of Two Survival Functions Based On Weighted Differences Of Kaplan-Meier Curves, Hajime Uno, Lu Tian, Brian Claggett, L. J. Wei
Harvard University Biostatistics Working Paper Series
With censored event time observations, the logrank test is the most popular tool for testing the equality of two underlying survival distributions. Although this test is asymptotically distribution-free, it may not be powerful when the proportional hazards assumption is violated. Various other novel testing procedures have been proposed, which generally are derived by assuming a class of specific alternative hypotheses with respect to the hazard functions. The test considered by Pepe and Fleming (1989) is based on a linear combination of weighted differences of two Kaplan-Meier curves over time and is a natural tool to assess the difference of two …
Subsemble: An Ensemble Method For Combining Subset-Specific Algorithm Fits, Stephanie Sapp, Mark J. Van Der Laan, John Canny
Subsemble: An Ensemble Method For Combining Subset-Specific Algorithm Fits, Stephanie Sapp, Mark J. Van Der Laan, John Canny
U.C. Berkeley Division of Biostatistics Working Paper Series
Ensemble methods using the same underlying algorithm trained on different subsets of observations have recently received increased attention as practical prediction tools for massive datasets. We propose Subsemble: a general subset ensemble prediction method, which can be used for small, moderate, or large datasets. Subsemble partitions the full dataset into subsets of observations, fits a specified underlying algorithm on each subset, and uses a clever form of V-fold cross-validation to output a prediction function that combines the subset-specific fits. We give an oracle result that provides a theoretical performance guarantee for Subsemble. Through simulations, we demonstrate that Subsemble can be …
Targeted Maximum Likelihood Estimation For Dynamic And Static Longitudinal Marginal Structural Working Models, Maya L. Petersen, Joshua Schwab, Susan Gruber, Nello Blaser, Michael Schomaker, Mark J. Van Der Laan
Targeted Maximum Likelihood Estimation For Dynamic And Static Longitudinal Marginal Structural Working Models, Maya L. Petersen, Joshua Schwab, Susan Gruber, Nello Blaser, Michael Schomaker, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
This paper describes a targeted maximum likelihood estimator (TMLE) for the parameters of longitudinal static and dynamic marginal structural models. We consider a longitudinal data structure consisting of baseline covariates, time-dependent intervention nodes, intermediate time-dependent covariates, and a possibly time dependent outcome. The intervention nodes at each time point can include a binary treatment as well as a right-censoring indicator. Given a class of dynamic or static interventions, a marginal structural model is used to model the mean of the intervention specific counterfactual outcome as a function of the intervention, time point, and possibly a subset of baseline covariates. Because …
Varying Index Coefficient Models, Shujie Ma, Peter Xuekun Song
Varying Index Coefficient Models, Shujie Ma, Peter Xuekun Song
The University of Michigan Department of Biostatistics Working Paper Series
It has been a long history of utilizing interactions in regression analysis to investigate interactive effects of covariates on response variables. In this paper we aim to address two kinds of new challenges resulted from the inclusion of such high-order effects in the regression model for complex data. The first kind arises from a situation where interaction effects of individual covariates are weak but those of combined covariates are strong, and the other kind pertains to the presence of nonlinear interactive effects. Generalizing the single index coefficient regression model (Xia and Li, 1999), we propose a new class of semiparametric …
Balancing Score Adjusted Targeted Minimum Loss-Based Estimation, Samuel D. Lendle, Bruce Fireman, Mark J. Van Der Laan
Balancing Score Adjusted Targeted Minimum Loss-Based Estimation, Samuel D. Lendle, Bruce Fireman, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Adjusting for a balancing score is sufficient for bias reduction when estimating causal effects including the average treatment effect and effect among the treated. Estimators that adjust for the propensity score in a nonparametric way, such as matching on an estimate of the propensity score, can be consistent when the estimated propensity score is not consistent for the true propensity score but converges to some other balancing score. We call this property the balancing score property, and discuss a class of estimators that have this property. We introduce a targeted minimum loss-based estimator (TMLE) for a treatment specific mean with …
Optimal Tests Of Treatment Effects For The Overall Population And Two Subpopulations In Randomized Trials, Using Sparse Linear Programming, Michael Rosenblum, Han Liu, En-Hsu Yen
Optimal Tests Of Treatment Effects For The Overall Population And Two Subpopulations In Randomized Trials, Using Sparse Linear Programming, Michael Rosenblum, Han Liu, En-Hsu Yen
Johns Hopkins University, Dept. of Biostatistics Working Papers
We propose new, optimal methods for analyzing randomized trials, when it is suspected that treatment effects may differ in two predefined subpopulations. Such sub-populations could be defined by a biomarker or risk factor measured at baseline. The goal is to simultaneously learn which subpopulations benefit from an experimental treatment, while providing strong control of the familywise Type I error rate. We formalize this as a multiple testing problem and show it is computationally infeasible to solve using existing techniques. Our solution involves a novel approach, in which we first transform the original multiple testing problem into a large, sparse linear …
Estimating Effects On Rare Outcomes: Knowledge Is Power, Laura B. Balzer, Mark J. Van Der Laan
Estimating Effects On Rare Outcomes: Knowledge Is Power, Laura B. Balzer, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Many of the secondary outcomes in observational studies and randomized trials are rare. Methods for estimating causal effects and associations with rare outcomes, however, are limited, and this represents a missed opportunity for investigation. In this article, we construct a new targeted minimum loss-based estimator (TMLE) for the effect of an exposure or treatment on a rare outcome. We focus on the causal risk difference and statistical models incorporating bounds on the conditional risk of the outcome, given the exposure and covariates. By construction, the proposed estimator constrains the predicted outcomes to respect this model knowledge. Theoretically, this bounding provides …
Más-O-Menos: A Simple Sign Averaging Method For Discrimination In Genomic Data Analysis, Sihai Dave Zhao, Giovanni Parmigiani, Curtis Huttenhower, Levi Waldron
Más-O-Menos: A Simple Sign Averaging Method For Discrimination In Genomic Data Analysis, Sihai Dave Zhao, Giovanni Parmigiani, Curtis Huttenhower, Levi Waldron
Harvard University Biostatistics Working Paper Series
No abstract provided.
An Application Of Machine Learning Methods To The Derivation Of Exposure-Response Curves For Respiratory Outcomes, Ekaterina Eliseeva, Alan E. Hubbard, Ira B. Tager
An Application Of Machine Learning Methods To The Derivation Of Exposure-Response Curves For Respiratory Outcomes, Ekaterina Eliseeva, Alan E. Hubbard, Ira B. Tager
U.C. Berkeley Division of Biostatistics Working Paper Series
Analyses of epidemiological studies of the association between short-term changes in air pollution and health outcomes have not sufficiently discussed the degree to which the statistical models chosen for these analyses reflect what is actually known about the true data-generating distribution. We present a method to estimate population-level ambient air pollution (NO2) exposure-health (wheeze in children with asthma) response functions that is not dependent on assumptions about the data-generating function that underlies the observed data and which focuses on a specific scientific parameter of interest (the marginal adjusted association of exposure on probability of wheeze, over a grid of possible …
Penalized Smoothed Partial Rank Estimator For The Nonparametric Transformation Survival Model With High-Dimensional Covariates, Wei Dai, Yi Li
Penalized Smoothed Partial Rank Estimator For The Nonparametric Transformation Survival Model With High-Dimensional Covariates, Wei Dai, Yi Li
The University of Michigan Department of Biostatistics Working Paper Series
Microarray technology has the potential to lead to a better understanding of biological processes and diseases such as cancer. When failure time outcomes are also available, one might be interested in relating gene expression profiles to the survival outcome such as time to cancer recurrence or time to death. This is statistically challenging because the number of covariates greatly exceeds the number of observations. While the majority of work has focused on regularized Cox regression model and accelerated failure time model, they may be restrictive in practice. We relax the model assumption and and consider a nonparametric transformation model that …
Structured Functional Principal Component Analysis, Haochang Shou, Vadim Zipunnikov, Ciprian Crainiceanu, Sonja Greven
Structured Functional Principal Component Analysis, Haochang Shou, Vadim Zipunnikov, Ciprian Crainiceanu, Sonja Greven
Johns Hopkins University, Dept. of Biostatistics Working Papers
Motivated by modern observational studies, we introduce a class of functional models that expands nested and crossed designs. These models account for the natural inheritance of correlation structure from sampling design in studies where the fundamental sampling unit is a function or image. Inference is based on functional quadratics and their relationship with the underlying covariance structure of the latent processes. A computationally fast and scalable estimation procedure is developed for ultra-high dimensional data. Methods are illustrated in three examples: high-frequency accelerometer data for daily activity, pitch linguistic data for phonetic analysis, and EEG data for studying electrical brain activity …
Penalized Function-On-Function Regression, Andrada E. Ivanescu, Ana-Maria Staicu, Fabian Scheipl, Sonja Greven
Penalized Function-On-Function Regression, Andrada E. Ivanescu, Ana-Maria Staicu, Fabian Scheipl, Sonja Greven
Johns Hopkins University, Dept. of Biostatistics Working Papers
We propose a general framework for smooth regression of a functional response on one or multiple functional predictors. Using the mixed model representation of penalized regression expands the scope of function on function regression to many realistic scenarios. In particular, the approach can accommodate a densely or sparsely sampled functional response as well as multiple functional predictors that are observed: 1) on the same or different domains than the functional response; 2) on a dense or sparse grid; and 3) with or without noise. It also allows for seamless integration of continuous or categorical covariates and provides approximate confidence intervals …