Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- COBRA (107)
- Central Bank of Nigeria (26)
- University of Kentucky (22)
- Southern Methodist University (15)
- Virginia Commonwealth University (10)
-
- Georgia Southern University (9)
- University of Louisville (8)
- Clemson University (7)
- Michigan Technological University (7)
- University of Central Florida (7)
- Kennesaw State University (6)
- University of Arkansas, Fayetteville (6)
- City University of New York (CUNY) (5)
- Technological University Dublin (5)
- The University of Akron (5)
- University of Nevada, Las Vegas (5)
- The Texas Medical Center Library (4)
- University of Nebraska - Lincoln (4)
- California Polytechnic State University, San Luis Obispo (3)
- East Tennessee State University (3)
- James Madison University (3)
- Marshall University (3)
- University of New Hampshire (3)
- University of New Mexico (3)
- University of Texas Rio Grande Valley (3)
- Western Michigan University (3)
- Bucknell University (2)
- Chapman University (2)
- Claremont Colleges (2)
- Dartmouth College (2)
- Keyword
-
- Statistics (21)
- Regression (10)
- Machine Learning (8)
- Simulation (7)
- Bayesian inference (6)
-
- COVID-19 (6)
- Causal inference (6)
- Counting process (6)
- Model selection (6)
- Estimating equation (5)
- Machine learning (5)
- Prediction (5)
- Survival analysis (5)
- Censored data (4)
- Counterfactual (4)
- Cross-validation (4)
- Genetics (4)
- Linear regression (4)
- Logistic regression (4)
- Missing Data (4)
- Quantile regression (4)
- Semiparametric model (4)
- Variable Selection (4)
- Bayesian Statistics (3)
- Bayesian analysis (3)
- Bias (3)
- Censoring (3)
- Classification (3)
- Confidence intervals (3)
- Confounding (3)
- Publication Year
- Publication
-
- U.C. Berkeley Division of Biostatistics Working Paper Series (38)
- Harvard University Biostatistics Working Paper Series (27)
- CBN Journal of Applied Statistics (JAS) (26)
- Theses and Dissertations--Statistics (21)
- The University of Michigan Department of Biostatistics Working Paper Series (15)
-
- Electronic Theses and Dissertations (12)
- Theses and Dissertations (12)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (10)
- COBRA Preprint Series (9)
- College of Graduate Studies: Theses & Dissertations (8)
- Statistical Science Theses and Dissertations (8)
- All Dissertations (7)
- Dissertations, Master's Theses and Master's Reports (7)
- SMU Data Science Review (7)
- UW Biostatistics Working Paper Series (7)
- Data Science and Data Mining (6)
- Articles (5)
- Graduate Theses and Dissertations (5)
- Williams Honors College, Honors Research Projects (5)
- Reactor Campaign (TRP) (4)
- Dissertations (3)
- Dissertations and Theses (Open Access) (3)
- Published and Grey Literature from PhD Candidates (3)
- Theses, Dissertations and Capstones (3)
- Al-Bahir (2)
- CHIP Documents (2)
- Conference papers (2)
- Dartmouth College Ph.D Dissertations (2)
- Dissertations, 2014-2019 (2)
- Dissertations, Theses, and Capstone Projects (2)
- Publication Type
Articles 331 - 350 of 350
Full-Text Articles in Statistical Models
Tree-Based Multivariate Regression And Density Estimation With Right-Censored Data , Annette M. Molinaro, Sandrine Dudoit, Mark J. Van Der Laan
Tree-Based Multivariate Regression And Density Estimation With Right-Censored Data , Annette M. Molinaro, Sandrine Dudoit, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
We propose a unified strategy for estimator construction, selection, and performance assessment in the presence of censoring. This approach is entirely driven by the choice of a loss function for the full (uncensored) data structure and can be stated in terms of the following three main steps. (1) Define the parameter of interest as the minimizer of the expected loss, or risk, for a full data loss function chosen to represent the desired measure of performance. Map the full data loss function into an observed (censored) data loss function having the same expected value and leading to an efficient estimator …
Inference For The Population Total From Probability-Proportional-To-Size Samples Based On Predictions From A Penalized Spline Nonparametric Model, Hui Zheng, Rod Little
Inference For The Population Total From Probability-Proportional-To-Size Samples Based On Predictions From A Penalized Spline Nonparametric Model, Hui Zheng, Rod Little
The University of Michigan Department of Biostatistics Working Paper Series
Inference about the finite population total from probability-proportional-to-size (PPS) samples is considered. In previous work (Zheng and Little, 2003), penalized spline (p-spline) nonparametric model-based estimators were shown to generally outperform the Horvitz-Thompson (HT) and generalized regression (GR) estimators in terms of the root mean squared error. In this article we develop model-based, jackknife and balanced repeated replicate variance estimation methods for the p-spline based estimators. Asymptotic properties of the jackknife method are discussed. Simulations show that p-spline point estimators and their jackknife standard errors lead to inferences that are superior to HT or GR based inferences. This suggests that nonparametric …
Supervised Detection Of Regulatory Motifs In Dna Sequences, Sunduz Keles, Mark J. Van Der Laan, Sandrine Dudoit, Biao Xing, Michael B. Eisen
Supervised Detection Of Regulatory Motifs In Dna Sequences, Sunduz Keles, Mark J. Van Der Laan, Sandrine Dudoit, Biao Xing, Michael B. Eisen
U.C. Berkeley Division of Biostatistics Working Paper Series
Identification of transcription factor binding sites (regulatory motifs) is a major interest in contemporary biology. We propose a new likelihood based method, COMODE, for identifying structural motifs in DNA sequences. Commonly used methods (e.g. MEME, Gibbs sampler) model binding sites as families of sequences described by a position weight matrix (PWM) and identify PWMs that maximize the likelihood of observed sequence data under a simple multinomial mixture model. This model assumes that the positions of the PWM correspond to independent multinomial distributions with four cell probabilities. We address supervising the search for DNA binding sites using the information derived from …
Semiparametric Regression Models With Missing Data: The Mathematics In The Work Of Robins Et Al., Menggang Yu, Bin Nan
Semiparametric Regression Models With Missing Data: The Mathematics In The Work Of Robins Et Al., Menggang Yu, Bin Nan
The University of Michigan Department of Biostatistics Working Paper Series
This review is an attempt to understand the landmark papers of Robins, Rotnitzky, and Zhao (1994) and Robins and Rotnitzky (1992). We revisit their main results and corresponding proofs using the theory outlined in the monograph by Bickel, Klaassen, Ritov, and Wellner (1993). We also discuss an illustrative example to show the details of applying these theoretical results.
Penalized Spline Nonparametric Mixed Models For Inference About A Finite Population Mean From Two-Stage Samples, Hui Zheng, Rod Little
Penalized Spline Nonparametric Mixed Models For Inference About A Finite Population Mean From Two-Stage Samples, Hui Zheng, Rod Little
The University of Michigan Department of Biostatistics Working Paper Series
Samplers often distrust model-based approaches to survey inference due to concerns about model misspecification when applied to large samples from complex populations. We suggest that the model-based paradigm can work very successfully in survey settings, provided models are chosen that take into account the sample design and avoid strong parametric assumptions. The Horvitz-Thompson (HT) estimator is a simple design-unbiased estimator of the finite population total in probability sampling designs. From a modeling perspective, the HT estimator performs well when the ratios of the outcome values and the inclusion probabilities are exchangeable. When this assumption is not met, the HT estimator …
A Semiparametric Model Selection Criterion With Applications To The Marginal Structural Model, M. Alan Brookhart, Mark J. Van Der Laan
A Semiparametric Model Selection Criterion With Applications To The Marginal Structural Model, M. Alan Brookhart, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Estimators for the parameter of interest in semiparametric models often depend on a guessed model for the nuisance parameter. The choice of the model for the nuisance parameter can affect both the finite sample bias and efficiency of the resulting estimator of the parameter of interest. In this paper we propose a finite sample criterion based on cross validation that can be used to select a nuisance parameter model from a list of candidate models. We show that expected value of this criterion is minimized by the nuisance parameter model that yields the estimator of the parameter of interest with …
Semiparametric Receiver Operating Characteristic Analysis To Evaluate Biomarkers For Disease, Tianxi Cai, Margaret S. Pepe
Semiparametric Receiver Operating Characteristic Analysis To Evaluate Biomarkers For Disease, Tianxi Cai, Margaret S. Pepe
UW Biostatistics Working Paper Series
The receiver operating characteristic (ROC) curve is a popular method for characterizing the accuracy of diagnostic tests when test results are not binary. Various methodologies for estimating and comparing ROC curves have been developed. One approach, due to Pepe, uses a parametric regression model with the baseline function specified up to a finite-dimensional parameter. In this article we extend the regression models by allowing arbitrary nonparametric baseline functions. We also provide asymptotic distribution theory and procedures for making statistical inference. We illustrate our approach with dataset from a prostate cancer biomarker study. Simulation studies suggest that the extra flexibility inherent …
Semi-Parametric Regression For The Area Under The Receiver Operating Characteristic Curve, Lori E. Dodd, Margaret S. Pepe
Semi-Parametric Regression For The Area Under The Receiver Operating Characteristic Curve, Lori E. Dodd, Margaret S. Pepe
UW Biostatistics Working Paper Series
Medical advances continue to provide new and potentially better means for detecting disease. Such is true in cancer, for example, where biomarkers are sought for early detection and where improvements in imaging methods may pick up the initial functional and molecular changes associated with cancer development. In other binary classification tasks, computational algorithms such as Neural Networks, Support Vector Machines and Evolutionary Algorithms have been applied to areas as diverse as credit scoring, object recognition, and peptide-binding prediction. Before a classifier becomes an accepted technology, it must undergo rigorous evaluation to determine its ability to discriminate between states. Characterization of …
Checking Assumptions In Latent Class Regression Models Via A Markov Chain Monte Carlo Estimation Approach: An Application To Depression And Socio-Economic Status, Elizabeth Garrett, Richard Miech, Pamela Owens, William W. Eaton, Scott L. Zeger
Checking Assumptions In Latent Class Regression Models Via A Markov Chain Monte Carlo Estimation Approach: An Application To Depression And Socio-Economic Status, Elizabeth Garrett, Richard Miech, Pamela Owens, William W. Eaton, Scott L. Zeger
Johns Hopkins University, Dept. of Biostatistics Working Papers
Latent class regression models are useful tools for assessing associations between covariates and latent variables. However, evaluation of key model assumptions cannot be performed using methods from standard regression models due to the unobserved nature of latent outcome variables. This paper presents graphical diagnostic tools to evaluate whether or not latent class regression models adhere to standard assumptions of the model: conditional independence and non-differential measurement. An integral part of these methods is the use of a Markov Chain Monte Carlo estimation procedure. Unlike standard maximum likelihood implementations for latent class regression model estimation, the MCMC approach allows us to …
Recurrent Events Analysis In The Presence Of Time Dependent Covariates And Dependent Censoring, Maja Miloslavsky, Sunduz Keles, Mark J. Van Der Laan, Steve Butler
Recurrent Events Analysis In The Presence Of Time Dependent Covariates And Dependent Censoring, Maja Miloslavsky, Sunduz Keles, Mark J. Van Der Laan, Steve Butler
U.C. Berkeley Division of Biostatistics Working Paper Series
Recurrent events models have lately received a lot of attention in the literature. The majority of approaches discussed show the consistency of parameter estimates under the assumption that censoring is independent of the recurrent events process of interest conditional on the covariates included into the model. We provide an overview of available recurrent events analysis methods, and present an inverse probability of censoring weighted estimator for the regression parameters in the Andersen-Gill model that is commonly used for recurrent event analysis. This estimator remains consistent under informative censoring if the censoring mechanism is estimated consistently, and generally improves on the …
New Statitstical Methods For The Estimation Of The Mean And Standard Deviation From Normally Distributed Censored Samples, Abou El-Makarim Abd El-Alim Aboueissa
New Statitstical Methods For The Estimation Of The Mean And Standard Deviation From Normally Distributed Censored Samples, Abou El-Makarim Abd El-Alim Aboueissa
Dissertations
The main objective of this dissertation is to estimate the mean /x and standard deviation cr of a normal population from left-censored samples. We have developed new methods for calculating estimates for the mean and standard deviation of a normal population from left-censored samples. Some of these methods based on traditional estimating procedures. A new method of obtaining the Cohen maximum likelihood estimates for fx and cr without the aid of an auxiliary table will be introduced. This new method will be used to extend Cohen table of estimating the Cohen A-parameter that is required for calculating the maximum likelihood …
Robust Residuals And Diagnostics In Autoregressive Time Series, Kirk W. Anderson
Robust Residuals And Diagnostics In Autoregressive Time Series, Kirk W. Anderson
Dissertations
One of the goals of model diagnostics is outlier detection. In particular, we would like to use the residuals, appropriately standardized, to “flag” outliers. Hopefully, our (robust) procedure has yielded a fit that resists undue influence by outlying points, while simultaneously drawing attention to these interesting points via residual analysis. In this study we consider several different methods of standardizing the residuals resulting from autoregression. A large sample approximation for the variance of rank-based first order autoregressive time series residuals is developed. This provides studentized residuals, specific to the time series model and estimation procedure. Simulation studies are presented that …
Locally Efficient Estimation With Bivariate Right Censored Data , Christopher M. Quale, Mark J. Van Der Laan, James M. Robins
Locally Efficient Estimation With Bivariate Right Censored Data , Christopher M. Quale, Mark J. Van Der Laan, James M. Robins
U.C. Berkeley Division of Biostatistics Working Paper Series
Estimation for bivariate right censored data is a problem that has had much study over the past 15 years. In this paper we propose a new class of estimators for the bivariate survivor function based on locally efficient estimation. The locally efficient estimator takes bivariate estimators Fn and Gn of the distributions of the time variables T1,T2 and the censoring variables C1,C2, respectively, and maps them to the resulting estimator. If Fn and Gn are consistent estimators of F and G, respectively, then the resulting estimator will be nonparametrically efficient (thus the term ``locally efficient''). However, if either Fn or …
Accelerated Hazards Model: Method, Theory And Applications, Ying Qing Chen, Nicholas P. Jewell, Jingrong Yang
Accelerated Hazards Model: Method, Theory And Applications, Ying Qing Chen, Nicholas P. Jewell, Jingrong Yang
U.C. Berkeley Division of Biostatistics Working Paper Series
In an accelerated hazards model, the hazard functions of a failure time are related through the time scale-change, which is often a function of covariates and associated parameters. When the hazard functions have special properties, such as monotonicity in time, the parameters may be clinically meaningful in measuring a treatment effect. This paper reviews methodological and theoretical development of this model. Applications of the accelerated hazards model including sample size calculation in clinical trials, are also explored.
Locally Efficient Estimation Of Regression Parameters Using Current Status Data, Chris Andrews, Mark J. Van Der Laan, James M. Robins
Locally Efficient Estimation Of Regression Parameters Using Current Status Data, Chris Andrews, Mark J. Van Der Laan, James M. Robins
U.C. Berkeley Division of Biostatistics Working Paper Series
In biostatistics applications interest often focuses on the estimation of the distribution of a time-variable T. If one only observes whether or not T exceeds an observed monitoring time C, then the data structure is called current status data, also known as interval censored data, case I. We consider this data structure extended to allow the presence of both time-independent covariates and time-dependent covariate processes that are observed until the monitoring time. We assume that the monitoring process satisfies coarsening at random.
Our goal is to estimate the regression parameter beta of the regression model T = Z*beta+epsilon where the …
Why Prefer Double Robust Estimates? Illustration With Causal Point Treatment Studies, Romain Neugebauer, Mark J. Van Der Laan
Why Prefer Double Robust Estimates? Illustration With Causal Point Treatment Studies, Romain Neugebauer, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
In point treatment marginal structural models with treatment A, outcome Y and covariates W, causal parameters can be estimated under the assumption of no unobserved confounders. Three estimates can be used: the G-computation, Inverse Probability of Treatment Weighted (IPTW) or Double Robust (DR) estimates. The properties of the IPTW and DR estimates are known under an assumption on the treatment mechanism that we name "Experimental Treatment Assignment" (ETA) assumption. We show that the DR estimating function is unbiased when the ETA assumption is violated if the model used to regress Y on A and W is correctly specified. The practical …
Semiparametric Regression Analysis On Longitudinal Pattern Of Recurrent Gap Times, Ying Qing Chen, Mei-Cheng Wang, Yijian Huang
Semiparametric Regression Analysis On Longitudinal Pattern Of Recurrent Gap Times, Ying Qing Chen, Mei-Cheng Wang, Yijian Huang
U.C. Berkeley Division of Biostatistics Working Paper Series
In longitudinal studies, individual subjects may experience recurrent events of the same type over a relatively long period of time. The longitudinal pattern of the gaps between the successive recurrent events is often of great research interest. In this article, the probability structure of the recurrent gap times is first explored in the presence of censoring. According to the discovered structure, we introduce the proportional reverse-time hazards models with unspecified baseline functions to accommodate heterogeneous individual underlying distributions, when the ongitudinal pattern parameter is of main interest. Inference procedures are proposed and studied by way of proper riskset construction. The …
Inference For Proportional Mean Residual Life Model In The Presence Of Censoring, Ying Q. Chen, Nicholas P. Jewell
Inference For Proportional Mean Residual Life Model In The Presence Of Censoring, Ying Q. Chen, Nicholas P. Jewell
U.C. Berkeley Division of Biostatistics Working Paper Series
As a function of time t, mean residual life is defined as remaining life expectancy of a subject given its survival to t. It plays an important role in many research areas to characterise stochastic behavior of survival over time. Similar to the Cox proportional hazard model, the proportional mean residual life model were proposed in statistical literature to study association between the mean residual life and individual subject's explanatory covariates. In this article, we will study this model and develop appropriate inference procedures in presence of censoring. Numerical studies including simulation and real data analysis are presented as well.
Marginal Regression Of Gaps Between Recurrent Events, Yijian Huang, Ying Qing Chen
Marginal Regression Of Gaps Between Recurrent Events, Yijian Huang, Ying Qing Chen
U.C. Berkeley Division of Biostatistics Working Paper Series
Recurrent event data typically exhibit the phenomenon of intra-individual correlation, owing to not only observed covariates but also random effects. In many applications, the population can be reasonably postulated as a heterogeneous mixture of individual renewal processes, and the inference of interest is the effect of individual-level covariates. In this article, we suggest and investigate a marginal proportional hazards model for gaps between recurrent events. A connection is established between observed gap times and clustered survival data, however, with informative cluster size. We then derive a novel and general inference procedure for the latter, based on a functional formulation of …
Simulation Of Mathematical Models In Genetic Analysis, Dinesh Govindal Patel
Simulation Of Mathematical Models In Genetic Analysis, Dinesh Govindal Patel
All Graduate Theses and Dissertations, Spring 1920 to Summer 2023
In recent years a new field of statistics has become of importance in many branches of experimental science. This is the Monte Carlo Method, so called because it is based on simulation of stochastic processes. By stochastic process, it is meant some possible physical process in the real world that has some random or stochastic element in its structure. This is the subject which may appropriately be called the dynamic part of statistics or the statistics of "change," in contrast with the static statistical problems which have so far been the more systematically studied. Many obvious examples of such processes …