Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- COBRA (103)
- Central Bank of Nigeria (26)
- University of Kentucky (6)
- Georgia Southern University (5)
- Rochester Institute of Technology (5)
-
- Southern Methodist University (5)
- The British University in Egypt (5)
- Marshall University (4)
- Stephen F. Austin State University (4)
- University of Louisville (4)
- University of Connecticut (3)
- Virginia Commonwealth University (3)
- Claremont Colleges (2)
- East Tennessee State University (2)
- University at Albany, State University of New York (2)
- University of Arkansas, Fayetteville (2)
- University of Central Florida (2)
- University of Malaya (2)
- University of Nebraska - Lincoln (2)
- Western Michigan University (2)
- Bellarmine University (1)
- Bryant University (1)
- California Polytechnic State University, San Luis Obispo (1)
- Chapman University (1)
- City University of New York (CUNY) (1)
- Clemson University (1)
- Florida Institute of Technology (1)
- Georgetown University Law Center (1)
- Illinois Math and Science Academy (1)
- Illinois State University (1)
- Keyword
-
- Statistics (11)
- Counting process (6)
- Model selection (6)
- Prediction (6)
- Causal inference (5)
-
- Censored data (5)
- Cross-validation (5)
- Estimating equation (5)
- Semiparametric model (5)
- Simulation (5)
- Counterfactual (4)
- Estimation (4)
- Genetics (4)
- Nigeria (4)
- Risk (4)
- Bayesian (3)
- Bayesian inference (3)
- COVID-19 (3)
- Censoring (3)
- Gene expression (3)
- Linear regression (3)
- Longitudinal data (3)
- Loss function (3)
- Nonparametric regression (3)
- Regression (3)
- Asymptotic Theory (2)
- Baseball (2)
- Bayesian methods (2)
- Bias (2)
- Bootstrap (2)
- Publication Year
- Publication
-
- U.C. Berkeley Division of Biostatistics Working Paper Series (39)
- Harvard University Biostatistics Working Paper Series (27)
- CBN Journal of Applied Statistics (JAS) (26)
- The University of Michigan Department of Biostatistics Working Paper Series (13)
- Electronic Theses and Dissertations (10)
-
- Johns Hopkins University, Dept. of Biostatistics Working Papers (10)
- COBRA Preprint Series (7)
- UW Biostatistics Working Paper Series (7)
- Theses and Dissertations--Statistics (6)
- Articles (5)
- Basic Science Engineering (5)
- College of Graduate Studies: Theses & Dissertations (5)
- Theses and Dissertations (5)
- Theses, Dissertations and Capstones (4)
- Statistical Science Theses and Dissertations (3)
- Data Science and Data Mining (2)
- Dissertations (2)
- Electronic Theses & Dissertations (2024 - present) (2)
- Graduate Theses and Dissertations (2)
- Honors Scholar Theses (2)
- SMU Data Science Review (2)
- Student Works (2010-2019) (2)
- Al-Bahir (1)
- All Dissertations (1)
- All Graduate Theses and Dissertations, Spring 1920 to Summer 2023 (1)
- Business and Economics Honors Papers (1)
- CHIP Documents (1)
- CMC Senior Theses (1)
- Community & Environmental Health Faculty Publications (1)
- Department of Management: Faculty Publications (1)
- Publication Type
Articles 121 - 150 of 216
Full-Text Articles in Statistical Models
Nonparametric Regression With Missing Outcomes Using Weighted Kernel Estimating Equations, Lu Wang, Andrea Rotnitzky, Xihong Lin
Nonparametric Regression With Missing Outcomes Using Weighted Kernel Estimating Equations, Lu Wang, Andrea Rotnitzky, Xihong Lin
Harvard University Biostatistics Working Paper Series
No abstract provided.
Statistical Modelling And Inference For A Class Of Bivariate And Related Distributions., Ng Choung Min
Statistical Modelling And Inference For A Class Of Bivariate And Related Distributions., Ng Choung Min
Student Works (2010-2019)
This thesis considers bivariate extension of the Meixner class of distributions by the method of generalized trivariate reduction so that the marginal distributions have different parameters; in particular, a new bivariate negative binomial (BNB) distribution is examined. Different marginal parameters allow flexibility in statistical modelling and simulation studies when different marginal distributions and a specified correlation are required. The multivariate extension of this class of distributions is also given. Specifically, various interesting properties of the proposed BNB distribution, such as canonical expansion and quadrant dependence are examined. In addition, potential applications of the proposed distribution, as a bivariate mixed Poisson …
Some Problems Of Outliers In Circular Data., Ali H.M. Abuzaid
Some Problems Of Outliers In Circular Data., Ali H.M. Abuzaid
Student Works (2010-2019)
This study considers three problems of outliers in circular statistics. The first problem is an attempt to use the standard outlier detection procedures for linear data set by approximating circular variables by linear variables. This is possible for large values of concentration parameter. Series of simulation studies are carried out to specify the accepted value of the concentration parameter so that the von Mises distribution can be approximated by normal distribution. The second is the problem of outliers in circular samples. Two numerical tests of discordancy are proposed to identify outliers. The test statistics are based on the summation of …
Sequence Comparison And Stochastic Model Based On Multi-Order Markov Models, Xiang Fang
Sequence Comparison And Stochastic Model Based On Multi-Order Markov Models, Xiang Fang
Department of Statistics: Dissertations, Theses, and Student Research
This dissertation presents two statistical methodologies developed on multi-order Markov models. First, we introduce an alignment-free sequence comparison method, which represents a sequence using a multi-order transition matrix (MTM). The MTM contains information of multi-order dependencies and provides a comprehensive representation of the heterogeneous composition within a sequence. Based on the MTM, a distance measure is developed for pair-wise comparison of sequences. The new method is compared with the traditional maximum likelihood (ML) method, the complete composition vector (CCV) method and the improved version of the complete composition vector (ICCV) method using simulated sequences. We further illustrate the application of …
Shrinkage Estimation Of Expression Fold Change As An Alternative To Testing Hypotheses Of Equivalent Expression, Zahra Montazeri, Corey M. Yanofsky, David R. Bickel
Shrinkage Estimation Of Expression Fold Change As An Alternative To Testing Hypotheses Of Equivalent Expression, Zahra Montazeri, Corey M. Yanofsky, David R. Bickel
COBRA Preprint Series
Research on analyzing microarray data has focused on the problem of identifying differentially expressed genes to the neglect of the problem of how to integrate evidence that a gene is differentially expressed with information on the extent of its differential expression. Consequently, researchers currently prioritize genes for further study either on the basis of volcano plots or, more commonly, according to simple estimates of the fold change after filtering the genes with an arbitrary statistical significance threshold. While the subjective and informal nature of the former practice precludes quantification of its reliability, the latter practice is equivalent to using a …
A Class Of Semiparametric Mixture Cure Survival Models With Dependent Censoring, Megan Othus, Yi Li, Ram C. Tiwari
A Class Of Semiparametric Mixture Cure Survival Models With Dependent Censoring, Megan Othus, Yi Li, Ram C. Tiwari
Harvard University Biostatistics Working Paper Series
No abstract provided.
Analysis Of Randomized Comparative Clinical Trial Data For Personalized Treatment Selections, Tianxi Cai, Lu Tian, Peggy H. Wong, L. J. Wei
Analysis Of Randomized Comparative Clinical Trial Data For Personalized Treatment Selections, Tianxi Cai, Lu Tian, Peggy H. Wong, L. J. Wei
Harvard University Biostatistics Working Paper Series
No abstract provided.
Correlated Binary Regression Using Orthogonalized Residuals, Richard C. Zink, Bahjat F. Qaqish
Correlated Binary Regression Using Orthogonalized Residuals, Richard C. Zink, Bahjat F. Qaqish
COBRA Preprint Series
This paper focuses on marginal regression models for correlated binary responses when estimation of the association structure is of primary interest. A new estimating function approach based on orthogonalized residuals is proposed. This procedure allows a new representation and addresses some of the difficulties of the conditional-residual formulation of alternating logistic regressions of Carey, Zeger & Diggle (1993). The new method is illustrated with an analysis of data on impaired pulmonary function.
Group Comparison Of Eigenvalues And Eigenvectors Of Diffusion Tensors, Armin Schwartzman, Robert F. Dougherty, Jonathan E. Taylor
Group Comparison Of Eigenvalues And Eigenvectors Of Diffusion Tensors, Armin Schwartzman, Robert F. Dougherty, Jonathan E. Taylor
Harvard University Biostatistics Working Paper Series
No abstract provided.
The Strength Of Statistical Evidence For Composite Hypotheses With An Application To Multiple Comparisons, David R. Bickel
The Strength Of Statistical Evidence For Composite Hypotheses With An Application To Multiple Comparisons, David R. Bickel
COBRA Preprint Series
The strength of the statistical evidence in a sample of data that favors one composite hypothesis over another may be quantified by the likelihood ratio using the parameter value consistent with each hypothesis that maximizes the likelihood function. Unlike the p-value and the Bayes factor, this measure of evidence is coherent in the sense that it cannot support a hypothesis over any hypothesis that it entails. Further, when comparing the hypothesis that the parameter lies outside a non-trivial interval to the hypotheses that it lies within the interval, the proposed measure of evidence almost always asymptotically favors the correct hypothesis …
Calibrating Parametric Subject-Specific Risk Estimation, Tianxi Cai, Lu Tian, Hajime Uno, Scott D. Solomon, L. J. Wei
Calibrating Parametric Subject-Specific Risk Estimation, Tianxi Cai, Lu Tian, Hajime Uno, Scott D. Solomon, L. J. Wei
Harvard University Biostatistics Working Paper Series
No abstract provided.
Evaluating Subject-Level Incremental Values Of New Markers For Risk Classification Rule, Tianxi Cai, Lu Tian, Donald M. Lloyd-Jones, L. J. Wei
Evaluating Subject-Level Incremental Values Of New Markers For Risk Classification Rule, Tianxi Cai, Lu Tian, Donald M. Lloyd-Jones, L. J. Wei
Harvard University Biostatistics Working Paper Series
No abstract provided.
Measurement Error Caused By Spatial Misalignment In Environmental Epidemiology, Alexandros Gryparis, Christopher J. Paciorek, Ariana Zeka, Joel Schwartz, Brent A. Coull
Measurement Error Caused By Spatial Misalignment In Environmental Epidemiology, Alexandros Gryparis, Christopher J. Paciorek, Ariana Zeka, Joel Schwartz, Brent A. Coull
Harvard University Biostatistics Working Paper Series
No abstract provided.
Practical Large-Scale Spatio-Temporal Modeling Of Particulate Matter Concentrations, Christopher J. Paciorek, Jeff D. Yanosky, Robin C. Puett, Francine Laden, Helen H. Suh
Practical Large-Scale Spatio-Temporal Modeling Of Particulate Matter Concentrations, Christopher J. Paciorek, Jeff D. Yanosky, Robin C. Puett, Francine Laden, Helen H. Suh
Harvard University Biostatistics Working Paper Series
The last two decades have seen intense scientific and regulatory interest in the health effects of particulate matter (PM). Influential epidemiological studies that characterize chronic exposure of individuals rely on monitoring data that are sparse in space and time, so they often assign the same exposure to participants in large geographic areas and across time. We estimate monthly PM during 1988-2002 in a large spatial domain for use in studying health effects in the Nurses' Health Study. We develop a conceptually simple spatio-temporal model that uses a rich set of covariates. The model is used to estimate concentrations of PM10 …
Confidence Intervals For Negative Binomial Random Variables Of High Dispersion, David Shilane, Alan E. Hubbard, S N. Evans
Confidence Intervals For Negative Binomial Random Variables Of High Dispersion, David Shilane, Alan E. Hubbard, S N. Evans
U.C. Berkeley Division of Biostatistics Working Paper Series
This paper considers the problem of constructing confidence intervals for the mean of a Negative Binomial random variable based upon sampled data. When the sample size is large, we traditionally rely upon a Normal distribution approximation to construct these intervals. However, we demonstrate that the sample mean of highly dispersed Negative Binomials exhibits a slow convergence to the Normal in distribution as a function of the sample size. As a result, standard techniques (such as the Normal approximation and bootstrap) that construct confidence intervals for the mean will typically be too narrow and significantly undercover in the case of high …
Supervised Distance Matrices: Theory And Applications To Genomics, Katherine S. Pollard, Mark J. Van Der Laan
Supervised Distance Matrices: Theory And Applications To Genomics, Katherine S. Pollard, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
We propose a new approach to studying the relationship between a very high dimensional random variable and an outcome. Our method is based on a novel concept, the supervised distance matrix, which quantifies pairwise similarity between variables based on their association with the outcome. A supervised distance matrix is derived in two stages. The first stage involves a transformation based on a particular model for association. In particular, one might regress the outcome on each variable and then use the residuals or the influence curve from each regression as a data transformation. In the second stage, a choice of distance …
Confidence Intervals For The Population Mean Tailored To Small Sample Sizes, With Applications To Survey Sampling, Michael Rosenblum, Mark J. Van Der Laan
Confidence Intervals For The Population Mean Tailored To Small Sample Sizes, With Applications To Survey Sampling, Michael Rosenblum, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
The validity of standard confidence intervals constructed in survey sampling is based on the central limit theorem. For small sample sizes, the central limit theorem may give a poor approximation, resulting in confidence intervals that are misleading. We discuss this issue and propose methods for constructing confidence intervals for the population mean tailored to small sample sizes.
We present a simple approach for constructing confidence intervals for the population mean based on tail bounds for the sample mean that are correct for all sample sizes. Bernstein's inequality provides one such tail bound. The resulting confidence intervals have guaranteed coverage probability …
Using Regression Models To Analyze Randomized Trials: Asymptotically Valid Hypothesis Tests Despite Incorrectly Specified Models, Michael Rosenblum, Mark J. Van Der Laan
Using Regression Models To Analyze Randomized Trials: Asymptotically Valid Hypothesis Tests Despite Incorrectly Specified Models, Michael Rosenblum, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Regression models are often used to test for cause-effect relationships from data collected in randomized trials or experiments. This practice has deservedly come under heavy scrutiny, since commonly used models such as linear and logistic regression will often not capture the actual relationships between variables, and incorrectly specified models potentially lead to incorrect conclusions. In this paper, we focus on hypothesis test of whether the treatment given in a randomized trial has any effect on the mean of the primary outcome, within strata of baseline variables such as age, sex, and health status. Our primary concern is ensuring that such …
Super Learner, Mark J. Van Der Laan, Eric C. Polley, Alan E. Hubbard
Super Learner, Mark J. Van Der Laan, Eric C. Polley, Alan E. Hubbard
U.C. Berkeley Division of Biostatistics Working Paper Series
Previous articles (van der Laan and Dudoit (2003); van der Laan et al. (2006); Sinisi et al. (2007)) advertised and theoretically validated the use of cross-validation to select among many candidate estimators to compute a so called super learner which outperforms any of the given candidate estimators. The theoretical basis was provided for this super learner based on oracle results for the cross-validation selector (e.g., van der Laan and Dudoit (2003); van der Laan et al. (2006)) and in Sinisi et al. (2007). In addition, these papers contained a practical demonstration of the adaptivity of this so called super learner …
Semiparametric Regression Of Multi-Dimensional Genetic Pathway Data: Least Squares Kernel Machines And Linear Mixed Models, Dawei Liu, Xihong Lin, Debashis Ghosh
Semiparametric Regression Of Multi-Dimensional Genetic Pathway Data: Least Squares Kernel Machines And Linear Mixed Models, Dawei Liu, Xihong Lin, Debashis Ghosh
Harvard University Biostatistics Working Paper Series
No abstract provided.
Spatial Cluster Detection For Censored Outcome Data, Andrea J. Cook, Diane Gold, Yi Li
Spatial Cluster Detection For Censored Outcome Data, Andrea J. Cook, Diane Gold, Yi Li
Harvard University Biostatistics Working Paper Series
No abstract provided.
Predicting Future Responses Based On Possibly Misspecified Working Models, Tianxi Cai, Lu Tian, Scott D. Solomon, L.J. Wei
Predicting Future Responses Based On Possibly Misspecified Working Models, Tianxi Cai, Lu Tian, Scott D. Solomon, L.J. Wei
Harvard University Biostatistics Working Paper Series
No abstract provided.
Multiple Tests Of Association With Biological Annotation Metadata, Sandrine Dudoit, Sunduz Keles, Mark J. Van Der Laan
Multiple Tests Of Association With Biological Annotation Metadata, Sandrine Dudoit, Sunduz Keles, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
We propose a general and formal statistical framework for the multiple tests of associations between known fixed features of a genome and unknown parameters of the distribution of variable features of this genome in a population of interest. The known fixed gene-annotation profiles, corresponding to the fixed features of the genome, may concern Gene Ontology (GO) annotation, pathway membership, regulation by particular transcription factors, nucleotide sequences, or protein sequences. The unknown gene-parameter profiles, corresponding to the variable features of the genome, may be, for example, regression coefficients relating genome-wide transcript levels or DNA copy numbers to possibly censored biological and …
A Fine-Scale Linkage Disequilibrium Measure Based On Length Of Haplotype Sharing, Yan Wang, Lue Ping Zhao, Sandrine Dudoit
A Fine-Scale Linkage Disequilibrium Measure Based On Length Of Haplotype Sharing, Yan Wang, Lue Ping Zhao, Sandrine Dudoit
U.C. Berkeley Division of Biostatistics Working Paper Series
High-throughput genotyping technologies for single nucleotide polymorphisms (SNP) have enabled the recent completion of the International HapMap Project (Phase I), which has stimulated much interest in studying genome-wide linkage disequilibrium (LD) patterns. Conventional LD measures, such as D' and r-square, are two-point measurements, and their relationship with physical distance is highly noisy. We propose a new LD measure, defined in terms of the correlation coefficient for shared haplotype lengths around two loci, thereby borrowing information from multiple loci. A U-statistic-based estimator of the new LD measure, which takes into consideration the dependence structure of the observed data, is developed and …
Population Intervention Models In Causal Inference, Alan E. Hubbard, Mark J. Van Der Laan
Population Intervention Models In Causal Inference, Alan E. Hubbard, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Marginal structural models (MSM) provide a powerful tool for estimating the causal effect of a] treatment variable or risk variable on the distribution of a disease in a population. These models, as originally introduced by Robins (e.g., Robins (2000a), Robins (2000b), van der Laan and Robins (2002)), model the marginal distributions of treatment-specific counterfactual outcomes, possibly conditional on a subset of the baseline covariates, and its dependence on treatment. Marginal structural models are particularly useful in the context of longitudinal data structures, in which each subject's treatment and covariate history are measured over time, and an outcome is recorded at …
A Pseudolikelihood Approach For Simultaneous Analysis Of Array Comparative Genomic Hybridizations (Acgh), David A. Engler, Gayatry Mohapatra, David N. Louis, Rebecca Betensky
A Pseudolikelihood Approach For Simultaneous Analysis Of Array Comparative Genomic Hybridizations (Acgh), David A. Engler, Gayatry Mohapatra, David N. Louis, Rebecca Betensky
Harvard University Biostatistics Working Paper Series
DNA sequence copy number has been shown to be associated with cancer development and progression. Array-based Comparative Genomic Hybridization (aCGH) is a recent development that seeks to identify the copy number ratio at large numbers of markers across the genome. Due to experimental and biological variations across chromosomes and across hybridizations, current methods are limited to analyses of single chromosomes. We propose a more powerful approach that borrows strength across chromosomes and across hybridizations. We assume a Gaussian mixture model, with a hidden Markov dependence structure, and with random effects to allow for intertumoral variation, as well as intratumoral clonal …
Semiparametric Normal Transformation Models For Spatially Correlated Survival Data, Yi Li, Xihong Lin
Semiparametric Normal Transformation Models For Spatially Correlated Survival Data, Yi Li, Xihong Lin
Harvard University Biostatistics Working Paper Series
There is an emerging interest in modeling spatially correlated survival data in biomedical and epidemiological studies. In this paper, we propose a new class of semiparametric normal transformation models for right censored spatially correlated survival data. This class of models assumes that survival outcomes marginally follow a Cox proportional hazard model with unspecified baseline hazard, and their joint distribution is obtained by transforming survival outcomes to normal random variables, whose joint distribution is assumed to be multivariate normal with a spatial correlation structure. A key feature of the class of semiparametric normal transformation models is that it provides a rich …
Direct Effect Models, Mark J. Van Der Laan, Maya L. Petersen
Direct Effect Models, Mark J. Van Der Laan, Maya L. Petersen
U.C. Berkeley Division of Biostatistics Working Paper Series
The causal effect of a treatment on an outcome is generally mediated by several intermediate variables. Estimation of the component of the causal effect of a treatment that is mediated by a given intermediate variable (the indirect effect of the treatment), and the component that is not mediated by that intermediate variable (the direct effect of the treatment) is often relevant to mechanistic understanding and to the design of clinical and public health interventions. Under the assumption of no-unmeasured confounders for treatment and the intermediate variable, Robins & Greenland (1992) define an individual direct effect as the counterfactual effect of …
Application Of A Multiple Testing Procedure Controlling The Proportion Of False Positives To Protein And Bacterial Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan
Application Of A Multiple Testing Procedure Controlling The Proportion Of False Positives To Protein And Bacterial Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Simultaneously testing multiple hypotheses is important in high-dimensional biological studies. In these situations, one is often interested in controlling the Type-I error rate, such as the proportion of false positives to total rejections (TPPFP) at a specific level, alpha. This article will present an application of the E-Bayes/Bootstrap TPPFP procedure, presented in van der Laan et al. (2005), which controls the tail probability of the proportion of false positives (TPPFP), on two biological datasets. The two data applications include firstly, the application to a mass-spectrometry dataset of two leukemia subtypes, AML and ALL. The protein data measurements include intensity and …
Cross-Validating And Bagging Partitioning Algorithms With Variable Importance, Annette M. Molinaro, Mark J. Van Der Laan
Cross-Validating And Bagging Partitioning Algorithms With Variable Importance, Annette M. Molinaro, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
We present a cross-validated bagging scheme in the context of partitioning algorithms. To explore the benefits of the various bagging scheme, we compare via simulations the predictive ability of single Classification and Regression (CART) Tree with several previously suggested bagging schemes and with our proposed approach. Additionally, a variable importance measure is explained and illustrated.