Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- COBRA (21)
- Central Bank of Nigeria (20)
- Southern Methodist University (7)
- California Polytechnic State University, San Luis Obispo (4)
- City University of New York (CUNY) (4)
-
- The University of Akron (4)
- Purdue University (3)
- University of Louisville (3)
- Virginia Commonwealth University (3)
- Claremont Colleges (2)
- Murray State University (2)
- University of New Hampshire (2)
- University of Southern Maine (2)
- Belmont University (1)
- Bowling Green State University (1)
- Brigham Young University (1)
- California State University, San Bernardino (1)
- Colby College (1)
- East Tennessee State University (1)
- Georgia Southern University (1)
- Illinois State University (1)
- Kennesaw State University (1)
- LSU New Orleans (1)
- Lindenwood University (1)
- Michigan Technological University (1)
- Missouri State University (1)
- Montclair State University (1)
- Northern Illinois University (1)
- Old Dominion University (1)
- Portland State University (1)
- Keyword
-
- Statistics (12)
- Machine Learning (6)
- Regression (4)
- Data Science (3)
- Hierarchical models (3)
-
- Machine learning (3)
- Nigeria (3)
- R (3)
- Artificial Intelligence (2)
- Bayesian analysis (2)
- Breast Cancer (2)
- COVID-19 (2)
- Classification (2)
- Clustering (2)
- Decision Trees (2)
- EM algorithm (2)
- Economic Growth (2)
- Education (2)
- Forecasting (2)
- GARCH (2)
- Missing Data (2)
- Ordinal data (2)
- Ordinal regression (2)
- Poverty (2)
- ROC curve (2)
- Simulation (2)
- Statistical Modeling (2)
- Stock market (2)
- Text mining (2)
- Time series (2)
- Publication Year
- Publication
-
- CBN Journal of Applied Statistics (JAS) (20)
- SMU Data Science Review (6)
- The University of Michigan Department of Biostatistics Working Paper Series (5)
- UW Biostatistics Working Paper Series (5)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (4)
-
- Theses and Dissertations (4)
- Williams Honors College, Honors Research Projects (4)
- COBRA Preprint Series (3)
- Electronic Theses and Dissertations (3)
- Harvard University Biostatistics Working Paper Series (3)
- Statistics (3)
- Dissertations, Theses, and Capstone Projects (2)
- Honors Projects (2)
- Honors Theses (2)
- Honors Theses and Capstones (2)
- Publications and Research (2)
- The Summer Undergraduate Research Fellowship (SURF) Symposium (2)
- All Graduate Reports and Creative Projects, Fall 2023 to Present (1)
- Articles (1)
- Business and Economics Honors Papers (1)
- CMC Senior Theses (1)
- Data Science Undergraduate Honors Theses (1)
- Department of Mathematics Faculty Scholarship and Creative Works (1)
- Department of Statistics: Faculty Publications (1)
- Dissertations (1)
- Dissertations, Master's Theses and Master's Reports (1)
- Doctor of Business Administration Dissertations (1)
- Electronic Theses, Projects, and Dissertations (1)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (1)
- Graduate Theses/Dissertations (1)
- Publication Type
- File Type
Articles 91 - 112 of 112
Full-Text Articles in Categorical Data Analysis
A Unified Approach To Modeling Multivariate Binary Data Using Copulas Over Partitions, Bruce J. Swihart, Brian Caffo, Ciprian Crainiceanu
A Unified Approach To Modeling Multivariate Binary Data Using Copulas Over Partitions, Bruce J. Swihart, Brian Caffo, Ciprian Crainiceanu
Johns Hopkins University, Dept. of Biostatistics Working Papers
Many seemingly disparate approaches for marginal modeling have been developed in recent years. We demonstrate that many current approaches for marginal modeling of correlated binary outcomes produce likelihoods that are equivalent to the proposed copula-based models herein. These general copula models of underlying latent threshold random variables yield likelihood based models for marginal fixed effects estimation and interpretation in the analysis of correlated binary data. Moreover, we propose a nomenclature and set of model relationships that substantially elucidates the complex area of marginalized models for binary data. A diverse collection of didactic mathematical and numerical examples are given to illustrate …
Nonparametric Regression With Missing Outcomes Using Weighted Kernel Estimating Equations, Lu Wang, Andrea Rotnitzky, Xihong Lin
Nonparametric Regression With Missing Outcomes Using Weighted Kernel Estimating Equations, Lu Wang, Andrea Rotnitzky, Xihong Lin
Harvard University Biostatistics Working Paper Series
No abstract provided.
Software Internationalization: A Framework Validated Against Industry Requirements For Computer Science And Software Engineering Programs, John Huân Vũ
Master's Theses
View John Huân Vũ's thesis presentation at http://youtu.be/y3bzNmkTr-c.
In 2001, the ACM and IEEE Computing Curriculum stated that it was necessary to address "the need to develop implementation models that are international in scope and could be practiced in universities around the world." With increasing connectivity through the internet, the move towards a global economy and growing use of technology places software internationalization as a more important concern for developers. However, there has been a "clear shortage in terms of numbers of trained persons applying for entry-level positions" in this area. Eric Brechner, Director of Microsoft Development Training, suggested …
The Em Algorithm For Group Testing Regression Models Under Matrix Pooling, Christopher R. Bilder, Boan Zhang
The Em Algorithm For Group Testing Regression Models Under Matrix Pooling, Christopher R. Bilder, Boan Zhang
Department of Statistics: Faculty Publications
No abstract provided.
Multilevel Latent Class Models With Dirichlet Mixing Distribution, Chongzhi Di, Karen Bandeen-Roche
Multilevel Latent Class Models With Dirichlet Mixing Distribution, Chongzhi Di, Karen Bandeen-Roche
Johns Hopkins University, Dept. of Biostatistics Working Papers
Latent class analysis (LCA) and latent class regression (LCR) are widely used for modeling multivariate categorical outcomes in social sciences and biomedical studies. Standard analyses assume data of different respondents to be mutually independent, excluding application of the methods to familial and other designs in which participants are clustered. In this paper, we develop multilevel latent class model, in which subpopulation mixing probabilities are treated as random effects that vary among clusters according to a common Dirichlet distribution. We apply the Expectation-Maximization (EM) algorithm for model fitting by maximum likelihood (ML). This approach works well, but is computationally intensive when …
Measurement Error Caused By Spatial Misalignment In Environmental Epidemiology, Alexandros Gryparis, Christopher J. Paciorek, Ariana Zeka, Joel Schwartz, Brent A. Coull
Measurement Error Caused By Spatial Misalignment In Environmental Epidemiology, Alexandros Gryparis, Christopher J. Paciorek, Ariana Zeka, Joel Schwartz, Brent A. Coull
Harvard University Biostatistics Working Paper Series
No abstract provided.
Choice And Optimization Of Forecasting Models For Container Port Throughput, Qingcheng Xue
Choice And Optimization Of Forecasting Models For Container Port Throughput, Qingcheng Xue
World Maritime University Dissertations
No abstract provided.
Causal Comparisons In Randomized Trials Of Two Active Treatments: The Effect Of Supervised Exercise To Promote Smoking Cessation, Jason Roy, Joseph W. Hogan
Causal Comparisons In Randomized Trials Of Two Active Treatments: The Effect Of Supervised Exercise To Promote Smoking Cessation, Jason Roy, Joseph W. Hogan
COBRA Preprint Series
In behavioral medicine trials, such as smoking cessation trials, two or more active treatments are often compared. Noncompliance by some subjects with their assigned treatment poses a challenge to the data analyst. Causal parameters of interest might include those defined by subpopulations based on their potential compliance status under each assignment, using the principal stratification framework (e.g., causal effect of new therapy compared to standard therapy among subjects that would comply with either intervention). Even if subjects in one arm do not have access to the other treatment(s), the causal effect of each treatment typically can only be identified from …
A Nonstationary Negative Binomial Time Series With Time-Dependent Covariates: Enterococcus Counts In Boston Harbor, E. Andres Houseman, Brent Coull, James P. Shine
A Nonstationary Negative Binomial Time Series With Time-Dependent Covariates: Enterococcus Counts In Boston Harbor, E. Andres Houseman, Brent Coull, James P. Shine
Harvard University Biostatistics Working Paper Series
Boston Harbor has had a history of poor water quality, including contamination by enteric pathogens. We conduct a statistical analysis of data collected by the Massachusetts Water Resources Authority (MWRA) between 1996 and 2002 to evaluate the effects of court-mandated improvements in sewage treatment. Motivated by the ineffectiveness of standard Poisson mixture models and their zero-inflated counterparts, we propose a new negative binomial model for time series of Enterococcus counts in Boston Harbor, where nonstationarity and autocorrelation are modeled using a nonparametric smooth function of time in the predictor. Without further restrictions, this function is not identifiable in the presence …
Semi-Parametric Single-Index Two-Part Regression Models, Xiao-Hua Zhou, Hua Liang
Semi-Parametric Single-Index Two-Part Regression Models, Xiao-Hua Zhou, Hua Liang
UW Biostatistics Working Paper Series
In this paper, we proposed a semi-parametric single-index two-part regression model to weaken assumptions in parametric regression methods that were frequently used in the analysis of skewed data with additional zero values. The estimation procedure for the parameters of interest in the model was easily implemented. The proposed estimators were shown to be consistent and asymptotically normal. Through a simulation study, we showed that the proposed estimators have reasonable finite-sample performance. We illustrated the application of the proposed method in one real study on the analysis of health care costs.
The Proportional Odds Model For Assessing Rater Agreement With Multiple Modalities, Elizabeth Garrett-Mayer, Steven N. Goodman, Ralph H. Hruban
The Proportional Odds Model For Assessing Rater Agreement With Multiple Modalities, Elizabeth Garrett-Mayer, Steven N. Goodman, Ralph H. Hruban
Johns Hopkins University, Dept. of Biostatistics Working Papers
In this paper, we develop a model for evaluating an ordinal rating systems where we assume that the true underlying disease state is continuous in nature. Our approach in motivated by a dataset with 35 microscopic slides with 35 representative duct lesions of the pancreas. Each of the slides was evaluated by eight raters using two novel rating systems (PanIN illustrations and PanIN nomenclature),where each rater used each systems to rate the slide with slide identity masked between evaluations. We find that the two methods perform equally well but that differentiation of higher grade lesions is more consistent across raters …
Semiparametric Binary Regression Under Monotonicity Constraints, Moulinath Banerjee, Pinaki Biswas, Debashis Ghosh
Semiparametric Binary Regression Under Monotonicity Constraints, Moulinath Banerjee, Pinaki Biswas, Debashis Ghosh
The University of Michigan Department of Biostatistics Working Paper Series
Summary: We study a binary regression model where the response variable $\Delta$ is the indicator of an event of interest (for example, the incidence of cancer) and the set of covariates can be partitioned as $(X,Z)$ where $Z$ (real valued) is the covariate of primary interest and $X$ (vector valued) denotes a set of control variables. For any fixed $X$, the conditional probability of the event of interest is assumed to be a monotonic function of $Z$. The effect of the control variables is captured by a regression parameter $\beta$. We show that the baseline conditional probability function (corresponding to …
Binary Isotonic Regression Procedures, With Application To Cancer Biomarkers, Debashis Ghosh, Moulinath Banerjee, Pinaki Biswas
Binary Isotonic Regression Procedures, With Application To Cancer Biomarkers, Debashis Ghosh, Moulinath Banerjee, Pinaki Biswas
The University of Michigan Department of Biostatistics Working Paper Series
There is a lot of interest in the development and characterization of new biomarkers for screening large populations for disease. In much of the literature on diagnostic testing, increased levels of a biomarker correlate with increased disease risk. However, parametric forms are typically used to associate these quantities. In this article, we specify a monotonic relationship between biomarker levels with disease risk. This leads to consideration of a nonparametric regression model for a single biomarker. Estimation results using isotonic regression-type estimators and asymptotic results are given. We also discuss confidence set estimation in this setting and propose three procedures for …
A Bayesian Hierarchical Approach To Multirater Correlated Roc Analysis, Tim Johnson, Valen Johnson
A Bayesian Hierarchical Approach To Multirater Correlated Roc Analysis, Tim Johnson, Valen Johnson
The University of Michigan Department of Biostatistics Working Paper Series
In a common ROC study design, several readers are asked to rate diagnostics of the same cases processed under different modalities. We describe a Bayesian hierarchical model that facilitates the analysis of this study design by explicitly modeling the three sources of variation inherent to it. In so doing, we achieve substantial reductions in the posterior uncertainty associated with estimates of the differences in areas under the estimated ROC curves and corresponding reductions in the mean squared error (MSE) of these estimates. Based on simulation studies, both the widths of confidence intervals and MSE of estimates of differences in the …
A Bayesian Chi-Squared Test For Goodness Of Fit, Valen Johnson
A Bayesian Chi-Squared Test For Goodness Of Fit, Valen Johnson
The University of Michigan Department of Biostatistics Working Paper Series
This article describes an extension of classical x 2 goodness-of-fit tests to Bayesian model assessment. The extension, which essentially involvesevaluating Pearson's goodness-of-fit statistic at a parameter value drawn from its posterior distribution, has the important property that it is asymptoti-cally distributed as a x2 random variable on K-1 degrees of freedom, indepen-dently of the dimension of the underlying parameter vector. By averaging over the posterior distribution of this statistic, a global goodness-of-fit diagnostic is obtained. Advantages of this diagnostic{which may be interpreted as the area under an ROC curve{include ease of interpretation, computational conve-nience, and favorable power properties. The proposed …
Marginalized Transition Models For Longitudinal Binary Data With Ignorable And Nonignorable Dropout, Brenda F. Kurland, Patrick J. Heagerty
Marginalized Transition Models For Longitudinal Binary Data With Ignorable And Nonignorable Dropout, Brenda F. Kurland, Patrick J. Heagerty
UW Biostatistics Working Paper Series
We extend the marginalized transition model of Heagerty (2002) to accommodate nonignorable monotone dropout. Using a selection model, weakly identified dropout parameters are held constant and their effects evaluated through sensitivity analysis. For data missing at random (MAR), efficiency of inverse probability of censoring weighted generalized estimating equations (IPCW-GEE) is as low as 40% compared to a likelihood-based marginalized transition model (MTM) with comparable modeling burden. MTM and IPCW-GEE regression parameters both display misspecification bias for MAR and nonignorable missing data, and both reduce bias noticeably by improving model fit
Marginal Modeling Of Multilevel Binary Data With Time-Varying Covariates, Diana Miglioretti, Patrick Heagerty
Marginal Modeling Of Multilevel Binary Data With Time-Varying Covariates, Diana Miglioretti, Patrick Heagerty
UW Biostatistics Working Paper Series
We propose and compare two approaches for regression analysis of multilevel binary data when clusters are not necessarily nested: a GEE method that relies on a working independence assumption coupled with a three-step method for obtaining empirical standard errors; and a likelihood-based method implemented using Bayesian computational techniques. Implications of time-varying endogenous covariates are addressed. The methods are illustrated using data from the Breast Cancer Surveillance Consortium to estimate mammography accuracy from a repeatedly screened population.
Hierarchical Bivariate Time Series Models: A Combined Analysis Of The Effects Of Particulate Matter On Morbidity And Mortality, Francesca Dominici, Antonella Zanobetti, Scott L. Zeger, Joel Schwartz, Jonathan M. Samet
Hierarchical Bivariate Time Series Models: A Combined Analysis Of The Effects Of Particulate Matter On Morbidity And Mortality, Francesca Dominici, Antonella Zanobetti, Scott L. Zeger, Joel Schwartz, Jonathan M. Samet
Johns Hopkins University, Dept. of Biostatistics Working Papers
In this paper we develop a hierarchical bivariate time series model to characterize the relationship between particulate matter less than 10 microns in aerodynamic diameter (PM10) and both mortality and hospital admissions for cardiovascular diseases. The model is applied to time series data on mortality and morbidity for 10 metropolitan areas in the United States from 1986 to 1993. We postulate that these time series should be related through a shared relationship with PM10.
At the first stage of the hierarchy, we fit two seemingly unrelated Poisson regression models to produce city-specific estimates of the log relative rates of mortality …
Supervised Detection Of Regulatory Motifs In Dna Sequences, Sunduz Keles, Mark J. Van Der Laan, Sandrine Dudoit, Biao Xing, Michael B. Eisen
Supervised Detection Of Regulatory Motifs In Dna Sequences, Sunduz Keles, Mark J. Van Der Laan, Sandrine Dudoit, Biao Xing, Michael B. Eisen
U.C. Berkeley Division of Biostatistics Working Paper Series
Identification of transcription factor binding sites (regulatory motifs) is a major interest in contemporary biology. We propose a new likelihood based method, COMODE, for identifying structural motifs in DNA sequences. Commonly used methods (e.g. MEME, Gibbs sampler) model binding sites as families of sequences described by a position weight matrix (PWM) and identify PWMs that maximize the likelihood of observed sequence data under a simple multinomial mixture model. This model assumes that the positions of the PWM correspond to independent multinomial distributions with four cell probabilities. We address supervising the search for DNA binding sites using the information derived from …
Partial Auc Estimation And Regression, Lori E. Dodd, Margaret S. Pepe
Partial Auc Estimation And Regression, Lori E. Dodd, Margaret S. Pepe
UW Biostatistics Working Paper Series
Accurate disease diagnosis is critical for health care. New diagnostic and screening tests must be evaluated for their abilities to discriminate disease from non-diseased states. The partial area under the ROC curve (partial AUC) is a measure of diagnostic test accuracy. We present an interpretation of the partial AUC that gives rise to a new non-parametric estimator. This estimator is more robust than existing estimators, which make parametric assumptions. We show that the robustness is gained with only a moderate loss in efficiency. We describe a regression modelling framework for making inference about covariate effects on the partial AUC. Such …
Visualization Methods: A Comparative Study Of New, Traditional And Robust Procedures, Kimberly Crimin
Visualization Methods: A Comparative Study Of New, Traditional And Robust Procedures, Kimberly Crimin
Dissertations
Two major goals in discriminant analysis are discrimination and classification. In discrimination, the goal is to describe graphically (visualization) different features of several known groups. In classification, the goal is to allocate unknown observations to one of several known groups. We have developed new visualization procedures based on traditional estimating procedures and also on robust estimating procedures. We have further developed robust classification procedures. We propose several robust classification procedures based on coordinatewise and affine equivariant, rank-based robust estimates. Empirical studies are performed over many different error distributions. These studies result in empirical efficiencies of the robust and traditional procedures. …
The Maine Coast, A Statistical Source, Maine Coastal Program
The Maine Coast, A Statistical Source, Maine Coastal Program
Maine Collection
The Maine Coast, A Statistical Source
Maine Coastal Program, Natural Resource Planning Division, Maine State Planning Office , Augusta, Maine
First Printing June 1978 - Second Printing September 1978
Contents: Preface / Introduction / Chapter 1 - Demography / Chapter 2 - Land Use and Taxation / Chapter 3 - Economy / Chapter 4 - Housing / Chapter 5 - Transportation / Chapter 6 - Education / Chapter 7 - Recreation / Chapter 8 - Social Services / Chapter 9 - Natural Resources / References / Index