Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Medicine and Health Sciences (23)
- Public Health (22)
- Statistical Models (21)
- Statistical Methodology (16)
- Statistical Theory (16)
-
- Longitudinal Data Analysis and Time Series (15)
- Epidemiology (14)
- Genetics and Genomics (12)
- Life Sciences (12)
- Survival Analysis (11)
- Bioinformatics (10)
- Computational Biology (10)
- Clinical Epidemiology (9)
- Genetics (8)
- Biostatistics (7)
- Applied Mathematics (6)
- Clinical Trials (6)
- Numerical Analysis and Computation (6)
- Multivariate Analysis (4)
- Disease Modeling (3)
- Diseases (3)
- Health Services Research (3)
- Vital and Health Statistics (3)
- Applied Statistics (2)
- Biometry (2)
- Design of Experiments and Sample Surveys (2)
- Medical Specialties (2)
- Keyword
-
- Genetics (5)
- Current status data (3)
- Hierarchical models (3)
- Bioinformatics (2)
- Diagnostic test (2)
-
- Edgeworth expansion (2)
- Nonparametric maximum likelihood estimation (2)
- ROC curve (2)
- Sensitivity (2)
- Skewness (2)
- Stroke (2)
- Annual percent change (APC) (1)
- ADL (1)
- Age-adjusted cancer rates (1)
- Aging (1)
- Air pollution (1)
- As-treated analysis; Per-protocol analysis; Causal inference; Instrumental variables; Principal stratification; Propensity scores (1)
- Asymptotic bias and variance; Clustered survival data; Efficiency; Estimating equation; Kernel smoothing; Marginal model; Sandwich estimator (1)
- Asymptotic bias; EM algorithm; Maximum likelihood estimator; Measurement error; Structural modeling; Transitional Models (1)
- Asymptotic efficiency; Conditional score method; Functional modeling; Measurement error; Longitudinal data; Semiparametric inference; Transition models (1)
- Asymptotics; Augmented kernel estimating equations; Double robustness; Efficiency; Inverse probability weighted kernel estimating equations; Kernel smoothing (1)
- B-splines (1)
- Bayesian Methods (1)
- Bayesian analysis (1)
- Bayesian model (1)
- Bayesian modeling (1)
- Bias (1)
- Binary outcome (1)
- Binary outcomes; Copulas; Marginal likelihood; Multivariate logit; Multivariate probit: Stable distributions (1)
- Binary regression (1)
- Publication Year
Articles 1 - 30 of 49
Full-Text Articles in Categorical Data Analysis
Default Priors For The Intercept Parameter In Logistic Regressions, Philip S. Boonstra, Ryan P. Barbaro, Ananda Sen
Default Priors For The Intercept Parameter In Logistic Regressions, Philip S. Boonstra, Ryan P. Barbaro, Ananda Sen
The University of Michigan Department of Biostatistics Working Paper Series
In logistic regression, separation refers to the situation in which a linear combination of predictors perfectly discriminates the binary outcome. Because finite-valued maximum likelihood parameter estimates do not exist under separation, Bayesian regressions with informative shrinkage of the regression coefficients offer a suitable alternative. Little focus has been given on whether and how to shrink the intercept parameter. Based upon classical studies of separation, we argue that efficiency in estimating regression coefficients may vary with the intercept prior. We adapt alternative prior distributions for the intercept that downweight implausibly extreme regions of the parameter space rendering less sensitivity to separation. …
Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang
Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang
COBRA Preprint Series
Non-negative matrix factorization (NMF) is a widely used machine learning algorithm for dimension reduction of large-scale data. It has found successful applications in a variety of fields such as computational biology, neuroscience, natural language processing, information retrieval, image processing and speech recognition. In bioinformatics, for example, it has been used to extract patterns and profiles from genomic and text-mining data as well as in protein sequence and structure analysis. While the scientific performance of NMF is very promising in dealing with high dimensional data sets and complex data structures, its computational cost is high and sometimes could be critical for …
Models For Hsv Shedding Must Account For Two Levels Of Overdispersion, Amalia Magaret
Models For Hsv Shedding Must Account For Two Levels Of Overdispersion, Amalia Magaret
UW Biostatistics Working Paper Series
We have frequently implemented crossover studies to evaluate new therapeutic interventions for genital herpes simplex virus infection. The outcome measured to assess the efficacy of interventions on herpes disease severity is the viral shedding rate, defined as the frequency of detection of HSV on the genital skin and mucosa. We performed a simulation study to ascertain whether our standard model, which we have used previously, was appropriately considering all the necessary features of the shedding data to provide correct inference. We simulated shedding data under our standard, validated assumptions and assessed the ability of 5 different models to reproduce the …
Computational Model For Survey And Trend Analysis Of Patients With Endometriosis : A Decision Aid Tool For Ebm, Salvo Reina, Vito Reina, Franco Ameglio, Mauro Costa, Alessandro Fasciani
Computational Model For Survey And Trend Analysis Of Patients With Endometriosis : A Decision Aid Tool For Ebm, Salvo Reina, Vito Reina, Franco Ameglio, Mauro Costa, Alessandro Fasciani
COBRA Preprint Series
Endometriosis is increasingly collecting worldwide attention due to its medical complexity and social impact. The European community has identified this as a “social disease”. A large amount of information comes from scientists, yet several aspects of this pathology and staging criteria need to be clearly defined on a suitable number of individuals. In fact, available studies on endometriosis are not easily comparable due to a lack of standardized criteria to collect patients’ informations and scarce definitions of symptoms. Currently, only retrospective surgical stadiation is used to measure pathology intensity, while the Evidence Based Medicine (EBM) requires shareable methods and correct …
The Net Reclassification Index (Nri): A Misleading Measure Of Prediction Improvement With Miscalibrated Or Overfit Models, Margaret Pepe, Jin Fang, Ziding Feng, Thomas Gerds, Jorgen Hilden
The Net Reclassification Index (Nri): A Misleading Measure Of Prediction Improvement With Miscalibrated Or Overfit Models, Margaret Pepe, Jin Fang, Ziding Feng, Thomas Gerds, Jorgen Hilden
UW Biostatistics Working Paper Series
The Net Reclassification Index (NRI) is a very popular measure for evaluating the improvement in prediction performance gained by adding a marker to a set of baseline predictors. However, the statistical properties of this novel measure have not been explored in depth. We demonstrate the alarming result that the NRI statistic calculated on a large test dataset using risk models derived from a training set is likely to be positive even when the new marker has no predictive information. A related theoretical example is provided in which a miscalibrated risk model that includes an uninformative marker is proven to erroneously …
A Unified Approach To Non-Negative Matrix Factorization And Probabilistic Latent Semantic Indexing, Karthik Devarajan, Guoli Wang, Nader Ebrahimi
A Unified Approach To Non-Negative Matrix Factorization And Probabilistic Latent Semantic Indexing, Karthik Devarajan, Guoli Wang, Nader Ebrahimi
COBRA Preprint Series
Non-negative matrix factorization (NMF) by the multiplicative updates algorithm is a powerful machine learning method for decomposing a high-dimensional nonnegative matrix V into two matrices, W and H, each with nonnegative entries, V ~ WH. NMF has been shown to have a unique parts-based, sparse representation of the data. The nonnegativity constraints in NMF allow only additive combinations of the data which enables it to learn parts that have distinct physical representations in reality. In the last few years, NMF has been successfully applied in a variety of areas such as natural language processing, information retrieval, image processing, speech recognition …
A Unified Approach To Modeling Multivariate Binary Data Using Copulas Over Partitions, Bruce J. Swihart, Brian Caffo, Ciprian Crainiceanu
A Unified Approach To Modeling Multivariate Binary Data Using Copulas Over Partitions, Bruce J. Swihart, Brian Caffo, Ciprian Crainiceanu
Johns Hopkins University, Dept. of Biostatistics Working Papers
Many seemingly disparate approaches for marginal modeling have been developed in recent years. We demonstrate that many current approaches for marginal modeling of correlated binary outcomes produce likelihoods that are equivalent to the proposed copula-based models herein. These general copula models of underlying latent threshold random variables yield likelihood based models for marginal fixed effects estimation and interpretation in the analysis of correlated binary data. Moreover, we propose a nomenclature and set of model relationships that substantially elucidates the complex area of marginalized models for binary data. A diverse collection of didactic mathematical and numerical examples are given to illustrate …
Powerful Snp Set Analysis For Case-Control Genome Wide Association Studies, Michael C. Wu, Peter Kraft, Michael P. Epstein, Deanne M. Taylor, Stephen J. Chanock, David J. Hunter, Xihong Lin
Powerful Snp Set Analysis For Case-Control Genome Wide Association Studies, Michael C. Wu, Peter Kraft, Michael P. Epstein, Deanne M. Taylor, Stephen J. Chanock, David J. Hunter, Xihong Lin
Harvard University Biostatistics Working Paper Series
No abstract provided.
The Linkset Model For 2^N Contingency Tables, Mikel Aickin
The Linkset Model For 2^N Contingency Tables, Mikel Aickin
COBRA Preprint Series
Abstract The linkset model is defined for parametrizing the general 2^n contingency table. The linkset parameters are designed to represent latent influences that promote the co-occurrences of binary events beyond that explained by chance. Linkages involving 2 through n binary variables are included in this parametrization. The intent of this process is to elucidate the patterns of linkage, no matter how complex they might be, rather than to fit simplifying models. The relationship between linkset parameters and the natural parameters for a 2n table are derived, and large sample inference methods are provided. Examples are given from medical diagnostics, survival …
Nonparametric Regression With Missing Outcomes Using Weighted Kernel Estimating Equations, Lu Wang, Andrea Rotnitzky, Xihong Lin
Nonparametric Regression With Missing Outcomes Using Weighted Kernel Estimating Equations, Lu Wang, Andrea Rotnitzky, Xihong Lin
Harvard University Biostatistics Working Paper Series
No abstract provided.
A New Class Of Minimum Power Divergence Estimators With Applications To Cancer Surveillance, Nirian Martin, Yi Li
A New Class Of Minimum Power Divergence Estimators With Applications To Cancer Surveillance, Nirian Martin, Yi Li
Harvard University Biostatistics Working Paper Series
No abstract provided.
A Small Sample Correction For Estimating Attributable Risk In Case-Control Studies, Daniel B. Rubin
A Small Sample Correction For Estimating Attributable Risk In Case-Control Studies, Daniel B. Rubin
U.C. Berkeley Division of Biostatistics Working Paper Series
The attributable risk, often called the population attributable risk, is in many epidemiological contexts a more relevant measure of exposure-disease association than the excess risk, relative risk, or odds ratio. When estimating attributable risk with case-control data and a rare disease, we present a simple correction to the standard approach making it essentially unbiased, and also less noisy. As with analogous corrections given in Jewell (1986) for other measures of association, the adjustment often won't make a substantial difference unless the sample size is very small or point estimates are desired within fine strata, but we discuss the possible utility …
Multilevel Latent Class Models With Dirichlet Mixing Distribution, Chongzhi Di, Karen Bandeen-Roche
Multilevel Latent Class Models With Dirichlet Mixing Distribution, Chongzhi Di, Karen Bandeen-Roche
Johns Hopkins University, Dept. of Biostatistics Working Papers
Latent class analysis (LCA) and latent class regression (LCR) are widely used for modeling multivariate categorical outcomes in social sciences and biomedical studies. Standard analyses assume data of different respondents to be mutually independent, excluding application of the methods to familial and other designs in which participants are clustered. In this paper, we develop multilevel latent class model, in which subpopulation mixing probabilities are treated as random effects that vary among clusters according to a common Dirichlet distribution. We apply the Expectation-Maximization (EM) algorithm for model fitting by maximum likelihood (ML). This approach works well, but is computationally intensive when …
Measurement Error Caused By Spatial Misalignment In Environmental Epidemiology, Alexandros Gryparis, Christopher J. Paciorek, Ariana Zeka, Joel Schwartz, Brent A. Coull
Measurement Error Caused By Spatial Misalignment In Environmental Epidemiology, Alexandros Gryparis, Christopher J. Paciorek, Ariana Zeka, Joel Schwartz, Brent A. Coull
Harvard University Biostatistics Working Paper Series
No abstract provided.
A Censored Multinomial Regression Model For Perinatal Mother To Child Transmission Of Hiv, Charlotte C. Gard, Elizabeth R. Brown
A Censored Multinomial Regression Model For Perinatal Mother To Child Transmission Of Hiv, Charlotte C. Gard, Elizabeth R. Brown
UW Biostatistics Working Paper Series
In studies designed to estimate rates of perinatal mother to child transmission of HIV, HIV assays are scheduled at multiple points in time. Still infection status for some infants at some time points is often unknown, particularly when interim analyses are conducted. Logistic regression and Cox proportional hazards regression are commonly used to estimate covariate-adjusted transmission rates, but their methods for handling missing data may be inadequate. Here, we propose using censored multinomial regression models to estimate cumulative and conditional rates of HIV transmission. Through simulation, we show that the proposed methods perform better than standard logistic models in terms …
Estimating Time-To-Event From Longitudinal Categorical Data Using Random Effects Markov Models: Application To Multiple Sclerosis Progression, Micha Mandel, Rebecca A. Betensky
Estimating Time-To-Event From Longitudinal Categorical Data Using Random Effects Markov Models: Application To Multiple Sclerosis Progression, Micha Mandel, Rebecca A. Betensky
Harvard University Biostatistics Working Paper Series
No abstract provided.
Structural Inference In Transition Measurement Error Models For Longitudinal Data, Wenqin Pan, Xihong Lin, Donglin Zeng
Structural Inference In Transition Measurement Error Models For Longitudinal Data, Wenqin Pan, Xihong Lin, Donglin Zeng
Harvard University Biostatistics Working Paper Series
No abstract provided.
Estimation In Semiparametric Transition Measurement Error Models For Longitudinal Data, Wenqin Pan, Donglin Zeng, Xihong Lin
Estimation In Semiparametric Transition Measurement Error Models For Longitudinal Data, Wenqin Pan, Donglin Zeng, Xihong Lin
Harvard University Biostatistics Working Paper Series
No abstract provided.
Nonparametric Regression Using Local Kernel Estimating Equations For Correlated Failure Time Data, Zhangsheng Yu, Xihong Lin
Nonparametric Regression Using Local Kernel Estimating Equations For Correlated Failure Time Data, Zhangsheng Yu, Xihong Lin
Harvard University Biostatistics Working Paper Series
No abstract provided.
Causal Inference In Hybrid Intervention Trials Involving Treatment Choice, Qi Long, Rod Little, Xihong Lin
Causal Inference In Hybrid Intervention Trials Involving Treatment Choice, Qi Long, Rod Little, Xihong Lin
Harvard University Biostatistics Working Paper Series
No abstract provided.
A Comparison Of Methods For Estimating The Causal Effect Of A Treatment In Randomized Clinical Trials Subject To Noncompliance, Rod Little, Qi Long, Xihong Lin
A Comparison Of Methods For Estimating The Causal Effect Of A Treatment In Randomized Clinical Trials Subject To Noncompliance, Rod Little, Qi Long, Xihong Lin
Harvard University Biostatistics Working Paper Series
No abstract provided.
A Flexible General Class Of Marginal And Conditional Random Intercept Models For Binary Outcomes Using Mixtures Of Normals, Brian Caffo, Ming-Wen An, Charles A. Rohde
A Flexible General Class Of Marginal And Conditional Random Intercept Models For Binary Outcomes Using Mixtures Of Normals, Brian Caffo, Ming-Wen An, Charles A. Rohde
Johns Hopkins University, Dept. of Biostatistics Working Papers
Random intercept models for binary data are useful tools for addressing between subject heterogeneity. Unlike linear models, the non-linearity of link functions used for binary data force a distinction between marginal and conditional interpretations. This distinction is blurred in probit models with a normally distributed random intercept because the resulting model implies a probit marginal link as well. That is, this model is closed in the sense that the distribution associated with the marginal and conditional link functions and the random effect distribution are all of the same family. In this manuscript we explore another family of random intercept models …
Causal Comparisons In Randomized Trials Of Two Active Treatments: The Effect Of Supervised Exercise To Promote Smoking Cessation, Jason Roy, Joseph W. Hogan
Causal Comparisons In Randomized Trials Of Two Active Treatments: The Effect Of Supervised Exercise To Promote Smoking Cessation, Jason Roy, Joseph W. Hogan
COBRA Preprint Series
In behavioral medicine trials, such as smoking cessation trials, two or more active treatments are often compared. Noncompliance by some subjects with their assigned treatment poses a challenge to the data analyst. Causal parameters of interest might include those defined by subpopulations based on their potential compliance status under each assignment, using the principal stratification framework (e.g., causal effect of new therapy compared to standard therapy among subjects that would comply with either intervention). Even if subjects in one arm do not have access to the other treatment(s), the causal effect of each treatment typically can only be identified from …
A Computationally Tractable Multivariate Random Effects Model For Clustered Binary Data, Brent A. Coull, E. Andres Houseman, Rebecca A. Betensky
A Computationally Tractable Multivariate Random Effects Model For Clustered Binary Data, Brent A. Coull, E. Andres Houseman, Rebecca A. Betensky
Harvard University Biostatistics Working Paper Series
No abstract provided.
A Diagnostic Test For The Mixing Distribution In A Generalised Linear Mixed Model, Eric J. Tchetgen, Brent A. Coull
A Diagnostic Test For The Mixing Distribution In A Generalised Linear Mixed Model, Eric J. Tchetgen, Brent A. Coull
Harvard University Biostatistics Working Paper Series
We introduce a diagnostic test for the mixing distribution in a generalised linear mixed model. The test is based on the difference between the marginal maximum likelihood and conditional maximum likelihood estimates of a subset of the fixed effects in the model. We derive the asymptotic variance of this difference, and propose a test statistic that has a limiting chi-square distribution under the null hypothesis that the mixing distribution is correctly specified. For the important special case of the logistic regression model with random intercepts, we evaluate via simulation the power of the test in finite samples under several alternative …
Feature-Specific Penalized Latent Class Analysis For Genomic Data, E. Andres Houseman, Brent A. Coull, Rebecca A. Betensky
Feature-Specific Penalized Latent Class Analysis For Genomic Data, E. Andres Houseman, Brent A. Coull, Rebecca A. Betensky
Harvard University Biostatistics Working Paper Series
No abstract provided.
A Nonstationary Negative Binomial Time Series With Time-Dependent Covariates: Enterococcus Counts In Boston Harbor, E. Andres Houseman, Brent Coull, James P. Shine
A Nonstationary Negative Binomial Time Series With Time-Dependent Covariates: Enterococcus Counts In Boston Harbor, E. Andres Houseman, Brent Coull, James P. Shine
Harvard University Biostatistics Working Paper Series
Boston Harbor has had a history of poor water quality, including contamination by enteric pathogens. We conduct a statistical analysis of data collected by the Massachusetts Water Resources Authority (MWRA) between 1996 and 2002 to evaluate the effects of court-mandated improvements in sewage treatment. Motivated by the ineffectiveness of standard Poisson mixture models and their zero-inflated counterparts, we propose a new negative binomial model for time series of Enterococcus counts in Boston Harbor, where nonstationarity and autocorrelation are modeled using a nonparametric smooth function of time in the predictor. Without further restrictions, this function is not identifiable in the presence …
A User-Friendly Introduction To Link-Probit-Normal Models, Brian S. Caffo, Michael Griswold
A User-Friendly Introduction To Link-Probit-Normal Models, Brian S. Caffo, Michael Griswold
Johns Hopkins University, Dept. of Biostatistics Working Papers
Probit-normal models have attractive properties compared to logit-normal models. In particular, they allow for easy specification of marginal links of interest while permitting a conditional random effects structure. Moreover, programming fitting algorithms for probit-normal models can be trivial with the use of well-developed algorithms for approximating multivariate normal quantiles. In typical settings, the data cannot distinguish between probit and logit conditional link functions. Therefore, if marginal interpretations are desired, the default conditional link should be the most convenient one. We refer to models with a probit conditional link an arbitrary marginal link and a normal random effect distribution as link-probit-normal …
Semi-Parametric Single-Index Two-Part Regression Models, Xiao-Hua Zhou, Hua Liang
Semi-Parametric Single-Index Two-Part Regression Models, Xiao-Hua Zhou, Hua Liang
UW Biostatistics Working Paper Series
In this paper, we proposed a semi-parametric single-index two-part regression model to weaken assumptions in parametric regression methods that were frequently used in the analysis of skewed data with additional zero values. The estimation procedure for the parameters of interest in the model was easily implemented. The proposed estimators were shown to be consistent and asymptotically normal. Through a simulation study, we showed that the proposed estimators have reasonable finite-sample performance. We illustrated the application of the proposed method in one real study on the analysis of health care costs.
The Proportional Odds Model For Assessing Rater Agreement With Multiple Modalities, Elizabeth Garrett-Mayer, Steven N. Goodman, Ralph H. Hruban
The Proportional Odds Model For Assessing Rater Agreement With Multiple Modalities, Elizabeth Garrett-Mayer, Steven N. Goodman, Ralph H. Hruban
Johns Hopkins University, Dept. of Biostatistics Working Papers
In this paper, we develop a model for evaluating an ordinal rating systems where we assume that the true underlying disease state is continuous in nature. Our approach in motivated by a dataset with 35 microscopic slides with 35 representative duct lesions of the pancreas. Each of the slides was evaluated by eight raters using two novel rating systems (PanIN illustrations and PanIN nomenclature),where each rater used each systems to rate the slide with slide identity masked between evaluations. We find that the two methods perform equally well but that differentiation of higher grade lesions is more consistent across raters …