Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Statistical Methodology (27)
- Statistical Theory (27)
- Survival Analysis (17)
- Medicine and Health Sciences (16)
- Longitudinal Data Analysis and Time Series (14)
-
- Public Health (13)
- Applied Mathematics (7)
- Design of Experiments and Sample Surveys (7)
- Numerical Analysis and Computation (7)
- Epidemiology (6)
- Categorical Data Analysis (5)
- Clinical Epidemiology (5)
- Clinical Trials (5)
- Medical Specialties (5)
- Genetics (4)
- Genetics and Genomics (4)
- Life Sciences (4)
- Health Services Research (3)
- Multivariate Analysis (3)
- Biostatistics (2)
- Disease Modeling (2)
- Diseases (2)
- Microarrays (2)
- Endocrinology, Diabetes, and Metabolism (1)
- Experimental Analysis of Behavior (1)
- Health Psychology (1)
- Institutional and Historical (1)
- Keyword
-
- Prediction (5)
- Longitudinal data (4)
- Model selection (4)
- Generalized estimating equations (3)
- Hierarchical models (3)
-
- Sensitivity (3)
- Air pollution (2)
- Causal inference (2)
- Censored data (2)
- Clustered data (2)
- Cross validation (2)
- Cross-validation (2)
- Disease screening (2)
- Double robustness (2)
- Estimating equation (2)
- Frailty (2)
- Health expenditures (2)
- Joint model (2)
- Log-normal (2)
- Mixture model (2)
- Nonparametric regression (2)
- Q-Q plots (2)
- Regression (2)
- Regression models (2)
- Regression splines (2)
- Semiparametric model (2)
- Skewed distributions (2)
- Smoking (2)
- Smoothing splines (2)
- Specificity (2)
- Publication
-
- Johns Hopkins University, Dept. of Biostatistics Working Papers (15)
- The University of Michigan Department of Biostatistics Working Paper Series (12)
- UW Biostatistics Working Paper Series (11)
- U.C. Berkeley Division of Biostatistics Working Paper Series (8)
- Harvard University Biostatistics Working Paper Series (6)
- Publication Type
Articles 31 - 54 of 54
Full-Text Articles in Statistical Models
Stochastic Models Based On Molecular Hybridization Theory For Short Oligonucleotide Microarrays, Zhijin Wu, Richard Leblanc, Rafael A. Irizarry
Stochastic Models Based On Molecular Hybridization Theory For Short Oligonucleotide Microarrays, Zhijin Wu, Richard Leblanc, Rafael A. Irizarry
Johns Hopkins University, Dept. of Biostatistics Working Papers
High density oligonucleotide expression arrays are a widely used tool for the measurement of gene expression on a large scale. Affymetrix GeneChip arrays appear to dominate this market. These arrays use short oligonucleotides to probe for genes in an RNA sample. Due to optical noise, non-specific hybridization, probe-specific effects, and measurement error, ad-hoc measures of expression, that summarize probe intensities, can lead to imprecise and inaccurate results. Various researchers have demonstrated that expression measures based on simple statistical models can provide great improvements over the ad-hoc procedure offered by Affymetrix. Recently, physical models based on molecular hybridization theory, have been …
Efficient Semiparametric Marginal Estimation For Longitudinal/Clustered Data, Naisyin Wang, Raymond J. Carroll, Xihong Lin
Efficient Semiparametric Marginal Estimation For Longitudinal/Clustered Data, Naisyin Wang, Raymond J. Carroll, Xihong Lin
The University of Michigan Department of Biostatistics Working Paper Series
We consider marginal generalized semiparametric partially linear models for clustered data. Lin and Carroll (2001a) derived the semiparametric efficinet score funtion for this problem in the mulitvariate Gaussian case, but they were unable to contruct a semiparametric efficient estimator that actually achieved the semiparametric information bound. We propose such an estimator here and generalize the work to marginal generalized partially liner models. Asymptotic relative efficincies of the estimation or throughout are investigated. The finite sample performance of these estimators is evaluated through simulations and illustrated using a longtiudinal CD4 count data set. Both theoretical and numerical results indicate that properly …
Equivalent Kernels Of Smoothing Splines In Nonparametric Regression For Clustered/Longitudinal Data, Xihong Lin, Naisyin Wang, Alan H. Welsh, Raymond J. Carroll
Equivalent Kernels Of Smoothing Splines In Nonparametric Regression For Clustered/Longitudinal Data, Xihong Lin, Naisyin Wang, Alan H. Welsh, Raymond J. Carroll
The University of Michigan Department of Biostatistics Working Paper Series
We compare spline and kernel methods for clustered/longitudinal data. For independent data, it is well known that kernel methods and spline methods are essentially asymptotically equivalent (Silverman, 1984). However, the recent work of Welsh, et al. (2002) shows that the same is not true for clustered/longitudinal data. First, conventional kernel methods fail to account for the within- cluster correlation, while spline methods are able to account for this correlation. Second, kernel methods and spline methods were found to have different local behavior, with conventional kernels being local and splines being non-local. To resolve these differences, we show that a smoothing …
Histospline Method In Nonparametric Regression Models With Application To Clustered/Longitudinal Data, Raymond J. Carroll, Peter Hall, Tatiyana V. Apanasovich, Xihong Lin
Histospline Method In Nonparametric Regression Models With Application To Clustered/Longitudinal Data, Raymond J. Carroll, Peter Hall, Tatiyana V. Apanasovich, Xihong Lin
The University of Michigan Department of Biostatistics Working Paper Series
Kernel and smoothing methods for nonparametric function and curve estimation have been particularly successful in "standard" settings, where function values are observed subject to independent errors. However, when aspects of the function are known parametrically, or where the sampling scheme has significant structure, it can be quite difficult to adapt standard methods in such a way that they retain good statistical performance and continue to enjoy easy computability and good numerical properties. In particular, when using local linear modeling it is often awkward to both respect the sampling scheme and produce an estimator with good variance properties, without resorting to …
Measuring Treatment Effects Using Semiparametric Models, Zhuo Yu, Mark J. Van Der Laan
Measuring Treatment Effects Using Semiparametric Models, Zhuo Yu, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
In order to estimate the causal effect of treatments on an outcome of interest, one has to account for the effect of confounding factors which covary with the treatments and also contribute to the outcome of interest. In this paper, we use the semiparametric regression model to estimate the causal parameters. We assume the causal effect of the treatments can be described by the parametric component of the semiparametric regression model. Following the general methodology which was developed in van der Laan and Robins (2002) we give the orthogonal complement of the nuisance tangent space which identifies all the estimating …
A Varying-Coefficient Cox Model For The Effect Of Age At A Marker Event On Age At Menopause, Bin Nan, Xihong Lin, Lynda D. Lisabeth, Sioban D. Harlow
A Varying-Coefficient Cox Model For The Effect Of Age At A Marker Event On Age At Menopause, Bin Nan, Xihong Lin, Lynda D. Lisabeth, Sioban D. Harlow
The University of Michigan Department of Biostatistics Working Paper Series
. It is of recent interest in reproductive health research to investigate the validity of a marker event for the onset of menopausal transition and to estimate age at menopause using age at the marker event. We propose a varying coefficient Cox model to investigate the association between age at a marker event, denned as a specific bleeding pattern change, and age at menopause, where both events are subject to censoring and their association varies with age at the marker event. Estimation proceeds using the regression spline method. The proposed method is applied to the Tremin Trust Data to evaluate …
Asymptotically Optimal Model Selection Method With Right Censored Outcomes, Sunduz Keles, Mark J. Van Der Laan, Sandrine Dudoit
Asymptotically Optimal Model Selection Method With Right Censored Outcomes, Sunduz Keles, Mark J. Van Der Laan, Sandrine Dudoit
U.C. Berkeley Division of Biostatistics Working Paper Series
Over the last two decades, non-parametric and semi-parametric approaches that adapt well known techniques such as regression methods to the analysis of right censored data, e.g. right censored survival data, became popular in the statistics literature. However, the problem of choosing the best model (predictor) among a set of proposed models (predictors) in the right censored data setting have not gained much attention. In this paper, we develop a new cross-validation based model selection method to select among predictors of right censored outcomes such as survival times. The proposed method considers the risk of a given predictor based on the …
Tree-Based Multivariate Regression And Density Estimation With Right-Censored Data , Annette M. Molinaro, Sandrine Dudoit, Mark J. Van Der Laan
Tree-Based Multivariate Regression And Density Estimation With Right-Censored Data , Annette M. Molinaro, Sandrine Dudoit, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
We propose a unified strategy for estimator construction, selection, and performance assessment in the presence of censoring. This approach is entirely driven by the choice of a loss function for the full (uncensored) data structure and can be stated in terms of the following three main steps. (1) Define the parameter of interest as the minimizer of the expected loss, or risk, for a full data loss function chosen to represent the desired measure of performance. Map the full data loss function into an observed (censored) data loss function having the same expected value and leading to an efficient estimator …
An Extended General Location Model For Causal Inference From Data Subject To Noncompliance And Missing Values, Yahong Peng, Rod Little, Trivellore E. Raghuanthan
An Extended General Location Model For Causal Inference From Data Subject To Noncompliance And Missing Values, Yahong Peng, Rod Little, Trivellore E. Raghuanthan
The University of Michigan Department of Biostatistics Working Paper Series
Noncompliance is a common problem in experiments involving randomized assignment of treatments, and standard analyses based on intention-to treat or treatment received have limitations. An attractive alternative is to estimate the Complier-Average Causal Effect (CACE), which is the average treatment effect for the subpopulation of subjects who would comply under either treatment (Angrist, Imbens and Rubin, 1996, henceforth AIR). We propose an Extended General Location Model to estimate the CACE from data with non-compliance and missing data in the outcome and in baseline covariates. Models for both continuous and categorical outcomes and ignorable and latent ignorable (Frangakis and Rubin, 1999) …
Inference For The Population Total From Probability-Proportional-To-Size Samples Based On Predictions From A Penalized Spline Nonparametric Model, Hui Zheng, Rod Little
Inference For The Population Total From Probability-Proportional-To-Size Samples Based On Predictions From A Penalized Spline Nonparametric Model, Hui Zheng, Rod Little
The University of Michigan Department of Biostatistics Working Paper Series
Inference about the finite population total from probability-proportional-to-size (PPS) samples is considered. In previous work (Zheng and Little, 2003), penalized spline (p-spline) nonparametric model-based estimators were shown to generally outperform the Horvitz-Thompson (HT) and generalized regression (GR) estimators in terms of the root mean squared error. In this article we develop model-based, jackknife and balanced repeated replicate variance estimation methods for the p-spline based estimators. Asymptotic properties of the jackknife method are discussed. Simulations show that p-spline point estimators and their jackknife standard errors lead to inferences that are superior to HT or GR based inferences. This suggests that nonparametric …
Estimation Of Parameters In Replicated Time Series Regression Models, Genming Shi
Estimation Of Parameters In Replicated Time Series Regression Models, Genming Shi
Mathematics & Statistics Theses & Dissertations
The time series regression model was widely studied in the literature by several authors. However, statistical analysis of replicated time series regression models has received little attention. In this thesis, we study the application of quasi-least squares, a relatively new method, to estimate the parameters in replicated time series models with general ARMA( p, q) correlation structure. We also study several established methods for estimating the parameters in those models, including the maximum likelihood, method of moments, and the GEE method. Asymptotic comparisons of the methods are made bV fixing the number of repeated measurements in each series, and …
Cultural And Psychological Influences On Diabetic Adherence, Keikilani Mcmillin-Williams
Cultural And Psychological Influences On Diabetic Adherence, Keikilani Mcmillin-Williams
Loma Linda University Electronic Theses, Dissertations & Projects
Diabetes mellitus is a serious disease that poses a particular healthcare challenge because progression is considered controllable (Cox, et al, 1985; Vinicor, et al, 1996) yet treatment adherence, and thus outcome, is very poor (Gonder-Frederick, Cox, & Ritterband, 2002; Goodall, 1991). Culture is a lethal risk factor for diabetic contraction and treatment maintenance. Latinos within the United States are two-to-three times more likely to develop complications and die than non-Latinos (Haffner et al, 1996; Rubin, Peyrot, & Saudek, 1991) and are less likely to adhere to treatment (Lipton, Losey, Giachello, Mendez, & Girotti, 1998). Efforts to eliminate health disparities have …
Supervised Detection Of Regulatory Motifs In Dna Sequences, Sunduz Keles, Mark J. Van Der Laan, Sandrine Dudoit, Biao Xing, Michael B. Eisen
Supervised Detection Of Regulatory Motifs In Dna Sequences, Sunduz Keles, Mark J. Van Der Laan, Sandrine Dudoit, Biao Xing, Michael B. Eisen
U.C. Berkeley Division of Biostatistics Working Paper Series
Identification of transcription factor binding sites (regulatory motifs) is a major interest in contemporary biology. We propose a new likelihood based method, COMODE, for identifying structural motifs in DNA sequences. Commonly used methods (e.g. MEME, Gibbs sampler) model binding sites as families of sequences described by a position weight matrix (PWM) and identify PWMs that maximize the likelihood of observed sequence data under a simple multinomial mixture model. This model assumes that the positions of the PWM correspond to independent multinomial distributions with four cell probabilities. We address supervising the search for DNA binding sites using the information derived from …
Linear Models For Microarray Data Analysis: Hidden Similarities And Differences, M. Kathleen Kerr
Linear Models For Microarray Data Analysis: Hidden Similarities And Differences, M. Kathleen Kerr
UW Biostatistics Working Paper Series
In the past several years many linear models have been proposed for analyzing two-color microarray data. As presented in the literature, many of these models appear dramatically different. However, many of these models are reformulations of the same basic approach to analyzing microarray data. This paper demonstrates the equivalence of some of these models. Attention is directed at choices in microarray data analysis that have a larger impact on the results than the choice of linear model.
Mixtures Of Varying Coefficient Models For Longitudinal Data With Discrete Or Continuous Non-Ignorable Dropout, Joseph W. Hogan, Xihong Lin, Benjamin A. Herman
Mixtures Of Varying Coefficient Models For Longitudinal Data With Discrete Or Continuous Non-Ignorable Dropout, Joseph W. Hogan, Xihong Lin, Benjamin A. Herman
The University of Michigan Department of Biostatistics Working Paper Series
The analysis of longitudinal repeated measures data is frequently complicated by missing data due to informative dropout. We describe a mixture model for joint distribution for longitudinal repeated measures, where the dropout distribution may be continuous and the dependence between response and dropout is semiparametric. Specifically, we assume that responses follow a varying coefficient random effects model conditional on dropout time, where the regression coefficients depend on dropout time through unspecified nonparametric functions that are estimated using step functions when dropout time is discrete (e.g., for panel data) and using smoothing splines when dropout time is continuous. Inference under the …
Semiparametric Regression Models With Missing Data: The Mathematics In The Work Of Robins Et Al., Menggang Yu, Bin Nan
Semiparametric Regression Models With Missing Data: The Mathematics In The Work Of Robins Et Al., Menggang Yu, Bin Nan
The University of Michigan Department of Biostatistics Working Paper Series
This review is an attempt to understand the landmark papers of Robins, Rotnitzky, and Zhao (1994) and Robins and Rotnitzky (1992). We revisit their main results and corresponding proofs using the theory outlined in the monograph by Bickel, Klaassen, Ritov, and Wellner (1993). We also discuss an illustrative example to show the details of applying these theoretical results.
Estimating The Accuracy Of Polymerase Chain Reaction-Based Tests Using Endpoint Dilution, Jim Hughes, Patricia Totten
Estimating The Accuracy Of Polymerase Chain Reaction-Based Tests Using Endpoint Dilution, Jim Hughes, Patricia Totten
UW Biostatistics Working Paper Series
PCR-based tests for various microorganisms or target DNA sequences are generally acknowledged to be highly "sensitive" yet the concept of sensitivity is ill-defined in the literature on these tests. We propose that sensitivity should be expressed as a function of the number of target DNA molecules in the sample (or specificity when the target number is 0). However, estimating this "sensitivity curve" is problematic since it is difficult to construct samples with a fixed number of targets. Nonetheless, using serially diluted replicate aliquots of a known concentration of the target DNA sequence, we show that it is possible to disentangle …
Penalized Spline Nonparametric Mixed Models For Inference About A Finite Population Mean From Two-Stage Samples, Hui Zheng, Rod Little
Penalized Spline Nonparametric Mixed Models For Inference About A Finite Population Mean From Two-Stage Samples, Hui Zheng, Rod Little
The University of Michigan Department of Biostatistics Working Paper Series
Samplers often distrust model-based approaches to survey inference due to concerns about model misspecification when applied to large samples from complex populations. We suggest that the model-based paradigm can work very successfully in survey settings, provided models are chosen that take into account the sample design and avoid strong parametric assumptions. The Horvitz-Thompson (HT) estimator is a simple design-unbiased estimator of the finite population total in probability sampling designs. From a modeling perspective, the HT estimator performs well when the ratios of the outcome values and the inclusion probabilities are exchangeable. When this assumption is not met, the HT estimator …
A Semiparametric Model Selection Criterion With Applications To The Marginal Structural Model, M. Alan Brookhart, Mark J. Van Der Laan
A Semiparametric Model Selection Criterion With Applications To The Marginal Structural Model, M. Alan Brookhart, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Estimators for the parameter of interest in semiparametric models often depend on a guessed model for the nuisance parameter. The choice of the model for the nuisance parameter can affect both the finite sample bias and efficiency of the resulting estimator of the parameter of interest. In this paper we propose a finite sample criterion based on cross validation that can be used to select a nuisance parameter model from a list of candidate models. We show that expected value of this criterion is minimized by the nuisance parameter model that yields the estimator of the parameter of interest with …
Rank Regression In Stability Analysis, Ying Qing Chen, Annpey Pong, Biao Xing
Rank Regression In Stability Analysis, Ying Qing Chen, Annpey Pong, Biao Xing
U.C. Berkeley Division of Biostatistics Working Paper Series
Stability data are often collected to determine the shelf-life of certain characteristics of a pharmaceutical product, for example, a drug's potency over time. Statistical approaches such as the linear regression models are considered as appropriate to analyze the stability data. However, most of these regression models in both theory and practice rely heavily on their underlying parametric assumptions, such as normality of the continuous characteristics or their transformations. In this article, we propose and study some rank-based regression procedures for the stability data when the linear regression models are semiparametric with unspecified error structure. Numerical studies including Monte Carlo simulations …
Semiparametric Receiver Operating Characteristic Analysis To Evaluate Biomarkers For Disease, Tianxi Cai, Margaret S. Pepe
Semiparametric Receiver Operating Characteristic Analysis To Evaluate Biomarkers For Disease, Tianxi Cai, Margaret S. Pepe
UW Biostatistics Working Paper Series
The receiver operating characteristic (ROC) curve is a popular method for characterizing the accuracy of diagnostic tests when test results are not binary. Various methodologies for estimating and comparing ROC curves have been developed. One approach, due to Pepe, uses a parametric regression model with the baseline function specified up to a finite-dimensional parameter. In this article we extend the regression models by allowing arbitrary nonparametric baseline functions. We also provide asymptotic distribution theory and procedures for making statistical inference. We illustrate our approach with dataset from a prostate cancer biomarker study. Simulation studies suggest that the extra flexibility inherent …
Semi-Parametric Regression For The Area Under The Receiver Operating Characteristic Curve, Lori E. Dodd, Margaret S. Pepe
Semi-Parametric Regression For The Area Under The Receiver Operating Characteristic Curve, Lori E. Dodd, Margaret S. Pepe
UW Biostatistics Working Paper Series
Medical advances continue to provide new and potentially better means for detecting disease. Such is true in cancer, for example, where biomarkers are sought for early detection and where improvements in imaging methods may pick up the initial functional and molecular changes associated with cancer development. In other binary classification tasks, computational algorithms such as Neural Networks, Support Vector Machines and Evolutionary Algorithms have been applied to areas as diverse as credit scoring, object recognition, and peptide-binding prediction. Before a classifier becomes an accepted technology, it must undergo rigorous evaluation to determine its ability to discriminate between states. Characterization of …
Checking Assumptions In Latent Class Regression Models Via A Markov Chain Monte Carlo Estimation Approach: An Application To Depression And Socio-Economic Status, Elizabeth Garrett, Richard Miech, Pamela Owens, William W. Eaton, Scott L. Zeger
Checking Assumptions In Latent Class Regression Models Via A Markov Chain Monte Carlo Estimation Approach: An Application To Depression And Socio-Economic Status, Elizabeth Garrett, Richard Miech, Pamela Owens, William W. Eaton, Scott L. Zeger
Johns Hopkins University, Dept. of Biostatistics Working Papers
Latent class regression models are useful tools for assessing associations between covariates and latent variables. However, evaluation of key model assumptions cannot be performed using methods from standard regression models due to the unobserved nature of latent outcome variables. This paper presents graphical diagnostic tools to evaluate whether or not latent class regression models adhere to standard assumptions of the model: conditional independence and non-differential measurement. An integral part of these methods is the use of a Markov Chain Monte Carlo estimation procedure. Unlike standard maximum likelihood implementations for latent class regression model estimation, the MCMC approach allows us to …
Partial Auc Estimation And Regression, Lori E. Dodd, Margaret S. Pepe
Partial Auc Estimation And Regression, Lori E. Dodd, Margaret S. Pepe
UW Biostatistics Working Paper Series
Accurate disease diagnosis is critical for health care. New diagnostic and screening tests must be evaluated for their abilities to discriminate disease from non-diseased states. The partial area under the ROC curve (partial AUC) is a measure of diagnostic test accuracy. We present an interpretation of the partial AUC that gives rise to a new non-parametric estimator. This estimator is more robust than existing estimators, which make parametric assumptions. We show that the robustness is gained with only a moderate loss in efficiency. We describe a regression modelling framework for making inference about covariate effects on the partial AUC. Such …