Open Access. Powered by Scholars. Published by Universities.®

Physical Sciences and Mathematics Commons

Open Access. Powered by Scholars. Published by Universities.®

Articles 1 - 18 of 18

Full-Text Articles in Physical Sciences and Mathematics

Robust Likelihood-Based Analysis Of Multivariate Data With Missing Values, Rod Little, An Hyonggin Dec 2003

Robust Likelihood-Based Analysis Of Multivariate Data With Missing Values, Rod Little, An Hyonggin

The University of Michigan Department of Biostatistics Working Paper Series

The model-based approach to inference from multivariate data with missing values is reviewed. Regression prediction is most useful when the covariates are predictive of the missing values and the probability of being missing, and in these circumstances predictions are particularly sensitive to model misspecification. The use of penalized splines of the propensity score is proposed to yield robust model-based inference under the missing at random (MAR) assumption, assuming monotone missing data. Simulation comparisons with other methods suggest that the method works well in a wide range of populations, with little loss of efficiency relative to parametric models when the latter …


Weighting Adjustments For Unit Nonresponse With Multiple Outcome Variables, Sonya L. Vartivarian, Rod Little Nov 2003

Weighting Adjustments For Unit Nonresponse With Multiple Outcome Variables, Sonya L. Vartivarian, Rod Little

The University of Michigan Department of Biostatistics Working Paper Series

Weighting is a common form of unit nonresponse adjustment in sample surveys where entire questionnaires are missing due to noncontact or refusal to participate. Weights are inversely proportional to the probability of selection and response. A common approach computes the response weight adjustment cells based on covariate information. When the number of cells thus created is too large, a coarsening method such as response propensity stratification can be applied to reduce the number of adjustment cells. Simulations in Vartivarian and Little (2002) indicate improved efficiency and robustness of weighting adjustments based on the joint classification of the sample by two …


To Model Or Not To Model? Competing Modes Of Inference For Finite Population Sampling, Rod Little Nov 2003

To Model Or Not To Model? Competing Modes Of Inference For Finite Population Sampling, Rod Little

The University of Michigan Department of Biostatistics Working Paper Series

Finite population sampling is perhaps the only area of statistics where the primary mode of analysis is based on the randomization distribution, rather than on statistical models for the measured variables. This article reviews the debate between design and model-based inference. The basic features of the two approaches are illustrated using the case of inference about the mean from stratified random samples. Strengths and weakness of design-based and model-based inference for surveys are discussed. It is suggested that models that take into account the sample design and make weak parametric assumptions can produce reliable and efficient inferences in surveys settings. …


Maximum Likelihood Estimation Of Ordered Multinomial Parameters , Nicholas P. Jewell, Jack Kalbfleisch Oct 2003

Maximum Likelihood Estimation Of Ordered Multinomial Parameters , Nicholas P. Jewell, Jack Kalbfleisch

The University of Michigan Department of Biostatistics Working Paper Series

The pool-adjacent violator-algorithm (Ayer et al., 1955) has long been known to give the maximum likelihood estimator of a series of ordered binomial parameters, based on an independent observation from each distribution (see, Barlow et al., 1972). This result has immediate application to estimation of a survival distribution based on current survival status at a set of monitoring times. This paper considers an extended problem of maximum likelihood estimation of a series of ‘ordered’ multinomial parameters pi = (p1i, p2i, . . . , pmi) for 1 < = I < = k, where ordered means that pj1 < = pj2 < = .. . < = pjk for each j with 1 < = j < = m-1. The data consist of k independent observations X1, . . . ,Xk where Xi has a multinomial distribution with probability parameter pi and known index ni > = 1. By making use of variants of the pool adjacent violator algorithm, …


A Population Pharmacokinetic Model With Time-Dependent Covariates Measured With Errors, Lang Lil, Xihong Lin, Mort B. Brown, Suneel Gupta, Kyung-Hoon Lee Oct 2003

A Population Pharmacokinetic Model With Time-Dependent Covariates Measured With Errors, Lang Lil, Xihong Lin, Mort B. Brown, Suneel Gupta, Kyung-Hoon Lee

The University of Michigan Department of Biostatistics Working Paper Series

We propose a population pharmacokinetic (PK) model with time-dependent covariates measured with errors. This model is used to model S-oxybutynin's kinetics following an oral administration of Ditropan, and allows the distribution rate to depend on time-dependent covariates blood pressure and heart rate, which are measured with errors. We propose two two-step estimation methods: the second order two-step method with numerical solutions of differential equations (2orderND), and the second order two-step method with closed form approximate solutions of differential equations (2orderAD). The proposed methods are computationally easy and require fitting a linear mixed model at the first step and a nonlinear …


Equivalent Kernels Of Smoothing Splines In Nonparametric Regression For Clustered/Longitudinal Data, Xihong Lin, Naisyin Wang, Alan H. Welsh, Raymond J. Carroll Sep 2003

Equivalent Kernels Of Smoothing Splines In Nonparametric Regression For Clustered/Longitudinal Data, Xihong Lin, Naisyin Wang, Alan H. Welsh, Raymond J. Carroll

The University of Michigan Department of Biostatistics Working Paper Series

We compare spline and kernel methods for clustered/longitudinal data. For independent data, it is well known that kernel methods and spline methods are essentially asymptotically equivalent (Silverman, 1984). However, the recent work of Welsh, et al. (2002) shows that the same is not true for clustered/longitudinal data. First, conventional kernel methods fail to account for the within- cluster correlation, while spline methods are able to account for this correlation. Second, kernel methods and spline methods were found to have different local behavior, with conventional kernels being local and splines being non-local. To resolve these differences, we show that a smoothing …


Histospline Method In Nonparametric Regression Models With Application To Clustered/Longitudinal Data, Raymond J. Carroll, Peter Hall, Tatiyana V. Apanasovich, Xihong Lin Sep 2003

Histospline Method In Nonparametric Regression Models With Application To Clustered/Longitudinal Data, Raymond J. Carroll, Peter Hall, Tatiyana V. Apanasovich, Xihong Lin

The University of Michigan Department of Biostatistics Working Paper Series

Kernel and smoothing methods for nonparametric function and curve estimation have been particularly successful in "standard" settings, where function values are observed subject to independent errors. However, when aspects of the function are known parametrically, or where the sampling scheme has significant structure, it can be quite difficult to adapt standard methods in such a way that they retain good statistical performance and continue to enjoy easy computability and good numerical properties. In particular, when using local linear modeling it is often awkward to both respect the sampling scheme and produce an estimator with good variance properties, without resorting to …


Efficient Semiparametric Marginal Estimation For Longitudinal/Clustered Data, Naisyin Wang, Raymond J. Carroll, Xihong Lin Sep 2003

Efficient Semiparametric Marginal Estimation For Longitudinal/Clustered Data, Naisyin Wang, Raymond J. Carroll, Xihong Lin

The University of Michigan Department of Biostatistics Working Paper Series

We consider marginal generalized semiparametric partially linear models for clustered data. Lin and Carroll (2001a) derived the semiparametric efficinet score funtion for this problem in the mulitvariate Gaussian case, but they were unable to contruct a semiparametric efficient estimator that actually achieved the semiparametric information bound. We propose such an estimator here and generalize the work to marginal generalized partially liner models. Asymptotic relative efficincies of the estimation or throughout are investigated. The finite sample performance of these estimators is evaluated through simulations and illustrated using a longtiudinal CD4 count data set. Both theoretical and numerical results indicate that properly …


A Varying-Coefficient Cox Model For The Effect Of Age At A Marker Event On Age At Menopause, Bin Nan, Xihong Lin, Lynda D. Lisabeth, Sioban D. Harlow Sep 2003

A Varying-Coefficient Cox Model For The Effect Of Age At A Marker Event On Age At Menopause, Bin Nan, Xihong Lin, Lynda D. Lisabeth, Sioban D. Harlow

The University of Michigan Department of Biostatistics Working Paper Series

. It is of recent interest in reproductive health research to investigate the validity of a marker event for the onset of menopausal transition and to estimate age at menopause using age at the marker event. We propose a varying coefficient Cox model to investigate the association between age at a marker event, denned as a specific bleeding pattern change, and age at menopause, where both events are subject to censoring and their association varies with age at the marker event. Estimation proceeds using the regression spline method. The proposed method is applied to the Tremin Trust Data to evaluate …


An Extended General Location Model For Causal Inference From Data Subject To Noncompliance And Missing Values, Yahong Peng, Rod Little, Trivellore E. Raghuanthan Aug 2003

An Extended General Location Model For Causal Inference From Data Subject To Noncompliance And Missing Values, Yahong Peng, Rod Little, Trivellore E. Raghuanthan

The University of Michigan Department of Biostatistics Working Paper Series

Noncompliance is a common problem in experiments involving randomized assignment of treatments, and standard analyses based on intention-to treat or treatment received have limitations. An attractive alternative is to estimate the Complier-Average Causal Effect (CACE), which is the average treatment effect for the subpopulation of subjects who would comply under either treatment (Angrist, Imbens and Rubin, 1996, henceforth AIR). We propose an Extended General Location Model to estimate the CACE from data with non-compliance and missing data in the outcome and in baseline covariates. Models for both continuous and categorical outcomes and ignorable and latent ignorable (Frangakis and Rubin, 1999) …


On The Formation Of Weighting Adjustment Cells For Unit Nonresponse, Sonya Vartivarian, Rod Little Aug 2003

On The Formation Of Weighting Adjustment Cells For Unit Nonresponse, Sonya Vartivarian, Rod Little

The University of Michigan Department of Biostatistics Working Paper Series

A method is proposed for weighting adjustments for unit nonresponse based on a crossclassification by the estimated propensity to respond and by the predicted mean of a survey outcome. Simulations to assess the performance of the method are described.


Inference For The Population Total From Probability-Proportional-To-Size Samples Based On Predictions From A Penalized Spline Nonparametric Model, Hui Zheng, Rod Little Aug 2003

Inference For The Population Total From Probability-Proportional-To-Size Samples Based On Predictions From A Penalized Spline Nonparametric Model, Hui Zheng, Rod Little

The University of Michigan Department of Biostatistics Working Paper Series

Inference about the finite population total from probability-proportional-to-size (PPS) samples is considered. In previous work (Zheng and Little, 2003), penalized spline (p-spline) nonparametric model-based estimators were shown to generally outperform the Horvitz-Thompson (HT) and generalized regression (GR) estimators in terms of the root mean squared error. In this article we develop model-based, jackknife and balanced repeated replicate variance estimation methods for the p-spline based estimators. Asymptotic properties of the jackknife method are discussed. Simulations show that p-spline point estimators and their jackknife standard errors lead to inferences that are superior to HT or GR based inferences. This suggests that nonparametric …


Maximization By Parts In Likelihood Inference, Peter Xuekun Song, Yanqin Fan, Jack Kalbfleisch Jun 2003

Maximization By Parts In Likelihood Inference, Peter Xuekun Song, Yanqin Fan, Jack Kalbfleisch

The University of Michigan Department of Biostatistics Working Paper Series

This paper presents and examines a new algorithm for solving a score equation for the maximum likelyhood estimate in certain problems of practical interest. The method circumvents the need to compute second order derivaties of the full likelihood function. It exploits the structure of certain models that yield a natural decomposition of a very complicated likelihood function. In this decomposition, the first part is a log likelihood from a simply analyzed model and the second part is used to update estimates from the first. Convergence properties of this fixed point algorithm are examined and asymptotics are derived for estimators obtained …


Cluster Stability Scores For Microarray Data In Cancer Studies, Mark Smolkin, Debashis Ghosh Jun 2003

Cluster Stability Scores For Microarray Data In Cancer Studies, Mark Smolkin, Debashis Ghosh

The University of Michigan Department of Biostatistics Working Paper Series

A potential benefit of profiling of tissue samples using microarrays is the generation of molecular fingerprints that will define subtypes of disease. Hierarchical clustering has been the primary analytical tool used to define disease subtypes from microarray experiments in cancer settings. Assessing cluster reliability poses a major complication in analyzing output from these procedures. While much work has been done on assessing the global question of number of clusters in a dataset, relatively little research exists on assessing stability of individual clusters. A potential benefit of profiling of tissue samples using microarrays is the generation of molecular fingerprints that will …


Mixtures Of Varying Coefficient Models For Longitudinal Data With Discrete Or Continuous Non-Ignorable Dropout, Joseph W. Hogan, Xihong Lin, Benjamin A. Herman May 2003

Mixtures Of Varying Coefficient Models For Longitudinal Data With Discrete Or Continuous Non-Ignorable Dropout, Joseph W. Hogan, Xihong Lin, Benjamin A. Herman

The University of Michigan Department of Biostatistics Working Paper Series

The analysis of longitudinal repeated measures data is frequently complicated by missing data due to informative dropout. We describe a mixture model for joint distribution for longitudinal repeated measures, where the dropout distribution may be continuous and the dependence between response and dropout is semiparametric. Specifically, we assume that responses follow a varying coefficient random effects model conditional on dropout time, where the regression coefficients depend on dropout time through unspecified nonparametric functions that are estimated using step functions when dropout time is discrete (e.g., for panel data) and using smoothing splines when dropout time is continuous. Inference under the …


Semiparametric Regression Models With Missing Data: The Mathematics In The Work Of Robins Et Al., Menggang Yu, Bin Nan May 2003

Semiparametric Regression Models With Missing Data: The Mathematics In The Work Of Robins Et Al., Menggang Yu, Bin Nan

The University of Michigan Department of Biostatistics Working Paper Series

This review is an attempt to understand the landmark papers of Robins, Rotnitzky, and Zhao (1994) and Robins and Rotnitzky (1992). We revisit their main results and corresponding proofs using the theory outlined in the monograph by Bickel, Klaassen, Ritov, and Wellner (1993). We also discuss an illustrative example to show the details of applying these theoretical results.


Penalized Spline Nonparametric Mixed Models For Inference About A Finite Population Mean From Two-Stage Samples, Hui Zheng, Rod Little Mar 2003

Penalized Spline Nonparametric Mixed Models For Inference About A Finite Population Mean From Two-Stage Samples, Hui Zheng, Rod Little

The University of Michigan Department of Biostatistics Working Paper Series

Samplers often distrust model-based approaches to survey inference due to concerns about model misspecification when applied to large samples from complex populations. We suggest that the model-based paradigm can work very successfully in survey settings, provided models are chosen that take into account the sample design and avoid strong parametric assumptions. The Horvitz-Thompson (HT) estimator is a simple design-unbiased estimator of the finite population total in probability sampling designs. From a modeling perspective, the HT estimator performs well when the ratios of the outcome values and the inclusion probabilities are exchangeable. When this assumption is not met, the HT estimator …


Selective Multiple Imputation Of Keys For Statistical Disclosure Control In Microdata, Rod Little, Fang Liu Jan 2003

Selective Multiple Imputation Of Keys For Statistical Disclosure Control In Microdata, Rod Little, Fang Liu

The University of Michigan Department of Biostatistics Working Paper Series

The fundamental tension in statistical disclosure control (SDC) of microdata is the trade-off between the protection of individual respondents and the release of enough information for statistical inferences. We consider microdata that include key variables that contain identifying information and target variables that include sensitive information. Releasing the original data may expose some individuals in the sample to high risk of disclosure; deleting key variables is a common approach, but this loses information for some statistical analysis. This paper proposes selective multiple imputation of key variables (SMIKe) as an alternative SDC technique between those two extremes, and applies SMIKe to …