Open Access. Powered by Scholars. Published by Universities.®
Design of Experiments and Sample Surveys Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Statistical Theory (19)
- Statistical Methodology (18)
- Statistical Models (13)
- Medicine and Health Sciences (8)
- Public Health (8)
-
- Biostatistics (7)
- Clinical Trials (7)
- Epidemiology (6)
- Longitudinal Data Analysis and Time Series (5)
- Microarrays (5)
- Survival Analysis (5)
- Genetics and Genomics (4)
- Life Sciences (4)
- Bioinformatics (3)
- Clinical Epidemiology (3)
- Computational Biology (3)
- Genetics (3)
- Applied Statistics (2)
- Categorical Data Analysis (2)
- Health Services Research (2)
- Institutional and Historical (2)
- Multivariate Analysis (2)
- Vital and Health Statistics (2)
- Applied Mathematics (1)
- Biometry (1)
- Disease Modeling (1)
- Diseases (1)
- Keyword
-
- Missing data (3)
- Air pollution (2)
- Block design (2)
- Case-crossover design (2)
- Clinical trials (2)
-
- Diagnostic tests (2)
- Microarray (2)
- Missing at random (2)
- Sampling weights (2)
- Selection bias (2)
- Surrogate outcome (2)
- Weighting (2)
- Adaptive design; asymptotic normality; canonical distribution; clinical trial; group-sequential testing; targeted maximum likelihood methodology (1)
- Adaptive designs; Average treatment effect; Cluster randomized trials; Pair-matching; Randomized trials; Targeted minimum loss-based estimation (TMLE) (1)
- Area under the curve (1)
- Asymptotic linearity (1)
- Auxiliary variables (1)
- Balanced repeated replication (1)
- Bayesian inference (1)
- Bayesian methods; design-based inference; sampling weights (1)
- Bernstein's inequality; central limit theorem; confidence interval; influence curve; normal distribution; survey sampling (1)
- Between cluster variation (1)
- Blocked factorial (1)
- Case-cohort design (1)
- Causal inference (1)
- Causal inference; Censoring by death; Missing data; Potential outcomes; Principal stratification; Quantum mechanics (1)
- Censored linear regression (1)
- Cluster randomized trial (1)
- Clustering (1)
- Composite sampling (1)
- Publication Year
Articles 1 - 30 of 39
Full-Text Articles in Design of Experiments and Sample Surveys
A Simulation Study Of Diagnostics For Bias In Non-Probability Samples, Philip S. Boonstra, Roderick Ja Little, Brady T. West, Rebecca R. Andridge, Fernanda Alvarado-Leiton
A Simulation Study Of Diagnostics For Bias In Non-Probability Samples, Philip S. Boonstra, Roderick Ja Little, Brady T. West, Rebecca R. Andridge, Fernanda Alvarado-Leiton
The University of Michigan Department of Biostatistics Working Paper Series
A non-probability sampling mechanism is likely to bias estimates of parameters with respect to a target population of interest. This bias poses a unique challenge when selection is 'non-ignorable', i.e. dependent upon the unobserved outcome of interest, since it is then undetectable and thus cannot be ameliorated. We extend a simulation study by Nishimura et al. [International Statistical Review, 84, 43--62 (2016)], adding a recently published statistic, the so-called 'standardized measure of unadjusted bias', which explicitly quantifies the extent of bias under the assumption that a specified amount of non-ignorable selection exists. Our findings suggest that this new …
Models For Hsv Shedding Must Account For Two Levels Of Overdispersion, Amalia Magaret
Models For Hsv Shedding Must Account For Two Levels Of Overdispersion, Amalia Magaret
UW Biostatistics Working Paper Series
We have frequently implemented crossover studies to evaluate new therapeutic interventions for genital herpes simplex virus infection. The outcome measured to assess the efficacy of interventions on herpes disease severity is the viral shedding rate, defined as the frequency of detection of HSV on the genital skin and mucosa. We performed a simulation study to ascertain whether our standard model, which we have used previously, was appropriately considering all the necessary features of the shedding data to provide correct inference. We simulated shedding data under our standard, validated assumptions and assessed the ability of 5 different models to reproduce the …
Adaptive Pair-Matching In The Search Trial And Estimation Of The Intervention Effect, Laura Balzer, Maya L. Petersen, Mark J. Van Der Laan
Adaptive Pair-Matching In The Search Trial And Estimation Of The Intervention Effect, Laura Balzer, Maya L. Petersen, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
In randomized trials, pair-matching is an intuitive design strategy to protect study validity and to potentially increase study power. In a common design, candidate units are identified, and their baseline characteristics used to create the best n/2 matched pairs. Within the resulting pairs, the intervention is randomized, and the outcomes measured at the end of follow-up. We consider this design to be adaptive, because the construction of the matched pairs depends on the baseline covariates of all candidate units. As consequence, the observed data cannot be considered as n/2 independent, identically distributed (i.i.d.) pairs of units, as current practice assumes. …
Statistical Inference For Data Adaptive Target Parameters, Mark J. Van Der Laan, Alan E. Hubbard, Sara Kherad Pajouh
Statistical Inference For Data Adaptive Target Parameters, Mark J. Van Der Laan, Alan E. Hubbard, Sara Kherad Pajouh
U.C. Berkeley Division of Biostatistics Working Paper Series
Consider one observes n i.i.d. copies of a random variable with a probability distribution that is known to be an element of a particular statistical model. In order to define our statistical target we partition the sample in V equal size sub-samples, and use this partitioning to define V splits in estimation-sample (one of the V subsamples) and corresponding complementary parameter-generating sample that is used to generate a target parameter. For each of the V parameter-generating samples, we apply an algorithm that maps the sample in a target parameter mapping which represent the statistical target parameter generated by that parameter-generating …
Estimation And Testing In Targeted Group Sequential Covariate-Adjusted Randomized Clinical Trials, Antoine Chambaz, Mark J. Van Der Laan
Estimation And Testing In Targeted Group Sequential Covariate-Adjusted Randomized Clinical Trials, Antoine Chambaz, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
This article is devoted to the construction and asymptotic study of adaptive group sequential covariate-adjusted randomized clinical trials analyzed through the prism of the semiparametric methodology of targeted maximum likelihood estimation (TMLE). We show how to build, as the data accrue group-sequentially, a sampling design which targets a user-supplied optimal design. We also show how to carry out a sound TMLE statistical inference based on such an adaptive sampling scheme (therefore extending some results known in the i.i.d setting only so far), and how group-sequential testing applies on top of it. The procedure is robust (i.e., consistent even if the …
Lot Quality Assurance Sampling (Lqas) And The Mozambique Malaria Indicator Surveys, Caitlin Biedron, Marcello Pagano, Bethany L. Hedt, Albert Kilian, Amy Ratcliffe, Samuel Mabunda, Joseph J. Valadez
Lot Quality Assurance Sampling (Lqas) And The Mozambique Malaria Indicator Surveys, Caitlin Biedron, Marcello Pagano, Bethany L. Hedt, Albert Kilian, Amy Ratcliffe, Samuel Mabunda, Joseph J. Valadez
Harvard University Biostatistics Working Paper Series
No abstract provided.
Confidence Intervals For The Population Mean Tailored To Small Sample Sizes, With Applications To Survey Sampling, Michael Rosenblum, Mark J. Van Der Laan
Confidence Intervals For The Population Mean Tailored To Small Sample Sizes, With Applications To Survey Sampling, Michael Rosenblum, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
The validity of standard confidence intervals constructed in survey sampling is based on the central limit theorem. For small sample sizes, the central limit theorem may give a poor approximation, resulting in confidence intervals that are misleading. We discuss this issue and propose methods for constructing confidence intervals for the population mean tailored to small sample sizes.
We present a simple approach for constructing confidence intervals for the population mean based on tail bounds for the sample mean that are correct for all sample sizes. Bernstein's inequality provides one such tail bound. The resulting confidence intervals have guaranteed coverage probability …
Detailed Version: Analyzing Direct Effects In Randomized Trials With Secondary Interventions: An Application To Hiv Prevention Trials, Michael A. Rosenblum, Nicholas P. Jewell, Mark J. Van Der Laan, Stephen Shiboski, Ariane Van Der Straten, Nancy Padian
Detailed Version: Analyzing Direct Effects In Randomized Trials With Secondary Interventions: An Application To Hiv Prevention Trials, Michael A. Rosenblum, Nicholas P. Jewell, Mark J. Van Der Laan, Stephen Shiboski, Ariane Van Der Straten, Nancy Padian
U.C. Berkeley Division of Biostatistics Working Paper Series
This is the detailed technical report that accompanies the paper “Analyzing Direct Effects in Randomized Trials with Secondary Interventions: An Application to HIV Prevention Trials” (an unpublished, technical report version of which is available online at http://www.bepress.com/ucbbiostat/paper223).
The version here gives full details of the models for the time-dependent analysis, and presents further results in the data analysis section. The Methods for Improving Reproductive Health in Africa (MIRA) trial is a recently completed randomized trial that investigated the effect of diaphragm and lubricant gel use in reducing HIV infection among susceptible women. 5,045 women were randomly assigned to either the …
Analyzing Direct Effects In Randomized Trials With Secondary Interventions , Michael Rosenblum, Nicholas P. Jewell, Mark J. Van Der Laan, Stephen Shiboski, Ariane Van Der Straten, Nancy Padian
Analyzing Direct Effects In Randomized Trials With Secondary Interventions , Michael Rosenblum, Nicholas P. Jewell, Mark J. Van Der Laan, Stephen Shiboski, Ariane Van Der Straten, Nancy Padian
U.C. Berkeley Division of Biostatistics Working Paper Series
The Methods for Improving Reproductive Health in Africa (MIRA) trial is a recently completed randomized trial that investigated the effect of diaphragm and lubricant gel use in reducing HIV infection among susceptible women. 5,045 women were randomly assigned to either the active treatment arm or not. Additionally, all subjects in both arms received intensive condom counselling and provision, the "gold standard" HIV prevention barrier method. There was much lower reported condom use in the intervention arm than in the control arm, making it difficult to answer important public health questions based solely on the intention-to-treat analysis. We adapt an analysis …
Principal Stratification Designs To Estimate Input Data Missing Due To Death, Constantine E. Frangakis, Donald B. Rubin, Ming-Wen An, Ellen Mackenzie
Principal Stratification Designs To Estimate Input Data Missing Due To Death, Constantine E. Frangakis, Donald B. Rubin, Ming-Wen An, Ellen Mackenzie
Johns Hopkins University, Dept. of Biostatistics Working Papers
We consider studies of cohorts of individuals after a critical event, such as an injury, with the following characteristics. First, the studies are designed to measure “input” variables, which describe the period before the critical event, and to characterize the distribution of the input variables in the cohort. Second, the studies are designed to measure “output” variables, primarily mortality after the critical event, and to characterize the predictive (conditional) distribution of mortality given the input variables in the cohort. Such studies often possess the complication that the input data are missing for those who die shortly after the critical event …
Multiple Imputation In The Presence Of Outliers, Michael Elliott
Multiple Imputation In The Presence Of Outliers, Michael Elliott
The University of Michigan Department of Biostatistics Working Paper Series
We consider the problem of obtaining population-based inference in the presence of missing data and outliers in the context of estimating obesity prevalence and body-mass index (BMI) measures from the Healthy For Life Study. Identifying multiple outliers in a multivariate setting is problematic because of problems such as masking, in which groups of outliers inflate the covariance matrix in a fashion that prevents their identification when included, and swamping, in which outliers skew covariances in a fashion that make non-outling observations appear to be outliers. We develop a latent class model that assumes each observation belongs to one of $K$ …
A General Framework For Statistical Performance Comparison Of Evolutionary Computation Algorithms, David Shilane, Jarno Martikainen, Sandrine Dudoit, Seppo Ovaska
A General Framework For Statistical Performance Comparison Of Evolutionary Computation Algorithms, David Shilane, Jarno Martikainen, Sandrine Dudoit, Seppo Ovaska
U.C. Berkeley Division of Biostatistics Working Paper Series
This paper proposes a statistical methodology for comparing the performance of evolutionary computation algorithms. A two-fold sampling scheme for collecting performance data is introduced, and these data are analyzed using bootstrap-based multiple hypothesis testing procedures. The proposed method is sufficiently flexible to allow the researcher to choose how performance is measured, does not rely upon distributional assumptions, and can be extended to analyze many other randomized numeric optimization routines. As a result, this approach offers a convenient, flexible, and reliable technique for comparing algorithms in a wide variety of applications.
2^K Factorials In Blocks Of Size 2, With Application To Two-Color Microarray Experiments, Kathleen F. Kerr
2^K Factorials In Blocks Of Size 2, With Application To Two-Color Microarray Experiments, Kathleen F. Kerr
UW Biostatistics Working Paper Series
When a two-level design must be run in blocks of size two, there is a unique blocking scheme that enables estimation of all the main effects. Unfortunately this design does not enable estimation of any two-factor interactions. When the experimental goal is to estimate all main effects and two-factor interactions, it is necessary to combine replicates of the experiment that use different blocking schemes. In this paper we identify such designs for up to eight factors that enable estimation of all main effects and two-factor interactions with the fewest number of replications. In addition, we give a construction for general …
New Statistical Paradigms Leading To Web-Based Tools For Clinical/Translational Science, Knut M. Wittkowski
New Statistical Paradigms Leading To Web-Based Tools For Clinical/Translational Science, Knut M. Wittkowski
COBRA Preprint Series
As the field of functional genetics and genomics is beginning to mature, we become confronted with new challenges. The constant drop in price for sequencing and gene expression profiling as well as the increasing number of genetic and genomic variables that can be measured makes it feasible to address more complex questions. The success with rare diseases caused by single loci or genes has provided us with a proof-of-concept that new therapies can be developed based on functional genomics and genetics.
Common diseases, however, typically involve genetic epistasis, genomic pathways, and proteomic pattern. Moreover, to better understand the underlying biologi-cal …
Designs In Partially Controlled Studies: Messages From A Review, Fan Li, Constantine E. Frangakis
Designs In Partially Controlled Studies: Messages From A Review, Fan Li, Constantine E. Frangakis
Johns Hopkins University, Dept. of Biostatistics Working Papers
The ability to evaluate effects of factors on outcomes is increasingly important for a class of studies that control some but not all of the factors. Although important advances have been made in methods of analysis for such partially controlled studies,work on designs for such studies has been relatively limited. To help understand why, we review main designs that have been used for such partially controlled studies. Based on the review, we give two complementary reasons that explain the limited work on such designs, and suggest a new direction in this area.
Referent Selection Strategies In Case-Crossover Analyses Of Air Pollution Exposure Data: Implications For Bias, Holly Janes, Lianne Sheppard, Thomas Lumley
Referent Selection Strategies In Case-Crossover Analyses Of Air Pollution Exposure Data: Implications For Bias, Holly Janes, Lianne Sheppard, Thomas Lumley
UW Biostatistics Working Paper Series
The case-crossover design has been widely used to study the association between short term air pollution exposure and the risk of an acute adverse health event. The design uses cases only, and, for each individual, compares exposure just prior to the event with exposure at other control, or “referent” times. By making within-subject comparisons, time invariant confounders are controlled by design. Even more important in the air pollution setting is that, by matching referents to the index time, time varying confounders can also be controlled by design. Yet, the referent selection strategy is important for reasons other than control of …
Censored Linear Regression For Case-Cohort Studies, Bin Nan, Menggang Yu, Jack Kalbfleisch
Censored Linear Regression For Case-Cohort Studies, Bin Nan, Menggang Yu, Jack Kalbfleisch
The University of Michigan Department of Biostatistics Working Paper Series
Right censored data from a classical case-cohort design and a stratified case-cohort design are considered. In the classical case-cohort design, the subcohort is obtained as a simple random sample of the entire cohort, whereas in the stratified design, the subcohort is selected by independent Bernoulli sampling with arbitrary selection probabilities. For each design and under a linear regression model, methods for estimating the regression parameters are proposed and analyzed. These methods are derived by modifying the linear ranks tests and estimating equations that arise from full-cohort data using methods that are similar to the "pseudo-likelihood" estimating equation that has been …
Non-Parametric Estimation Of Roc Curves In The Absence Of A Gold Standard, Xiao-Hua Zhou, Pete Castelluccio, Chuan Zhou
Non-Parametric Estimation Of Roc Curves In The Absence Of A Gold Standard, Xiao-Hua Zhou, Pete Castelluccio, Chuan Zhou
UW Biostatistics Working Paper Series
In evaluation of diagnostic accuracy of tests, a gold standard on the disease status is required. However, in many complex diseases, it is impossible or unethical to obtain such the gold standard. If an imperfect standard is used as if it were a gold standard, the estimated accuracy of the tests would be biased. This type of bias is called imperfect gold standard bias. In this paper we develop a maximum likelihood (ML) method for estimating ROC curves and their areas of ordinal-scale tests in the absence of a gold standard. Our simulation study shows the proposed estimates for the …
New Estimating Methods For Surrogate Outcome Data, Bin Nan
New Estimating Methods For Surrogate Outcome Data, Bin Nan
The University of Michigan Department of Biostatistics Working Paper Series
Surrogate outcome data arise frequently in medical research. The true outcomes of interest are expensive or hard to ascertain, but measurements of surrogate outcomes (or more generally speaking, the correlates of the true outcomes) are usually available. In this paper we assume that the conditional expectation of the true outcome given covariates is known up to a finite dimensional parameter. When the true outcome is missing at random, the e±cient score function for the parameter in the conditional mean model has a simple form, which is similar to the generalized estimating functions. There is no integral equation involved as in …
Bayesian Geostatistical Design, Peter J. Diggle, Soren Lophaven
Bayesian Geostatistical Design, Peter J. Diggle, Soren Lophaven
Johns Hopkins University, Dept. of Biostatistics Working Papers
This paper describes the use of model-based geostatistics for choosing the optimal set of sampling locations, collectively called the design, for a geostatistical analysis. Two types of design situations are considered. These are retrospective design, which concerns the addition of sampling locations to, or deletion of locations from, an existing design, and prospective design, which consists of choosing optimal positions for a new set of sampling locations. We propose a Bayesian design criterion which focuses on the goal of efficient spatial prediction whilst allowing for the fact that model parameter values are unknown. The results show that in this situation …
Does Weighting For Nonresponse Increase The Variance Of Survey Means?, Rod Little, Sonya L. Vartivarian
Does Weighting For Nonresponse Increase The Variance Of Survey Means?, Rod Little, Sonya L. Vartivarian
The University of Michigan Department of Biostatistics Working Paper Series
Nonresponse weighting is a common method for handling unit nonresponse in surveys. A widespread view is that the weighting method is aimed at reducing nonresponse bias, at the expense of an increase in variance. Hence, the efficacy of weighting adjustments becomes a bias-variance trade-off. This note suggests that this view is an oversimplification -- nonresponse weighting can in fact lead to a reduction in variance as well as bias. A covariate for a weighting adjustment must have two characteristics to reduce nonresponse bias - it needs to be related to the probability of response, and it needs to be related …
Causal Inference In Hybrid Intervention Trials Involving Treatment Choice, Qi Long, Rod Little, Xihong Lin
Causal Inference In Hybrid Intervention Trials Involving Treatment Choice, Qi Long, Rod Little, Xihong Lin
The University of Michigan Department of Biostatistics Working Paper Series
Randomized allocation of treatments is a cornerstone of experimental design, but has drawbacks when a limited set of individuals are willing to be randomized, or the act of randomization undermines the success of the treatment. Choice-based experimental designs allow a subset of the participants to choose their treatments. We discuss here causal inferences for experimental designs where some participants are randomly allocated to treatments and others receive their treatment preference. This paper was motivated by the “Women Take Pride” (WTP) study (Janevic et al., 2001), a doubly randomized preference trail (DRPT) to assess behavioral interventions for women with heart disease. …
Multiple Imputation For Interval Censored Data With Auxiliary Variables, Chiu-Hsieh Hsu, Jeremy Taylor, Susan Murray
Multiple Imputation For Interval Censored Data With Auxiliary Variables, Chiu-Hsieh Hsu, Jeremy Taylor, Susan Murray
The University of Michigan Department of Biostatistics Working Paper Series
We propose a nonparametric multiple imputation scheme, NPMLE imputation, for the analysis of interval censored survival data. Features of the method are that it converts interval-censored data problems to complete data or right censored data problems to which many standard approaches can be used, and the measures of uncertainty are easily obtained. In addition to the event time of primary interest, there are frequently other auxiliary variables that are associated with the event time. For the goal of estimating the marginal survival distribution, these auxiliary variables may provide some additional information about the event time for the interval censored observations. …
Optimal Sample Size For Multiple Testing: The Case Of Gene Expression Microarrays, Peter Muller, Giovanni Parmigiani, Christian Robert, Judith Rousseau
Optimal Sample Size For Multiple Testing: The Case Of Gene Expression Microarrays, Peter Muller, Giovanni Parmigiani, Christian Robert, Judith Rousseau
Johns Hopkins University, Dept. of Biostatistics Working Papers
We consider the choice of an optimal sample size for multiple comparison problems. The motivating application is the choice of the number of microarray experiments to be carried out when learning about differential gene expression. However, the approach is valid in any application that involves multiple comparisons in a large number of hypothesis tests. We discuss two decision problems in the context of this setup: the sample size selection and the decision about the multiple comparisons. We adopt a decision theoretic approach,using loss functions that combine the competing goals of discovering as many ifferentially expressed genes as possible, while keeping …
Overlap Bias In The Case-Crossover Design, With Application To Air Pollution Exposures, Holly Janes, Lianne Sheppard, Thomas Lumley
Overlap Bias In The Case-Crossover Design, With Application To Air Pollution Exposures, Holly Janes, Lianne Sheppard, Thomas Lumley
UW Biostatistics Working Paper Series
The case-crossover design uses cases only, and compares exposures just prior to the event times to exposures at comparable control, or “referent” times, in order to assess the effect of short-term exposure on the risk of a rare event. It has commonly been used to study the effect of air pollution on the risk of various adverse health events. Proper selection of referents is crucial, especially with air pollution exposures, which are shared, highly seasonal, and often have a long term time trend. Hence, careful referent selection is important to control for time-varying confounders, and in order to ensure that …
Robust Likelihood-Based Analysis Of Multivariate Data With Missing Values, Rod Little, An Hyonggin
Robust Likelihood-Based Analysis Of Multivariate Data With Missing Values, Rod Little, An Hyonggin
The University of Michigan Department of Biostatistics Working Paper Series
The model-based approach to inference from multivariate data with missing values is reviewed. Regression prediction is most useful when the covariates are predictive of the missing values and the probability of being missing, and in these circumstances predictions are particularly sensitive to model misspecification. The use of penalized splines of the propensity score is proposed to yield robust model-based inference under the missing at random (MAR) assumption, assuming monotone missing data. Simulation comparisons with other methods suggest that the method works well in a wide range of populations, with little loss of efficiency relative to parametric models when the latter …
Marginalized Transition Models For Longitudinal Binary Data With Ignorable And Nonignorable Dropout, Brenda F. Kurland, Patrick J. Heagerty
Marginalized Transition Models For Longitudinal Binary Data With Ignorable And Nonignorable Dropout, Brenda F. Kurland, Patrick J. Heagerty
UW Biostatistics Working Paper Series
We extend the marginalized transition model of Heagerty (2002) to accommodate nonignorable monotone dropout. Using a selection model, weakly identified dropout parameters are held constant and their effects evaluated through sensitivity analysis. For data missing at random (MAR), efficiency of inverse probability of censoring weighted generalized estimating equations (IPCW-GEE) is as low as 40% compared to a likelihood-based marginalized transition model (MTM) with comparable modeling burden. MTM and IPCW-GEE regression parameters both display misspecification bias for MAR and nonignorable missing data, and both reduce bias noticeably by improving model fit
Weighting Adjustments For Unit Nonresponse With Multiple Outcome Variables, Sonya L. Vartivarian, Rod Little
Weighting Adjustments For Unit Nonresponse With Multiple Outcome Variables, Sonya L. Vartivarian, Rod Little
The University of Michigan Department of Biostatistics Working Paper Series
Weighting is a common form of unit nonresponse adjustment in sample surveys where entire questionnaires are missing due to noncontact or refusal to participate. Weights are inversely proportional to the probability of selection and response. A common approach computes the response weight adjustment cells based on covariate information. When the number of cells thus created is too large, a coarsening method such as response propensity stratification can be applied to reduce the number of adjustment cells. Simulations in Vartivarian and Little (2002) indicate improved efficiency and robustness of weighting adjustments based on the joint classification of the sample by two …
To Model Or Not To Model? Competing Modes Of Inference For Finite Population Sampling, Rod Little
To Model Or Not To Model? Competing Modes Of Inference For Finite Population Sampling, Rod Little
The University of Michigan Department of Biostatistics Working Paper Series
Finite population sampling is perhaps the only area of statistics where the primary mode of analysis is based on the randomization distribution, rather than on statistical models for the measured variables. This article reviews the debate between design and model-based inference. The basic features of the two approaches are illustrated using the case of inference about the mean from stratified random samples. Strengths and weakness of design-based and model-based inference for surveys are discussed. It is suggested that models that take into account the sample design and make weak parametric assumptions can produce reliable and efficient inferences in surveys settings. …
Inference For The Population Total From Probability-Proportional-To-Size Samples Based On Predictions From A Penalized Spline Nonparametric Model, Hui Zheng, Rod Little
Inference For The Population Total From Probability-Proportional-To-Size Samples Based On Predictions From A Penalized Spline Nonparametric Model, Hui Zheng, Rod Little
The University of Michigan Department of Biostatistics Working Paper Series
Inference about the finite population total from probability-proportional-to-size (PPS) samples is considered. In previous work (Zheng and Little, 2003), penalized spline (p-spline) nonparametric model-based estimators were shown to generally outperform the Horvitz-Thompson (HT) and generalized regression (GR) estimators in terms of the root mean squared error. In this article we develop model-based, jackknife and balanced repeated replicate variance estimation methods for the p-spline based estimators. Asymptotic properties of the jackknife method are discussed. Simulations show that p-spline point estimators and their jackknife standard errors lead to inferences that are superior to HT or GR based inferences. This suggests that nonparametric …