Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

2005

Discipline
Institution
Keyword
Publication
Publication Type

Articles 151 - 180 of 279

Full-Text Articles in Statistics and Probability

G-Computation Estimation Of Nonparametric Causal Effects On Time-Dependent Mean Outcomes In Longitudinal Studies, Romain Neugebauer, Mark J. Van Der Laan Jul 2005

G-Computation Estimation Of Nonparametric Causal Effects On Time-Dependent Mean Outcomes In Longitudinal Studies, Romain Neugebauer, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Two approaches to Causal Inference based on Marginal Structural Models (MSM) have been proposed. They provide different representations of causal effects with distinct causal parameters. Initially, a parametric MSM approach to Causal Inference was developed: it relies on correct specification of a parametric MSM. Recently, a new approach based on nonparametric MSM was introduced. This later approach does not require the assumption of a correctly specified MSM and thus is more realistic if one believes that correct specification of a parametric MSM is unlikely in practice. However, this approach was described only for investigating causal effects on mean outcomes collected …


Good Measures On Cantor Space, Ethan Akin Jul 2005

Good Measures On Cantor Space, Ethan Akin

Mathematics and Statistics Faculty Research & Creative Works

While there is, up to homeomorphism, only one Cantor space, i.e. one zero-dimensional, perfect, compact, nonempty metric space, there are many measures on Cantor space which are not topologically equivalent. The clopen values set for a full, nonatomic measure μ is the countable dense subset {μ(U): U is clopen} of the unit interval. It is a topological invariant for the measure. For the class of good measures, it is a complete invariant. A full, nonatomic measure μ is good if whenever U, V are clopen sets with μ(U) < μ(V), there exists W a clopen subset of V such that μ(W) = μ(U). These measures have interesting dynamical properties. They are exactly the measures which arise from uniquely ergodic minimal systems on Cantor space. For some of them there is a unique generic measure-preserving homeomorphism. That is, within the Polish group of such homeomorphisms there is a dense, G δ conjugacy class. ©2004 American Mathematical Society.


A Note On The Construction Of Counterfactuals And The G-Computation Formula, Zhuo Yu, Mark J. Van Der Laan Jun 2005

A Note On The Construction Of Counterfactuals And The G-Computation Formula, Zhuo Yu, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Robins' causal inference theory assumes existence of treatment specific counterfactual variables so that the observed data augmented by the counterfactual data will satisfy a consistency and a randomization assumption. In this paper we provide an explicit function that maps the observed data into a counterfactual variable which satisfies the consistency and randomization assumptions. This offers a practically useful imputation method for counterfactuals. Gill & Robins [2001]'s construction of counterfactuals can be used as an imputation method in principle, but it is very hard to implement in practice. Robins [1987] shows that the counterfactual distribution can be identified from the observed …


Cross-Validated Bagged Learning, Mark J. Van Der Laan, Sandra E. Sinisi, Maya L. Petersen Jun 2005

Cross-Validated Bagged Learning, Mark J. Van Der Laan, Sandra E. Sinisi, Maya L. Petersen

U.C. Berkeley Division of Biostatistics Working Paper Series

Many applications aim to learn a high dimensional parameter of a data generating distribution based on a sample of independent and identically distributed observations. For example, the goal might be to estimate the conditional mean of an outcome given a list of input variables. In this prediction context, Breiman (1996a) introduced bootstrap aggregating (bagging) as a method to reduce the variance of a given estimator at little cost to bias. Bagging involves applying the estimator to multiple bootstrap samples, and averaging the result across bootstrap samples. In order to deal with the curse of dimensionality, typical practice has been to …


On Additive Regression Of Expectancy, Ying Qing Chen Jun 2005

On Additive Regression Of Expectancy, Ying Qing Chen

UW Biostatistics Working Paper Series

Regression models have been important tools to study the association between outcome variables and their covariates. The traditional linear regression models usually specify such an association by the expectations of the outcome variables as function of the covariates and some parameters. In reality, however, interests often focus on their expectancies characterized by the conditional means. In this article, a new class of additive regression models is proposed to model the expectancies. The model parameters carry practical implication, which may allow the models to be useful in applications such as treatment assessment, resource planning or short-term forecasting. Moreover, the new model …


Spatio-Temporal Point Processes: Methods And Applications, Peter J. Diggle Jun 2005

Spatio-Temporal Point Processes: Methods And Applications, Peter J. Diggle

Johns Hopkins University, Dept. of Biostatistics Working Papers

No abstract provided.


A Partial Likelihood For Spatio-Temporal Point Processes, Peter J. Diggle Jun 2005

A Partial Likelihood For Spatio-Temporal Point Processes, Peter J. Diggle

Johns Hopkins University, Dept. of Biostatistics Working Papers

Spatio-temporal point process data arise in many fields of application. An intuitively natural way to specify a model for a spatio-temporal point process is through its conditional intensity at location x and time t, given the history of the process up to time t. Typically, this results in an analytically intractable likelihood. Likelihood-based inference therefore relies on Monte Carlo methods which are computationally intensive and require careful tuning to each application. We propose a partial likelihood alternative which is computationally straightforward and can be applied routinely. We apply the method to data from the 2001 foot-and-mouth epidemic in the UK, …


Polydesigns And Causal Inference, Fan Li, Constantine E. Frangakis Jun 2005

Polydesigns And Causal Inference, Fan Li, Constantine E. Frangakis

Johns Hopkins University, Dept. of Biostatistics Working Papers

In an increasingly common class of studies, the goal is to evaluate causal effects of treatments that are only partially controlled by the investigator. In such studies there are two conflicting features: (1) a model on the full cohort design and data can identify the causal effects of interest, but can be sensitive to extreme regions of that design's data, where model specification can have more impact; and (2) models on a reduced design (i.e., a subset of the full data), e.g., conditional likelihood on matched subsets of data, can avoid such sensitivity, but do not generally identify the causal …


Computational Optical Biopsy, Yi Li, Ming Jiang, Ge Wang Jun 2005

Computational Optical Biopsy, Yi Li, Ming Jiang, Ge Wang

Mathematics and Statistics Faculty Publications

Optical molecular imaging is based on fluorescence or bioluminescence, and hindered by photon scattering in the tissue, especially in patient studies. Here we propose a computational optical biopsy (COB) approach to localize and quantify a light source deep inside a subject. In contrast to existing optical biopsy techniques, our scheme is to collect optical signals directly from a region of interest along one or multiple biopsy paths in a subject, and then compute features of an underlying light source distribution. In this paper, we formulate this inverse problem in the framework of diffusion approximation, demonstrate the solution uniqueness properties in …


An Empirical Process Limit Theorem For Sparsely Correlated Data, Thomas Lumley Jun 2005

An Empirical Process Limit Theorem For Sparsely Correlated Data, Thomas Lumley

UW Biostatistics Working Paper Series

We consider data that are dependent, but where most small sets of observations are independent. By extending Bernstein's inequality we prove a strong law of law numbers and an empirical process central limit theorem under bracketing entropy conditions.


Model Choice In Time Series Studies Of Air Pollution And Mortality, Roger D. Peng, Francesca Dominici, Thomas A. Louis Jun 2005

Model Choice In Time Series Studies Of Air Pollution And Mortality, Roger D. Peng, Francesca Dominici, Thomas A. Louis

Johns Hopkins University, Dept. of Biostatistics Working Papers

Multi-city time series studies of particulate matter (PM) and mortality and morbidity have provided evidence that daily variation in air pollution levels is associated with daily variation in mortality counts. These findings served as key epidemiological evidence for the recent review of the United States National Ambient Air Quality Standards (NAAQS) for PM. As a result, methodological issues concerning time series analysis of the relation between air pollution and health have attracted the attention of the scientific community and critics have raised concerns about the adequacy of current model formulations. Time series data on pollution and mortality are generally analyzed …


Multivariate Analysis Of Ecologicai Data Using Canco, Anne M. Parkhurst Jun 2005

Multivariate Analysis Of Ecologicai Data Using Canco, Anne M. Parkhurst

Department of Statistics: Faculty Publications

This book is about understanding and applying multivariate statistical methods useful for analyzing complex ecological problems. Researchers and students seeking to improve their ability to collect and analyze data from field observations and experiments, especially those interested in the response or variation of biotic communities to environmental conditions or experimental manipulation, will find this handbook helpful. The methods discussed are widely used in plant community ecology, as well as other areas in biology.

The book is tutorial in nature. The authors provide advice on how to best apply the multivariate statistical methods using the CANOCO for Windows, a licensed computer …


Fit-To-Fight: Waist Vs. Waist/Height Measurements To Determine An Individual's Fitness Level A Study In Statistical Regression And Analysis, Steven J. Swiderski Jun 2005

Fit-To-Fight: Waist Vs. Waist/Height Measurements To Determine An Individual's Fitness Level A Study In Statistical Regression And Analysis, Steven J. Swiderski

Theses and Dissertations

Air Force members are to be tested for fitness by measuring their abdominal circumference, counting the number of sit-ups and push-ups they can accomplish, and the time it takes them to run 1 and miles. The abdominal measurement is a "one-size-fits-all" fitness standard. This research determines that a person's waist-to-height ratio is a better measurement than the waist measurement to estimate an individual's fitness level. This research estimates that all of the variables used to proxy fitness (Gender, Age, Height, Waist Circumference, Waist-to-Height Ratio, Push-Ups, and Sit-Ups) are statistically significant and do represent good estimators of physical fitness. This research …


The Navigation Potential Of Signals Of Opportunity-Based Time Difference Of Arrival Measurements, Kenneth A. Fisher Jun 2005

The Navigation Potential Of Signals Of Opportunity-Based Time Difference Of Arrival Measurements, Kenneth A. Fisher

Theses and Dissertations

This research introduces the concept of navigation potential, NP, to quantify the intrinsic ability to navigate using a given signal. NP theory is a new, information theory-like concept that provides a theoretical performance limit on estimating navigation parameters from a received signal that is modeled through a stochastic mapping of the transmitted signal and measurement noise. NP theory is applied to SOP-based TDOA systems in general as well as for the Gaussian case. Furthermore, the NP is found for a received signal consisting of the transmitted signal, multiple delayed and attenuated replicas of the transmitted signal, and measurement noise. Multipath-based …


Rank-Based Methods For Repeated Measures Data Under Exchangeable Errors, John Kloke Jun 2005

Rank-Based Methods For Repeated Measures Data Under Exchangeable Errors, John Kloke

Dissertations

Rank-based estimation methods provide alternatives to least squares. Estimators derived via least squares are generally not robust to aberrant observations.Rank-based methods for linear models generalize traditional Wilcoxon procedures in the simple location models and are robust.

In the usual linear model it is assumed that the errors are independent. In the case of repeated measures data several observations are taken on each experimental unit. In the case of longitudinal data the measures are taken on the same subject over time. As such an independence assumption does not seem valid. A common solution to this is to make an assumption on …


A Linear Regression Framework For Receiver Operating Characteristic(Roc) Curve Analysis, Zheng Zhang, Margaret S. Pepe May 2005

A Linear Regression Framework For Receiver Operating Characteristic(Roc) Curve Analysis, Zheng Zhang, Margaret S. Pepe

UW Biostatistics Working Paper Series

In the field of medical diagnostic testing, the receiver operating characteristics(ROC) curve has long been used as a standard statistical tool to assess the accuracy of tests that yield continuous results. Although previous research in this area focused mostly on estimating the ROC curve, recently it has been recognized that the accuracy of a given test may fluctuate depending on certain factors, which motivates modelling covariate effects on the ROC curve. Comparing the corresponding ROC curves between two or more tests is a special case of covariate effect modelling. In this manuscript, we introduce a linear regression framework to model …


Attributable Risk Function In The Proportional Hazards Model, Ying Qing Chen, Chengcheng Hu, Yan Wang May 2005

Attributable Risk Function In The Proportional Hazards Model, Ying Qing Chen, Chengcheng Hu, Yan Wang

UW Biostatistics Working Paper Series

As an epidemiological parameter, the population attributable fraction is an important measure to quantify the public health attributable risk of an exposure to morbidity and mortality. In this article, we extend this parameter to the attributable fraction function in survival analysis of time-to-event outcomes, and further establish its estimation and inference procedures based on the widely used proportional hazards models. Numerical examples and simulations studies are presented to validate and demonstrate the proposed methods.


Structure And Dynamics Of Soluble Guanylyl Cyclase, Kentaro Sugino May 2005

Structure And Dynamics Of Soluble Guanylyl Cyclase, Kentaro Sugino

Theses

Soluble guanylyl cyclase (sGC) is one of the key enzymes involved in many fundamental biological processes including vasodilatation. It can be allosterically activated by synthetic compound such as YC-l. Recently, the 3D structure of adenylyl cyclase (AC), which is a homologue of sGC, was determined. Using AC as template and homology modeling, the 3D structure of sGC is predicted. Prior experimental work has suggested two binding modes of YC- 1. In the current investigation, molecular dynamics simulations (MD) were conducted to seek more detail of molecular mechanism of sGC activation.

From these MD simulations, a tentative mechanism of sGC activation …


New Statistical Paradigms Leading To Web-Based Tools For Clinical/Translational Science, Knut M. Wittkowski May 2005

New Statistical Paradigms Leading To Web-Based Tools For Clinical/Translational Science, Knut M. Wittkowski

COBRA Preprint Series

As the field of functional genetics and genomics is beginning to mature, we become confronted with new challenges. The constant drop in price for sequencing and gene expression profiling as well as the increasing number of genetic and genomic variables that can be measured makes it feasible to address more complex questions. The success with rare diseases caused by single loci or genes has provided us with a proof-of-concept that new therapies can be developed based on functional genomics and genetics.

Common diseases, however, typically involve genetic epistasis, genomic pathways, and proteomic pattern. Moreover, to better understand the underlying biologi-cal …


Determining The Optimum Number Of Increments In Composite Sampling, John Ellis Hathaway May 2005

Determining The Optimum Number Of Increments In Composite Sampling, John Ellis Hathaway

Theses and Dissertations

Composite sampling can be more cost effective than simple random sampling. This paper considers how to determine the optimum number of increments to use in composite sampling. Composite sampling terminology and theory are outlined and a model is developed which accounts for different sources of variation in compositing and data analysis. This model is used to define and understand the process of determining the optimum number of increments that should be used in forming a composite. The blending variance is shown to have a smaller range of possible values than previously reported when estimating the number of increments in a …


Estimating Function Based Cross-Validation And Learning, Mark J. Van Der Laan, Daniel Rubin May 2005

Estimating Function Based Cross-Validation And Learning, Mark J. Van Der Laan, Daniel Rubin

U.C. Berkeley Division of Biostatistics Working Paper Series

Suppose that we observe a sample of independent and identically distributed realizations of a random variable. Given a model for the data generating distribution, assume that the parameter of interest can be characterized as the parameter value which makes the population mean of a possibly infinite dimensional estimating function equal to zero. Given a collection of candidate estimators of this parameter, and specification of the vector estimating function, we propose cross-validation criteria for selecting among these estimators. This cross-validation criteria is defined as the Euclidean norm of the empirical mean over the validation sample of the estimating function at the …


Prognosis Of Stage Ii Colon Cancer By Non-Neoplastic Mucosa Gene Expresssion Profiling, Alain Barrier, Sandrine Dudoit, Et Al. May 2005

Prognosis Of Stage Ii Colon Cancer By Non-Neoplastic Mucosa Gene Expresssion Profiling, Alain Barrier, Sandrine Dudoit, Et Al.

U.C. Berkeley Division of Biostatistics Working Paper Series

Aims. This study assessed the possibility to build a prognosis predictor, based on non-neoplastic mucosa microarray gene expression measures, in stage II colon cancer patients. Materials and Methods. Non-neoplastic colonic mucosa mRNA samples from 24 patients (10 with a metachronous metastasis, 14 with no recurrence) were profiled using the Affymetrix HGU133A GeneChip. The k-nearest neighbor method was used for prognosis prediction using microarray gene expression measures. Leave-one-out cross-validation was used to select the number of neighbors and number of informative genes to include in the predictor. Based on this information, a prognosis predictor was proposed and its accuracy estimated by …


Colon Cancer Prognosis Prediction By Gene Expression Profiling, Alain Barrier, Sandrine Dudoit, Et Al. May 2005

Colon Cancer Prognosis Prediction By Gene Expression Profiling, Alain Barrier, Sandrine Dudoit, Et Al.

U.C. Berkeley Division of Biostatistics Working Paper Series

Aims. This study assessed the possibility to build a prognosis predictor, based on microarray gene expression measures, in stage II and III colon cancer patients. Materials and Methods. Tumour (T) and non-neoplastic mucosa (NM) mRNA samples from 18 patients (9 with a recurrence, 9 with no recurrence) were profiled using the Affymetrix HGU133A GeneChip. The k-nearest neighbour method was used for prognosis prediction using T and NM gene expression measures. Six-fold cross-validation was applied to select the number of neighbours and the number of informative genes to include in the predictors. Based on this information, one T-based and one NM-based …


Multiple Imputation For Correcting Verification Bias, Ofer Harel, Xiao-Hua Zhou May 2005

Multiple Imputation For Correcting Verification Bias, Ofer Harel, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

In the case in which all subjects are screened using a common test, and only a subset of these subjects are tested using a golden standard test, it is well documented that there is a risk for bias, called verification bias. When the test has only two levels (e.g. positive and negative) and we are trying to estimate the sensitivity and specificity of the test, one is actually constructing a confidence interval for a binomial proportion. Since it is well documented that this estimation is not trivial even with complete data, we adopt Multiple imputation (MI) framework for verification bias …


On The Power Function Of Bayesian Tests With Application To Design Of Clinical Trials: The Fixed-Sample Case, Lyle Broemeling, Dongfeng Wu May 2005

On The Power Function Of Bayesian Tests With Application To Design Of Clinical Trials: The Fixed-Sample Case, Lyle Broemeling, Dongfeng Wu

Journal of Modern Applied Statistical Methods

Using a Bayesian approach to clinical trial design is becoming more common. For example, at the MD Anderson Cancer Center, Bayesian techniques are routinely employed in the design and analysis of Phase I and II trials. It is important that the operating characteristics of these procedures be determined as part of the process when establishing a stopping rule for a clinical trial. This study determines the power function for some common fixed-sample procedures in hypothesis testing, namely the one and two-sample tests involving the binomial and normal distributions. Also considered is a Bayesian test for multi-response (response and toxicity) in …


Two Sides Of The Same Coin: Bootstrapping The Restricted Vs. Unrestricted Model, Panagiotis Mantalos May 2005

Two Sides Of The Same Coin: Bootstrapping The Restricted Vs. Unrestricted Model, Panagiotis Mantalos

Journal of Modern Applied Statistical Methods

The properties of the bootstrap test for restrictions are studied in two versions: 1) bootstrapping under the null hypothesis, restricted, and 2) bootstrapping under the alternative hypothesis, unrestricted. This article demonstrates the equivalence of these two methods, and illustrates the small sample properties of the Wald test for testing Granger-Causality in a stable stationary VAR system by Monte Carlo methods. The analysis regarding the size of the test reveals that, as expected, both bootstrap tests have actual sizes that lie close to the nominal size. Regarding the power of the test, the Wald and bootstrap tests share the same power …


Manifestation Of Differences In Item-Level Characteristics In Scale-Level Measurement Invariance Tests Of Multi-Group Confirmatory Factor Analyses, Bruno D. Zumbo, Kim H. Koh May 2005

Manifestation Of Differences In Item-Level Characteristics In Scale-Level Measurement Invariance Tests Of Multi-Group Confirmatory Factor Analyses, Bruno D. Zumbo, Kim H. Koh

Journal of Modern Applied Statistical Methods

If a researcher applies the conventional tests of scale-level measurement invariance through multi-group confirmatory factor analysis of a PC matrix and MLE to test hypotheses of strong and full measurement invariance when the researcher has a rating scale response format wherein the item characteristics are different for the two groups of respondents, do these scale-level analyses reflect (or ignore) differences in item threshold characteristics? Results of the current study demonstrate the inadequacy of judging the suitability of a measurement instrument across groups by only investigating the factor structure of the measure for the different groups with a PC matrix and …


Right-Tailed Testing Of Variance For Non-Normal Distributions, Michael C. Long, Ping Sa May 2005

Right-Tailed Testing Of Variance For Non-Normal Distributions, Michael C. Long, Ping Sa

Journal of Modern Applied Statistical Methods

A new test of variance for non-normal distribution with fewer restrictions than the current tests is proposed. Simulation study shows that the new test controls the Type I error rate well, and has power performance comparable to the competitors. In addition, it can be used without restrictions.


Multiple Imputation For Missing Ordinal Data, Ling Chen, Marian Toma-Drane, Robert F. Valois, J. Wanzer Drane May 2005

Multiple Imputation For Missing Ordinal Data, Ling Chen, Marian Toma-Drane, Robert F. Valois, J. Wanzer Drane

Journal of Modern Applied Statistical Methods

Simulations were used to compare complete case analysis of ordinal data with including multivariate normal imputations. MVN methods of imputation were not as good as using only complete cases. Bias and standard errors were measured against coefficients estimated from logistic regression and a standard data set.


Coverage Properties Of Optimized Confidence Intervals For Proportions, John P. Wendell, Sharon P. Cox May 2005

Coverage Properties Of Optimized Confidence Intervals For Proportions, John P. Wendell, Sharon P. Cox

Journal of Modern Applied Statistical Methods

Wardell (1997) provided a method for constructing confidence intervals on a proportion that modifies the Clopper-Pearson (1934) interval by allowing for the upper and lower binomial tail probabilities to be set in a way that minimizes the interval width. This article investigates the coverage properties of these optimized intervals. It is found that the optimized intervals fail to provide coverage at or above the nominal rate over some portions of the binomial parameter space but may be useful as an approximate method.