Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

2011

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 121 - 150 of 406

Full-Text Articles in Statistics and Probability

Targeted Maximum Likelihood Estimation Of Natural Direct Effect, Wenjing Zheng, Mark J. Van Der Laan Jul 2011

Targeted Maximum Likelihood Estimation Of Natural Direct Effect, Wenjing Zheng, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

In many causal inference problems, one is interested in the direct causal effect of an exposure on an outcome of interest that is not mediated by certain intermediate variables. Robins and Greenland (1992) and Pearl (2000) formalized the definition of two types of direct effects (natural and controlled) under the counterfactual framework. Since then, identifiability conditions for these effects have been studied extensively. By contrast, considerably fewer efforts have been invested in the estimation problem of the natural direct effect. In this article, we propose a semiparametric efficient, multiply robust estimator for the natural direct effect of a binary treatment …


On The Covariate-Adjusted Estimation For An Overall Treatment Difference With Data From A Randomized Comparative Clinical Trial, Lu Tian, Tianxi Cai, Lihui Zhao, L. J. Wei Jul 2011

On The Covariate-Adjusted Estimation For An Overall Treatment Difference With Data From A Randomized Comparative Clinical Trial, Lu Tian, Tianxi Cai, Lihui Zhao, L. J. Wei

Harvard University Biostatistics Working Paper Series

No abstract provided.


General Recognition Theory Extended To Include Response Times: Predictions For A Class Of Parallel Systems, Joseph W. Houpt, James T. Townsend, Noah H. Silbert Jul 2011

General Recognition Theory Extended To Include Response Times: Predictions For A Class Of Parallel Systems, Joseph W. Houpt, James T. Townsend, Noah H. Silbert

Psychology Faculty Publications

No abstract provided.


A Statistical Test For The Capacity Coefficient, Joseph W. Houpt, James T. Townsend Jul 2011

A Statistical Test For The Capacity Coefficient, Joseph W. Houpt, James T. Townsend

Psychology Faculty Publications

No abstract provided.


From Deep Space 9 To The Gamma Quadrant!, James T. Townsend, Joseph W. Houpt Jul 2011

From Deep Space 9 To The Gamma Quadrant!, James T. Townsend, Joseph W. Houpt

Psychology Faculty Publications

No abstract provided.


Targeted Minimum Loss Based Estimation Based On Directly Solving The Efficient Influence Curve Equation, Paul Chaffee, Mark J. Van Der Laan Jul 2011

Targeted Minimum Loss Based Estimation Based On Directly Solving The Efficient Influence Curve Equation, Paul Chaffee, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Applying targeted maximum likelihood estimation to longitudinal data can be computationally intensive. As the number of time points and/or number of intermediate factors grows, the computation resources consumed by these algorithms likewise increases. Different TMLE algorithms have different computational speeds and implementation challenges; there may also be efficiency differences of the corresponding estimators. The algorithm we describe here proceeds by solving the empirical efficient influence curve equation directly using numerical computation methods, rather than indirectly (by solving a score equation), which is the usual route. We believe that this estimator is the simplest of the TMLE procedures to implement in …


Assessing Medicare Beneficiaries’ Strength‐Of‐Preference Scores For Health Care Options: How Engaging Does The Elicitation Technique Need To Be?, Trafford Crump, Hilary A. Llewellyn-Thomas Jul 2011

Assessing Medicare Beneficiaries’ Strength‐Of‐Preference Scores For Health Care Options: How Engaging Does The Elicitation Technique Need To Be?, Trafford Crump, Hilary A. Llewellyn-Thomas

Dartmouth Scholarship

The objective was to determine if participants’ strength‐of‐preference scores for elective health care interventions at the end‐of‐life (EOL) elicited using a non‐engaging technique are affected by their prior use of an engaging elicitation technique.


Computational Insight With Monte Carlo Simulations, Boyan Kostadinov Jul 2011

Computational Insight With Monte Carlo Simulations, Boyan Kostadinov

Publications and Research

We introduce Monte Carlo simulations for estimating areas by playing a game of "darts". We also introduce simulations of random walks. We use compact, vectorized programming, based on the R language, for all computer simulations and visualizations, aimed at high school students. This presentation is based on the Invited, prime time lecture given at the summer camp for gifted high school students at City College of New York, July 13, 2011.


Variable Importance Analysis With The Multipim R Package, Stephan J. Ritter, Nicholas P. Jewell, Alan E. Hubbard Jul 2011

Variable Importance Analysis With The Multipim R Package, Stephan J. Ritter, Nicholas P. Jewell, Alan E. Hubbard

U.C. Berkeley Division of Biostatistics Working Paper Series

We describe the R package multiPIM, including statistical background, functionality and user options. The package is for variable importance analysis, and is meant primarily for analyzing data from exploratory epidemiological studies, though it could certainly be applied in other areas as well. The approach taken to variable importance comes from the causal inference field, and is different from approaches taken in other R packages. By default, multiPIM uses a double robust targeted maximum likelihood estimator (TMLE) of a parameter akin to the attributable risk. Several regression methods/machine learning algorithms are available for estimating the nuisance parameters of the models, including …


Reduced Bayesian Hierarchical Models: Estimating Health Effects Of Simultaneous Exposure To Multiple Pollutants, Jennifer F. Bobb, Francesca Dominici, Roger D. Peng Jul 2011

Reduced Bayesian Hierarchical Models: Estimating Health Effects Of Simultaneous Exposure To Multiple Pollutants, Jennifer F. Bobb, Francesca Dominici, Roger D. Peng

Johns Hopkins University, Dept. of Biostatistics Working Papers

Quantifying the health effects associated with simultaneous exposure to many air pollutants is now a research priority of the US EPA. Bayesian hierarchical models (BHM) have been extensively used in multisite time series studies of air pollution and health to estimate health effects of a single pollutant adjusted for potential confounding of other pollutants and other time-varying factors. However, when the scientific goal is to estimate the impacts of many pollutants jointly, a straightforward application of BHM is challenged by the need to specify a random-effect distribution on a high-dimensional vector of nuisance parameters, which often do not have an …


A Unified Approach To Non-Negative Matrix Factorization And Probabilistic Latent Semantic Indexing, Karthik Devarajan, Guoli Wang, Nader Ebrahimi Jul 2011

A Unified Approach To Non-Negative Matrix Factorization And Probabilistic Latent Semantic Indexing, Karthik Devarajan, Guoli Wang, Nader Ebrahimi

COBRA Preprint Series

Non-negative matrix factorization (NMF) by the multiplicative updates algorithm is a powerful machine learning method for decomposing a high-dimensional nonnegative matrix V into two matrices, W and H, each with nonnegative entries, V ~ WH. NMF has been shown to have a unique parts-based, sparse representation of the data. The nonnegativity constraints in NMF allow only additive combinations of the data which enables it to learn parts that have distinct physical representations in reality. In the last few years, NMF has been successfully applied in a variety of areas such as natural language processing, information retrieval, image processing, speech recognition …


Multiple Testing Of Local Maxima For Detection Of Unimodal Peaks In 1d, Armin Schwartzman, Yulia Gavrilov, Robert J. Adler Jul 2011

Multiple Testing Of Local Maxima For Detection Of Unimodal Peaks In 1d, Armin Schwartzman, Yulia Gavrilov, Robert J. Adler

Harvard University Biostatistics Working Paper Series

No abstract provided.


Targeted Methods For Finding Quantitative Trait Loci, Hui Wang, Sherri Rose, Mark J. Van Der Laan Jul 2011

Targeted Methods For Finding Quantitative Trait Loci, Hui Wang, Sherri Rose, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Conventional genetic mapping methods typically assume parametric models with Gaussian errors, and obtain parameter estimates through maximum likelihood estimation. We propose a general semiparametric model to map quantitative trait loci (QTL) in experimental crosses. In contrast with widely-used interval mapping (IM) derived methods, our model requires fewer assumptions and also accommodates various machine learning algorithms. Estimation using both targeted maximum likelihood and collaborative targeted maximum likelihood methods is compared to a composite interval mapping (CIM) approach. We demonstrate with simulations and real data analyses that, on average, our semiparametric targeted learning approach produces less biased QTL effect estimates than those …


Assessing The Effect Of Wal-Mart In Rural Utah Areas, Angela Nelson Jul 2011

Assessing The Effect Of Wal-Mart In Rural Utah Areas, Angela Nelson

Theses and Dissertations

Walmart and other “big box” stores seek to expand in rural markets, possibly due to cheap land and lack of zoning laws. In August 2000, Walmart opened a store in Ephraim, a small rural town in central Utah. It is of interest to understand how Walmart's entrance into the local market changes the sales tax revenue base for Ephraim and for the surrounding municipalities. It is thought that small “Mom and Pop” stores go out of business because they cannot compete with Walmart's prices, leading to a decrease in variety, selection, convenience, and most importantly, sales tax revenue base in …


An Introduction To Bayesian Methodology Via Winbugs And Proc Mcmc, Heidi Lula Lindsey Jul 2011

An Introduction To Bayesian Methodology Via Winbugs And Proc Mcmc, Heidi Lula Lindsey

Theses and Dissertations

Bayesian statistical methods have long been computationally out of reach because the analysis often requires integration of high-dimensional functions. Recent advancements in computational tools to apply Markov Chain Monte Carlo (MCMC) methods are making Bayesian data analysis accessible for all statisticians. Two such computer tools are Win-BUGS and SASR 9.2's PROC MCMC. Bayesian methodology will be introduced through discussion of fourteen statistical examples with code and computer output to demonstrate the power of these computational tools in a wide variety of settings.


Gene Set Analysis For Longitudinal Gene Expression Data, Ke Zhang, Haiyan Wang, Arne C. Bathke, Solomon W. Harrar, Hans-Peter Piepho, Youping Deng Jul 2011

Gene Set Analysis For Longitudinal Gene Expression Data, Ke Zhang, Haiyan Wang, Arne C. Bathke, Solomon W. Harrar, Hans-Peter Piepho, Youping Deng

Statistics Faculty Publications

BACKGROUND: Gene set analysis (GSA) has become a successful tool to interpret gene expression profiles in terms of biological functions, molecular pathways, or genomic locations. GSA performs statistical tests for independent microarray samples at the level of gene sets rather than individual genes. Nowadays, an increasing number of microarray studies are conducted to explore the dynamic changes of gene expression in a variety of species and biological scenarios. In these longitudinal studies, gene expression is repeatedly measured over time such that a GSA needs to take into account the within-gene correlations in addition to possible between-gene correlations.

RESULTS: We provide …


Helin Institutions' Collection Statistics From Fy 10 To Fy 11, Martha Rice Sanders Jul 2011

Helin Institutions' Collection Statistics From Fy 10 To Fy 11, Martha Rice Sanders

HELIN Collection Statistics

Statistical information about the total number of item and holdings (serials) records held by each HELIN member institution as of June 30, 2010, and June 30, 2011. Gives the percentage of growth for each institution. Additionally, a chart and statistics for the number of item records held by each HELIN member institution as of June 30, 2011. A Chart of e-book collection totals and the libraries to which they belong. Finally, a chart of serials holdings for both paper (plus microform, etc.) and electronic journals, including the CRIARL libraries.


A Study On Facility Planning Using Discrete Event Simulation: Case Study Of A Grain Delivery Terminal, Sarah M. Asio Jul 2011

A Study On Facility Planning Using Discrete Event Simulation: Case Study Of A Grain Delivery Terminal, Sarah M. Asio

Department of Industrial and Management Systems Engineering: Dissertations, Theses, and Student Research

The application of traditional approaches to the design of efficient facilities can be tedious and time consuming when uncertainty and a number of constraints exist. Queuing models and mathematical programming techniques are not able to capture the complex interaction between resources, the environment and space constraints for dynamic stochastic processes. In the following study discrete event simulation is applied to the facility planning process for a grain delivery terminal. The discrete event simulation approach has been applied to studies such as capacity planning and facility layout for a gasoline station and evaluating the resource requirements for a manufacturing facility. To …


How To Combine Independent Data Sets For The Same Quantity, Theodore P. Hill, Jack Miller Jul 2011

How To Combine Independent Data Sets For The Same Quantity, Theodore P. Hill, Jack Miller

Research Scholars in Residence

This paper describes a new mathematical method called conflation for consolidating data from independent experiments that measure the same physical quantity. Conflation is easy to calculate and visualize and minimizes the maximum loss in Shannon information in consolidating several independent distributions into a single distribution. A formal mathematical treatment of conflation has recently been published. For the benefit of experimenters wishing to use this technique, in this paper we derive the principal basic properties of conflation in the special case of normally distributed (Gaussian) data. Examples of applications to measurements of the fundamental physical constants and in high energy physics …


Comparing Hall Of Fame Baseball Players Using Most Valuable Player Ranks, Paul Kvam Jul 2011

Comparing Hall Of Fame Baseball Players Using Most Valuable Player Ranks, Paul Kvam

Department of Math & Statistics Faculty Publications

We propose a rank-based statistical procedure for comparing performances of top major league baseball players who performed in different eras. The model is based on using the player ranks from voting results for the most valuable player awards in the American and National Leagues. The current voting procedure has remained the same since 1932, so the analysis regards only data for players whose career blossomed after that time. Because the analysis is based on quantiles, its basis is nonparametric and relies on a simple link function. Results are stratified by fielding position, and we compare 73 Hall of Fame players …


A Bayesian Model Averaging Approach For Observational Gene Expression Studies, Xi Kathy Zhou, Fei Liu, Andrew J. Dannenberg Jun 2011

A Bayesian Model Averaging Approach For Observational Gene Expression Studies, Xi Kathy Zhou, Fei Liu, Andrew J. Dannenberg

COBRA Preprint Series

Identifying differentially expressed (DE) genes associated with a sample characteristic is the primary objective of many microarray studies. As more and more studies are carried out with observational rather than well controlled experimental samples, it becomes important to evaluate and properly control the impact of sample heterogeneity on DE gene finding. Typical methods for identifying DE genes require ranking all the genes according to a pre-selected statistic based on a single model for two or more group comparisons, with or without adjustment for other covariates. Such single model approaches unavoidably result in model misspecification, which can lead to increased error …


When Does Combining Markers Improve Classification Performance And What Are Implications For Practice?, Aasthaa Bansal, Margaret Sullivan Pepe Jun 2011

When Does Combining Markers Improve Classification Performance And What Are Implications For Practice?, Aasthaa Bansal, Margaret Sullivan Pepe

UW Biostatistics Working Paper Series

When an existing standard marker does not have sufficient classification accuracy on its own, new markers are sought with the goal of yielding a combination with better performance. The primary criterion for selecting new markers is that they have good performance on their own and preferably be uncorrelated with the standard. Most often linear combinations are considered. In this paper we investigate the increment in performance that is possible by combining a novel continuous marker with a moderately performing standard continuous marker under a variety of biologically motivated models for their joint distribution. We find that an uncorrelated continuous marker …


Hierarchical Probit Models For Ordinal Ratings Data, Allison M. Butler Jun 2011

Hierarchical Probit Models For Ordinal Ratings Data, Allison M. Butler

Theses and Dissertations

University students often complete evaluations of their courses and instructors. The evaluation tool typically contains questions about the course and the instructor on an ordinal Likert scale. We assess instructor effectiveness while adjusting for known confounders. We present a probit regression model with a latent variable to measure the instructor effectiveness accounting for student specific covariates, such as student grade in the course, high school and university GPA, and ACT score.


Targeted Maximum Likelihood Estimation Of Conditional Relative Risk In A Semi-Parametric Regression Model, Cathy Tuglus, Kristin E. Porter, Mark J. Van Der Laan Jun 2011

Targeted Maximum Likelihood Estimation Of Conditional Relative Risk In A Semi-Parametric Regression Model, Cathy Tuglus, Kristin E. Porter, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

The conditional relative risk is an important measure in medical and epidemiological studies when the outcome of interest is binary (i.e. disease vs. no disease). When the outcome is common, estimation of conditional relative risk and related parameters can be problematic, especially when the exposure or covariates are continuous. We propose a new estimation procedure based on targeted maximum likelihood methodology that targets the parameters relating to the conditional relative risk for common outcomes under a log-linear, or multiplicative, semi-parametric model. In this paper, we present three possible targeted maximum likelihood estimators for relative risk parameters implied by such a …


Copy Number Variants In Candidate Genes Are Genetic Modifiers Of Hirschsprung Disease, Qian Jiang, Yen Yi Ho, Li Hao, Courtney Nichols Berrios, Aravinda Chakravarti Jun 2011

Copy Number Variants In Candidate Genes Are Genetic Modifiers Of Hirschsprung Disease, Qian Jiang, Yen Yi Ho, Li Hao, Courtney Nichols Berrios, Aravinda Chakravarti

Faculty Publications

Hirschsprung disease (HSCR) is a neurocristopathy characterized by absence of intramural ganglion cells along variable lengths of the gastrointestinal tract. The HSCR phenotype is highly variable with respect to gender, length of aganglionosis, familiality and the presence of additional anomalies. By molecular genetic analysis, a minimum of 11 neuro-developmental genes (RET, GDNF, NRTN, SOX10, EDNRB, EDN3, ECE1, ZFHX1B, PHOX2B, KIAA1279, TCF4) are known to harbor rare, high-penetrance mutations that confer a large risk to the bearer. In addition, two other genes (RET, NRG1) harbor common, low-penetrance polymorphisms that contribute only partially to risk and can act as genetic modifiers. To …


Probabilistic Assessment Of Drought Characteristics Using A Hidden Markov Model, Ganeshchandra Mallya, Shivam Tripathi, Sergey Kirshner, Rao S. Govindaraju Jun 2011

Probabilistic Assessment Of Drought Characteristics Using A Hidden Markov Model, Ganeshchandra Mallya, Shivam Tripathi, Sergey Kirshner, Rao S. Govindaraju

2011 Symposium on Data-Driven Approaches to Droughts

Droughts are evaluated using drought indices that measure the departure of meteorological and hydrological variables such as precipitation and stream flow from their long-term averages. While there are many drought indices proposed in the literature, most of them use pre-defined thresholds for identifying drought classes ignoring the inherent uncertainties in characterizing droughts. In this study, a hidden Markov model (HMM) [1] is developed for probabilistic classification of drought states. The HMM captures space and time dependence in the data. The proposed model is applied to assess drought characteristics in Indiana using monthly precipitation and stream flow data. The comparison of …


Super Learner Based Conditional Density Estimation With Application To Marginal Structural Models, Ivan Diaz Munoz, Mark J. Van Der Laan Jun 2011

Super Learner Based Conditional Density Estimation With Application To Marginal Structural Models, Ivan Diaz Munoz, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

In this paper we present a histogram-like estimator of a conditional density that uses super learner crossvalidation to estimate the histogram probabilities, as well as the optimal number and position of the bins. This estimator is an alternative to kernel density estimators when the dimension of the problem is large. We demonstrate its applicability to estimation of Marginal Structural Model (MSM) parameters in which an initial estimator of the treatment %mechanism is needed. MSM estimation based on the proposed density estimator results in less biased estimates, when compared to estimates based on a misspecified parametric model.


Comparing Roc Curves Derived From Regression Models, Venkatraman E. Seshan, Mithat Gonen, Colin B. Begg Jun 2011

Comparing Roc Curves Derived From Regression Models, Venkatraman E. Seshan, Mithat Gonen, Colin B. Begg

Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series

In constructing predictive models, investigators frequently assess the incremental value of a predictive marker by comparing the ROC curve generated from the predictive model including the new marker with the ROC curve from the model excluding the new marker. Many commentators have noticed empirically that a test of the two ROC areas often produces a non-significant result when a corresponding Wald test from the underlying regression model is significant. A recent article showed using simulations that the widely-used ROC area test [1] produces exceptionally conservative test size and extremely low power [2]. In this article we show why the ROC …


On Causal Mediation Analysis With A Survival Outcome, Eric J. Tchetgen Tchetgen Jun 2011

On Causal Mediation Analysis With A Survival Outcome, Eric J. Tchetgen Tchetgen

Harvard University Biostatistics Working Paper Series

Suppose that having established a marginal total effect of a point exposure on a time-to-event outcome, an investigator wishes to decompose this effect into its direct and indirect pathways, also know as natural direct and indirect effects, mediated by a variable known to occur after the exposure and prior to the outcome. This paper proposes a theory of estimation of natural direct and indirect effects in two important semiparametric models for a failure time outcome. The underlying survival model for the marginal total effect and thus for the direct and indirect effects, can either be a marginal structural Cox proportional …


Semiparametric Estimation Of Models For Natural Direct And Indirect Effects, Eric J. Tchetgen Tchetgen, Ilya Shpitser Jun 2011

Semiparametric Estimation Of Models For Natural Direct And Indirect Effects, Eric J. Tchetgen Tchetgen, Ilya Shpitser

Harvard University Biostatistics Working Paper Series

In recent years, researchers in the health and social sciences have become increasingly interested in mediation analysis. Specifically, upon establishing a non-null total effect of an exposure, investigators routinely wish to make inferences about the direct (indirect) pathway of the effect of the exposure not through (through) a mediator variable that occurs subsequently to the exposure and prior to the outcome. Natural direct and indirect effects are of particular interest as they generally combine to produce the total effect of the exposure and therefore provide insight on the mechanism by which it operates to produce the outcome. A semiparametric theory …